Projects
ML and Quant Research Projects
Research, publications, and applied projects spanning machine learning and quantitative finance.
ML / AI Research
Jun 2026
ML / AI~60 secondsResumeRadar, AI Resume Scorer
A free, live web tool that scores a resume the way a technical recruiter would in roughly 60 seconds. Users upload a PDF and receive an explainable 0 to 100 score across open-source contributions, self-directed projects, production experience, and technical skills, along with concrete recommendations for improvement. The system enriches the score with signals pulled from the candidate's GitHub, including stars, primary languages, and authored versus total commits.
~60 seconds Runtime4 dimensions Score axesFree CostLive DeploymentStreamlitLLMGitHub APIProductApr 2026
ML / AI6 (7B to 70B)Reasoning Model Failure Analysis, LLM Interpretability
A controlled LLM evaluation pipeline spanning six reasoning models from 7B to 70B parameters, designed to disentangle reasoning length effects from forced re-entry interventions. The study measured a 36-point accuracy decline in Llama-distilled models while Qwen-distilled models remained robust. Multi-GPU inference was conducted with a bfloat16 KV cache on 4x GH200 GPUs.
6 (7B to 70B) Models evaluated36 pts down Llama degradationstable Qwen degradation4x GH200 HardwarePyTorchHuggingFaceSlurmMulti-GPUKSE 2024, Published
ML / AI94% accAdversarial Robustness via Entropy Based Feature Selection in RL
An entropy-based feature selection framework for reinforcement learning agents that achieved 94 and 95 percent accuracy on Lunar Lander and Bipedal Walker under adversarial perturbations, outperforming KL Divergence and Joint Entropy baselines across Gym environments. The public reconstruction now ships standalone entropy and mutual-information diagnostics with adaptive histogram binning and a majority-class baseline.
94% acc Lunar Lander95% acc Bipedal WalkerKL, Joint-H Baselines beatenKSE 2024 VenueOpenAI GymPyTorchRLPublicationMutual InformationCISS 2025, Published
ML / AIFluorescenceMouse Brain Cell Segmentation in Fluorescence Microscopy
A deep learning segmentation pipeline for high-noise fluorescence microscopy images, comprising a CNN architecture and a custom preprocessing routine for automated cell boundary detection.
Fluorescence ModalityCell boundary TaskCISS 2025 VenueOpenCVPyTorchCNNsPublicationCISS 2025, Published
ML / AI<100 msVirtual Yoga Instructor with Real Time Feedback
A real-time pose estimation and corrective feedback system using normalized joint angle features and repetition counting. The system achieves sub-100 ms latency on standard hardware and remains robust to variations in body size and camera angle.
<100 ms LatencyNorm. joint angles FeaturesCISS 2025 VenueOpenCVMediaPipePyTorchPublicationJun 2026
ML / AICLIVLM Evaluation Harness, Multimodal Model CLI
A command-line evaluation harness for vision-language models that emits structured JSON and CSV logs per run. The tool is designed for reproducible multimodal benchmarking with explicit configuration, deterministic seeds, and per-sample traceability. Metrics now include exact-match and SQuAD-style token-F1 with per-example F1 traceability and a constant-baseline reference.
CLI SurfaceJSON, CSV OutputsVLMs / MLLMs TargetPythonCLIVLMEvaluationToken-F1Jun 2026
ML / AICartPoleRobust RL Observation Noise Benchmark
A CartPole observation-noise benchmark with a DQN baseline and extensible noise processes, designed for controlled evaluation of reinforcement learning robustness under sensor perturbations. The suite now covers the evaluation loop, DQN agent, and noise processes, with input-validation hardening.
CartPole EnvironmentDQN BaselinePluggable NoisePyTorchOpenAI GymDQNBenchmarkReproducibilityJun 2026
ML / AI0.7591AI for Construction Safety, Evidence Package
A construction-safety vision evaluation reaching F1 0.7591, packaged with paired statistical tests, decision-threshold sensitivity analysis, and end-to-end claim traceability from raw predictions to reported numbers. Sparse-category rates are now reported with Wilson, Agresti-Coull, and exact Clopper-Pearson confidence intervals.
0.7591 F1Paired TestsThreshold sweep SensitivityComputer VisionVLMEvaluationSafetyConfidence IntervalsJun 2026
ML / AIMulti-outputPrivacy-Preserving Career Prediction Benchmark
A privacy-preserving multi-output career-prediction benchmark built on synthetic data, paired with an entropy-based feature selection routine that controls disclosure while preserving downstream predictive utility. The benchmark now includes McNemar and Cochran-Q significance tests, Cohen's kappa agreement, and per-class diagnostics.
Multi-output TaskSynthetic DataEntropy-based SelectionPythonMulti-OutputPrivacyFeature SelectionSignificance Testing
Quantitative Research
Jan 2026
Quant100K antitheticOptions Pricing Engine and Greeks Computation
Black-Scholes closed-form and Monte Carlo pricers with 100K antithetic paths for European equity options. Delta, Gamma, and Vega are computed both analytically and via finite differences across strike and maturity grids. Sensitivities now extend to first- through third-order Greeks (vanna, vomma, charm, speed, zomma, color), each cross-checked against finite differences.
100K antithetic MC pathsΔ, Γ, Vega + higher-order GreeksBS, MC PricersPythonNumPySciPyMonte CarloHigher-Order GreeksDec 2025
QuantAAPL/MSFT, KO/PEP, XOM/CVXStatistical Pairs Trading Backtest
An Engle-Granger market-neutral strategy applied to AAPL/MSFT, KO/PEP, and XOM/CVX over daily data from 2015 to 2023. The backtest incorporates walk-forward cointegration screening, out-of-sample hedge ratios, and a 1 bp transaction cost assumption. Sharpe ratio, maximum drawdown, turnover, and signal decay are reported. Mean-reversion diagnostics now include the Hurst exponent, a Lo-MacKinlay variance-ratio test, and a Kalman time-varying hedge ratio.
AAPL/MSFT, KO/PEP, XOM/CVX Pairs2015 to 2023 daily Window1 bp CostsWalk-forward MethodPythonstatsmodelsyfinanceBacktestingKalman FilterMay 2025
QuantSPY + 5 namesGARCH Volatility Modeling and Stochastic Time Series
A GARCH(1,1) implementation applied to SPY and five single-name equities, validated using AIC, BIC, and Ljung-Box diagnostics. GARCH, LSTM, and rolling volatility baselines were benchmarked across the 2020 and 2022 stress periods through walk-forward error analysis. Forecast evaluation now adds Mincer-Zarnowitz calibration, HAC (Newey-West) Diebold-Mariano tests, and range-based Parkinson/Garman-Klass volatility estimators.
SPY + 5 names Universe2020, 2022 Stress periodsAIC/BIC, Ljung-Box DiagnosticsPythonstatsmodelsGARCHTime SeriesDiebold-MarianoDec 2023
Quant10+ yrs dailyLSTM Based Financial Time Series Forecasting
A stacked LSTM trained on over ten years of daily equity price and volume data, evaluated out-of-sample on the 2022 to 2023 period against ARIMA, GARCH, and random walk baselines using walk-forward error decomposition. Forecast comparisons now report Diebold-Mariano significance (with a Newey-West variance option) and the Theil U2 statistic.
10+ yrs daily Training data2022 to 2023 Held outARIMA, GARCH, RW BaselinesPyTorchLSTMTime SeriesyfinanceDiebold-Mariano
Tools and Lists
Smaller artifacts: a reproducibility template and curated reading lists supporting the research above.
- Template
Reproducible ML Experiments Template
A drop-in skeleton for config-driven, per-run-logged ML experiments. Standardizes config, seeds, and output layout so runs are comparable and traceable from day one. It now captures git provenance (commit, branch, dirty state) into environment.json, validates configs, maintains a run index with config fingerprints, and ships a CI workflow running pytest, ruff, and mypy.
- Curated list
Awesome LLM Reasoning Evaluation
A curated list of benchmarks, papers, and tools focused specifically on evaluating reasoning behavior in large language models.
- Curated list
Awesome VLM Evaluation
A curated list of benchmarks, datasets, papers, and tools for evaluating vision-language models and multimodal LLMs.