Sentinel Intruder Detection System
An intruder-detection system fusing simulated Wokwi hardware sensors with computer vision (YOLO person + weapon detection, fire, and face recognition), wired to a Streamlit dashboard with real-time email alerts.
84 things I've grown, from clinical LLM systems to causal studies to weekend ML experiments. New projects sprout here on their own as I push them to GitHub. The big blooms are below; pick a patch to wander through the rest.
semantic search · your phrase and every project are embedded with OpenAI, then ranked by cosine similarity ✦
An intruder-detection system fusing simulated Wokwi hardware sensors with computer vision (YOLO person + weapon detection, fire, and face recognition), wired to a Streamlit dashboard with real-time email alerts.
A multimodal medical-record companion unifying RAG, document understanding, speech, and vision. Consensus extraction across LLMs hit 85.1% micro-F1 with sub-2s latency.
Benchmarked KIVI quantization, TopK sparsity, SnapKV eviction & MLA on Llama-2-7B with Triton kernels, 4× cache compression, 1.93× faster decode, 3.1× peak throughput.
End-to-end causal inference study of the Moertel (1990) colon cancer trial (n=929), evaluating average and heterogeneous treatment effects, mediation pathways, confounding bias, and external validity through transportability analyses.
A multi-agent CrewAI system for U.S. federal legal analysis: semantic USC retrieval, precedent search, elements analysis, and draft generation.
AI food app for video-to-recipe extraction, pantry matching, and personalized health coaching. React + Vite + Framer Motion on Netlify.
A multimodal medical-record companion unifying RAG, document understanding, speech, and vision. Consensus extraction across LLMs hit 85.1% micro-F1 with sub-2s latency.
A multi-agent CrewAI system for U.S. federal legal analysis: semantic USC retrieval, precedent search, elements analysis, and draft generation.
Full-stack RAG pipeline turning natural language into SQL and visualizations for sports performance analytics.
AI food app for video-to-recipe extraction, pantry matching, and personalized health coaching. React + Vite + Framer Motion on Netlify.
Converts cooking videos into structured recipes via a multi-stage vision-language pipeline: frame extraction, visual understanding, LLM reasoning.
Student-built dashboard for Columbia MSDS course reviews, live Google-Sheets data, AI-summarized reviews, rankings, and side-by-side comparisons.
AI coaching that maps fitness queries to personalized exercises via two-stage retrieval + LLM re-ranking. FastAPI, PostgreSQL, Claude.
AI diary turning journal entries into personalized verses + music recs. DistilRoBERTa emotion, K-Means over 867 songs, FAISS RAG, GPT-4o-mini.
AI blog-post generator with DALL·E images, SEO-optimized content with customizable tone, length, and generated visuals.
Educational medical-image analysis with Gemini Vision, upload or camera capture, safe non-diagnostic insights with built-in disclaimers.
Explains technical concepts through personalized analogies based on your interests. Interactive and friendly.
Defensive benchmark: do agent architectures reduce prompt-injection susceptibility? Inert canary tokens only
A false-positive rate for LLM safety behaviour: does the model refuse legitimate, sensitive-adjacent requests?
The accuracy-vs-cost Pareto frontier for local LLM inference on Apple silicon, energy included
Live cross-model benchmark: reasoning accuracy on known-answer problems, real streaming latency and throughput, and structured-output validity across GPT-4o, Claude Sonnet 5, Gemini 2.5 Flash and Grok-4.3. Bring your own keys.
Adaptive reasoning-depth controller: early-exit chain-of-thought when the answer stabilizes, vs fixed budgets and self-consistency.
Production-style LLM gateway in pure Python stdlib: auth, semantic cache, rate limiting, retries, fallbacks, circuit breaker, cost tracking. Load-tested through a brownout.
Speculative decoding benchmark: draft size x gamma x task sweep with closed-form speedup validated by token-level Monte Carlo, plus a real Colab notebook.
LLM inference profiler: TTFT, inter-token latency, throughput, memory and quality across FP16/INT8/INT4, batch size, context length, KV caching. Dashboard included.
Delta debugging (ddmin) for LLM prompts: shrink a failing prompt to the minimal sub-prompt that still reproduces the failure.
Defensive LLM safety evaluation: refusal robustness across harm categories and framings, over-refusal on benign lookalikes, consistency. No attack content.
Cost-, latency-, and quality-aware LLM model router benchmarked against always-large and always-small baselines over a 4k-request workload.
Regression testing for LLM reasoning: declarative test specs, consistency checks, and a CI gate that fails deploys on category regressions.
Fault-injection harness for structured LLM outputs: JSON mode vs function calling vs grammar constraints vs a staged repair pipeline.
Long-context stress tests: lost-in-the-middle heatmaps, multi-hop joins, instruction retention, distractors, and the latency and KV-memory bill.
Does agreement across reasoning chains predict correctness? Calibration curves, ECE, and five selection strategies vs chain count.
A multi-agent CrewAI system for U.S. federal legal analysis: semantic USC retrieval, precedent search, elements analysis, and draft generation.
An agent memory layer benchmarked on facts that change: contradiction resolution and recency vs naive similarity
Debug multi-agent systems like distributed systems: happens-before graphs and deadlock/livelock/lost-update detection
Compare agent architectures at matched cost, not matched accuracy: the cost/accuracy frontier
Defensive benchmark: do agent architectures reduce prompt-injection susceptibility? Inert canary tokens only
Record-and-replay behavioural diffs for agent changes: path divergence vs outcome equivalence
Full-stack RAG pipeline turning natural language into SQL and visualizations for sports performance analytics.
Student-built dashboard for Columbia MSDS course reviews, live Google-Sheets data, AI-summarized reviews, rankings, and side-by-side comparisons.
AI diary turning journal entries into personalized verses + music recs. DistilRoBERTa emotion, K-Means over 867 songs, FAISS RAG, GPT-4o-mini.
TF-IDF + Linear SVM on the WELFake dataset with a real-time credibility-prediction interface.
A summariser where every sentence carries its source span and unsupported ones are flagged, with a validated NLI verifier
End-to-end causal inference study of the Moertel (1990) colon cancer trial (n=929), evaluating average and heterogeneous treatment effects, mediation pathways, confounding bias, and external validity through transportability analyses.
ML & predictive analytics framework for identifying high-risk child-welfare cases using NCANDS data.
Generate the conditions that break your detector, measure the per-edit taxonomy, then train on the failures
Benchmarking causal discovery under realistic measurement error, with a reliability-correction fix
A mediation analysis where the reader turns the untestable assumptions: DML effects with an interactive sensitivity dial
End-to-end causal inference study of the Moertel (1990) colon cancer trial (n=929), evaluating average and heterogeneous treatment effects, mediation pathways, confounding bias, and external validity through transportability analyses.
Visual analysis of how diet and lifestyle contribute to colorectal cancer risk, built in R/Shiny.
Shape bucketing as an optimisation over a real request-length distribution: DP-optimal buckets vs the recompilation storm
ML & predictive analytics framework for identifying high-risk child-welfare cases using NCANDS data.
Predicts used-car prices (India), 65.6% R² with XGBoost + Optuna, SHAP explainability across 4 ensembles.
Gradient Boosting on physicochemical properties with SHAP, feature importance, and PDPs for interpretability.
Production-style pipeline + Gradient Boosting for loan approval with SHAP explanations and an interactive UI.
XGBoost regression on California housing with an interactive what-if explorer.
Gradient Boosting on clinical indicators with SHAP for global and per-patient explanations.
Classifies sonar signals as rock or mine via cross-validated automated model selection.
Predicts repeat-purchase behavior for food-delivery businesses and surfaces key retention drivers.
EDA, feature engineering, SMOTE for imbalance, and tuned models exploring heart-disease risk factors.
Diagnose out-of-memory errors that are fragmentation, not capacity: a caching-allocator replay, a fragmentation-aware OOM detector, and a size-class fix that removes the OOM without adding memory
Catastrophic forgetting as a dose-response law, inverted into a calculator for the largest safe fine-tune
Map where annotators disagree over preference space, and how that reshapes reward models
Rewrite text so it means the same but no longer sounds like you: the anonymity-meaning frontier
Offline RL where the headline is the confidence interval's width and the power analysis, not a fragile point estimate
Does any data ordering beat shuffling at fixed compute, once scoring cost is counted? An honest negative
Ranked grasp hypotheses with stated reasons and an abstention option, plus the success-vs-coverage curve
Learn control from pairwise trajectory preferences with a direct objective, no reward model
AutoML that refuses features on cost grounds, not just accuracy: joint accuracy-cost Pareto selection with a per-feature argument
Synthetic control on a tiny donor pool with honest placebo-permutation inference for small places
Auto-propose and run the negative-control falsification tests an observational analysis should have failed
A logging simulator with known ground truth that maps where off-policy evaluation estimators lie
A leakage audit for image datasets: masked-variant ablations that fingerprint what the dataset really teaches
Recover true labels and annotator reliabilities from noisy votes with Dawid-Skene, and verify only the suspicious labels
A model that knows when a question is underspecified: answer-entropy ambiguity detection with clarifying questions
What ordinary handling does to an image watermark: survivability cards for a block-DCT scheme (defensive)
Active learning for hand-eye calibration: D-optimal pose selection reaches target accuracy in far fewer poses
A gallery of small environments where a plausible reward produces a degenerate policy, plus an honest look at detection
Dataset-free per-rep form feedback from your own best reps, validated on exactly-known degradations
Cluster, name, and validate a classifier's failure modes, then prove the names identify held-out errors
ML & predictive analytics framework for identifying high-risk child-welfare cases using NCANDS data.
Predicts used-car prices (India), 65.6% R² with XGBoost + Optuna, SHAP explainability across 4 ensembles.
Production-style pipeline + Gradient Boosting for loan approval with SHAP explanations and an interactive UI.
XGBoost regression on California housing with an interactive what-if explorer.
Gradient Boosting on clinical indicators with SHAP for global and per-patient explanations.
Predicts repeat-purchase behavior for food-delivery businesses and surfaces key retention drivers.
EDA, feature engineering, SMOTE for imbalance, and tuned models exploring heart-disease risk factors.
Predict the grokking (delayed-generalisation) jump from the first 10% of training, via gradient-noise dynamics
Forecast when a tabular model will decay from early drift dynamics, with calibrated conformal intervals
When does gradient boosting stop being the right model? Learning-curve crossovers across 20 tabular datasets, with a meta-feature recommender
Diachronic term drift on arXiv with a seed-noise null and a held-out predictive test
Does agreement across reasoning chains predict correctness? Calibration curves, ECE, and five selection strategies vs chain count.
Automated keratoconus detection using SVM and deep neural networks.
Classifies CT scans into Normal/Cyst/Stone/Tumor via VGG19 & ResNet50 transfer learning, 99.2% accuracy.
Automated cataract detection using CNNs and transfer learning.
VGG16 transfer learning + fine-tuning on GTSRB, 98% accuracy across 43 classes.
ResNet50 transfer learning across 38 plant-disease classes with training, evaluation, and inference tools.
Deep neural networks combining static + dynamic analysis to detect malicious Android apps.
An intruder-detection system fusing simulated Wokwi hardware sensors with computer vision (YOLO person + weapon detection, fire, and face recognition), wired to a Streamlit dashboard with real-time email alerts.
A multimodal medical-record companion unifying RAG, document understanding, speech, and vision. Consensus extraction across LLMs hit 85.1% micro-F1 with sub-2s latency.
Converts cooking videos into structured recipes via a multi-stage vision-language pipeline: frame extraction, visual understanding, LLM reasoning.
Educational medical-image analysis with Gemini Vision, upload or camera capture, safe non-diagnostic insights with built-in disclaimers.
Automated keratoconus detection using SVM and deep neural networks.
Classifies CT scans into Normal/Cyst/Stone/Tumor via VGG19 & ResNet50 transfer learning, 99.2% accuracy.
Automated cataract detection using CNNs and transfer learning.
VGG16 transfer learning + fine-tuning on GTSRB, 98% accuracy across 43 classes.
ResNet50 transfer learning across 38 plant-disease classes with training, evaluation, and inference tools.
A free touch sensor: audio auto-labels contact, a vision model learns it, then the microphone is thrown away
The smallest file that keeps the detector right: per-image JPEG bitrate policy vs a global quality
Benchmarked KIVI quantization, TopK sparsity, SnapKV eviction & MLA on Llama-2-7B with Triton kernels, 4× cache compression, 1.93× faster decode, 3.1× peak throughput.
Search which layers actually need their bits: evolutionary per-layer mixed-precision quantization vs uniform
Live cross-model benchmark: reasoning accuracy on known-answer problems, real streaming latency and throughput, and structured-output validity across GPT-4o, Claude Sonnet 5, Gemini 2.5 Flash and Grok-4.3. Bring your own keys.
LLM inference profiler: TTFT, inter-token latency, throughput, memory and quality across FP16/INT8/INT4, batch size, context length, KV caching. Dashboard included.
Deep neural networks combining static + dynamic analysis to detect malicious Android apps.
Secure storage of encrypted medical/review data using ECDSA signatures, Proof of Work, and hash chaining.
A modular confusion/diffusion pipeline with swappable primitives (perm, spectral, xor, latin, inn), and a metrics harness (entropy, correlation, NPCR, UACI).
An intruder-detection system fusing simulated Wokwi hardware sensors with computer vision (YOLO person + weapon detection, fire, and face recognition), wired to a Streamlit dashboard with real-time email alerts.
A free touch sensor: audio auto-labels contact, a vision model learns it, then the microphone is thrown away
every project embedded and projected to 2D, so similar work sits close together. drop in something you care about and see where you land.
mapping the galaxy… 🌌