Public CV

Saif Ryan Gangaram

Data Scientist 2 at Los Alamos National Laboratory and Computer Science Ph.D. student at the University of New Mexico, doing active research in AI/ML, high performance computing, Big Data, and systems, with a focus on scalable systems and algorithms for AI/ML training and inference.

Home

Contact

Education

University of New Mexico, Department of Computer Science

Ph.D. in Computer Science, expected 2028

Focus: high performance computing, Big Data, AI/ML, signal processing, scientific data systems

Research advisor: Dr. Amanda Bienz

University at Buffalo, Institute for AI and Data Science

Master of Professional Studies in Data Sciences and Applications, 2023

University at Buffalo, College of Arts and Sciences

Bachelor of Arts in Mathematics, Computing and Applied Mathematics, 2021

Research Interests

Professional Profile

My work sits at the intersection of scientific computing, Big Data, data-intensive software engineering, high performance computing, applied machine learning, and signal analysis. I have built production-oriented research tools for large sensor and experimental datasets, database-backed analysis interfaces, HPC-enabled workflows, streaming inference pipelines, generative AI and deep learning models, and validation tooling for reproducible technical analysis. My model work includes supervised learning, unsupervised learning, clustering, regression, and large-scale training on terabyte-class datasets with models that can reach millions of parameters. I am doing active research at Los Alamos National Laboratory and the University of New Mexico in AI/ML, HPC, Big Data, and systems, with a current emphasis on scalable systems and scalable algorithms for AI/ML training and inference. I am also building PolytopeOS and the Polytope Compiler, a from-scratch research operating system and compiler toolchain for scientific and engineering workloads.

Experience

Data Scientist 2, Los Alamos National Laboratory

2026 - Present
  • Conduct active research across AI/ML, HPC, Big Data, and systems, with an emphasis on scalable systems and algorithms for AI/ML training and inference.
  • Develop data acquisition, analysis, automation, and validation workflows for engineering test and scientific computing environments.
  • Apply statistical modeling, time-frequency analysis, denoising, stationarity testing, and signal characterization to high-volume technical datasets.
  • Build Python, C++, Dask, Slurm, NetCDF, and ML-enabled workflows for scalable analysis, batch execution, microbatched inference, and reproducible deployment.
  • Develop Big Data, generative AI, supervised, and unsupervised modeling workflows for clustering, regression, anomaly detection, and deep learning on terabyte-scale datasets.
  • Built an internal agentic AI workflow harness from scratch with Codex and Claude Code-style software assistance, multi-agent orchestration, approved internal LLM API integration, and tool-enabled automation for controlled enterprise environments.
  • Integrated enterprise collaboration context, meeting-summary workflows, policy and IT support interactions, document generation, and email/workflow drafting to streamline recurring technical and management tasks.
  • Received strong stakeholder interest for broader phased organizational deployment in approved internal environments.
  • Harden large-file processing pipelines with chunked execution, manifest tracking, integrity validation, runtime instrumentation, and atomic output publishing.
  • Optimize inference and data movement for large collections of multi-terabyte technical files where throughput and low-latency operation are core algorithmic requirements.
  • Improve cross-environment portability through local/HPC execution paths, environment-driven configuration, and reproducible workflow documentation.

Software Engineer, Automation and AI/ML, Space Dynamics Laboratory

2023 - 2026
  • Developed scientific data pipelines, database-backed web tools, and automation workflows for aerospace and remote-sensing research settings.
  • Integrated PostgreSQL, SQLite, HDF5, NetCDF, Python, C#, C++, Fortran, MATLAB, Flask, JavaScript, and dashboard tooling into analysis workflows.
  • Built searchable data interfaces with visualization, selection, export, compression, schema inspection, and reproducible model-input generation features.
  • Applied machine learning, anomaly detection, clustering, regression, simulation, image processing, and verification methods to large sensor, atmospheric, and telemetry-style datasets.
  • Designed multithreaded communication and automation components using REST, TCP/UDP interfaces, synchronization primitives, and testable configuration workflows.
  • Collaborated with multidisciplinary teams using Git, CI/CD practices, technical documentation, requirements analysis, integration testing, and iterative stakeholder feedback.

Teaching Assistant, University at Buffalo

2019 - 2020
  • Supported computer science labs and recitations for more than 45 students.
  • Assisted with grading, exam administration, student mentoring, and occasional lecture support.

Selected Projects

Hyperion Z: Self-Hosted Agentic AI Operations Console

Developed a local-first, open-source agentic harness using Deno, TypeScript, vanilla JavaScript, and WebSocket, with no frontend framework, bundler, or required cloud dependency beyond model APIs. Completed Phase 2 with parallel OpenAI and Anthropic agent sessions, token-by-token streaming, live tool and error events, file context injection, persistent cross-session memory, AI-assisted email drafting with tone control, and tmux session management with command suggestions based on live pane output. Added a full mock mode so the interface remains explorable without API keys. Phase 3 is planned to add multi-user authentication, CalDAV integration, MCP server support, and persistent sessions.

github.com/srgangaram-swe/Hyperion

Crucible: Local-First Training Data Refinery

Built a local-first data refinery and data-centric research platform for distributed training, now at Phase 8 of 8. The pipeline spans batch and streaming ingestion into bronze/silver/gold Parquet layers, statistical quality gates with quarantine, exact and MinHash/LSH near-duplicate removal, content-addressed dataset versions with lineage, a point-in-time-correct feature layer, deterministic training shards with a CPU/DDP/FSDP reference transformer trainer and exact checkpoint resume, an idempotent DAG orchestrator with a FastAPI metadata service and Streamlit dashboard, and a seed-controlled ablation harness for deduplication, mixing, quality gating, and data-scaling studies. v1.0 hardening and release are in progress.

github.com/srgangaram-swe/Crucible

Internal Agentic AI Workflow Harness

Built an internal agentic AI harness with Codex and Claude Code-style development assistance, multi-agent orchestration, approved internal LLM API integration, tool execution, enterprise collaboration context retrieval, meeting summarization, document generation, and email/workflow drafting. Designed the system for controlled internal environments, stakeholder-driven adoption, and phased expansion across a larger technical organization.

SynaptiCast: Causal Neural Telemetry Decoding and BCI Research Platform

Built a production-style ML research platform for causal neural telemetry decoding, probabilistic forecasting, and closed-loop BCI systems research. The platform generates reproducible synthetic motor-intent data, preprocesses neural signals, trains classical, CNN, and transformer decoders, runs masked, contrastive, and next-window self-supervised pretraining, evaluates embeddings, and emits publication-ready reports and benchmark sweeps with CI and laptop-friendly defaults. Sprint 1, a leakage-safe causal bidirectional send/receive telemetry benchmark, is complete; the roadmap continues through governed public electrophysiology datasets with held-out-subject evaluation, BIDS-native multimodal neuroimaging, a typed closed-loop BCI control plane, and a deterministic real-time edge runtime.

github.com/srgangaram-swe/SynaptiCast

Dork LLM: End-to-End LLM Systems Platform

Built a compact, production-style LLM systems platform with a tiny GPT trained from scratch, transformer internals, reusable evaluation harnesses for models and agents, cited RAG over documents, an agentic research assistant, FastAPI and Streamlit serving, Docker, CI, tests, typed configs, and local-first offline fallback paths. The latest release delivered a truthful model-backed chat vertical slice with corrected ML invariants, a streaming API, an accessible UI, and full-stack CI; the roadmap continues through statistical rigor and uncertainty, deep learning systems, a grounded production platform, and a reproducible public release.

github.com/srgangaram-swe/dork-llm

SignaLattice: Calibrated Probabilistic Market Forecasting

Built a calibrated probabilistic market forecasting platform spanning point-in-time data ingestion with schema validation and offline fallback, leakage-aware time-series and cross-sectional feature engineering, walk-forward machine learning with time-series-safe splits, vectorized long-only and long/short backtesting with costs and slippage, risk and scenario analytics, local experiment tracking with dataset hashes and git metadata, tests, CI, Docker, and generated reports. v0.2 shipped calibrated probabilistic forecasts with proper scoring rules, causal deep learning, tradability evidence, and a model-risk go/no-go report; the roadmap adds a bitemporal point-in-time data mesh, Bayesian cross-asset scenario analysis, a ForecastOps control plane, and shadow evaluation toward a reproducible v1.0.

github.com/srgangaram-swe/Signalattice

AlphaForge: Regime-Aware Quantitative ML Research Platform

Built an end-to-end, regime-aware quantitative machine learning research platform for multi-horizon alpha signals on public or synthetic OHLCV data. The system includes causal feature engineering, a custom Gaussian HMM regime engine used strictly causally, embargoed walk-forward validation with purged K-Fold and combinatorial purged CV, deflated and probabilistic Sharpe ratios with probability-of-backtest-overfitting statistics, out-of-sample-only costed backtests, portfolio construction with regime-aware exposure, risk analytics, reports, a Streamlit dashboard, a FastAPI service, and a C++17 limit-order-book execution core with pybind11 bindings. Recent milestones delivered a self-financing backtest with defensible execution timing, cash accounting, fills, and costs, and neural temporal alpha models with an inspectable training protocol and first-class evaluation plots; institutional portfolio construction with point-in-time covariance and factor risk, and release-grade research governance, are next on the roadmap.

github.com/srgangaram-swe/AlphaForge

Learning Systems Atlas: Evidence-Driven ML Systems Portfolio

Building a reproducible, production-grade portfolio of supervised, unsupervised, deep, and reinforcement learning systems, from mathematical foundations and classical estimators to neural models and decision policies. Every reference experiment has typed configurations, explicit seed streams, leakage controls, a naive baseline, held-out evaluation, versioned artifacts, diagnostic plots, and unit plus integration tests. Sprints 1 and 2 are complete, covering strict experiment contracts, a transactional runner with SHA-256 manifests, and from-scratch supervised systems including CART, random forests, gradient boosting, k-NN, kernel SVM, and naive Bayes; the roadmap continues through unsupervised learning, a from-scratch autograd engine and PyTorch infrastructure, reinforcement learning, benchmarks and documentation, an algorithms-and-proofs library, and distributed ML systems and scale engineering.

github.com/srgangaram-swe/learning-systems-atlas

Conductor: Agentic AI Dashboard for Software Development

Developing a personal agentic AI dashboard tool for software development and workflow optimization. The project focuses on coordinating development tasks, surfacing project state, supporting debugging and documentation workflows, and helping developers move from intent to verified implementation.

PolytopeOS and the Polytope Compiler

Building a from-scratch research operating system and compiler toolchain with a constraint-aware resource plane for scientific and engineering workloads. PolytopeOS treats workload intent, reproducibility, and resource budgets as first-class kernel and runtime concepts: declarative CPU, GPU, memory, latency, energy, and locality objectives with observable scheduling decisions, hermetic reproducible workspaces, and topology and accelerator awareness for HPC, distributed training and inference, and low-jitter workloads, on a memory-safe Rust systems core. The Polytope Compiler is developed alongside the OS with explicit intermediate representations and diagnostics.

github.com/srgangaram-swe/polytope-os

Parallel Principal Component Analysis with MPI

Designed distributed PCA with rank-local sharding, global mean and covariance reductions, eigenpair broadcast, local projection, optional whitening, metadata logging, and scalability benchmarks.

Large-File Scientific ML Inference

Built overlap-aware slab processing and microbatching for large scientific data, improving memory behavior, output validation, manifest tracking, and reproducible inference with Dask, PyTorch, and NetCDF-oriented data products. The workflow is designed for fast inference across large collections of multi-terabyte files, including workloads that may require near-real-time throughput.

Signal Processing and Workflow Reliability

Developed public-safe signal characterization workflows using spectral analysis, spectrograms, power estimates, stationarity tests, runtime tracing, and validation reports to improve confidence in noisy scientific data products.

Real-Time Human Emotion Recognition

Developed a CNN-based computer vision pipeline using public image datasets, augmentation, regularization, feature extraction, and real-time video inference.

Financial Risk Modeling

Developed probability-of-default models with gradient boosting, neural networks, feature engineering, Bayesian hyperparameter optimization, stress testing, and error analysis.

Scientific Database and Web Tooling

Designed database-backed research tools for search, plotting, export, schema inspection, data selection, and reproducible modeling workflows.

Quantitative Finance and Portfolio Optimization

Built a Python-based financial analysis system using portfolio objects, numerical analysis, and optimization methods to evaluate performance metrics and risk-aware allocation strategies.

Epidemiological SIR and Network Modeling

Implemented SIR and network-based epidemic models using public health datasets, numerical integration, and error analysis to study spread dynamics at regional scales.

Recognition and Presentations

Technical Skills

Languages: Python, C++, C#, SQL, Bash, R, Java, JavaScript, HTML, CSS, Julia, MATLAB, Scala, Fortran, TypeScript, Rust

ML, LLMs, and Data Science: PyTorch, TensorFlow, Keras, Scikit-learn, NumPy, Pandas, generative AI, Codex, Claude Code, LLM APIs, agentic AI systems, multi-agent orchestration, tool-calling workflows, RAG, vector search, LLM evaluation harnesses, transformer implementation, supervised learning, unsupervised learning, clustering, regression, CNNs, U-Net-style segmentation, LSTMs, Transformers, anomaly detection, forecasting

HPC, Big Data, and Systems: MPI, Dask, Slurm, distributed workflows, scalable analytics, distributed training data workflows, multithreading, synchronization primitives, socket programming, Docker, Linux, Windows, Git, CI/CD

Data and Visualization: PostgreSQL, SQLite, DuckDB, BigQuery, HDF5, NetCDF, Parquet, TDMS-style scientific data, vector databases, OHLCV market data, dataset versioning, lineage, data quality gates, deduplication, terabyte-scale datasets, distributed data processing, Matplotlib, Seaborn, Tableau, Grafana, technical reporting

Analysis: Statistical inference, regression, Bayesian analysis, time series, signal processing, spectral methods, numerical computation, validation, error analysis, quantitative ML, walk-forward validation, backtesting, portfolio construction, risk analytics

Software Practices: Requirements analysis, integration testing, workflow hardening, documentation, reproducible batch execution, FastAPI, Streamlit, LLM-assisted development with human validation, enterprise workflow automation

Public Information Note

This public CV intentionally omits sensitive access details, restricted system names, operational parameters, non-public datasets, and program-specific details.