University of New Mexico Computer Science

Saif Ryan Gangaram

Data Scientist 2 at LANL and UNM Computer Science Ph.D. student doing active research in AI/ML, HPC, Big Data, and systems, with a focus on scalable systems and algorithms for AI/ML training and inference.

Current Role
Data Scientist 2, LANL
Research Areas
AI/ML, HPC, Big Data, Systems
Ph.D. Program
CS, UNM
Portrait of Saif Ryan Gangaram

About

Scientific computing for large, noisy, consequential data.

I am a Computer Science Ph.D. student at the University of New Mexico with a background in mathematics, data science, software engineering, and applied machine learning. I am doing active research at both Los Alamos National Laboratory and UNM in AI/ML, high performance computing, Big Data, and systems, and much of my current work centers on scalable systems and scalable algorithms for AI/ML training and inference.

I am advised by Dr. Amanda Bienz at UNM. Professionally, I currently work as a Data Scientist 2 at Los Alamos National Laboratory, after previously working as a Software Engineer in Automation and AI/ML at Space Dynamics Laboratory and as a Teaching Assistant at the University at Buffalo.

I am especially interested in scalable algorithms, high-throughput data systems, time-frequency analysis, robust ML inference workflows, and the engineering practices that turn research code into dependable software. I am also building PolytopeOS and the Polytope Compiler, a from-scratch research operating system and compiler toolchain. This public site intentionally keeps project descriptions technical, high level, and suitable for open audiences.

Research

Research Themes

High Performance Scientific Computing

Parallel algorithms, distributed workflows, MPI-based numerical methods, Dask pipelines, and batch execution on HPC systems.

Scalable AI/ML Systems

Scalable systems and scalable algorithms for AI/ML training and inference: distributed training, high-throughput and low-latency inference, performance engineering, and low-level systems foundations including compilers and operating systems.

Big Data and Scalable Analytics

Current research includes Big Data methods for large scientific datasets, distributed processing, scalable storage formats, and analytics workflows for terabyte-scale data collections.

Signal Processing and Time Series Analysis

Denoising, spectral analysis, stationarity testing, transient detection, and quality checks for high-volume sensor data.

AI/ML for Scientific Data

Generative AI, supervised and unsupervised learning, clustering, regression, deep learning, anomaly detection, segmentation, forecasting, and validation for large scientific datasets.

Reliable Data Systems

Reproducible ingestion, schema design, workflow hardening, data validation, and visualization for large experimental collections.

Selected Work

Projects

Open Source | Deno, TypeScript, WebSocket

Hyperion Z

Phase 2 complete: a self-hosted, local-first agentic AI operations console that runs OpenAI and Anthropic agents in parallel, streams tokens and tool events live, and provides file context, persistent memory, terminal control, AI-assisted email drafting, and a no-key mock mode.

View open-source repository
Data-Centric AI | Parquet, DuckDB, Training Data

Crucible

A local-first data refinery and data-centric research platform for distributed training, now at Phase 8 of 8: batch and streaming ingestion into bronze/silver/gold Parquet layers, statistical quality gates with quarantine, exact and MinHash/LSH near-duplicate removal, content-addressed dataset versions with lineage, point-in-time feature layers, deterministic training shards with a CPU/DDP/FSDP reference trainer, DAG orchestration with a metadata API and dashboard, and multi-seed data ablations. v1.0 hardening is in progress.

View open-source repository
LANL | Agentic AI, LLM APIs, Workflow Automation

Internal Agentic AI Workflow Harness

Built a Codex and Claude Code-style internal agentic AI harness from scratch for approved enterprise LLM APIs, with multi-agent orchestration, software-development assistance, collaboration context retrieval, meeting summarization, email and workflow drafting, and document generation for controlled internal environments.

ML Research | PyTorch, Transformers, BCI

SynaptiCast

A production-style research platform for causal neural telemetry decoding, probabilistic forecasting, and closed-loop BCI systems, with classical and CNN/transformer decoders, self-supervised pretraining, and publication-ready reports. Sprint 1's leakage-safe causal bidirectional telemetry benchmark is complete, with a roadmap spanning governed public electrophysiology data, multimodal neuroimaging, a closed-loop BCI control plane, and a deterministic real-time edge runtime.

View open-source repository
LLM Systems | PyTorch, RAG, FastAPI

Dork LLM

A statistically rigorous, full-stack LLM systems lab with a tiny GPT trained from scratch, reusable model and agent evaluation suites, cited RAG over documents, an agentic research assistant, and local-first offline fallbacks. The latest release ships a truthful model-backed chat vertical slice with a streaming API, accessible UI, and full-stack CI, with a roadmap through statistical rigor and uncertainty, deep learning systems, and a grounded production platform.

View open-source repository
Quant ML | Probabilistic Forecasting, Backtesting

SignaLattice

Calibrated probabilistic market forecasting, from point-in-time data to decision-ready evidence. v0.2 shipped calibrated forecasts with proper scoring rules, causal deep learning, walk-forward tradability evidence, and a model-risk go/no-go report, on top of reproducible ingestion, leakage-aware features, and costed backtesting. The roadmap adds a bitemporal point-in-time data mesh, Bayesian cross-asset scenario analysis, and a ForecastOps control plane.

View open-source repository
Quant ML | Alpha Signals, Risk Analytics

AlphaForge

A regime-aware quantitative ML research platform for multi-horizon alpha signals: leakage-safe walk-forward validation with purged K-Fold and combinatorial purged CV, a causal Gaussian HMM regime engine, deflated Sharpe and probability-of-backtest-overfitting statistics, a self-financing backtest with defensible cash accounting and fills, neural temporal alpha models with first-class evaluation tooling, and a C++17 limit-order-book execution core. Institutional portfolio construction and factor risk are next on the roadmap.

View open-source repository
ML Systems | Python, Reproducible Research

Learning Systems Atlas

A reproducible, production-grade portfolio of supervised, unsupervised, deep, and reinforcement learning systems built around evidence rather than demos. Sprints 1 and 2 are complete: strict experiment contracts, a transactional runner with SHA-256 manifests, and from-scratch supervised systems spanning decision trees, random forests, gradient boosting, k-NN, kernel SVM, and naive Bayes. The roadmap continues through unsupervised learning, a from-scratch autograd engine, reinforcement learning, an algorithms-and-proofs library, and distributed ML systems and scale engineering.

View open-source repository
Personal Project | Agentic AI, Developer Tools

Conductor

Developing an agentic AI dashboard for software development and workflow optimization, with a focus on coordinating tasks, surfacing project state, supporting debugging workflows, and improving how developers move from intent to verified changes.

Systems | Rust, Compilers, Operating Systems

PolytopeOS

A from-scratch research operating system with a constraint-aware resource plane for scientific and engineering workloads: declarative CPU, GPU, memory, latency, and energy objectives, reproducible hermetic workspaces, topology and accelerator awareness for HPC and distributed training and inference, and a memory-safe Rust systems core. Developed alongside the Polytope Compiler, a from-scratch toolchain with explicit intermediate representations and diagnostics.

View open-source repository
UNM | MPI, NumPy, mpi4py

Parallel Principal Component Analysis

Implemented distributed PCA with rank-local data sharding, collective covariance aggregation, eigenpair broadcast, optional standardization, and scalability logging for large datasets.

Scientific ML | Python, Dask, PyTorch

Large-File Inference Workflow

Built a slabbed and microbatched inference path for large scientific files, replacing dense full-file processing with overlap-aware chunking, manifest tracking, and fast validation outputs for high-throughput and near-real-time workloads.

Graduate Project | Computer Vision

Real-Time Emotion Recognition

Trained and evaluated CNN-based models using large public image datasets, applying augmentation, regularization, feature extraction, and real-time video processing techniques.

Graduate Project | Data Mining

Financial Risk Modeling

Developed supervised learning pipelines for probability-of-default estimation with feature engineering, hyperparameter tuning, stress testing, and error analysis.

Software Engineering | Databases

Scientific Data Access Tools

Designed database-backed analysis tools for large scientific collections, with searchable metadata, plotting, export workflows, and reproducible input generation for modeling studies.

AI/ML | Image and Signal Data

Segmentation and Detection Models

Applied CNNs, U-Net style segmentation, clustering, and classical image processing to improve detection and denoising in scientific imagery and signal-derived products.

Scientific Software | Workflow Reliability

Reproducible Pipeline Hardening

Built instrumentation, validation checks, runtime tracing, and portable execution paths for scientific data workflows, improving reproducibility across local, Linux, and HPC-style environments.

Background

Experience and Education

2026 - Present

Data Scientist 2, Los Alamos National Laboratory

Conducts active research and development across AI/ML, HPC, Big Data, and systems, including data acquisition, analysis, automation, signal processing, supervised and unsupervised ML, generative AI, and HPC-integrated workflows for engineering test and validation environments. Built an internal agentic AI workflow harness for approved LLM APIs, multi-agent software assistance, enterprise collaboration context retrieval, meeting summarization, and document/email automation in controlled internal environments.

2023 - 2026

Software Engineer, Space Dynamics Laboratory

Built data pipelines, database-backed web tools, scientific analysis workflows, and automation software for aerospace and remote-sensing research settings.

2024 - 2025

Technical Recognition and Presentation

Received internal recognition for user-centered engineering, data tooling, documentation, and technical communication; presented telemetry anomaly-detection methods at a public technical venue.

2019 - 2020

Teaching Assistant, University at Buffalo

Supported lab and recitation instruction, grading, exam administration, and student mentoring for undergraduate computer science coursework.

University of New Mexico

Ph.D. in Computer Science, expected 2028; advisor: Dr. Amanda Bienz

University at Buffalo

M.P.S. in Data Sciences and Applications, 2023

University at Buffalo

B.A. in Mathematics, Computing and Applied Mathematics, 2021

Methods

Technical Areas

Python C++ C# SQL MATLAB Julia PyTorch TensorFlow Scikit-learn Codex Claude Code LLM APIs Agentic AI Multi-Agent Systems RAG LLM Evaluation FastAPI Streamlit Quant ML Backtesting Portfolio Construction Risk Analytics NumPy Pandas Dask Slurm MPI PostgreSQL SQLite NetCDF HDF5 Parquet DuckDB Data Quality Dataset Versioning Deduplication Docker Linux Git Time Series Spectral Analysis Visualization

Contact

Open to research conversations and technical collaboration.

For public research, academic, or professional communication, email is the best way to reach me.

Public Information Note

The descriptions on this site are intentionally limited to public-safe topics. They omit sensitive access details, restricted system names, operational parameters, non-public datasets, and program-specific details.