High Performance Scientific Computing
Parallel algorithms, distributed workflows, MPI-based numerical methods, Dask pipelines, and batch execution on HPC systems.
University of New Mexico Computer Science
Data Scientist 2 at LANL and UNM Computer Science Ph.D. student doing active research in AI/ML, HPC, Big Data, and systems, with a focus on scalable systems and algorithms for AI/ML training and inference.
About
I am a Computer Science Ph.D. student at the University of New Mexico with a background in mathematics, data science, software engineering, and applied machine learning. I am doing active research at both Los Alamos National Laboratory and UNM in AI/ML, high performance computing, Big Data, and systems, and much of my current work centers on scalable systems and scalable algorithms for AI/ML training and inference.
I am advised by Dr. Amanda Bienz at UNM. Professionally, I currently work as a Data Scientist 2 at Los Alamos National Laboratory, after previously working as a Software Engineer in Automation and AI/ML at Space Dynamics Laboratory and as a Teaching Assistant at the University at Buffalo.
I am especially interested in scalable algorithms, high-throughput data systems, time-frequency analysis, robust ML inference workflows, and the engineering practices that turn research code into dependable software. I am also building PolytopeOS and the Polytope Compiler, a from-scratch research operating system and compiler toolchain. This public site intentionally keeps project descriptions technical, high level, and suitable for open audiences.
Research
Parallel algorithms, distributed workflows, MPI-based numerical methods, Dask pipelines, and batch execution on HPC systems.
Scalable systems and scalable algorithms for AI/ML training and inference: distributed training, high-throughput and low-latency inference, performance engineering, and low-level systems foundations including compilers and operating systems.
Current research includes Big Data methods for large scientific datasets, distributed processing, scalable storage formats, and analytics workflows for terabyte-scale data collections.
Denoising, spectral analysis, stationarity testing, transient detection, and quality checks for high-volume sensor data.
Generative AI, supervised and unsupervised learning, clustering, regression, deep learning, anomaly detection, segmentation, forecasting, and validation for large scientific datasets.
Reproducible ingestion, schema design, workflow hardening, data validation, and visualization for large experimental collections.
Selected Work
Phase 2 complete: a self-hosted, local-first agentic AI operations console that runs OpenAI and Anthropic agents in parallel, streams tokens and tool events live, and provides file context, persistent memory, terminal control, AI-assisted email drafting, and a no-key mock mode.
View open-source repositoryA local-first data refinery and data-centric research platform for distributed training, now at Phase 8 of 8: batch and streaming ingestion into bronze/silver/gold Parquet layers, statistical quality gates with quarantine, exact and MinHash/LSH near-duplicate removal, content-addressed dataset versions with lineage, point-in-time feature layers, deterministic training shards with a CPU/DDP/FSDP reference trainer, DAG orchestration with a metadata API and dashboard, and multi-seed data ablations. v1.0 hardening is in progress.
View open-source repositoryBuilt a Codex and Claude Code-style internal agentic AI harness from scratch for approved enterprise LLM APIs, with multi-agent orchestration, software-development assistance, collaboration context retrieval, meeting summarization, email and workflow drafting, and document generation for controlled internal environments.
A production-style research platform for causal neural telemetry decoding, probabilistic forecasting, and closed-loop BCI systems, with classical and CNN/transformer decoders, self-supervised pretraining, and publication-ready reports. Sprint 1's leakage-safe causal bidirectional telemetry benchmark is complete, with a roadmap spanning governed public electrophysiology data, multimodal neuroimaging, a closed-loop BCI control plane, and a deterministic real-time edge runtime.
View open-source repositoryA statistically rigorous, full-stack LLM systems lab with a tiny GPT trained from scratch, reusable model and agent evaluation suites, cited RAG over documents, an agentic research assistant, and local-first offline fallbacks. The latest release ships a truthful model-backed chat vertical slice with a streaming API, accessible UI, and full-stack CI, with a roadmap through statistical rigor and uncertainty, deep learning systems, and a grounded production platform.
View open-source repositoryCalibrated probabilistic market forecasting, from point-in-time data to decision-ready evidence. v0.2 shipped calibrated forecasts with proper scoring rules, causal deep learning, walk-forward tradability evidence, and a model-risk go/no-go report, on top of reproducible ingestion, leakage-aware features, and costed backtesting. The roadmap adds a bitemporal point-in-time data mesh, Bayesian cross-asset scenario analysis, and a ForecastOps control plane.
View open-source repositoryA regime-aware quantitative ML research platform for multi-horizon alpha signals: leakage-safe walk-forward validation with purged K-Fold and combinatorial purged CV, a causal Gaussian HMM regime engine, deflated Sharpe and probability-of-backtest-overfitting statistics, a self-financing backtest with defensible cash accounting and fills, neural temporal alpha models with first-class evaluation tooling, and a C++17 limit-order-book execution core. Institutional portfolio construction and factor risk are next on the roadmap.
View open-source repositoryA reproducible, production-grade portfolio of supervised, unsupervised, deep, and reinforcement learning systems built around evidence rather than demos. Sprints 1 and 2 are complete: strict experiment contracts, a transactional runner with SHA-256 manifests, and from-scratch supervised systems spanning decision trees, random forests, gradient boosting, k-NN, kernel SVM, and naive Bayes. The roadmap continues through unsupervised learning, a from-scratch autograd engine, reinforcement learning, an algorithms-and-proofs library, and distributed ML systems and scale engineering.
View open-source repositoryDeveloping an agentic AI dashboard for software development and workflow optimization, with a focus on coordinating tasks, surfacing project state, supporting debugging workflows, and improving how developers move from intent to verified changes.
A from-scratch research operating system with a constraint-aware resource plane for scientific and engineering workloads: declarative CPU, GPU, memory, latency, and energy objectives, reproducible hermetic workspaces, topology and accelerator awareness for HPC and distributed training and inference, and a memory-safe Rust systems core. Developed alongside the Polytope Compiler, a from-scratch toolchain with explicit intermediate representations and diagnostics.
View open-source repositoryImplemented distributed PCA with rank-local data sharding, collective covariance aggregation, eigenpair broadcast, optional standardization, and scalability logging for large datasets.
Built a slabbed and microbatched inference path for large scientific files, replacing dense full-file processing with overlap-aware chunking, manifest tracking, and fast validation outputs for high-throughput and near-real-time workloads.
Trained and evaluated CNN-based models using large public image datasets, applying augmentation, regularization, feature extraction, and real-time video processing techniques.
Developed supervised learning pipelines for probability-of-default estimation with feature engineering, hyperparameter tuning, stress testing, and error analysis.
Designed database-backed analysis tools for large scientific collections, with searchable metadata, plotting, export workflows, and reproducible input generation for modeling studies.
Applied CNNs, U-Net style segmentation, clustering, and classical image processing to improve detection and denoising in scientific imagery and signal-derived products.
Built instrumentation, validation checks, runtime tracing, and portable execution paths for scientific data workflows, improving reproducibility across local, Linux, and HPC-style environments.
Background
Conducts active research and development across AI/ML, HPC, Big Data, and systems, including data acquisition, analysis, automation, signal processing, supervised and unsupervised ML, generative AI, and HPC-integrated workflows for engineering test and validation environments. Built an internal agentic AI workflow harness for approved LLM APIs, multi-agent software assistance, enterprise collaboration context retrieval, meeting summarization, and document/email automation in controlled internal environments.
Built data pipelines, database-backed web tools, scientific analysis workflows, and automation software for aerospace and remote-sensing research settings.
Received internal recognition for user-centered engineering, data tooling, documentation, and technical communication; presented telemetry anomaly-detection methods at a public technical venue.
Supported lab and recitation instruction, grading, exam administration, and student mentoring for undergraduate computer science coursework.
Ph.D. in Computer Science, expected 2028; advisor: Dr. Amanda Bienz
M.P.S. in Data Sciences and Applications, 2023
B.A. in Mathematics, Computing and Applied Mathematics, 2021
Methods
Contact
For public research, academic, or professional communication, email is the best way to reach me.
The descriptions on this site are intentionally limited to public-safe topics. They omit sensitive access details, restricted system names, operational parameters, non-public datasets, and program-specific details.