Portfolio · AI platforms

Mohammadreza “Hamid” Matiny

AI Infrastructure & MLOps Engineer specializing in LLM serving, GPU orchestration, and production observability.

Currently: Shipping AI infrastructure and data platforms across 30 public repositories.

Platform engineering for models that have to ship.

I’m Mohammadreza (“Hamid”) Matiny — an AI infrastructure and MLOps engineer with 4 years across deep learning systems, data pipelines, software backend systems, and cloud infrastructure on GCP and AWS. My work sits where model code meets the platform: serving contracts, GPU scheduling, experiment provenance, and the observability that makes production behavior trustworthy.

Today I build AI platform infrastructure — LLM serving, GPU orchestration, fleet telemetry, and computer-vision pipelines — alongside agentic-AI security work (policy-as-code gateways, tool-abuse controls, audit trails). That work spans 30 public repositories, 4 of them production-shaped platforms with real CI/CD and Terraform-provisioned infrastructure — and where it matters, measured results instead of estimates: one lakehouse run processed 42,972 records at a 91.8% acceptance rate with zero post-gate rejections.

I care about honest maturity signals: tagged releases when they’re real, architecture diagrams that match the repo, and metrics that come from measured runs — not marketing copy.

4 years Deep learning systems, data pipelines, software backend, and cloud infrastructure
4 platforms Production-shaped systems: multi-backend LLM serving, streaming lakehouses, CV detection
30 repos Public repositories spanning LLM infra, data engineering, and computer vision

Case studies, not just repos.

Architecture, engineering decisions, and honest maturity — verified against public GitHub sources.

Vulcan

Active on main

Multi-backend LLM serving and GPU-orchestration platform — one contract across vLLM, Triton, Ray Serve, KServe, and BentoML.

Tagged v1.2.0 across 22 phases: advanced GPU serving (GPTQ/AWQ/FP8, TensorRT-LLM templates), cost-per-token tracking, training backends, LoRA/PEFT, DVC exports, pluggable MLflow/W&B tracking, and a tool-grounded LangGraph advisor with non-fabrication CI.

Architecture

Clients and a LangGraph advisor hit a routing gateway that speaks a unified model-serving contract (/health, /metrics, /v1/infer). Backends — BentoML, Ray Serve, Triton (+ TensorRT-LLM), vLLM (+ GPTQ/AWQ/FP8 packs), KServe — plug in behind that contract. Training jobs (Ray Train, FSDP/DDP, DeepSpeed, LoRA/PEFT) share a TrainingJobSpec; MLflow and W&B track runs; DVC versions exports. GPU infra (Kueue / Karpenter / MIG) is validated in CI without burning real GPU cost. SageMaker and Bedrock are selectable managed paths. Observability (Prometheus, Grafana, Tempo, cost-per-token) grounds every advisor number in real evidence.

Key decisions

  • Unified serving contract instead of per-backend APIs (ADR-001)
  • GPU cost-safety: validate-only infra in CI; no invented tokens/s (ADR-002, ADR-007)
  • Kueue multi-tenant scheduling + MIG partitioning strategy (ADR-003, ADR-004)
  • LangGraph advisor is tool-grounded — every stated number must appear in Prometheus/benchmark evidence (ADR-014)
  • Pluggable experiment tracking: MLflow self-hosted + W&B offline-only in CI (ADR-013)
  • vLLM
  • Triton
  • Ray Serve
  • KServe
  • BentoML
  • Kueue
  • Karpenter
  • TensorRT-LLM
  • LoRA/PEFT
  • DVC
  • MLflow
  • W&B
  • LangGraph
  • SageMaker
  • Bedrock

System shape

Vulcan architecture Clients and LangGraph advisor connect to a gateway and model-serving contract, which fans out to BentoML, Ray Serve, Triton, vLLM, and KServe. Training, GPU infra, and observability feed the same system. Clients console / API Gateway routing :9007 Advisor · LangGraph tool-grounded only Model serving contract /health · /metrics · /v1/infer BentoML Ray Serve Triton + TRT-LLM vLLM GPTQ/AWQ/FP8 KServe Training contract Ray · FSDP · LoRA Tracking + DVC MLflow · W&B GPU · Kueue/Karpenter/MIG Observability Managed SM · Bedrock

View repository →

Argus

Active on main

Production-shaped fleet telemetry platform — Kafka/Ray/Flink ingest, Iceberg + Dagster lakehouse, drift detection, OPA-backed incidents, and a read-only AI copilot.

Tagged v1.0.0 — CHANGELOG calls it the first production-shaped release (Phases 0–15). 45 Docker Compose services, 41 test files (including Kafka integration tests), and 6 active CI workflows (ci, docker-build, semgrep, e2e-nightly, load-nightly, chaos-nightly).

Architecture

Redpanda/MSK → Ray ingest → stream-processor QA gate (Flink option) → Iceberg + Trino lakehouse → Dagster/MLflow orchestration → drift-monitor (KS tests, embeddings, Evidently) → OPA-backed incident-engine (circuit breakers) → api-gateway (OIDC/Keycloak, OPA RBAC) → Next.js dashboard, plus a read-only Qdrant-RAG AI copilot with its own eval harness. Same container images run via Docker Compose locally or Terraform + Helm (one chart per service) + Argo CD app-of-apps on EKS.

Key decisions

  • Contract-first streaming path with an explicit QA gate before lakehouse writes
  • Iceberg + Dagster for reproducible lakehouse materialization
  • OPA policy for incident decisions — not prompt-only automation
  • Copilot is read-only against telemetry and runbooks, backed by an eval harness
  • Documented scope cuts (KNOWN_GAPS.md) instead of overclaiming — no service mesh/mTLS, Vault-backed secrets, or column-level lineage yet
  • Kafka
  • Redpanda
  • Ray
  • Flink
  • Iceberg
  • Trino
  • Dagster
  • OpenTelemetry
  • OPA
  • Argo CD
  • Terraform
  • Next.js
  • Qdrant

System shape

Argus pipeline architecture Streaming path from Kafka or Redpanda through Ray, Flink QA, Iceberg lakehouse, Dagster, drift monitor, incident engine, OpenTelemetry, dashboard, and AI copilot. Kafka / Redpanda Ray Flink QA gate Iceberg lakehouse Dagster Drift monitor Incident engine · OPA OpenTelemetry Dashboard AI copilot

View repository →

PRISM

Active on main

Multi-warehouse fleet-intelligence platform — camera/sensor ingest, a PySpark lakehouse with dbt gold models, and OpenCV/ONNX defect detection with human review.

Tagged v1.2.0 — all 20 phases (0–19) complete, including a golden-path chaos e2e test against the live Compose stack. 17-service Docker Compose stack, 19 test files, 2 CI workflows (lint/test, Terraform validate + release packaging).

Architecture

Camera/sensor ingest lands in a bronze zone, then splits: a PySpark medallion lakehouse (bronze → silver → gold, dbt-modeled) on one side, an OpenCV/ONNX YOLO-family CV service on the other, routing low-confidence findings to a Django control-plane review queue. Gold data fans out through one activation contract to both Redshift and Snowflake, mirrors to Azure Databricks/ADLS for DR, and feeds a Vue 3 + Three.js digital-twin cockpit. An incident-engine runs per-asset circuit breakers on OPA/Rego trip policies; Dagster orchestrates the lakehouse and drift-monitor; a tool-grounded AI copilot answers only from evidence it can cite.

Key decisions

  • Two-layer Pydantic → Pandera validation gate before any bronze promotion
  • Per-asset circuit breakers driven by OPA/Rego trip policies, not hardcoded thresholds
  • Drift-monitor baselines never build from synthetic scenario data — health stays non-ready until a real baseline earns it
  • Copilot is tool-grounded — every answer must cite real evidence (ADR-004)
  • Cloud paths (AWS + Azure Terraform) are plan/validate/checkov-only in CI; human apply only (ADR-001)
  • PySpark
  • Databricks
  • dbt
  • OpenCV
  • ONNX / YOLO
  • Django
  • Vue 3 + Three.js
  • OPA / Rego
  • Dagster
  • Terraform
  • Snowflake
  • Redshift

System shape

PRISM architecture Fleet ingest into a bronze zone, splitting into an OpenCV/ONNX CV service with human review and a PySpark lakehouse with dbt gold models, both feeding a control plane and activation gateway that serve Redshift, Snowflake, a digital-twin cockpit, and an AI copilot. Fleet ingest Bronze zone CV service OpenCV + ONNX/YOLO PySpark lakehouse bronze → silver → gold Review queue Gold · dbt models Control plane Django + RBAC Activation gateway → Redshift + Snowflake Cockpit + AI copilot

View repository →

FORGE

Active on main

Offline AV perception & auto-labeling platform — 2D/3D detection, tracking, sensor fusion, and active-learning pseudo-labeling over a versioned Parquet data lake.

Tagged v0.2.0 — all 11 phases complete (ingest through visualize plus productionization). 15 test files, 2 CI workflows. Detection heads are randomly initialized research baselines, not trained on real labels — documented honestly in KNOWN_GAPS.md rather than overclaimed.

Architecture

The forge CLI runs an 8-stage pipeline end to end locally: ingest (nuScenes → Parquet lake, DVC, Hydra configs) → detect2d (Faster R-CNN) and detect3d (PointNet-style) → track (SORT: Kalman + Hungarian IoU) → fuse (calibrated projection + IoU) → label (trust scoring + active learning) → evaluate (BEV distance, mAP against held-out ground truth) → curate (LanceDB dedup) → visualize (rerun.io / Foxglove MCAP). A parallel cloud path — S3 → Lambda → SQS → DynamoDB → EventBridge → Step Functions → ECS Fargate → Glue (11 tables) → Athena — is Terraform-defined and structurally verified in CI, but intentionally never applied against live AWS.

Key decisions

  • Ground-truth labels used only for evaluation, never as a pipeline input — no label leakage
  • Every pipeline stage round-trips through a versioned Parquet lake instead of ad hoc intermediate files
  • Cloud orchestration built and CI-verified as code, deliberately never deployed — same cost-safety policy as Vulcan, PRISM, and hydra-data-factory
  • Ray distributed execution wired for the two heaviest stages (detect2d/detect3d); the rest run local by design
  • Random-init detectors are labeled as smoke-tested baselines, not tuned models — no inflated accuracy claims
  • PyTorch Lightning
  • Faster R-CNN
  • SORT
  • Ray
  • LanceDB
  • MLflow
  • W&B
  • DVC
  • Hydra
  • Terraform
  • Parquet

System shape

FORGE architecture Local pipeline from ingest through 2D and 3D detection, tracking, sensor fusion, active-learning labeling, then evaluation, curation, and visualization. A parallel cloud infrastructure path is Terraform-verified but never deployed against live AWS. nuScenes → Parquet Cloud infra Terraform · verified, never applied detect2d Faster R-CNN detect3d PointNet-style track SORT: Kalman + IoU fuse calibrated projection label trust scoring + active learning evaluate · mAP vs GT curate · LanceDB dedup visualize · rerun/MCAP

View repository →

aegis

Active on main

AI-native defense-in-depth gateway for LLM apps and agents — prompt injection, data exfiltration, and tool/MCP abuse, with a tamper-evident audit trail.

Tagged v0.3.1 — all 12 build-order stages complete. 11-service Docker Compose stack, 53 test files, 3 CI workflows (ci, release, security). Publishes its own adaptive red-team result instead of a vendor catch-rate demo: real-model hardening cuts round-1 bypass rate (10.8% → 9.2%) but overall bypass rate under sustained adaptive attack stays flat at ~48%.

Architecture

A Go gateway fronts the defended chat pipeline: requests pass through Python input-defense detectors, a Go policy-engine evaluating CEL policy packs, and a provider-agnostic Go model-router, then Python output-defense detectors (plus an LLM judge) before any response is released. A separate agent-gate (Go) enforces tool/MCP call permissions with taint tracking. Every enforcement decision is Ed25519-signed into a Postgres-backed audit trail. A continuous red-team engine runs adaptive campaigns directly against input/output defense — deliberately bypassing policy-engine — and publishes bypass-rate evidence to a React/TS dashboard.

Key decisions

  • Security decisions fail closed (gateway/policy-engine/agent-gate outage returns 502, never releases an unchecked response) — only observability fails open, documented explicitly in FAILURE_MODES.md
  • No static default credentials anywhere in the repo — dashboard and gateway keys are generated at container startup or via a credential script
  • Publishes its own adaptive red-team bypass rates honestly, including the unflattering result that hardening barely moves sustained bypass rate — continuous monitoring over a one-time "solved" claim
  • Red-team probes input/output defense directly, deliberately bypassing policy-engine, to measure detector effectiveness in isolation
  • Tamper-evident audit trail via Ed25519-signed receipts, not plain logs
  • Go
  • Python
  • TypeScript
  • CEL policy-as-code
  • Prompt-Guard
  • Toxic-BERT
  • spaCy NER
  • PostgreSQL
  • Redis
  • React

System shape

aegis architecture A Go gateway routes requests through input defense, a CEL policy engine, and a model router before output defense releases a response. A parallel agent-gate enforces tool permissions. Every decision is signed into an audit trail shown on a dashboard, while a red-team engine continuously probes input and output defense directly. App / SDK Gateway (Go) Input defense Python detectors Policy engine Go + CEL Model router provider-agnostic Agent-gate tool / MCP taint Output defense detectors + judge Audit Ed25519-signed receipts Dashboard Red-team engine adaptive campaigns

View repository →

Production-validated AV telemetry lakehouse — contract validation, DLQ isolation, Terraform-provisioned AWS path.

Production-validated pipeline with measured run results: 42,972 records ingested · 91.8% acceptance · 0% post-gate rejections.

Architecture

Ingest mock fleet JSON → Pydantic + Pandera contract validation with DLQ isolation for rejects → PyArrow/Parquet (Snappy, Hive partitioning) into S3 → Glue catalog. Dual orchestration: local Airflow + MLflow, and AWS Step Functions + Lambda. Infra fully provisioned with Terraform (S3, Glue, IAM).

Key decisions

  • Dual-path orchestration (Airflow local / Step Functions cloud) sharing one transformation core
  • Hard contract gate with DLQ isolation — rejects never contaminate the lake
  • Hive-partitioned Parquet for query-friendly AV telemetry
  • Everything provisioned as Terraform — no console-only resources
  • Pydantic
  • Pandera
  • PyArrow
  • Parquet
  • Terraform
  • S3
  • Glue
  • Airflow
  • Step Functions
  • MLflow

System shape

hydra-data-factory architecture Fleet JSON ingest through Pydantic and Pandera validation with DLQ isolation, then Parquet to S3 with Glue catalog, orchestrated by Airflow or Step Functions via Terraform. Fleet JSON ingest Pydantic + Pandera contract gate DLQ isolation Parquet · Hive S3 + Glue Airflow + MLflow Step Functions Terraform

View repository →

Where the work actually concentrates.

Drawn from 30 public repositories spanning LLM infra, data engineering, and computer vision — organized by the systems I ship, not an exhaustive badge wall.

LLM Serving & Inference

  • vLLM
  • Triton Inference Server
  • TensorRT-LLM
  • Ray Serve
  • KServe
  • BentoML
  • GPTQ / AWQ / FP8
  • LoRA / PEFT
  • Bedrock
  • SageMaker

GPU Orchestration & Cloud

  • Kubernetes
  • Kueue
  • Karpenter
  • NVIDIA MIG
  • AWS
  • GCP
  • Terraform
  • Helm
  • Argo CD
  • Docker

MLOps & Experiment Tracking

  • MLflow
  • Weights & Biases
  • DVC
  • GitHub Actions
  • Dagster
  • Kubeflow
  • Prometheus
  • Grafana
  • OpenTelemetry

Computer Vision & Perception

  • OpenCV
  • ONNX
  • YOLO / ByteTrack
  • 2D/3D detection & tracking
  • Sensor fusion (radar / LiDAR / camera)
  • PyTorch
  • Active learning / pseudo-labeling
  • Edge AI

Data & Streaming Systems

  • Kafka / Redpanda
  • Apache Flink
  • Apache Iceberg
  • PyArrow / Parquet
  • Pandera
  • Pydantic
  • Airflow
  • Step Functions
  • Ray
  • PySpark / Databricks
  • dbt

Security & Agentic Systems

  • Prompt-injection defense
  • Policy-as-code (OPA)
  • LangGraph
  • PydanticAI
  • Temporal
  • Audit trails
  • Red-teaming patterns

Let’s talk platforms.

Open to conversations about AI infrastructure, LLM serving, MLOps platforms, and agentic security — especially roles where production reliability matters as much as model quality.

Available globally · open to remote