SYDNEY · PLATFORM / AI INFRA
I build and run I build and run the platforms that run models in prod.
Bartu Ogur Platform & AI Infrastructure Engineer
I work at the layer everything else runs on — Kubernetes, AWS, and AI/LLM infrastructure — and keep it fast, secure, and standing on its own.
Capability
- platform
- Platform & Kubernetes. Production Kubernetes from bare metal up — kubeadm on Proxmox and EKS on AWS, Rook-Ceph storage, Traefik / MetalLB ingress, GitOps, and a Prometheus / Thanos / Grafana / Loki stack. Foundations that scale, self-heal, and give you peace of mind.
- cloud
- Cloud & AWS. AWS Solutions Architect (Professional) and primary owner of a multi-account production estate — EKS, VPC, identity. The FinOps behind it — Savings Plans, Reserved Instances, right-sizing — cut the annual bill ~30% and holds it down year over year, with no hit to delivery.
- ai · llm
- AI / LLM infrastructure. Large models in production: multi-GPU inference pools (vLLM) behind a LiteLLM gateway, plus RAG (Qdrant) and agentic orchestration (LangChain / MCP). I tune throughput and concurrency to balance three things at once — reasoning depth, per-user responsiveness, and how many users a cluster can serve — then turn those models into agentic platforms whole teams use daily.
- security
- Security. I came up through security, so I think about how systems get attacked before how they ship. Web-application pentesting on an automated stack — OWASP ZAP, Burp Suite, Nuclei, Nessus, Metasploit — alongside blue-team hardening, SSO everywhere I can reach, and secure remote access (OpenVPN, plus SOCKS5 WireGuard via Dante). The instinct has saved production more than once.
Selected work
all work →A company-wide agentic AI platform
60+ active users Sole owner of the secure, highly available AI platform 60+ staff depend on — rebuilt from a single-user open-source agent into a multi-tenant system with isolated per-user workspaces.Multi-Host GPU Inference Pool
~120 tok/s/user (~350 agg) throughput Designed and tuned a pool of mixed GPUs across several machines that serves a stack of self-hosted models — an 80B flagship, a faster mid-size model, and embeddings — fast enough for a whole team to use at once, without paying per-seat cloud bills.Production Kubernetes Platform, Built From Scratch
7 nodes Designed and operates the 7-node, self-managed Kubernetes platform that runs the company's AI, development, and production workloads end to end.Security as a Discipline: AppSec to Estate Hardening
5+ yrs (2 titled) doing security 5+ years owning security for a RegTech group — came up through offensive security, built a repeatable pre-release security gate, and now own hardening across the whole platform as it grows.About
I'm a platform, AI, and infrastructure engineer in Sydney. I work at the layer everything else runs on — the systems that have to stay up, scale, and recover on their own. I came up through security and support, so I think about how things fail before how they ship, and I move into unfamiliar territory fast: the distance between “I haven't done this” and “it's in production” tends to be short. I taught myself Kubernetes, Docker, and AWS on a product-support line and built from there. Off the clock I'm either under a barbell — strength training and bodybuilding, 5+ years — or upgrading the homelab: a Proxmox + Kubernetes lab that's quietly turning into a small data centre, where I prototype and break things before they reach production.
Contact
Hiring, building, or want to talk infrastructure? bartu@bartuogur.com