Enes Uzun
I'm an AI infrastructure engineer at Ford Otosan in Istanbul. I run the GPU cluster and LLM serving stack that 15,000 employees and 250 developers use, and I write here about inference, GPUs and coding agents.
Notes
- The first note is on its way.
Selected work
- An LLM gateway that keeps confidential data on-premRoutes each request to open models on our own GPUs or to Azure OpenAI, depending on what the data is allowed to do.
- triton-server-hpaScales NVIDIA Triton on Kubernetes by GPU utilisation instead of CPU, with GPU time-slicing so it runs on one card.
- Root-causing agent stalls in jcodeThree bugs in a Rust coding agent, found while running it against self-hosted vLLM and SGLang. All fixed upstream.