Role Description DigitalOcean is expanding its AI Infrastructure layer to support the next generation of AI-driven applications. We are seeking a Senior Engineer 2 to join our AI Inference Data Plane team. In this role, you will be a key technical leader responsible for designing, developing, and delivering high-scale, resilient data plane services that power our "Inference as a Service" offering. Technical Leadership: Act as a technical leader on the team, driving the end-to-end design, development, and delivery of critical data plane components hosting large generative AI models. System Design: Architect and refine system design proposals for our high-scale, multi-tenant AI inference cloud ecosystem, ensuring they meet rigorous availability and resiliency standards. Performance Optimization: Implement and optimize distributed inference hosting using techniques like tensor/data parallelism, KV cache optimizations, and smart routing. Collaboration: Work cross-functionally with Product Managers, customer-facing teams, and other engineering teams to align technical roadmaps with customer needs. Distributed Serving at Scale: Build on Kubernetes-native distributed inference frameworks like llm-d (or alternatives such as NVIDIA Dynamo, Ray Serve) to deliver prefill/decode disaggregation, KV-cache-aware routing, tiered prefix caching, and wide expert parallelism for MoE models. Flow Control flow control and fairness across tenants; autoscaling inference pools; and moving gigabytes of KV-cache between prefill and decode instances with negligible overhead. Open Source Contributions: Contribute upstream to llm-d, vLLM, and the inference gateway ecosystem, and represent DigitalOcean in these communities. Mentorship: Coach and mentor junior engineers, fostering a culture of technical excellence and continuous improvement. Operational Excellence: Maintain and operate critical, high-scale services, utilizing observability tools and defining SLOs to ensure superior platform health. Qualifications Hands-on experience hosting large language or multimodal models using inference engines like vLLM, SGLang, or TensorRT. Familiarity with distributed inference serving frameworks such as llm-d, NVIDIA Dynamo, or Ray Serve. Hands-on experience with vLLM or alternatives (SGLang, TensorRT-LLM, TGI, Modular MAX), including internals like continuous batching, paged attention, and prefix caching. Understanding of why cluster-scale serving is hard: KV-cache locality is partitioned across workers, naive round-robin routing destroys cache hit rates and tail latency, and disaggregated prefill/decode requires fast cross-pod KV transfer (e.g., NIXL). Merged contributions to vLLM, llm-d, SGLang, or similar projects strongly preferred. Knowledge of common LLM architectures and optimization techniques (e.g., continuous batching, quantization). Expert-level proficiency in GoLang or Python and familiarity with gRPC. Proven experience shipping customer-facing software products and running critical services in a high-scale environment similar to DigitalOcean. Experience integrating and building with open-source software. Requirements Compensation Range: $139,200 - $174,000 This is a remote role. JR: 2026-7624 #LI-Remote Benefits Competitive array of benefits to support employee well-being. Reimbursement for relevant conferences, training, and education. Access to LinkedIn Learning's 10,000+ courses for continued growth and development. Equity compensation to eligible employees, including equity grants upon hire and the option to participate in our Employee Stock Purchase Program.
Read Less