Artificial Intelligence ,000 – ,000 + Equity
Head of AI Infrastructure & LLM Systems
Location: Seattle, WA / Remote Search Type: Retained Executive Search Posted: Just posted
Role Overview & Executive Context
Directing large-scale GPU cluster infrastructure, distributed LLM inference pipelines, model optimization, and high-performance computing (HPC) engineering for next-generation generative AI models.
Key Responsibilities
- Architect scalable GPU compute clusters (H100/B200) for large language model pre-training
- Optimize inference latency, memory bandwidth, and distributed PyTorch / vLLM execution
- Lead a team of 15+ senior infrastructure, MLOps, and systems engineers
- Collaborate with AI Research Scientists to deploy foundation models to production
Required Qualifications & Leadership Experience
- 8+ years in distributed systems engineering and high-performance compute architecture
- Hands-on mastery of CUDA, Triton, PyTorch, Ray, Kubernetes, and InfiniBand networking
- Track record of operating 1,000+ GPU production clusters for LLM inference or training
- MS or PhD in Computer Science, Computer Engineering, or related quantitative field
Confidential Application
Your identity remains strictly confidential until authorized.