TALENTMATCH .NET
Artificial Intelligence ,000 – ,000 + Equity

Head of AI Infrastructure & LLM Systems

Location: Seattle, WA / Remote Search Type: Retained Executive Search Posted: Just posted

Role Overview & Executive Context

Directing large-scale GPU cluster infrastructure, distributed LLM inference pipelines, model optimization, and high-performance computing (HPC) engineering for next-generation generative AI models.

Key Responsibilities

  • Architect scalable GPU compute clusters (H100/B200) for large language model pre-training
  • Optimize inference latency, memory bandwidth, and distributed PyTorch / vLLM execution
  • Lead a team of 15+ senior infrastructure, MLOps, and systems engineers
  • Collaborate with AI Research Scientists to deploy foundation models to production

Required Qualifications & Leadership Experience

  • 8+ years in distributed systems engineering and high-performance compute architecture
  • Hands-on mastery of CUDA, Triton, PyTorch, Ray, Kubernetes, and InfiniBand networking
  • Track record of operating 1,000+ GPU production clusters for LLM inference or training
  • MS or PhD in Computer Science, Computer Engineering, or related quantitative field

Confidential Application

Your identity remains strictly confidential until authorized.

SECURE PORTAL ACCESS