← Back to Jobs

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
Crusoe
San Francisco, CA; Sunnyvale, CA
Full-time
$179,000 - $218,000 base + bonus and equity
Posted Aug 10, 2026
Emerging Technologies & Specialty RolesGPU architecturepredictive maintenanceNVLinkInfiniBanddirect-to-chip cooling
Job Description
Serve as Crusoe’s technical authority for GPU platforms across data center engineering and operations. This specialist translates upcoming NVIDIA and AMD silicon roadmaps into facility power, cooling, rack-spacing, tooling, sparing, and operating requirements for large AI clusters.
The role builds predictive health monitoring from fleet telemetry, creates precision repair procedures and diagnostics, leads complex hardware root-cause analysis, and helps facilities prepare for next-generation 2,000-watt-class accelerators and direct-to-chip liquid cooling.
Requirements
• Ten or more years in hardware engineering, systems architecture, or data center infrastructure
• Expert knowledge of NVIDIA Hopper, Blackwell, or Rubin and AMD Instinct architectures
• Deep understanding of NVLink, NVSwitch, InfiniBand, HBM, PCIe, and large-scale GPU clusters
• Ability to translate silicon heat-load profiles into CDU sizing, secondary-loop, power, and rack requirements
• Python, Go, or Bash experience building telemetry and health-check tooling with platforms such as DCGM or ROCm
• Experience using failure telemetry for predictive maintenance, spares planning, and field-service workflows
• Deep practical knowledge of direct-to-chip liquid cooling and high-density thermal management
• Bachelor’s or master’s degree in electrical engineering, computer engineering, or a related field
