Partner with NVIDIA’s software, research, architecture, and product teams to align technical requirements and strategic priorities, fostering the AI ecosystem on RTX and DGX PCs.
Build and optimize the local AI inference stack for RTX, RTX Pro, and DGX GPUs, with a focus on performance, stability, and scalability across diverse hardware architectures.
Design and develop modern inference runtimes and execution stacks using frameworks such as llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX, supporting LLM, vision-language, TTS, ASR, and diffusion-based AI workloads.
Perform end-to-end optimization of AI models, data pipelines, and inference runtimes to maximize performance on current and next-generation GPU architectures. Apply model optimization techniques, including quantization, pruning, sparsity, and distillation, to enable efficient deployment of large models on local and edge devices.
Conduct system-level debugging, performance tuning, and performance-accuracy trade-off analysis; develop infrastructure for performance and accuracy sweeps; analyse results to identify gaps and drive fixes; and establish engineering guidelines to accelerate bring-up and ensure production readiness of new models and inference backends.