Back to Search Results
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Bengaluru, KA, India
Career Level: Mid-Senior Level
Industries: Technology, Software, IT, Electronics

Description



WHAT YOU DO AT AMD CHANGES EVERYTHING 

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.  Together, we advance your career.  



The Role

We're building GPUs to run the most demanding AI in the world — Transformers, LLMs, and large vision models — and we need an experienced hands-on engineer who can tell us how fast those workloads will run before the silicon exists. This is an AI performance projection/modeling role: your job is to model AI workloads on our GPU architecture, predict their performance with confidence, and use those projections to shape the hardware we build.

You'll work across the full modeling spectrum — analytical models for fast architectural what-ifs, and architectural and cycle-accurate simulators for precision — choosing the right level of fidelity for the question at hand. Because these workloads live on the GPU, you'll need real GPU architecture depth to do it.

If you think in rooflines, arithmetic intensity, tensor/pipeline/data parallelism, and KV-cache bandwidth budgets — and you want your models to change the microarchitecture, not just describe it — this is the role.

 

What You'll Do

Project AI workload performance (your core mission)

  • Build and use models — analytical, architectural, and cycle-accurate — to project the performance of generative AI and LLM workloads (attention, GEMMs, collectives, training and inference, vision models etc) on our GPU architecture.
  • Choose the right fidelity for the question: analytical bounds and rooflines for rapid trade-off exploration, architectural and cycle-accurate simulation for precise answers.
  • Find where performance will be won or lost — compute-engine occupancy, L1/L2 hierarchy, memory bandwidth, SoC fabric — and quantify the impact of each.

Characterize workloads & shape the hardware

  • Characterize AI workloads end to end — trace and profile execution, extract structural footprints, and translate them into microarchitectural bottlenecks and performance bounding boxes.
  • Run architecture experiments to evaluate GPU configurations, scaling limits, and trade-offs against real AI workloads, backing every recommendation with sensitivity analysis and data.
  • Extend and maintain the cycle-accurate / cycle-approximate simulators for the GPU subsystems AI workloads depend on — compute engines, cache hierarchies, memory subsystems, interconnects.

Prove it on silicon (post-silicon)

  • Run AI workloads on early silicon, capture telemetry, and correlate results back to your pre-silicon models to sharpen fidelity.
  • Root-cause AI performance mismatches on physical silicon using low-level performance counters, register dumps, and Linux tracing.

Collaborate across the stack

  • Partner with compiler, runtime, and framework teams to turn model insights into optimizations, and to ground your projections in how workloads actually execute.
  • Act as the bridge between architecture/RTL and software, translating AI workload requirements into concrete architectural constraints.

What You'll Bring

  • AI workload expertise: Deep, hands-on understanding of LLM and Transformer performance — scaling techniques, parallel training/inference strategies (tensor, pipeline, data parallelism), kernel execution graphs — and how they map to hardware.
  • Performance modeling breadth: Hands-on experience with analytical performance models and execution-driven / architectural / cycle-accurate simulators, and sound judgment about which to reach for. Track record of turning models into decisions.
  • GPU architecture depth: Strong command of GPU execution pipelines, SIMD/SIMT models, cache hierarchy, memory technologies, and high-bandwidth interconnects (PCIe, custom fabrics).
  • Engineering craft: Expert modern C++ and production-grade Python.
  • Execution: Excellent structured problem-solving; able to turn noisy simulation data into crisp, actionable findings for engineering stakeholders.
  • Nice to have: Experience with NPU / AI-accelerator architecture, or modeling AI workloads on dedicated inference/training accelerators.

 

#LI-AA1

 



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website