Back to Search Results
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Bengaluru, KA, India
Career Level: Associate
Industries: Technology, Software, IT, Electronics

Description



WHAT YOU DO AT AMD CHANGES EVERYTHING 

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.  Together, we advance your career.  



 

Role Overview

We are seeking an experienced SMTS Software Quality Engineer to join the HIP Runtime Quality Engineering team within the ROCm ecosystem. In this role, you will focus on defect triage, root‑cause analysis, and scalable test automation for the HIP runtime, ensuring robustness, portability, and performance across AMD and NVIDIA GPU backends.

The role sits at the intersection of runtime systems, GPU APIs, and CI-driven quality assurance, working closely with HIP runtime developers, compiler teams, and system software engineers to improve test coverage and accelerate issue resolution across a fast-evolving open‑source stack.

HIP is a C++ runtime API and kernel language designed for CUDA portability and heterogeneous GPU programming in the ROCm platform.

Key Responsibilities Defect Triage & Root Cause Analysis
  • Own end-to-end triage of HIP runtime failures observed in CI, nightly, and release testing pipelines.
  • Analyze failures across HIP runtime APIs, including memory management, streams, events, kernels, graphs, and device management.
  • Perform log analysis, repro reduction, and backtrace debugging across Linux and Windows environments.
  • Correlate failures across:
    • ROCr system runtime
    • HIP runtime
    • Driver / firmware interfaces
    • Compiler (hipcc / clang) interactions
  • File high‑quality bug reports with clear reproduction steps, root cause hypotheses, and impact assessment.
Test Automation & Framework Development
  • Design, enhance, and maintain automated test frameworks for HIP runtime validation using existing HIP test infrastructure (Catch2-based and custom harnesses).
  • Expand functional, negative, stress, and regression coverage for:
    • HIP Runtime APIs
    • Graph APIs
    • Memory copy variants (1D/2D/3D, async/sync)
    • Multi-GPU and peer-to-peer use cases
  • Integrate tests into CI/CD pipelines with actionable failure reporting.
  • Develop Python, C++, and shell-based automation for test execution, result aggregation, and trend analysis.
  • Drive test flakiness reduction and improve signal-to-noise ratio in large-scale CI.
Cross-Team Collaboration
  • Partner with HIP runtime developers to:
    • Review test gaps for new features and APIs
    • Validate bug fixes and performance-sensitive code paths
  • Collaborate with:
    • Compiler and language runtime teams
    • ROCr system runtime teams
    • Release engineering and CI infrastructure teams
  • Contribute test fixes and enhancements upstream in open-source HIP/ROCm repositories.
Quality Metrics & Continuous Improvement
  • Define and track quality metrics such as failure escape rate, test coverage, and mean-time-to-resolution (MTTR).
  • Identify systemic issues and propose data-driven quality improvements.
  • Participate in release readiness assessments for HIP runtime drops.
Required Qualifications Technical Skills
  • Strong experience in software quality engineering, test automation, or SDET roles.
  • Solid proficiency in C++ with debugging experience (gdb/lldb).
  • Working knowledge of HIP or CUDA runtime concepts, including kernels, streams, memory models, and device execution.
  • Experience with Linux-based development and debugging.
  • Familiarity with CI systems (Jenkins, GitHub Actions, GitLab CI, or similar).
  • Experience writing and maintaining automated tests for low-level systems or runtime libraries.
Debugging & Analysis Skills
  • Ability to reason about runtime abstractions, concurrency, and asynchronous execution.
  • Strong skills in failure reproduction, minimization, and log-driven debugging.
  • Experience triaging issues across hardware–software boundaries is a plus.
Preferred Qualifications
  • Experience with GPU programming models (HIP, CUDA, OpenCL).
  • Familiarity with ROCm stack components (HIP, ROCr, compilers, drivers).
  • Experience testing graph APIs, async execution models, or multi-GPU systems.
  • Prior contributions to open-source system software or test frameworks.
  • Knowledge of Python-based test orchestration and reporting tools.
  • Knowledge and application of AI tools in validation and testing.
What Success Looks Like
  • Faster, higher-quality issue triage with actionable root cause insights.
  • Improved automation coverage for HIP runtime APIs and edge cases.
  • Reduced CI noise and faster feedback for developers.
  • Strong collaboration with runtime and compiler engineers to ship stable, high-quality HIP releases.

 

#LI-NR1



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website