AI Engineer – Model Optimization & Acceleration
AMD — Bengaluru, Karnataka
₹15–32 LPA (HireSetu estimate) · 0–3 yrs exp · Freshers eligible · Full-time · onsite
HireSetu listing reference adea46 — confirm requirements below, then apply on the employer site.
Quick answer: This is post-training work: take a model that already works, export it, get it onto an accelerator, then cut its latency, memory footprint and cost per inference. Training or fine-tuning experience alone won't map onto it.
From AMD's job posting
WHAT YOU DO AT AMD CHANGES EVERYTHING
At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.
Job Title: AI Engineer – Model Optimization & Acceleration
Location: Bangalore, India
Role Overview
Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU).
Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive.
Key Responsibilities
- Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech
- Port models across frameworks (e.g., PyTorch → ONNX → runtimes)
- Deploy on hardware accelerators (GPU/NPU) and optimize performance
- Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion)
- Apply quantization and model compression (FP32 → lower precision)
- Profile and debug system and model performance
Required Skills
- Strong in PyTorch (or similar), ONNX (or equivalent)
- Proficient in Python and C++
- Experience with GPU/hardware acceleration (CUDA/ROCm or similar)
- Solid understanding of deep learning models (transformers, CNNs)
- Knowledge of optimization, quantization, and performance tuning
Good to Have
- Edge AI or embedded deployment
- Generative or multi-modal AI systems
- Distributed inference or streaming pipelines
#LI-PK1
Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.
This posting is for an existing vacancy.
Text from AMD's official job posting, formatted by HireSetu. The employer's own page is the final word on details.
HireSetu's note
Speeding up models someone else trained
The core of the listing is a chain: a PyTorch model, exported through ONNX, executed on a device-specific runtime, then made faster. You won't be tuning loss curves; you'll be in profilers. Both Python and C++ are named, and C++ is where the runtime and kernel side sits, so expect to read code below the framework layer, not just call it.
Precision and memory are the levers
Quantization down from FP32 is listed plainly and it's the part candidates underrate. Know what breaks when precision drops: calibration, per-channel against per-tensor scaling, and the accuracy hit you accept for the speed you gain. Learn to tell a compute-bound kernel from a bandwidth-bound one, because that decides whether fusion or batching will help at all. Have one real example ready of a technique you applied, and what it did to latency and memory.
A week of profiling and re-measuring
Expect long stretches of benchmarking. Run the model, profile it end to end, find the operator or memory copy eating the time, change it, measure again, then confirm the output still matches the reference within tolerance. Moving between frameworks brings its own failure class: ops that won't export, shape mismatches, dynamic axes that don't survive conversion. Debugging is listed as a skill for a reason, and this is what it means here.
Which accelerator stack you'd work on
AMD builds its own software stack for its accelerators, so the practical question is whether your day-to-day targets ROCm on AMD silicon or a CUDA-based environment, and whether you sit on the GPU, NPU or CPU path. Ask that early. Also ask which domain the specific team serves, robotics, healthcare or automotive, because deployment constraints differ sharply between them even when the optimization work looks identical.
Drafted with AI from AMD's job posting and checked by Prashanth Jakkula before publishing. See our Editorial Policy.
How to apply
- Open AMD's careers site with the Apply button.
- Follow the application steps on the site and upload your resume.
- Keep the confirmation email or application number for follow-up.
No fees, ever. HireSetu never charges candidates, and no genuine employer asks for money to apply, interview or join. If anyone asks you to pay, stop and report the job.
Similar openings on HireSetu
- 2027 Campus Hire – System Engineer at Qualcomm — Bengaluru — VLSI · Full-time
- 2027 Campus Hire Associate Engineer, Software at Qualcomm — Hyderabad — VLSI · Full-time
- Silicon Engineer, Platform and Devices, University Graduate, 2026 at Google — Bengaluru — VLSI · Full-time
- CAD/EDA Methodology Engineer at Qualcomm — Bengaluru — VLSI · Full-time
- Engineer — ASIC Physical Design at Arm — Bengaluru — VLSI · Full-time
View AI Engineer – Model Optimization & Acceleration on HireSetu · More AMD jobs · VLSI jobs · Jobs in Bengaluru · Career Insights · FAQ · Companies