주요업무
[Mission of the Role]
We’re hiring an Embedded AI Engineer to productize MV-Gen2 multi-camera, Transformer-based vision perception models (e.g., object detection / lane / occupancy), and E2E models(e.g., path planner, and control) across multiple automotive SoCs—Renesas, TI (Jacinto/TDA4), Qualcomm, and NVIDIA Orin/Thor.
You will lead INT8 quantization (PTQ/QAT), model porting, runtime/accelerator optimization, and—critically—provide HW-aware redesign guidance to internal DL model teams so architectures remain deployable and efficient under SoC constraints.
This role is a unique opportunity to work on high-impact, cutting-edge research that directly contributes to the development of next-generation autonomous driving systems.
[Key Responsibilities]
The selected candidate will be responsible for designing, developing, and optimizing deep learning models for ADAS/Autonomy.
1) Quantization & Deployment (Multi-Camera Transformers)
• Convert and deploy multi-camera Transformer perception models to embedded inference stacks using PTQ/QAT, calibration pipelines, and accuracy recovery strategies (mixed precision, selective quantization, layer-wise sensitivity).
• Build robust export/packaging flow (e.g., PyTorch → ONNX → target runtime) and resolve conversion/runtime issues (unsupported ops, precision constraints, graph transforms).
2) HW-Aware Model Redesign Feedback (Internal Technical Leadership)
• Provide actionable feedback to internal DL engineers on SoC-friendly Transformer design :
• Reduce attention/FFN compute & activation memory, manage token/BEV grid size, and avoid/replace ops that break target toolchains.
• Propose architecture patterns that preserve accuracy while meeting embedded constraints (bandwidth, SRAM/DDR pressure, operator support).
• Codify guidelines into “design rules” (allowed/avoid ops, preferred blocks, quantization-robust practices) and drive adoption through reviews and decision logs.
3) Multi-SoC Runtime & Accelerator Optimization
• Optimize latency/throughput/memory for real-time perception across NPU/GPU/DLA/DSP backends :
• Renesas: DRP-AI INT8 PTQ flow (calibration-based static quantization) and translation pipeline integration.
• NVIDIA Orin: TensorRT build/engine tuning, DLA constraints(INT8/FP16) and GPU–DLA workload partitioning.
• Qualcomm: ONNX Runtime QNN Execution Provider/ QNN SDK-based acceleration strategy and deployment validation.
TI: TIDL quantization modes (PTQ/QAT) and “deployable-by-design” constraints.
4) Tooling & Performance Infrastructure
• Establish repeatable profiling + regression harness for latency/memory/accuracy across SoCs; publish performance dashboards and release artifacts.
• Collaborate with platform/firmware teams to unblock runtime integration and memory/IO bottlenecks.