Provide actionable feedback to internal DL engineers on SoC-friendly Transformer design
Reduce attention/FFN compute & activation memory, manage token/BEV grid size, and avoid/replace ops that break target toolchains.
Propose architecture patterns that preserve accuracy while meeting embedded constraints (bandwidth, SRAM/DDR pressure, operator support).
Codify guidelines into “design rules” (allowed/avoid ops, preferred blocks, quantization-robust practices) and drive adoption through reviews and decision logs.
Optimize latency/throughput/memory for real-time perception across NPU/GPU/DLA/DSP backends