llvm -keys=cpu -mcpu=core-avx2
Challenge #1: How can the generator understand the target hardware beyond a target string?
Add model-facing hardware information through two channels.
HW_GPU_EDGE
Challenge #2: How can one generator reuse tuning knowledge from multiple source devices?
Interpretation: Improved validity over TLM, up to 1.5x ish speedup for specific workloads, around 1.1x speedup for representative subgraphs
Interpretation: Hardware grounding mainly restores CPU validity. LoRA experts and workload-aware routing improve latency