Skip to content

Roadmap

Upcoming features and improvements for the vLLM QAIC plugin.

Last updated: June 2026

AOT Mode

Feature Status Description
LMCache integration โœ… Ongoing KV cache sharing across fleet via LMCache
llm-d orchestration โœ… Ongoing Kubernetes-native distributed inference
Additional model architectures โœ… Ongoing Expanding validated model list

PYT (Eager / JIT) Mode

Feature Status Description
Fused HexNN kernels โœ… Ongoing PagedAttention, MLA, GatedLinearAttention
torch.compile support ๐Ÿ“‹ Planned Wider model coverage via compiled graphs
Triton kernel support ๐Ÿ“‹ Planned Execution of Triton Kernels
Text-only LLM support โœ… Ongoing Extending beyond VLMs
Speculative decoding ๐Ÿ“‹ Planned SpD support in eager mode

Status Legend

Symbol Meaning
โœ… Ongoing Active development
๐Ÿ“‹ Planned On the roadmap, not yet started
๐Ÿงช Experimental Available but not production-ready