← terug naar overzicht

Voorbij statische RAG: een adaptief raamwerk met drie metrieken voor efficiënte inferentie met lange context op gangbare GPU's

onderzoek 📅 2026-09-17
arXiv:2609.17564v1 Announce Type: new Abstract: Deploying retrieval-augmented generation (RAG) on commodity GPUs such as the NVIDIA T4 (16 GB VRAM) exposes a practical failure mode we call the Compression Paradox: neural prompt compression can add key-value (KV) cache contention and preprocessing latency that outweigh generation-time savings, while skipping compression can cause out-of-memory (OOM) failures on long contexts. We identify two distinct failure mechanisms when a vLLM-served LLM and

🔗 lees originele bron