On 7 October 2026 JFrog disclosed a critical vulnerability in LMCache, open-source software that speeds up LLM servers such as vLLM. CVE-2026-105192 scores CVSS 9.8 when the multiprocess server is bound to a routable address, and no fixed version is available yet. An unauthenticated attacker who can reach the ZeroMQ port can run code as the LMCache process user — often root on official container images.
In multiprocess mode LMCache runs as a standalone cache server that LLM workers reach over ZeroMQ. By default it listens only on localhost. It becomes reachable from other hosts when operators set a routable address — common in multi-node clusters. The project’s own Kubernetes example starts the server listening on every interface. LMCache inside a single vLLM process does not open the port at all.
The ZeroMQ socket has no authentication. One message type is unpacked with Python `pickle`, which can carry and execute code on decode. The server unpacks while reading arguments, before checking message type, so a crafted message runs the sender’s code immediately. Affected versions: 0.3.9 (October 2025) through 0.5.5 (latest stable), plus 0.5.6 release candidates and the development branch. Discoverer: Yuval Moravchick, JFrog Security Research.
Until a patch ships, JFrog advises not assigning the multiprocess server a routable address and keeping the port on the local machine or a trusted cluster network. A firewall reduces risk but does not remove it — any host that can still connect can run code. LMCache has not published its own security advisory, and JFrog provides no way to tell whether a server was already attacked.
The day before the CVE went public, six additional GitHub reports alleged unauthenticated multi-tenant cache access and network services without login — no CVEs, no maintainer confirmation. A related vLLM DoS (CVE-2026-105756, CVSS 6.5) in the LMCache multiprocess connector is already fixed in vLLM 0.30.0. The pattern — unauthenticated socket to pickle — echoes the ShadowMQ findings across AI inference frameworks in November 2025.
For organisations running AI agent and LLM infrastructure in production, this is an immediate exposure question: a cache server accidentally bound to 0.0.0.0 becomes an unpatched RCE surface. Inventory every vLLM/LMCache deployment, remove routable binds and segment cluster networks.
LLM cache tiers are often forgotten in vulnerability management because they are treated as “internal performance components” rather than attack surface. In practice they sit next to GPU nodes, frequently run as root in containers, and can reach models, prompts and sometimes secrets injected via environment variables. A pickle RCE there is not just DoS — it can mean theft of model weights, prompt injection against other workers and a pivot into the cluster. The ShadowMQ lesson from 2025 still holds: serialisation formats that execute code do not belong behind unauthenticated sockets.
Nordic AI labs, SaaS firms and public-sector pilots testing vLLM should especially review Helm charts and Terraform modules that copied the LMCache Kubernetes example. Default listening on all interfaces is convenient in the lab and dangerous in production. Put ZeroMQ behind NetworkPolicy, mTLS, or at least hostNetwork=false with a strict pod CIDR. Document which team owns the cache node — otherwise patching becomes “someone else’s problem” when advisories finally arrive.
What IT and security leads should do now
Inventory LMCache 0.3.9–0.5.5 (and RCs). If multiprocess mode is used: ensure ZeroMQ listens only on localhost or private cluster nets; remove routable binds. Upgrade vLLM to ≥0.30.0 for CVE-2026-105756. Monitor unexpected processes/connections from LMCache containers. Plan to apply the vendor patch as soon as it ships — until then network isolation is the only effective control. Also review whether pickle is used in other internal AI services and replace with safer protocols where possible.