llama.cpp builds b7492 through the latest b9060 contains a use-after-free vulnerability in llama-server affecting six tokenization endpoints (/tokenize, /detokenize, /infill, /apply-template, /rerank, and /anthropic/count_tokens) that bypass the task queue and access ctx_server.vocab directly on HTTP worker threads. Attackers can exploit a time-of-check-time-of-use race condition where the main thread destroys and frees vocab after the synchronization lock is released but before the handler finishes using it, causing a crash or potential code execution when --sleep-idle-seconds is configured.
{
"cna_assigner": "VulnCheck",
"cwe_ids": [
"CWE-367",
"CWE-416"
],
"osv_generated_from": "https://github.com/CVEProject/cvelistV5/tree/main/cves/2026/43xxx/CVE-2026-43632.json",
"unresolved_ranges": [
{
"extracted_events": [
{
"introduced": "b7492"
},
{
"last_affected": "b9060"
}
],
"source": "AFFECTED_FIELD"
},
{
"extracted_events": [
{
"introduced": "b7492"
},
{
"last_affected": "b9060"
}
],
"source": "CPE_FIELD"
},
{
"extracted_events": [
{
"introduced": "b7492"
}
],
"source": "DESCRIPTION"
}
]
}