{"api_version":"1","generated_at":"2026-09-17T06:07:12+00:00","cve":"CVE-2026-89811","urls":{"html":"https://cve.report/CVE-2026-89811","api":"https://cve.report/api/cve/CVE-2026-89811.json","docs":"https://cve.report/api","cve_org":"https://www.cve.org/CVERecord?id=CVE-2026-89811","nvd":"https://nvd.nist.gov/vuln/detail/CVE-2026-89811"},"summary":{"title":"drm/amdkfd: Add TLB flush after MES queue eviction/suspension","description":"In the Linux kernel, the following vulnerability has been resolved:\n\ndrm/amdkfd: Add TLB flush after MES queue eviction/suspension\n\nMES (Micro Engine Scheduler) does not perform heavy-weight TLB\ninvalidation after unmapping queues, unlike HWS which does this\nautomatically. This causes a race condition where in-flight DMA\ndescriptors can access memory that has been unmapped, leading to page\nfaults and GPU queue hangs during SVM page migration.\n\nThe issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest\nfailures on gfx1151 (Strix Point) with XNACK mode 1 enabled - the GPU\ncompute queue hangs with packets submitted but never consumed.\n\nAdd kfd_flush_tlb() calls after MES queue removal in two locations:\n- evict_process_queues_cpsch(): after all queues removed during eviction\n- suspend_queues(): after debug/criu queue suspension (with mem_fence barrier)\n\nThis ensures all in-flight memory accesses from unmapped queues are\nflushed before memory is freed or migrated.\n\n(cherry picked from commit f5c4f88e0f9c45a8fb9dfac0c1df726c95e41b77)","state":"PUBLISHED","assigner":"Linux","published_at":"2026-09-16 11:16:46","updated_at":"2026-09-16 15:18:10"},"problem_types":[],"metrics":[{"version":"3.1","source":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","type":"Secondary","score":"8.8","severity":"HIGH","vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","data":{"version":"3.1","vectorString":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","baseScore":8.8,"baseSeverity":"HIGH","attackVector":"LOCAL","attackComplexity":"LOW","privilegesRequired":"LOW","userInteraction":"NONE","scope":"CHANGED","confidentialityImpact":"HIGH","integrityImpact":"HIGH","availabilityImpact":"HIGH"}},{"version":"3.1","source":"CNA","type":"DECLARED","score":"8.8","severity":"HIGH","vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","data":{"baseScore":8.8,"baseSeverity":"HIGH","vectorString":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","version":"3.1"}}],"references":[{"url":"https://git.kernel.org/stable/c/94e25cb6ab7f4f025bcdcd8ea79fda30f12843a4","name":"https://git.kernel.org/stable/c/94e25cb6ab7f4f025bcdcd8ea79fda30f12843a4","refsource":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","tags":[],"title":"","mime":"","httpstatus":"","archivestatus":"0"},{"url":"https://git.kernel.org/stable/c/6aae545c87087e4dc96e205452243813bfa73ba2","name":"https://git.kernel.org/stable/c/6aae545c87087e4dc96e205452243813bfa73ba2","refsource":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","tags":[],"title":"","mime":"","httpstatus":"","archivestatus":"0"},{"url":"https://git.kernel.org/stable/c/e230c546ed93741833ab5babb0b202f9c2052b45","name":"https://git.kernel.org/stable/c/e230c546ed93741833ab5babb0b202f9c2052b45","refsource":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","tags":[],"title":"","mime":"","httpstatus":"","archivestatus":"0"},{"url":"https://www.cve.org/CVERecord?id=CVE-2026-89811","name":"CVE Program record","refsource":"CVE.ORG","tags":["canonical"]},{"url":"https://nvd.nist.gov/vuln/detail/CVE-2026-89811","name":"NVD vulnerability detail","refsource":"NVD","tags":["canonical","analysis"]}],"affected":[{"source":"CNA","vendor":"Linux","product":"Linux","version":"affected 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2 e230c546ed93741833ab5babb0b202f9c2052b45 git","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"affected 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2 6aae545c87087e4dc96e205452243813bfa73ba2 git","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"affected 1da177e4c3f41524e886b7f1b8a0c1fc7321cac2 94e25cb6ab7f4f025bcdcd8ea79fda30f12843a4 git","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"affected 6.18.51 semver","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"affected 7.2.5 semver","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"unaffected 6.18.51 6.18.* semver","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"unaffected 7.2.5 7.2.* semver","platforms":[]},{"source":"CNA","vendor":"Linux","product":"Linux","version":"unaffected 7.3-rc2 * original_commit_for_fix","platforms":[]}],"timeline":[],"solutions":[],"workarounds":[],"exploits":[],"credits":[],"nvd_cpes":[],"vendor_comments":[],"enrichments":{"kev":null,"epss":null,"legacy_qids":[]},"source_records":{"cve_program":{"containers":{"cna":{"affected":[{"defaultStatus":"unaffected","product":"Linux","programFiles":["drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c"],"repo":"https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git","vendor":"Linux","versions":[{"lessThan":"e230c546ed93741833ab5babb0b202f9c2052b45","status":"affected","version":"1da177e4c3f41524e886b7f1b8a0c1fc7321cac2","versionType":"git"},{"lessThan":"6aae545c87087e4dc96e205452243813bfa73ba2","status":"affected","version":"1da177e4c3f41524e886b7f1b8a0c1fc7321cac2","versionType":"git"},{"lessThan":"94e25cb6ab7f4f025bcdcd8ea79fda30f12843a4","status":"affected","version":"1da177e4c3f41524e886b7f1b8a0c1fc7321cac2","versionType":"git"},{"lessThan":"6.18.51","status":"affected","version":"0","versionType":"semver"},{"lessThan":"7.2.5","status":"affected","version":"0","versionType":"semver"}]},{"defaultStatus":"affected","product":"Linux","programFiles":["drivers/gpu/drm/amd/amdkfd/kfd_device_queue_manager.c"],"repo":"https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git","vendor":"Linux","versions":[{"lessThanOrEqual":"6.18.*","status":"unaffected","version":"6.18.51","versionType":"semver"},{"lessThanOrEqual":"7.2.*","status":"unaffected","version":"7.2.5","versionType":"semver"},{"lessThanOrEqual":"*","status":"unaffected","version":"7.3-rc2","versionType":"original_commit_for_fix"}]}],"cpeApplicability":[{"nodes":[{"cpeMatch":[{"criteria":"cpe:2.3:o:linux:linux_kernel:*:*:*:*:*:*:*:*","versionEndExcluding":"6.18.51","vulnerable":true},{"criteria":"cpe:2.3:o:linux:linux_kernel:*:*:*:*:*:*:*:*","versionEndExcluding":"7.2.5","vulnerable":true},{"criteria":"cpe:2.3:o:linux:linux_kernel:*:*:*:*:*:*:*:*","versionEndExcluding":"7.3-rc2","vulnerable":true}],"negate":false,"operator":"OR"}]}],"descriptions":[{"lang":"en","value":"In the Linux kernel, the following vulnerability has been resolved:\n\ndrm/amdkfd: Add TLB flush after MES queue eviction/suspension\n\nMES (Micro Engine Scheduler) does not perform heavy-weight TLB\ninvalidation after unmapping queues, unlike HWS which does this\nautomatically. This causes a race condition where in-flight DMA\ndescriptors can access memory that has been unmapped, leading to page\nfaults and GPU queue hangs during SVM page migration.\n\nThe issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest\nfailures on gfx1151 (Strix Point) with XNACK mode 1 enabled - the GPU\ncompute queue hangs with packets submitted but never consumed.\n\nAdd kfd_flush_tlb() calls after MES queue removal in two locations:\n- evict_process_queues_cpsch(): after all queues removed during eviction\n- suspend_queues(): after debug/criu queue suspension (with mem_fence barrier)\n\nThis ensures all in-flight memory accesses from unmapped queues are\nflushed before memory is freed or migrated.\n\n(cherry picked from commit f5c4f88e0f9c45a8fb9dfac0c1df726c95e41b77)"}],"metrics":[{"cvssV3_1":{"baseScore":8.8,"baseSeverity":"HIGH","vectorString":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","version":"3.1"},"scenarios":[{"lang":"en","value":"AV:L - The bug is reached only via local /dev/kfd paths: AMDKFD_IOC_CREATE_QUEUE plus SVM/userptr activity (AMDKFD_IOC_SVM, munmap of USERPTR BOs) or KFD_IOC_DBG_TRAP_SUSPEND_QUEUES, which call evict_process_queues_cpsch()/suspend_queues(); there is no network, adjacent-wireless, or physical-device path.\nAC:L - The attacker controls both sides of the race: they keep DMA in flight on their own MES compute/SDMA queues while concurrently forcing eviction or suspend via SVM migration, userptr MMU notifiers, TTM pressure, or self-debug suspend, matching the MultiThreadMigrationTest trigger with no victim timing or uncontrollable layout required.\nPR:L - kfd_open() and the CREATE_QUEUE/SVM/DBG_TRAP ioctls perform no capability checks; a local unprivileged user with typical render/video-group access to /dev/kfd and /dev/dri/renderD* on desktops, ROCm nodes, and multi-tenant AMD GPU clouds is sufficient, not init-namespace root.\nUI:N - The attacker opens their own KFD and DRM render-node descriptors, submits compute work, and triggers queue eviction or suspend in that same process; no separate victim action such as mounting a filesystem or opening a crafted file is required.\nS:C - MES queue removal omits the heavyweight PASID TLB invalidation HWS performs, so in-flight GPU DMA can use stale GPUVM translations to physical pages after they are unmapped, freed, or migrated, bypassing the GPUVM/PASID DMA isolation boundary that is supposed to confine GPU access.\nC:H - In-flight GPU DMA through stale GPUVM TLB entries can read physical pages after they leave the process mapping and are recycled to the page allocator, kernel, or another process, giving an arbitrary memory-disclosure primitive.\nI:H - The same stale translations allow GPU DMA writes into those recycled physical pages, enabling arbitrary memory corruption and control-flow hijacking through a device DMA write primitive.\nA:H - The issue is documented to cause GPU page faults and compute-queue hangs (packets submitted but never consumed), which can stall the shared GPU and force amdgpu reset, denying compute service to all co-located workloads."}]}],"providerMetadata":{"dateUpdated":"2026-09-16T14:38:56.152Z","orgId":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","shortName":"Linux"},"references":[{"url":"https://git.kernel.org/stable/c/e230c546ed93741833ab5babb0b202f9c2052b45"},{"url":"https://git.kernel.org/stable/c/6aae545c87087e4dc96e205452243813bfa73ba2"},{"url":"https://git.kernel.org/stable/c/94e25cb6ab7f4f025bcdcd8ea79fda30f12843a4"}],"title":"drm/amdkfd: Add TLB flush after MES queue eviction/suspension","x_generator":{"engine":"bippy-1.2.0"}}},"cveMetadata":{"assignerOrgId":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","assignerShortName":"Linux","cveId":"CVE-2026-89811","datePublished":"2026-09-16T10:30:44.274Z","dateReserved":"2026-09-11T19:38:34.768Z","dateUpdated":"2026-09-16T14:38:56.152Z","state":"PUBLISHED"},"dataType":"CVE_RECORD","dataVersion":"5.2"},"nvd":{"publishedDate":"2026-09-16 11:16:46","lastModifiedDate":"2026-09-16 15:18:10","problem_types":[],"metrics":{"cvssMetricV31":[{"source":"416baaa9-dc9f-4396-8d5f-8c081fb06d67","type":"Secondary","cvssData":{"version":"3.1","vectorString":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","baseScore":8.8,"baseSeverity":"HIGH","attackVector":"LOCAL","attackComplexity":"LOW","privilegesRequired":"LOW","userInteraction":"NONE","scope":"CHANGED","confidentialityImpact":"HIGH","integrityImpact":"HIGH","availabilityImpact":"HIGH"},"exploitabilityScore":2,"impactScore":6}]},"configurations":[]},"legacy_mitre":{"record":{"CveYear":"2026","CveId":"89811","Ordinal":"1","Title":"drm/amdkfd: Add TLB flush after MES queue eviction/suspension","CVE":"CVE-2026-89811","Year":"2026"},"notes":[{"CveYear":"2026","CveId":"89811","Ordinal":"1","NoteData":"In the Linux kernel, the following vulnerability has been resolved:\n\ndrm/amdkfd: Add TLB flush after MES queue eviction/suspension\n\nMES (Micro Engine Scheduler) does not perform heavy-weight TLB\ninvalidation after unmapping queues, unlike HWS which does this\nautomatically. This causes a race condition where in-flight DMA\ndescriptors can access memory that has been unmapped, leading to page\nfaults and GPU queue hangs during SVM page migration.\n\nThe issue manifests as KFDSVMRangeTest.MultiThreadMigrationTest\nfailures on gfx1151 (Strix Point) with XNACK mode 1 enabled - the GPU\ncompute queue hangs with packets submitted but never consumed.\n\nAdd kfd_flush_tlb() calls after MES queue removal in two locations:\n- evict_process_queues_cpsch(): after all queues removed during eviction\n- suspend_queues(): after debug/criu queue suspension (with mem_fence barrier)\n\nThis ensures all in-flight memory accesses from unmapped queues are\nflushed before memory is freed or migrated.\n\n(cherry picked from commit f5c4f88e0f9c45a8fb9dfac0c1df726c95e41b77)","Type":"Description","Title":"drm/amdkfd: Add TLB flush after MES queue eviction/suspension"}]}}}