潛藏21年的核心漏洞!
摘要
這份報告探討四個 Linux kernel 本機權限提升(LPE)漏洞 — DirtyAH6 (CVE-2026-80844)、 TUNderflow (CVE-2026-81000)、 PPPoEject (CVE-2026-68121) 和 DiagSpill (CVE-2026-74469) — 由先前發表 CIFSwitch 和 OVSwrap [1] 的同一位研究者一併揭露。這些底層漏洞在 kernel 中存活了 10 到 21 年。其中三個需要 unprivileged user namespaces,而 DiagSpill 完全不需要任何特殊 capability。在非常特定的條件下, DirtyAH6 和 DiagSpill 的 corruption primitives 可從遠端觸發。我們分析根本原因、上游修復程式以及攻擊策略,並將此漏洞類別與先前的 OOB-write LPE 研究進行比較 [2] [3] 。
1. 背景與發現方法論
「這四個漏洞是透過結合作者先前研究中的兩種技術而發現的:其一是利用圖形的安全相關核心物件追蹤(如 CIFSwitch),其二是用於以『幾何方式』推理記憶體狀態的工具(如 OVSwrap) [1] 。
表 1 總結了這四個漏洞。
| 名稱 | CVE | 子系統 | 漏洞類別 | 需要 User NS |
|---|---|---|---|---|
| DirtyAH6 | CVE-2026-80844 | IPsec AH6 / XFRM (IPv6) | 因未檢查 segments_left 導致的 OOB memmove | 是 |
| TUNderflow | CVE-2026-81000 | TUN/TAP + netkit + OVS | headroom 計算中的整數下溢 | 是 |
| PPPoEject | CVE-2026-68121 | PPPoE send path | skb head 的 Use-after-free | 是 |
| DiagSpill | CVE-2026-74469 | SCTP / sctp_diag | 16-bit counter wraparound → ~8 MiB OOB copy | 否 |
表 1. 四個 LPE 漏洞概覽 [1] 。
2. 根本原因分析
2.1 DirtyAH6:未檢查的 Routing-Header 指標算術
在計算或驗證 authentication data 之前,ipv6_rearrange_rthdr() 會正規化 IPv6 欄位,包括 routing-header addresses。該函數從 hdrlen 推導出 address count,然後將指標前進 (segments - segments_left),但未驗證 segments_left ≤ segments [1] 。一個帶有 hdrlen=2 和 segments_left=255 的 raw IPv6 HDRINCL packet 會使指標向後移動 4,064 bytes,並將 4,064-byte length 傳遞給 memmove()。在作為 IPv6 路由器/閘道並以傳輸模式採用 AH 的目標上,此漏洞可從遠端觸發導致 crash;透過目標上的 grooming,作者在實驗室中展示了 remote root [1] 。修復方法是簡單的 bounds check(程式 1)。
- segments = rthdr->hdrlen >> 1; // Derive segment count: hdrlen is in 8-byte units, so divide by 2 to get address count
- if (segments_left > segments) // Validate the routing header's remaining-segments field against the real count
- return -EINVAL; // Reject malformed header through the existing AH6 error path before any pointer math
- rthdr->segments_left = 0; // Header is well-formed; consume all segments (set to 0) before rearranging addresses
程式 1. DirtyAH6 修復
2.2 TUNderflow:Headroom Integer Underflow
tun_set_headroom() 將 receive headroom 直接儲存到 tun->align,而 tun_get_user() 重複使用該欄位來決定要在 skb head 中保留多少 packet data。一個配置了 4,096 bytes headroom 的 netkit device,在 VXLAN device 和 Open vSwitch datapath 之下,可能會將 4,160 bytes 傳播到 raw TUN port。SKB_MAX_HEAD(4160) 隨後發生下溢:負的 good_linear 值溢位被轉成一個極大的正 size_t;prepad + linear 與 len - linear 也發生溢位;而 tun_alloc_skb() 讓 skb->data 超出其 4,096 bytes 配置範圍 64 bytes [1] 。修復方法(程式 2)將儲存的 headroom 限制在 one-page skb-head budget 和最大的合法 16-bit offset,為 raw-TUN byte 或完整的 TAP Ethernet header 保留空間。
- max_headroom = min_t(size_t, SKB_MAX_HEAD(0), U16_MAX - 1); // Cap headroom at the smaller of the one-page skb-head budget and the max 16-bit header offset
- if ((tun->flags & TUN_TYPE_MASK) == IFF_TAP) // Distinguish TAP (Ethernet) devices from plain TUN devices
- max_headroom -= ETH_HLEN + NET_IP_ALIGN; // For TAP, reserve space for the 14-byte Ethernet header plus 2-byte IP alignment
- else
- max_headroom -= 1; // For TUN, reserve 1 byte for the raw protocol/version byte kept in the head
- tun->align = clamp_t(int, new_hr, NET_SKB_PAD, max_headroom);// Clamp the user-supplied headroom into [NET_SKB_PAD, max_headroom] before storing it in tun->align
程式 2. TUNderflow 修復
修復的第二部分防禦性地確保保留的 bytes 在使用前確實存在(程式 3)。
- case IFF_TUN: // Handle plain TUN (IP tunnel) packet reception
- if (tun->flags & IFF_NO_PI) { // IFF_NO_PI means no packet-information header precedes the payload
- u8 ip_version; // Variable to hold the IP version nibble read from the first payload byte
- if (!pskb_may_pull(skb, 1)) { // Ensure at least 1 byte is linear (contiguous) in the skb head; pulls from fragments if needed
- err = -EINVAL; // Packet is malformed/truncated: record invalid-argument error
- goto drop; // Jump to the shared drop path and discard the packet
- }
- ip_version = skb->data[0] >> 4; // Safely read the high nibble of the first byte = IP version (4 or 6)
- // ... remaining protocol dispatch ...
- }
- // ...
- break; // End of IFF_TUN handling
- case IFF_TAP: // Handle TAP (Ethernet bridge) packet reception
- if (!pskb_may_pull(skb, ETH_HLEN)) { // Ensure a full 14-byte Ethernet header is linear in the skb head
- err = -ENOMEM; // Record out-of-memory error (header cannot be made contiguous)
- drop_reason = SKB_DROP_REASON_HDR_TRUNC; // Annotate the drop with a dedicated "header truncated" reason for tracing
- goto drop; // Discard the packet through the shared drop path
- }
程式 3. TUNderflow 修復
2.3 PPPoEject:dev_hard_header() 期間的陳舊 skb-Head Pointer
pppoe_sendmsg() 建構了一個 skb,複製了 payload,然後要求 lower device 建構其 hardware header — 同時保持指向 skb head 的 pointer。device callback 可能會呼叫 pskb_expand_head(),這會釋放舊的 head。透過在 FUSE 上 block payload copy,同時將第一個 GRE/IP6GRE port 新增到空的 team 或 bonding device,reallocation 「ejected」了舊的 skb head,而 PPPoE 仍持有指向它的 pointer;後續的 header 和 length writes 使用了這個 stale pointer [1] 。修復方法透過 skb 的 network-header offset 重新載入 header,並教導 pskb_expand_head() 更新該 offset(程式 4)。
- dev_hard_header(skb, dev, ETH_P_PPP_SES, // Ask the lower device to build the Ethernet header for a PPPoE session frame;
- po->pppoe_pa.remote, NULL, // destination = peer MAC, source = device MAC, protocol = PPPoE session;
- total_len); // this callback may reallocate and free the skb head via pskb_expand_head()
- ph = pppoe_hdr(skb); // After header creation, re-derive the PPPoE header pointer through skb's network-header
- // offset (which pskb_expand_head() now updates) instead of reusing the stale pointer
- memcpy(ph, &hdr, sizeof(struct pppoe_hdr)); // Copy the session/version/type and length fields into the (valid) PPPoE header location
程式 4. PPPoEject 修復
圖 1 將脆弱的互動建模為 sequence diagram。
圖 1. PPPoEject use-after-free 觸發序列。
2.4 DiagSpill:16-bit Transport-Counter Wraparound
一個 SCTP association 最多可持有 65,536 個 peer transports,但 transport_count 是一個 16-bit field,因此第 65,536 個 transport 會使 counter 回零。sctp_diag 隨後未保留任何 peer payload,但將完整的 transport list 複製到 Netlink reply 中,導致大約 8 MiB 溢出 buffer 末端 [1] 。不需要 capabilities 或 user namespaces — 只需要 SCTP 和 sctp_diag 可用性。如果啟用了 ASCONF/ADD-IP 並搭配 SCTP-AUTH 或 net.sctp.addip_noauth_enable=1(預設皆為停用),malicious peer 可以從遠端新增足夠的 transports,而本地的 sock_diag consumer(例如 ss)會觸發 overwrite,導致 remote crash/DoS;作者認為這裡沒有通往完整 remote root 的路徑 [1] 。圖 2 顯示 overflow 序列。
圖 2. DiagSpill counter-wraparound overflow 序列
修復方法將 association 上限設為 U16_MAX unique peers,同時仍允許查詢已知的 addresses(程式 5)。
- if (asoc->peer.transport_count == U16_MAX) // If the 16-bit peer counter already holds 65535, refuse to add another UNIQUE transport
- return NULL; // return failure so the caller cannot create the 65,536th entry that would wrap to 0
- peer = sctp_transport_new(asoc->base.net, addr, gfp); // Counter is below the limit: allocate and initialize a new transport for this address
程式 5. DiagSpill 修復
3. 攻擊策略
這四個都是 memory-corruption bugs,其提升至 root 需要針對每個目標進行 grooming;發布的 PoCs 針對特定的 distro/kernel/CPU/memory 組合進行了調整 [1] 。DirtyAH6 透過獨立的 AH6 routing-header memmove() OOB corrupts skb_shared_info,然後引導後續的 ESP decrypt 寫入 file-backed fragment,將 pam_rootok.so 替換為 pam_permit.so,使 su 產生 root shell — 這是在 DirtyFrag [4] 中也分析過的 page-cache technique family。TUNderflow 將 file-backed pipe buffers 安排在 malformed TUN packet 旁邊,使 OVS out-of-bounds write 設定 PIPE_BUF_FLAG_CAN_MERGE,之後 pipe write 修補 /etc/pam.d/su。PPPoEject 將填充的 fd tables 競速進入 freed skb head,並將 live file entry 重新導向到攻擊者構造的假 struct file;關閉該 fd 會呼叫 controlled kernel callback 安裝 root credentials。DiagSpill 將 sock_diag overwrite groom 到 page tables,對應 host memory,重寫 credential object,安裝 sudoers rule,並開啟 root shell [1] 。這種多階段模式(heap shaping → controlled corruption → credential overwrite)反映了 ksmbd CVE-2025-37947 的方法論,其中 page-allocator grooming 和 msg_msg collisions 在 overwrite cred structures 之前產生了 arbitrary read/write [2] 。Copy Fail 同樣顯示,對 shared page-cache memory 的一次小型 deterministic write 就足以實現完整 LPE [3] 。
4. 受影響版本與緩解措施
DirtyAH6 和 PPPoEject 影響從 2.6.12 開始的 kernels;TUNderflow 從 4.6 開始;DiagSpill 從 4.7 開始 [1] 。包含所有四個修復的第一批 stable releases 是 5.10.270、5.15.221、6.1.188、6.6.157、6.12.109、6.18.50 和 7.2.4。在無法修補的情況下,停用 unprivileged user namespaces 會移除前三個的一般使用者路徑 — 但不適用於具有 CAP_NET_ADMIN 的 containers,也不適用於 DiagSpill,只要 SCTP 和 sctp_diag 可用,它就仍然可被利用。停用未使用的功能(AH6、TUN、PPPoE、SCTP/sctp_diag)也會移除底層漏洞 [1] 。測試中 AppArmor 和 SELinux 未能阻擋攻擊。
5. 結論
這四重奏說明 networking subsystems 中長期存在的 kernel code — routing-header parsing、headroom accounting、hardware-header construction 和 diagnostic counters — 仍然會產生強大的 LPE primitives。修復程式是小規模、局部化的驗證,這是此漏洞類別的典型特徵 [1] [2] 。從防禦角度來看,最緊急的重點是 DiagSpill:一個 capability-free、可從遠端觸發的 ~8 MiB overflow,它能在 user-namespace restrictions 下存活,使 SCTP module exposure 成為具體的強化決策,而非理論上的考量。