摘要

這份報告探討四個 Linux kernel 本機權限提升(LPE)漏洞 — DirtyAH6 (CVE-2026-80844)、 TUNderflow (CVE-2026-81000)、 PPPoEject (CVE-2026-68121) 和 DiagSpill (CVE-2026-74469) — 由先前發表 CIFSwitch 和 OVSwrap [1] 的同一位研究者一併揭露。這些底層漏洞在 kernel 中存活了 10 到 21 年。其中三個需要 unprivileged user namespaces,而 DiagSpill 完全不需要任何特殊 capability。在非常特定的條件下, DirtyAH6 和 DiagSpill 的 corruption primitives 可從遠端觸發。我們分析根本原因、上游修復程式以及攻擊策略,並將此漏洞類別與先前的 OOB-write LPE 研究進行比較 [2] [3] 。

潛藏21年的核心漏洞!四個Linux權限提升如何躲過所有人的眼睛? | 資訊安全新聞

1. 背景與發現方法論

「這四個漏洞是透過結合作者先前研究中的兩種技術而發現的:其一是利用圖形的安全相關核心物件追蹤(如 CIFSwitch),其二是用於以『幾何方式』推理記憶體狀態的工具(如 OVSwrap) [1] 。

表 1 總結了這四個漏洞。

名稱 CVE 子系統 漏洞類別 需要 User NS
DirtyAH6 CVE-2026-80844 IPsec AH6 / XFRM (IPv6) 因未檢查 segments_left 導致的 OOB memmove 是
TUNderflow CVE-2026-81000 TUN/TAP + netkit + OVS headroom 計算中的整數下溢 是
PPPoEject CVE-2026-68121 PPPoE send path skb head 的 Use-after-free 是
DiagSpill CVE-2026-74469 SCTP / sctp_diag 16-bit counter wraparound → ~8 MiB OOB copy 否

表 1. 四個 LPE 漏洞概覽 [1] 。

2. 根本原因分析

2.1 DirtyAH6:未檢查的 Routing-Header 指標算術

在計算或驗證 authentication data 之前,ipv6_rearrange_rthdr() 會正規化 IPv6 欄位,包括 routing-header addresses。該函數從 hdrlen 推導出 address count,然後將指標前進 (segments - segments_left),但未驗證 segments_left ≤ segments [1] 。一個帶有 hdrlen=2 和 segments_left=255 的 raw IPv6 HDRINCL packet 會使指標向後移動 4,064 bytes,並將 4,064-byte length 傳遞給 memmove()。在作為 IPv6 路由器/閘道並以傳輸模式採用 AH 的目標上,此漏洞可從遠端觸發導致 crash;透過目標上的 grooming,作者在實驗室中展示了 remote root [1] 。修復方法是簡單的 bounds check(程式 1)。

  1. segments = rthdr->hdrlen >> 1; // Derive segment count: hdrlen is in 8-byte units, so divide by 2 to get address count
  2. if (segments_left > segments) // Validate the routing header's remaining-segments field against the real count
  3. return -EINVAL; // Reject malformed header through the existing AH6 error path before any pointer math
  4. rthdr->segments_left = 0; // Header is well-formed; consume all segments (set to 0) before rearranging addresses

程式 1. DirtyAH6 修復

2.2 TUNderflow:Headroom Integer Underflow

tun_set_headroom() 將 receive headroom 直接儲存到 tun->align,而 tun_get_user() 重複使用該欄位來決定要在 skb head 中保留多少 packet data。一個配置了 4,096 bytes headroom 的 netkit device,在 VXLAN device 和 Open vSwitch datapath 之下,可能會將 4,160 bytes 傳播到 raw TUN port。SKB_MAX_HEAD(4160) 隨後發生下溢:負的 good_linear 值溢位被轉成一個極大的正 size_t;prepad + linear 與 len - linear 也發生溢位;而 tun_alloc_skb() 讓 skb->data 超出其 4,096 bytes 配置範圍 64 bytes [1] 。修復方法(程式 2)將儲存的 headroom 限制在 one-page skb-head budget 和最大的合法 16-bit offset,為 raw-TUN byte 或完整的 TAP Ethernet header 保留空間。

  1. max_headroom = min_t(size_t, SKB_MAX_HEAD(0), U16_MAX - 1); // Cap headroom at the smaller of the one-page skb-head budget and the max 16-bit header offset
  2. if ((tun->flags & TUN_TYPE_MASK) == IFF_TAP) // Distinguish TAP (Ethernet) devices from plain TUN devices
  3. max_headroom -= ETH_HLEN + NET_IP_ALIGN; // For TAP, reserve space for the 14-byte Ethernet header plus 2-byte IP alignment
  4. else
  5. max_headroom -= 1; // For TUN, reserve 1 byte for the raw protocol/version byte kept in the head
  6. tun->align = clamp_t(int, new_hr, NET_SKB_PAD, max_headroom);// Clamp the user-supplied headroom into [NET_SKB_PAD, max_headroom] before storing it in tun->align

程式 2. TUNderflow 修復

修復的第二部分防禦性地確保保留的 bytes 在使用前確實存在(程式 3)。

  1. case IFF_TUN: // Handle plain TUN (IP tunnel) packet reception
  2. if (tun->flags & IFF_NO_PI) { // IFF_NO_PI means no packet-information header precedes the payload
  3. u8 ip_version; // Variable to hold the IP version nibble read from the first payload byte
  4. if (!pskb_may_pull(skb, 1)) { // Ensure at least 1 byte is linear (contiguous) in the skb head; pulls from fragments if needed
  5. err = -EINVAL; // Packet is malformed/truncated: record invalid-argument error
  6. goto drop; // Jump to the shared drop path and discard the packet
  7. }
  8. ip_version = skb->data[0] >> 4; // Safely read the high nibble of the first byte = IP version (4 or 6)
  9. // ... remaining protocol dispatch ...
  10. }
  11. // ...
  12. break; // End of IFF_TUN handling
  13. case IFF_TAP: // Handle TAP (Ethernet bridge) packet reception
  14. if (!pskb_may_pull(skb, ETH_HLEN)) { // Ensure a full 14-byte Ethernet header is linear in the skb head
  15. err = -ENOMEM; // Record out-of-memory error (header cannot be made contiguous)
  16. drop_reason = SKB_DROP_REASON_HDR_TRUNC; // Annotate the drop with a dedicated "header truncated" reason for tracing
  17. goto drop; // Discard the packet through the shared drop path
  18. }

程式 3. TUNderflow 修復

2.3 PPPoEject:dev_hard_header() 期間的陳舊 skb-Head Pointer

pppoe_sendmsg() 建構了一個 skb,複製了 payload,然後要求 lower device 建構其 hardware header — 同時保持指向 skb head 的 pointer。device callback 可能會呼叫 pskb_expand_head(),這會釋放舊的 head。透過在 FUSE 上 block payload copy,同時將第一個 GRE/IP6GRE port 新增到空的 team 或 bonding device,reallocation 「ejected」了舊的 skb head,而 PPPoE 仍持有指向它的 pointer;後續的 header 和 length writes 使用了這個 stale pointer [1] 。修復方法透過 skb 的 network-header offset 重新載入 header,並教導 pskb_expand_head() 更新該 offset(程式 4)。

  1. dev_hard_header(skb, dev, ETH_P_PPP_SES, // Ask the lower device to build the Ethernet header for a PPPoE session frame;
  2. po->pppoe_pa.remote, NULL, // destination = peer MAC, source = device MAC, protocol = PPPoE session;
  3. total_len); // this callback may reallocate and free the skb head via pskb_expand_head()
  4. ph = pppoe_hdr(skb); // After header creation, re-derive the PPPoE header pointer through skb's network-header
  5. // offset (which pskb_expand_head() now updates) instead of reusing the stale pointer
  6. memcpy(ph, &hdr, sizeof(struct pppoe_hdr)); // Copy the session/version/type and length fields into the (valid) PPPoE header location

程式 4. PPPoEject 修復

圖 1 將脆弱的互動建模為 sequence diagram。

sequenceDiagram participant U as Unprivileged Process participant P as pppoe_sendmsg() participant D as Lower Dev (team/GRE) participant F as FUSE-backed Payload participant H as skb Head U->>P: sendmsg() with PPPoE payload P->>H: alloc skb; copy payload start P->>F: block on payload copy (FUSE stall) U->>D: add first GRE port to empty team dev D->>H: dev_hard_header() callback H->>H: pskb_expand_head(): free old head, allocate new head P->>H: write PPPoE hdr via STALE pointer (use-after-free) Note over P,H: OOB header + length write on freed head

圖 1. PPPoEject use-after-free 觸發序列。

2.4 DiagSpill:16-bit Transport-Counter Wraparound

一個 SCTP association 最多可持有 65,536 個 peer transports,但 transport_count 是一個 16-bit field,因此第 65,536 個 transport 會使 counter 回零。sctp_diag 隨後未保留任何 peer payload,但將完整的 transport list 複製到 Netlink reply 中,導致大約 8 MiB 溢出 buffer 末端 [1] 。不需要 capabilities 或 user namespaces — 只需要 SCTP 和 sctp_diag 可用性。如果啟用了 ASCONF/ADD-IP 並搭配 SCTP-AUTH 或 net.sctp.addip_noauth_enable=1(預設皆為停用),malicious peer 可以從遠端新增足夠的 transports,而本地的 sock_diag consumer(例如 ss)會觸發 overwrite,導致 remote crash/DoS;作者認為這裡沒有通往完整 remote root 的路徑 [1] 。圖 2 顯示 overflow 序列。

sequenceDiagram participant M as Malicious SCTP Peer participant S as SCTP Association participant C as sctp_diag (sock_diag) participant N as Netlink Reply Buffer M->>S: ASCONF ADD-IP x65536 (transports) S->>S: transport_count (16-bit) wraps 65535 to 0 C->>S: build diag reply: reserve = transport_count * sizeof(sockaddr_storage) = 0 C->>N: copy full 65536-entry transport list into ~0-sized payload N->>N: ~8 MiB out-of-bounds write past reply buffer

圖 2. DiagSpill counter-wraparound overflow 序列

修復方法將 association 上限設為 U16_MAX unique peers,同時仍允許查詢已知的 addresses(程式 5)。

  1. if (asoc->peer.transport_count == U16_MAX) // If the 16-bit peer counter already holds 65535, refuse to add another UNIQUE transport
  2. return NULL; // return failure so the caller cannot create the 65,536th entry that would wrap to 0
  3. peer = sctp_transport_new(asoc->base.net, addr, gfp); // Counter is below the limit: allocate and initialize a new transport for this address

程式 5. DiagSpill 修復

3. 攻擊策略

這四個都是 memory-corruption bugs,其提升至 root 需要針對每個目標進行 grooming;發布的 PoCs 針對特定的 distro/kernel/CPU/memory 組合進行了調整 [1] 。DirtyAH6 透過獨立的 AH6 routing-header memmove() OOB corrupts skb_shared_info,然後引導後續的 ESP decrypt 寫入 file-backed fragment,將 pam_rootok.so 替換為 pam_permit.so,使 su 產生 root shell — 這是在 DirtyFrag [4] 中也分析過的 page-cache technique family。TUNderflow 將 file-backed pipe buffers 安排在 malformed TUN packet 旁邊,使 OVS out-of-bounds write 設定 PIPE_BUF_FLAG_CAN_MERGE,之後 pipe write 修補 /etc/pam.d/su。PPPoEject 將填充的 fd tables 競速進入 freed skb head,並將 live file entry 重新導向到攻擊者構造的假 struct file;關閉該 fd 會呼叫 controlled kernel callback 安裝 root credentials。DiagSpill 將 sock_diag overwrite groom 到 page tables,對應 host memory,重寫 credential object,安裝 sudoers rule,並開啟 root shell [1] 。這種多階段模式(heap shaping → controlled corruption → credential overwrite)反映了 ksmbd CVE-2025-37947 的方法論,其中 page-allocator grooming 和 msg_msg collisions 在 overwrite cred structures 之前產生了 arbitrary read/write [2] 。Copy Fail 同樣顯示,對 shared page-cache memory 的一次小型 deterministic write 就足以實現完整 LPE [3] 。

4. 受影響版本與緩解措施

DirtyAH6 和 PPPoEject 影響從 2.6.12 開始的 kernels;TUNderflow 從 4.6 開始;DiagSpill 從 4.7 開始 [1] 。包含所有四個修復的第一批 stable releases 是 5.10.270、5.15.221、6.1.188、6.6.157、6.12.109、6.18.50 和 7.2.4。在無法修補的情況下,停用 unprivileged user namespaces 會移除前三個的一般使用者路徑 — 但不適用於具有 CAP_NET_ADMIN 的 containers,也不適用於 DiagSpill,只要 SCTP 和 sctp_diag 可用,它就仍然可被利用。停用未使用的功能(AH6、TUN、PPPoE、SCTP/sctp_diag)也會移除底層漏洞 [1] 。測試中 AppArmor 和 SELinux 未能阻擋攻擊。

5. 結論

這四重奏說明 networking subsystems 中長期存在的 kernel code — routing-header parsing、headroom accounting、hardware-header construction 和 diagnostic counters — 仍然會產生強大的 LPE primitives。修復程式是小規模、局部化的驗證,這是此漏洞類別的典型特徵 [1] [2] 。從防禦角度來看,最緊急的重點是 DiagSpill:一個 capability-free、可從遠端觸發的 ~8 MiB overflow,它能在 user-namespace restrictions 下存活,使 SCTP module exposure 成為具體的強化決策,而非理論上的考量。