1. 簡介

共享記憶體通訊(Shared Memory Communications,SMC) 定義於 RFC 7609,是一種 kernel-transparent 的傳輸方式,讓標準 TCP socket 在雙方 peer 都宣告支援後,能透明地切換到 共享記憶體 或 RDMA 資料路徑 ,否則退回一般 TCP。其 SMC-D variant 過去只能透過 IBM Z 的 Internal Shared Memory(ISM)裝置觸及,這限制了多年來的獨立程式碼審查。建構於 Direct Internal Buffer Sharing(DIBS) 抽象層之上的虛擬傳輸 dibs_loopback 問世後,讓 SMC-D 能在一般 x86 硬體上完全透過 loopback 執行,因此 XBOW 發現了 CVE-2026-72018,這是 Linux kernel 的 SMC-D 驅動程式中的越界寫入,並將一個微弱的 primitive 轉變成可用的 root exploit [1] 。這個案例之所以有啟發性,正是因為其底層 primitive——在一個半可控的 offset 上固定 16 位元組的零寫入——通常被認為太微弱而不值得追究 [2] ,然而它本身即足以在不需任何輔助資訊洩漏的情況下取得 root shell。

XBOW證明零寫入漏洞也能摧毀SMC-D loopback的防線 | 資訊安全新聞

2. 威脅模型與漏洞路徑

連線建立使用短暫的 CLC handshake(Proposal/Accept/Confirm),其中 peer 傳送四個由攻擊者控制的欄位來描述其接收緩衝區: gid 、 token 、 dmbe_idx 與 dmbe_size 。由於 SMC-D 原始的威脅模型假設 peer 是透過硬體中介緩衝區通訊的可信主機,當程式碼變得可被任意本地或網路 peer 觸及後,從未加入任何邊界檢查。

  1. // net/smc/af_smc.c:739-750
  2. static void smcd_conn_save_peer_info(struct smc_sock *smc,
  3. struct smc_clc_msg_accept_confirm *clc)
  4. {
  5. int bufsize = smc_uncompress_bufsize(clc->d0.dmbe_size); // [1] decode peer-supplied size code into bytes
  6. smc->conn.peer_rmbe_idx = clc->d0.dmbe_idx; // [2] store peer-chosen buffer element index, unchecked
  7. smc->conn.peer_token = ntohll(clc->d0.token); // [3] store peer-chosen buffer token, unchecked
  8. smc->conn.peer_rmbe_size = bufsize - sizeof(struct smcd_cdc_msg);
  9. atomic_set(&smc->conn.peer_rmbe_space, smc->conn.peer_rmbe_size);
  10. smc->conn.tx_off = bufsize * smc->conn.peer_rmbe_idx; // [4] offset = attacker bufsize * attacker index, no range check
  11. }

步驟 [4] 這行將兩個都源自 handshake 訊息的值相乘,產生傳輸 offset tx_off ,而在其後整個呼叫鏈中未施加任何上限。

  1. // net/smc/smc_tx.c:303-314
  2. int smcd_tx_ism_write(struct smc_connection *conn, void *data, size_t len,
  3. u32 offset, int signal)
  4. {
  5. int rc;
  6. rc = smc_ism_write(conn->lgr->smcd, conn->peer_token,
  7. conn->peer_rmbe_idx, signal,
  8. conn->tx_off + offset, // [1] unchecked offset forwarded into the ISM layer
  9. data, len); // [2] data/len describe the 32-byte CDC header, not peer input
  10. return rc;
  11. }
  1. // drivers/dibs/dibs_loopback.c:234-257
  2. static int dibs_lo_move_data(struct dibs_dev *dibs, u64 dmb_tok,
  3. unsigned int idx, bool sf, unsigned int offset,
  4. void *data, unsigned int size)
  5. {
  6. struct dibs_lo_dmb_node *rmb_node = NULL, *tmp_node;
  7. // [1] lookup target buffer by token; rmb_node left NULL on miss
  8. if (!rmb_node) { read_unlock_bh(&ldev->dmb_ht_lock); return -EINVAL; } // [2] only a NULL-token check, no length check
  9. memcpy((char *)rmb_node->cpu_addr + offset, data, size); // [3] write proceeds at cpu_addr+offset regardless of buffer length
  10. return 0;
  11. }

驅動程式只驗證存在符合 token 的緩衝區;從未驗證 offset + size 是否維持在 rmb_node->len 之內。目的緩衝區是單一固定大小的 kernel folio,因此上述步驟 [4] 計算出、超過該 folio 長度的任何 offset 都會落在相鄰的 heap 區域。從 CLC 欄位到 memcpy 之間不存在任何邊界檢查,因此只要 dmbe_idx 大到足以將 offset 推超過 dmb_node->len,就會把 32 位元組的 CDC header 寫入相鄰的 kernel heap [1] 。

3. 從遠端 Primitive 到本地 Exploit

原本的 write-what primitive 需要真正的遠端 peer,這在過去意味著需要真實 ISM 硬體的 VM-escape 框架。loopback 傳輸讓「client」與「server」兩種角色能在單一主機上執行,將此 bug 重新定位為本地權限提升的候選,任何能建立 SMC-D 連線並操控自身 loopback 流量的使用者皆可觸及。

由於 loopback 一律宣告 dmbe_idx = 0 ,攻擊者必須攔截並改寫傳輸中的 CLC 訊息。相關的 on-wire 結構帶有三個關鍵欄位:

  1. // net/smc/smc_clc.h
  2. struct smcd_clc_msg_accept_confirm_common {
  3. __be64 gid; // [1] sender/device identifier, not attacked directly
  4. __be64 token; // [2] selects which DMB (buffer) receives the write
  5. u8 dmbe_idx; // [3] forged non-zero to push offset past the buffer boundary
  6. #if defined(__LITTLE_ENDIAN_BITFIELD)
  7. u8 reserved3 : 4,
  8. dmbe_size : 4; // [4] forged to control the multiplier (bufsize) in tx_off
  9. #endif
  10. u16 reserved4;
  11. __be32 linkid;
  12. } __packed;

一條 nftables NFQUEUE 規則將 loopback 封包導向 userspace,在那裡解析 CLC Accept/Confirm 訊息,改寫 token / dmbe_idx / dmbe_size ,重新計算 IPv4/TCP checksum,再將封包重新注入——僅需 CAP_NET_ADMIN,而這本來就是註冊 SMC-D 裝置所必需的。

4. 將零寫入轉變為 Root

在 send-only 的連線上,producer cursor 為非零,但 consumer cursor 與 32 位元組 CDC header 尾端的保留位元組維持為零,因此在選定的、對齊 folio 的 offset 上提供一段可靠的 16 位元組零值:

  1. // net/smc/smc_cdc.h
  2. struct smcd_cdc_msg {
  3. struct smc_wr_rx_hdr common; // [1] type field, not part of the zero run
  4. u8 res1[7]; // [2] reserved, not exploited
  5. union smcd_cdc_cursor prod; // [3] producer cursor, nonzero, not exploited
  6. union smcd_cdc_cursor cons; // [4] consumer cursor, reliably zero on send-only path
  7. u8 res3[8]; // [5] reserved, reliably zero -- together [4]+[5] form the 16-byte primitive
  8. } __aligned(8);

此 primitive 選定的目標是 kernel 的 cred 結構,其身份欄位連續排列:

  1. // include/linux/cred.h
  2. struct cred {
  3. atomic_long_t usage; // [1] refcount, must not be zeroed or the object is freed prematurely
  4. kuid_t uid; // [2] real UID, left untouched by this exploit
  5. kgid_t gid;
  6. kuid_t suid; // [3] saved UID -- first field inside the 16-byte write window
  7. kgid_t sgid; // [4] saved GID
  8. kuid_t euid; // [5] effective UID -- zeroing this is what grants root on permission checks
  9. kgid_t egid; // [6] effective GID
  10. kuid_t fsuid;
  11. kgid_t fsgid;
  12. };

若寫入的起始位置正好落在 suid ,這 16 個零位元組恰好涵蓋 suid 、 sgid 、 euid 與 egid 。僅將 euid 歸零即已足夠,因為 kernel 在權限檢查時會將有效 UID 為 0 的 process 視為 root;隨後由該 process 呼叫 setresuid(0,0,0) 便使權限變更永久化,並取得穩定的 root shell。由於寫入的值是常數,且目標是記憶體內的欄位而非程式碼指標,因此不需要任何指標洩漏或位址揭露。

5. Exploit 可靠性

有三個條件決定成功與否:(1) 確定寫入落在哪個實體緩衝區,作法是擷取 warm-up connection 的 token,並在寫入連線上重複使用;(2) 先安排緩衝區配置,再大量建立僅含單一認證的子程序,使它們的 cred 物件被精心佈置,緊貼在目標緩衝區之後;以及 (3) 避免破壞 exploit 自身所使用的 SMC 連線追蹤結構,作法是早期嘗試造成 kernel panic 後擴大 offset 掃描範圍。

sequenceDiagram participant A as Attacker process (uid=1000) participant NFQ as NFQUEUE/MITM (local, CAP_NET_ADMIN) participant K as Kernel SMC-D / DIBS loopback participant C as Sprayed cred objects A->>K: Open warm-up SMC-D connection K-->>A: Allocate target DMB, return token A->>C: Fork many uid=1000 children with fresh cred (capset) A->>K: Open write connection (CLC Proposal) K->>NFQ: CLC Accept/Confirm packet (loopback) NFQ->>NFQ: Rewrite token, dmbe_idx, dmbe_size; fix checksums NFQ->>K: Reinject forged CLC packet K->>K: smcd_conn_save_peer_info computes tx_off (unchecked) A->>K: Trigger CDC send (16 zero bytes at tx_off) K->>C: memcpy writes zeros past DMB folio into adjacent cred C->>C: Child polls own euid, detects zero C->>A: Signal success, call setresuid(0,0,0) A->>A: Root shell obtained

由於結果取決於開機時的 heap 佈局,此 exploit 在 100 次獨立開機中進行測量,其中 22 次成功,首次成功出現在第 7 次開機 [1] ;測試在 Ubuntu 24.04(kernel 7.1.0-rc6/x86_64)上進行,並停用緩解措施。廠商公告列出 kernel 版本 6.12.97、6.18.40 與 7.1.5 及之後版本為不受影響 [2] 。

6. 結論

此漏洞屬於一個眾所周知的類別:peer 控制的 index 被用於 offset 運算而沒有邊界檢查。此案例的獨特之處在於,寫入 primitive 本身被認為太微弱而無法實際利用——在粗略可控的 offset 上固定 16 個零位元組——通常會被降級處理,轉而偏好能提供任意值或位址控制的 primitive。這個攻擊策略顛倒了常見的思路,從『我們能寫什麼』轉向『零值在哪裡有價值』,並落在 cred 結構的身份欄位上,其中 euid 為零本身就代表特權。這顯示出,即使是因硬體移植到軟體而遺留下來的邊界檢查缺失(此處以 SMC-D 類別標記),仍然可被利用,即便最初看起來該原始機制似乎無法利用。這再次強調,任何因移除專用硬體限制而新近可觸及的子系統,都應在明確的惡意同儕模型下重新審計。