| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: use hlist_del_init_rcu for state_cache and state_cache_input
Commit 14acf9652e56 ("xfrm: defensively unhash xfrm_state lists in
__xfrm_state_delete") converted bydst/bysrc/byseq/byspi from
hlist_del_rcu() to hlist_del_init_rcu() so that a second
__xfrm_state_delete() on the same object becomes a no-op rather than a
write through LIST_POISON pprev. It missed state_cache and
state_cache_input, which kept hlist_del_rcu():
- hlist_del_rcu() leaves pprev = LIST_POISON2 (non-NULL), so
hlist_unhashed() returns false.
- hlist_del_init_rcu() leaves pprev = NULL, so hlist_unhashed()
returns true.
A second __xfrm_state_delete() therefore enters __hlist_del() on the
already-deleted state_cache/state_cache_input nodes and does
WRITE_ONCE(*pprev, next) through LIST_POISON2 — a write use-after-free
once the slab is reused. The corruption can in turn cause a subsequent
hlist_for_each_entry_rcu traversal to follow a dangling next pointer,
producing the read use-after-free reported in xfrm_input_state_lookup().
Switch state_cache and state_cache_input to hlist_del_init_rcu() to
match the other four lists, closing the write use-after-free and, with
it, the read use-after-free it spawns. |
| In the Linux kernel, the following vulnerability has been resolved:
exec: Cleanup POSIX timers right after de_thread()
A per-thread CPU timer holds a reference to the PID of the thread it is
attached to and, while it is armed, its node is queued in that thread's
posix_cputimers. The task is looked up by that PID.
When a non-leader thread exec()s, de_thread() changes which task owns
that PID. pid_task(timer->it.cpu.pid, PIDTYPE_PID) then returns NULL,
but the node is still queued on tsk, which is alive. timer_lock_sighand()
takes a failed lookup to mean that the node is already dequeued, so it
has nothing to undo.
begin_new_exec() calls posix_cpu_timers_exit(me) right after
exec_task_namespaces() and that removes the leftover node, so the state
normally stays invisible. But bprm->point_of_no_return is set before
de_thread(), so if unshare_files(), set_mm_exe_file(), exec_mmap() or
exec_task_namespaces() fails, the task dies before it gets there.
exit_itimers() then frees the k_itimer while its node is still queued,
and reaping tsk later erases that freed node from the rbtree.
In short:
the non-leader thread B the parent
timer_create(CLOCK_THREAD_CPUTIME_ID)
timer_settime()
arm_timer() // the node is queued on B
execve()
de_thread(B)
exchange_tids(B, leader) // B's PID now belongs to the leader
release_task(leader)
__exit_signal(leader)
posix_cpu_timers_exit(leader) // cleans leader's queue, not B's
__unhash_process(leader) // that PID has no task anymore
exec_mmap()
mmap_read_lock_killable(old_mm)
kill(B, SIGKILL)
// -EINTR
get_signal()
do_exit()
exit_itimers()
posix_timer_delete()
posix_cpu_timer_del()
posix_timer_unhash_and_free() // freed while still queued
wait4()
release_task(B)
posix_cpu_timers_exit(B)
cleanup_timerqueue()
timerqueue_del() // use-after-free
Move the POSIX timer cleanup right after de_thread() before any of the
later failure conditions brings the task into do_exit().
[ tglx: Move the cleanup right after de_thread() ] |
| In the Linux kernel, the following vulnerability has been resolved:
net: lock the socket in sock_gettstamp()
sk->sk_flags must only be changed while holding the socket lock,
because sock_set_flag() and sock_reset_flag() use non atomic
operations (__set_bit() and __clear_bit()).
sock_gettstamp() is one of the last places where a bit of sk->sk_flags
is changed from a syscall without owning the socket lock, through
sock_enable_timestamp(sk, SOCK_TIMESTAMP).
sk_set_memalloc() and sk_clear_memalloc() also change sk->sk_flags
without the socket lock, but their callers (nbd, iscsi_tcp, nvme-tcp,
sunrpc, wireguard) need a careful audit, this will be addressed in a
separate patch.
Jungwoo Lee and Wongi Lee reported an UDP socket use-after-free
caused by this bug: a SIOCGSTAMPNS_NEW ioctl racing with bind()
can cancel the SOCK_RCU_FREE bit that udp_lib_get_port() just set,
because both threads perform a read-modify-write on the same word.
CPU 0 (bind) CPU 1 (SIOCGSTAMPNS_NEW)
-------------------------------- ----------------------------
read sk_flags = F read sk_flags = F
compute F | BIT(SOCK_RCU_FREE) compute F | BIT(SOCK_TIMESTAMP)
store F | BIT(SOCK_RCU_FREE)
sk_add_node_rcu(sk, ...)
store F | BIT(SOCK_TIMESTAMP)
After the lost update, SOCK_RCU_FREE is clear while the socket is
visible to lockless UDP receive lookups. sk_destruct() then frees
the socket immediately instead of waiting for a RCU grace period,
while the receive path still holds a reference-less pointer to it:
BUG: KASAN: slab-use-after-free in ipv4_pktinfo_prepare+0x30/0x410
Read of size 8 at addr ffff888008806610 by task exploit/207
CPU: 0 UID: 1000 PID: 207 Comm: exploit Not tainted 6.12.95+ #1
ipv4_pktinfo_prepare+0x30/0x410
udp_queue_rcv_one_skb+0x51c/0x1180
udp_unicast_rcv_skb+0x109/0x350
ip_protocol_deliver_rcu+0x14b/0x310
ip_local_deliver_finish+0x29d/0x390
ip_local_deliver+0x24d/0x2a0
Only grab the socket lock when SOCK_TIMESTAMP has to be set,
to keep the common case lockless. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: virt_wifi: don't transfer operstate before register
virt_wifi_newlink() calls netif_stacked_transfer_operstate() before
register_netdevice(). If the lower device is dormant, that queues the
new netdev on lweventlist while it is still uninitialized. If
registration fails after that, for example because of an invalid name
such as "bad/name", free_netdev() immediately frees the object. A
later linkwatch_fire_event() then use-after-frees the list entry.
Move the transfer to after netdev_upper_dev_link(), as macvlan and
ipvlan already do. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/core: Reject unregistering netdevs in ib_get_eth_speed
ib_device_get_netdev() intentionally returns a referenced net_device even
when it is unregistering, so matching and cleanup callers can still find
the association. The reference keeps struct net_device allocated, but does
not guarantee that the device remains operational.
ib_get_eth_speed() uses the returned device operationally by invoking its
ethtool callback. Although that call is made under RTNL, the function does
not verify the registration state first. An asynchronous RDMA port query
can therefore call into a netdev after NETDEV_UNREGISTER and ndo_uninit
have completed.
Check for NETREG_REGISTERED while holding RTNL and return -ENODEV for a
device which is being unregistered. Keeping RTNL across the check and the
ethtool operation prevents unregister from starting between them.
Keep the speed fallback and warning under RTNL as well, so the warning can
safely read netdev->name. Drop the netdev reference before releasing RTNL
once all accesses to the device are complete. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: cfg80211: don't filter by BSS type when removing stale entries
When an assoc AP switches to a channel that already has a BSS entry,
cfg80211_update_assoc_bss_entry() removes that entry before rehashing
the real one, since the two would otherwise collide in the BSS rbtree.
The lookup for that entry also required it to match the connection's BSS
type, so an entry advertising e.g. the IBSS capability bit was left in
place, and the following cfg80211_rehash_bss() then ran into it:
WARN_ON(!cmp)
Changing the type shouldn't really happen, but can be triggered by a
rogue AP/device, so drop the check and remove any entries matching
the comparison. |
| In the Linux kernel, the following vulnerability has been resolved:
IB/isert: wait for deferred control PDU completions before releasing the connection
isert_send_done() hands ISTATE_SEND_TASKMGTRSP, ISTATE_SEND_REJECT and
ISTATE_SEND_TEXTRSP completions off to isert_comp_wq and returns. The work
item then runs isert_completion_put() -> isert_put_cmd(), which reads
isert_conn->conn and takes conn->cmd_lock.
Nothing orders that work item against teardown. isert_wait_conn() queues
isert_release_work, which frees isert_conn, and iscsit_close_connection()
frees the iscsit_conn right after it returns, so the queued work can run
against freed memory.
Count the deferred control PDU completions per connection and let
isert_wait_conn() wait for them before the release work is queued.
ISTATE_SEND_LOGOUTRSP is deliberately not counted: that branch runs
iscsit_logout_post_handler(), which ends up waiting for
conn->conn_wait_comp, and that completion is only sent by
iscsit_close_connection() after it has called iscsit_wait_conn().
Waiting for it here would deadlock. Its wait stays the existing
isert_wait4logout().
The splat below is from a kernel with tracing printk()s and an msleep(200)
injected into isert_do_control_comp() to widen the window:
BUG: KASAN: slab-use-after-free in isert_put_cmd+0x53d/0x620
Read of size 8 at addr ffff8881054f1038 by task kworker/u17:1/182
CPU: 0 UID: 0 PID: 182 Comm: kworker/u17:1 Tainted: G B 7.2.0-rc5-TWIDE-gb8babf08acc7 #1 PREEMPT(lazy)
Tainted: [B]=BAD_PAGE
Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
Workqueue: isert_comp_wq isert_do_control_comp
Call Trace:
<TASK>
dump_stack_lvl+0x53/0x70
print_report+0xd0/0x630
? __pfx__raw_spin_lock_irqsave+0x10/0x10
? _raw_spin_unlock_irqrestore+0x3e/0x70
? isert_put_cmd+0x53d/0x620
kasan_report+0xce/0x100
? isert_put_cmd+0x53d/0x620
isert_put_cmd+0x53d/0x620
? isert_completion_put+0x305/0x330
? isert_do_control_comp+0x2ef/0x310
process_one_work+0x633/0x1030
? assign_work+0x11d/0x370
worker_thread+0x45b/0xd10
? __pfx_worker_thread+0x10/0x10
? __pfx_worker_thread+0x10/0x10
kthread+0x2c6/0x3b0
? recalc_sigpending+0x15c/0x1e0
? __pfx_kthread+0x10/0x10
ret_from_fork+0x36e/0x5a0
? __pfx_ret_from_fork+0x10/0x10
? __switch_to+0x572/0xdd0
? __pfx_kthread+0x10/0x10
ret_from_fork_asm+0x1a/0x30
</TASK>
Allocated by task 48:
kasan_save_stack+0x33/0x60
kasan_save_track+0x14/0x30
__kasan_kmalloc+0x8f/0xa0
__kmalloc_cache_noprof+0x158/0x370
isert_cma_handler+0x1e3/0x2ae0
cma_cm_event_handler+0x3e/0x240
cma_ib_req_handler+0x17d9/0x4490
cm_process_work+0x41/0x330
cm_work_handler+0x5727/0xc160
process_one_work+0x633/0x1030
worker_thread+0x45b/0xd10
kthread+0x2c6/0x3b0
ret_from_fork+0x36e/0x5a0
ret_from_fork_asm+0x1a/0x30
Freed by task 184:
kasan_save_stack+0x33/0x60
kasan_save_track+0x14/0x30
kasan_save_free_info+0x3b/0x60
__kasan_slab_free+0x43/0x70
kfree+0x121/0x380
iscsit_close_connection+0x7cf/0x1e60
iscsit_take_action_for_connection_exit+0x1b6/0x360
iscsi_target_tx_thread+0x472/0x690
kthread+0x2c6/0x3b0
ret_from_fork+0x36e/0x5a0
ret_from_fork_asm+0x1a/0x30 |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: cfg80211: don't free driver-owned scan requests
When an interface goes down while a scan is running, cfg80211 completes
the scan towards userspace and frees the scan request. However, the
driver can be convinced that it owns the request, since the cancellation
is (intended to be) asynchronous.
The WARN_ON() in the netdev notifier was meant to catch this, but it's
not actually avoidable, so it triggers and we get a UAF in scan_done().
There doesn't seem to be a great way around it, so just track that the
driver is still convinced it owns the request, and then just free it on
completion if it was already cancelled. Also remove the warnings since
they can trigger in the intended architecture. |
| In the Linux kernel, the following vulnerability has been resolved:
esp: downgrade zerocopy managed frags before mutating skb frags
On the out-of-place output path (esp->inplace == false) ESP rewrites the
skb frag array: esp_output_head() appends a trailer frag and
esp_output_tail() replaces the frags with a destination page, both
referenced with get_page().
When the skb carries zerocopy managed frags (SKBFL_MANAGED_FRAG_REFS) the
payload frags are owned by the ubuf and must not be referenced or
unreferenced individually, but ESP mutates the frag array without ever
downgrading the skb. This breaks the managed-frag invariant two ways:
- esp_ssg_unref() walks the source scatterlist and drops a page
reference for every frag, including the ubuf-owned payload frags,
pushing their refcount below the GUP pin bias while the pages are
still pinned, i.e. a use-after-free of the zerocopy pages;
- esp_output_tail() installs its destination page as frag 0 with
get_page() but leaves SKBFL_MANAGED_FRAG_REFS set, so
skb_release_data() takes the skip_unref branch and never drops that
reference, leaking the x->xfrag page at packet rate.
Fix this the way every other frag-mutating site does (__ip_append_data(),
__ip6_append_data(), tcp_sendmsg_locked()) and call
skb_zcopy_downgrade_managed() before ESP touches the frag array: it takes
a real reference on each existing frag and clears SKBFL_MANAGED_FRAG_REFS,
so the per-frag unref in esp_ssg_unref() and the frag release in
skb_release_data() are both balanced and no mixed-ownership frag array is
left behind. |
| In the Linux kernel, the following vulnerability has been resolved:
cifs: Fix server use-after-free in cifs_chan_skip_or_disable()
When a secondary channel is no longer supported by the server,
cifs_chan_skip_or_disable() drops the channel reference with
cifs_put_tcp_session() and then continues to use the server pointer by
calling cifs_signal_cifsd_for_reconnect() on it and reading its
primary_server pointer. cifs_put_tcp_session() can drop the last
reference of the channel and tear it down, so both the channel and the
primary server (whose reference is also dropped by
cifs_put_tcp_session()) can be freed before they are signaled for
reconnect.
Signal the channel and the primary server and capture the primary
server pointer before dropping the channel reference with
cifs_put_tcp_session(). |
| In the Linux kernel, the following vulnerability has been resolved:
smb: client: fix use-after-free of iface in cifs_try_adding_channels()
cifs_try_adding_channels() iterates ses->iface_list with
list_for_each_entry_safe_from(), which captures the next entry
(niface) under iface_lock. The loop body then drops iface_lock for
the whole duration of cifs_ses_add_channel().
A concurrent interface refresh (SMB3_request_interfaces() ->
parse_server_interfaces()) marks all ifaces inactive and removes and
frees any that are not re-advertised via list_del() + kref_put(),
where release_iface() is a bare kfree(). Since niface typically has
no channel holding a reference, the list reference is its last and it
can be freed inside the unlocked window. On continue, the iterator
advance step then dereferences niface->iface_head.next, and the loop
body reads iface->rdma_capable/is_active, both on freed memory.
Fix this by never keeping an unreferenced list pointer across the
unlocked window. Each channel attempt now re-scans the list from the
head under iface_lock, takes a kref on the selected candidate, and
passes only that referenced candidate to cifs_ses_add_channel().
weight_fulfilled still tracks selection progress, so restarting the
scan preserves the original weighted distribution and the
weight_fulfilled-before-kref_put ordering on the failure path.
Add a per-pass attempts cap so a flapping interface refresh cannot
keep the inner loop spinning within a single tries increment. |
| In the Linux kernel, the following vulnerability has been resolved:
drm/msm: RCU-free the scheduler-containing ring and VM objects
Both struct msm_ringbuffer and struct msm_gem_vm embed a struct
drm_gpu_scheduler. msm_ringbuffer_destroy() and the VM free callback
msm_gem_vm_free() call drm_sched_fini() on the embedded scheduler and then
free the containing object with plain kfree().
drm_sched_fence_get_timeline_name() returns fence->sched->name, and the
scheduler fence keeps a .release callback so it is not ops-detached on
signalling. A finished fence exported to userspace (the submit out-fence, or
a VM_BIND fence, via sync_file / drm_syncobj) keeps pointing at the embedded
scheduler after the ring/VM is freed, so a later get_timeline_name() --
reachable unprivileged through SYNC_IOC_FILE_INFO -- dereferences freed slab
memory (KASAN slab-use-after-free read).
Per the dma-fence lifetime contract the exporter must keep the data backing a
signalled fence alive for an RCU grace period. Free the scheduler-containing
objects with kfree_rcu() instead of kfree().
Patchwork: https://patchwork.freedesktop.org/patch/750234/ |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/core: fix refcount bug in iwpm_get_nlmsg_request()
iwpm_get_nlmsg_request() initializes refcount _after_ list_add_tail()
making it accessible to global list where another CPU can kref_get()
on nlmsg_request causing a refcount "addition on 0" bug. Fix this
by initializing kref _before_ list_add_tail() so refcount for
nlmsg_request can be incremented/decremented normally. In addition,
also initialize every field before list_add_tail(). |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/ucma: Serialize join and leave on copy_to_user failure
rdma_join_multicast() queues RoCE work that later reads the ucma_multicast
through event->param.ud.private_data, then list_add()s the CMA multicast
at the head of id_priv->mc_list. rdma_leave_multicast() matches only by
sockaddr and destroys the first hit.
ucma_process_join() used to drop ctx->mutex after a successful join and
retake it only if copy_to_user() failed. Two concurrent JOIN_MCAST calls
with the same address can therefore insert a second CMA entry before the
first thread's leave. leave then cancels the newer work and the older
worker still dereferences the ucma_multicast that the first thread frees.
Keep ctx->mutex held from rdma_join_multicast() through copy_to_user() and,
on -EFAULT, through rdma_leave_multicast() so leave cannot miss this join.
Do not leave if join itself failed: that path never published this address
on mc_list, and a leave-by-addr would destroy an earlier successful join. |
| In the Linux kernel, the following vulnerability has been resolved:
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
Kijo analyzed another race in the POSIX CPU timer code:
Commit bf635681c906 converted cpu_timer::firing from a tristate value to a
boolean. This lost the distinction between "not owned by the firing list"
and "still owned, but delivery was canceled". The resulting race is:
expiry handler timer_settime() timer_delete()
-------------- --------------- --------------
collect timer onto
private firing list
firing = true
observes firing = true
firing = false
return TIMER_RETRY
wait for handler
observes firing = false
finish deletion
unhash and free timer
resume list traversal
read freed elist.next
-> UAF
The firing bit is clearly the wrong indicator since that commit.
Check whether the timer is queued on the expiry list or not instead. If it
is queued clear the firing bit to prevent signal delivery as before and
return TIMER_RETRY so the caller unlocks the timer which allows the expiry
code to make progress and remove it from the list. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: libipw: reject too-short association responses
libipw_handle_assoc_resp() reads the capability, status and aid fields
of the 30-byte association response prefix and then computes the
information element length as
stats->len - sizeof(*frame)
stats->len is a u16 and sizeof() has type size_t, so the subtraction is
evaluated as size_t and wraps instead of going negative. Truncating
that to the u16 length parameter of libipw_parse_info_param() turns a
frame shorter than the fixed fields into a length near 64 KiB, and the
parser then reads past the receive buffer.
Both the ipw2100 and ipw2200 management receive paths reach this
function having established only that the frame carries the generic
24-byte three-address header.
Reject the frame before any fixed field is touched.
Found by an AI-assisted review of length arithmetic in management frame
parsers. Verified with a KUnit case under Generic KASAN on arm64 under
QEMU; I do not have the hardware, so it is not tested on a real device. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
syzbot reported a suspicious RCU usage warning in ip6_pkt_drop():
WARNING: suspicious RCU usage in ip6_pkt_drop
include/net/addrconf.h:389 suspicious rcu_dereference_check() usage!
Call Trace:
__in6_dev_get_safely include/net/addrconf.h:389 [inline]
ip6_pkt_drop+0x596/0x610 net/ipv6/route.c:4620
ip6_pkt_discard+0x1c/0x30 net/ipv6/route.c:4651
xfrm_trans_reinject+0x324/0x630 net/xfrm/xfrm_input.c:806
process_one_work kernel/workqueue.c:3322 [inline]
process_scheduled_works+0xa8e/0x14e0 kernel/workqueue.c:3405
worker_thread+0xa47/0xfb0 kernel/workqueue.c:3486
When commit 4f4920669d21 ("xfrm: Reinject transport-mode packets through
workqueue") converted xfrm_trans_reinject from a tasklet to a workqueue,
the reinjection loop ceased running in softirq context. Workqueue workers
run in process context where local_bh_disable() does not enter an RCU
read-side critical section under CONFIG_PREEMPT_RCU.
Because finish callbacks (such as ip6_rcv_finish) expect to run under an
RCU read lock (performing route lookups, l3mdev lookups, and accessing
RCU-protected data structures), invoking them in workqueue context without
rcu_read_lock() triggers RCU lockdep warnings.
Furthermore, packets queued to the workqueue via xfrm_trans_queue_net()
may carry non-refcounted (noref) dst entries (e.g. from ip_route_input_noref).
Additionally, on netdevice unregistration, dst_dev_put() replaces dst->dev
with blackhole_netdev, so dst entries do not keep skb->dev alive while
queued in the workqueue.
Fix these issues by:
1. Calling skb_dst_force(skb) in xfrm_trans_queue_net() while still in the
caller's RCU section to ensure dst is reference-counted before queuing.
2. Holding a reference on skb->dev via dev_hold()/dev_put() across workqueue
deferral so skb->dev remains valid during finish() callback processing.
3. Acquiring rcu_read_lock() around the finish callback invocation loop in
xfrm_trans_reinject(). |
| In the Linux kernel, the following vulnerability has been resolved:
hwmon: (w83791d) remove fan/pwm 4-5 sysfs group on remove
When the fan/pwm 4-5 pins are not used as GPIO, w83791d_probe()
creates the w83791d_group_fanpwm45 sysfs group on the I2C client
device.
The probe error path removes this group when a later initialization
step fails, but the normal remove path only removes w83791d_group.
As a result, the optional fan/pwm 4-5 sysfs files can remain after the
driver is unbound.
The callbacks associated with these files access the driver data,
which is devm allocated and released after driver unbind. Leaving the
sysfs files behind can therefore result in accesses to stale driver
data.
Remove w83791d_group_fanpwm45 during normal teardown as well.
This issue was found by manual code inspection. |
| In the Linux kernel, the following vulnerability has been resolved:
mips: select CONFIG_WEAK_REORDERING_BEYOND_LLSC from CONFIG_EYEQ
On I6500 CPU cores, lld and scd give no ordering guarantees (same as all
other instructions). To respect the assumption that arch_cmpxchg() is
fully ordered, we must inject sync instructions above and below our
lld/scd loops using the already in place WEAK_REORDERING_BEYOND_LLSC
infrastructure.
Otherwise, bad things can happen:
[ 34.054496] CPU 3 Unable to handle kernel paging request at virtual address 0000000000000000, epc == a80000080838e01c, ra == a80000080838dfc4
[ 34.054559] Oops[#1]:
[ 34.069561] CPU: 3 UID: 0 PID: 170 Comm: pipe_race Not tainted 7.2.0-rc6-01553-gb73c35220968-dirty #103 VOLUNTARY
[ 34.079932] Hardware name: Mobile EyeQ5 MP5 Evaluation board
[ 34.085592] $ 0 : 0000000000000000 0000000000000001 0000000000000000 0000000000000000
[ 34.093616] $ 4 : a800000808ee2618 000000000b7a879d 0000000000001000 0000000000000000
[ 34.101638] $ 8 : 0000000000e3f2c9 0000000000000000 a800000808a2a9f8 0000000000000000
[ 34.109660] $12 : a8000008139ffcd8 ffffffff84080018 a80000080837fae0 7878787878787878
[ 34.117682] $16 : a800000807e82940 0000000000001000 0000000000000000 0000000000000000
[ 34.125704] $20 : a800000802920e00 a8000008139ffdf8 a800000802649400 0000000000e3f2c9
[ 34.133726] $24 : 0000000000000006 00000001200406e0
[ 34.141783] $28 : a8000008139fc000 a8000008139ffd10 0000000000e3f2c8 a80000080838dfc4
[ 34.149837] epc : a80000080838e01c anon_pipe_read+0xd4/0x428
[ 34.155697] ra : a80000080838dfc4 anon_pipe_read+0x7c/0x428
[ 34.161549] Status: 140000e3 KX SX UX KERNEL EXL IE
[ 34.166551] Cause : 40800408 (ExcCode 02)
[ 34.170574] BadVA : 0000000000000000
[ 34.174161] PrId : 0001b028 (MIPS I6500)
[ 34.178183] Process pipe_race (pid: 170, threadinfo=000000005ca35720, task=00000000e1013890, tls=000000014ebbb780)
[ 34.188568] Stack : a800000802649400 0000000000000000 0000000000000000 a8000008139ffdd0
[ 34.196623] 0000000000000fba a800000808ee0000 0000000000000001 a8000008130c3e80
[ 34.204676] a8000008080d1280 a8000008139ffd58 a8000008139ffd58 1dbd2b22ea1dd500
[ 34.212729] a800000802649400 a800000808ee0000 ffffffffffffffea 0000000000000001
[ 34.220783] 0000000000001000 0000000000000000 00000001200ae518 ffffffffffffffff
[ 34.228836] 000000fffbe0e530 a80000080837edf4 000000fffbe0e530 0000000000000000
[ 34.236890] 0000000000000000 0000000000000000 000000014ebb55a0 0000000000001000
[ 34.244943] 0000000000000001 a800000802649400 0000000000000000 0000000000000000
[ 34.252996] 0000000000000000 0000400400000000 0000000000000000 1dbd2b22ea1dd500
[ 34.261049] 00000000140000e3 a800000802649400 a800000802649400 a800000808ee0000
[ 34.269103] ...
[ 34.271568] Call Trace:
[ 34.274026] [<a80000080838e01c>] anon_pipe_read+0xd4/0x428
[ 34.279533] [<a80000080837edf4>] vfs_read+0x25c/0x318
[ 34.284607] [<a80000080837faac>] ksys_read+0x104/0x138
[ 34.289763] [<a80000080802b9cc>] syscall_common+0x44/0x68
[ 34.295187]
[ 34.296689] Code: f84000cf 02209825 de020010 <dc420000> d8400004 02002825 0040f809 02802025 f84000c3
[ 34.306504]
[ 34.308099] ---[ end trace 0000000000000000 ]---
My initial reproducer was the xdp-tools test suite. A standalone
reproducer would be an lld/scd loop that, when the read is reordered by
the CPU, triggers a fault. We can achieve this from userspace by
stressing an anonymous pipe, which uses a mutex. Program used:
// SPDX-License-Identifier: GPL-2.0
// pipe_race.c - reproducer for MIPS LL/SC reordering vs fs/pipe.c
//
// Two userspace processes on an anonymous pipe:
// parent = writer: tight write() loop
// child = reader: tight read() loop
#define _GNU_SOURCE
#include <assert.h>
#include <errno.h>
#include <sched.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/types.h>
#include
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
KVM: PPC: Book3S HV: fix use-after-free in kvmhv_emulate_tlbie_all_lpid()
kvmhv_emulate_tlbie_all_lpid() iterates the nested-guest IDR and drops
mmu_lock before calling kvmhv_emulate_tlbie_lpid(), but does not hold a
reference on the kvm_nested_guest pointer obtained from the IDR. A
concurrent vCPU issuing a single-LPID tlbie (is=2, ric=2) can race
through kvmhv_flush_nested() -> kvmhv_remove_nested() -> idr_remove /
--refcnt -> kvmhv_release_nested() -> kfree(gp) in that window, leaving
the iterating vCPU with a dangling pointer. The subsequent
mutex_lock(&gp->tlb_lock) and accesses to gp->shadow_pgtable,
gp->shadow_lpid and gp->l1_host all touch freed memory. The free path
is fully L1-controlled.
Fix this by incrementing gp->refcnt inside the loop before dropping
mmu_lock, mirroring what kvmhv_get_nested() does, and releasing the
reference with kvmhv_put_nested() after the per-guest work completes.
This is the same get/put discipline already used at every other
call site that drops mmu_lock while holding a nested-guest pointer. |