| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
seg6: set IPSKB_L3SLAVE from IP6SKB_L3SLAVE on IPIP decapsulation
When an SRv6 packet arrives on an interface enslaved to a VRF,
vrf_ip6_rcv() sets IP6SKB_L3SLAVE in IP6CB, but decap_and_validate()
has never set IPSKB_L3SLAVE in IPCB. The bit stayed clear in the
common case, and with CONFIG_IPV6_MIP6 the leftover frag_max_size of
a reassembled outer packet could even set it, with no VRF involved.
Commit 44930446dde4 ("ipv6: seg6: clear IPv4 control block on IPIP
decapsulation") then made the unreliable bit reliably clear.
The effect of the missing flag is visible with End.DX4 when a
delivery to a local address of the node reaches the socket lookup.
For example, a UDP socket bound to the enslaved ingress interface
does not receive any of the decapsulated packets, while an unbound
socket outside the VRF does.
This contradicts Documentation/networking/vrf.rst: by default the
scope of an unbound UDP or TCP socket is limited to the default VRF.
Set IPSKB_L3SLAVE for IPv4 in decap_and_validate(), which already does
the same for IPv6. The socket lookup then matches the decapsulated
packet like any other packet received on that enslaved interface. Such
a packet matches an unbound UDP or TCP socket only when
udp_l3mdev_accept or tcp_l3mdev_accept is set. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: ath11k: cleanup arsta in ath11k_mac_peer_cleanup_all()
When mac80211 removes a sta, it calls .sta_state() which in turn calls
ath11k_mac_station_remove(). In that function we clean up both peers &
arsta related resources.
But when the firmware crashes, ath11k calls ieee80211_restart_hw(), which
assumes that all driver related resources are cleaned up beforehand. This
cleanup is supposedly done by ath11k_mac_peer_cleanup_all() but does not
in fact free arsta->rx_stats / tx_stats.
Extract the arsta cleanup from ath11k_mac_station_remove() into a
new ath11k_mac_station_cleanup() and call it from both there and
ath11k_mac_peer_cleanup_all().
This should handle kmemleaks reports like:
unreferenced object 0xffffff801ae66400 (size 1024):
comm "hostapd", pid 1306, jiffies 4295011565
hex dump (first 32 bytes):
00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................
00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 ................
backtrace (crc d61c08ec):
kmemleak_alloc+0x3c/0x50
__kmalloc_cache_noprof+0x2b0/0x3e0
ath11k_mac_op_sta_state+0x1dc/0xb10
drv_sta_state+0xac/0x6f8
sta_info_insert_rcu+0x314/0x5e0
sta_info_insert+0x14/0x38
ieee80211_add_station+0x10c/0x1a0
nl80211_new_station+0x3e8/0x680
genl_family_rcv_msg_doit+0xc0/0x120
genl_rcv_msg+0x1b4/0x258
netlink_rcv_skb+0x4c/0x108
genl_rcv+0x38/0x60
netlink_unicast+0x190/0x278
netlink_sendmsg+0x15c/0x370
____sys_sendmsg+0x120/0x290
___sys_sendmsg+0x70/0xa0
Tested-on: QCN9074 hw1.0 PCI WLAN.HK.2.9.0.1-01977-QCAHKSWPL_SILICONZ-1 |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: mac80211: don't start a ROC while scanning
The ROC work can be pending when a scan starts (which requires
ROC list to be empty, but that's possible), and then a new ROC
can be added to the list and the work will pick it up.
Avoid starting that ROC if a scan made it between things, as
otherwise we'll hit a warning later:
WARNING: net/mac80211/offchannel.c:404 at ieee80211_start_next_roc+0x256/0x2d0
Workqueue: events_unbound cfg80211_wiphy_work
Call Trace:
__ieee80211_scan_completed+0x4fd/0xe40 net/mac80211/scan.c:537
ieee80211_scan_work+0x472/0x1ff0 net/mac80211/scan.c:1193
cfg80211_wiphy_work+0x410/0x570 net/wireless/core.c:513 |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: virtio: reset device before deleting virtqueues
virtsnd_remove() and virtsnd_freeze() delete the virtqueues before
resetting the device. del_vqs() frees the vring backing, but does not
provide a generic device quiesce operation. In particular, modern
virtio-pci keeps enabled queues active until the device is reset.
Reset the device before deleting the virtqueues so it can no longer
access the vring memory when that memory is released. This also covers
probe failures after DRIVER_OK, which unwind through virtsnd_remove(). |
| In the Linux kernel, the following vulnerability has been resolved:
net: skbuff: do not leave stale header offsets after pskb_carve()
pskb_carve_inside_header() and pskb_carve_inside_nonlinear() remove
the first bytes of a packet and reallocate skb->head.
All the headers that were present before the operation are gone,
but both functions call skb_headers_offset_update(skb, 0), which
is a no-op : skb->mac_header, skb->network_header,
skb->transport_header and skb->csum_start keep their old values and
now describe bytes which are no longer there.
Both helpers size the new head from the old skb_end_offset(), so the
stale offsets still land inside the new allocation. They point past
skb_tail_pointer() though, to bytes that were never initialized.
pskb_carve_inside_nonlinear() is the worst case, because it leaves a
zombie skb with an empty linear part (skb->data ==
skb_tail_pointer(skb), skb_headlen(skb) == 0), while
skb_mac_header_was_set() is still true and skb->mac_header is way
ahead of skb->data.
The only user of pskb_extract() is rds_tcp_data_recv(), and the
carved skb is queued on tinc->ti_skb_list. When the RDS incoming
message is released, rds_tcp_inc_free() calls skb_queue_purge(),
which frees the skbs with SKB_DROP_REASON_QUEUE_PURGE. This is
visible from drop_monitor, which then tries to pull back to the
(bogus) mac header :
skbuff: __skb_pull(len=234)
skb len=6968 data_len=6968 headroom=0 headlen=0 tailroom=0
end-tail=384 mac=(234,14) mac_len=14 net=(248,40) trans=288
shinfo(txflags=0 nr_frags=1 gso(size=1428 type=16 segs=5))
csum(0x100120 start=288 offset=16 ip_summed=3 complete_sw=0 valid=1 level=0)
hash(0x7b446c6c sw=0 l4=1) proto=0x86dd pkttype=0 iif=60
kernel BUG at ./include/linux/skbuff.h:2847!
Add skb_carve_reset_headers() to mark the mac and transport headers
as not set, reset the network header, clear skb->mac_len, and drop
a now meaningless CHECKSUM_PARTIAL (csum_start no longer describes
anything).
Invalidate the inner offsets as well. Unlike mac_header and
transport_header they have no "unset" sentinel, so a leftover
non-zero value still looks like a real header. Zero
skb->inner_mac_header, skb->inner_network_header,
skb->inner_transport_header, skb->inner_protocol and
skb->encapsulation, so that all the header state is invalidated in
one place.
v2: fixed an inaccurate changelog. The stale offsets stay inside the
new skb->head, which is never smaller than the old one, they
simply point past skb_tail_pointer() to bytes that are gone.
Thanks to Xuanqiang Luo for insisting on this.
Also invalidate the inner header state, as suggested by the
netdev AI review :
https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260911114922.621937-1-edumazet%40google.com |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: hda: trace PCM open only after assigning a stream
Stream assignment can fail when hardware streams are exhausted.
Move the tracepoint after the NULL check because its payload accesses
the assigned stream tag.
Detected by static analysis and reviewed with AI-assisted source auditing. |
| In the Linux kernel, the following vulnerability has been resolved:
drm/xe/shrinker: Take a runtime PM ref before shrinking non-system memory
__xe_shrinker_walk() walks the SYSTEM and TT LRUs without a runtime PM
reference. Shrinking a bo outside system memory invalidates its GPU
mappings, which needs the device resumed, so while it is runtime
suspended the page table zap trips an assert and the TLB invalidation
returns -ENODEV:
WARNING: drivers/gpu/drm/xe/xe_bo.c:770 at xe_bo_move_notify+0x1fc/0x450 [xe]
xe_bo_shrink+0x20f/0x2b0 [xe]
__xe_shrinker_walk+0x174/0x410 [xe]
xe_shrinker_scan+0x10c/0x1e0 [xe]
do_shrink_slab+0x176/0x7e0
drop_caches_sysctl_handler+0x9c/0xf0
Take a reference before walking a memory type other than XE_PL_SYSTEM
and stop there if it cannot be acquired. Reuse the shrinker's existing
acquire path, which resumes the device directly where reclaim allows
that and otherwise queues the PM worker for a later scan. Stop the walk
once the scan target is met, so a satisfied scan does not wake the
device. System memory is still reclaimed while the device is suspended.
Gate this on xe_device_is_l2_flush_optimized(), the same condition under
which xe_bo_trigger_rebind() issues the invalidation for a non-fault-mode
vm, so reclaim is unaffected elsewhere. The System CCS copy already has
its own reference in xe_bo_shrink().
Only a non-fault-mode vm can reach this, since a fault-mode vm requires
LR mode and that holds a runtime PM reference for the vm's lifetime.
Reproduced with igt@xe_madvise@dontneed-before-exec while the GPU is
runtime suspended.
v2: simplify needs_rpm check. (Matt)
retarget Fixes tag since the issue occurs with the non-fault-mode
path added by 4e7ebff69aed.
v3: handle this in xe_shrinker.c instead of xe_bo.c (Thomas)
v4: stop the walk once the scan target is met. (Sashiko)
v5: rebase on the freed page accounting fix. (Sashiko)
v6: reuse the shrinker acquire path so runtime pm can be resumed
directly instead of always queueing a worker. (Thomas)
v7: replace xe_pm_runtime_put() with xe_shrinker_runtime_pm_put(). (Thomas)
(cherry picked from commit 628f92b28bf4c371c10207daf6fc4caee0c0db2e) |
| In the Linux kernel, the following vulnerability has been resolved:
drm: Fix drm_pending_vblank_event leak in error path for out_fence_ptr
When an out_fence_ptr is provided but DRM_MODE_PAGE_FLIP_EVENT is not
set, a drm_pending_vblank_event will be allocated. If later, there is an
allocation failure or another failure at setup_out_fence(), that event
will not have base.fence set and it will not be released at
complete_signaling().
Release the event and set crtc_state->event to NULL just like in the
DRM_MODE_PAGE_FLIP_EVENT case when there is a failure at
drm_event_reserve_init(). That is, prepare_signaling() releases the
event and there is nothing to be done at complete_signaling(). Use
drm_event_cancel_free() as that will also undo drm_event_reserve_init()
in case it has been called. |
| In the Linux kernel, the following vulnerability has been resolved:
dmaengine: Fix device kref underflow in dma_chan_put()
dma_chan_get() takes chan->device->ref only on the slow path:
/* no kref on fast path */
if (chan->client_count) {
__module_get(owner);
chan->client_count++;
return 0;
}
if (!try_module_get(owner))
return -ENODEV;
if (!dma_device_get(chan->device)) { // calls kref_get_unless_zero()
dma_chan_put() drops the ref unconditionally, so every fast-path
get/put pair drops one extra device reference.
The bug fires when two conditions hold together: a non-private
provider has a persistent client holding chan->client_count > 0
and another client cycles dmaengine_get()/dmaengine_put().
When the kref hits zero, the subsequent dma_find_channel() returns
NULL even though the provider module is still loaded.
Fix this by dropping device->ref only on the last put, matching the
single slow-path get. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/rtrs-clt: Fix CQ pool leak when connect is interrupted
The client borrows shared CQ credits in the ADDR_RESOLVED handler via
ib_cq_pool_get(), before the peer is connected. create_cm() can return
-ERESTARTSYS from wait_event_interruptible_timeout() without destroying
the CM ID. The init_conns() and stop-and-destroy paths then call
destroy_con_cq_qp() while cq is still NULL (no PUT) and only afterwards
rdma_destroy_id().
CMA serializes the handler against rdma_destroy_id() with handler_mutex,
but that does not order the GET against destroy_con_cq_qp(). If
ADDR_RESOLVED has already passed the DESTROYING check, it can take
con_mutex, GET credits, and then lose the con to kfree. Device
unregister later hits WARN_ON(cq->cqe_used) in ib_cq_pool_cleanup().
Set a per-connection flag under con_mutex before CQ/QP teardown so a
racing ADDR_RESOLVED cannot borrow credits after teardown has begun. |
| In the Linux kernel, the following vulnerability has been resolved:
netlink: do not free nlk->groups while lockless readers can use it
netlink_realloc_groups() uses krealloc() under netlink_table_grab().
Whenever NLGRPSZ(groups) lands in a different kmalloc bucket, the old
bitmap is freed immediately.
Two readers of nlk->groups / nlk->ngroups do not hold the netlink
table lock:
1) sk_diag_dump_groups(). Hashed (bound) sockets are dumped from the
rhashtable walk in __netlink_diag_dump(), which only holds RCU.
Only the mc_list part of the dump takes nl_table_lock.
2) netlink_native_seq_show() (/proc/net/netlink), whose walk has been
lockless since commit 21e4902aea80 ("netlink: Lockless lookup with
RCU grace period in socket release").
Both can read a freed buffer, and sk_diag_dump_groups() can also read
past the end of the old (smaller) buffer if it happens to load the old
@groups pointer together with the new @ngroups value, copying the
result into a NETLINK_DIAG_GROUPS attribute.
This is the same class of bug that commit f773608026ee ("netlink:
access nlk groups safely in netlink bind and getname") fixed for bind()
and getname(); these two readers were missed. Simply grabbing the table
lock in sk_diag_dump_groups() is not an option, because it is also
called with nl_table_lock already held from the mc_list section of the
dump.
Make the lockless readers safe instead:
- Allocate a new bitmap and free the old one after an RCU grace period,
instead of relying on the implicit kfree() done by krealloc().
- Publish @groups before @ngroups, both with release semantics, and have
the lockless readers load @ngroups first. A reader can then never pair
the new (bigger) size with the old (smaller) buffer, and a reader
picking up the new pointer while still seeing the old size is
guaranteed to see the initialized bitmap.
netlink_realloc_groups() is called from process context (bind() and
setsockopt()), so kfree_rcu_mightsleep() can be used, once the table
has been released. |
| In the Linux kernel, the following vulnerability has been resolved:
Bluetooth: btintel_pcie: fix off-by-one bounds check in RX submit
btintel_pcie_submit_rx() used frbd_index > rxq->count to guard the
FRBD array access, allowing frbd_index == rxq->count to pass through
and index one element past the end of the array. Change the check to
>= rxq->count so every out-of-range index is rejected.
This issue was reported by Claude Mythos. |
| In the Linux kernel, the following vulnerability has been resolved:
cachefiles: Fix error return when vfs_mkdir() fails
When vfs_mkdir() fails, the error code is not extracted from the
returned error pointer. This causes mkdir_error to be reached with
ret=0, which leads to returning ERR_PTR(0) (NULL) instead of a
proper error pointer.
Fix this by extracting the error code from the error pointer when
vfs_mkdir() fails. |
| In the Linux kernel, the following vulnerability has been resolved:
x86/kprobes: Fix crash when probing CS CALL instructions
When using eBPF to probe CS CALL instructions within a function,
a crash can be triggered.
The eBPF tool probes offset 257 of the __hrtimer_run_queues()
function:
<__hrtimer_run_queues+249>: nopl 0x0(%rax,%rax,1)
<__hrtimer_run_queues+254>: mov %r14,%rdi
<__hrtimer_run_queues+257>: cs call <__x86_indirect_thunk_r12>
<__hrtimer_run_queues+263>: mov %eax,%r12d
<__hrtimer_run_queues+266>: xchg %ax,%ax
<__hrtimer_run_queues+268>: mov %r13,%rdi
Which triggers this crash:
BUG: unable to handle page fault for address: 00000000000f41c9
#PF: supervisor write access in kernel mode
#PF: error_code(0x0002) - not-present page
PGD 0 P4D 0
Oops: 0002 [#1] SMP NOPTI
CPU: 1 PID: 0 Comm: swapper/1 Kdump: loaded Tainted: P
RIP: 0010:__hrtimer_run_queues+0x106/0x230
Note that __hrtimer_run_queues+0x106 is __hrtimer_run_queues+262, which is
at the 6th byte of the above CS CALL instruction. Since the CS CALL
instruction occupies 6 bytes, the exception occurred in the middle of that
call instruction.
The root cause is that when using eBPF tools to probe in the middle of a
function, a kprobe with INT3 is used as the underlying implementation.
During single-step emulation of the original CALL instruction,
int3_emulate_call() assumes that the probed CALL instruction is 5 bytes
long. However, the actual CS-prefixed CALL instruction occupies 6 bytes,
so it constructs an incorrect exception return address. When the CPU
returns from the kprobe handler, the next instruction to be executed is at
the address of the last byte of that CS CALL instruction. Coincidentally,
starting from that address, the CPU fetches and decodes a completely
different instruction, which ultimately triggers a kernel crash.
Fix the issue by using the actual instruction length obtained from
the instruction decoder when constructing the exception return
address, rather than relying on the hardcoded CALL_INSN_SIZE macro.
[ mingo: Refined the changelog ] |
| In the Linux kernel, the following vulnerability has been resolved:
Bluetooth: hci_qca: Do not write to the serial port after it is closed
hci_uart_close() closes the serdev port if HCI_QUIRK_NON_PERSISTENT_SETUP
is set (for example, for the WCN399x family). A failed hci_dev_open_sync()
following a successful qca_setup() calls hdev->close() but not
hdev->shutdown(), so the port is closed while power->vregs_on is left true.
qca_serdev_remove() then passes its power->vregs_on test and calls
qca_power_off(), which writes to the closed port unconditionally.
Seen on a WCN3988 by unbinding the driver after a controller failure. The
trace below is from a 7.0.0 based kernel, where qca_power_off() was still
named qca_power_shutdown():
Unable to handle kernel NULL pointer dereference at virtual address
0000000000000038
Call trace:
tty_set_termios+0x50/0x238 (P)
ttyport_set_baudrate+0x84/0xc0
serdev_device_set_baudrate+0x24/0x40
qca_power_shutdown+0x158/0x1fc [hci_uart]
qca_serdev_remove+0x54/0x68 [hci_uart]
serdev_drv_remove+0x1c/0x2c
device_remove+0x4c/0x80
device_release_driver_internal+0x1cc/0x224
device_driver_detach+0x18/0x24
unbind_store+0xb4/0xc0
Check HCI_UART_PROTO_READY, which hci_uart_close() clears in the same place
it closes the port, before writing to it. The regulator disable is left
unconditional so the controller is still powered down.
The dangling serport->tty that turns this into a use-after-free is
addressed in a separate patch. |
| In the Linux kernel, the following vulnerability has been resolved:
dmaengine: mmp_pdma: fix wrong sg length in mmp_pdma_prep_slave_sg()
In mmp_pdma_prep_slave_sg(), for_each_sg() iterates the scatterlist
putting each entry into 'sg', but the entry length is read from 'sgl'
(the list head) instead of 'sg' (the current entry):
for_each_sg(sgl, sg, sg_len, i) {
addr = sg_dma_address(sg);
avail = sg_dma_len(sgl); /* should be 'sg' */
Consequently 'avail' is always the length of the first entry. For
multi-sg lists this causes out-of-bounds reads when a later entry is
shorter than the first, and silent data loss when it is longer.
Single-sg or uniformly-sized lists happen to mask the issue. |
| In the Linux kernel, the following vulnerability has been resolved:
net: bcmgenet: restore the hardware filters on open
bcmgenet_hfb_init() runs INIT_LIST_HEAD() on priv->rxnfc_list, which drops
every rule off the list, and bcmgenet_open() calls it on each ifup. Every
rule the user configured is silently lost:
# ethtool -N eth0 flow-type ether dst $MAC action 0
Added rule with ID 0
# ethtool -n eth0 | grep -c Filter:
1
# ip link set eth0 down && ip link set eth0 up
# ethtool -n eth0 | grep -c Filter:
0
Initialise the lists once at probe and restore the rules on open, as
bcmgenet_resume() already does. |
| In the Linux kernel, the following vulnerability has been resolved:
drm/vc4: Use managed KMS polling to fix UAF on unbind
vc4_kms_load() calls drm_kms_helper_poll_init() but the driver provides
no matching drm_kms_helper_poll_fini(). The output poll work stays
scheduled after unbind and runs on the freed drm_device:
# modprobe vc4; rmmod vc4; sleep 10
BUG: KASAN: slab-use-after-free in delayed_work_timer_fn
BUG: KASAN: slab-use-after-free in drm_client_dev_hotplug [drm]
Workqueue: events output_poll_execute [drm_kms_helper]
Allocated by task 171: __devm_drm_dev_alloc
Freed by task 262 (rmmod): drm_dev_put / component_del
Use drmm_kms_helper_poll_init() so polling is finalized with the device,
as other drivers do. |
| In the Linux kernel, the following vulnerability has been resolved:
ALSA: 6fire: fix OOB write from device-reported iso length
usb6fire_pcm_in_urb_handler() sizes each outgoing isochronous packet as
(actual_length - 4) / (in_n_analog << 2) * (out_n_analog << 2) + 4, where
actual_length is the unsigned length the device reported for the matching
IN packet. A packet completed with status 0 and actual_length < 4 wraps
the subtraction to 0x7fffffec; a zero-length isochronous packet is legal
on the bus, and the preceding loop rejects only non-zero status. The sum
reaches memset() on out_urb->buffer, a 4832-byte object from
kcalloc(PCM_MAX_PACKET_SIZE, PCM_N_PACKETS_PER_URB).
Even without the wrap the result is out of bounds: at 88.2/96 kHz the
4-in/6-out scaling turns a full 420-byte IN packet into 628, so eight
packets span 5024 bytes of that buffer. usb_submit_urb() rejects an
over-long descriptor only after the memset() and the
usb6fire_pcm_playback() copy of user PCM data have run.
Guard the subtraction as the sibling usb6fire_pcm_capture() already does,
and limit the frame count to what fits in rt->out_packet_size, the OUT
endpoint's wMaxPacketSize. This bounds total_length by the buffer size
while keeping each packet length aligned to a whole output frame.
BUG: KASAN: out-of-bounds in usb6fire_pcm_in_urb_handler (sound/usb/6fire/pcm.c:338)
Write of size 18446744073709551456 at addr ffff88802a3d0000 by task vhci_rx/5018
Call Trace:
dump_stack_lvl (lib/dump_stack.c:94 lib/dump_stack.c:120)
print_report (mm/kasan/report.c:378 mm/kasan/report.c:482)
kasan_report (mm/kasan/report.c:595)
kasan_check_range (mm/kasan/generic.c:186 mm/kasan/generic.c:200)
__asan_memset (mm/kasan/shadow.c:84)
usb6fire_pcm_in_urb_handler (sound/usb/6fire/pcm.c:338)
__usb_hcd_giveback_urb (drivers/usb/core/hcd.c:1657)
usb_hcd_giveback_urb (drivers/usb/core/hcd.c:1741)
vhci_rx_loop (drivers/usb/usbip/vhci_rx.c:107 drivers/usb/usbip/vhci_rx.c:242)
kthread (kernel/kthread.c:436)
ret_from_fork (arch/x86/kernel/process.c:158)
ret_from_fork_asm (arch/x86/entry/entry_64.S:245)
Allocated by task 10:
__kmalloc_cache_noprof (mm/slub.c:5563)
usb6fire_pcm_init (sound/usb/6fire/pcm.c:560 sound/usb/6fire/pcm.c:595)
usb6fire_chip_probe (sound/usb/6fire/chip.c:133)
usb_probe_interface (drivers/usb/core/driver.c:399)
The buggy address belongs to the object at ffff88802a3d0000
which belongs to the cache kmalloc-8k of size 8192
The buggy address is located 0 bytes inside of
4832-byte region [ffff88802a3d0000, ffff88802a3d12e0)
Kernel panic - not syncing: Fatal exception in interrupt |
| In the Linux kernel, the following vulnerability has been resolved:
drm/msm/dp: skip PUSH_IDLE when the link was never enabled
msm_dp_display_atomic_enable() returns early when link training fails,
leaving ->power_on false and the main link down.
msm_dp_display_atomic_disable() nevertheless writes DP_STATE_CTRL_PUSH_IDLE
and waits for an idle-pattern completion that cannot arrive, so every failed
enable is followed by "PUSH_IDLE pattern timedout".
Every other step of the teardown is already gated on that flag:
msm_dp_display_disable(), called from .atomic_post_disable(), returns early
on !power_on. The PUSH_IDLE write is the only one that is not, so the
controller's runtime-PM reference is then dropped without the link having
been taken down.
On glymur (Snapdragon X2 Elite) the consequence is not a warning. The SoC
does not survive it: TrustZone force-stops the SOCCP and ADSP remote
processors and the machine resets silently about 50 ms later, with no oops
and no panic. On an ASUS Zenbook A16 (UX3607OA), whose eDP panel does not
currently train, this reproduces without any compositor or GPU involvement:
# eDP enable has already failed with "Failed link training (rc=-104)"
echo 1 > /sys/class/graphics/fb0/blank
[535.645455] === marker ===
[535.694833] qcom_q6v5_pas d00000.remoteproc: fatal error received: \
sys_m_smsm.c:512:TZ force stop
[535.694875] remoteproc remoteproc0: crash detected in soccp: type fatal error
[535.728857] qcom_q6v5_pas 6800000.remoteproc: fatal error received: \
sys_m_smsm.c:783:err fatal notification received from TZ
<SoC reset>
Gate the PUSH_IDLE write on ->power_on so the disable path is consistent
with the rest of the teardown. With this applied the same sequence is
harmless and the machine stays up; without it, it resets every time.
The unconditional write dates back to the original DP driver
(c943b4948b58 ("drm/msm/dp: add displayPort driver support")), but the
surrounding code has been restructured several times since, so no Fixes:
tag is offered.
Note that the eDP link-training failure that exposes this on the A16 is a
separate problem in the glymur eDP PHY and is reported separately; this
change is about not damaging the machine when training fails, for whatever
reason.
Tested on ASUS Zenbook A16 (UX3607OA), Snapdragon X2 Elite Extreme, on
linux-next next-20260803 and next-20260807. The machine has since been
running next-20260807 with this patch as its daily driver.
Patchwork: https://patchwork.freedesktop.org/patch/745167/ |