Ubuntu Desktop and PREEMPT_RT kernel issue

Hi,

I am testing Ubuntu Desktop with the RT kernel on a Rubik Pi 3. After interacting with the GUI (e.g. moving the mouse cursor), the system becomes unresponsive - the desktop freezes, network is lost, and drm error messages appear on the debug tty interface (logs below).

Is it a known issue or is there a recommended configuration for running the desktop environment together with the RT kernel?

Environment

Board: Rubik Pi 3
OS: Ubuntu Desktop 24.04.4
Kernel: 6.8.0-1071-qcom-rt
Date tested: March 28, 2026

Kernel cmdline:

/proc/cmdline 
BOOT_IMAGE=/boot/vmlinuz-6.8.0-1071-qcom-rt root=UUID=131450ff-95bc-4791-b611-70855201b0cd ro console=ttyMSM0,115200n8 pcie_pme=nomsi earlycon console=ttyMSM0,115200n8 pcie_pme=nomsi earlycon irqaffinity=0-1,5-7 rcu_nocbs=2-4 isolcpus=2-4 nohz_full=2-4 pcie_aspm=off fsck.mode=force fsck.repair=yes

Kernel logs:

[  278.294313] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[  278.781129] irq 232: nobody cared (try booting with the "irqpoll" option)
[  278.781509] handlers:
[  278.781511] [<00000000eadffe51>] irq_default_primary_handler threaded [<00000000bce6dcae>] msm_irq [msm]
[  278.781668] Disabling IRQ #232
[  289.223808] Error sending AMC RPMH requests (-110)

[  118.085129] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[  128.945330] irq 232: nobody cared (try booting with the "irqpoll" option)
[  128.945693] handlers:
[  128.945695] [<00000000f73c08a6>] irq_default_primary_handler threaded [<00000000ec8fed8d>] msm_irq [msm]
[  128.945852] Disabling IRQ #232
[  129.957141] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[  132.741149] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0: hangcheck detected gpu lockup rb 0!
[  132.741322] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     completed fence: 10220
[  132.741468] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     submitted fence: 10221
[  133.764148] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0: hangcheck detected gpu lockup rb 0!
[  133.764316] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     completed fence: 10220
[  133.764464] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     submitted fence: 10221
[  134.789148] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0: hangcheck detected gpu lockup rb 0!
[  134.789317] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     completed fence: 10220
[  134.789464] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     submitted fence: 10221
[  135.749154] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0: hangcheck detected gpu lockup rb 0!
[  135.749328] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     completed fence: 10220
[  135.749475] msm_dpu ae01000.display-controller: [drm:hangcheck_handler [msm]] *ERROR* 6.3.5.0:     submitted fence: 10221
[  140.970456] msm_dpu ae01000.display-controller: [drm:recover_worker [msm]] *ERROR* 6.3.5.0: hangcheck recover!

We are currently syncing this issue internally.

Could you please share the specific use case or requirements for using the PREEMPT_RT kernel?

Thanks for the reply.
The PREEMPT_RT kernel is required because the system is used as an Ethercat master for servomotor control in a robotic application, so deterministic latency is needed. At the same time DE is used to provide robot UI.

Currently we are using an x86-64 Intel-based platform for that, but we are considering moving to ARM, and performance-wise QCS6490 looks very promising and Rubik Pi3 board has all the interfaces we need.

Hi, please let me know if you managed to reproduce the issue and if there are any updates on it

The latest rt kernel is 6.8.0-1078-qcom-rt
Does the issue still occur on 1078 rt kernel?

Updated to the latest rt kernel (6.8.0-1078-qcom-rt) and the issue is still there. Got this a couple seconds after moving mouse cursor

[  338.377002] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[  338.802766] irq 232: nobody cared (try booting with the "irqpoll" option)
[  338.803159] handlers:
[  338.803161] [<00000000c8865134>] irq_default_primary_handler threaded [<00000000f327da01>] msm_irq [msm]
[  338.803318] Disabling IRQ #232
[  348.398839] Error sending AMC RPMH requests (-110)
[  359.149900] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
[  359.149920] rcu:     0-....: (1 GPs behind) idle=2fac/1/0x4000000000000002 softirq=0/0 fqs=2077 rcuc=5259 jiffies(starved)
[  359.149929] rcu:              hardirqs   softirqs   csw/system
[  359.149932] rcu:      number:     5439          0            0
[  359.149936] rcu:     cputime:     4456          0            0   ==> 10504(ms)
[  359.149941] rcu:     (detected by 2, t=5252 jiffies, g=18681, q=1730 ncpus=8)
[  422.171850] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
[  422.171869] rcu:     0-....: (1 GPs behind) idle=2fac/1/0x4000000000000000 softirq=0/0 fqs=8264 rcuc=21014 jiffies(starved)
[  422.171879] rcu:              hardirqs   softirqs   csw/system
[  422.171881] rcu:      number:    39331          0            0
[  422.171885] rcu:     cputime:    31201          0            0   ==> 73524(ms)
[  422.171890] rcu:     (detected by 2, t=21007 jiffies, g=18681, q=2506 ncpus=8)

I tried on my side, but I cannot replicate your issue.
I can use mouse smoothly on desktop + rt kernel.

Guessing you’re using old version of the software.
Please flash x03 version and try it again:

I just noticed the cmdline you were having: irqaffinity=0-1,5-7 rcu_nocbs=2-4 isolcpus=2-4 nohz_full=2-4
this looks like different from QC’s doc: rcu_nocbs=7 isolcpus=7 irqaffinity=0-6
Where did you get those params?
Can you update to QC’s recommendation, and try whether the issue still exists?

Flashed the latest ubuntu image but the issue persits. Actually we’ve found that setting /dev/cpu_dma_latency to some value lower than let’s say 1000 triggers the issue. Our software was setting it as it’s a common practice for RT systems.

Ok, now we tried disabling it in our app (so /dev/cpu_dma_latency was set to its default) and at first glance it helped, but then we ran glmark2 in background and our Ethercat communication immediately suffered from latency spikes and frame drops. Mouse movements were also causing latency spikes.

It’s just a way to specify which cores should be isolated, I believe QC provided these values as an example. However I’ve tried using QC config to be sure, and it didn’t change the situation a bit.

I don’t think QC’s recommendation is just an example.
Anyway, I’m using that recommended config from QC, and cannot replicate your issue.

I think the “common practice” from redhat is not for arm64 system.
And those kernel settings, including through sysctl, or init script will affect the behavior.
You might have settings somewhere in sysctl, init script, or app start script, which cause the issue.
That’s why I cannot replicate your issue.

You can flash the system, install rt kernel and don’t change any setting or config, and do not install any other software or deb packages. I don’t think there’s mouse issue in this condition.

Regarding cmdline params it could be as you say, and if so they should explicitly say that you are allowed to run only one RT task specifically on cpu core 7. And that sounds a bit strange to me to be honest. Nevertheless I set it as they say in order to exclude it from possible causes.

Regarding /dev/cpu_dma_latency it could be platform dependent, so now I leave it as it is by default.

I’ve found a way to cause system hang that I believe would be easy for you to reproduce. And it appears only on rt kernel.

Run stress-ng --hdd 4 --hdd-bytes 100M -t 2m --metrics-brief
Then perform some GUI activity like opening and using terminal or app center or browser for some time. On my system I get

[ 2454.998178] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[ 2481.174988] [drm:dpu_encoder_frame_done_timeout:2459] [dpu error]enc31 frame done timeout
[ 2481.573685] irq 231: nobody cared (try booting with the "irqpoll" option)
[ 2481.574040] handlers:
[ 2481.574042] [<00000000e34eee1f>] irq_default_primary_handler threaded [<0000000006de429c>] msm_irq [msm]
[ 2481.574202] Disabling IRQ #231
[ 2491.191320] mmc1: Timeout waiting for hardware cmd interrupt.
[ 2491.191326] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========
[ 2491.191329] mmc1: sdhci: Sys addr:  0x00000000 | Version:  0x00007202
[ 2491.191332] mmc1: sdhci: Blk size:  0x00000040 | Blk cnt:  0x00000000
[ 2491.191335] mmc1: sdhci: Argument:  0x92001c10 | Trn mode: 0x00000013
[ 2491.191338] mmc1: sdhci: Present:   0x03d800f0 | Host ctl: 0x0000001f
[ 2491.191340] mmc1: sdhci: Power:     0x0000000f | Blk gap:  0x00000000
[ 2491.191342] mmc1: sdhci: Wake-up:   0x00000000 | Clock:    0x00000007
[ 2491.191344] mmc1: sdhci: Timeout:   0x0000000d | Int stat: 0x00000101
[ 2491.191346] mmc1: sdhci: Int enab:  0x03ff110b | Sig enab: 0x03ff110b
[ 2491.191349] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00000000
[ 2491.191351] mmc1: sdhci: Caps:      0x322dc8b2 | Caps_1:   0x0000808f
[ 2491.191353] mmc1: sdhci: Cmd:       0x0000341a | Max curr: 0x00000000
[ 2491.191355] mmc1: sdhci: Resp[0]:   0x00001010 | Resp[1]:  0x00000000
[ 2491.191357] mmc1: sdhci: Resp[2]:   0x00000000 | Resp[3]:  0x00000000
[ 2491.191359] mmc1: sdhci: Host ctl2: 0x0000000b
[ 2491.191361] mmc1: sdhci: ADMA Err:  0x00000000 | ADMA Ptr: 0x0000000ffffff218
[ 2491.191363] mmc1: sdhci_msm: ----------- VENDOR REGISTER DUMP -----------
[ 2491.191365] mmc1: sdhci_msm: DLL sts: 0x000001dc | DLL cfg:  0x00b7642c | DLL cfg2: 0x0000a800
[ 2491.191367] mmc1: sdhci_msm: DLL cfg3: 0x00000010 | DLL usr ctl:  0x2c010800 | DDR cfg: 0x80040873
[ 2491.191369] mmc1: sdhci_msm: Vndr func: 0x00018a9c | Vndr func2 : 0xf88218a8 Vndr func3: 0x02626040
[ 2491.191371] mmc1: sdhci: ============================================
[ 2491.191437] brcmfmac: brcmf_sdio_htclk: HT Avail request error: -110
[ 2491.964243] Error sending AMC RPMH requests (-110)
[ 2501.434994] mmc1: Timeout waiting for hardware cmd interrupt.
[ 2501.435011] mmc1: sdhci: ============ SDHCI REGISTER DUMP ===========
[ 2501.435016] mmc1: sdhci: Sys addr:  0x00000000 | Version:  0x00007202
[ 2501.435023] mmc1: sdhci: Blk size:  0x00000040 | Blk cnt:  0x00000000
[ 2501.435028] mmc1: sdhci: Argument:  0x92001c10 | Trn mode: 0x00000013
[ 2501.435032] mmc1: sdhci: Present:   0x03d800f0 | Host ctl: 0x0000001f
[ 2501.435036] mmc1: sdhci: Power:     0x0000000f | Blk gap:  0x00000000
[ 2501.435041] mmc1: sdhci: Wake-up:   0x00000000 | Clock:    0x00000007
[ 2501.435044] mmc1: sdhci: Timeout:   0x0000000d | Int stat: 0x00000101
[ 2501.435049] mmc1: sdhci: Int enab:  0x03ff110b | Sig enab: 0x03ff110b
[ 2501.435053] mmc1: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00000000
[ 2501.435057] mmc1: sdhci: Caps:      0x322dc8b2 | Caps_1:   0x0000808f
[ 2501.435061] mmc1: sdhci: Cmd:       0x0000341a | Max curr: 0x00000000
[ 2501.435065] mmc1: sdhci: Resp[0]:   0x00001010 | Resp[1]:  0x00000000
[ 2501.435069] mmc1: sdhci: Resp[2]:   0x00000000 | Resp[3]:  0x00000000
[ 2501.435072] mmc1: sdhci: Host ctl2: 0x0000000b
[ 2501.435077] mmc1: sdhci: ADMA Err:  0x00000000 | ADMA Ptr: 0x0000000000000000
[ 2501.435080] mmc1: sdhci_msm: ----------- VENDOR REGISTER DUMP -----------
[ 2501.435085] mmc1: sdhci_msm: DLL sts: 0x000001dc | DLL cfg:  0x00b7642c | DLL cfg2: 0x0000a800
[ 2501.435090] mmc1: sdhci_msm: DLL cfg3: 0x00000010 | DLL usr ctl:  0x2c010800 | DDR cfg: 0x80040873
[ 2501.435096] mmc1: sdhci_msm: Vndr func: 0x00018a9c | Vndr func2 : 0xf88218a8 Vndr func3: 0x02626040
[ 2501.435099] mmc1: sdhci: ============================================
[ 2501.435174] brcmfmac: brcmf_sdio_htclk: HT Avail request error: -110
[ 2502.043734] rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
[ 2502.043743] rcu:     0-....: (1 GPs behind) idle=8064/1/0x4000000000000000 softirq=0/0 fqs=2625 rcuc=5252 jiffies(starved)
[ 2502.043748] rcu:              hardirqs   softirqs   csw/system
[ 2502.043749] rcu:      number:       63          0            0
[ 2502.043751] rcu:     cputime:     8285          0            0   ==> 10504(ms)
[ 2502.043754] rcu:     (detected by 4, t=5252 jiffies, g=73377, q=8218 ncpus=8)

Ah, I should add that Ubuntu Server is running without issues even with custom cmdline params and /dev/cpu_dma_latency set to 0. And from that perspective it looks like drm/dpu kernel module issue.

Thanks for sharing the test case!
I tried on RB3 Gen2, which also uses QCS6490 chipset. The result is the same as you described.
So I guess it’s common issue for 6490 rt kernel, rather than a “RUBIK Pi 3” specific issue.
I’ll report this to upstream ubuntu forum. Let’s see how it goes.

1 Like

I posted this issue to upstream: Bug #2161035 “Realtime (rt) kernel freeze crash then reboot on h...” : Bugs : Ubuntu on Qualcomm IoT platform
Let’s see how it goes.

1 Like