Hmi adaptive rendering method, device and electronic equipment

CN122715221APending Publication Date: 2026-09-08CHINA FAW CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610888326.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

这种切换缺乏中间态的平滑过渡,导致界面突兀变化,容易打断驾驶员的思维流,造成认知摩擦,影响驾驶安全

Benefits of technology

[0042]This invention aims to achieve the effect of generating HMI interfaces that match the driver's current psychological state in real time using generative AI technology, while strictly protecting the user's biometric privacy. This results in an extremely personalized experience that is different for each person and each time. Through generative AI models, the system is no longer limited to a few preset UI images, but instead "draws" the interface in real time based on the driver's real-time state (stress, fatigue).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122715221A_ABST
    Figure CN122715221A_ABST
Patent Text Reader

Abstract

This invention provides an HMI adaptive rendering method, apparatus, and electronic device, relating to the field of intelligent cockpit technology. The method includes: real-time acquisition of a driver's facial video stream and voice stream; face detection and feature extraction of the original video frames using a lightweight convolutional neural network to generate a cognitive state vector; inputting the cognitive state vector into a pre-trained diffusion model to determine the interface rendering mode; and performing real-time rendering of the HMI interface according to the determined interface rendering mode. The HMI adaptive rendering method, apparatus, and electronic device provided by this invention enable the realization of complex generative visual effects even on automotive-grade chips with limited computing power, reducing the threshold and cost of technology implementation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cockpit technology, and more particularly to HMI adaptive rendering methods, devices, and electronic devices. Background Technology

[0002] As the core of human-vehicle interaction, the human-machine interface (HMI) technology of the intelligent cockpit is evolving from static preset to dynamic adaptive. Current mainstream solutions can be divided into two categories: static preset HMI and rule-based adaptive HMI.

[0003] Static preset HMI is the most widely adopted solution in the market. During the design phase, the development team pre-creates several fixed UI themes (such as "Sport Mode", "Comfort Mode" and "Night Mode"). The system switches between them simply based on the vehicle's driving mode or time. Its hardware foundation is the cockpit domain controller (SoC) driving the display screen to render images.

[0004] Some high-end models have introduced simple adaptive logic. For example, when the vehicle speed exceeds a threshold, the central control screen will hide some entertainment functions; or when the driver monitoring system (DMS) detects that the driver has closed their eyes, the system will issue a warning.

[0005] There are still three major pain points in the existing technology.

[0006] First, the interactive experience is rigid and lacks context awareness.

[0007] Existing HMIs cannot deeply perceive the driver's real-time psychological state (such as anxiety, fatigue, and excitement). For example, when the driver is in a state of high-pressure anxiety, the system still displays a brightly colored and information-heavy interface, which not only fails to relieve stress but also increases cognitive load and leads to distraction. This "one-size-fits-all" experience is far from the level that an intelligent cockpit should be.

[0008] Next, there is a sharp contradiction between privacy protection and artificial intelligence.

[0009] To achieve more refined self-adaptation, the system needs to collect highly sensitive and privacy-sensitive data such as high-definition facial videos, eye-tracking trajectories, and even voiceprints. Existing solutions face a dilemma: uploading raw data to the cloud for analysis poses a serious risk of privacy breaches; while complete localization to protect privacy is limited by the computing power of automotive-grade chips, making it difficult to run complex generative AI models to render high-quality dynamic interfaces in real time. This contradiction has become a major bottleneck restricting the development of intelligent cockpits.

[0010] In addition, the response mechanism is slow and rigid.

[0011] Rule-based systems typically employ black-and-white, on / off switching; for example, detecting driver fatigue might trigger music playback. This lack of smooth transitions between intermediate states leads to abrupt interface changes, easily disrupting the driver's thought process, causing cognitive friction, and ultimately impacting driving safety. Summary of the Invention

[0012] The purpose of this invention is to provide an HMI adaptive rendering method, device, electronic device, and storage medium that enables complex generative visual effects to be achieved even on automotive-grade chips with limited computing power, thereby reducing the threshold and cost of technology implementation.

[0013] This invention provides the following solution:

[0014] According to one aspect of the present invention, an HMI adaptive rendering method is provided, the HMI adaptive rendering method comprising:

[0015] Real-time acquisition of the driver's facial video and voice streams;

[0016] A lightweight convolutional neural network is used to perform face detection and feature extraction on the original video frames to generate a cognitive state vector.

[0017] The cognitive state vector is input into a pre-trained diffusion model to determine the interface rendering mode;

[0018] Perform real-time rendering of the HMI interface according to the determined interface rendering mode.

[0019] Optionally, real-time acquisition of the driver's facial video stream and voice stream, including:

[0020] The driver's facial video stream and voice stream are collected in real time at a predetermined frequency and then input into the vehicle-mounted neural network processing unit.

[0021] Optionally, a lightweight convolutional neural network is deployed on the neural network processing unit.

[0022] Optionally, the lightweight convolutional neural network is used to perform face detection and feature extraction on the original video frames to generate a cognitive state vector, including:

[0023] Face detection and a preset number of key points are located on the original video frames. The original pixel data is discarded immediately after feature extraction is completed, without any form of caching or uploading.

[0024] Information such as facial micro-expressions, eyelid closure degree, and speech fundamental frequency is mapped into a cognitive state vector.

[0025] Optional facial micro-expressions include frowning and downturned corners of the mouth.

[0026] Optionally, the cognitive state vector is input into a pre-trained diffusion model to determine the interface rendering mode, including:

[0027] The cognitive state vector is used as a conditional control parameter and input into the diffusion model;

[0028] If the pressure level in the conditional control parameters is greater than the preset pressure level threshold, the noise reduction mode is activated.

[0029] If the distractibility parameter in the conditional control is less than the preset distractibility threshold, the focus mode is activated.

[0030] Optionally, in the noise reduction mode, color saturation is reduced and information hierarchy is simplified.

[0031] Optionally, in the focus mode, key driving information can be magnified.

[0032] According to a second aspect of the present invention, an HMI adaptive rendering apparatus is provided, which performs HMI adaptive rendering using the method described above, the HMI adaptive rendering apparatus comprising:

[0033] The acquisition module is used to acquire the driver's facial video stream and voice stream in real time;

[0034] The extraction module is used to perform face detection and feature extraction on the original video frames using a lightweight convolutional neural network, and generate a cognitive state vector.

[0035] The determination module is used to input the cognitive state vector into the pre-trained diffusion model to determine the interface rendering mode;

[0036] The rendering module is used to perform real-time rendering of the HMI interface according to a defined interface rendering mode.

[0037] According to three aspects of the present invention, an electronic device is provided, the electronic device comprising:

[0038] Processor, communication interface, memory, and communication bus.

[0039] The processor, communication interface, and memory communicate with each other via a communication bus.

[0040] The memory stores a computer program that, when executed by the processor, enables the processor to perform the steps of the HMI adaptive rendering method described above.

[0041] The above solution achieves the following beneficial technical effects:

[0042] This invention aims to achieve the effect of generating HMI interfaces that match the driver's current psychological state in real time using generative AI technology, while strictly protecting the user's biometric privacy. This results in an extremely personalized experience that is different for each person and each time. Through generative AI models, the system is no longer limited to a few preset UI images, but instead "draws" the interface in real time based on the driver's real-time state (stress, fatigue).

[0043] This invention utilizes a "dynamic privacy grading" mechanism to directly convert highly sensitive raw video streams into low-sensitivity "state vectors" on the vehicle's NPU. Only the desensitized vector data participates in subsequent rendering logic, ensuring that the original facial data never leaves the local machine, thus achieving "data not leaving the vehicle," mitigating privacy leakage risks, and meeting increasingly stringent global data compliance requirements. By decoupling complex physiological feature extraction from interface generation, and using a lightweight encoder to extract features on the vehicle, the cloud or high-performance GPU is only responsible for rendering and generation based on feature vectors. This enables complex generative visual effects to be achieved even on automotive-grade chips with limited computing power, lowering the barriers and costs of technology implementation. Attached Figure Description

[0044] Figure 1 This is a flowchart of an HMI adaptive rendering method provided in one or more embodiments of the present invention;

[0045] Figure 2 This is a flowchart of an HMI adaptive rendering method provided in one or more embodiments of the present invention;

[0046] Figure 3 This is a flowchart of the operation extraction method provided in one or more embodiments of the HMI adaptive rendering method of the present invention;

[0047] Figure 4 This is a flowchart of the determination operation in the HMI adaptive rendering method provided by one or more embodiments of the present invention;

[0048] Figure 5 This is a structural diagram of an HMI adaptive rendering apparatus provided in one or more embodiments of the present invention;

[0049] Figure 6 This is a structural diagram of an electronic device provided in one or more embodiments of the present invention. Detailed Implementation

[0050] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Figure 1 This is a flowchart of an HMI adaptive rendering method provided in one or more embodiments of the present invention. See also... Figure 1 The HMI adaptive rendering method includes the following steps:

[0052] S11 captures the driver's facial video stream and voice stream in real time.

[0053] S12 uses a lightweight convolutional neural network to perform face detection and feature extraction on the original video frames, generating a cognitive state vector.

[0054] S13, input the cognitive state vector into the pre-trained diffusion model to determine the interface rendering mode.

[0055] S14, execute real-time rendering of the HMI interface according to the determined interface rendering mode.

[0056] As described above in this invention, existing interface rendering solutions often require uploading facial images of people inside the cabin to the cloud.

[0057] Such operations would be suspected of infringing on user privacy in practical applications, and therefore their application has encountered great resistance.

[0058] The technical solution provided by this invention can ensure that the user's original facial image is extracted inside the vehicle cabin, ensuring that sensitive data does not leave the vehicle cabin and will not be accessed by any third-party data platform.

[0059] Moreover, this embodiment enables complex generative visual effects to be achieved even on automotive-grade chips with limited computing power, reducing the threshold and cost of technology implementation.

[0060] Specifically, the system first uses cameras and microphone arrays in the cockpit to capture facial video and voice streams of the driver.

[0061] It is important to note that the video and audio stream acquisition process described here is a synchronous acquisition process. That is to say, the acquired video and audio streams are strictly synchronized in time.

[0062] After completing the multimodal data acquisition described above, the acquired video and audio streams are input into a pre-set convolutional neural network model.

[0063] Using this pre-built convolutional neural network model, it is possible to extract facial features and voiceprint features of the driver.

[0064] After processing by the convolutional neural network model, the model outputs a 128-dimensional cognitive state vector. This vector contains feature data extracted from the video and audio streams. However, it is impossible to completely reproduce the original video and audio streams from these extracted feature data alone.

[0065] In other words, after processing by the above convolutional neural network model, the originally sensitive video and audio streams can no longer be reproduced, thus achieving the desensitization of sensitive data.

[0066] After extracting the aforementioned feature data and generating a cognitive state vector that meets the requirements, the generated cognitive state vector is input into a diffusion model.

[0067] The diffusion model defines a forward process that gradually diffuses data into random noise, while the reverse process iteratively recovers the data from the random noise. To date, it has achieved great success in fields such as image generation, video, and audio.

[0068] By inputting the cognitive state vector into the diffusion model, the rendering mode of HMI rendering can be obtained through further reasoning and calculation of the cognitive state vector.

[0069] Typically, in this embodiment, the HMI rendering mode can be either a noise reduction mode or a focus mode.

[0070] In noise reduction mode, color saturation is reduced and information hierarchy is simplified during the interface rendering process.

[0071] In focus mode, key driving information will be magnified during the interface rendering process. For example, the speed font can be increased by 50%, and the navigation arrow can be highlighted.

[0072] Through the above processing, the HMI interface rendering can present different visual styles according to the driver's psychological state in the cockpit, so that the rendering effect is suitable for the driver's current psychological state.

[0073] Moreover, the above processing does not involve a large amount of computation, which can be well adapted to the computing resources available in the cockpit. This allows complex generative visual effects to be achieved even in a cockpit with limited computing resources, reducing the threshold and cost of technology implementation.

[0074] Figure 2 This is a flowchart of an HMI adaptive rendering method provided in one or more embodiments of the present invention. See also... Figure 2 The HMI adaptive rendering method includes the following steps:

[0075] S21 collects the driver's facial video stream and voice stream in real time at a predetermined frequency and inputs them into the vehicle-mounted neural network processing unit.

[0076] S22 uses a lightweight convolutional neural network to perform face detection and feature extraction on the original video frames, generating a cognitive state vector.

[0077] S23, input the cognitive state vector into the pre-trained diffusion model to determine the interface rendering mode.

[0078] S24, execute real-time rendering of the HMI interface according to the determined interface rendering mode.

[0079] This embodiment is based on the foregoing embodiments of the present invention and further provides the execution process of the HMI adaptive rendering method.

[0080] In this embodiment, the driver's facial video stream and speech stream are captured by a microphone array equipped in the cockpit and an in-vehicle camera.

[0081] The video and audio streams are captured synchronously. This means that facial images and spoken audio at the same moment are strictly aligned on the timeline.

[0082] Furthermore, the video images and spoken audio within the cockpit are captured at a fixed frequency. Specifically, in this embodiment, both video images and spoken audio are captured at a frequency of 30fps.

[0083] After the facial video images and spoken voice are collected, they are input into the vehicle's NPU.

[0084] The vehicle-mounted NPU has the following characteristics: it is loaded with a convolutional neural network, and it can process the input signal based on the convolutional neural network loaded on it.

[0085] Figure 3 This is a flowchart illustrating the extraction operations in the HMI adaptive rendering method provided by one or more embodiments of the present invention. See also... Figure 3 A lightweight convolutional neural network is used to perform face detection and feature extraction on the original video frames to generate a cognitive state vector. The process includes the following steps:

[0086] S31 performs face detection and localization of a preset number of key points on the original video frames. The original pixel data is discarded immediately after feature extraction is completed, without any form of caching or uploading.

[0087] S32 maps information such as facial micro-expressions, eyelid closure degree, and speech fundamental frequency into a cognitive state vector.

[0088] This embodiment is based on the foregoing embodiments of the present invention and provides a further detailed description of the cognitive state vector generation process.

[0089] The process of generating cognitive state vectors can be more specifically divided into two sub-steps: feature extraction and cognitive state vector generation.

[0090] The feature extraction process is performed using a convolutional neural network built into the system. Specifically, the convolutional neural network is integrated into an NPU processor.

[0091] More specifically, after the original video frames are fed into the convolutional neural network frame by frame, the convolutional neural network will perform face detection on the input video frames.

[0092] Secondly, the convolutional neural network also needs to locate 68 key points. These key points are typically key points on a human face image. After extracting the location data of these key points, further calculations are performed on this location data to obtain the key parameters in the cognitive state vector.

[0093] Furthermore, video frames processed by the convolutional neural network are immediately discarded, without further caching or uploading of the original video frames. Because these frames are discarded immediately, no leakage of sensitive user information is possible. Moreover, since the processing result of the convolutional neural network—the cognitive state vector—cannot be used to reconstruct the user's facial image, this image processing via the convolutional neural network achieves one-step desensitization of sensitive data.

[0094] After completing the extraction of the above feature data, the extracted key information, such as facial micro-expressions, eyelid closure degree, and speech fundamental frequency, is further mapped to complete the generation of the cognitive state vector.

[0095] Figure 4 This is a flowchart illustrating the determination operation in the HMI adaptive rendering method provided by one or more embodiments of the present invention. See also... Figure 4 The cognitive state vector is input into a pre-trained diffusion model to determine the interface rendering mode, including the following steps:

[0096] S41, the cognitive state vector is used as a conditional control parameter and input into the diffusion model.

[0097] S42, if the pressure level in the condition control parameters is greater than the preset pressure level threshold, activate the noise reduction mode.

[0098] S43, if the distraction in the conditional control parameters is less than the preset distraction threshold, activate the focus mode.

[0099] In addition to the convolutional neural network mentioned earlier, the adaptive rendering process provided in this invention also uses another model: the diffusion model.

[0100] In the foregoing embodiments of the present invention, the convolutional neural network is used to extract corresponding feature data from the input video stream and speech stream, and to generate a cognitive state vector based on the extracted feature data.

[0101] In this embodiment, the diffusion model serves to take the input cognitive state vector as a conditional control parameter and determine the rendering mode to be started when the next rendering operation is performed based on the specific parameter values ​​in the cognitive state vector.

[0102] Specifically, if the pressure level in the condition control parameters is greater than the preset pressure level threshold, then the noise reduction mode needs to be activated during the rendering process.

[0103] In noise reduction mode, color saturation needs to be reduced and information hierarchy simplified. This display rendering process significantly reduces noise in the rendered interface. For example, the color saturation of the rendered interface can be reduced to below 30%. Another example is that only driving parameters such as vehicle speed and battery level can be displayed on the actual interface, while other parameters are removed.

[0104] If the distractibility parameter in the conditional control is less than the preset distractibility threshold, then the focus mode needs to be activated.

[0105] In focus mode, key driving information needs to be amplified. Amplifying key driving information allows the driver to focus on driving itself, while reducing attention to other information irrelevant to driving.

[0106] Figure 5 This is a structural diagram of an HMI adaptive rendering apparatus provided in one or more embodiments of the present invention. The HMI adaptive rendering apparatus performs HMI adaptive rendering using the method described above. See also... Figure 5 The HMI adaptive rendering appliance includes:

[0107] The acquisition module 51 is used to acquire the driver's facial video stream and voice stream in real time.

[0108] The extraction module 52 is used to perform face detection and feature extraction on the original video frames through a lightweight convolutional neural network to generate a cognitive state vector.

[0109] The determination module 53 is used to input the cognitive state vector into the pre-trained diffusion model to determine the interface rendering mode.

[0110] The rendering module 54 is used to perform real-time rendering of the HMI interface according to the determined interface rendering mode.

[0111] It is worth noting that although only some basic functional modules are disclosed in the embodiments of this invention, it does not mean that the composition of this system is limited to the above-mentioned basic functional modules. On the contrary, what this embodiment intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can arbitrarily add one or more functional modules in combination with existing technology to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. The fact that this embodiment only discloses a few basic functional modules should not be considered as the scope of protection of the claims of this invention being limited to the disclosed basic functional modules. At the same time, for the convenience of description, the above device is described separately according to its functions as various units and modules. Of course, in implementing this invention, the functions of each unit and module can be implemented in one or more software and / or hardware.

[0112] like Figure 6 As shown, the present invention also provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the HMI adaptive rendering method.

[0113] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 6 The structure shown in this embodiment of the invention includes an electronic device comprising one or more processors 610 and a memory 620; the processors 610 in this electronic device may be one or more. Figure 6 Taking a processor 610 as an example; memory 620 is used to store one or more programs; the one or more programs are executed by the one or more processors 610, so that the one or more processors 610 implement the HMI adaptive rendering method as described in any one of the embodiments of the present invention.

[0114] The electronic device may also include an input device 630 and an output device 640.

[0115] The processor 610, memory 620, input device 630, and output device 640 in this electronic device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0116] The memory 620 in this electronic device serves as a computer-readable storage medium, capable of storing one or more programs. These programs can be software programs, computer-executable programs, or modules, such as the program instructions / modules corresponding to the HMI adaptive rendering method provided in this embodiment. The processor 610 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 620, thereby implementing the HMI adaptive rendering method described in the above embodiment.

[0117] Memory 620 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, memory 620 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, memory 620 may further include memory remotely located relative to processor 610, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0118] Input device 630 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 640 may include display devices such as a display screen.

[0119] The present invention also provides a computer-readable storage medium, comprising: storing a computer program executable by a vehicle, wherein when the computer program is run on the vehicle, the vehicle performs the steps of the HMI adaptive rendering method.

[0120] Specifically, the computer storage medium in this embodiment of the invention can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be—but is not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for adaptive rendering of HMI, characterized in that, The HMI adaptive rendering method includes: Real-time acquisition of the driver's facial video and voice streams; A lightweight convolutional neural network is used to perform face detection and feature extraction on the original video frames to generate a cognitive state vector. The cognitive state vector is input into a pre-trained diffusion model to determine the interface rendering mode; Perform real-time rendering of the HMI interface according to the determined interface rendering mode.

2. The method of claim 1, wherein, Real-time acquisition of the driver's facial video stream and voice stream, including: The driver's facial video stream and voice stream are collected in real time at a predetermined frequency and then input into the vehicle-mounted neural network processing unit.

3. The method of claim 2, wherein, The neural network processing unit is equipped with a lightweight convolutional neural network.

4. The method of claim 1, wherein, The lightweight convolutional neural network is used to perform face detection and feature extraction on the original video frames, generating a cognitive state vector, including: Face detection and a preset number of key points are located on the original video frames. The original pixel data is discarded immediately after feature extraction is completed, without any form of caching or uploading. Information such as facial micro-expressions, eyelid closure degree, and speech fundamental frequency is mapped into a cognitive state vector.

5. The method of claim 4, wherein, Facial micro-expressions include frowning and downturned corners of the mouth.

6. The method of claim 1, wherein, The cognitive state vector is input into a pre-trained diffusion model to determine the interface rendering mode, including: The cognitive state vector is used as a conditional control parameter and input into the diffusion model; If the pressure level in the conditional control parameters is greater than the preset pressure level threshold, the noise reduction mode is activated. If the distractibility parameter in the conditional control is less than the preset distractibility threshold, the focus mode is activated.

7. The method of claim 6, wherein, In the noise reduction mode, color saturation is reduced and information hierarchy is simplified.

8. The method according to claim 6, characterized in that, In the focus mode, key driving information is magnified.

9. An HMI adaptive rendering device, characterized in that, HMI adaptive rendering is performed using the method described in any one of claims 1-8, wherein the HMI adaptive rendering apparatus comprises: The acquisition module is used to acquire the driver's facial video stream and voice stream in real time; The extraction module is used to perform face detection and feature extraction on the original video frames using a lightweight convolutional neural network, and generate a cognitive state vector. The determination module is used to input the cognitive state vector into the pre-trained diffusion model to determine the interface rendering mode; The rendering module is used to perform real-time rendering of the HMI interface according to a defined interface rendering mode.

10. An electronic device, characterized in that, The electronic device includes: Processor, communication interface, memory, and communication bus. The processor, communication interface, and memory communicate with each other via a communication bus. The memory stores a computer program that, when executed by the processor, enables the processor to perform the steps of the HMI adaptive rendering method according to any one of claims 1 to 8.