In-vehicle multimedia information fusion display method
Patent Information
- Application Number
- CN202511747406.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-11-26
AI Technical Summary
该问题不仅破坏了信息显示的时域连续性和视觉稳定性,还可能在复杂驾驶情境下干扰驾驶员的注意分配,从而降低整体行驶安全性与信息融合可靠性
1、本发明通过在车载多媒体信息融合显示系统中引入多模态数据融合感知、认知负荷校正、布局重构动态控制及认知反馈环自平衡机制,实现了信息展示与驾驶员认知状态之间的自适应协调与稳定耦合。通过多模态时间基准对齐与可信度加权的融合分析方法,有效抑制了单模态噪声和感知延迟对认知负荷评估的干扰,显著提升了驾驶员状态识别的实时性与准确性。基于认知负荷水平动态触发布局重构控制函数,能够在不同驾驶情境下自适应调节界面信息密度和优先级,实现信息简化与回补的平衡控制,从而减轻驾驶认知负担并维持操作流畅性,同时,通过对界面调整频率、注视点变化率的持续监测量化布局重构震荡强度,并结合认知负荷变化构建认知反馈环失稳风险模型,对潜在的循环性震荡趋势进行提前识别与干预。
Smart Images

Figure CN121375822B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia information fusion display technology, and more specifically, to a method for in-vehicle multimedia information fusion display. Background Technology
[0002] In existing in-vehicle multimedia information fusion display systems, driver state perception and adaptive interface layout are crucial foundations for achieving intelligent human-machine interaction. Systems typically rely on multi-source features such as physiological signals, gaze hotspot distribution, and voice interaction frequency to estimate the driver's cognitive load level in real time and dynamically adjust the interface display density and information priority based on this estimation. However, when multimodal data contains noise, occlusion, or delay, the cognitive load assessment module is prone to misjudgment. When the system mistakenly believes the driver is under high load, it triggers a layout reconstruction algorithm to simplify the interface content or hide information to alleviate driving stress. However, this layout change directly affects the driver's visual attention distribution and information acquisition path, causing them to frequently shift their gaze and operational focus to compensate for the hidden information. Since these fluctuations in gaze behavior are re-collected by the system and input into the cognitive load model, the model interprets them as attention drift or information overload, further reducing the amount of information displayed. As this chain of "cognitive assessment—layout adjustment—attentional feedback—reassessment" continues to cycle, a typical positive feedback loop forms within the system. This amplifies the cognitive assessment error round by round, ultimately leading to frequent redrawing or flickering of the interface layout within a short period, and periodic oscillations in information density. This cyclical instability caused by the interaction between cognitive load miscalculation and the layout adaptive algorithm essentially reflects a dynamic feedback control defect in the human-machine coupling layer of the in-vehicle information fusion system, namely the "cognitive feedback loop instability" problem. This problem not only disrupts the temporal continuity and visual stability of information display but may also interfere with the driver's attention allocation in complex driving situations, thereby reducing overall driving safety and the reliability of information fusion. Summary of the Invention
[0003] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method for fusion display of in-vehicle multimedia information to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: The method for fusion display of in-vehicle multimedia information includes the following steps: The driver's multimodal data stream is collected in real time by in-vehicle sensing devices, and the asynchronous modes are aligned with the time base to generate a time-consistent multimodal data stream. An embedded lightweight neural network model is used to perform fusion analysis on multimodal data streams to obtain an estimate of the driver's immediate cognitive load. The credibility of each modality is calculated, and the immediate cognitive load estimate is corrected based on the credibility of the modality to obtain the cognitive load level. The layout reconstruction control function is triggered based on the corrected cognitive load level to dynamically adjust the information display strategy. The background system continuously monitors the frequency of interface adjustments and the rate of change of gaze point, calculates the layout reconstruction oscillation intensity coefficient, and identifies the layout reconstruction oscillation risk of the in-vehicle multimedia system. To address the risk of layout reconfiguration oscillations, a cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconfiguration oscillation intensity coefficient, further identifying the cognitive feedback loop instability risk of in-vehicle multimedia systems. To address the risk of cognitive feedback loop instability, a self-balancing strategy is implemented to maintain a stable layout.
[0005] In a preferred embodiment, the step of using an embedded lightweight neural network model to fuse and analyze multimodal data streams to obtain an estimate of the driver's immediate cognitive load is as follows: Each modal data undergoes feature standardization and normalization before being input into the embedded lightweight neural network model; The embedded lightweight neural network model adopts a lightweight multimodal fusion network structure, including a modal feature encoding layer, a modal attention fusion layer, a temporal fusion layer, and an output regression layer; To reduce single-mode anomalies and noise interference, the reliability is calculated for each mode. ,in Uncertainty is a measure of uncertainty. The variance can be calculated from the modal output variance: ,in The variance of the modal output within the current time window. Let be the deep feature vector of the i-th mode. It is the stability constant; The cognitive load level is obtained by correcting the instantaneous cognitive load estimate based on the reliability of the modality. ,in For cognitive load level, This represents the immediate cognitive load estimate, where n is the total number of modalities.
[0006] In a preferred embodiment, a layout reconfiguration control function is triggered based on the corrected cognitive load level: ,in This is the current interface layout state. This information priority matrix defines the importance of different functional modules. This represents the layout adjustment strategy function. This indicates the interface layout state for the next moment.
[0007] In a preferred embodiment, the design logic of the layout adjustment strategy function is as follows: The layout adjustment strategy is divided into three modes based on the cognitive load level range: High-load area: If so, the information simplification strategy will be triggered; Medium load zone: If so, maintain a stable layout strategy; Low load area: If so, the information replenishment strategy will be executed.
[0008] In a preferred embodiment, the information simplification strategy is based on a priority matrix. High-priority modules are retained; low-priority modules are hidden; a smooth fade-out animation and transparency transition mechanism are used to avoid abrupt changes in the interface; and the state of hidden modules is recorded at the same time. The information recovery strategy: when the driver's cognitive load is below a threshold. In this process, some hidden information modules are gradually restored; a delayed echo and spatial smooth transition strategy is used to prevent a sudden increase in information density; and highly interactive and reference modules are restored first. The stable layout strategy is to maintain the current layout and freeze unnecessary animations and layout adjustments.
[0009] In a preferred embodiment, the specific calculation formula for the layout reconstruction oscillation intensity coefficient is as follows: ,in To reconstruct the oscillation intensity coefficient, Adjust the frequency for the interface. The rate of change of fixation point. , These represent the preset proportional coefficients for the interface adjustment frequency and the rate of change of gaze point, respectively. , All are greater than 0.
[0010] In a preferred embodiment, the layout reconstruction oscillation intensity coefficient is compared with a preset layout reconstruction oscillation intensity coefficient threshold. If the layout reconstruction oscillation intensity coefficient is greater than the layout reconstruction oscillation intensity coefficient threshold, it indicates that the in-vehicle multimedia system has a layout reconstruction oscillation risk.
[0011] In a preferred embodiment, the cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconstruction oscillation intensity coefficient: ,in This is a risk index for the instability of the cognitive feedback loop. Cognitive load reference curves generated for historical driving behavior. , These represent the difference between the cognitive load level and the cognitive load reference curve generated from historical driving behavior, and the preset proportional coefficient of the layout reconstruction oscillation intensity coefficient, respectively. , All are greater than 0.
[0012] In a preferred embodiment, the cognitive feedback loop instability risk index is compared with a preset cognitive feedback loop instability risk index threshold. If the cognitive feedback loop instability risk index is greater than the cognitive feedback loop instability risk index threshold, it indicates that the current in-vehicle multimedia system has a cognitive feedback loop instability risk.
[0013] The technical effects and advantages of this invention are as follows: 1. This invention achieves adaptive coordination and stable coupling between information display and driver cognitive state by introducing multimodal data fusion perception, cognitive load correction, dynamic layout reconstruction control, and a cognitive feedback loop self-balancing mechanism into the in-vehicle multimedia information fusion display system. Through a fusion analysis method using multimodal time reference alignment and credibility weighting, the interference of single-modal noise and perception delay on cognitive load assessment is effectively suppressed, significantly improving the real-time performance and accuracy of driver state recognition. Based on a layout reconstruction control function dynamically triggered by cognitive load levels, the interface information density and priority can be adaptively adjusted under different driving scenarios to achieve a balance between information simplification and replenishment, thereby reducing the cognitive burden on the driver and maintaining operational smoothness. Simultaneously, by continuously monitoring the interface adjustment frequency and gaze point change rate, the intensity of layout reconstruction oscillations is quantified, and a cognitive feedback loop instability risk model is constructed in conjunction with cognitive load changes, enabling early identification and intervention of potential cyclical oscillation trends.
[0014] 2. This invention not only improves the temporal continuity of in-vehicle information display and the comfort of human-computer interaction, but also enhances the cognitive stability and information fusion reliability of the system in complex dynamic scenarios, significantly improving the overall driving safety and the consistency of human-computer interaction experience. Attached Figure Description
[0015] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example: Figure 1 The present invention provides a method for fusion and display of in-vehicle multimedia information, comprising the following steps: The vehicle-mounted sensing devices (including cameras, infrared eye trackers, skin conductance sensors, voice acquisition microphones, etc.) are used to collect multimodal data streams from the driver in real time, and the asynchronous modes are aligned with the time reference to generate a time-consistent multimodal data stream. An embedded lightweight neural network model is used to perform fusion analysis on multimodal data streams to obtain an estimate of the driver's immediate cognitive load. The credibility of each modality is calculated, and the immediate cognitive load estimate is corrected based on the credibility of the modality to obtain the cognitive load level. The layout reconstruction control function is triggered based on the corrected cognitive load level to dynamically adjust the information display strategy. The background system continuously monitors the frequency of interface adjustments and the rate of change of gaze point, calculates the layout reconstruction oscillation intensity coefficient, and identifies the layout reconstruction oscillation risk of the in-vehicle multimedia system. To address the risk of layout reconfiguration oscillations, a cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconfiguration oscillation intensity coefficient, further identifying the cognitive feedback loop instability risk of in-vehicle multimedia systems. To address the risk of cognitive feedback loop instability, a self-balancing strategy is implemented to maintain a stable layout.
[0018] The driver's multimodal data stream is collected in real time by in-vehicle sensing devices, and the asynchronous modes are aligned with the time base to generate a time-consistent multimodal data stream. In this embodiment of the invention, the in-vehicle sensing device includes a camera, an infrared eye tracker, a skin conductance sensor band, and a voice acquisition microphone; the multimodal data stream includes visual modalities: driver's facial image, gaze direction, and expression recognition video stream; physiological modalities: heart rate, skin conductance (GSR), eye movement signals, etc.; behavioral modalities: steering wheel angle, pedal pressure, and hand movement data; voice modalities: in-vehicle voice commands and voice emotion characteristics; and environmental modalities: vehicle speed, in-vehicle brightness, and noise level, etc. A unified time server generates a global timestamp during the acquisition of data from each modality, ensuring that different modal data have an alignable time index during the acquisition phase. For any two modes and modality Define its time deviation for: ,in For modality Data collection timestamps For modality Data collection timestamp; if (in If a time synchronization threshold is set for the system (e.g., 50ms), then an asynchronous phenomenon is determined; and a time deviation is detected in real time through a sliding window to form a time deviation sequence. The mode with the highest sampling frequency and the most stable signal (usually the behavioral or visual mode) is used as the time reference mode. For non-reference modes Linear interpolation is performed based on the sampling interval and time deviation to generate a model similar to the reference mode. Synchronization timing: ,in For modality With reference mode Synchronous timing, For modality The timestamp of the kth data collection session. For modality With reference mode Time deviation; After time alignment, all modalities are entered into a unified buffer to generate a time-consistent multimodal data stream. ,in For visual modal data, For physiological purposes, For behavioral modal data, For speech modal data, This is environmental modal data.
[0019] An embedded lightweight neural network model is used to perform fusion analysis on multimodal data streams to obtain an estimate of the driver's immediate cognitive load. The credibility of each modality is calculated, and the immediate cognitive load estimate is corrected based on the credibility of the modality to obtain the cognitive load level. In this embodiment of the invention, an embedded lightweight neural network model is used to fuse and analyze multimodal data streams to obtain an estimate of the driver's real-time cognitive load, as follows: Each modal data undergoes feature standardization and normalization before being input into the embedded lightweight neural network model (common feature standardization and normalization methods include Min-Max normalization, Z-Score normalization, etc.). The embedded lightweight neural network model adopts a lightweight multimodal fusion network (LMFNet) structure, including a modal feature encoder layer: each modality extracts deep feature vectors separately through a lightweight convolutional or GRU encoder layer; the modal feature encoder layer is used for modal feature extraction. ,in Modal data of the i-th mode, For modal feature encoding function, Let be the deep feature vector of the i-th mode; It should be noted that deep feature vectors include visual features (eye movement frequency, blink rate, gaze point distribution, etc.) and physiological features (heart rate, skin conductivity, etc.). Behavioral characteristics (steering wheel micro-movement frequency, pedal adjustment speed, etc.); Speech features (speech rate, intonation changes, etc.); Environmental characteristics (noise level, light intensity, etc.).
[0020] The Attention Fusion Layer is used to calculate modal weight coefficients based on a cross-modal attention mechanism, as follows: ,in Let be the modal weight coefficient for the i-th mode. Let n be the deep feature vector of the j-th mode, and n be the total number of modes; and generate a fused feature vector. : ; Temporal Fusion Layer: Uses gated recurrent units (GRUs) to integrate temporal context information; Output Layer: Outputs an estimate of the driver's immediate cognitive load. ,in This is an estimate of the immediate cognitive load. This is a gated loop unit.
[0021] In this invention, the real-time cognitive load estimate refers to a quantitative indicator of the driver's current psychological and attentional resource consumption level, calculated in real-time using an embedded lightweight neural network model based on multimodal data collected by in-vehicle sensing devices. This estimate reflects the driver's stress level in processing external information at the current moment and is the core input signal for the system's dynamic adjustment of human-computer interaction. The significance of the real-time cognitive load estimate lies in its ability to provide rapid response feedback to the driver's cognitive state, enabling the system to capture changes in the driver's attention on a millisecond timescale, thereby achieving real-time adaptive information display. However, due to the potential for noise interference, occlusion, delay, or modal mismatch during the acquisition of multimodal data, the real-time estimate often contains uncertainties or biases.
[0022] In this embodiment of the invention, to reduce single-mode anomalies or noise interference, the system calculates the reliability for each mode. ,in Uncertainty is a measure of uncertainty. The variance can be calculated from the modal output variance: ,in The variance of the modal output within the current time window is used to characterize the degree of fluctuation. It is a stability constant, usually taken as 0.1 to 1.0, used to prevent the denominator from approaching 0, and to control the upper limit of uncertainty (representing the threshold of modal output fluctuation allowed by the system).
[0023] The cognitive load level is obtained by correcting the instantaneous cognitive load estimate based on the reliability of the modality. ,in Cognitive load level.
[0024] This invention introduces the concept of cognitive load level, which is a stable cognitive assessment result obtained by dynamically weighting and correcting the real-time cognitive load estimate based on the credibility of each modality. Cognitive load level not only reflects the driver's overall cognitive state but also embodies the system's robust suppression capability against anomalous data after multimodal information fusion. Its significance lies in the fact that, through a credibility weighting mechanism, the influence weight of anomalous modalities can be automatically reduced, while the contribution of high-confidence modalities is strengthened, making the cognitive load assessment result more stable, continuous, and interpretable. As a key control variable, cognitive load level is directly used to trigger the layout reconfiguration control function, enabling adaptive adjustment of interface information density, content hierarchy, and interaction form. This ensures driving safety while improving the accuracy of information display and the balance of interactive experience. In other words, the real-time cognitive load estimate is an "instantaneous quantity" that rapidly reflects changes in driving status, while the weighted and corrected cognitive load level is a "stable quantity" used by the system for decision-making and control. Together, they constitute the dynamic basis of the human-machine cognitive coupling closed loop.
[0025] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. The layout reconstruction control function is triggered based on the corrected cognitive load level to dynamically adjust the information display strategy. In this embodiment of the invention, a layout reconfiguration control function is triggered based on the corrected cognitive load level: ,in This refers to the current interface layout state (including module positions, display hierarchy, information density, etc.). This information priority matrix defines the importance of different functional modules. This represents the layout adjustment strategy function. This sets the interface layout for the next moment. Design logic of layout adjustment strategy function: The layout adjustment strategy is divided into three modes based on the cognitive load level range: High-load area: If so, the information simplification strategy will be triggered; Medium load zone: If so, maintain a stable layout strategy; Low load area: If so, the information replenishment strategy will be executed; It should be noted that, The threshold for high cognitive load level The above formulas are all dimensionless calculations for the low cognitive load level threshold. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here.
[0026] The information simplification strategy (when the load is too high): based on the priority matrix. High-priority modules (such as navigation, collision warning, and vehicle speed information) are retained; low-priority modules (such as entertainment information and environmental parameters) are hidden; a smooth fade-out animation and transparency transition mechanism are used to avoid sudden interface changes; at the same time, the state of the hidden modules is recorded so that they can be restored when the load decreases.
[0027] The information recovery strategy (when the load is too low): When the driver's cognitive load is below a threshold... In this process, some hidden information modules are gradually restored; a delayed echo and spatial smooth transition strategy is used to prevent a sudden increase in information density; and highly interactive or reference modules (such as energy consumption statistics and music control) are restored first.
[0028] The stable layout strategy (with cognitive load in a moderate range) is as follows: maintain the current layout stable and freeze unnecessary animations and layout adjustments.
[0029] The background system continuously monitors the frequency of interface adjustments and the rate of change of gaze point, calculates the layout reconstruction oscillation intensity coefficient, and identifies the layout reconstruction oscillation risk of the in-vehicle multimedia system. The specific calculation formula for the layout reconstruction oscillation intensity coefficient in this embodiment of the invention is as follows: ,in To reconstruct the oscillation intensity coefficient, The frequency of interface adjustments can be calculated by comparing the number of interface adjustments within a fixed time period with the fixed time period. The fixation point change rate (can be calculated by comparing the number of fixation point changes within a fixed time period with the fixed time period). , These represent the preset proportional coefficients for the interface adjustment frequency and the rate of change of gaze point, respectively. , All are greater than 0; In this invention, the layout reconstruction oscillation intensity coefficient is a dynamic indicator used to quantify the coupling relationship between the frequency of layout changes in the in-vehicle multimedia interface within a certain time window and the degree of fluctuation in the driver's gaze behavior. This coefficient comprehensively reflects the degree of visual disturbance and information stability imbalance generated by the system during adaptive interface reconstruction. Specifically, when the system frequently adjusts the interface layout according to changes in cognitive load (e.g., information hiding, module movement, or changes in display density), the driver's gaze point often shifts rapidly and non-linearly, forming obvious attention path oscillations. When such oscillation behavior is continuously collected and fed back into the system's cognitive load model, it triggers new layout reconstruction actions, thus forming a cyclically amplified oscillation. The layout reconstruction oscillation intensity coefficient is the core parameter for measuring and identifying this cyclical oscillation process.
[0030] The significance of this coefficient lies primarily in two aspects: First, it provides an objective quantitative standard for measuring the dynamic stability of the interface, enabling real-time monitoring of whether interface adjustments are too frequent or excessive, thus identifying potential "display instability" trends in advance. Second, it can serve as an input to the cognitive feedback loop risk model, working in conjunction with the cognitive load level to construct a comprehensive assessment of the system's dynamic coupling state, achieving closed-loop risk perception from both "human-caused fluctuations" and "interface response" dimensions. In other words, the layout reconstruction oscillation intensity coefficient is not only a core quantitative indicator for identifying the system's operational stability but also a crucial control basis for driving the in-vehicle multimedia display system to achieve self-balancing regulation and cognitive feedback stabilization.
[0031] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. , The settings should be tailored to the specific circumstances. For example, an expert-empowered approach could be adopted, where experts in relevant fields are invited to determine the pre-defined proportions for each indicator through professional opinion surveys and comprehensive evaluations. , The initial value can be 0.5 or 0.5.
[0032] The layout reconstruction oscillation intensity coefficient is compared with the preset layout reconstruction oscillation intensity coefficient threshold. If the layout reconstruction oscillation intensity coefficient is greater than the layout reconstruction oscillation intensity coefficient threshold, it indicates that the in-vehicle multimedia system has a layout reconstruction oscillation risk.
[0033] To address the risk of layout reconfiguration oscillations, a cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconfiguration oscillation intensity coefficient, further identifying the cognitive feedback loop instability risk of in-vehicle multimedia systems. In this embodiment of the invention, a cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconstruction oscillation intensity coefficient: ,in This is a risk index for the instability of the cognitive feedback loop. Cognitive load reference curves generated for historical driving behavior. , These represent the difference between the cognitive load level and the cognitive load reference curve generated from historical driving behavior, and the preset proportional coefficient of the layout reconstruction oscillation intensity coefficient, respectively. , All are greater than 0; In this invention, the cognitive feedback loop instability risk index is a core indicator used to quantify the potential instability of an in-vehicle multimedia system caused by the coupled feedback between dynamic misestimation of cognitive load and interface adaptive layout oscillations during human-machine interaction. This index comprehensively considers the deviation between the current driver's cognitive load level and the historical cognitive load reference curve, as well as the amplification effect of the layout reconstruction oscillation intensity coefficient in the feedback chain, thereby reflecting the overall dynamic coordination between cognitive response and interface control. When the index value is high, it indicates that the current adaptive adjustment speed or amplitude of the system has exceeded the acceptable range of the driver's cognitive state, posing a risk of feedback amplification due to misjudgment or delay; conversely, when the index is low, the system's layout adjustment is well synchronized with the driver's cognitive rhythm, and the operation is within a stable range.
[0034] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. , The settings should be tailored to the specific circumstances. For example, an expert-empowered approach could be adopted, where experts in relevant fields are invited to determine the pre-defined proportions for each indicator through professional opinion surveys and comprehensive evaluations. , The initial value can be 0.5 or 0.5.
[0035] The cognitive feedback loop instability risk index is compared with the preset cognitive feedback loop instability risk index threshold. If the cognitive feedback loop instability risk index is greater than the cognitive feedback loop instability risk index threshold, it indicates that the current in-vehicle multimedia system has a cognitive feedback loop instability risk.
[0036] To address the risk of cognitive feedback loop instability, a self-balancing strategy is implemented to maintain a stable layout.
[0037] In this embodiment of the invention, the self-balancing strategy adjusts the interface refresh rate and animation rate to achieve dynamic frequency control; a progressive information display mechanism is adopted to reduce cognitive abrupt changes; and the neural network model parameter updates are frozen under high-risk conditions to prevent error amplification.
[0038] This invention achieves adaptive coordination and stable coupling between information display and driver cognitive state by introducing multimodal data fusion perception, cognitive load correction, dynamic layout reconstruction control, and a cognitive feedback loop self-balancing mechanism into an in-vehicle multimedia information fusion display system. Through a fusion analysis method using multimodal time reference alignment and credibility weighting, the interference of single-modal noise and perception delay on cognitive load assessment is effectively suppressed, significantly improving the real-time performance and accuracy of driver state recognition. Based on a layout reconstruction control function dynamically triggered by cognitive load levels, the interface information density and priority can be adaptively adjusted under different driving scenarios, achieving a balance between information simplification and replenishment, thereby reducing the cognitive burden on the driver and maintaining operational smoothness. Simultaneously, by continuously monitoring the interface adjustment frequency and gaze point change rate, the intensity of layout reconstruction oscillations is quantified, and a cognitive feedback loop instability risk model is constructed in conjunction with cognitive load changes, enabling early identification and intervention of potential cyclical oscillation trends.
[0039] This invention not only improves the temporal continuity of in-vehicle information display and the comfort of human-computer interaction, but also enhances the cognitive stability and information fusion reliability of the system in complex dynamic scenarios, significantly improving overall driving safety and the consistency of human-computer interaction experience.
[0040] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0041] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0042] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for fusion and display of in-vehicle multimedia information, characterized in that: Includes the following steps: The driver's multimodal data stream is collected in real time by in-vehicle sensing devices, and the asynchronous modes are aligned with the time base to generate a time-consistent multimodal data stream. An embedded lightweight neural network model is used to perform fusion analysis on multimodal data streams to obtain an estimate of the driver's immediate cognitive load. The credibility of each modality is calculated, and the immediate cognitive load estimate is corrected based on the credibility of the modality to obtain the cognitive load level. The layout reconstruction control function is triggered based on the corrected cognitive load level to dynamically adjust the information display strategy. The background system continuously monitors the frequency of interface adjustments and the rate of change of gaze point, calculates the layout reconstruction oscillation intensity coefficient, and identifies the layout reconstruction oscillation risk of the in-vehicle multimedia system. To address the risk of layout reconfiguration oscillations, a cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconfiguration oscillation intensity coefficient, further identifying the cognitive feedback loop instability risk of in-vehicle multimedia systems. To address the risk of cognitive feedback loop instability, a self-balancing strategy is implemented to maintain a stable layout.
2. The in-vehicle multimedia information fusion display method according to claim 1, characterized in that: The method utilizes an embedded lightweight neural network model to fuse and analyze multimodal data streams, obtaining an estimate of the driver's real-time cognitive load, as detailed below: Each modal data undergoes feature standardization and normalization before being input into the embedded lightweight neural network model; The embedded lightweight neural network model adopts a lightweight multimodal fusion network structure, including a modal feature encoding layer, a modal attention fusion layer, a temporal fusion layer, and an output regression layer; To reduce single-mode anomalies and noise interference, the reliability is calculated for each mode. ,in Uncertainty is a measure of uncertainty. The variance can be calculated from the modal output variance: ,in The variance of the modal output within the current time window. Let be the deep feature vector of the i-th mode. It is the stability constant; The cognitive load level is obtained by correcting the instantaneous cognitive load estimate based on the reliability of the modality. ,in For cognitive load level, This represents the immediate cognitive load estimate, where n is the total number of modalities.
3. The in-vehicle multimedia information fusion display method according to claim 2, characterized in that: The layout reconfiguration control function is triggered based on the corrected cognitive load level: ,in This is the current interface layout state. This information priority matrix defines the importance of different functional modules. This represents the layout adjustment strategy function. This indicates the interface layout state for the next moment.
4. The in-vehicle multimedia information fusion display method according to claim 3, characterized in that: Design logic of layout adjustment strategy function: The layout adjustment strategy is divided into three modes based on the cognitive load level range: High-load area: If so, the information simplification strategy will be triggered; Medium load zone: If so, maintain a stable layout strategy; Low load area: If so, the information replenishment strategy will be executed.
5. The in-vehicle multimedia information fusion display method according to claim 4, characterized in that: The information simplification strategy is based on the priority matrix. High-priority modules are retained; low-priority modules are hidden; a smooth fade-out animation and transparency transition mechanism are used to avoid abrupt changes in the interface; and the state of hidden modules is recorded at the same time. The information recovery strategy: when the driver's cognitive load is below a threshold. In this process, some hidden information modules are gradually restored; a delayed echo and spatial smooth transition strategy is used to prevent a sudden increase in information density. Prioritize restoring modules with high interactivity and those that provide reference information; The stable layout strategy is to maintain the current layout and freeze unnecessary animations and layout adjustments.
6. The in-vehicle multimedia information fusion display method according to claim 5, characterized in that: The specific formula for calculating the layout reconstruction oscillation intensity coefficient is as follows: ,in To reconstruct the oscillation intensity coefficient, Adjust the frequency for the interface. The rate of change of fixation point. , These represent the preset proportional coefficients for the interface adjustment frequency and the rate of change of gaze point, respectively. , All are greater than 0.
7. The in-vehicle multimedia information fusion display method according to claim 6, characterized in that: The layout reconstruction oscillation intensity coefficient is compared with the preset layout reconstruction oscillation intensity coefficient threshold. If the layout reconstruction oscillation intensity coefficient is greater than the layout reconstruction oscillation intensity coefficient threshold, it indicates that the in-vehicle multimedia system has a layout reconstruction oscillation risk.
8. The in-vehicle multimedia information fusion display method according to claim 7, characterized in that: The cognitive feedback loop instability risk model is constructed based on the cognitive load level and the layout reconstruction oscillation intensity coefficient. ,in This is a risk index for the instability of the cognitive feedback loop. Cognitive load reference curves generated for historical driving behavior. , These represent the difference between the cognitive load level and the cognitive load reference curve generated from historical driving behavior, and the preset proportional coefficient of the layout reconstruction oscillation intensity coefficient, respectively. , All are greater than 0.
9. The in-vehicle multimedia information fusion display method according to claim 8, characterized in that: The cognitive feedback loop instability risk index is compared with the preset cognitive feedback loop instability risk index threshold. If the cognitive feedback loop instability risk index is greater than the cognitive feedback loop instability risk index threshold, it indicates that the current in-vehicle multimedia system has a cognitive feedback loop instability risk.
Citation Information
Patent Citations
Measuring cognitive load
US20100217097A1
System for task and notification handling in a connected car
US20130038437A1