An intelligent cockpit interaction method and system, a vehicle and a storage medium

By integrating multimodal arbitration of voice, gestures, and bioelectrical signals into the intelligent cockpit system, and dynamically adjusting the priority of command execution, the system solves the problems of low recognition accuracy and insufficient personalized response in existing multimodal interaction schemes, and achieves an efficient and safe intelligent interaction experience.

CN122634499APending Publication Date: 2026-08-25CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610800613.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing intelligent cockpit interaction solutions lack dynamic adjustment in handling multimodal command conflicts, resulting in low recognition accuracy, inability to provide personalized responses based on user habits and driving scenarios, and failure to consider the driver's physiological state and vehicle operating conditions, leading to low interaction efficiency.

Method used

By acquiring users' voice, gestures, and bioelectric signals, calculating confidence levels, and performing weighted fusion, the system dynamically adjusts the execution priority of instructions based on the current interaction scenario, establishes a multimodal arbitration mechanism, and achieves personalized responses and matching of security requirements.

Benefits of technology

It improves the accuracy of command recognition in complex driving environments, reduces the risk of single-modal misrecognition and command conflict, enhances the intelligence of cockpit interaction and driving comfort, and ensures driving safety and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122634499A_ABST
    Figure CN122634499A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent cockpits, and discloses an intelligent cockpit interaction method, system, vehicle and storage medium, the method comprising the following steps: acquiring voice signals, gesture signals and bioelectric signals of a user, and respectively calculating voice confidence, gesture confidence and bioelectric confidence; performing weighted calculation on the voice confidence, gesture confidence and bioelectric confidence to obtain an arbitration score; determining a to-be-executed instruction based on the arbitration score; determining a current interaction scene, and correcting the instruction execution priority of the to-be-executed instruction based on the current interaction scene to obtain a final control instruction; and delivering the final control instruction to a device corresponding to the cockpit for execution. Through the multi-modal interaction arbitration mechanism of fusing voice, gesture and bioelectric signals, and dynamically correcting the instruction priority in combination with the current interaction scene, the application can improve interaction accuracy and reliability, not only realizes scene-adaptive personalized response, but also improves driving comfort and user satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart cockpit technology, specifically to a smart cockpit interaction method, system, vehicle, and storage medium. Background Technology

[0002] With the continuous improvement of the intelligence level of new energy vehicles, intelligent cockpits have become the core carrier of human-vehicle interaction. Currently, mainstream intelligent cockpit systems generally integrate multiple interaction methods such as voice control, touch control, and gesture recognition, aiming to provide drivers with a more convenient and natural operating experience. For example, functions such as adjusting the air conditioning via voice commands, controlling navigation via touchscreen, and switching music via gestures have gradually become common configurations.

[0003] However, existing intelligent cockpit interaction solutions have many shortcomings in practical applications. For example, although existing solutions integrate multiple modalities such as voice, touch, and gestures, the handling of command conflicts between these modalities is relatively simple, such as using preset priorities to resolve multiple command conflicts. This type of method cannot be dynamically adjusted according to user habits and driving scenarios, resulting in limited command recognition accuracy. In addition, existing solutions mainly rely on external signals such as voice and gestures, failing to consider the driver's own physiological state and the actual operating conditions of the vehicle, resulting in poor interaction efficiency and difficulty in achieving truly personalized intelligent interaction.

[0004] In summary, there is an urgent need for an intelligent interaction solution based on multiple scenarios such as external environment, user status, and vehicle operation, in order to improve cockpit interaction efficiency and user experience. Summary of the Invention

[0005] This invention provides an intelligent cockpit interaction method, system, vehicle, and storage medium to solve the problem of low interaction efficiency in existing intelligent cockpits, which seriously restricts the level of intelligent interaction and user experience, as mentioned in the above-mentioned technical background.

[0006] In a first aspect, the present invention provides an intelligent cockpit interaction method, the method comprising: Acquire the user's voice signal, gesture signal, and bioelectrical signal, and calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively; The arbitration score is obtained by weighting the confidence scores of speech, gesture, and bioelectricity. The instructions to be executed are determined based on the arbitration score; Determine the current interaction scenario, and based on the current interaction scenario, adjust the execution priority of the instruction to be executed to obtain the final control instruction; The final control command is issued to the corresponding equipment in the cockpit for execution.

[0007] This invention calculates and weights the confidence levels of three modalities—voice, gesture, and bioelectricity—to comprehensively determine user intent. This effectively reduces the risk of erroneous execution caused by misidentification of a single modality or command conflicts, and improves the accuracy of command recognition in complex driving environments. Simultaneously, based on the real-time interaction scenario, such as rain, low battery, or fatigued driving, the execution priority of commands to be executed is dynamically adjusted. This allows the cockpit control strategy to meet the safety needs and user expectations of different scenarios, avoiding the rigidity of commands caused by fixed priority modes. This not only achieves personalized responses that adapt to different scenarios and improves the intelligence of cockpit interaction, but also enhances driving comfort and user satisfaction.

[0008] In one optional implementation, before issuing the final control command to the corresponding device in the cockpit for execution, the smart cockpit interaction method further includes: The remaining battery power, ambient temperature, and user comfort index are obtained. The user comfort index is calculated by fusing at least one parameter among seat pressure distribution, skin temperature and humidity, and breathing rate. Based on the remaining battery power, ambient temperature, and user comfort index, the optimal operating power of each device in the cabin is calculated. After correcting the equipment power parameters in the final control command based on the optimal operating power, the final control command is issued to the corresponding equipment in the cockpit for execution.

[0009] Before issuing the final control command to the corresponding equipment in the cockpit, this invention establishes a linkage optimization model based on the remaining battery power, ambient temperature difference, and user comfort. This not only enables dynamic and precise adjustment of equipment power but also achieves a dynamic balance between comfort and energy consumption, ensuring safety while simultaneously achieving optimal energy saving.

[0010] In one optional implementation, the optimal operating power of each device in the cabin is calculated based on the remaining battery power, ambient temperature, and user comfort index, including: Determine the current temperature difference based on the ambient temperature; Determine the power compensation coefficient based on the current temperature difference and the remaining battery power; Obtain the initial operating power of each device in the cockpit; Based on the power compensation coefficient, the current temperature difference, and the user comfort index, the initial operating power is adjusted to obtain the corresponding optimal operating power.

[0011] This invention dynamically adjusts power based on ambient temperature, battery status, and user experience, ensuring that the calculation results simultaneously match environmental energy consumption characteristics, battery power supply capacity, and the user's actual needs, thereby achieving a precise balance between comfort and energy consumption.

[0012] In one optional implementation, the confidence scores for speech, gestures, and bioelectricity are calculated separately, including: Acoustic and semantic features are extracted from the speech signal, and feature fusion calculation is performed to obtain the speech confidence score. Three-dimensional trajectory features and kinematic features are extracted from the gesture signal, and feature fusion calculation is performed to obtain the gesture confidence score; After normalizing the bioelectric signals, a weighted calculation is performed to obtain the bioelectric confidence level; wherein, the bioelectric signals include at least one of heart rate signals, heart rate variability signals, skin conductance response signals, and grip strength signals.

[0013] This invention achieves accurate quantitative evaluation of the interaction intent of each modality through a multi-feature fusion computing architecture of voice, gesture, and bioelectricity, providing a reliable and unified decision-making basis for subsequent multimodal weighted arbitration.

[0014] In one alternative implementation, determining the instruction to be executed based on the arbitration score includes: When the arbitration score is higher than the second preset threshold, the instruction to be executed is the control instruction corresponding to the voice signal, gesture signal and bioelectric signal; When the arbitration score is between the first preset threshold and the second preset threshold, the instruction to be executed is the control instruction confirmed by the user from the set of alternative control instructions. The set of alternative control instructions consists of multiple high-probability cockpit control instructions selected based on the confidence level of each signal, the current interaction scenario filtering, and the user's historical interaction preferences. When the arbitration score is lower than the first preset threshold, the instruction to be executed is a fuzzy control instruction. The fuzzy control instruction is a control instruction generated by combining historical interaction data, the current interaction scenario and user preferences to perform intent association reasoning and ambiguity resolution. The second preset threshold is greater than the first preset threshold.

[0015] This invention divides the arbitration score into three gradient intervals, forming a full-link processing flow that directly executes explicit instructions, intelligently confirms fuzzy instructions, and infers the intent of invalid instructions. This addresses the pain point of coexisting erroneous execution and interaction failure from the root, achieving a balance between interaction accuracy, efficiency, and user experience.

[0016] In one alternative implementation, determining the current interaction scenario includes: Acquire environmental perception data, user status data, and vehicle operation data; After preprocessing the environmental perception data, user status data and vehicle operation data respectively, the external environment features in the environmental perception data, the driver status features in the user status data and the vehicle operation features in the vehicle operation data are extracted. The first weight of the external environment features, the second weight of the driver state features, and the third weight of the vehicle operation features are obtained respectively; wherein, the first weight is greater than the second weight, the second weight is greater than the third weight, and the sum of the weights of the first weight, the second weight and the third weight is a preset value; The scene comprehensive judgment value is obtained by multiplying the external environment features by the first weight, the driver state features by the second weight, and the vehicle operation features by the third weight, respectively, and then summing them. The corresponding current interaction scene is determined based on the scene comprehensive judgment value.

[0017] This invention achieves intelligent scene determination with high accuracy, high real-time performance, and high security by constructing a scene recognition system that integrates the three-dimensional weighted fusion of the environment, user, and vehicle, and allocates safety-oriented weights. This provides core foundational support for the execution of the entire intelligent cockpit interaction strategy.

[0018] In one optional implementation, based on the current interaction scenario, the instruction execution priority of the instruction to be executed is modified to obtain the final control instruction, including: Obtain the scene type of the current interaction scenario. The scene type includes at least one of the following: rainy day scenario, low battery scenario, fatigue driving scenario, high temperature scenario, low temperature scenario, commuting scenario, highway scenario, and parent-child scenario. The preset priority correction rules are matched according to the scenario type. The priority correction rules are used to adjust the execution priority of control commands in the categories of safety, energy consumption control, driving comfort and non-essential entertainment. The original execution priority of the instruction to be executed is adjusted according to the priority correction rules to obtain the corrected execution priority; Based on the revised execution priority, the instructions to be executed are sorted, delayed, or masked to generate the final control instructions.

[0019] This invention achieves precise matching between instruction execution strategies and core scenario requirements by constructing a full-link priority correction mechanism that includes scenario type matching, dynamic weighting of four types of instructions, and hierarchical instruction processing.

[0020] In one optional implementation, after the final control command is issued to the corresponding device in the cockpit for execution, the smart cockpit interaction method further includes: Real-time collection of user feedback on executed commands, including confirmation and rejection feedback; The weights of each signal used in the weighted calculation are dynamically adjusted based on feedback information, and the control strategy corresponding to the current interaction scenario is optimized.

[0021] This invention achieves continuous iterative improvement of system performance and deep adaptation of personalized experience by constructing a full-link self-learning mechanism that integrates instruction execution, user feedback, and two-way closed-loop optimization. It is the core support for the entire intelligent cockpit interaction system to evolve and upgrade from passive execution to proactive operation.

[0022] In one optional implementation, the weights of each signal used in the weighted calculation are dynamically adjusted based on feedback information, including: When the feedback is a confirmation, increase the weight of the voice signal and decrease the weight of the gesture signal. When the feedback is negative, increase the weight of the bioelectric signal; The weights of each signal after adjustment are normalized so that the sum of the weights of each signal is a preset value.

[0023] This invention, through the design of a precise weight update mechanism with feedback type-oriented weight adjustment and normalized constraints, not only ensures the self-evolution capability of the multimodal arbitration system, but also achieves controllable optimization direction, long-term system stability, and rapid adaptation to user habits.

[0024] Secondly, the present invention provides an intelligent cockpit interaction system, the system comprising: The data acquisition module is used to acquire the user's voice signal, gesture signal, and bioelectrical signal, and to calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively. The fusion arbitration module is used to perform weighted calculations on voice confidence, gesture confidence, and bioelectric confidence to obtain an arbitration score; The instruction determination module is used to determine the instructions to be executed based on the arbitration score; The scene optimization module is used to determine the current interaction scene and, based on the current interaction scene, modify the execution priority of the instruction to be executed to obtain the final control instruction; The command issuance module is used to issue final control commands to the corresponding equipment in the cockpit for execution.

[0025] The intelligent cockpit interaction system of this invention features a data acquisition module that simultaneously collects voice, gesture, and bioelectrical signals and calculates their respective confidence levels. A fusion arbitration module employs a weighted fusion strategy to comprehensively judge user intent, significantly reducing the probability of misidentification and command conflicts compared to single-modal or fixed-priority methods, thus ensuring interaction reliability in complex driving environments. Simultaneously, a scene optimization module identifies the current interaction scene in real time and dynamically adjusts the execution priority of pending commands based on scene type. This allows the system to distinguish the urgency of safety, energy consumption, and entertainment commands, ensuring that critical commands are responded to first, while unnecessary commands are delayed or blocked. This not only enhances driving safety and user comfort but also supports personalization and continuous optimization, significantly improving driving comfort and user satisfaction while ensuring a higher level of intelligent cockpit interaction.

[0026] Thirdly, the present invention provides a vehicle, the vehicle including a controller, the controller including a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the intelligent cockpit interaction method of the first aspect or any corresponding embodiment described above.

[0027] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the intelligent cockpit interaction method of the first aspect or any corresponding embodiment described above.

[0028] The intelligent cockpit interaction method and system provided by this invention effectively reduces the risk of single-modal misidentification or command conflict by integrating voice, gesture and bioelectric signals for multimodal confidence weighted arbitration, and improves the accuracy of command recognition in complex driving environments, thereby ensuring the accuracy and robustness of interaction. Furthermore, by dynamically correcting the command execution priority in combination with the current interaction scenario, unnecessary interference can be limited, driving distraction can be reduced, and driving safety can be ensured. At the same time, scene adaptive optimization is achieved, thereby improving the intelligence of cockpit interaction, driving comfort and user satisfaction. Attached Figure Description

[0029] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the first type of intelligent cockpit interaction method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a second process for an intelligent cockpit interaction method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the architecture of the intelligent cockpit interaction system; Figure 4 This is a schematic diagram of the multimodal arbitration algorithm. Figure 5 This is a schematic diagram of the scene recognition and dynamic control process; Figure 6 This is a structural block diagram of an intelligent cockpit interaction system according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the hardware structure of a vehicle according to an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] According to an embodiment of the present invention, an embodiment of an intelligent cockpit interaction method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] This embodiment provides an intelligent cockpit interaction method. Figure 1 This is a schematic flowchart of the first type of intelligent cockpit interaction method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Acquire the user's voice signal, gesture signal, and bioelectrical signal, and calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively.

[0034] It should be noted that the voice signal in this embodiment refers to the sound wave signal emitted by the user through speech. It is the most direct and efficient active interaction method, carrying clear instructions such as "turn on the air conditioner" or "navigate to the company." It can be collected through a distributed microphone array and the driver's / passenger's voice can be picked up directionally using beamforming technology, while filtering out engine noise, wind noise, road noise, and other environmental interference to output a clear digital voice signal. Gesture signals refer to non-voice commands issued by the user through specific hand / arm movements. They are suitable for scenarios where voice is inconvenient (such as when making or receiving phone calls) and include waving, clenching a fist, swiping, pointing, etc. They can be collected by an in-vehicle TOF (Time of Flight) camera, which emits infrared light and receives reflected signals, calculates the three-dimensional spatial coordinates and motion trajectory of the hand, and extracts the shape, speed, direction, and other features of the gesture to achieve contactless gesture recognition. Bioelectric signals refer to weak electrical signals or physiological response signals generated by the user's body. They are a passive and implicit interactive method that can reflect the user's true physiological state and potential intentions. They are not easily affected by environmental interference and can be collected by dedicated sensors integrated into vehicle components. For example, steering wheel grip force sensors can be used to detect the size and changes of the user's grip force, seat heart rate electrodes embedded in the seat back can be used to detect heart rate and heart rate variability, or skin conductance sensors integrated in the steering wheel or armrest can be used to detect skin conductance responses (used to reflect emotional tension).

[0035] To further explain, in this embodiment, the confidence level is a standardized value that quantifies the authenticity of the user's interaction intent (its specific value can be adaptively adjusted according to actual needs, such as a value range of 0-1). The higher the value, the greater the probability that the corresponding signal corresponds to a real operation instruction.

[0036] Step S102: The confidence scores for speech, gesture, and bioelectricity are weighted and calculated to obtain the arbitration score.

[0037] In this embodiment, when multiple modalities issue commands simultaneously or command conflicts occur, this step aims to reflect the reliability differences of different modalities through weight allocation, and finally output a standardized comprehensive decision score, namely the arbitration score.

[0038] Step S103: Determine the instructions to be executed based on the arbitration score.

[0039] In this embodiment, the number of instructions to be executed is not limited and can be adaptively determined according to the actual situation. For example, suppose at a certain moment the voice signal is the user saying "turn on the air conditioner", and the calculated voice confidence is 0.8 (clear recognition); the gesture signal is the user unconsciously waving their hand, and the calculated gesture confidence is 0.3 (misidentified as "turn off the air conditioner"); the bioelectric signal is the user's grip strength is stable and their emotions are calm, and the calculated bioelectric confidence is 0.9 (indicating no intention of emergency operation); the weights of each signal are voice 0.5, gesture 0.2, and bioelectric 0.3; then the arbitration score = 0.8 × 0.5 + 0.3 × 0.2 + 0.9 × 0.3 = 0.73, which belongs to the high confidence interval, and the system directly executes the "turn on the air conditioner" instruction.

[0040] Step S104: Determine the current interaction scenario, and based on the current interaction scenario, modify the execution priority of the instruction to be executed to obtain the final control instruction.

[0041] In this embodiment, the specific method for determining the current interaction scenario and its content can be adaptively adjusted according to actual needs. For example, by acquiring environmental information (such as weather), vehicle operation information (such as vehicle speed and battery), and user status data (such as fatigue), it is possible to accurately identify driving scenarios such as rainy days, low battery, fatigued driving, high temperature, and parent-child interaction.

[0042] Step S105: Issue the final control command to the corresponding equipment in the cockpit for execution.

[0043] In this embodiment, the specific type of device can be determined based on actual needs, such as a screen, air conditioner, etc.

[0044] The intelligent cockpit interaction method provided in this embodiment of the invention calculates and weights the confidence levels of three modalities—voice, gesture, and bioelectricity—to comprehensively determine user intent. This effectively reduces the risk of erroneous execution caused by misidentification of a single modality or command conflicts, and improves the accuracy of command recognition in complex driving environments. Simultaneously, based on the real-time identified current interaction scenario, such as rain, low battery, or fatigued driving, the execution priority of commands to be executed is dynamically adjusted. This allows the cockpit control strategy to meet the safety needs and user expectations of different scenarios, avoiding command rigidity caused by fixed priority modes. This not only achieves personalized responses that adapt to different scenarios and improves the intelligence of cockpit interaction, but also enhances driving comfort and user satisfaction.

[0045] This embodiment provides an intelligent cockpit interaction method. Figure 2 This is a schematic diagram of a second type of intelligent cockpit interaction method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Acquire the user's voice signal, gesture signal, and bioelectrical signal, and calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively.

[0046] Specifically, step S201 above calculates speech confidence, gesture confidence, and bioelectric confidence, including: Step a1: Extract acoustic and semantic features from the speech signal and perform feature fusion calculation to obtain the speech confidence level.

[0047] In this embodiment, acoustic features, namely acoustic recognition scores, are used to reflect speech clarity; semantic features, namely cockpit-specific semantic matching scores, are used to reflect the rationality of instructions. By combining the two, homophones, accent interference, and invalid speech in non-cockpit scenarios can be effectively filtered out, such as accurately distinguishing between "open the sunroof" and "open the car window" in a noisy environment.

[0048] Step a2: Extract three-dimensional trajectory features and kinematic features from the gesture signal, and perform feature fusion calculation to obtain the gesture confidence score.

[0049] In this embodiment, the three-dimensional trajectory feature is the three-dimensional depth trajectory captured by the TOF camera, and the kinematic features include velocity, acceleration, attitude angle, etc. By fusing the two, similar gestures (such as "waving" and "waving") can be accurately distinguished, solving the problem of easy misjudgment of two-dimensional gestures.

[0050] Step a3: After normalizing the bioelectric signals, a weighted calculation is performed to obtain the bioelectric confidence level; wherein, the bioelectric signals include at least one of heart rate signals, heart rate variability signals, skin conductance response signals, and grip strength signals.

[0051] In this embodiment, physiological signals such as heart rate variability (HRV), galvanic skin response (GSR), and steering wheel grip strength are introduced into the interaction confidence system. Normalized weighted calculations quantify the user's physiological stability and the intensity of their operational intent, filling the gap in existing technologies that cannot capture implicit user intentions. Simultaneously, it effectively filters out unconscious misoperations (such as occasional hand tremors or grip strength changes due to tension while driving), and when voice and gesture signals are interfered with, physiological signals assist in determining the user's true intention, further reducing the overall false trigger rate.

[0052] In this embodiment of the invention, a multi-feature fusion computing architecture for each of the three modalities of voice, gesture, and bioelectricity is used to achieve accurate quantitative evaluation of the interaction intent of each modality, providing a reliable and unified decision-making basis for subsequent multimodal weighted arbitration.

[0053] Step S202 involves weighting the speech confidence, gesture confidence, and bioelectrical confidence to obtain the arbitration score. For details, please refer to [link to relevant documentation]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.

[0054] Step S203: Determine the instructions to be executed based on the arbitration score.

[0055] In this embodiment, step S203 includes: Step b1: When the arbitration score is higher than the second preset threshold, the instruction to be executed is the control instruction corresponding to the voice signal, gesture signal and bioelectric signal.

[0056] In this embodiment, the instruction to be executed is not the instruction of each of the three modes, but rather the valid instruction that is pointed to by the signals of the three modes together, or confirmed after weighted arbitration.

[0057] Step b2: When the arbitration score is between the first preset threshold and the second preset threshold, the instruction to be executed is the control instruction confirmed by the user from the alternative control instruction set. The alternative control instruction set consists of multiple high-probability cockpit control instructions selected based on the confidence level of each signal, the current interaction scenario filtering, and the user's historical interaction preferences.

[0058] It should be noted that in this embodiment, the confidence ranking is used to prioritize displaying the instructions with the highest support for each modal signal; the current interaction scenario filtering is used to automatically exclude instructions that do not conform to the current scenario (such as not displaying the parking mode option during vehicle driving, and prioritizing the display of defogging-related instructions in rainy weather); and the user preference is based on historical interaction data to prioritize displaying frequently used function instructions.

[0059] In this embodiment, this step aims to generate 2-3 high-probability alternative instructions for the user to quickly select, so that the user does not need to re-enter the complete instruction, but only needs to confirm to complete the operation.

[0060] Step b3: When the arbitration score is lower than the first preset threshold, the instruction to be executed is a fuzzy control instruction. The fuzzy control instruction is a control instruction generated after combining historical interaction data, the current interaction scenario and user preferences to perform intent association reasoning and ambiguity resolution; wherein, the second preset threshold is greater than the first preset threshold.

[0061] It should be noted that in this embodiment, intent association reasoning and ambiguity resolution technology are used to accurately understand the user's implicit needs. For example, when the user says "adjust it", the system will combine the current air conditioning temperature, music volume and the user's historical operation habits to infer whether to adjust the air conditioning temperature or the music volume. The determination of this type of command can greatly reduce the number of times the user has to repeatedly input, thereby improving the smoothness of the cockpit interaction.

[0062] In this embodiment, the fusion of three-dimensional information of "historical interaction, current scene and user preference" is applied to the intent reasoning of low confidence commands, which solves the problem that traditional systems cannot recognize ambiguous commands (such as "adjust it" or "too noisy").

[0063] Furthermore, this embodiment divides the arbitration score into three gradient intervals, forming a full-link processing flow of direct execution of explicit instructions, intelligent confirmation of ambiguous instructions, and inference of intent for invalid instructions, namely: High confidence interval (arbitration score > second preset threshold): This threshold is trained based on 100,000 sets of real vehicle interaction data to ensure that clear intention commands can be executed quickly without additional confirmation and with low response latency; Medium confidence interval (first preset threshold < arbitration score ≤ second preset threshold): False trigger signals are filtered through a secondary confirmation mechanism, greatly reducing the false execution rate; Low confidence interval (arbitration score ≤ first preset threshold): Instead of directly discarding low-quality signals, we use contextual reasoning to uncover potential intentions, thereby improving the success rate of interactions in low confidence scenarios.

[0064] In this embodiment of the invention, by dividing the arbitration score into three gradient intervals, a balance can be achieved between interaction accuracy, interaction efficiency, and user experience.

[0065] Step S204: Determine the current interaction scenario, and based on the current interaction scenario, modify the execution priority of the instruction to be executed to obtain the final control instruction.

[0066] In this embodiment, determining the current interaction scenario in step S204 includes: Step c1: Obtain environmental perception data, user status data, and vehicle operation data.

[0067] In this embodiment, environmental perception data refers to natural and road environment data inside and outside the vehicle that affect driving safety and driving experience; user status data refers to the driver's physiological state, mental state, and behavioral characteristics data, which are used to reflect the driver's driving ability and attention level; vehicle operation data refers to the vehicle's own operating status and energy consumption data, which are used to reflect the vehicle's capability boundaries and driving conditions; the specific content of each type of data and its acquisition method can be adaptively determined according to actual needs and conventional methods in the field.

[0068] Step c2 involves preprocessing the environmental perception data, user status data, and vehicle operation data respectively, and then extracting the external environment features from the environmental perception data, the driver status features from the user status data, and the vehicle operation features from the vehicle operation data.

[0069] In this embodiment, the specific content of the extracted data features is not limited. For example, external environment features are used to quantify the impact of the external environment on driving safety and driving experience, including rainfall intensity and light intensity; driver state features are used to quantify the driver's physiological state, mental state and attention level, including fatigue index, concentration and emotional tension; vehicle operation features are used to quantify the vehicle's own operating conditions and capability boundaries, including remaining battery charge (SOC), driving speed level and energy consumption level.

[0070] Step c3: Obtain the first weight of the external environment features, the second weight of the driver state features, and the third weight of the vehicle operation features respectively; wherein, the first weight is greater than the second weight, the second weight is greater than the third weight, and the sum of the weights of the first weight, the second weight and the third weight is a preset value.

[0071] In this embodiment, weights are assigned based on the "degree of impact on driving safety" (i.e., environment > user > vehicle). External environmental factors that directly affect driving safety (such as rain, strong light, and low temperature) are used as the core basis for scene determination, ensuring that safety-related scenarios are identified first and corresponding strategies are triggered. For example, in a heavy rain scenario, even if the vehicle has sufficient battery power and the driver is in good condition, the system will still prioritize determining it as a rainy scene and immediately implement safety strategies such as windshield defrosting, high-contrast head-up display (HUD) display, and restriction of non-emergency notifications to avoid driving risks caused by fogged windows and blurred vision.

[0072] It should be noted that the specific values ​​of the weights in this embodiment can be flexibly calibrated using real vehicle test data to adapt to the positioning of different models and the needs of different user groups (for example, the weight of vehicle operation data can be appropriately increased for sporty models, and the weight of user comfort can be appropriately increased for family-oriented models).

[0073] Step c4: Multiply the external environment features by the first weight, the driver state features by the second weight, and the vehicle operation features by the third weight, and then sum them to obtain the scene comprehensive judgment value. Based on the scene comprehensive judgment value, determine the corresponding current interaction scene.

[0074] In this embodiment, a comprehensive weighted judgment value can be used to identify complex scenarios with multiple overlapping factors (such as rain, fatigued driving, high temperature, low battery, high speed, and family activities), generating targeted complex control strategies rather than simply adding up single strategies. For example, in the combined scenario of rain and low battery, the system will prioritize ensuring the full power operation of safety devices such as the windshield defroster and HUD display, while appropriately reducing the power consumption of non-essential devices such as air conditioning and rear entertainment systems, achieving an optimal balance between safety and range. In high temperature and commuting scenarios, the air conditioning will be turned on in advance for pre-cooling, and the air conditioning power will be dynamically adjusted according to road conditions.

[0075] In this embodiment of the invention, a scene recognition system with three-dimensional weighted fusion of environment, user and vehicle, and safety-oriented weight allocation is constructed, which realizes intelligent scene determination with high accuracy, high real-time performance and high security, and provides core foundational support for the execution of the entire intelligent cockpit interaction strategy.

[0076] In this embodiment, step S204 above modifies the instruction execution priority of the instruction to be executed based on the current interaction scenario to obtain the final control instruction, including: Step d1: Obtain the scene type of the current interaction scenario. The scene type includes at least one of the following: rainy day scenario, low battery scenario, fatigue driving scenario, high temperature scenario, low temperature scenario, commuting scenario, highway scenario, and parent-child scenario.

[0077] Note that the scenario types in this embodiment are only examples and can be added or removed as needed for actual applications.

[0078] Step d2: Match the preset priority correction rules according to the scenario type. The priority correction rules are used to adjust the execution priority of control commands for safety, energy consumption control, driving comfort and non-essential entertainment.

[0079] In this embodiment, safety-related control commands are directly related to driving safety and have the highest global authority, such as commands for windshield defrosting, brake assist, lane departure warning, fatigue alert, and lighting control; energy consumption management-related control commands refer to commands that affect vehicle range and energy efficiency, such as commands for adjusting air conditioning power, adjusting screen brightness, charging management, and shutting down unnecessary equipment; safety-related control commands refer to commands that improve the comfort of passengers, such as commands for adjusting air conditioning temperature, seat heating / ventilation, music volume adjustment, and window control; and non-essential entertainment-related control commands refer to entertainment commands unrelated to driving safety and basic experience, such as commands for video playback, games, online live streaming, and rear-seat entertainment systems.

[0080] Step d3: Adjust the original execution priority of the instruction to be executed according to the priority correction rule to obtain the corrected execution priority.

[0081] In this embodiment, the priority adjustment rule follows the principle of "safety first, energy consumption adaptation, and user experience priority". For example, in rainy / fatigue driving scenarios, safety-related instructions have the highest priority, while unnecessary entertainment instructions are automatically blocked; in high / low temperature scenarios, driving comfort instructions are given priority to quickly adjust the cabin environment.

[0082] Step d4: Sort, postpone, or mask the instructions to be executed according to the modified execution priority to generate the final control instructions.

[0083] In this embodiment, multiple pending instructions can be processed simultaneously. The system intelligently sorts, postpones, or blocks instructions based on a modified priority to avoid system lag and user confusion caused by instructions being executed in a cluster. For example, in commuting scenarios, the system automatically executes instructions in the order of "navigation instructions > air conditioning adjustment > music playback > message notifications"; in family scenarios, it prioritizes instructions such as rear entertainment controls and child lock activation. This reduces manual intervention required by users due to unreasonable instruction execution order, significantly shortening the average completion time of a single interaction and greatly improving user satisfaction.

[0084] In this embodiment of the invention, a full-link priority correction mechanism is constructed, which includes scenario type matching, dynamic weighting of four types of instructions, and hierarchical processing of instructions, thereby achieving precise matching between instruction execution strategies and core scenario requirements.

[0085] Step S205: Obtain the remaining battery power, ambient temperature, and user comfort index. The user comfort index is calculated by fusing at least one parameter among seat pressure distribution, skin temperature and humidity, and breathing frequency.

[0086] In this embodiment, a user comfort index based on physiological parameters is introduced into the energy consumption optimization system. By quantifying the user's physical sensation through objective data such as seat pressure distribution, skin temperature and humidity, and breathing rate, the system replaces the traditional subjective temperature setting, ensuring that power adjustment accurately matches the user's actual needs.

[0087] Step S206: Calculate the optimal operating power of each device in the cabin based on the remaining battery power, ambient temperature, and user comfort index.

[0088] In this embodiment, the optimal operating power can be adaptively adjusted based on actual needs. For example, seat heating can be turned off only when SOC < 30%, while seat cushion heating is retained. The screen brightness decreases linearly with the power consumption rather than suddenly dimming, in order to minimize the impact of energy saving on the user experience. When the user's comfort index is low, the power of environmental control devices such as air conditioning and seats is automatically increased. When the user is in a comfortable state, the power consumption of the devices can be gradually reduced, taking into account both energy saving and user experience.

[0089] Specifically, step S206 includes: Step e1: Determine the current temperature difference based on the ambient temperature.

[0090] Step e2: Determine the power compensation coefficient based on the current temperature difference and the remaining battery power.

[0091] In this embodiment, the ambient temperature difference and the remaining battery power are coupled to generate a power compensation coefficient, enabling dynamic power allocation under different temperature environments. For example, under the same SOC, the air conditioning power compensation coefficient will be appropriately increased in low-temperature winter environments to avoid insufficient heating due to excessive energy saving; while in high-temperature summer environments, the cooling power and battery life consumption need to be balanced.

[0092] Step e3: Obtain the initial operating power of each device in the cockpit.

[0093] In this embodiment, the initial operating power is the default power of each device at the factory, which can be determined based on the device's manufacturing manual.

[0094] Step e4: Adjust the initial operating power according to the power compensation coefficient, the current temperature difference and the user comfort index to obtain the corresponding optimal operating power.

[0095] In this embodiment of the invention, the power is dynamically adjusted by considering ambient temperature, battery status, and user comfort. This ensures that the calculation results match the environmental energy consumption characteristics, battery power supply capacity, and the user's actual comfort needs, thereby achieving a precise balance between comfort and energy consumption.

[0096] Step S207: After correcting the equipment power parameters in the final control command based on the optimal operating power, the final control command is sent to the corresponding equipment in the cockpit for execution.

[0097] In this embodiment, before issuing the final control command to the corresponding equipment in the cockpit, by establishing a linkage optimization model of the remaining battery power, ambient temperature difference, and user comfort, not only can the dynamic and fine adjustment of equipment power be realized, but also the dynamic balance between comfort and energy consumption can be achieved, ensuring safety while simultaneously meeting the optimal energy saving requirements.

[0098] It should be noted that, addressing the shortcomings of traditional cockpit interaction systems—fixed weights, rigid strategies, inability to self-evolve, and a one-size-fits-all approach—this application improves the intelligent and personalized interactive experience of cockpit interaction through a self-learning mechanism based on user feedback. Therefore, after issuing the final control command to the corresponding device in the cockpit for execution, the intelligent cockpit interaction method of this embodiment further includes: Step f1: Collect user feedback information on executed instructions in real time. The feedback information includes confirmation feedback and rejection feedback.

[0099] In this embodiment, the specific content of the feedback information can be adaptively adjusted based on actual needs. For example, if a user executes a command to lower the temperature and then manually / voice-checks to raise the air conditioner temperature, the corresponding feedback information will be a rejection response.

[0100] Step f2: Dynamically adjust the weights of each signal used in the weighted calculation based on the feedback information, and optimize the control strategy corresponding to the current interaction scenario.

[0101] In this embodiment, step f2 above, which dynamically adjusts the weights of each signal used in the weighted calculation based on feedback information, includes: Step f21: When the feedback information is a confirmation feedback, increase the weight of the voice signal and decrease the weight of the gesture signal.

[0102] In this embodiment, when the user confirms the executed command, the weight of the corresponding signal is increased for voice because voice is the user's most proactive and explicit expression of intent. The user's confirmation of the command's correctness indicates that the voice recognition was accurate. Strengthening the voice weight allows for faster and more direct execution of clear voice commands in the future, reducing unnecessary secondary confirmations. The weight of the corresponding signal is decreased for gesture because gestures are the most ambiguous modality (e.g., unconscious waving, adjusting the steering wheel, or picking up objects can all be misrecognized). The user's confirmation of the command's correctness often means that the gesture signal either did not participate in the decision-making process or was consistent with the voice but less reliable. Appropriately reducing the gesture weight reduces the risk of future false triggers from the outset.

[0103] Step f22: When the feedback information is negative feedback, increase the weight of the bioelectric signal.

[0104] In this embodiment, when the user rejects the executed instruction, the weight of the corresponding signal is increased for the bioelectric signal because the bioelectric signal is an unconscious physiological response of the user and cannot be faked. Once the user rejects the instruction, it indicates that there was a misjudgment in the voice or gesture. At this time, strengthening the weight of the bioelectric signal can enable it to play a stronger cross-validation role in future multimodal arbitration, timely filter out the erroneous signals of the dominant modality, and help enhance the error correction capability of the system.

[0105] Step f23: Normalize the weights of each signal after adjustment so that the sum of the weights of each signal is a preset value.

[0106] In this embodiment, the normalization process mandates that the sum of all signal weights after adjustment remain at a preset value (e.g., 1), which solves the system imbalance problem caused by the "unlimited growth / decay of a certain modality weight" that is prone to occur in traditional dynamic weights. For example, it will not cause the voice weight to rise to 100% due to multiple consecutive voice confirmations by the user, completely blocking gestures and bioelectric signals; nor will it cause the bioelectric weight to be too high due to multiple rejections, over-relying on physiological signals and ignoring the user's active commands.

[0107] Specifically, this embodiment designs a precise weight update mechanism with feedback type-oriented weight adjustment and normalization constraints, which not only ensures the self-evolution capability of the multimodal arbitration system, but also achieves controllable optimization direction, long-term system stability, and rapid adaptation to user habits.

[0108] In this embodiment of the invention, a full-link self-learning mechanism is constructed, which includes instruction execution, user feedback, and two-way closed-loop optimization. This mechanism enables continuous iterative improvement of system performance and deep adaptation of personalized experience, serving as the core support for the entire intelligent cockpit interaction system to evolve from passive execution to proactive upgrade.

[0109] In practical applications, existing intelligent cockpit interactions in new energy vehicles suffer from technical defects such as fragmented interaction (i.e., voice / touch / gesture interactions lack a unified arbitration mechanism, and there is no priority decision logic when multiple commands conflict), lack of scene perception (i.e., temperature / seat adjustment relies on preset rules and does not consider the dynamic coupling between the user's real-time physiological state and external environmental parameters), extensive energy management (i.e., the energy consumption of devices such as screen brightness / air conditioning power is not dynamically correlated with the battery SOC value), and poor service continuity (i.e., user preference data is not synchronized across devices, and historical behavior learning does not consider scene context). This paper proposes an intelligent cockpit interaction system based on multimodal perception and dynamic scene adaptation. This system achieves dynamic optimization of cockpit device control strategies by integrating environmental perception, user state recognition, and vehicle operation data, and establishes an intelligent linkage mechanism between energy consumption and driving range.

[0110] It should be noted that the core technologies and their effects of the above system are as follows: 1. Multimodal Interaction Unified Arbitration and Dynamic Weight Adaptation: By constructing a three-modal confidence-weighted arbitration mechanism for voice signals, gesture signals, and bioelectric signals, and combining it with preset basic weights (such as 0.4 for voice signals, 0.3 for gesture signals, and 0.3 for bioelectric signals), the reinforcement learning module adjusts and normalizes the weights in real time based on user confirmation / rejection feedback, and sets three-level decision thresholds to complete the instruction judgment.

[0111] Technical effects: It completely solves the problem of multiple command conflicts and no priority decision-making, and realizes unified scheduling of multimodal interaction; command execution is more in line with user intent, the success rate of multimodal conflict resolution is 92.3%, the average response latency is only 220ms, and the false trigger rate is <0.5 times / thousand kilometers.

[0112] 2. Dynamic Scene Adaptation Based on Environment, User, and Vehicle: By integrating three types of data—external environment, in-vehicle user physiological state, and vehicle operation—a dynamic coupling model of scene, physiology, and vehicle is established, and exclusive control strategies are generated for typical scenarios such as rainy days and low battery.

[0113] Technical benefits: By breaking free from the limitations of fixed preset rules, the cabin equipment can be precisely adjusted to match real-time scenarios and user physiological needs, significantly improving driving comfort and safety.

[0114] It should be noted that the scenarios in this embodiment are common types of car use, such as: rainy days, low car battery, hot / cold weather, driver fatigue, daily short-distance commuting, long-distance highway driving, and having children in the car. The system can comprehensively judge and determine the current scenario by obtaining the above three types of data from sensors. For example, it can check the external environment, such as whether it is raining and the temperature; check the driver's condition, such as whether he is tired or uncomfortable; and check the current vehicle condition, such as the battery level, driving time, and speed. Among these, the external environment has the greatest impact on safety and is taken first; the driver's condition is the second; and finally, the vehicle condition is considered. The combination of these three factors can automatically identify the current scenario.

[0115] To further clarify, the scenario-physiology-vehicle dynamic coupling model in this embodiment binds three types of information—outside weather, driver's physical condition, and vehicle battery level / driving status—to automatically correspond to a set of cockpit control schemes, rather than using a fixed and unchanging pattern. This model collects various data on the external environment, driver's condition, and vehicle in real time; cleans up the messy data and removes interfering information; calculates a comprehensive value based on the importance of environment > driver > vehicle to determine the current scenario; matches each scenario with corresponding control methods for devices such as air conditioning, screens, and seats; and continuously adjusts and optimizes this matching logic based on the driver's acceptance / disapproval of control commands.

[0116] To further explain, since the needs of different scenarios are different, such as prioritizing defogging to ensure safety on rainy days, saving power to ensure range when the battery is low, reducing interference when the driver is fatigued, and gentle control when traveling with children; if a fixed mode is used, it cannot adapt to these special situations. The exclusive control strategy designed in this embodiment can make the cockpit control more in line with the current driving environment, making driving safer and sitting more comfortable.

[0117] 3. Optimization of cabin energy consumption in conjunction with battery SOC and comfort index: By establishing energy consumption optimization models for cabin equipment such as screens, air conditioning, and seat heating, battery SOC, ambient temperature difference, and user comfort index are used as core adjustment parameters, and differentiated power constraint strategies are adopted for different equipment.

[0118] Technical effects: It enables refined control of cabin energy consumption, reducing ineffective energy consumption while ensuring comfort. The NEDC (New European Driving Cycle) driving range optimization rate is 7.8%, balancing driving experience and the driving range capability of new energy vehicles.

[0119] It should be noted that the energy consumption optimization model in this embodiment is a power regulator after the instruction is determined and before the device executes it. That is, the system first uses a multimodal arbitration algorithm to determine what operation needs to be performed (such as turning on the air conditioner or adjusting the screen brightness); then it substitutes three real-time data points—the current battery level, the in-vehicle temperature, and the user's comfort level—into the model to automatically calculate the most suitable operating power for the air conditioner, screen, and seat heating; finally, it sends the instruction with the adjusted power to the cabin equipment for execution, thus avoiding unnecessary power consumption by the equipment.

[0120] Furthermore, the energy consumption optimization model in this embodiment is strongly correlated with different interaction scenarios. That is, different scenarios use different adjustment methods. For example, on rainy days or when the driver is fatigued, safety is prioritized and the power of safety devices such as defogging and HUD is not reduced. When the vehicle's battery is low, the power of the air conditioner and screen is automatically reduced to save power and maintain range. When the weather is too hot or too cold, the power of the air conditioner is automatically increased to prioritize comfort. For daily commutes or long-distance highway trips, the power is finely adjusted to balance power consumption and user experience based on the duration of use.

[0121] Furthermore, the constraint strategies in this embodiment include: tiered power consumption constraints, where low power levels necessitate power saving while sufficient power levels allow for relaxed restrictions; comfort constraints, where increased power consumption for devices like air conditioners increases user discomfort and decreases power consumption when user comfort is achieved; device differentiation constraints, where air conditioners, screens, and seats are adjusted separately, with air conditioners being the most power-intensive and subject to strict control, while seats have lower power consumption and are not strictly limited; and temperature compensation constraints, where greater temperature differences from the standard 25°C result in greater power compensation for the devices. Note that these constraint strategies are based on the following: range is a core requirement for new energy vehicles, necessitating tiered power consumption control; cabin adjustments ultimately aim to enhance user comfort; different devices have different power consumption and functions, requiring a one-size-fits-all approach; and ambient temperature directly impacts user experience, necessitating dynamic compensation.

[0122] 4. Personalized learning of user preferences and continuous cross-device service: User profiles are built through data aggregation and association rule mining, and stored in an encrypted TEE (Trusted Execution Environment) area to achieve cross-device synchronization and context-based learning of user preferences.

[0123] Technical effects: It solves the problems of asynchronous preference data and lack of service continuity. The intelligent interaction and control of the cockpit continuously adapts to the user's personalized habits, and the convenience and intelligence level of long-term use are significantly improved.

[0124] In one specific embodiment, Figure 3 This is a schematic diagram of the intelligent cockpit interaction system architecture. As shown in the diagram, the architecture includes: an environmental perception unit, a user status unit, a vehicle data unit, and a central decision processor. The environmental perception unit includes external sensors (millimeter-wave radar for rainfall data, temperature and humidity sensors) and internal sensors (TOF camera for gesture data, infrared array for occupant distribution). The user status unit includes visual perception (micro-expression recognition, posture tracking) and bioelectrical signals (seat heart rate detection electrodes for heart rate monitoring, steering wheel grip force sensor for grip force detection). The vehicle data unit includes data from the three-electric system (battery SOC value, motor output power) and navigation system data (remaining range, route changes).

[0125] It should be noted that the environmental perception unit, user status unit, and vehicle data unit constitute the perception layer. The environmental perception data, user status data, and vehicle data output by this layer are the core input sources for cockpit interaction. These three types of data first undergo timestamp alignment and fusion preprocessing by the central decision processor, then are input into a multimodal arbitration algorithm for command decision-making. Finally, the system, in conjunction with the scenario adaptation and energy consumption optimization model, outputs control commands, forming a complete closed-loop interaction process. The specific process is as follows: 1. First, determine the basic control commands: The system uses a multimodal arbitration algorithm to determine what operation the user wants to perform, such as "turn on the air conditioner, increase the screen brightness, or turn on the seat heating", based on voice, gestures, and the driver's physical state, and obtains the basic control commands.

[0126] 2. Scenario adaptation and execution priority setting: The system combines the weather outside the vehicle, the driver's status, and the vehicle's remaining battery power to identify the current scenario (such as rainy days, low battery, and fatigued driving) and determine the execution priority, that is: safety-related instructions are executed first, and non-essential instructions are postponed when the battery is low.

[0127] 3. Energy consumption optimization calculates the operating power of equipment: The system calculates the most suitable power consumption of air conditioning, screen and seat by substituting the vehicle's remaining power, ambient temperature and driver comfort into the energy consumption model. For example, if the power is low, the power of air conditioning is reduced and the power of seat heating is increased when the weather is cold.

[0128] 4. Integrate final output instructions and closed-loop optimization: The system integrates the previous execution priority and equipment power into the basic instructions and issues them to the cockpit equipment for execution; at the same time, it receives the user's approval / disapproval of the instructions and adjusts the subsequent scenario and energy consumption parameters accordingly, so that the interaction becomes more and more in line with the user's habits over time.

[0129] To further explain, the environmental perception data is used in this embodiment as follows: 1. External data (such as rainfall, temperature and humidity, light intensity): used to determine the external scene (rain / high temperature / strong light), directly trigger the adjustment of the cabin environment equipment (such as defogging, heating, brightness compensation), and at the same time serve as the external constraint condition for the multimodal arbitration algorithm to adjust the execution priority of instructions.

[0130] 2. In-vehicle data (such as gesture trajectory and occupant distribution): Extract gesture features and calculate gesture confidence, which serve as the core input parameters of the multimodal arbitration algorithm. At the same time, it is used to locate the interaction object and avoid accidental triggering by non-target occupants.

[0131] To further explain, the user status data is used in this embodiment as follows: 1. Visual data (such as micro-expressions and postures): Recognize user emotions and operational intentions, calculate user state confidence, and assist in correcting the intention of interaction commands.

[0132] 2. Bioelectric data (such as heart rate and grip strength): Calculate bioelectric confidence level and generate user comfort index to provide core parameters for energy consumption optimization model.

[0133] Note that the confidence levels of the two types of data (also known as signals) mentioned above are directly input into the multimodal arbitration algorithm to participate in the arbitration score calculation. Furthermore, this part does not include user voice data because the user state data here only collects the driver's facial expressions, posture, heart rate, and steering wheel grip strength to determine if the driver is fatigued, in a good mood, and comfortable; it only reflects the driver's physical state. Voice data / signals, on the other hand, are commands given directly by the user (such as "turn on the air conditioner"), which are independent command inputs and will be given separately to the multimodal arbitration algorithm, not included in the user state data. In short, in this embodiment, voice data is used to directly give commands, while user state data is used to assist in judging the driver's state and correcting commands; the two types of data are used separately.

[0134] To further explain, the specific use of vehicle data in this embodiment is as follows: 1. Three-electric data (such as battery SOC, motor power): Construct energy consumption constraint boundaries. The SOC value directly participates in the power adjustment of the cockpit equipment and also serves as the basis for adjusting the operating condition weights of the arbitration algorithm.

[0135] 2. Navigation data (such as remaining mileage and route elevation): Predict driving scenarios and optimize interaction and energy consumption strategies in advance to ensure the stability of battery life and interaction.

[0136] It should be noted that the perception layer data and the multimodal arbitration algorithm in this embodiment have the following relationship: 1. Data Input Association: The speech confidence, gesture confidence, and bioelectric confidence output by the perception layer are the only input sources for the confidence weighted calculation of the multimodal arbitration algorithm. The algorithm calculates the arbitration score with a basic weight of 0.4×speech + 0.3×gesture + 0.3×bioelectric.

[0137] 2. Decision Constraint Association: Environmental perception data (scene) and vehicle data (operating condition) serve as external constraints, dynamically adjusting the weight coefficients and decision thresholds of the arbitration algorithm. For example, in rainy scenarios, the weight of safety-related instructions is increased, and in low SOC scenarios, the score threshold of unnecessary interaction instructions is reduced.

[0138] 3. Execution of closed-loop correlation: The command results output by the arbitration algorithm are combined with the scene and energy consumption data of the perception layer to generate the final control action; the user's confirmation / rejection feedback on the command is sent back to the perception layer and algorithm module for reinforcement learning weight update, realizing full-link adaptation of perception-decision-execution-feedback.

[0139] Note that confirmation feedback indicates that the user accepts the execution of the scenario instruction; rejection feedback indicates that the user chooses to exit or change the scenario instruction being executed.

[0140] Furthermore, the multimodal arbitration algorithm in this embodiment is specifically as follows: def modality_arbitration(voice_conf, gesture_conf, bio_conf): # Calculate the confidence weights for each modality score = 0.4*voice_conf + 0.3*gesture_conf + 0.3*bio_conf if score>0.7: return execute_action() elif 0.5 <score<= 0.7: return request_confirmation() else: return log_ambiguity() In this system, a confidence-weighted decision-making mechanism is adopted, and the arbitration function modality_arbitration() is defined as: arbitration score = 0.4 × voice confidence + 0.3 × gesture confidence + 0.3 × bioelectrical signal confidence bio_conf.

[0141] It should be noted that the detailed process of the three types of confidence quantification in this embodiment includes: 1. Voice Confidence Quantization: Extract the acoustic model score (0~1) of the speech recognition engine; calculate the semantic cosine similarity (0~1) between the command text and the cockpit command library; fuse the results, i.e.: Voice confidence = 0.6 × acoustic score + 0.4 × semantic matching degree, normalized to 0~1. Note that this confidence quantization adds cockpit scene semantic constraints, which is a scene-based improvement.

[0142] 2. Gesture Confidence Quantization (TOF Trajectory + Kinematic Features): A TOF camera acquires 3D depth point clouds and temporal gesture trajectories; the Dynamic Time Warping (DTW) algorithm is used to calculate the trajectory matching degree (0~1) between the real-time trajectory and the standard template; kinematic features such as velocity, acceleration, and attitude angles are extracted, and Euclidean distance similarity (0~1) is calculated; the results are then fused: Gesture confidence = 0.5 × trajectory matching degree + 0.5 × kinematic feature similarity, normalized to 0~1. Note that this confidence quantization uses 3D depth trajectory and cockpit space constraints, which has advantages in resisting light interference and reducing false recognition, representing an improvement for automotive-grade scenarios.

[0143] 3. Bioelectric Confidence Quantification: Physiological signals such as heart rate, HRV, GSR, and grip strength are collected; each individual signal is normalized to 0-1 to reflect physiological stability and intent intensity; the signals are then fused and calculated as follows: Bioelectric Confidence = 0.4 × HRV + 0.3 × GSR + 0.3 × Grip Strength, normalized to 0-1. Note that this confidence quantification is used in multimodal decision-making, representing an innovative application scenario.

[0144] To further explain, the specific meaning of the decision threshold execution rule in this embodiment includes: 1. Arbitration score > 0.7: Direct execution of instructions (confidence threshold obtained through training with 100,000 sets of interactive data).

[0145] 2. 0.5 < Arbitration Score ≤ 0.7: Trigger a secondary confirmation mechanism, displaying alternative commands via the HUD. Note that the alternative commands are 2-3 high-probability cabin control commands (such as air conditioning temperature adjustment, seat heating, and volume adjustment) selected based on confidence ranking, scenario filtering, and user historical preferences. Users can quickly confirm these commands through a single modal.

[0146] 3. Arbitration score ≤ 0.5: Record the fuzzy command and initiate contextual analysis. Note that this contextual analysis combines historical interactions, the current scenario, and user preferences to perform intent association reasoning and ambiguity resolution, rather than processing fuzzy commands in isolation. This helps ensure a higher success rate for interactions in low-confidence scenarios. For details on the multimodal arbitration algorithm, please refer to [link to relevant documentation]. Figure 4 .

[0147] Furthermore, this embodiment also includes dynamic weight adjustment, namely: the system has a built-in reinforcement learning module that updates the weight coefficients in real time based on historical confirmation / rejection records. For example, when the user confirms a valid command, the voice weight increases by 10% and the gesture weight decreases by 5%; when the user rejects a command, the bioelectric signal weight increases by 5%. Note that normalization is performed after weight adjustment to ensure that Σweight=1.

[0148] It should be noted that the specific algorithm of the energy consumption optimization model in this embodiment is as follows: P_opt = [P_screen×0.8k, P_ac×1.2e^(-0.1C), P_seat×1.0] in: k = SOC% × (1+0.05ΔT) ΔT = 25℃ - Ambient temperature C = User comfort index (0-10 scale).

[0149] It should be explained that P_screen is the factory default base operating power of the vehicle screen; P_ac is the factory default base operating power of the vehicle air conditioning; and P_seat is the factory default base operating power of the seat heating / ventilation.

[0150] In this embodiment, the system acquires three key data points in real time: vehicle remaining battery power (SOC), current ambient temperature, and user comfort index C. Then, it calculates intermediate parameters: temperature difference ΔT = 25℃ minus the current ambient temperature; battery compensation coefficient k = current battery percentage × (1 + 0.05 × temperature difference). These parameters are then substituted into formulas to calculate: optimal power for the in-vehicle screen = screen base operating power × 0.8 × k; optimal power for the in-vehicle air conditioner = air conditioner base operating power × 1.2 × coefficient automatically adjusted based on comfort; seat power remains unchanged at the base power. Finally, the system calculates the most suitable power consumption for each device and sends it to the cabin equipment for execution, achieving both power saving and comfort.

[0151] It should be explained that in this embodiment, the SOC coupling coefficient k is activated when SOC < 20% and the power saving mode is enabled by 0.6 times the air conditioner's base power; when SOC > 80% the restriction is lifted and the air conditioner's power is increased by 1.2 times; this temperature compensation factor achieves a 5% coefficient adjustment for every 1°C temperature difference.

[0152] To further explain, the comfort index C in this embodiment is calculated by integrating multiple parameters such as seat pressure distribution, skin temperature and humidity, and breathing rate; for every unit increase in this index, the air conditioning power decreases by 10% according to the index curve.

[0153] To further explain, the device differentiation adjustment in this embodiment includes: linear adjustment of screen brightness (e.g., brightness decreases by 8% for every 10% decrease in SOC) and step control of seat heating (e.g., only seat cushion heating is turned on when SOC < 30%).

[0154] In a specific embodiment, such as rain scene optimization, the input parameters include: rainfall intensity (>50mm / h), driver status (blinking frequency >20 times / minute), and battery SOC (65%); the output actions include: 1. Activating the windshield electric heater (priority A); 2. Switching the HUD to high contrast mode (raindrop compensation algorithm); 3. Restricting non-emergency notifications (attention protection mode); 4. Switching the air conditioner to external circulation (to prevent glass fogging).

[0155] In this embodiment, combined with Figure 3The system architecture diagram is as follows (i.e., Environmental Perception Unit → User Status Unit → Vehicle Data Unit → Central Decision Processor → Device Execution). Input parameters are collected by the perception layer (corresponding to the Environmental Perception Unit in the architecture diagram), namely: environmental perception data (Environmental Perception Unit), rainfall intensity > 50 mm / h (millimeter-wave radar acquisition), user status data (User Status Unit), driver blink frequency > 20 times / minute (visual perception acquisition), vehicle data (Vehicle Data Unit), and battery SOC = 65% (three-electric system acquisition). Based on these input parameters, the entire cockpit interaction process—from perception acquisition to data fusion, multimodal arbitration, scene decision-making, and action output—is completed. The specific implementation is as follows: Step 1: Data Acquisition and Upload at the Perception Layer (corresponding to lower-level modules in the architecture diagram): The millimeter-wave radar in the environmental perception unit detects rainfall intensity, determines it as a heavy rainfall scenario, and uploads the data to the central decision processor; the OV2740 camera in the user status unit detects blink frequency, determines that the driver's attention is reduced and they are prone to fatigue, and uploads the user status characteristics; the three-electric system in the vehicle data unit obtains that the battery SOC is 65%, determines it as a normal energy consumption condition, and uploads the vehicle operating parameters.

[0156] Step 2, Central Decision Processor Data Fusion: The central decision processor performs timestamp alignment, feature fusion, and scene labeling on the three types of data to generate composite scene features of rainy day - insufficient attention - normal battery level, which serve as the core basis for multimodal arbitration and scene adaptation.

[0157] Step 3: Multimodal arbitration algorithm participates in intent determination: When there is no user active instruction, the system triggers active safety interaction based on scene confidence; the scene confidence is calculated by weighting environmental data (0.6) + user status data (0.3) + vehicle data (0.1). If the score is >0.7, the active control instruction is triggered directly without secondary confirmation.

[0158] Note that user-initiated commands are control requests actively issued by the user, such as saying "turn on the air conditioning" to the car or using gestures to change music. These are user-initiated commands and are distinct from commands automatically triggered by the system based on scenario assessment (such as automatic defogging in rainy weather). Furthermore, the 0.6, 0.3, and 0.1 values ​​in the scenario confidence score are fixed weights determined based on their importance to driving safety, combined with extensive real-vehicle testing. For example, environmental data (rain, strong light, temperature) directly affects driving safety and has the highest importance, hence a weight of 0.6; user status data (driver fatigue, discomfort) indirectly affects safety and has a lower importance, hence a weight of 0.3; vehicle data (battery level, speed) has the least direct impact on driving safety, hence a weight of 0.1. The calculation uses the confidence scores of each of the three data types, multiplied by their respective weights, and then summed to obtain the final scenario confidence score. A score greater than 0.7 triggers the system to automatically control the system.

[0159] Step 4: Dynamic scene adaptation and energy consumption optimization linkage: Scene adaptation, such as heavy rainfall → prioritize the execution of defogging, vision enhancement, and attention protection strategies; Energy consumption optimization, such as SOC=65% being within the normal range, do not activate power saving limits to ensure that safety equipment operates at full power.

[0160] Step 5: Output Control Actions (Executed by Cockpit Equipment): The central decision processor sends instructions to the cockpit execution unit, outputting four actions: activating the windshield electric heater (priority A, essential for defogging in rainy weather), switching the HUD to high-contrast mode (raindrop compensation algorithm, improving visibility), limiting non-emergency notifications (attention protection mode, avoiding distraction), and switching the air conditioning to external circulation (to prevent windshield fogging and ensure driving safety). Note that for details on scene recognition and dynamic control, please refer to [link to relevant documentation]. Figure 5 .

[0161] Step 6, Feedback and Closed-Loop Optimization: After the device executes, the perception layer will continuously collect environmental and user status data. If the frequency of rainfall / blinking returns to normal, the system will automatically exit the rainy day optimization mode. At the same time, the interaction record of this scene will be stored in the user profile for multimodal arbitration weight iteration.

[0162] In summary, the correspondence between input parameters and output actions in this embodiment is shown in Table 1.

[0163] Table 1

[0164] In this embodiment, the above-mentioned intelligent cockpit interaction system mainly includes the following processes in practical applications: I. Hardware configuration example.

[0165] In this embodiment, this part includes the following: 1. Sensor network deployment, as detailed below: 1.1 The environmental perception unit includes an external sensor group and an internal perception module. The external sensor group is a multispectral sensor array (model AMS AS7341L) integrated on the inside of the windshield, enabling simultaneous monitoring of rainfall (1600nm near-infrared band) and ultraviolet intensity. A temperature and humidity composite sensor (Sensirion SCD40) is embedded in the roof shark fin antenna module (i.e., the vehicle-mounted integrated antenna module), with a sampling frequency of 10Hz and a measurement range of -40℃ to 85℃. The internal perception module consists of a capacitive grip force sensor (TE Connectivity FSG15N1A) embedded at the 3 / 9 o'clock position on the steering wheel, with a range of 0-50N and a linearity of ±1.5%. A flexible piezoelectric film array (8×8, 20mm spacing) is integrated into the seat back, supporting posture pressure distribution monitoring and heart rate detection.

[0166] 1.2 The vision processing system includes: a driver's side monitoring unit and a cockpit surround view system. The driver's side monitoring unit uses an OV2740 camera (2 megapixels, f / 1.8 aperture) with a 940nm infrared LED to achieve pupil tracking in low-light conditions; the binocular cameras have a 65mm spacing, with a depth detection accuracy of ±3mm@0.5m. The cockpit surround view system consists of four AR0233 global shutter cameras (120° FOV) providing panoramic coverage, mounted on the A-pillars and roof console; the image processing board uses an onboard FPGA (Xilinx Zynq UltraScale+) for real-time distortion correction with a latency of <8ms.

[0167] 2. Computing and control platform, as detailed below: 2.1 Main control chipset: Qualcomm SA8155P automotive-grade SoC: running Linux RTOS, responsible for multimodal data fusion and decision-making algorithms; 8-core Kryo 485 CPU (up to 2.4GHz); Adreno 640 GPU, used for visual feature extraction (throughput 1.2TOPS); Cambricon MLU220 intelligent accelerator card, used for deploying biosignal processing models (INT8 quantization, energy efficiency ratio 3.2TOPS / W).

[0168] 2.2 The communication architecture includes: in-vehicle network and short-range wireless. The in-vehicle network is a CAN FD bus, transmitting vehicle electrical, electronic control, and powertrain data (500kbps base rate, 2Mbps burst mode); Gigabit Ethernet is used for raw data transmission from cameras (TSN Time-Sensitive Networking Protocol). The short-range wireless is a BLE 5.1 ​​Mesh network for connecting biosensors (maximum 32 nodes, 20ms broadcast interval); a UWB positioning module (Decawave DW3000) is used for spatial positioning during gesture interaction (accuracy ±5cm).

[0169] II. Software Implementation Process.

[0170] In this embodiment, this part includes the following: 1. Data processing pipeline, see Table 2 for details.

[0171] Table 2

[0172] 2. Implementation of key software modules, as detailed below: 2.1 Real-time Data Fusion Engine: Utilizing the Apache Kafka stream processing framework, a three-layer data pipeline is constructed, including: a high-speed channel (latency <50ms), a standard channel (200ms time window), and a batch processing channel (updated hourly). The high-speed channel employs zero-copy technology to directly access the DMA buffer, deploys a lightweight protocol parser (ProtoBuf encoding), and supports priority preemption mechanisms (such as interrupting the current task due to collision warnings). The standard channel aggregates multi-source data based on a sliding window algorithm, performs timestamp alignment and data validity verification, and uses Kalman filtering to eliminate sensor noise. The batch processing channel constructs user profile feature vectors, such as [driving style, device usage frequency, physiological response baseline], uses the FP-Growth algorithm to mine cross-scenario association rules, generates personalized configuration templates, and encrypts and stores them in the TEE area.

[0173] 2.2 Optimization of Dynamic Arbitration Algorithm: A reinforcement learning mechanism is introduced to update the weight coefficients. The specific algorithm is as follows: class WeightUpdater: def __init__(self): self.alpha = 0.4 # Initial speech weights self.beta = 0.3 # Initial gesture weight self.gamma = 0.3 # Initial weights of the bioelectrical signal def update_weights(self, user_feedback): # Adjust weights based on user confirmation / rejection actions if user_feedback == 'confirm': self.alpha *= 1.1 self.beta *= 0.95 elif user_feedback == 'reject': self.gamma += 0.05 # Weight normalization processing total = self.alpha + self.beta + self.gamma self.alpha / = total self.beta / = total self.gamma / = total The key technical features of the weight update mechanism include the initial weight allocation basis and the online learning strategy. The initial weight allocation basis is: voice 0.4 (dominated by high-frequency, essential commands), gestures 0.3 (spatial operational advantages), and bioelectrical signals 0.3 (capturing latent needs). The online learning strategy defines a feedback gain factor α=0.85 (attenuating historical influences), employs an ε-greedy strategy to balance exploration and utilization (ε=0.15), and performs weekly weight convergence tests (variance threshold σ² < 0.05).

[0174] 2.3 Implementation of energy consumption optimization model: Establish equipment power constraint matrix, see Table 3 for details.

[0175] Table 3

[0176] III. System Integration Testing.

[0177] In this embodiment, after the system completes hardware deployment, software integration, and algorithm programming, it undergoes comprehensive integration testing according to automotive-grade standards. Building upon existing reliability and road testing, five new specialized tests are added: multimodal arbitration, scenario adaptation, energy consumption optimization, data security, and hardware-software synergy. These tests comprehensively verify the system's functional integrity, real-time performance, stability, and security, meeting the requirements for pre-installed mass production of new energy vehicles. Specifically, these tests include: 1. Basic Reliability and Environmental Adaptability Testing (Improved and Completed): Building upon the original extreme temperature and electromagnetic compatibility tests, automotive-grade environmental and mechanical reliability tests are added to verify the stable operation of the sensing and computing layers under complex automotive conditions. This section includes the following: 1.1 Extreme Environment Testing: High-temperature operation, continuous operation in an 85℃ environmental chamber for 72 hours, with infrared sensor and bioelectrode signal attenuation rate <3%; Low-temperature testing, cold start test at -40℃, bioelectrode impedance change within ±10%, system start-up time <3s; Temperature and humidity cycling, 100 cycles of -40℃~85℃, 10%~90% RH, with no system crashes or data loss.

[0178] 1.2 Mechanical vibration and shock test: Vehicle vibration, random vibration test from 10Hz to 2000Hz, the connection between the sensor and the main control board is not loose; Mechanical shock, 100g / 6ms impact test, the system data acquisition and decision-making functions are normal.

[0179] 1.3 Electromagnetic Compatibility (EMC) Test: The camera wiring harness is double-shielded with a shielding coverage of >90%; the biosignal acquisition circuit has a common-mode choke with an attenuation of >30dB from 100MHz to 1GHz; it meets the GB / T 18387 and ISO 11452 automotive EMC standards, and there is no false triggering of commands due to electromagnetic interference.

[0180] 2. Core Algorithm Integration Testing (New and Improved): Specialized closed-loop testing will be conducted on the multimodal arbitration algorithm, dynamic scene adaptation, and energy consumption optimization model to verify that the algorithm's functionality and performance meet the standards. This section includes the following: 2.1 Multimodal Arbitration Special Test: The test content includes multi-command conflict resolution, confidence-based decision-making, secondary confirmation, contextual analysis, and adaptive weight update; the test indicators are a multimodal conflict resolution success rate of ≥92.3%, a command false trigger rate of <0.5 times / thousand kilometers, and a weight iteration response time of <100ms; the test results are a fuzzy command context analysis accuracy of ≥89% and a secondary confirmation interaction success rate of ≥95%.

[0181] 2.2 Dynamic Scene Adaptation Special Test: The test content is automatic adaptation to 12 typical scenarios such as rain, strong light, low temperature, low battery, and driver fatigue; the test indicators are scene recognition accuracy ≥98% and scene switching response latency <220ms; the test results are full scene adaptation without omissions or lag, and the control command matching degree with the scene is 100%.

[0182] 2.3 Energy Consumption Optimization Model Special Test: The test content includes the linkage between battery SOC and cabin equipment power, comfort-energy consumption balance, and graded power saving mode; the test indicators are power saving rate ≥35% when SOC<20%, and NEDC range optimization rate ≥7.8%; the test results show that the equipment power adjustment is accurate, the comfort index fluctuation is <5%, and there is no excessive power saving that leads to a decline in experience.

[0183] 3. Hardware / Software Collaboration and Real-Time Testing (New and Improved): Verify the end-to-end data transmission, timing synchronization, and real-time response capabilities from the perception layer to the decision-making layer and the execution layer. This section includes the following: 3.1 Real-time data transmission: CAN FD / Gigabit Ethernet data transmission delay <50ms, camera data return delay <8ms; multi-source data timestamp alignment error <10ms, no data misalignment or frame loss.

[0184] 3.2 End-to-end execution latency: The average latency of the entire link from perception and data acquisition to algorithm decision-making to device execution is 220ms, which meets the real-time requirements of automotive grade.

[0185] 3.3 Multi-task concurrency stability: Simultaneously running multimodal perception, algorithm decision-making, device control, and data storage tasks, it can run continuously for 72 hours without crashes or lag.

[0186] 4. Data Security and Privacy Testing (New and Improved): Security testing is conducted on user biometrics and interaction preference data to ensure compliance with vehicle data security regulations. This section includes the following: 4.1 Biometric data is transmitted and stored using AES-256 encryption, with the key automatically rotating every 30 minutes, ensuring no data leakage.

[0187] 4.2 Sensitive commands (windows, doors) have a 100% pass rate for dual verification, with no illegal operations.

[0188] 4.3 User profile data is encrypted and stored in the TEE secure area, with a 100% unauthorized access interception rate.

[0189] 5. Real-world road testing (improvements and refinements): Building upon existing road tests, the testing was expanded to include multiple road conditions, scenarios, and long mileage, covering urban areas, highways, mountain roads, rainy weather, and low temperatures. A total of 100,000 kilometers of real-vehicle road testing was completed, the results of which are shown in Table 4. The road test conclusions are: the system operates stably under all road conditions and scenarios, with smooth interaction and significant energy-saving effects, meeting the requirements for mass production in vehicles.

[0190] Table 4

[0191] IV. Supplementary Implementation Details

[0192] In this embodiment, this part includes the following: 1. Thermal Management Strategy: When the battery SOC < 30%, a graded power-saving mode is activated; Level 1 power saving turns off the rear entertainment system (power saving of 85W); Level 2 power saving limits the maximum airflow of the air conditioner (power reduction of 40%); Level 3 power saving switches to monocular vision mode (processing power consumption reduction of 35%).

[0193] 2. Human-computer interaction optimization: The steering wheel vibration intensity in the haptic feedback system is linearly related to the vehicle's lane departure angle (0.5g~2.5g acceleration adjustable); at the same time, the seat side wing support airbags dynamically adjust the pressure according to the turning G-value (range 5-25kPa).

[0194] 3. Data security mechanism: Biometric data encryption adopts AES-256+CBC mode, and the key is rotated every 30 minutes; sensitive operations require dual verification. Security commands (such as window control) require a combination of voice and gesture confirmation.

[0195] This embodiment also provides an intelligent cockpit interaction system for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0196] This embodiment also provides an intelligent cockpit interaction system, such as Figure 6 As shown, it includes: The data acquisition module 601 is used to acquire the user's voice signal, gesture signal and bioelectric signal, and calculate the voice confidence, gesture confidence and bioelectric confidence respectively.

[0197] The fusion arbitration module 602 is used to perform weighted calculations on voice confidence, gesture confidence, and bioelectric confidence to obtain an arbitration score.

[0198] The instruction determination module 603 is used to determine the instruction to be executed based on the arbitration score.

[0199] The scene optimization module 604 is used to determine the current interaction scene and, based on the current interaction scene, correct the execution priority of the instruction to be executed to obtain the final control instruction.

[0200] The command issuing module 605 is used to issue final control commands to the corresponding equipment in the cockpit for execution.

[0201] In some alternative implementations, the data acquisition module 601 includes: The first calculation submodule is used to extract acoustic and semantic features from the speech signal and perform feature fusion calculation to obtain the speech confidence level.

[0202] The second calculation submodule is used to extract three-dimensional trajectory features and kinematic features from the gesture signal, and perform feature fusion calculation to obtain the gesture confidence score.

[0203] The third calculation submodule is used to normalize the bioelectric signals and then perform weighted calculations to obtain the bioelectric confidence score; wherein the bioelectric signals include at least one of heart rate signals, heart rate variability signals, skin conductance response signals and grip strength signals.

[0204] In some alternative implementations, the instruction determination module 603 includes: The first instruction determination submodule is used to determine the control instructions corresponding to voice signals, gesture signals, and bioelectric signals when the arbitration score is higher than the second preset threshold.

[0205] The second instruction determination submodule is used to determine the instruction to be executed when the arbitration score is between the first preset threshold and the second preset threshold. The instruction to be executed is the control instruction confirmed by the user from the alternative control instruction set. The alternative control instruction set consists of multiple high-probability cockpit control instructions selected based on the confidence level of each signal, the current interaction scenario filtering, and the user's historical interaction preferences.

[0206] The third instruction determination submodule is used to determine the instruction to be executed when the arbitration score is lower than the first preset threshold. The instruction to be executed is a fuzzy control instruction. The fuzzy control instruction is a control instruction generated after combining historical interaction data, current interaction scenario and user preferences to perform intent association reasoning and ambiguity resolution. The second preset threshold is greater than the first preset threshold.

[0207] In some alternative implementations, the scene optimization module 604 includes: The first scenario determination submodule is used to acquire environmental perception data, user status data, and vehicle operation data.

[0208] The second scenario determination submodule is used to preprocess environmental perception data, user status data, and vehicle operation data respectively, and then extract the external environment features from the environmental perception data, the driver status features from the user status data, and the vehicle operation features from the vehicle operation data.

[0209] The third scenario determination submodule is used to obtain the first weight of the external environment features, the second weight of the driver state features, and the third weight of the vehicle operation features respectively; wherein, the first weight is greater than the second weight, the second weight is greater than the third weight, and the sum of the weights of the first weight, the second weight and the third weight is a preset value.

[0210] The fourth scenario determination submodule is used to multiply the external environment features by the first weight, the driver's state features by the second weight, and the vehicle operation features by the third weight, and then sum them to obtain a comprehensive scenario determination value. Based on the comprehensive scenario determination value, the corresponding current interaction scenario is determined.

[0211] In some optional implementations, the scene optimization module 604 further includes: The first scenario optimization submodule is used to obtain the scenario type of the current interaction scenario. The scenario type includes at least one of the following: rainy day scenario, low battery scenario, fatigue driving scenario, high temperature scenario, low temperature scenario, commuting scenario, highway scenario, and parent-child scenario.

[0212] The second scenario optimization submodule is used to match preset priority correction rules according to scenario type. The priority correction rules are used to adjust the execution priority of control commands for safety, energy consumption control, driving comfort and non-essential entertainment.

[0213] The third scenario optimization submodule is used to adjust the original execution priority of the instruction to be executed according to the priority correction rules to obtain the corrected execution priority.

[0214] The fourth scenario optimization submodule is used to sort, delay, or mask the instructions to be executed according to the corrected execution priority, and generate the final control instructions.

[0215] In some alternative implementations, the system further includes: The energy consumption correction module is used to obtain the remaining battery power, ambient temperature, and user comfort index. The user comfort index is calculated by fusing at least one parameter among seat pressure distribution, skin temperature and humidity, and breathing rate. Based on the remaining battery power, ambient temperature, and user comfort index, the optimal operating power of each device in the cabin is calculated. After correcting the device power parameters in the final control command based on the optimal operating power, the final control command is sent to the corresponding device in the cabin for execution.

[0216] The user feedback module is used to collect user feedback information on executed commands in real time, including confirmation feedback and rejection feedback; based on the feedback information, the weights of each signal used in the weighted calculation are dynamically adjusted, and the control strategy corresponding to the current interaction scenario is optimized.

[0217] In some optional implementations, the energy consumption correction module includes: a correction submodule, used to determine the current temperature difference based on the ambient temperature; determine the power compensation coefficient based on the current temperature difference and the remaining battery power; obtain the initial operating power of each device in the cabin; and adjust each initial operating power according to the power compensation coefficient, the current temperature difference, and the user comfort index to obtain the corresponding optimal operating power.

[0218] In some optional implementations, the user feedback module includes: an adjustment submodule, used to increase the weight of the voice signal and decrease the weight of the gesture signal when the feedback information is confirmation feedback; to increase the weight of the bioelectric signal when the feedback information is rejection feedback; and to normalize the weights of the adjusted signals so that the sum of the weights of the signals is a preset value.

[0219] The intelligent cockpit interaction system provided in this embodiment of the invention can execute the intelligent cockpit interaction method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0220] The intelligent cockpit interaction system in this embodiment of the invention features a data acquisition module that simultaneously collects voice, gesture, and bioelectrical signals and calculates their respective confidence levels. A fusion arbitration module employs a weighted fusion strategy to comprehensively judge user intent, significantly reducing the probability of misidentification and command conflicts compared to single-modality or fixed-priority methods, thus ensuring the reliability of interaction in complex driving environments. Simultaneously, a scene optimization module identifies the current interaction scene in real time and dynamically adjusts the execution priority of commands based on the scene type. This allows the system to distinguish the urgency of safety, energy consumption, and entertainment commands, ensuring that critical commands are responded to first, while unnecessary commands are delayed or blocked. This not only enhances driving safety and user comfort but also supports personalization and continuous optimization, significantly improving driving comfort and user satisfaction while ensuring a higher level of intelligent cockpit interaction.

[0221] Figure 7 This is a schematic diagram of a vehicle structure provided in an embodiment of the present invention. See below for details. Figure 7 The diagram illustrates a structural schematic suitable for implementing a vehicle according to an embodiment of the present invention. The vehicle may include a processor (e.g., a central processing unit, graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory ROM 702 or a program loaded from memory 708 into a random access memory RAM 703. The RAM 703 also stores various programs and data required for vehicle operation. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0222] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows the vehicle to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 Vehicles with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0223] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 709, or installed from a memory 708, or installed from a ROM 702. When the computer program is executed by the processor 701, it performs the functions defined in the intelligent cockpit interaction method of the embodiments of the present invention. Figure 7 The vehicle shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.

[0224] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the intelligent cockpit interaction method shown in the above embodiments is implemented.

[0225] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A smart cockpit interaction method, characterized in that, The method includes: Acquire the user's voice signal, gesture signal, and bioelectrical signal, and calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively; The arbitration score is obtained by weighting the confidence scores of the speech, gestures, and bioelectricity. The instructions to be executed are determined based on the arbitration score; Determine the current interaction scenario, and based on the current interaction scenario, modify the instruction execution priority of the instruction to be executed to obtain the final control instruction; The final control command is issued to the corresponding equipment in the cockpit for execution.

2. The intelligent cockpit interaction method according to claim 1, characterized in that, Before issuing the final control command to the corresponding equipment in the cockpit for execution, the method further includes: The remaining battery power, ambient temperature, and user comfort index are obtained. The user comfort index is calculated based on at least one parameter among seat pressure distribution, skin temperature and humidity, and respiratory rate. Based on the remaining battery power, the ambient temperature, and the user comfort index, calculate the optimal operating power of each device in the cabin; After correcting the device power parameters in the final control command based on the optimal operating power, the final control command is sent to the corresponding device in the cockpit for execution.

3. The intelligent cockpit interaction method according to claim 2, characterized in that, The calculation of the optimal operating power of each device in the cabin based on the remaining battery power, the ambient temperature, and the user comfort index includes: Determine the current temperature difference based on the ambient temperature; Based on the current temperature difference and the remaining battery power, determine the power compensation coefficient; Obtain the initial operating power of each device in the cockpit; Based on the power compensation coefficient, the current temperature difference, and the user comfort index, the initial operating power is adjusted to obtain the corresponding optimal operating power.

4. The intelligent cockpit interaction method according to claim 1, characterized in that, The calculation of speech confidence, gesture confidence, and bioelectric confidence includes: Acoustic and semantic features are extracted from the speech signal, and feature fusion calculation is performed to obtain the speech confidence score. Three-dimensional trajectory features and kinematic features are extracted from the gesture signal, and feature fusion calculation is performed to obtain the gesture confidence score; After normalizing the bioelectric signals, a weighted calculation is performed to obtain the bioelectric confidence level; wherein the bioelectric signals include at least one of heart rate signals, heart rate variability signals, skin conductance response signals, and grip strength signals.

5. The intelligent cockpit interaction method according to claim 1, characterized in that, The determination of the instruction to be executed based on the arbitration score includes: When the arbitration score is higher than the second preset threshold, the instruction to be executed is a control instruction corresponding to a voice signal, a gesture signal, and a bioelectric signal; When the arbitration score is between the first preset threshold and the second preset threshold, the instruction to be executed is a control instruction confirmed by the user from the set of alternative control instructions. The set of alternative control instructions consists of multiple high-probability cockpit control instructions selected based on the confidence level of each signal, the current interaction scenario filtering, and the user's historical interaction preferences. When the arbitration score is lower than the first preset threshold, the instruction to be executed is a fuzzy control instruction. The fuzzy control instruction is a control instruction generated after combining historical interaction data, current interaction scenario and user preferences to perform intent association reasoning and ambiguity resolution. Wherein, the second preset threshold is greater than the first preset threshold.

6. The intelligent cockpit interaction method according to claim 1, characterized in that, Determining the current interaction scenario includes: Acquire environmental perception data, user status data, and vehicle operation data; After preprocessing the environmental perception data, user status data, and vehicle operation data respectively, the external environment features in the environmental perception data, the driver status features in the user status data, and the vehicle operation features in the vehicle operation data are extracted. The first weight of the external environment feature, the second weight of the driver state feature, and the third weight of the vehicle operation feature are obtained respectively; wherein, the first weight is greater than the second weight, the second weight is greater than the third weight, and the sum of the weights of the first weight, the second weight and the third weight is a preset value; The external environment features are multiplied by the first weight, the driver state features are multiplied by the second weight, and the vehicle operation features are multiplied by the third weight, and then summed to obtain a scene comprehensive judgment value. The corresponding current interaction scene is determined based on the scene comprehensive judgment value.

7. The intelligent cockpit interaction method according to any one of claims 1 to 6, characterized in that, The step of modifying the instruction execution priority of the instruction to be executed based on the current interaction scenario to obtain the final control instruction includes: Obtain the scene type of the current interaction scene, which includes at least one of the following: rain scene, low battery scene, fatigue driving scene, high temperature scene, low temperature scene, commuting scene, highway scene, and parent-child scene; According to the scenario type, a preset priority correction rule is matched, and the priority correction rule is used to adjust the execution priority of safety, energy consumption control, driving comfort and non-essential entertainment control commands. The original execution priority of the instruction to be executed is adjusted according to the priority correction rule to obtain the corrected execution priority; The instructions to be executed are sorted, delayed, or masked according to the modified execution priority to generate the final control instructions.

8. The intelligent cockpit interaction method according to claim 1, characterized in that, After issuing the final control command to the corresponding equipment in the cockpit for execution, the method further includes: Real-time collection of user feedback on executed instructions, including confirmation feedback and rejection feedback; The weights of each signal used in the weighted calculation are dynamically adjusted based on the feedback information, and the control strategy corresponding to the current interaction scenario is optimized.

9. The intelligent cockpit interaction method according to claim 8, characterized in that, The step of dynamically adjusting the weights of each signal used in the weighted calculation based on the feedback information includes: When the feedback information is a confirmation feedback, the weight of the voice signal is increased and the weight of the gesture signal is decreased; When the feedback information is a negative feedback, the weight of the bioelectric signal is increased; The weights of each signal after adjustment are normalized so that the sum of the weights of each signal is a preset value.

10. An intelligent cockpit interaction system, characterized in that, The system includes: The data acquisition module is used to acquire the user's voice signal, gesture signal, and bioelectrical signal, and to calculate the voice confidence score, gesture confidence score, and bioelectrical confidence score respectively. The fusion arbitration module is used to perform weighted calculations on the voice confidence, the gesture confidence, and the bioelectric confidence to obtain an arbitration score; The instruction determination module is used to determine the instruction to be executed based on the arbitration score; The scene optimization module is used to determine the current interaction scene and, based on the current interaction scene, modify the instruction execution priority of the instruction to be executed to obtain the final control instruction; The command issuing module is used to issue the final control command to the corresponding equipment in the cockpit for execution.

11. A vehicle, characterized in that, The vehicle includes a controller, which includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the intelligent cockpit interaction method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the intelligent cockpit interaction method according to any one of claims 1 to 9.