A Smart Cockpit Interaction Management Method

By receiving and tagging multimodal interaction data in the intelligent cockpit system, and combining it with vehicle status parameters for filtering and conflict detection, the redundancy and chaos in interaction management in existing technologies are solved, achieving more accurate intent recognition and safer task scheduling.

CN120832023BActive Publication Date: 2025-12-02XIAMEN JINLONG CAR ACCESSORIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511326772.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-02
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing intelligent cockpit systems are prone to problems such as redundant navigation requests and chaotic driving safety management due to improper multimodal interactive data processing and task scheduling during driving.

Method used

By receiving and tagging voice, gesture, and facial expression interaction data, and combining this with vehicle status parameters for correlation filtering, conflict detection, priority classification, and task scheduling, the accuracy and security of interaction intentions are ensured.

Benefits of technology

It effectively avoids accidental triggering and task confusion, improves the traceability of interactive information and system stability, and ensures driving safety and real-time task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832023B_ABST
    Figure CN120832023B_ABST
Patent Text Reader

Abstract

This invention provides an intelligent cockpit interaction management method, relating to the field of data processing technology. The method includes: receiving voice interaction data, gesture interaction data, and facial expression interaction data collected in the cockpit; performing intent parsing to extract triggering conditions and target functions; combining vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters to perform correlation filtering; performing conflict detection to determine and remove interaction intents with abnormal triggering conditions or conflicting target functions as invalid; classifying tasks according to the category of target functions to generate a task candidate queue; confirming tasks in the task candidate queue and generating a valid task queue for tasks that meet the triggering conditions; and performing task scheduling according to the valid task queue, interrupting low-priority tasks and switching execution when a higher-priority task input is detected. This invention improves the accuracy of cockpit interaction management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an intelligent cockpit interaction management method. Background Technology

[0002] Existing smart cockpit systems typically employ multimodal interaction methods to enhance the human-machine interface experience, including voice recognition, touch operation, gesture recognition, and facial recognition. In most solutions, the in-vehicle control unit receives user voice commands via a voice recognition module and combines this with gesture or facial expression information captured by in-vehicle cameras to perform multimodal data fusion, thereby enabling operation of the entertainment system, navigation system, and environmental control modules. For example, a user can use the voice command "play music" in the in-vehicle infotainment system; the system will then combine head orientation or gesture movements to confirm the user's intent and trigger the corresponding function accordingly.

[0003] However, in specific application scenarios, existing technologies are prone to defects in data processing and task scheduling at the interaction management level. When the vehicle is in motion, if the driver and passengers discuss a place name, the system may incorrectly interpret the voice segment as "navigation destination input," thus triggering unnecessary navigation tasks in the information management process. This mis-triggering leads to redundant navigation requests in the background task queue, affecting the system's management and allocation of priority tasks (such as driving safety prompts and vehicle status warnings). Therefore, existing technologies are insufficient in the filtering of interactive data, task determination, and scheduling strategies, especially in the management and decision-making stages involving multi-source information, which can easily cause task flow chaos and fail to meet the requirements of accuracy and real-time performance in interactive information management in driving scenarios. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent cockpit interaction management method, which aims to solve the problems mentioned in the background art.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] A smart cockpit interaction management method, the method comprising:

[0007] It receives voice interaction data, gesture interaction data, and facial expression interaction data collected in the cockpit, and adds source markers and timestamps to each data point to obtain the raw multimodal interaction data;

[0008] Based on the raw data of multimodal interaction, intent parsing is performed to extract triggering conditions and target functions, forming a candidate set of interaction intents;

[0009] Based on the candidate set of interaction intents, combined with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a correlation screening is performed to obtain the first set of interaction intents.

[0010] For the first set of interactive intents, perform conflict detection, and determine and remove interactive intents that have abnormal triggering conditions or conflicting target functions as invalid, thus obtaining the second set of interactive intents.

[0011] Based on the second set of interactive intents, priority is categorized according to the type of target function to generate a task candidate queue;

[0012] The task candidate queue is checked, and tasks that meet the triggering conditions are generated into a valid task queue.

[0013] Task scheduling is performed based on the available task queue. When a higher-priority task input is detected, the lower-priority task is interrupted and execution is switched.

[0014] Preferably, based on the candidate set of interaction intentions, and in conjunction with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a relevance screening is performed to obtain a first set of interaction intentions, including:

[0015] The interaction intent candidate set is continuously compared with the in-vehicle noise level parameter. When the noise level continuously exceeds the preset noise threshold within the preset time window, the corresponding voice interaction data is determined to be invalid and removed, and the first candidate set is generated.

[0016] The duration is determined based on the first candidate set and the driver's gaze area parameters. When the driver's gaze direction continuously deviates from the preset angle range of the central control interface for more than the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, and a second candidate set is generated.

[0017] The second candidate set is further subdivided and judged based on the vehicle's operating status parameters. When the operating status is high-speed straight driving, safety control-related interaction intentions are retained first. When the operating status is low-speed driving, navigation and safety-related interaction intentions are retained at the same time. When the operating status is parked, entertainment-related interaction intentions are allowed to be retained. Finally, the first set of interaction intentions is generated.

[0018] Preferably, for the first set of interactive intents, conflict detection is performed, and interactive intents with abnormal triggering conditions or conflicting target functions are determined to be invalid and removed, resulting in a second set of interactive intents, including:

[0019] The interaction intent is compared with the timestamp of each interaction intent in the first set of interaction intents. If the timestamp is not continuous or exceeds the preset time interval, the interaction intent is determined to be invalid and removed, and a set of interaction intents filtered by time is generated.

[0020] Consistency is determined based on the triggering conditions of different modalities in the time-filtered set of interaction intents. When voice interaction data and gesture interaction data have the same target function but conflicting triggering conditions, interaction intents that match the vehicle operating status parameters are retained first, and interaction intents that do not meet the conditions are removed to generate a set of interaction intents filtered by modality.

[0021] Based on the input source of the modally filtered set of interaction intentions, when there is a target conflict between the driver's input and the passenger's input, the interaction intention of the driver's input is retained first, the interaction intention of the passenger's input is removed, and a second set of interaction intentions is generated.

[0022] Preferably, based on the second set of interactive intents, the task candidate queue is generated by prioritizing the target functions according to their categories, including:

[0023] Based on the target functions of the second set of interactive intents, an initial classification is performed, with interactive intents involving safety prompts and safety controls classified as high priority, interactive intents involving navigation settings and route planning classified as medium priority, and interactive intents involving entertainment playback and information display classified as low priority, thus generating an initial task candidate queue.

[0024] The initial task candidate queue is associated with the vehicle operating status parameters. When the operating status is emergency braking or collision warning, all non-safety-related tasks are downgraded as a whole, and a dynamically adjusted task candidate queue is generated.

[0025] The task candidate queue is dynamically adjusted and arranged sequentially according to priority to form an execution order, resulting in the final task candidate queue.

[0026] Preferably, a continuous comparison is performed based on the candidate set of interaction intentions and the in-vehicle noise level parameters, including:

[0027] Voice interaction data is extracted from the candidate set of interaction intents, and multiple frames are sampled within a preset time window in combination with in-vehicle noise level parameters to generate a noise change sequence.

[0028] Based on the noise change sequence, a continuous threshold is determined. When the noise continuously exceeds the preset noise threshold for a period of time longer than the preset second duration, the corresponding voice interaction data is determined to be invalid and removed, thus obtaining the voice interaction data.

[0029] Based on the voice interaction data, when the time period is less than a preset duration, the corresponding voice interaction data is retained to form a first candidate set.

[0030] Preferably, the duration determination is based on the first candidate set and the driver's gaze region parameters, including:

[0031] Gesture interaction data is extracted from the first candidate set and compared frame by frame with the driver's gaze area parameters to generate gaze deviation records;

[0032] Based on the gaze deviation record, the continuous deviation time is counted. When the continuous deviation time exceeds the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, thus obtaining the gesture interaction data.

[0033] Based on the gesture interaction data, when the continuous deviation time is less than a preset first duration, the corresponding gesture interaction data is retained to form a second candidate set.

[0034] Preferably, the determination is further subdivided based on the second candidate set and vehicle operating status parameters, including:

[0035] Based on the second candidate set, and combined with the vehicle speed, steering angle and braking status in the vehicle operating status parameters, an operating status combination is formed;

[0036] Based on the combination of operating states, the vehicle is determined to be in a state of high-speed straight driving, low-speed driving, or parking. The interaction intentions corresponding to different states are classified and processed to generate a set of interaction intentions subdivided by operating states.

[0037] Based on the set of interaction intents subdivided by operating status, safety control-related interaction intents are retained when driving straight at high speed, navigation-related and safety control-related interaction intents are retained simultaneously when driving at low speed, and entertainment-related interaction intents are retained when parked, forming the first set of interaction intents.

[0038] Preferably, the comparison is performed based on the timestamps of each interaction intent in the first set of interaction intents, including:

[0039] Extract the timestamp of each interaction intent from the first set of interaction intents, calculate the time interval between adjacent interaction intents, and generate time interval data;

[0040] Threshold determination is performed based on time interval data. When the time interval is greater than the preset time interval, the interaction intent is determined to be abnormal and removed, forming a set of interaction intents filtered by interval.

[0041] The time window is determined based on the set of interaction intents filtered by interval. If an interaction intent does not fall within the preset time window, it is determined to be invalid and removed, thus generating a set of interaction intents filtered by time.

[0042] Preferably, consistency determination is performed based on the triggering conditions of different modalities in the time-filtered set of interaction intents, including:

[0043] Based on the time-filtered set of interaction intentions, extract voice interaction data and gesture interaction data to generate trigger condition pairs;

[0044] A consistency comparison is performed based on the triggering conditions. When the target functions are consistent but the triggering conditions are inconsistent, a modal conflict record is generated.

[0045] Based on the modal conflict records and combined with the vehicle operating status parameters, priority is determined. When the vehicle is in a high-speed driving state, voice interaction data is retained, and when the vehicle is in a low-speed driving or parked state, gesture interaction data is retained. Interaction intentions that do not conform to the vehicle status are eliminated, forming a set of interaction intentions that have been modally filtered.

[0046] Preferably, the determination is made based on the input source of the modally filtered set of interaction intents, including:

[0047] Input source tags are extracted from the modally filtered set of interaction intents to generate input source data;

[0048] Based on the input source data, determine whether the interaction intent comes from the driver or the passenger. When the driver input and the passenger input conflict, retain the driver input, determine that the passenger input is invalid and remove it, and generate a set of interaction intents filtered by source.

[0049] Based on the source-filtered set of interaction intents, when the occupant inputs information related to vehicle safety control, the interaction intent enters a pending confirmation state and is not executed directly. Instead, a confirmation request is sent to the driver. After the driver confirms, the interaction intent is added to the source-filtered set of interaction intents to form a second set of interaction intents.

[0050] The above-described solution of the present invention has at least the following beneficial effects:

[0051] First, by adding source tags and timestamps to the voice, gesture, and facial expression interaction data collected in the cockpit, accurate data identification and timing management can be achieved before multimodal data fusion. This mechanism ensures the traceability of each interaction message, thus avoiding recognition errors caused by chaotic data sources or disordered timing in traditional systems. For example, when the driver and passengers issue commands almost simultaneously, the system can distinguish the interacting entities through source tags and determine their order through timestamps, ensuring more accurate subsequent processing.

[0052] Next, intent parsing is performed based on the raw multimodal interaction data, simultaneously extracting triggering conditions and target functions to form a set of candidate interaction intents. Compared to existing technologies, this approach no longer relies on the results of a single modality but integrates multiple interaction information, making the system more robust in determining the user's true intent. For example, when a driver says the voice command "turn on navigation" accompanied by a gesture, the system can more accurately confirm that the driver intends to perform navigation, rather than other tasks, thus reducing the probability of false triggers.

[0053] Secondly, by combining vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters to perform relevance filtering on the candidate set of interaction intentions, erroneous operations caused by environmental interference and invalid input can be significantly reduced. When the noise level exceeds a threshold, the system automatically removes voice commands within that time period; when the driver's gaze direction continuously deviates from the central control screen, the system blocks corresponding gesture operations; and at high speeds, the system prioritizes retaining safety-related intentions. These processing logics enable the system to better meet the safety requirements and interaction habits of driving scenarios, improving the effectiveness of interaction data.

[0054] Furthermore, by performing conflict detection on the first set of interaction intents, contradictions between different modal inputs or conflicts between different user inputs can be effectively resolved. For example, when the target function corresponding to a voice command and a gesture operation is inconsistent, the system will select the interaction intent that is more in line with the driving conditions based on the operating status; when there is a conflict between the driver's and passenger's inputs, the system will prioritize retaining the driver's input, thereby ensuring the centralization of vehicle control. This conflict handling mechanism avoids the problem of chaotic task flow and enhances the stability and security of the system.

[0055] Furthermore, prioritizing the second set of interactive intents based on target function categories and generating a task candidate queue allows the system to have a clear execution order when handling multiple tasks. Safety prompts and control tasks are assigned the highest priority, navigation tasks the medium priority, and entertainment and information display tasks the low priority. This hierarchical strategy ensures that when resources are limited or tasks conflict, the system can prioritize responding to functions most critical to driving safety, preventing low-priority tasks from monopolizing system resources.

[0056] Finally, by confirming tasks in the candidate task queue and generating a valid task queue, and then implementing priority switching in task scheduling, the system can immediately interrupt low-priority tasks when a higher-priority input is detected. This mechanism ensures timely response in emergencies. For example, when the system is playing entertainment, if a collision warning command is detected, the system will immediately interrupt the entertainment task and prioritize the warning, thereby improving driving safety and the real-time performance of task scheduling. Attached Figure Description

[0057] Figure 1 This is a flowchart of an intelligent cockpit interaction management method provided by an embodiment of the present invention. Detailed Implementation

[0058] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0059] like Figure 1 As shown, an embodiment of the present invention proposes an intelligent cockpit interaction management method, the method comprising:

[0060] It receives voice interaction data, gesture interaction data, and facial expression interaction data collected in the cockpit, and adds source markers and timestamps to each data point to obtain the raw multimodal interaction data;

[0061] Based on the raw data of multimodal interaction, intent parsing is performed to extract triggering conditions and target functions, forming a candidate set of interaction intents;

[0062] Based on the candidate set of interaction intents, combined with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a correlation screening is performed to obtain the first set of interaction intents.

[0063] For the first set of interactive intents, perform conflict detection, and determine and remove interactive intents that have abnormal triggering conditions or conflicting target functions as invalid, thus obtaining the second set of interactive intents.

[0064] Based on the second set of interactive intents, the target functions are prioritized and a task candidate queue is generated. Among them, safety-related target functions are set to high priority, navigation target functions are set to medium priority, and entertainment or information display target functions are set to low priority.

[0065] The task candidate queue is checked, and tasks that meet the triggering conditions are generated into a valid task queue.

[0066] Task scheduling is performed based on the available task queue. When a higher-priority task input is detected, the lower-priority task is interrupted and execution is switched.

[0067] In this embodiment of the invention, the proposed intelligent cockpit interaction management method effectively avoids the common problems of false triggering and task confusion in existing methods during the reception and processing of multimodal interaction data. By adding source markers and timestamps to voice, gesture, and facial expression data during the acquisition stage, subsequent processing stages can accurately determine the validity and temporal continuity of the data, thereby improving the traceability and reliability of interaction information. For example, when a driver issues voice commands consecutively within a short period of time, this method can identify the sequential relationship of the commands through timestamps, avoiding mistaking earlier temporary operations for the current control task.

[0068] Next, by parsing the multimodal interaction data, the triggering conditions and target functions can be accurately extracted, forming a set of candidate interaction intentions. This process avoids the over-reliance on single-modal data and improves the comprehensiveness of user intention recognition. For example, when a user says the voice command "play music," this method not only recognizes the voice content but also combines gestures or facial expressions to further confirm the validity of the command, making the results more in line with the actual needs in driving scenarios.

[0069] Secondly, based on the candidate interaction intents, parameters such as vehicle operating status, in-vehicle noise level, and driver gaze area are introduced for relevance screening, significantly reducing erroneous executions caused by interference factors. When the vehicle is traveling at high speed, this method prioritizes identifying and retaining safety-related interaction intents, ensuring that driving safety information receives priority responses. For example, while the vehicle is traveling at high speed, if the driver casually discusses "going to a certain place" with passengers while simultaneously performing air conditioning adjustments by looking at the central control interface, this method will determine that the voice command is interfering and discard it, retaining only air conditioning-related operations.

[0070] Furthermore, by performing conflict detection on the first set of interaction intents, duplication or conflict between target functions can be avoided. This method eliminates intents with abnormal triggering conditions or contradictory target functions, thus forming a more concise and reliable set of interactions. For example, when a user simultaneously uses the voice command "open navigation" and the gesture operation "exit navigation," this method will combine the vehicle's operating status and triggering conditions to make a judgment, prioritizing the operation that matches the current driving state and eliminating invalid commands, thus avoiding confusion in the function execution of this method in the vehicle's infotainment system.

[0071] Furthermore, by prioritizing the second set of interactive intents, a task candidate queue can be generated based on the importance of different functions. Safety-related functions are set to the highest priority, navigation functions to medium priority, and entertainment and information display functions to the lowest priority. This classification mechanism gives the method a clear scheduling order when executing tasks. For example, when the method receives both a "safe distance reminder request" and a "play music" request from the driver, it will prioritize the safety reminder over playing music.

[0072] Finally, in the task confirmation and scheduling stages, the solution of this embodiment ensures that all tasks executed by this method are valid tasks that meet the triggering conditions. When a higher-priority task input is detected, this method automatically interrupts the lower-priority task and switches to execution. This scheduling strategy not only ensures the timeliness of high-priority tasks but also prevents low-priority tasks from interfering with driving safety. For example, if the method detects a collision warning signal input while the driver is viewing vehicle status information, it will immediately interrupt the vehicle information display and prioritize the execution of the collision warning task, thereby significantly improving driving safety and the reliability of the interactive method.

[0073] This includes receiving voice interaction data, gesture interaction data, and facial expression interaction data collected within the cockpit, and adding source markers and timestamps to each data point to obtain the raw multimodal interaction data, specifically including:

[0074] Voice input from the driver and passengers is collected via an in-vehicle microphone module, while gestures and facial expressions are captured using cameras installed inside the vehicle. Before entering the voice recognition module, the voice data is tagged with a source marker indicating whether the voice originated from the driver's or passenger's microphone. Gesture and facial expression data are similarly tagged at the acquisition point to distinguish between driver and passenger actions. All data is automatically timestamped, generated by the vehicle control unit's unified clock to indicate the data acquisition time. This step ensures the consistency of data source and order in subsequent processing stages.

[0075] Specifically, based on the raw data of multimodal interaction, intent parsing is performed to extract triggering conditions and target functions, forming a candidate set of interaction intents, which includes:

[0076] Voice data is converted into text content by the speech recognition module, and then semantically analyzed by the intent recognition model to extract keywords and command types. Gesture data undergoes feature extraction through the action recognition module, such as recognizing finger swipes, fist clenching, and clicking. Facial expression data is extracted using the facial expression recognition module to extract emotion tags, such as gaze, blinking, and nodding. The processing unit fuses data from different modalities and, based on a preset rule base or a trained classification model, determines the corresponding triggering conditions (e.g., user gaze at the central control screen accompanied by voice commands) and the target function (e.g., navigation, playing music, adjusting air conditioning), and uses the determination results as a candidate set of interaction intents.

[0077] In a preferred embodiment of the present invention, based on the candidate set of interaction intentions, and in conjunction with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a relevance screening is performed to obtain a first set of interaction intentions, including:

[0078] The interaction intent candidate set is continuously compared with the in-vehicle noise level parameter. When the noise level continuously exceeds the preset noise threshold within the preset time window, the corresponding voice interaction data is determined to be invalid and removed, and the first candidate set is generated.

[0079] The duration is determined based on the first candidate set and the driver's gaze area parameters. When the driver's gaze direction continuously deviates from the preset angle range of the central control interface for more than the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, and a second candidate set is generated.

[0080] The second candidate set is further subdivided and judged based on the vehicle's operating status parameters. When the operating status is high-speed straight driving, safety control-related interaction intentions are retained first. When the operating status is low-speed driving, navigation and safety-related interaction intentions are retained at the same time. When the operating status is parked, entertainment-related interaction intentions are allowed to be retained. Finally, the first set of interaction intentions is generated.

[0081] In this embodiment of the invention, by introducing multi-dimensional parameters such as vehicle operating status, in-vehicle noise level, and driver gaze area into the candidate set of interaction intentions for correlation filtering, the accuracy and usability of interaction data can be effectively improved. In this way, the method not only relies on raw voice or gesture signals but also combines external environment and driver status for comprehensive judgment, thereby significantly reducing the probability of redundant and false triggering operations.

[0082] Next, when the noise level inside the vehicle is high, this method compares the data against a preset noise threshold and removes invalid voice data. For example, when the vehicle passes through a tunnel or the windows are open, causing the background noise to exceed the threshold, this method will identify that the voice input in that environment may be distorted, and thus automatically block these commands to avoid erroneously triggering navigation or entertainment tasks.

[0083] Secondly, by combining the driver's gaze area parameters to determine the duration, erroneous operations caused by driver inattention can be avoided. When the driver's gaze is detected to be deviating from the central control screen for an extended period, this method will determine that the gesture operation is ineffective and discard it. For example, if the driver inadvertently makes a hand gesture while turning their head to observe road conditions, this method will automatically ignore the input, thereby avoiding erroneous control responses.

[0084] Finally, by further subdividing the judgment based on the vehicle's operating status, this method can flexibly retain different types of interaction intents under different driving conditions. When the vehicle is traveling at high speed, this method only retains safety control-related intents, ensuring the absolute priority of driving safety; when the vehicle is traveling at low speed, this method allows for the simultaneous retention of navigation and safety-related intents to meet diverse driving needs; when the vehicle is parked, this method can appropriately retain entertainment-related intents to improve passenger comfort and entertainment experience. This hierarchical and differentiated approach allows the interaction method to dynamically adapt to driving scenarios, balancing safety and convenience.

[0085] The preset time window refers to multiple consecutive samples of the in-vehicle noise level within a fixed time period to dynamically reflect noise changes. The length of this time window can be set according to the vehicle application scenario, such as 1 second, 2 seconds, or 3 seconds. Within the time window, this method collects noise values ​​at fixed sampling intervals (e.g., once every 100 milliseconds) and stores these noise values ​​in chronological order to form a noise sequence. When this method needs to determine the validity of a segment of voice data, it checks the noise sampling values ​​within the corresponding time period and determines whether it is affected by noise interference based on the continuous performance within the time window. By using the time window mechanism, misjudgments caused by instantaneous noise fluctuations can be avoided.

[0086] The preset noise threshold serves as a baseline for determining the reliability of in-vehicle voice data. This threshold can be determined based on typical in-vehicle noise levels; for example, background noise at idle is typically around 50 decibels, while wind noise at high speeds may exceed 70 decibels. This method can set the noise threshold to a fixed value (e.g., 65 decibels) or use an adaptive approach, dynamically adjusting it based on the average noise level of historical samples plus a correction coefficient. If the noise value continuously sampled within a time window consistently exceeds this threshold, it indicates that the voice signal may be severely interfered with, and the corresponding voice interaction data will be marked as invalid. This setting balances voice recognition accuracy in both static and dynamic environments.

[0087] The preset angle range refers to the maximum allowable deviation angle between the driver's gaze direction and the reference direction of the vehicle's central control interface. This method uses an eye-tracking camera or infrared sensor mounted on the dashboard to detect the angle between the driver's gaze direction and the center point of the central control screen in real time. When this angle is less than or equal to the preset angle range (e.g., ±15° or ±20°), the driver's gaze is considered to be within the effective range; when the angle is greater than this range, the driver's gaze is considered to have deviated from the interaction area. This range can be set through experimental calibration, ensuring that the driver can operate naturally while maintaining their attention focused on the main interaction interface. By setting the angle range, random actions by the driver during non-interactive situations can be filtered out.

[0088] The preset first duration is used to limit the duration of driver gaze deviation. This time threshold can be determined using driving behavior experimental data, such as 500 milliseconds, 1 second, or 2 seconds. When the method detects that the driver's gaze direction continuously exceeds the preset angle range for a cumulative period exceeding this duration, it is determined that the driver has not been paying attention to the interactive interface during that period, and therefore the gesture interaction data corresponding to that period will be deemed invalid and discarded. If the deviation duration does not exceed this duration, it indicates that the driver has only briefly shifted their gaze, such as checking the rearview mirror, and the gesture data is still considered valid. By setting this duration, short-term deviation and long-term failure can be effectively distinguished.

[0089] In a preferred embodiment of the present invention, a conflict detection is performed on the first set of interactive intents, and interactive intents with abnormal triggering conditions or conflicting target functions are determined to be invalid and removed, resulting in a second set of interactive intents, including:

[0090] The interaction intent is compared with the timestamp of each interaction intent in the first set of interaction intents. If the timestamp is not continuous or exceeds the preset time interval, the interaction intent is determined to be invalid and removed, and a set of interaction intents filtered by time is generated.

[0091] Consistency is determined based on the triggering conditions of different modalities in the time-filtered set of interaction intents. When voice interaction data and gesture interaction data have the same target function but conflicting triggering conditions, interaction intents that match the vehicle operating status parameters are retained first, and interaction intents that do not meet the conditions are removed to generate a set of interaction intents filtered by modality.

[0092] Based on the input source of the modally filtered set of interaction intentions, when there is a target conflict between the driver's input and the passenger's input, the interaction intention of the driver's input is retained first, the interaction intention of the passenger's input is removed, and a second set of interaction intentions is generated.

[0093] In this embodiment of the invention, by performing conflict detection on the first set of interactive intents, the reliability and consistency of task selection can be further guaranteed, thereby avoiding functional conflicts or execution anomalies caused by the simultaneous existence of multiple interactive intents.

[0094] Next, by comparing the timestamps of the interaction intent, abnormal commands that are discontinuous or exceed the preset interval can be eliminated. For example, if the method receives an old command that is long overdue after a normal voice operation by the driver, it will remove the old command to prevent it from being confused with the new task, thereby improving the method's ability to control the timeliness of the task.

[0095] Secondly, in the process of determining the consistency of multimodal triggering conditions, this method can ensure that it makes a reasonable choice when there is a conflict between voice and gesture input. For example, when the driver gives the command "turn on navigation" by voice, but the gesture performs the operation of "exit navigation", this method will combine the vehicle's operating status to determine whether to prioritize the voice operation when driving at high speed and the gesture operation when parked, thereby avoiding ambiguity in task execution.

[0096] Finally, in the event of input source conflicts, this method can distinguish between the input intentions of the driver and passengers, and always prioritizes the driver's actions, thereby ensuring that control of the vehicle remains centralized with the driver. For example, when a passenger issues a command to change the vehicle's speed mode, this method automatically blocks that operation to avoid threatening driving safety. This mechanism ensures that the driver has the highest decision-making authority in critical interaction scenarios, thus significantly improving the safety and rationality of this method.

[0097] The preset time interval refers to the time threshold used to determine whether two adjacent interaction intentions belong to the same interaction cycle. This threshold is usually set based on the user's operating habits and interaction response requirements in driving scenarios. For example, when a driver issues voice commands or gestures continuously, the normal interval is often less than 1 or 2 seconds, while commands exceeding this duration are likely to belong to different operation cycles.

[0098] In specific processing, this method extracts the timestamp carried by each interaction intent in the first set of interaction intents and calculates the interval between two adjacent timestamps in turn. When the interval is detected to be greater than a preset time interval threshold, it is determined that the subsequent interaction intent is not continuous with the previous one, and may be a delayed input or an irrelevant historical input. Therefore, the interaction intent is marked as invalid and removed.

[0099] For example, if the preset time interval is 2 seconds, when the driver says "turn on music" and "turn up the volume" twice within 1 second, this method will treat them as consecutive interactive inputs and retain them simultaneously. However, if the driver issues the "turn off music" command after 5 seconds, this method will determine that the command exceeds the reasonable range of consecutive interaction and remove it from the set to prevent delayed operations from interfering with the task being performed.

[0100] By setting and applying preset time intervals, this method can ensure that the retained interaction intentions maintain continuity and logic in the time dimension, providing high-quality data input for subsequent consistency determination and conflict detection.

[0101] In a preferred embodiment of the present invention, a task candidate queue is generated by prioritizing the target functions according to the second set of interactive intents, including:

[0102] Based on the target functions of the second set of interactive intents, an initial classification is performed, with interactive intents involving safety prompts and safety controls classified as high priority, interactive intents involving navigation settings and route planning classified as medium priority, and interactive intents involving entertainment playback and information display classified as low priority, thus generating an initial task candidate queue.

[0103] The initial task candidate queue is associated with the vehicle operating status parameters. When the operating status is emergency braking or collision warning, all non-safety-related tasks are downgraded as a whole, and a dynamically adjusted task candidate queue is generated.

[0104] The task candidate queue is dynamically adjusted and arranged sequentially according to priority to form an execution order, resulting in the final task candidate queue.

[0105] In this embodiment of the invention, by prioritizing the second set of interactive intentions, a task candidate queue that conforms to driving safety logic can be constructed, enabling the method to complete scheduling in an orderly and efficient manner when processing multiple tasks.

[0106] Next, this method first prioritizes interaction intents related to safety prompts and controls, assigning them the highest priority; those related to navigation settings and route planning, the medium priority; and those related to entertainment playback and information display, the lowest priority. For example, when a user simultaneously triggers both a "close enough distance warning" and "play music" actions while driving, this method will prioritize the safety prompt and will not delay the response to the safety task due to the low-priority entertainment action.

[0107] Secondly, by associating the initial task candidate queue with the vehicle's operating status, the priority order of different functions can be dynamically adjusted. When the vehicle is in emergency braking or collision warning mode, this method automatically degrades all non-safety-related tasks, thereby focusing computational and interactive resources on the most critical safety tasks. For example, when the vehicle detects a potential collision risk, even if the user is setting up navigation, this method will prioritize interrupting the navigation task and immediately trigger a collision warning.

[0108] Finally, by sequentially arranging the dynamically adjusted task candidate queue, this method can establish a clear execution order based on priority. This mechanism not only ensures orderly task execution but also improves the flexibility and responsiveness of task switching. For example, when driving at low speeds, this method can handle low-to-medium priority entertainment tasks while executing navigation commands; however, in sudden emergencies, this method will immediately interrupt entertainment operations and quickly switch to high-priority task execution, thus ensuring driving safety while also considering user experience.

[0109] The process involves initial classification based on the target functions of the second set of interactive intents. Interactive intents related to security prompts and controls are classified as high priority, those related to navigation settings and route planning as medium priority, and those related to entertainment playback and information display as low priority. This generates an initial task candidate queue, specifically including:

[0110] Once the second set of interaction intents is formed, the target function corresponding to each interaction intent in the set is read sequentially and compared with a preset function category table. The function category table can be stored in the memory of the vehicle control unit, and the table records the typical function and priority correspondence. For example, collision warning, lane departure warning, and brake assist are defined as safety control functions; destination input, route replanning, and navigation start / exit are defined as navigation functions; audio playback, video information display, and in-vehicle entertainment control are defined as entertainment functions.

[0111] After comparison, the control unit assigns an initial priority label to each interaction intent and writes the tasks with priority labels into the candidate task queue in sequence. At this point, the tasks in the candidate queue are arranged according to the order of input, and the priority labels serve as the basis for subsequent dynamic adjustment and sorting. This classification mechanism ensures that tasks of different functional types are distinguishable during scheduling, providing basic data for subsequent dynamic adjustment and execution order arrangement.

[0112] Specifically, based on the association between the initial task candidate queue and vehicle operating status parameters, when the operating status is emergency braking or collision warning, all non-safety-related tasks are downgraded, generating a dynamically adjusted task candidate queue, which includes:

[0113] After the candidate task queue is generated, the vehicle control unit continuously monitors vehicle operating status parameters, which are provided in real time by speed sensors, braking sensors, and collision warning sensors. When an emergency braking state is detected, or a collision warning signal is issued, the vehicle control unit triggers a priority adjustment procedure. The adjustment procedure first iterates through all tasks in the initial candidate queue to determine whether they are safety-related tasks. If they are safety-related tasks, their high priority is maintained; if they are navigation or entertainment tasks, their priority is uniformly downgraded by one level.

[0114] For example, a navigation task initially classified as medium priority might be downgraded to low priority after adjustment; an entertainment task initially classified as low priority might be marked as lowest priority or suspended after adjustment, and no longer participate in the current scheduling. After processing, the vehicle control unit generates a dynamically adjusted task candidate queue. This queue ensures that in an emergency, all computing and interaction resources are focused on safety tasks, thereby preventing non-critical tasks from interfering with driving safety.

[0115] In a preferred embodiment of the present invention, the continuous comparison between the candidate set of interaction intentions and the in-vehicle noise level parameters includes:

[0116] Voice interaction data is extracted from the candidate set of interaction intents, and multiple frames are sampled within a preset time window in combination with in-vehicle noise level parameters to generate a noise change sequence.

[0117] Based on the noise change sequence, a continuous threshold is determined. When the noise continuously exceeds the preset noise threshold for a period of time longer than the preset second duration, the corresponding voice interaction data is determined to be invalid and removed, thus obtaining the voice interaction data.

[0118] Based on the voice interaction data, when the time period is less than a preset duration, the corresponding voice interaction data is retained to form a first candidate set.

[0119] In this embodiment of the invention, by continuously comparing voice interaction data with in-vehicle noise level parameters, interference from environmental noise can be effectively avoided. This method collects multiple frames of noise samples within a preset time window and generates a noise change sequence, so that the determination of voice input no longer depends on a single instantaneous data point, but is based on continuous analysis over time, thereby improving the stability and accuracy of the determination.

[0120] Next, when the noise level continuously exceeds the threshold within a certain period, this method will determine the voice data for that period as invalid and discard it, thereby avoiding false voice triggering in scenarios such as open windows, wind noise, or road echoes. For example, when a vehicle is traveling at high speed through a tunnel, the voice recognition module may capture irrelevant conversations from the driver. The mechanism of this embodiment can determine that these conversations do not meet the conditions for voice interaction and thus discard the data.

[0121] Secondly, when the period during which noise exceeds the threshold is less than the preset duration, this method retains the voice interaction data, thus ensuring that brief interference does not lead to the loss of valid voice commands. For example, when a driver issues the voice command "turn on navigation" with the window open, although there is wind noise for a short time, the method can still correctly recognize and execute the navigation command because the noise duration is short, thereby improving the reliability and fault tolerance of voice interaction in actual driving.

[0122] Specifically, voice interaction data is extracted from a candidate set of interaction intentions, and multi-frame sampling is performed within a preset time window in conjunction with in-vehicle noise level parameters to generate a noise change sequence, including:

[0123] First, voice-related data entries are extracted from the candidate set of interaction intentions; this data is collected by the vehicle's microphone. Then, within a preset time window (e.g., 1 second or 2 seconds), in-vehicle noise level parameters are acquired at fixed time intervals. Each acquired noise value is considered a sampling point; after collecting multiple sampling points consecutively, a noise sequence over time is formed. This sequence reflects the dynamic changes in noise level within the time window, providing a basis for subsequent validity determination.

[0124] Specifically, based on the noise change sequence, a continuous threshold determination is performed. When the noise continuously exceeds a preset noise threshold for a period longer than a preset second duration, the corresponding voice interaction data is determined to be invalid and discarded, resulting in voice interaction data, which specifically includes:

[0125] After obtaining the noise variation sequence, this method compares the noise value with a preset threshold one by one. If the noise value is found to be higher than the threshold multiple times consecutively, the duration of this continuous interval is recorded. When this duration exceeds a preset duration threshold (e.g., lasting more than two seconds), the voice data collected within this time period is deemed insufficiently valid, marked as invalid, and discarded. If the duration of continuous high noise is less than the preset threshold, the voice data is not directly discarded but proceeds to the next step for further filtering. Through this continuous judgment process, invalid voice input collected in long-term high-noise environments can be effectively filtered out.

[0126] The preset second duration is used to limit the duration threshold of in-vehicle noise interference, and is an important parameter for determining whether voice interaction data is severely affected by environmental noise. Its setting is based on typical vehicle usage scenarios, such as high-speed driving, tunnel echoes, open windows, or multiple passengers talking inside the vehicle.

[0127] Specifically, this method continuously collects noise values ​​within a preset time window and generates a noise change sequence. When a noise value is detected to be continuously higher than a preset noise threshold, the duration of this continuous interval is accumulated. If the duration exceeds a preset second duration, the voice interaction data within this period is determined to be invalid and removed from the candidate set. For example, if the second duration is set to 2 seconds, and the noise level is consistently higher than the threshold within 3 seconds, it is determined that the voice commands within this period may be distorted or cannot be reliably recognized, and therefore are filtered out.

[0128] The preset second duration can be determined through experimental calibration. It is typically adjusted between 0.5 and 3 seconds to strike a balance between "avoiding accidental deletion of valid voice" and "timely removal of invalid voice." In some embodiments, this duration can be dynamically adjusted based on the vehicle's operating status: in high-speed driving scenarios, the preset second duration can be set shorter to quickly block invalid input caused by strong wind noise; in parked scenarios, it can be set longer to ensure that the method retains as much voice interaction data as possible in a stable environment.

[0129] By introducing a preset second duration, this method can accurately control the effectiveness of voice data under continuous noise interference, thereby improving the reliability of voice interaction.

[0130] Specifically, based on the voice interaction data, when the time period is less than a preset duration, the corresponding voice interaction data is retained to form a first candidate set, which includes:

[0131] When the duration of noise exceeding the threshold is less than a preset duration, this method continues to check the validity of the voice interaction data within that time period. The checks may include whether the duration of the voice segment exceeds the minimum recognition duration, whether the voice signal is complete, and whether it contains recognizable command keywords. If the voice interaction data meets these conditions, it is retained as valid data and added to the first candidate set; otherwise, it is deemed invalid and discarded. This ensures that even under short-term noise interference, the driver's valid voice input can be preserved as much as possible.

[0132] In a preferred embodiment of the present invention, the duration determination based on the first candidate set and the driver's gaze region parameters includes:

[0133] Gesture interaction data is extracted from the first candidate set and compared frame by frame with the driver's gaze area parameters to generate gaze deviation records;

[0134] Based on the gaze deviation record, the continuous deviation time is counted. When the continuous deviation time exceeds the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, thus obtaining the gesture interaction data.

[0135] Based on the gesture interaction data, when the continuous deviation time is less than a preset first duration, the corresponding gesture interaction data is retained to form a second candidate set.

[0136] In this embodiment of the invention, by combining the driver's gaze region parameters to determine the duration of gesture interaction data, interference from unintentional actions or distractions by the driver can be effectively avoided. This method generates gaze deviation records by comparing the driver's gaze direction and gesture operations frame by frame, thus linking the validity of the gestures to the driver's visual focus.

[0137] Next, if the driver's gaze deviates from the central control interface for more than a preset time, this method will automatically determine that the corresponding gesture data is invalid and discard it. For example, when the driver turns his head to observe the road conditions behind him during a lane change, if his hand moves unintentionally, this method will ignore the action to avoid misinterpreting it as a control command.

[0138] Secondly, when the driver's gaze deviates from the central control interface only briefly, this method retains the corresponding gesture interaction data, thus ensuring the usability of gesture operations. For example, if the driver quickly glances at the rearview mirror and immediately returns to the central control interface to make a "switch music" gesture, this method can recognize the validity of the action, thereby ensuring the normal execution of entertainment control functions. This mechanism strikes a balance between safe driving and convenient interaction, significantly improving the naturalness and accuracy of human-computer interaction.

[0139] Specifically, gesture interaction data is extracted from the first candidate set and compared frame-by-frame with the driver's gaze region parameters to generate gaze deviation records, including:

[0140] First, interaction data related to gestures is selected from the first candidate set. This gesture data is typically collected by cameras installed in the central control area or near the steering wheel and processed by an action recognition algorithm. Simultaneously, the driver's gaze area parameters are provided in real-time by an eye-tracking camera or infrared sensor. This method correlates the gesture data and gaze area data frame-by-frame on a timeline to determine whether the driver's gaze falls within a preset range on the central control screen or interaction area when performing the gesture. If the gaze deviates from this range, the method marks it as "deviation" in the record; otherwise, it is marked as "valid." Finally, a continuous gaze deviation record is generated for subsequent judgment.

[0141] Specifically, based on the gaze deviation record, the continuous deviation time is counted. When the continuous deviation time exceeds a preset first duration, the corresponding gesture interaction data is determined to be invalid and discarded, resulting in gesture interaction data, which specifically includes:

[0142] After obtaining gaze deviation records, this method performs statistical analysis on the time segments marked as "deviations." The analysis method involves grouping consecutive deviation frames into a single time period and accumulating the duration of that period. When a consecutive deviation exceeds a preset duration threshold (e.g., more than 2 seconds), the gesture interaction data corresponding to that time period is deemed invalid and removed from the candidate set. This method ensures that gestures are only considered invalid when the driver has not gazed at the interaction area for an extended period, thus avoiding misjudgments caused by brief gaze shifts.

[0143] Specifically, based on gesture interaction data, when the continuous deviation time is less than a preset first duration, the corresponding gesture interaction data is retained to form a second candidate set, which includes:

[0144] When statistical results indicate that the gaze deviation time is less than a set threshold, this method considers the driver to have maintained attention to the interaction area for most of the time, and therefore, gesture interaction data related to this time period is considered valid. In this case, the method retains this gesture data and writes it into the second candidate set. This ensures that even if there is a brief gaze shift when the driver performs a gesture (such as quickly checking the rearview mirror or observing the road), the valid operation can still be recognized and retained by the method. In this way, the gesture data in the second candidate set has higher reliability and applicability.

[0145] In a preferred embodiment of the present invention, the process of further subdividing and determining based on the second candidate set and vehicle operating state parameters includes:

[0146] Based on the second candidate set, and combined with the vehicle speed, steering angle and braking status in the vehicle operating status parameters, an operating status combination is formed;

[0147] Based on the combination of operating states, the vehicle is determined to be in a state of high-speed straight driving, low-speed driving, or parking. The interaction intentions corresponding to different states are classified and processed to generate a set of interaction intentions subdivided by operating states.

[0148] Based on the set of interaction intents subdivided by operating status, safety control-related interaction intents are retained when driving straight at high speed, navigation-related and safety control-related interaction intents are retained simultaneously when driving at low speed, and entertainment-related interaction intents are retained when parked, forming the first set of interaction intents.

[0149] In this embodiment of the invention, by further subdividing and determining the second candidate set with vehicle operating state parameters, dynamic adaptation of the interaction intent under different driving conditions can be achieved, making the response of this method more in line with actual driving needs and safety constraints. This method combines vehicle speed, steering angle, and braking status to form an operating state combination, and determines whether the vehicle is currently in a high-speed straight-line, low-speed driving, or parked state.

[0150] Next, when the vehicle is traveling straight at high speed, this method prioritizes retaining safety control-related interaction intents, thereby ensuring that the driver receives rapid and clear safety prompts or control responses in high-speed scenarios. For example, on a highway, if the driver simultaneously issues commands for "air conditioning adjustment" and "distance warning," this method will prioritize responding to the distance warning to ensure driving safety.

[0151] Secondly, when the vehicle is traveling at low speeds, this method retains both navigation-related and safety-related interaction intents to provide more flexible navigation assistance functions while ensuring safety. For example, when driving in congested urban areas, the driver can simultaneously receive forward collision warning information and use voice input to adjust the navigation destination.

[0152] Finally, when the vehicle is parked, this method allows for the retention of entertainment-related interaction intentions, thereby enhancing the user's comfort and entertainment experience when the vehicle is stationary. For example, while waiting in a parking lot, users can play music or browse information using gestures or voice commands, and this method prioritizes ensuring the smooth operation of these functions. This state segmentation mechanism enables contextualized processing of interaction strategies, significantly improving the intelligence level of this method and the user experience.

[0153] Among them, based on the second candidate set, and combined with the vehicle operating state parameters such as vehicle speed, steering angle, and braking state, an operating state combination is formed, specifically including:

[0154] First, vehicle speed, steering angle, and braking status parameters are collected in real time from vehicle sensors. Vehicle speed is provided by a speed sensor, steering angle by a steering wheel angle sensor or wheel angle sensor, and braking status by a brake pedal travel sensor or brake pressure sensor. Then, these three parameters are correlated at the same time to form a parameter set, i.e., an operating state combination. For example, if the vehicle speed is 80 km / h, the steering angle is close to zero, and the braking status is not triggered at a certain moment, the operating state combination can be labeled as "high-speed straight-ahead state." In this way, a corresponding operating state combination can be generated at each moment for subsequent determination.

[0155] Specifically, the system determines whether the vehicle is in a high-speed straight-ahead, low-speed driving, or parked state based on the combination of operating states, and classifies the interaction intentions corresponding to different states to generate a set of interaction intentions subdivided by operating state, including:

[0156] After forming the operating state combination, this method compares the combined value with preset operating state determination rules. These rules can be defined as follows: when the vehicle speed is greater than a preset high-speed threshold and the steering angle is close to zero, and the brakes are not triggered, it is determined to be a high-speed straight-ahead state; when the vehicle speed is less than the high-speed threshold but greater than zero, it is determined to be a low-speed driving state; when the vehicle speed is equal to zero and the brakes are engaged, it is determined to be a parking state. Subsequently, the method filters and marks the interactive intents in the second candidate set according to different operating states. For example, in the high-speed straight-ahead state, only safety control-related interactive intents are retained, while navigation or entertainment-related interactive intents are marked as invalid; in the low-speed driving state, safety-related and navigation-related interactive intents are retained; in the parking state, safety, navigation, and entertainment-related interactive intents are allowed to be retained. Finally, a set of interactive intents subdivided by operating state is obtained.

[0157] Specifically, based on the set of interaction intentions subdivided by operating state, the first set of interaction intentions is formed by retaining safety control-related interaction intentions when driving straight at high speed, retaining both navigation-related and safety control-related interaction intentions when driving at low speed, and retaining entertainment-related interaction intentions when parked. This set includes:

[0158] After obtaining the set of interaction intents segmented by vehicle operating state, this method performs a final filtering based on the determined vehicle operating state. If the operating state is high-speed straight driving, only intents related to safety control are written into the first set of interaction intents, such as distance warning and lane departure warning. If the operating state is low-speed driving, intents related to navigation tasks and safety control are written into the first set of interaction intents, such as route updates and braking prompts. If the operating state is parked, entertainment-related intents are selected from the set and written into the first set of interaction intents, such as playing music and displaying multimedia information. In this way, the resulting first set of interaction intents can be flexibly adjusted according to the differences in vehicle operating state, providing high-quality input data for subsequent conflict detection and task scheduling.

[0159] In a preferred embodiment of the present invention, the comparison is performed based on the timestamps of each interaction intent in the first set of interaction intents, including:

[0160] Extract the timestamp of each interaction intent from the first set of interaction intents, calculate the time interval between adjacent interaction intents, and generate time interval data;

[0161] Threshold determination is performed based on time interval data. When the time interval is greater than the preset time interval, the interaction intent is determined to be abnormal and removed, forming a set of interaction intents filtered by interval.

[0162] The time window is determined based on the set of interaction intents filtered by interval. If an interaction intent does not fall within the preset time window, it is determined to be invalid and removed, thus generating a set of interaction intents filtered by time.

[0163] In this embodiment of the invention, by comparing the timestamps of each interaction intent in the first set of interaction intents, expired instructions or abnormal interval data can be effectively prevented from affecting the correct execution of the task. This method, by calculating the time interval between adjacent interaction intents, can determine whether the interaction input is continuous and reasonable, thereby ensuring the timeliness and consistency of the task.

[0164] Next, when the time interval exceeds a preset time interval, this method will determine the interaction intent as abnormal and remove it, thereby preventing historical residual operation signals from entering the task queue. For example, if the "start navigation" voice command issued by the driver several minutes ago is delayed in recognition due to network latency, and if the command is no longer reasonable in the current driving state, this method will automatically determine it as an invalid command and remove it.

[0165] Secondly, within the set of interaction intents filtered by intervals, this method further performs a unified time window determination to ensure that the retained interaction intents all belong to the same interaction cycle. For example, when the driver consecutively issues the commands "adjust the air conditioning temperature" and "play music," if the previous command exceeds the current time window, this method will discard it, retaining only the command matching the current window. This approach ensures continuity between interaction tasks, improving the accuracy and efficiency of task scheduling.

[0166] Specifically, the process involves extracting the timestamp of each interaction intent from the first set of interaction intents, calculating the time interval between adjacent interaction intents, and generating time interval data, which includes:

[0167] First, the timestamp of each interaction intent is extracted sequentially from the first set of interaction intents. This timestamp is automatically generated by the vehicle control unit during the data acquisition phase and is used to mark the moment the data was collected. Then, the timestamps of two adjacent interaction intents are compared sequentially to calculate the time interval between them. The interval values ​​of all adjacent interaction intents are recorded sequentially, ultimately generating a time interval data table. This data table visually reflects the distribution of each interaction intent along the timeline, providing a basis for subsequent validity determination.

[0168] Specifically, a threshold is determined based on time interval data. When the time interval exceeds a preset time interval, the interaction intent is judged as abnormal and removed, forming a set of interaction intents filtered by time interval, which includes:

[0169] After obtaining the time interval data, this method compares each time interval with a preset threshold. The threshold can be set according to the actual application scenario, such as two seconds or five seconds. When the interval between a given interactive intent and its preceding interactive intent exceeds the threshold, the interactive intent is considered to have exceeded the reasonable range of continuous input, possibly due to historical residual data or delayed input data, and is therefore judged as abnormal and removed. The remaining interactive intents that do not exceed the threshold are retained, forming a set of interval-filtered interactive intents. This step ensures that the retained interactive intents are continuous and related in time.

[0170] Specifically, a time window is determined based on the interval-filtered set of interaction intents. If an interaction intent does not fall within a preset time window, it is deemed invalid and removed, generating a time-filtered set of interaction intents, which includes:

[0171] After interval filtering, this method compares the filtered interaction intents with a preset time window. This time window limits the interaction intents to appearing within a specific time period, such as 1 to 3 seconds. If the timestamp of an interaction intent is outside this unified window, it is considered out of sync with other interaction commands, deemed invalid, and discarded. Only interaction intents falling within the unified time window are retained, ultimately generating a time-filtered set of interaction intents. This step ensures consistency in the time dimension among multimodal interaction inputs, preventing cross-time period interference data from entering subsequent processing flows.

[0172] In a preferred embodiment of the present invention, consistency determination is performed based on the triggering conditions of different modalities in a time-filtered set of interaction intents, including:

[0173] Based on the time-filtered set of interaction intentions, extract voice interaction data and gesture interaction data to generate trigger condition pairs;

[0174] A consistency comparison is performed based on the triggering conditions. When the target functions are consistent but the triggering conditions are inconsistent, a modal conflict record is generated.

[0175] Based on the modal conflict records and combined with the vehicle operating status parameters, priority is determined. When the vehicle is in a high-speed driving state, voice interaction data is retained, and when the vehicle is in a low-speed driving or parked state, gesture interaction data is retained. Interaction intentions that do not conform to the vehicle status are eliminated, forming a set of interaction intentions that have been modally filtered.

[0176] In this embodiment of the invention, by performing consistency determination on the triggering conditions of different modalities in a time-filtered set of interactive intents, contradictions and conflicts between multimodal inputs such as voice and gestures can be effectively avoided. This method, by forming voice and gesture triggering condition pairs, can perform consistency comparisons on different triggering conditions for the same target function, thereby determining whether modal conflicts exist.

[0177] Next, when the target functions are the same but the triggering conditions are different, this method generates a modal conflict record and determines the priority based on the vehicle's operating status. For example, when the driver inputs "turn on navigation" via voice and simultaneously performs a gesture to "exit navigation," this method will select based on the driving conditions: in high-speed driving, voice input will be prioritized to ensure that the driver's instructions are not interfered with by gestures; while in parked driving, gesture operations will be prioritized to accommodate the user's more natural interaction method.

[0178] Secondly, by eliminating interaction intentions that are inconsistent with the vehicle's status, this method maintains the consistency and rationality of interaction results. For example, when a driver simultaneously says "play music" and makes a "turn off music" gesture while driving at low speed, this method will prioritize selecting the more suitable operation based on the vehicle's operating status, avoiding interference with navigation or safety tasks. This mechanism ensures the harmonious unity of multimodal interactions, thereby improving the accuracy and stability of human-computer interaction.

[0179] Specifically, voice interaction data and gesture interaction data are extracted from a time-filtered set of interaction intentions to generate trigger condition pairs, including:

[0180] First, voice interaction data and gesture interaction data are extracted from the time-filtered set of interaction intents. Voice interaction data typically contains the recognized voice command text and its corresponding target function information, while gesture interaction data contains the recognized gesture actions and their mapped target function information. Then, the voice interaction data and gesture interaction data are paired according to their timestamps. If they belong to the same interaction period in time, they are combined into a trigger condition pair. Each trigger condition pair contains trigger conditions for the same target function in both modalities, providing input for subsequent consistency comparison.

[0181] Specifically, a consistency comparison is performed based on the triggering conditions. When the target functions are consistent but the triggering conditions are inconsistent, a modal conflict record is generated, which includes:

[0182] After generating the trigger condition pairs, this method compares the target functions expressed by the voice modality and the gesture modality one by one. When it detects that the two modalities point to the same target function, but their trigger conditions are contradictory, a modal conflict record is generated. For example, if the voice modal command is "start navigation," but the gesture modal command corresponds to the action "exit navigation," then it is determined that there is a trigger condition conflict between the two. In this case, this method will mark the conflict type in the record (such as "navigation start - navigation exit conflict") and save the conflict record for subsequent priority determination.

[0183] Specifically, based on modal conflict records and vehicle operating status parameters, priority is determined. Voice interaction data is retained when the vehicle is traveling at high speed, while gesture interaction data is retained when the vehicle is traveling at low speed or parked. Interaction intentions that do not conform to the vehicle's status are discarded, forming a modally filtered set of interaction intentions, which specifically includes:

[0184] After generating modal conflict records, this method prioritizes data based on vehicle operating state parameters, provided by vehicle speed, steering angle, and braking signals. When the vehicle is traveling at high speed, voice interaction data is prioritized because voice interaction is more suitable than gestures for drivers in high-speed scenarios. When the vehicle is traveling at low speed or parked, gesture interaction data is prioritized because gestures are safer and more natural. While prioritizing, interaction intentions that do not conform to the operating state are marked as invalid and discarded. Finally, this method rewrites the retained interaction intentions into a set, forming a modally filtered set of interaction intentions for subsequent processing.

[0185] In a preferred embodiment of the present invention, the determination is made based on the input source of the modally filtered set of interaction intents, including:

[0186] Input source tags are extracted from the modally filtered set of interaction intents to generate input source data;

[0187] Based on the input source data, determine whether the interaction intent comes from the driver or the passenger. When the driver input and the passenger input conflict, retain the driver input, determine that the passenger input is invalid and remove it, and generate a set of interaction intents filtered by source.

[0188] Based on the source-filtered set of interaction intents, when the occupant inputs information related to vehicle safety control, the interaction intent enters a pending confirmation state and is not executed directly. Instead, a confirmation request is sent to the driver. After the driver confirms, the interaction intent is added to the source-filtered set of interaction intents to form a second set of interaction intents.

[0189] In this embodiment of the invention, by determining the input source of the modally filtered set of interaction intentions, the interaction intentions of the driver and passengers can be effectively distinguished, thereby ensuring that key operational authority is always concentrated in the hands of the driver. This method identifies the source of interaction data by extracting input source markers, and prioritizes retaining driver input and discarding passenger input when conflicts occur, in order to avoid passenger misoperation affecting driving safety.

[0190] Next, when both the driver's and passenger's input commands are applied to the same target function, this method will automatically select the driver's command based on priority. For example, when a passenger issues a command to "change navigation destination," and the driver has already set the current navigation route via voice, this method will prioritize retaining the driver's settings, thereby ensuring that the vehicle navigation does not deviate due to passenger interference.

[0191] Secondly, when occupant input involves vehicle safety control functions, this method places the input in a pending confirmation state and sends a confirmation request to the driver. This mechanism prevents arbitrary occupant operations from directly affecting vehicle safety. For example, when an occupant attempts to issue a command to "turn off collision warning," the method will not execute it immediately but will prompt the driver for confirmation. Only after the driver's confirmation will the input be added to the interaction intent set. This design maximizes driving safety while preserving the possibility of occupant participation in specific situations.

[0192] Specifically, input source tags are extracted from the modally filtered set of interaction intents to generate input source data, which includes:

[0193] First, the input source tag for each interaction intent is extracted from the modally filtered set of interaction intents. This tag is generated during the interaction data acquisition phase. For example, voice data can be distinguished from driver and passenger input using multiple microphone arrays in the vehicle, while gesture and facial expression data can be determined from the driver's or passenger's side using camera position. After extraction, these input source tags are converted into input source data, which contains at least two types of labels: driver input and passenger input. In this way, this method can provide clear data basis for subsequent input conflict determination.

[0194] Specifically, the system determines whether the interaction intent originates from the driver or a passenger based on the input source data. When a conflict arises between driver and passenger input, the driver's input is retained, while the passenger's input is deemed invalid and discarded, generating a set of interaction intents filtered by source. This set specifically includes:

[0195] After generating the input source data, this method compares the target function corresponding to the driver's input and the passenger's input one by one. When a conflict is detected, i.e., different input commands exist for the same target function simultaneously, the method prioritizes and retains the driver's input, marking the passenger's input as invalid and discarding it. For example, if the driver inputs "turn on navigation" via voice while the passenger simultaneously inputs "turn off navigation" via touch, the method will retain the driver's "turn on navigation" intention and discard the passenger's "turn off navigation" intention. Through this step, a set of interaction intentions filtered by source can be formed, ensuring the driver's dominant position in vehicle operation.

[0196] Specifically, based on the source-filtered set of interaction intents, when the occupant input involves vehicle safety control, the interaction intent enters a pending confirmation state, is not executed directly, and a confirmation request is sent to the driver; after the driver confirms, the interaction intent is added to the source-filtered set of interaction intents, forming a second set of interaction intents, which specifically includes:

[0197] After determining the conflict between the driver and occupants, this method further examines the set of interaction intents filtered by the source. If the set contains occupant input that relates to vehicle safety control functions, such as disabling the collision warning or modifying the safe distance threshold, the input will not directly enter the execution process but will be placed in a pending confirmation state. This method will issue a confirmation request to the driver via voice broadcast or central control interface prompts. For example, this method may prompt: "The occupant requests to disable the collision warning; please confirm whether to execute." If the driver confirms within the set confirmation time, the occupant input is added to the set and becomes a valid interaction intent; if the driver does not confirm or explicitly refuses, the input will be permanently removed. Finally, after being processed by this confirmation mechanism, the set forms a second set of interaction intents, ensuring that safety-related tasks are always ultimately under the driver's control.

[0198] Embodiments of the present invention also provide an intelligent cockpit interaction management system, the system comprising:

[0199] The data acquisition module is used to receive voice interaction data, gesture interaction data and facial expression interaction data collected in the cockpit, and add source markers and timestamps to each data point to obtain the raw multimodal interaction data;

[0200] The intent parsing module is used to parse the original multimodal interaction data, extract the triggering conditions and target functions, and form a candidate set of interaction intents.

[0201] The relevance filtering module is used to perform relevance filtering based on the candidate set of interaction intents, combined with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, to obtain the first set of interaction intents.

[0202] The conflict detection module is used to perform conflict detection on the first set of interactive intents, and to determine and remove interactive intents that have abnormal triggering conditions or conflict with the target function as invalid, so as to obtain the second set of interactive intents.

[0203] The priority classification module is used to classify the target functions according to their priority based on the second set of interactive intents, and generate a task candidate queue. Among them, safety-related target functions are set as high priority, navigation target functions are set as medium priority, and entertainment or information display target functions are set as low priority.

[0204] The task confirmation module is used to confirm the tasks in the task candidate queue and generate a valid task queue from the tasks that meet the triggering conditions.

[0205] The task scheduling module is used to perform task scheduling based on the valid task queue. When a higher priority task input is detected, the lower priority task is interrupted and the execution is switched.

[0206] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0207] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0208] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0209] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for intelligent cockpit interaction management, characterized in that, The method includes: It receives voice interaction data, gesture interaction data, and facial expression interaction data collected in the cockpit, and adds source markers and timestamps to each data point to obtain the raw multimodal interaction data; Based on the raw data of multimodal interaction, intent parsing is performed to extract triggering conditions and target functions, forming a candidate set of interaction intents; Based on the candidate set of interaction intents, combined with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a correlation screening is performed to obtain the first set of interaction intents. For the first set of interactive intents, perform conflict detection, and determine and remove interactive intents that have abnormal triggering conditions or conflicting target functions as invalid, thus obtaining the second set of interactive intents. Based on the second set of interactive intents, priority is categorized according to the type of target function to generate a task candidate queue; The task candidate queue is checked, and tasks that meet the triggering conditions are generated into a valid task queue. Task scheduling is performed based on the available task queue. When a higher-priority task input is detected, the lower-priority task is interrupted and execution is switched. For the first set of interactive intents, conflict detection is performed. Interactive intents with abnormal triggering conditions or conflicting target functions are determined to be invalid and removed, resulting in the second set of interactive intents, including: The interaction intent is compared with the timestamp of each interaction intent in the first set of interaction intents. If the timestamp is not continuous or exceeds the preset time interval, the interaction intent is determined to be invalid and removed, and a set of interaction intents filtered by time is generated. Consistency is determined based on the triggering conditions of different modalities in the time-filtered set of interaction intents. When voice interaction data and gesture interaction data have the same target function but conflicting triggering conditions, interaction intents that match the vehicle operating status parameters are retained first, and interaction intents that do not meet the conditions are removed to generate a set of interaction intents filtered by modality. Based on the input source of the modally filtered set of interaction intentions, when there is a target conflict between the driver's input and the passenger's input, the interaction intention of the driver's input is retained first, the interaction intention of the passenger's input is removed, and a second set of interaction intentions is generated.

2. The intelligent cockpit interaction management method according to claim 1, characterized in that, Based on the candidate set of interaction intentions, combined with vehicle operating status parameters, in-vehicle noise level parameters, and driver gaze area parameters, a relevance screening is performed to obtain the first set of interaction intentions, including: The interaction intent candidate set is continuously compared with the in-vehicle noise level parameter. When the noise level continuously exceeds the preset noise threshold within the preset time window, the corresponding voice interaction data is determined to be invalid and removed, and the first candidate set is generated. The duration is determined based on the first candidate set and the driver's gaze area parameters. When the driver's gaze direction continuously deviates from the preset angle range of the central control interface for more than the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, and a second candidate set is generated. The second candidate set is further subdivided and judged based on the vehicle's operating status parameters. When the operating status is high-speed straight driving, safety control-related interaction intentions are retained first. When the operating status is low-speed driving, navigation and safety-related interaction intentions are retained at the same time. When the operating status is parked, entertainment-related interaction intentions are allowed to be retained. Finally, the first set of interaction intentions is generated.

3. The intelligent cockpit interaction management method according to claim 1, characterized in that, Based on the second set of interactive intents, tasks are prioritized according to the category of target function, and a candidate queue is generated, including: Based on the target functions of the second set of interactive intents, an initial classification is performed, with interactive intents involving safety prompts and safety controls classified as high priority, interactive intents involving navigation settings and route planning classified as medium priority, and interactive intents involving entertainment playback and information display classified as low priority, thus generating an initial task candidate queue. The initial task candidate queue is associated with the vehicle operating status parameters. When the operating status is emergency braking or collision warning, all non-safety-related tasks are downgraded as a whole, and a dynamically adjusted task candidate queue is generated. The task candidate queue is dynamically adjusted and arranged sequentially according to priority to form an execution order, resulting in the final task candidate queue.

4. The intelligent cockpit interaction management method according to claim 2, characterized in that, A continuous comparison is performed based on the candidate set of interaction intentions and the in-vehicle noise level parameters, including: Voice interaction data is extracted from the candidate set of interaction intents, and multiple frames are sampled within a preset time window in combination with in-vehicle noise level parameters to generate a noise change sequence. Based on the noise change sequence, a continuous threshold is determined. When the noise continuously exceeds the preset noise threshold for a period of time longer than the preset second duration, the corresponding voice interaction data is determined to be invalid and removed, thus obtaining the voice interaction data. Based on the voice interaction data, when the time period is less than a preset duration, the corresponding voice interaction data is retained to form a first candidate set.

5. The intelligent cockpit interaction management method according to claim 2, characterized in that, Duration determination is based on the first candidate set and the driver's gaze region parameters, including: Gesture interaction data is extracted from the first candidate set and compared frame by frame with the driver's gaze area parameters to generate gaze deviation records; Based on the gaze deviation record, the continuous deviation time is counted. When the continuous deviation time exceeds the preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, thus obtaining the gesture interaction data. Based on the gesture interaction data, when the continuous deviation time is less than a preset first duration, the corresponding gesture interaction data is retained to form a second candidate set.

6. The intelligent cockpit interaction management method according to claim 2, characterized in that, The determination is further subdivided based on the second candidate set and vehicle operating status parameters, including: Based on the second candidate set, and combined with the vehicle speed, steering angle and braking status in the vehicle operating status parameters, an operating status combination is formed; Based on the combination of operating states, the vehicle is determined to be in a state of high-speed straight driving, low-speed driving, or parking. The interaction intentions corresponding to different states are classified and processed to generate a set of interaction intentions subdivided by operating states. Based on the set of interaction intents subdivided by operating status, safety control-related interaction intents are retained when driving straight at high speed, navigation-related and safety control-related interaction intents are retained simultaneously when driving at low speed, and entertainment-related interaction intents are retained when parked, forming the first set of interaction intents.

7. The intelligent cockpit interaction management method according to claim 1, characterized in that, The comparison is performed based on the timestamps of each interaction intent in the first set of interaction intents, including: Extract the timestamp of each interaction intent from the first set of interaction intents, calculate the time interval between adjacent interaction intents, and generate time interval data; Threshold determination is performed based on time interval data. When the time interval is greater than the preset time interval, the interaction intent is determined to be abnormal and removed, forming a set of interaction intents filtered by interval. The time window is determined based on the set of interaction intents filtered by interval. If an interaction intent does not fall within the preset time window, it is determined to be invalid and removed, thus generating a set of interaction intents filtered by time.

8. The intelligent cockpit interaction management method according to claim 1, characterized in that, Consistency is determined based on the triggering conditions of different modalities within a time-filtered set of interaction intents, including: Based on the time-filtered set of interaction intentions, extract voice interaction data and gesture interaction data to generate trigger condition pairs; A consistency comparison is performed based on the triggering conditions. When the target functions are consistent but the triggering conditions are inconsistent, a modal conflict record is generated. Based on the modal conflict records and combined with the vehicle operating status parameters, priority is determined. When the vehicle is in a high-speed driving state, voice interaction data is retained, and when the vehicle is in a low-speed driving or parked state, gesture interaction data is retained. Interaction intentions that do not conform to the vehicle status are eliminated, forming a set of interaction intentions that have been modally filtered.

9. The intelligent cockpit interaction management method according to claim 1, characterized in that, The determination is based on the input source of the modally filtered set of interaction intents, including: Input source tags are extracted from the modally filtered set of interaction intents to generate input source data; Based on the input source data, determine whether the interaction intent comes from the driver or the passenger. When the driver input and the passenger input conflict, retain the driver input, determine that the passenger input is invalid and remove it, and generate a set of interaction intents filtered by source. Based on the source-filtered set of interaction intents, when the occupant inputs information related to vehicle safety control, the interaction intent enters a pending confirmation state and is not executed directly. Instead, a confirmation request is sent to the driver. After the driver confirms, the interaction intent is added to the source-filtered set of interaction intents to form a second set of interaction intents.

Citation Information

Patent Citations

  • A multi-mode depth fusion airborne cockpit man-machine interaction method

    CN109933272A

  • Interaction method and device, electronic equipment and computer readable storage medium

    CN120596202A