Intelligent cabin interaction management method

By tagging and timestamping voice, gesture, and facial expression data in the intelligent cockpit system, and combining vehicle status and driver parameters to filter conflicts, the problems of false triggering and task confusion in multimodal interaction data processing are resolved, thereby improving the accuracy of interaction management and driving safety.

CN120832023AActive Publication Date: 2025-10-24XIAMEN JINLONG CAR ACCESSORIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511326772.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-10-24
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing intelligent cockpit systems are prone to false triggering and task confusion during driving due to improper multimodal interactive data processing and task scheduling, which affects driving safety and the accuracy and real-time performance of information management.

Method used

By receiving and tagging voice, gesture, and facial expression data with source and timestamp, and combining vehicle status, noise level, and driver gaze area parameters, the system performs relevance filtering and conflict detection, priority classification, and task scheduling to ensure the accuracy of interaction intent and the efficient execution of tasks.

Benefits of technology

It improves the traceability and robustness of multimodal interaction data, reduces environmental interference and user input errors, and ensures driving safety and the stability and real-time performance of task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832023A_ABST
    Figure CN120832023A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent cabin interaction management method, and relates to the technical field of data processing, and the method comprises the steps: receiving voice interaction data, gesture interaction data and expression interaction data collected in a cabin; performing intention analysis, and extracting a trigger condition and a target function; in combination with the vehicle running state parameters, the in-vehicle noise level parameters and the driver watching area parameters, correlation screening is executed; executing conflict detection, and judging the interaction intentions with abnormal triggering conditions or target function conflicts as invalid and removing the interaction intentions; performing priority classification according to the category of the target function to generate a task candidate queue; performing task confirmation on the task candidate queue, and generating an effective task queue from the tasks meeting the trigger condition; task scheduling is executed according to the effective task queue, and when higher-priority task input is detected, low-priority tasks are interrupted and execution is switched; according to the invention, the accuracy of cabin interaction management is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an intelligent cockpit interaction management method. BACKGROUND

[0002] The existing intelligent cockpit system usually adopts a multi-modal interaction mode to improve the human-computer interaction experience, including voice recognition, touch operation, gesture recognition, and facial recognition, etc. In most solutions, the vehicle-mounted control unit receives the user's voice instructions through the voice recognition module, and at the same time combines the gesture or expression information collected by the in-vehicle camera to perform multi-modal data fusion, so as to realize the operation of the entertainment system, navigation system, and environmental control module. For example, the user can play music through the voice instruction "play music" in the car system, and the system will combine the head orientation or gesture action to confirm the user's intention, and accordingly trigger the corresponding function.

[0003] However, in specific application scenarios, the existing technology is prone to defects in data processing and task scheduling in the process of interaction management. When the vehicle is in the process of driving, if the driver and the passenger talk about a certain place name, the system may incorrectly analyze the voice segment as "navigation destination input", thereby triggering unnecessary navigation tasks in the information management process. This mis-triggering will cause redundant navigation requests in the background task queue, affecting the management and allocation of priority tasks (such as driving safety prompts and vehicle state warnings) by the system. As can be seen, the existing technology has deficiencies in the filtering of interaction data, the determination of tasks, and the scheduling strategy, especially in the management and decision-making links involving multi-source information, which is prone to cause task flow confusion, and does not meet the accuracy and real-time requirements of interaction information management in the driving scenario. SUMMARY

[0004] The purpose of the present application is to provide an intelligent cockpit interaction management method to solve the problems mentioned in the background art.

[0005] To solve the above technical problems, the technical solution of the present application is as follows: An intelligent cockpit interaction management method, the method comprising: receiving voice interaction data, gesture interaction data, and expression interaction data collected in the cockpit, and adding a source mark and a timestamp to each data to obtain multi-modal interaction original data; According to the multi-modal interaction original data, performing intention analysis, extracting trigger conditions and target functions, and forming an interaction intention candidate set; According to the interaction intention candidate set, combining vehicle running state parameters, in-vehicle noise level parameters, and driver gaze area parameters, performing relevance screening to obtain a first interaction intention set; For the first interaction intent set, conflict detection is performed, and interaction intents that trigger abnormal conditions or target function conflicts are determined to be invalid and removed to obtain a second interaction intent set; According to the second interaction intent set, the target functions are prioritized according to their categories to generate a task candidate queue; Task confirmation is performed on the task candidate queue, and tasks that meet the trigger conditions are generated to form an effective task queue; According to the effective task queue, task scheduling is performed, and when a higher priority task input is detected, the low priority task is interrupted and execution is switched.

[0006] Preferably, according to the interaction intent candidate set, in combination with the vehicle operating state parameter, the in-vehicle noise level parameter, and the driver's gaze area parameter, relevance screening is performed to obtain the first interaction intent set, including: According to the interaction intent candidate set and the in-vehicle noise level parameter, when the noise level continuously exceeds the preset noise threshold within the preset time window, the corresponding voice interaction data is determined to be invalid and removed to generate the first candidate set; According to the first candidate set and the driver's gaze area parameter, duration determination is performed, and when the driver's gaze direction continuously remains within a preset angle range away from the central control interaction interface for more than a preset first time period, the corresponding gesture interaction data is determined to be invalid and removed to generate the second candidate set; According to the second candidate set and the vehicle operating state parameter, subdivision determination is performed, and when the operating state is high-speed straight driving, safety control-related interaction intents are preferentially retained, when the operating state is low-speed driving, navigation and safety-related interaction intents are simultaneously retained, and when the operating state is parking, entertainment-related interaction intents are allowed to be retained, to finally generate the first interaction intent set.

[0007] Preferably, for the first interaction intent set, conflict detection is performed, and interaction intents that trigger abnormal conditions or target function conflicts are determined to be invalid and removed to obtain a second interaction intent set, including: According to the timestamps of the interaction intents in the first interaction intent set, when the timestamps are not continuous or exceed the preset time interval, the interaction intents are determined to be invalid and removed to generate an interaction intent set filtered by time; According to the trigger conditions of different modalities in the interaction intent set filtered by time, when the voice interaction data and the gesture interaction data have consistent target functions but conflicting trigger conditions, the interaction intent that matches the vehicle operating state parameter is preferentially retained, and the interaction intent that does not meet the condition is removed to generate an interaction intent set filtered by modality; According to the input source of the modal filtered interaction intention set, when the driver input and the passenger input have target conflicts, the interaction intention of the driver input is preferentially retained, and the interaction intention of the passenger input is removed, to generate a second interaction intention set.

[0008] Preferably, according to the second interaction intention set, the target functions are classified in priority, to generate a task candidate queue, including: According to the initial classification of the target functions of the second interaction intention set, the interaction intention related to safety prompts and safety control is divided into high priority, the interaction intention related to navigation setting and path planning is divided into medium priority, and the interaction intention related to entertainment playing and information display is divided into low priority, to generate an initial task candidate queue; According to the association of the initial task candidate queue and the vehicle running state parameter, when the running state is emergency braking or collision warning, all non-safety related tasks are downgraded as a whole, to generate a dynamically adjusted task candidate queue; According to the sequential arrangement of the dynamically adjusted task candidate queue, the execution order is formed according to the priority, to obtain a final task candidate queue.

[0009] Preferably, according to the continuous comparison of the interaction intention candidate set and the in-vehicle noise level parameter, including: According to the interaction intention candidate set, voice interaction data is extracted, and multi-frame sampling is performed in a preset time window in combination with the in-vehicle noise level parameter, to generate a noise change sequence; According to the continuous threshold determination of the noise change sequence, when the time period during which the noise continuously exceeds the preset noise threshold is greater than a preset second time length, the corresponding voice interaction data is determined as invalid and removed, to obtain voice interaction data; According to the voice interaction data, when the time period is less than the preset time length, the corresponding voice interaction data is retained, to form a first candidate set.

[0010] Preferably, according to the duration determination of the first candidate set and the driver gaze area parameter, including: According to the first candidate set, gesture interaction data is extracted, and frame-by-frame comparison is performed with the driver gaze area parameter, to generate a gaze deviation record; According to the gaze deviation record, continuous deviation time is counted, and when the continuous deviation time exceeds a preset first time length, the corresponding gesture interaction data is determined as invalid and removed, to obtain gesture interaction data; According to the gesture interaction data, when the continuous deviation time is less than the preset first time length, the corresponding gesture interaction data is retained, to form a second candidate set.

[0011] Preferably, according to the subdivision determination of the second candidate set and the vehicle running state parameter, including: According to the second candidate set, in combination with the vehicle running state parameters of vehicle speed, steering angle and braking state, a running state combination is formed; According to the running state combination, it is determined that the vehicle is in a high-speed straight driving, low-speed driving or parking state, and the interactive intentions corresponding to different states are classified and processed to generate an interactive intention set subdivided by running states; According to the interactive intention set subdivided by running states, the safety control related interactive intention is retained when driving at high speed, the navigation related and safety control related interactive intentions are retained when driving at low speed, and the entertainment related interactive intention is retained when parking, forming a first interactive intention set.

[0012] Preferably, the timestamps of each interactive intention in the first interactive intention set are compared, including: According to the first interactive intention set, the timestamp of each interactive intention is extracted, the time interval of adjacent interactive intentions is calculated, and time interval data is generated; According to the time interval data, threshold determination is performed, when the time interval is greater than the preset time interval, the interactive intention is determined as abnormal and is removed, forming an interactive intention set filtered by interval; According to the interactive intention set filtered by interval, time window determination is performed, when the interactive intention does not fall into the preset time window, the interactive intention is determined as invalid and is removed, generating an interactive intention set filtered by time.

[0013] Preferably, consistency determination is performed according to the trigger conditions of different modalities in the interactive intention set filtered by time, including: According to the interactive intention set filtered by time, voice interaction data and gesture interaction data are extracted, and trigger condition pairs are generated; According to the trigger condition pairs, consistency comparison is performed, when the target function is consistent but the trigger condition is inconsistent, a modality conflict record is generated; According to the modality conflict record, priority determination is performed in combination with the vehicle running state parameters, when the vehicle is in a high-speed driving state, voice interaction data is retained, when the vehicle is in a low-speed driving or parking state, gesture interaction data is retained, and interactive intentions that do not conform to the vehicle state are removed, forming an interactive intention set filtered by modality.

[0014] Preferably, determination is performed according to the input source of the interactive intention set filtered by modality, including: According to the interactive intention set filtered by modality, input source markers are extracted, and input source data is generated; According to the input source data, it is determined whether the interactive intention comes from the driver or the passenger, when the driver input and the passenger input conflict, the driver input is retained, the passenger input is determined as invalid and is removed, and an interactive intention set filtered by source is generated; According to the set of interaction intents screened by the source, when the occupant inputs related to the vehicle safety control, the interaction intent enters a to-be-confirmed state, is not directly executed, and a confirmation request is issued to the driver; after the driver confirms, the interaction intent is added to the set of interaction intents screened by the source to form a second set of interaction intents.

[0015] The above scheme of the present application at least includes the following beneficial effects: Firstly, by adding source marks and time stamps to the voice, gesture and expression interaction data collected in the cabin, accurate identification and timing management of the data can be realized before multi-modal data fusion. This mechanism ensures that each piece of interaction information is traceable, thereby avoiding recognition errors caused by chaotic data sources or timing errors in traditional systems. For example, when the driver and the occupant almost simultaneously issue instructions, the system can distinguish the interaction subjects through the source marks and determine their sequence through the time stamps, ensuring more accurate subsequent processing.

[0016] Then, on the basis of the multi-modal interaction original data, intent analysis can be performed to extract the trigger conditions and target functions simultaneously, forming a candidate set of interaction intents. Compared with the prior art, this way no longer relies on the results of a single mode, but integrates multiple interaction information, thereby making the system more robust in judging the user's true intention. For example, when the driver speaks a voice instruction of "turn on navigation" accompanied by a gesture action, the system can more accurately confirm that the driver indeed intends to navigate, rather than other tasks, thereby reducing the probability of false triggering.

[0017] Secondly, by combining vehicle operating state parameters, in-vehicle noise level parameters and driver gaze area parameters to perform relevance screening on the candidate set of interaction intents, false operations caused by environmental interference and invalid inputs can be significantly reduced. When the noise level exceeds the threshold, the system will automatically exclude voice instructions in that period; when the driver's gaze direction continuously deviates from the center screen, the system will shield the corresponding gesture operation; in the high-speed driving state, the system will preferentially retain safety-related intents. These processing logics make the system more in line with the safety needs and interaction habits of the driving scene, improving the effectiveness of the interaction data.

[0018] Furthermore, by performing conflict detection on the first set of interaction intents, contradictions between different modal inputs or conflicts between different user inputs can be effectively solved. For example, when the target functions corresponding to the voice instruction and the gesture operation are inconsistent, the system will select the interaction intent that is more in line with the driving conditions in combination with the operating state; when the driver and the occupant input conflict, the system will preferentially retain the driver input, thereby ensuring the concentration of vehicle control rights. This conflict processing mechanism avoids the problem of task flow confusion, enhancing the stability and safety of the system.

[0019] Further, the second interaction intent set is prioritized according to the target function category, and a task candidate queue is generated, which can make the system have a clear execution order when processing multiple tasks. The safety prompt and control class task is given the highest priority, the navigation task is given the medium priority, and the entertainment and information display class task is given the low priority. This hierarchical strategy ensures that the system can prioritize the functions most critical to driving safety when resources are limited or tasks conflict, avoiding low-priority tasks from preempting system resources.

[0020] Finally, by performing task confirmation on the task candidate queue and generating an effective task queue, and then implementing priority switching in task scheduling, the system can interrupt low-priority tasks immediately when a higher-priority input is detected. This mechanism can ensure timely response in emergency situations. For example, when the system is executing entertainment playback, once a collision warning instruction is detected, the system will immediately interrupt the entertainment task and execute the warning prompt first, thereby improving driving safety and the real-time performance of task scheduling. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a flow chart of an intelligent cockpit interaction management method provided by an embodiment of the application. DETAILED DESCRIPTION

[0022] Exemplary embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings; however, they are not limited thereto and can be embodied in various forms. It is to be understood that the embodiments described herein are merely described in order to more completely and specifically disclose the present disclosure and to convey the substance thereof to those skilled in the art.

[0023] As shown in Figure 1 An embodiment of the application proposes an intelligent cockpit interaction management method, which comprises: Receiving voice interaction data, gesture interaction data and expression interaction data collected in the cockpit, and adding source labels and time stamps to each piece of data to obtain multi-modal interaction raw data; According to the multi-modal interaction raw data, performing intent analysis, extracting trigger conditions and target functions, and forming an interaction intent candidate set; According to the interaction intent candidate set, combining vehicle operating state parameters, in-vehicle noise level parameters and driver gaze area parameters, performing relevance screening to obtain a first interaction intent set; For the first interaction intent set, performing conflict detection, determining the interaction intent with abnormal trigger conditions or conflicting target functions as invalid and removing it to obtain a second interaction intent set; Based on the second interaction intention set, priority classification is performed according to the category of the target function to generate a task candidate queue, wherein the safety-related target function is set to high priority, the navigation target function is set to medium priority, and the entertainment or information display target function is set to low priority; Confirm the tasks in the candidate task queue and generate the tasks that meet the trigger conditions into the valid task queue; Task scheduling is performed according to the valid task queue. When a higher priority task input is detected, the low priority task is interrupted and execution is switched.

[0024] In an embodiment of the present invention, the proposed intelligent cockpit interaction management method can effectively avoid the false triggering and task confusion problems commonly found in existing methods during the reception and processing of multimodal interaction data. By adding source tags and timestamps to voice, gesture, and expression data during the acquisition phase, subsequent processing steps can accurately determine the validity and temporal continuity of the data, thereby improving the traceability and reliability of the interaction information. For example, when a driver issues voice commands repeatedly within a short period of time, the method can identify the sequential relationship of the commands through timestamps, preventing earlier temporary operations from being mistaken for the current control task.

[0025] Next, by performing intent analysis on multimodal interaction data, trigger conditions and target functions can be accurately extracted to form a set of candidate interaction intents. This process avoids the method's over-reliance on single-modal data and improves the comprehensiveness of user intent recognition. For example, when a user speaks the voice command "play music," the method not only recognizes the voice content but also combines gestures or facial expressions to further confirm the validity of the command, making the results more in line with the actual needs of driving scenarios.

[0026] Secondly, based on the candidate interaction intentions, parameters such as vehicle operating status, in-vehicle noise level, and driver's gaze area are introduced for correlation screening, which greatly reduces erroneous execution caused by interference factors. When the vehicle is in a high-speed straight-ahead state, this method will prioritize identifying and retaining safety-related interaction intentions to ensure that driving safety information is responded to first. For example, while the vehicle is driving at high speed, if the driver casually discusses "going somewhere" with a passenger while performing air conditioning adjustment actions by looking at the central control interface, this method will determine that the voice command is interfering and eliminate it, retaining only air conditioning-related operations.

[0027] Further, by performing conflict detection on the first set of interaction intents, duplication or conflict between target functions can be avoided. The method can eliminate intents that have abnormal trigger conditions or contradict each other, thereby forming a more concise and reliable interaction set. For example, when the user uses the voice instruction "turn on navigation" and the gesture operation "exit navigation" at the same time, the method can make a judgment in combination with the running state of the vehicle and the trigger conditions, and preferentially retain the operation that conforms to the current driving state and eliminate invalid instructions, thereby avoiding confusion in function execution of the vehicle machine.

[0028] Further, by performing priority classification on the second set of interaction intents, a task candidate queue can be generated according to the importance of different functions. Safety-related functions are set to the highest priority, navigation functions are set to the medium priority, and entertainment and information display functions are set to the lowest priority. Such a classification mechanism enables the method to have a clear scheduling order when executing tasks. For example, when the method receives the driver's "safety distance prompt request" and "play music" operations at the same time, the method will preferentially execute the safety prompt rather than the music playing.

[0029] Finally, in the task confirmation and task scheduling links, the scheme of the embodiment of the application can ensure that the tasks executed by the method are all valid tasks that meet the trigger conditions. When a higher-priority task input is detected, the method will automatically interrupt the low-priority task and switch to execution. Such a scheduling strategy not only ensures the timeliness of high-priority tasks, but also prevents low-priority tasks from interfering with driving safety. For example, when the driver is watching vehicle status information, if the method simultaneously detects a collision warning signal input, the vehicle information display will be immediately interrupted, and the collision warning task will be preferentially executed, thereby significantly improving the driving safety and the reliability of the interaction method.

[0030] Among them, the voice interaction data, gesture interaction data and expression interaction data collected in the cabin are received, and a source mark and a time stamp are added to each data to obtain multi-modal interaction raw data, specifically including: The voice input of the driver and the passenger is collected through the vehicle-mounted microphone module, and the gesture action and facial expression information are obtained by using the camera arranged in the vehicle. The voice data is attached with a source mark before entering the voice recognition module, and the mark is used to indicate whether the voice comes from the microphone at the driver's position or the microphone at the passenger's position. The gesture data and the expression data are also marked with a source mark to distinguish whether the operation is of the driver or the passenger. All data are automatically added with a time stamp, and the time stamp is generated by a unified clock of the vehicle-mounted control unit and is used to indicate the collection time of the data. This step ensures that the source and order of the data remain consistent in the subsequent processing stage.

[0031] According to the multi-modal interaction original data, intention analysis is performed, trigger conditions and target functions are extracted, and an interaction intention candidate set is formed, specifically including: The voice data is converted into text content after passing through the voice recognition module, and is subjected to semantic analysis by the intention recognition model to extract keywords and command types. The gesture data is subjected to feature extraction by the action recognition module, such as recognizing finger sliding, clenching, clicking, and other operation modes. The expression data is subjected to emotion label extraction by the facial expression recognition module, such as gaze, blinking, nodding, and other behaviors. The processing unit fuses data from different modalities and determines their corresponding trigger conditions (such as the user gazing at the center screen accompanied by a voice instruction) and target functions (such as navigation, playing music, adjusting the air conditioner) according to a pre-set rule base or a trained classification model, and the determination result is taken as the interaction intention candidate set.

[0032] In a preferred embodiment of the present application, according to the interaction intention candidate set, in combination with the vehicle operating state parameter, the in-vehicle noise level parameter, and the driver gaze area parameter, correlation screening is performed to obtain a first interaction intention set, including: According to the interaction intention candidate set and the in-vehicle noise level parameter, when the noise level continuously exceeds the pre-set noise threshold within the pre-set time window, the corresponding voice interaction data is determined to be invalid and is removed, and a first candidate set is generated; According to the first candidate set and the driver gaze area parameter, when the driver's gaze direction continuously remains deviated from the central control interaction interface by more than a pre-set first time length within a pre-set angle range, the corresponding gesture interaction data is determined to be invalid and is removed, and a second candidate set is generated; According to the second candidate set and the vehicle operating state parameter, when the operating state is high-speed straight driving, safety control related interaction intentions are preferentially retained, when the operating state is low-speed driving, navigation and safety related interaction intentions are retained, and when the operating state is parking, entertainment related interaction intentions are allowed to be retained, and finally a first interaction intention set is generated.

[0033] In the embodiment of the present application, by introducing multi-dimensional parameters such as vehicle operating state, in-vehicle noise level, and driver gaze area based on the interaction intention candidate set for correlation screening, the accuracy and practicality of the interaction data can be effectively improved. In this way, the present method not only relies on original voice or gesture signals, but also combines external environment and driver state for comprehensive judgment, thereby significantly reducing the probability of redundant and false triggering operations.

[0034] Then, when the in-vehicle noise level is high, the method compares the preset noise threshold to eliminate invalid voice data. For example, when the vehicle passes through a tunnel or the window is opened, causing the background noise to exceed the threshold, the method identifies that the voice input in this environment may be distorted, thereby automatically shielding these instructions to avoid false triggering of navigation or entertainment tasks.

[0035] Secondly, by combining the driver's gaze area parameter for duration determination, false operations caused by the driver's distraction can be avoided. When the driver's long gaze direction deviates from the center screen, the method determines that the gesture operation lacks effectiveness and is eliminated. For example, when the driver inadvertently makes hand movements while turning to observe the road conditions, the method automatically ignores the input, thereby avoiding false control responses.

[0036] Finally, by combining the vehicle operating state for subdivision determination, the method can flexibly retain different types of interaction intentions in different driving conditions. When the vehicle is in a high-speed straight driving state, the method only retains safety control-related intentions to ensure the absolute priority of driving safety; when the vehicle is in a low-speed driving state, the method allows the simultaneous retention of navigation and safety-related intentions to meet the diversified needs of driving; when the vehicle is in a parking state, the method can appropriately retain entertainment intentions to improve the comfort and entertainment experience of the vehicle. This hierarchical and differentiated processing approach allows the interaction method to dynamically adapt to driving scenarios, balancing safety and convenience.

[0037] Among them, the preset time window refers to multiple consecutive samplings of the in-vehicle noise level within a fixed length of time, used to dynamically reflect the changes in noise. The length of the time window can be set according to the vehicle application scenario, such as 1 second, 2 seconds, or 3 seconds. Within the time window, the method will collect noise values at a fixed sampling interval (such as every 100 milliseconds), and store these noise values in chronological order to form a noise sequence. When the method needs to determine the validity of a segment of voice data, it checks the noise sampling values within the time period corresponding to the voice data and determines whether it is disturbed by noise based on the continuous performance within the time window. By using the time window mechanism, false judgments caused by instantaneous noise fluctuations can be avoided.

[0038] The preset noise threshold is used as a baseline for determining whether the in-vehicle voice data is reliable. The threshold can be determined in combination with the typical noise environment in the vehicle, for example, the background noise in the vehicle is generally about 50 decibels when the vehicle is idling, and the wind noise can exceed 70 decibels when driving at high speed. The method can set the noise threshold to a fixed value (such as 65 decibels), or use an adaptive method to dynamically adjust according to the average noise level of historical sampling plus a correction factor. When the continuously sampled noise values in the time window are always higher than the threshold, it indicates that the voice signal may be severely disturbed, and the corresponding voice interaction data will be marked as invalid. Through such setting, the voice recognition accuracy in static and dynamic environments can be considered.

[0039] The preset angle range refers to the maximum deviation angle allowed between the driver's gaze direction and the reference direction of the center control interaction interface of the vehicle. The method detects the included angle between the driver's gaze direction and the center point of the center control screen in real time through the eye tracking camera or infrared sensor installed on the instrument panel. When the included angle is less than or equal to the set angle range (for example, ± 15° or ± 20°), it is determined that the driver's gaze is within the effective range; when the included angle is greater than the angle range, it is considered that the driver's line of sight has deviated from the interaction area. The setting of the range can be obtained through experimental calibration, that is, while ensuring that the driver can operate naturally, it is ensured that the driver's attention is focused on the main interaction interface. By setting the angle range, the driver's random movements in non-interaction situations can be filtered out.

[0040] The preset first time length is used to limit the time threshold of the driver's gaze deviation. The time threshold can be determined through driving behavior experiment data, for example, 500 milliseconds, 1 second or 2 seconds. When the method detects that the driver's gaze direction continuously exceeds the preset angle range for a time that accumulates more than the time length, it is determined that the driver does not pay attention to the interaction interface in this period, and therefore the gesture interaction data corresponding to this period is determined to be invalid and is removed. If the deviation duration does not exceed the time length, it indicates that the driver only temporarily shifts his gaze, for example, to observe the rearview mirror, and the gesture data is still considered valid. By setting the time length, the short-time deviation and long-time invalidation situations can be effectively distinguished.

[0041] In a preferred embodiment of the present application, for the first set of interaction intents, conflict detection is performed, and the interaction intents that trigger abnormal conditions or target function conflicts are determined to be invalid and removed to obtain a second set of interaction intents, including: According to the timestamps of the interaction intents in the first set of interaction intents, when the timestamps are not continuous or exceed a preset time interval, the interaction intents are determined to be invalid and removed to generate a set of interaction intents filtered by time; Consistency determination is performed according to the trigger conditions of different modalities in the time-filtered interaction intention set. When the voice interaction data and the gesture interaction data have the same target function but conflicting trigger conditions, the interaction intention that matches the vehicle operating state parameter is retained, and the interaction intention that does not meet the condition is removed, to generate a modal-filtered interaction intention set. When the driver input and the passenger input have a target conflict, the interaction intention of the driver input is retained, and the interaction intention of the passenger input is removed, to generate a second interaction intention set.

[0042] In the embodiments of the present application, by performing conflict detection on the first interaction intention set, the reliability and consistency of task selection can be further guaranteed, thereby avoiding function conflicts or execution abnormalities caused by the simultaneous existence of multiple interaction intentions.

[0043] Then, by comparing the timestamps of the interaction intentions, abnormal instructions that are not continuous or exceed the preset interval can be excluded. For example, after a normal voice operation by the driver, if the method receives an old instruction that is separated for a long time due to delay, the old instruction will be removed to prevent it from being confused with a new task, thereby improving the ability of the method to control the timeliness of the task.

[0044] Secondly, in the consistency determination process of the multi-modal trigger condition, the method can make a reasonable choice when there is a conflict between voice and gesture input. For example, when the driver issues a "start navigation" instruction by voice, but the gesture performs an "exit navigation" operation, the method will determine the vehicle operating state and prioritize the voice operation when driving at high speed, and prioritize the gesture operation when parking, thereby avoiding ambiguity in task execution.

[0045] Finally, when there is a conflict in input source, the method can distinguish the input intentions of the driver and the passenger, and always prioritize the driver's operation, thereby ensuring that the control right of the vehicle is concentrated on the driver. For example, when the passenger issues an instruction to change the vehicle speed mode, the method will automatically shield the operation to avoid threatening the driving safety. This mechanism can ensure that the driver has the highest decision-making power in critical interaction scenarios, thereby significantly improving the safety and rationality of the method.

[0046] The preset time interval refers to a time threshold used to determine whether two adjacent interaction intentions belong to the same interaction period. The threshold is usually set according to the user's operation habits and interaction response requirements in the driving scenario. For example, when the driver continuously issues voice instructions or gesture operations, the normal interval is usually less than 1 second or 2 seconds, and instructions exceeding this time length are likely to belong to different operation periods.

[0047] In the specific processing, the method extracts the timestamp carried by each interaction intention in the first interaction intention set, and sequentially calculates the interval duration between adjacent two timestamps. When it is detected that the interval is greater than the preset time interval threshold, it is determined that the latter interaction intention does not have continuity with the former one, which may be a delayed input or an irrelevant historical input, so the interaction intention is marked as invalid and eliminated.

[0048] For example, if the preset time interval is 2 seconds, when the driver successively speaks the voice instructions "turn on music" and "turn up the volume" within 1 second, the method regards them as continuous interaction inputs and retains them at the same time. However, if the driver issues the instruction "turn off music" after 5 seconds, the method determines that the instruction is beyond the reasonable range of continuous interaction, and eliminates it from the set to prevent the delayed operation from interfering with the task being performed.

[0049] By setting and applying the preset time interval, the method can ensure that the retained interaction intentions maintain coherence and logic in the time dimension, providing high-quality data input for subsequent consistency determination and conflict detection.

[0050] In a preferred embodiment of the present application, according to the second interaction intention set, the target functions are prioritized according to the categories to generate a task candidate queue, including: According to the target functions of the second interaction intention set, the initial classification is performed, the interaction intentions related to safety prompts and safety control are divided into high priority, the interaction intentions related to navigation setting and path planning are divided into medium priority, and the interaction intentions related to entertainment playing and information display are divided into low priority, to generate an initial task candidate queue; According to the initial task candidate queue and the vehicle running state parameter, when the running state is emergency braking or collision warning, all non-safety related tasks are downgraded as a whole to generate a dynamically adjusted task candidate queue; According to the dynamically adjusted task candidate queue, the sequential arrangement is performed to form the execution order according to the priority, and the final task candidate queue is obtained.

[0051] In the embodiment of the present application, by prioritizing the second interaction intention set, a task candidate queue that meets the driving safety logic can be constructed, so that the method can orderly and efficiently complete the scheduling when processing multiple tasks.

[0052] Next, the method first classifies the interaction intents related to safety prompts and safety controls as the highest priority, classifies the interaction intents related to navigation settings and path planning as the medium priority, and classifies the interaction intents related to entertainment playing and information display as the lowest priority. For example, when the user triggers the "close distance prompt" and "play music" operations at the same time during driving, the method prioritizes the safety prompt and does not delay the response to the safety task due to the low-priority entertainment operation.

[0053] Secondly, by associating the initial task candidate queue with the vehicle operating state, the priority order of different functions can be dynamically adjusted. When the vehicle is in an emergency braking or collision warning state, the method automatically downgrades all non-safety-related tasks, thereby concentrating computing and interaction resources on the most critical safety tasks. For example, when the vehicle detects a potential collision risk, even if the user is setting a navigation, the method will prioritize interrupting the navigation task and immediately triggering a collision warning prompt.

[0054] Finally, by sequentially arranging the dynamically adjusted task candidate queue, the method can form a clear execution order according to the priority. This mechanism not only ensures the orderly execution of tasks, but also improves the flexibility and response speed of task switching. For example, when driving at low speed, the method can process medium and low-priority entertainment tasks while executing navigation instructions; but in the event of an emergency, the method will immediately interrupt the entertainment operation and quickly switch to high-priority task execution, thereby ensuring driving safety while considering user experience.

[0055] Among them, according to the initial classification of the target function of the second interaction intent set, the interaction intent related to safety prompts and safety controls is classified as high priority, the interaction intent related to navigation settings and path planning is classified as medium priority, and the interaction intent related to entertainment playing and information display is classified as low priority, to generate an initial task candidate queue, specifically including: After the second interaction intent set is formed, the target function corresponding to each interaction intent in the set is read in turn, and it is compared with the preset function category table. The function category table can be stored in the memory of the vehicle control unit, and the table records the correspondence between typical functions and priorities. For example, collision warning, lane deviation prompt, braking assistance, etc. are defined as safety control class functions; destination input, path re-planning, navigation start / exit, etc. are defined as navigation class functions; audio playing, video information display, in-vehicle entertainment control, etc. are defined as entertainment class functions.

[0056] After the comparison is completed, the control unit assigns an initial priority label to each interaction intent, and writes the tasks with the priority label into the candidate task queue in turn. At this time, the tasks in the candidate queue are arranged in the order of input, and the priority label is used as the basis for subsequent dynamic adjustment and sorting. Through this classification mechanism, the tasks of different functional types can be distinguished when scheduling, providing basic data for subsequent dynamic adjustment and execution order arrangement.

[0057] In the method, the initial task candidate queue is associated with a vehicle operating state parameter, and when the operating state is emergency braking or collision warning, all non-safety related tasks are downgraded as a whole to generate a dynamically adjusted task candidate queue, specifically including: After the candidate task queue is generated, the vehicle control unit continuously monitors the vehicle operating state parameter, which is provided by the speed sensor, brake sensor and collision warning sensor in real time. When it is detected that the vehicle is in an emergency braking state, or the collision warning method issues a warning signal, the vehicle control unit triggers the priority adjustment program. The adjustment program will first traverse all tasks in the initial candidate queue to determine whether they are safety related tasks. If they are safety related tasks, their high priority is maintained; if they are navigation tasks or entertainment tasks, their priority is uniformly downgraded by one level.

[0058] For example, the original navigation task is of medium priority in the initial classification, and after adjustment it will be downgraded to low priority; the entertainment task is of low priority in the initial classification, and after adjustment it will be marked as the lowest priority or suspended state, and will no longer participate in the current scheduling. After the processing is completed, the vehicle control unit generates a dynamically adjusted task candidate queue, which ensures that all computing and interaction resources are concentrated on safety tasks in an emergency state, thereby avoiding interference of non-critical tasks with driving safety.

[0059] In a preferred embodiment of the present application, the interaction intent candidate set is continuously compared with the in-vehicle noise level parameter, including: According to the interaction intent candidate set, voice interaction data is extracted, and multiple frame sampling is performed within a preset time window in combination with the in-vehicle noise level parameter to generate a noise change sequence; According to the noise change sequence, continuous threshold determination is performed, and when the time period during which the noise continuously exceeds the preset noise threshold is greater than a preset second length, the corresponding voice interaction data is determined as invalid and removed to obtain voice interaction data; According to the voice interaction data, the corresponding voice interaction data is retained when the time period is less than the preset length to form a first candidate set.

[0060] In the embodiment of the present application, by continuously comparing the voice interaction data and the in-vehicle noise level parameter, the interference of environmental noise on the interaction method can be effectively avoided. The method collects multiple noise samples in a preset time window and generates a noise change sequence, so that the determination of voice input is no longer dependent on single instantaneous data, but based on continuous analysis in time, thereby improving the stability and accuracy of the determination.

[0061] Then, when the noise level continuously exceeds the threshold value within a certain period of time, the method determines the voice data in this period as invalid and eliminates it, thereby avoiding voice mis-triggering in scenarios such as opening the window, wind noise, or road echo. For example, when the vehicle is driving at high speed through a tunnel, the voice recognition module may capture the driver's irrelevant conversation, and the mechanism of the embodiment of the present application can determine that it does not meet the voice interaction condition, and then eliminate the data.

[0062] Secondly, when the time period during which the noise exceeds the threshold value is less than the preset length of time, the method retains the voice interaction data, thereby ensuring that the loss of effective voice instructions will not be caused by temporary interference. For example, when the driver issues a "turn on navigation" voice command in the window opening state, although there is wind noise for a short time, since the noise duration is short, the method can still correctly recognize and execute the navigation instruction, thereby improving the reliability and fault tolerance of voice interaction in actual driving.

[0063] Among them, the voice interaction data is extracted according to the interaction intent candidate set, and multiple frames are sampled in a preset time window in combination with the in-vehicle noise level parameter to generate a noise change sequence, which specifically includes: First, the data entries related to voice are extracted from the interaction intent candidate set, which are collected by the vehicle-mounted microphone. Then, in a preset time window (such as 1 second or 2 seconds), the in-vehicle noise level parameter is obtained at a fixed time interval, and each collected noise value is taken as a sampling point. After continuously collecting multiple sampling points, a noise sequence that changes with time is formed. This sequence reflects the dynamic change of the noise level in the time window, providing a basis for subsequent effectiveness determination.

[0064] Among them, the noise change sequence is continuously threshold determined, when the time period during which the noise continuously exceeds the preset noise threshold value is greater than the preset second length of time, the corresponding voice interaction data is determined as invalid and eliminated, to obtain the voice interaction data, specifically including: After obtaining the noise change sequence, the method compares the noise value with the preset threshold one by one. If it is found that the noise value is higher than the threshold for many times in succession, the time length of this continuous interval is recorded. When the time length exceeds the preset time threshold (for example, lasting more than two seconds), it is determined that the voice data collected in this time period is insufficient in validity, which is marked as invalid and discarded. If the continuous high-noise time length is lower than the preset threshold, the voice data will not be discarded directly, but will be further screened in the next step. Through such a continuous determination process, invalid voice input collected in a long-term high-noise environment can be effectively filtered out.

[0065] The preset second time length is used to define the duration threshold of the in-vehicle noise interference, and is an important parameter for judging whether the voice interaction data is seriously affected by environmental noise. Its setting is based on the typical use scenarios of the vehicle, such as high-speed driving of the vehicle, tunnel echo, opening of the vehicle window, or conversation among multiple passengers in the vehicle.

[0066] Specifically, the method continuously collects noise values within a preset time window and generates a noise change sequence. When it is detected that the noise value is continuously higher than the preset noise threshold, the continuous duration of the interval is accumulated. If the continuous duration exceeds the preset second time length, it is determined that the voice interaction data in this period does not have validity, and is removed from the candidate set. For example, if the second time length is set to 2 seconds, when the noise level is always higher than the threshold within 3 seconds, it is determined that the voice command in this period may have been distorted or cannot be reliably recognized, and therefore it is filtered out.

[0067] The numerical value of the preset second time length can be determined by experimental calibration. It is usually adjusted between 0.5 seconds and 3 seconds to balance between "avoiding false deletion of valid voice" and "timely removing invalid voice". In some embodiments, the time length can be dynamically adjusted in combination with the vehicle running state: in the high-speed driving scenario, the preset second time length can be set to be shorter to quickly shield invalid input caused by strong wind noise; in the parking scenario, it can be set to be longer to ensure that the method retains voice interaction data as much as possible when the environment is stable.

[0068] By introducing the preset second time length, the method can accurately control the validity of voice data during continuous noise interference, thereby improving the reliability of voice interaction.

[0069] According to the voice interaction data, the corresponding voice interaction data is retained when the time period is less than the preset time length to form a first candidate set, specifically including: When the duration of noise exceeding the threshold is less than the preset length, the method continues to check the voice interaction data in the time period. The content of the check can include whether the duration of the voice segment is greater than the minimum recognition length, whether the voice signal is complete, whether it contains recognizable command keywords, etc. If the voice interaction data meets these conditions, it is retained as valid data and written into the first candidate set; if not, it is determined to be invalid and discarded. In this way, even under short-term noise interference, the valid voice input of the driver can still be retained as much as possible.

[0070] In a preferred embodiment of the application, the duration determination according to the first candidate set and the driver's gaze area parameter includes: According to the first candidate set, gesture interaction data is extracted and compared frame by frame with the driver's gaze area parameter to generate a gaze deviation record; According to the gaze deviation record, the continuous deviation time is counted, and when the continuous deviation time exceeds the preset first length, the corresponding gesture interaction data is determined to be invalid and discarded, obtaining gesture interaction data; According to the gesture interaction data, when the continuous deviation time is less than the preset first length, the corresponding gesture interaction data is retained to form a second candidate set.

[0071] In the embodiment of the application, by combining the driver's gaze area parameter to determine the duration of the gesture interaction data, the interference of the driver's unintentional action or distraction on the interaction result can be effectively avoided. The method generates a gaze deviation record by comparing the driver's gaze direction and gesture operation frame by frame, so that the validity of the gesture can be associated with the driver's line of sight focus.

[0072] Then, when the driver's gaze direction deviates from the central control interaction interface for more than a preset length, the method automatically determines that the corresponding gesture data is invalid and is discarded. For example, when the driver turns his head to observe the road conditions behind during lane changing, if the hand swings unintentionally, the method will ignore the action to avoid misidentifying it as a control instruction.

[0073] Secondly, when the driver's gaze direction only deviates from the central control interaction interface for a short time, the method retains the corresponding gesture interaction data, thereby ensuring the usability of the gesture operation. For example, when the driver quickly glances at the rearview mirror and then returns to the central control interface and makes a "switch music" gesture, the method can recognize the validity of the action, thereby ensuring the normal execution of the entertainment control function. This mechanism can balance safety driving and interaction convenience, significantly improving the naturalness and accuracy of human-computer interaction.

[0074] Wherein, the gesture interaction data is extracted according to the first candidate set, and frame-by-frame comparison is performed with the driver's gaze area parameter to generate a gaze deviation record, specifically including: First, the interaction data related to the gesture is screened out from the first candidate set. These gesture data are usually collected by the camera installed near the center control area or steering wheel and processed by the motion recognition algorithm. At the same time, the driver's gaze area parameter is provided by the eye tracking camera or infrared sensor in real time. This method will correspond the gesture data and gaze area data frame by frame on the time axis to determine whether the driver's line of sight falls within the preset range of the center control screen or the interaction area when performing the gesture action. If the line of sight deviates from the range, the method will mark it as "deviation" in the record, otherwise it will be marked as "effective". Finally, a continuous gaze deviation record is generated for subsequent judgment.

[0075] Wherein, the continuous deviation time is counted according to the gaze deviation record, and when the continuous deviation time exceeds a preset first duration, the corresponding gesture interaction data is determined to be invalid and removed, to obtain gesture interaction data, specifically including: After obtaining the gaze deviation record, the method will count the time segments marked as "deviation" in it. The counting method is: the continuously appearing deviation frames are merged into a time period, and the duration of this period is accumulated. When a continuous deviation time exceeds a preset duration threshold (for example, more than 2 seconds), the gesture interaction data corresponding to the time period is determined to be invalid and is removed from the candidate set. This determination method ensures that only when the driver does not gaze at the interaction area for a long time, the related gesture will be considered invalid, thereby avoiding false judgments due to temporary line of sight shift.

[0076] Wherein, according to the gesture interaction data, the corresponding gesture interaction data is retained when the continuous deviation time is less than the preset first duration to form a second candidate set, specifically including: When the statistical result shows that the gaze deviation time is less than the set threshold, the method considers that the driver maintains attention to the interaction area most of the time, so the gesture interaction data related to this period is considered valid. In this case, the method will retain these gesture data and write them into the second candidate set. This can ensure that when the driver performs gestures, even if there is a short gaze shift (such as quickly looking at the rearview mirror or observing the road), the valid operation can still be recognized and retained by the method. In this way, the gesture data in the second candidate set has higher reliability and applicability.

[0077] In a preferred embodiment of the present application, the second candidate set is further divided and determined according to the vehicle running state parameter, including: According to the second candidate set, the vehicle speed, steering angle and braking state in the vehicle running state parameter are combined to form a running state combination; According to the running state combination, it is determined that the vehicle is in a high-speed straight driving, low-speed driving or parking state, and the interaction intents corresponding to different states are classified and processed to generate an interaction intent set subdivided by the running state; According to the interaction intent set subdivided by the running state, the safety control related interaction intent is retained when driving at high speed, the navigation related and safety control related interaction intents are retained at the same time when driving at low speed, and the entertainment related interaction intent is retained in the parking state, forming a first interaction intent set.

[0078] In the embodiments of the present application, by subdividing the second candidate set and the vehicle running state parameters, the dynamic adaptation of the interaction intent in different driving conditions can be realized, so that the response of the method is more in line with the actual driving demand and safety constraints. The method forms a running state combination by combining the vehicle speed, steering angle and braking state, and determines that the vehicle is currently in a high-speed straight driving, low-speed driving or parking state.

[0079] Then, when the vehicle is in a high-speed straight driving state, the method will preferentially retain the safety control related interaction intent, so as to ensure that the driver can obtain fast and clear safety prompt or control response in the high-speed scene. For example, on the highway, if the driver simultaneously issues the "air conditioner adjustment" and "vehicle distance warning" instructions, the method will preferentially respond to the vehicle distance warning to ensure driving safety.

[0080] Secondly, when the vehicle is in a low-speed driving state, the method will simultaneously retain the navigation related and safety related interaction intents, so as to provide more flexible navigation auxiliary functions under the premise of ensuring safety. For example, when driving on a congested urban road, the driver can simultaneously receive the front collision warning information and use voice input to adjust the navigation destination.

[0081] Finally, when the vehicle is in a parking state, the method allows the retention of entertainment related interaction intents, thereby improving the comfort and entertainment experience of the user in the stationary state of the vehicle. For example, when parking in a parking lot, the user can play music, browse information through gestures or voice instructions, and the method will preferentially ensure the smoothness of these functions. This state subdivision mechanism realizes the scene processing of the interaction strategy, and significantly improves the intelligent level and user experience of the method.

[0082] Among them, according to the second candidate set, the vehicle speed, steering angle and braking state in the vehicle running state parameter are combined to form a running state combination, which specifically includes: Firstly, the vehicle speed, steering angle and brake state parameters are collected in real time from the vehicle sensors. The vehicle speed is provided by the speed sensor, the steering angle is provided by the steering wheel angle sensor or the wheel angle sensor, and the brake state is provided by the brake pedal stroke sensor or the brake pressure sensor. Then, the above three parameters at the same time are associated to form a parameter set, i.e. the running state combination. For example, the vehicle speed at a certain time is 80 kilometers per hour, the steering angle is close to zero, and the brake state is not triggered, then the running state combination can be marked as "high-speed straight running state". In this way, a corresponding running state combination can be generated at each moment for subsequent judgment.

[0083] Among them, according to the running state combination, the vehicle is determined to be in a high-speed straight running, low-speed running or parking state, and the interactive intent corresponding to different states is classified and processed to generate an interactive intent set subdivided by running state, including: After forming the running state combination, the method compares the combination value with the preset running state determination rule. The rule can be defined as: when the vehicle speed is greater than the preset high-speed threshold and the steering angle is close to zero and the brake is not triggered, it is determined to be in a high-speed straight running state; when the vehicle speed is less than the high-speed threshold and greater than zero, it is determined to be in a low-speed running state; when the vehicle speed is equal to zero and the brake is triggered, it is determined to be in a parking state. Then, the method will filter and mark the interactive intent in the second candidate set according to different running states. For example, in the high-speed straight running state, only the interactive intent related to safety control is retained, while the navigation or entertainment interactive intent is marked as invalid; in the low-speed running state, the interactive intent related to safety and navigation is retained; in the parking state, the interactive intent related to safety, navigation and entertainment is allowed to be retained. Finally, an interactive intent set subdivided by running state is obtained.

[0084] Among them, according to the interactive intent set subdivided by running state, the safety control related interactive intent is retained when running at high speed, the navigation related and safety control related interactive intent is retained when running at low speed, and the entertainment related interactive intent is retained when in parking state, to form the first interactive intent set, including: After obtaining the interaction intent set subdivided by running state, the method performs final screening according to the determined vehicle running state. If the running state is high-speed straight driving, only the intents related to safety control in the set are written into the first interaction intent set, such as vehicle distance warning, lane deviation prompt, etc. If the running state is low-speed driving, the intents related to navigation task and safety control in the set are written into the first interaction intent set at the same time, such as path update and braking prompt. If the running state is parking, the entertainment-related intents in the set are selected and written into the first interaction intent set, such as playing music and displaying multimedia information. In this way, the first interaction intent set formed can be flexibly adjusted according to the differences in vehicle running state, providing high-quality input data for subsequent conflict detection and task scheduling.

[0085] In a preferred embodiment of the present application, the timestamps of the interaction intents in the first interaction intent set are compared, including: The timestamps of each interaction intent are extracted from the first interaction intent set, the time interval between adjacent interaction intents is calculated, and time interval data is generated; According to the time interval data, threshold determination is performed. When the time interval is greater than the preset time interval, the interaction intent is determined to be abnormal and is removed, forming an interaction intent set filtered by interval; According to the interaction intent set filtered by interval, time window determination is performed. When the interaction intent does not fall within the preset time window, the interaction intent is determined to be invalid and is removed, generating an interaction intent set filtered by time.

[0086] In the embodiment of the present application, by comparing the timestamps of the interaction intents in the first interaction intent set, the influence of expired instructions or abnormal interval data on the correct execution of the task can be effectively avoided. By calculating the time interval between adjacent interaction intents, the method can determine whether the interaction input is continuous and reasonable, thereby ensuring the timeliness and consistency of the task.

[0087] Next, when the time interval is greater than the preset time interval, the method determines that the interaction intent is abnormal and is removed, thereby preventing historical residual operation signals from entering the task queue. For example, when the driver issued a "start navigation" voice command a few minutes ago, the command was delayed due to network delay and was not recognized. If the command is no longer reasonable under the current driving state, the method will automatically determine it to be an invalid command and remove it.

[0088] Secondly, in the interaction intent set filtered by interval, the method further performs unified time window determination, so as to ensure that the reserved interaction intents belong to the same interaction period. For example, when the driver successively issues the operations of "adjusting air conditioner temperature" and "playing music", if the former operation exceeds the current time window, the method will eliminate it and only reserve the operation instruction matched with the current window. This processing manner can ensure the continuity between the interaction tasks and improve the accuracy and execution efficiency of task scheduling.

[0089] Among them, the timestamp of each interaction intent is extracted according to the first interaction intent set, the time interval of adjacent interaction intents is calculated, and the time interval data is generated, specifically including: First, the timestamp carried by each interaction intent in the first interaction intent set is extracted in sequence. The timestamp is automatically generated by the vehicle-mounted control unit in the data collection stage and is used to mark the time when the data is collected. Then, the timestamps of two adjacent interaction intents are sequentially compared, the time interval between them is calculated, and the interval values of all adjacent interaction intents are sequentially recorded, and finally a time interval data table is generated. The data table can intuitively reflect the distribution of each interaction intent on the time axis and provide a basis for subsequent effectiveness determination.

[0090] Among them, threshold determination is performed according to the time interval data, when the time interval is greater than the preset time interval, the interaction intent is determined as abnormal and eliminated, forming the interaction intent set filtered by interval, specifically including: After obtaining the time interval data, the method compares each time interval with a preset threshold. The threshold can be set according to the actual application scene, for example, two seconds or five seconds. When the interval between a certain interaction intent and the previous interaction intent is greater than the threshold, it is considered that the interaction intent exceeds the reasonable range of continuous input, and it may be historical residual data or delayed input data, so it is determined as abnormal and eliminated. The remaining interaction intents that do not exceed the threshold are reserved to form the interaction intent set filtered by interval. This step ensures that the reserved interaction intents have continuity and relevance in time.

[0091] Among them, the time window determination is performed according to the interaction intent set filtered by interval, when the interaction intent does not fall into the preset time window, the interaction intent is determined as invalid and eliminated, and the interaction intent set filtered by time is generated, specifically including: After the interval screening is completed, the method compares the screened interaction intents with a preset time window. The time window is used to limit the interaction intents that must be concentrated within a certain time period, such as a range of 1 second to 3 seconds. When the timestamp of a certain interaction intent is not within the unified window range, it is considered that the intent is out of synchronization with other interaction instructions, and is determined to be invalid and is rejected. Only the interaction intents within the unified time window are retained, and a time-screened interaction intent set is finally generated. Through this step, consistency between multi-modal interaction inputs in the time dimension can be ensured, and interference data across time periods is avoided from entering the subsequent processing process.

[0092] In a preferred embodiment of the present application, consistency determination is performed according to the trigger conditions of different modalities in the time-screened interaction intent set, including: According to the time-screened interaction intent set, voice interaction data and gesture interaction data are extracted to generate trigger condition pairs. According to the trigger condition pairs, consistency comparison is performed, and when the target function is consistent but the trigger condition is inconsistent, a modality conflict record is generated. According to the modality conflict record, priority determination is performed in combination with vehicle operating state parameters, voice interaction data is retained when the vehicle is in a high-speed driving state, gesture interaction data is retained when the vehicle is in a low-speed driving or parking state, and interaction intents that do not meet the vehicle state are rejected to form a modality-screened interaction intent set.

[0093] In the embodiment of the present application, by performing consistency determination on the trigger conditions of different modalities in the time-screened interaction intent set, contradictions and conflicts between voice, gesture and other multi-modal inputs can be effectively avoided. By forming voice and gesture trigger condition pairs, the method can perform consistency comparison on different trigger conditions of the same target function, so as to determine whether there is a modality conflict.

[0094] Next, when the target function is consistent but the trigger condition is inconsistent, the method generates a modality conflict record and performs priority determination in combination with the vehicle operating state. For example, when the driver inputs "turn on navigation" by voice and simultaneously performs "exit navigation" by gesture, the method will select according to the driving condition: in a high-speed driving state, the voice input is preferentially retained to ensure that the driver's instruction is not disturbed by the gesture; and in a parking state, the gesture operation is preferentially retained to adapt to a more natural interaction mode of the user.

[0095] Secondly, by rejecting the interaction intention that does not conform to the vehicle state, the method can maintain the consistency and rationality of the interaction result. For example, when the driver says "play music" and makes a "turn off music" gesture at the same time while driving at low speed, the method will prefer to combine the running state to select a more suitable operation, avoiding the interference of navigation or safety tasks. This mechanism can ensure the harmony and unity of multi-modal interaction, thereby improving the accuracy and stability of human-computer interaction.

[0096] Among them, the voice interaction data and gesture interaction data are extracted from the time-filtered interaction intention set to generate trigger condition pairs, specifically including: First, the voice interaction data and gesture interaction data are extracted from the time-filtered interaction intention set. The voice interaction data usually contains the recognized voice instruction text and its corresponding target function information, and the gesture interaction data contains the recognized gesture action and its mapped target function information. Then, the voice interaction data and gesture interaction data are paired according to the time stamp, and if they belong to the same interaction period in time, they are combined into a trigger condition pair. Each trigger condition pair contains the trigger conditions of the same target function in two modalities, providing input for subsequent consistency comparison.

[0097] Among them, according to the consistency comparison of the trigger condition pair, when the target function is consistent but the trigger condition is inconsistent, a modal conflict record is generated, specifically including: After the generation of the trigger condition pair, the method will compare the target functions expressed by the voice and gesture modalities one by one. When it is detected that the target functions pointed by the two modalities are the same, but their trigger conditions are contradictory, a modal conflict record will be generated. For example, the voice modality instruction is "turn on navigation", while the gesture modality instruction corresponds to the action "exit navigation", then it is determined that there is a trigger condition conflict. In this case, the method will mark the conflict type (such as "navigation start-navigation exit conflict") in the record, and save the conflict record for subsequent priority determination.

[0098] Among them, according to the modal conflict record, the priority is determined by combining the vehicle running state parameter, when the vehicle is in high-speed driving state, the voice interaction data is retained, when the vehicle is in low-speed driving or parking state, the gesture interaction data is retained, and the interaction intention that does not conform to the vehicle state is rejected, forming the modal-filtered interaction intention set, specifically including: After generating the modal conflict record, the method combines the vehicle operating state parameters to determine the priority. The operating state parameters are provided by the vehicle speed, steering angle, and brake signal. When the vehicle is in a high-speed driving state, the voice interaction data is preferentially retained because the driver is more suitable for using voice interaction rather than gesture operation in a high-speed scenario; when the vehicle is in a low-speed driving or parking state, the gesture interaction data is preferentially retained because gesture operation is safer and more natural at this time. While determining the priority, the interaction intent that does not meet the operating state is marked as invalid and removed. Finally, the method rewrites the retained interaction intent into a set to form a modal-filtered interaction intent set for subsequent processing.

[0099] In a preferred embodiment of the application, the input source of the modal-filtered interaction intent set is determined, including: According to the input source data, it is determined whether the interaction intent comes from the driver or the passenger. When the driver input conflicts with the passenger input, the driver input is retained, and the passenger input is determined to be invalid and removed, to generate a source-filtered interaction intent set. According to the source-filtered interaction intent set, when the passenger input involves vehicle safety control, the interaction intent enters a pending state and is not directly executed, and a confirmation request is sent to the driver. After the driver confirms, the interaction intent is added to the source-filtered interaction intent set to form a second interaction intent set.

[0100] In the embodiment of the application, by determining the input source of the modal-filtered interaction intent set, the interaction intent of the driver and the passenger can be effectively distinguished, so that the key operation right is always concentrated in the hands of the driver. The method identifies the source of the interaction data by extracting the input source marker, preferentially retains the driver input when a conflict occurs, and removes the passenger input to avoid the impact of passenger misoperation on driving safety.

[0101] Next, when the input instructions of the driver and the passenger simultaneously act on the same target function, the method automatically selects the driver's instruction according to the priority. For example, when the passenger issues an instruction to "change the navigation destination", and the driver has set the current navigation path through voice, the method preferentially retains the driver's setting to ensure that the vehicle navigation does not deviate due to passenger interference.

[0102] ​Secondly, when the passenger input involves the vehicle safety control function, the method places the input in the to-be-confirmed state and sends a confirmation request to the driver. This mechanism can prevent the passenger's arbitrary operation from directly affecting the safety of the vehicle. For example, when the passenger attempts to issue an "off collision warning" instruction, the method does not immediately execute, but prompts the driver for confirmation, and only after the driver confirms will the input be added to the interaction intent set. This design can maximize driving safety while also preserving the possibility of passenger involvement in certain situations.

[0103] Among them, the input source label is extracted from the interaction intent set filtered by the mode, and the input source data is generated, specifically including: First, the input source label carried by each interaction intent in the interaction intent set filtered by the mode is extracted. The label is generated during the interaction data collection stage. For example, voice data can distinguish between driver and passenger inputs through multiple microphone arrays in the vehicle, and gesture data and expression data can be distinguished from the driver side or the passenger side through the camera position. After extraction, these input source labels are converted into input source data, which at least includes two types of labels: driver input and passenger input. In this way, the method can provide clear data basis for subsequent input conflict determination.

[0104] Among them, according to the input source data, it is determined whether the interaction intent comes from the driver or the passenger, when the driver input and the passenger input conflict, the driver input is retained, the passenger input is determined to be invalid and is removed, and the interaction intent set filtered by the source is generated, specifically including: After generating the input source data, the method compares the target functions corresponding to the driver input and the passenger input one by one. When a conflict is detected between the two, i.e., different input instructions exist for the same target function, the method will retain the driver input according to the priority rule, mark the passenger input as invalid and remove it. For example, when the driver inputs "turn on navigation" through voice input, and the passenger inputs "turn off navigation" through touch operation at the same time, the method will retain the "turn on navigation" intent of the driver and remove the "turn off navigation" intent of the passenger. Through this step, the interaction intent set filtered by the source can be formed, and the dominant position of the driver in vehicle operation can be ensured.

[0105] Among them, according to the interaction intent set filtered by the source, when the passenger input involves vehicle safety control, the interaction intent enters the to-be-confirmed state and is not directly executed, and a confirmation request is sent to the driver; after the driver confirms, the interaction intent is added to the interaction intent set filtered by the source to form a second interaction intent set, specifically including: After the driver and passenger conflict determination is completed, the method further checks the interaction intent set screened by the source. If there is a passenger input in the set and the input involves a vehicle safety control class function, such as turning off the collision warning or modifying the safety distance threshold, the input will not directly enter the execution flow, but will be placed in a pending confirmation state. The method will issue a confirmation request to the driver through voice broadcast or prompt on the central control interface. For example, the method can prompt: "Passenger requests to turn off the collision warning, please confirm whether to execute." If the driver confirms within the set confirmation time, the passenger input is added to the set as a valid interaction intent; if the driver does not confirm or explicitly refuses, the input will be permanently excluded. Finally, after the set is processed by this confirmation mechanism, a second interaction intent set is formed, ensuring that safety-related tasks are always under the control of the driver.

[0106] Embodiments of the present application also provide an intelligent cockpit interaction management system, the system comprising: A data acquisition module for receiving voice interaction data, gesture interaction data and expression interaction data collected in the cockpit, and adding source labels and timestamps to each data to obtain multi-modal interaction raw data; An intent analysis module for analyzing the intent according to the multi-modal interaction raw data, extracting the trigger condition and the target function, and forming an interaction intent candidate set; A relevance screening module for performing relevance screening according to the interaction intent candidate set, combining vehicle operating state parameters, in-vehicle noise level parameters and driver gaze area parameters, and obtaining a first interaction intent set; A conflict detection module for performing conflict detection on the first interaction intent set, determining the interaction intent with abnormal trigger condition or conflicting target function as invalid and excluding it, and obtaining a second interaction intent set; A priority classification module for performing priority classification according to the second interaction intent set, generating a task candidate queue, wherein safety-related target functions are set as high priority, navigation target functions are set as medium priority, and entertainment or information display target functions are set as low priority; A task confirmation module for performing task confirmation on the task candidate queue, generating an effective task queue for tasks that meet the trigger condition; A task scheduling module for performing task scheduling according to the effective task queue, interrupting low-priority tasks and switching execution when a higher-priority task input is detected.

[0107] It should be noted that the system corresponds to the above method, and all implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0108] The embodiment of the present application also provides a computing device, comprising a processor, a memory storing a computer program, the computer program being executed by the processor to perform the method as described above. All implementation manners in the above method embodiment are suitable for this embodiment and can achieve the same technical effects.

[0109] The embodiment of the present application also provides a computer readable storage medium storing instructions, which, when executed on a computer, cause the computer to perform the method as described above. All implementation manners in the above method embodiment are suitable for this embodiment and can achieve the same technical effects.

[0110] The above is the preferred embodiment of the present application, it should be pointed out that, for those skilled in the art, without departing from the principles of the present application, can make a number of improvements and refinements, these improvements and refinements should also be considered as the protection scope of the present application.

Claims

1. A smart cockpit interaction management method, characterized in that, The method comprises: Receiving voice interaction data, gesture interaction data and expression interaction data collected in the cabin, and adding source marks and time stamps to each piece of data to obtain multi-modal interaction original data; According to the multi-modal interaction original data, intention analysis is performed to extract trigger conditions and target functions, and an interaction intention candidate set is formed; According to the interaction intention candidate set, in combination with vehicle running state parameters, in-cabin noise level parameters and driver gaze area parameters, correlation screening is performed to obtain a first interaction intention set; For the first interaction intention set, conflict detection is performed, and interaction intentions with abnormal trigger conditions or conflicting target functions are determined as invalid and removed to obtain a second interaction intention set; According to the second interaction intention set, priority classification is performed according to the categories of target functions to generate a task candidate queue; Task confirmation is performed on the task candidate queue, and tasks that meet the trigger conditions are generated to form an effective task queue; According to the effective task queue, task scheduling is performed, and when a higher priority task input is detected, a low priority task is interrupted and execution is switched. 2.The intelligent cabin interaction management method of claim 1, wherein, According to the interaction intention candidate set, in combination with vehicle running state parameters, in-cabin noise level parameters and driver gaze area parameters, correlation screening is performed to obtain a first interaction intention set, comprising: According to the interaction intention candidate set and the in-cabin noise level parameter, continuity comparison is performed, when the noise level continuously exceeds the preset noise threshold within the preset time window, the corresponding voice interaction data is determined as invalid and removed to generate a first candidate set; According to the first candidate set and the driver gaze area parameter, duration determination is performed, when the driver's gaze direction continuously remains deviated from the central control interaction interface by a preset angle range for more than a preset first time length, the corresponding gesture interaction data is determined as invalid and removed to generate a second candidate set; According to the second candidate set and the vehicle running state parameter, subdivision determination is performed, when the running state is high-speed straight driving, safety control related interaction intentions are preferentially retained, when the running state is low-speed driving, navigation and safety related interaction intentions are retained, and when the running state is parking, entertainment related interaction intentions are allowed to be retained, finally generating the first interaction intention set. 3.The intelligent cabin interaction management method of claim 1, wherein, For the first interaction intention set, conflict detection is performed, and interaction intentions with abnormal trigger conditions or conflicting target functions are determined as invalid and removed to obtain a second interaction intention set, comprising: According to the time stamps of each interaction intention in the first interaction intention set, comparison is performed, when the time stamps are discontinuous or exceed the preset time interval, the interaction intention is determined as invalid and removed, and an interaction intention set screened by time is generated; According to the trigger conditions of different modalities in the interaction intention set screened by time, consistency determination is performed, when the voice interaction data and the gesture interaction data have consistent target functions but conflicting trigger conditions, the interaction intention matching the vehicle running state parameter is preferentially retained, and the interaction intention not meeting the condition is removed, and an interaction intention set screened by modalities is generated; According to the input source of the modal filtered interaction intention set, when the driver input and the passenger input have target conflicts, the interaction intention of the driver input is preferentially retained, and the interaction intention of the passenger input is removed, to generate a second interaction intention set. 4.The intelligent cabin interaction management method of claim 1, wherein, According to the second interaction intention set, the target functions are classified according to priority, to generate a task candidate queue, including: According to the initial classification of the target functions of the second interaction intention set, the interaction intentions related to safety prompts and safety control are divided into high priority, the interaction intentions related to navigation setting and path planning are divided into medium priority, and the interaction intentions related to entertainment playing and information display are divided into low priority, to generate an initial task candidate queue; According to the initial task candidate queue and the vehicle running state parameter, when the running state is emergency braking or collision warning, all non-safety related tasks are downgraded as a whole, to generate a dynamically adjusted task candidate queue; According to the dynamically adjusted task candidate queue, the execution order is formed according to the priority, to obtain a final task candidate queue. 5.The intelligent cabin interaction management method of claim 2, wherein, According to the continuous comparison of the interaction intention candidate set and the in-vehicle noise level parameter, including: According to the voice interaction data extracted from the interaction intention candidate set, the in-vehicle noise level parameter is sampled in a preset time window, to generate a noise change sequence; According to the continuous threshold judgment of the noise change sequence, when the time period during which the noise continuously exceeds the preset noise threshold is greater than a preset second time length, the corresponding voice interaction data is judged as invalid and removed, to obtain voice interaction data; According to the voice interaction data, when the time period is less than the preset time length, the corresponding voice interaction data is retained, to form a first candidate set. 6.The intelligent cabin interaction management method of claim 2, wherein, According to the duration judgment of the first candidate set and the driver gaze area parameter, including: According to the gesture interaction data extracted from the first candidate set, the driver gaze area parameter is compared frame by frame, to generate a gaze deviation record; According to the gaze deviation record, the continuous deviation time is counted, and when the continuous deviation time exceeds a preset first time length, the corresponding gesture interaction data is judged as invalid and removed, to obtain gesture interaction data; According to the gesture interaction data, when the continuous deviation time is less than the preset first time length, the corresponding gesture interaction data is retained, to form a second candidate set.

7. The intelligent cabin interaction management method of claim 2, wherein, According to the subdivision judgment of the second candidate set and the vehicle running state parameter, including: According to the second candidate set, the vehicle speed, the steering angle and the braking state in the vehicle running state parameter are combined to form a running state combination; According to the running state combination, it is judged that the vehicle is in a high-speed straight driving, low-speed driving or parking state, and the interaction intentions corresponding to different states are classified and processed, to generate an interaction intention set subdivided according to the running state; According to the interaction intention set subdivided according to the running state, the safety control related interaction intention is retained when driving at high speed, the navigation related and safety control related interaction intentions are retained when driving at low speed, and the entertainment related interaction intention is retained when parking, to form a first interaction intention set. 8.The intelligent cabin interaction management method of claim 3, wherein, According to the comparison of the time stamps of the interaction intentions in the first interaction intention set, including: According to the first interaction intention set, the timestamp of each interaction intention is extracted, the time interval of adjacent interaction intentions is calculated, and time interval data is generated; According to the time interval data, threshold determination is performed, when the time interval is greater than the preset time interval, the interaction intention is determined as abnormal and is eliminated, forming an interaction intention set filtered by interval; According to the interaction intention set filtered by interval, time window determination is performed, when the interaction intention does not fall into the preset time window, the interaction intention is determined as invalid and is eliminated, generating an interaction intention set filtered by time. 9.The intelligent cabin interaction management method of claim 3, wherein, According to the trigger conditions of different modalities in the interaction intention set filtered by time, consistency determination is performed, including: According to the interaction intention set filtered by time, voice interaction data and gesture interaction data are extracted, and trigger condition pairs are generated; According to the trigger condition pairs, consistency comparison is performed, when the target function is consistent but the trigger condition is inconsistent, a modal conflict record is generated; According to the modal conflict record, priority determination is performed in combination with vehicle operating state parameters, when the vehicle is in high-speed driving state, voice interaction data is retained, when the vehicle is in low-speed driving or parking state, gesture interaction data is retained, and interaction intentions that do not meet the vehicle state are eliminated, forming an interaction intention set filtered by modality. 10.The intelligent cabin interaction management method of claim 3, wherein, According to the input source of the interaction intention set filtered by modality, determination is performed, including: According to the interaction intention set filtered by modality, input source markers are extracted, and input source data is generated; According to the input source data, it is determined whether the interaction intention comes from the driver or the passenger, when the driver input and the passenger input conflict, the driver input is retained, the passenger input is determined invalid and is eliminated, generating an interaction intention set filtered by source; According to the interaction intention set filtered by source, when the passenger input involves vehicle safety control, the interaction intention enters a to-be-confirmed state and is not directly executed, and a confirmation request is sent to the driver; After the driver confirms, the interaction intention is added to the interaction intention set filtered by source, forming a second interaction intention set.

Citation Information

Patent Citations

  • A multi-mode depth fusion airborne cockpit man-machine interaction method

    CN109933272A

  • Heterogeneous data acquisition and interaction method and device, equipment and storage medium

    CN115391444A

  • Interaction method and device, electronic equipment and computer readable storage medium

    CN120596202A

  • Apparatus for predicting intention of user using multi modal information and method thereof

    KR1020110009614A

  • Control device and method with voice and / or gestural recognition for the interior lighting of a vehicle

    US20170270924A1