An intelligent cockpit multi-modal voice interaction system and method

By combining environmental audio, video, and in-vehicle perception parameters into a multimodal voice interaction system, the problems of voice recognition errors and operation conflicts have been solved, achieving accurate response and improved safety in complex environments.

CN121034318BActive Publication Date: 2026-05-01SHENZHEN SHENHANG HUACHUANG AUTOMOBILE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SHENHANG HUACHUANG AUTOMOBILE TECH CO LTD
Filing Date
2025-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing intelligent cockpit multimodal voice interaction systems are prone to voice recognition errors in environments with low signal-to-noise ratios or when the driver's lip movements are partially obscured. Furthermore, they lack the ability to handle semantic overlap, operational exclusivity, or inconsistencies in modal information between different candidate intentions, leading to execution errors or operational conflicts and increasing driving risks.

Method used

By collecting environmental audio and video information and combining it with vehicle interior environmental perception parameters, a suitable voice interaction input signal is generated. By combining the correspondence between speech phonemes and lip movements, a multimodal segment sequence is constructed. The final execution instruction is selected through conflict resolution functions and safety constraint rules to ensure that the system responds accurately in complex environments and reduces the risk of misoperation.

Benefits of technology

It effectively avoids accidental triggering of voice interaction mode, reduces recognition errors, shortens response time, reduces driving risks, improves system safety and controllability, and ensures accurate response in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034318B_ABST
    Figure CN121034318B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of speech processing, and discloses an intelligent cockpit multi-modal speech interaction system and method, which comprises a speech triggering unit, collects environmental audio and video information in the cockpit, combines with environmental perception parameters in the vehicle, judges whether to enter a speech interaction mode, generates a speech interaction input signal adapted to the current environment when the speech interaction triggering condition is established, a mouth shape analysis unit extracts acoustic features of the speech interaction input signal, synchronously analyzes the lip movement track of the driver in the video information, establishes the corresponding relationship between the speech phoneme and the mouth shape movement, forms joint analysis features, a candidate generation unit segments and aligns the joint analysis features, constructs a continuous multi-modal segment sequence, synchronizes the multi-modal segment sequence in time, projects to a predefined intention space, and further obtains a candidate intention set containing different candidate intentions, and improves the human-computer interaction experience of the intelligent cockpit.
Need to check novelty before this filing date? Find Prior Art

Description

A multimodal voice interaction system and method for intelligent cockpits Technical Field

[0001] This invention relates to the field of voice processing technology, and more specifically, to a smart cockpit multimodal voice interaction system and method. Background Technology

[0002] Existing multimodal voice interaction systems and methods for intelligent cockpits mainly suffer from the following problems:

[0003] In existing multimodal voice interaction systems for smart cockpits, voice command recognition typically relies on a single modality, such as audio or video signals, for candidate intent generation and final operational decisions. Such methods are prone to voice recognition errors in environments with low signal-to-noise ratios or when the driver's lip movements are partially obscured, leading to inaccurate candidate intents.

[0004] Traditional systems often employ the highest confidence principle to directly execute operations during candidate intent screening, neglecting potential semantic overlap, operational exclusivity, or modal inconsistencies between different candidate intents. When conflicts exist between multimodal information, the system cannot dynamically adjust or resolve these conflicts, easily leading to execution errors or operational conflicts. Existing technologies lack safety constraint assessments of dynamic cockpit context information, potentially resulting in inappropriate operations during dangerous driving or special cockpit conditions, increasing driving risks.

[0005] In view of this, the present invention proposes an intelligent cockpit multimodal voice interaction system and method to solve the above problems. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a smart cockpit multimodal voice interaction system, comprising:

[0007] The voice trigger unit collects ambient audio and video information in the cabin and combines it with the environmental perception parameters inside the vehicle to determine whether to enter the voice interaction mode. When the voice interaction trigger condition is met, it generates a voice interaction input signal that is adapted to the current environment.

[0008] The lip-reading analysis unit extracts acoustic features from the voice interaction input signal, simultaneously analyzes the driver's lip movement trajectory in the video information, establishes the correspondence between speech phonemes and lip movements, and forms joint analysis features;

[0009] The candidate generation unit segments and aligns the joint parsed features to construct a continuous multimodal segment sequence; by synchronizing the multimodal segment sequence in time and projecting it onto a predefined intent space, a set of candidate intents containing different candidate intents is obtained.

[0010] The intent arbitration unit, based on the candidate intent set, executes predefined dynamic rule arbitration logic, filters candidate intents layer by layer, and determines the final execution instruction;

[0011] The collaborative response unit controls the internal devices of the intelligent cockpit to perform corresponding operations according to the final execution command, updates the cockpit dynamic context information, generates multimodal collaborative response signals and outputs them.

[0012] Specifically, the environmental perception parameters include:

[0013] Background noise level (decibels), microphone array status, real-time vehicle speed, vehicle operation status, number and distribution of people inside the cockpit, operating status of cockpit equipment, light intensity inside the cockpit, and camera field of view.

[0014] Specifically, the method for acquiring the voice interaction input signal includes:

[0015] The system collects ambient audio and video information of the driver's face, and combines this with environmental perception parameters inside the vehicle for comprehensive evaluation to determine whether to enter the voice interaction mode. In the ambient audio, it detects whether there are audio segments that match the preset voice interaction wake-up words with a preset matching degree threshold. The audio segments that match the preset matching degree threshold are extracted and further generated into an audio trigger candidate set.

[0016] Extract the start and end times of all audio segments in the audio trigger candidate set as the audio trigger candidate time window; in the video information, select a time period with the same start and end time range as the audio trigger candidate time window as the video trigger candidate time window, and extract video frames within the video trigger candidate time window in which the driver faces the camera and has a lip opening and closing trajectory as the video trigger candidate set.

[0017] Using environmental perception parameters as constraints, the audio trigger candidate set and the video trigger candidate set are verified synchronously over a time period to determine whether the audio trigger and the video trigger are valid within the same time period. When the audio trigger candidate set and the video trigger candidate set are valid simultaneously within the same time period, it is determined that the voice interaction mode is entered, and a voice interaction input signal adapted to the current environment is generated.

[0018] Specifically, the method for extracting acoustic features from voice interaction input signals includes:

[0019] The voice interaction input signal is divided into frames according to a preset frame length and frame shift, and a window function is applied to each frame. The windowed voice interaction input signal is then subjected to a fast Fourier transform to obtain the spectral representation of the voice interaction input signal. Acoustic features are then calculated based on the spectral representation.

[0020] Specifically, the method for obtaining the joint parsing features includes:

[0021] Face detection is performed on the collected video information to obtain the driver's face bounding box; within the driver's face bounding box, the positions of key points in the lip region are located based on facial key points; the lip region boundary is constructed based on the two-dimensional coordinates of the lip region key point positions; the lip region boundary positions of continuous video frames in the video information are tracked, and geometric normalization is performed on the tracked lip region boundaries to convert them into a continuous lip image time series; lip motion trajectory features are extracted from the lip image time series to form a lip motion trajectory sequence.

[0022] The acquired acoustic features are aligned with the execution time of the lip movement trajectory sequence to establish the correspondence between speech phonemes and lip movements, forming joint analytical features.

[0023] Specifically, the method for constructing the multimodal fragment sequence includes:

[0024] The joint analytical features are segmented and divided into different multimodal segments based on continuous time windows. The joint analytical features within each multimodal segment are then aggregated to obtain a multimodal segment sequence.

[0025] Specifically, the method for obtaining the candidate intent set includes:

[0026] The multimodal segment sequence is time-aligned on a global time scale. The predefined intent space is a set of semantic categories, and each intent corresponds to a multimodal feature distribution region. The time-aligned multimodal segment sequence is projected into the predefined intent space to obtain the candidate intent and initial confidence corresponding to each multimodal segment. All projection results are collected to generate a candidate intent set containing different candidate intents.

[0027] Specifically, the method for determining the final execution instruction includes:

[0028] For the candidate intent set, calculate the overall confidence of each candidate intent; after the candidate intent is generated, introduce a conflict resolution function to adjust the overall confidence of candidate intents that may have conflicts, so that the overall confidence of candidate intents with high conflict intensity decreases.

[0029] Based on security constraint rules, it is determined whether the candidate intent is allowed to be executed under the current environmental perception parameters; the candidate intents that are allowed to be executed are screened layer by layer, and the candidate intent with the highest comprehensive confidence and that meets the security constraint rules is selected as the final execution instruction.

[0030] Specifically, the method for generating and outputting multimodal cooperative response signals includes:

[0031] Based on the final execution command, control signals are sent to the devices inside the smart cockpit, causing the corresponding devices to perform corresponding operations. The operation results of the corresponding devices are recorded in real time, and corresponding multimodal collaborative response signals are generated and output through the smart cockpit's interaction interface.

[0032] A multimodal voice interaction method for intelligent cockpits includes:

[0033] S1. Collect ambient audio and video information in the cabin, combine it with the environmental perception parameters inside the vehicle, and determine whether to enter the voice interaction mode. When the voice interaction trigger condition is met, generate a voice interaction input signal that is adapted to the current environment.

[0034] S2. Extract acoustic features from the voice interaction input signal, simultaneously analyze the driver's lip movement trajectory in the video information, establish the correspondence between speech phonemes and mouth movements, and form joint analytical features;

[0035] S3. Segment and align the joint parsed features to construct a continuous multimodal segment sequence; synchronize the multimodal segment sequence in time and project it onto a predefined intent space to obtain a set of candidate intents containing different candidate intents;

[0036] S4. Based on the candidate intent set, execute the predefined dynamic rule arbitration logic to filter candidate intents layer by layer and determine the final execution instruction;

[0037] S5. Based on the final execution command, control the internal devices of the intelligent cockpit to perform corresponding operations, update the cockpit dynamic context information, generate multimodal collaborative response signals and output them.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] This invention combines environmental audio, video, and vehicle interior environmental perception parameters to form a triple judgment mechanism, which can effectively avoid accidental triggering of voice interaction mode, while ensuring accurate response in complex environments. It compensates for the shortcomings of single voice recognition, and assists in correcting recognition results and reducing recognition errors in situations such as excessive noise, unclear pronunciation, or heavy dialect. The precise triggering logic avoids unnecessary interaction wake-ups, reducing driver distraction due to irrelevant prompts. At the same time, it quickly locks into valid interaction intentions through multimodal information, shortens system response preparation time, and allows drivers to quickly enter the process when interaction is needed, reducing the time spent with hands off the steering wheel or eyes off the road, and reducing driving risks.

[0040] By adjusting the overall confidence of candidate intentions through conflict resolution functions, potential conflicting intentions can be actively suppressed, making the final executed instructions more reasonable and avoiding erroneous operations caused by semantic or operational mutual exclusion. By associating real-time environmental perception parameters with candidate intentions through safety constraint rules, only candidate intentions that meet the safety constraint rules can be executed, fundamentally avoiding improper operations in dangerous driving scenarios, improving the safety and controllability of the system, and realizing intelligent, highly reliable and safe multimodal voice interaction. Attached Figure Description

[0041] Figure 1 is a schematic diagram of the structure of a smart cockpit multimodal voice interaction system according to the present invention;

[0042] Figure 2 is a schematic diagram of the process of a multimodal voice interaction method for an intelligent cockpit according to the present invention. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Example 1

[0045] Please refer to Figure 1. This embodiment provides an intelligent cockpit multimodal voice interaction system, which specifically includes the following steps:

[0046] The voice trigger unit collects ambient audio and video information in the cabin and combines it with the environmental perception parameters inside the vehicle to determine whether to enter the voice interaction mode. When the voice interaction trigger condition is met, it generates a voice interaction input signal that is adapted to the current environment.

[0047] The lip-reading analysis unit extracts acoustic features from the voice interaction input signal, simultaneously analyzes the driver's lip movement trajectory in the video information, establishes the correspondence between speech phonemes and lip movements, and forms joint analysis features;

[0048] The candidate generation unit segments and aligns the joint parsed features to construct a continuous multimodal segment sequence; by synchronizing the multimodal segment sequence in time and projecting it onto a predefined intent space, a set of candidate intents containing different candidate intents is obtained.

[0049] The intent arbitration unit, based on the candidate intent set, executes predefined dynamic rule arbitration logic, filters candidate intents layer by layer, and determines the final execution instruction;

[0050] The collaborative response unit controls the internal devices of the intelligent cockpit to perform corresponding operations according to the final execution command, updates the cockpit dynamic context information, generates multimodal collaborative response signals and outputs them.

[0051] Environmental perception parameters include:

[0052] In a multimodal voice interaction system for intelligent cockpits, environmental perception parameters refer to key parameters that reflect the state of the vehicle's interior and surrounding environment and may affect the accuracy of voice interaction triggering or processing effectiveness. These environmental perception parameters can be collected in real-time by various sensors deployed within the vehicle's cockpit, such as microphone arrays, cameras, and CAN bus sensors, to assist in determining whether to enter voice interaction mode. These parameters include the following:

[0053] Background noise level (decibels), microphone array status, real-time vehicle speed, vehicle operation status, number and distribution of people inside the cockpit, operating status of cockpit equipment, light intensity inside the cockpit, and camera field of view.

[0054] Background noise decibel level: The background noise intensity (such as engine noise at idle, tire / wind noise while driving, air conditioning noise, etc.) is collected in real time by the microphone array in the smart cockpit to determine whether the current environment is suitable for voice recognition; Microphone array status: Used to determine whether the microphone is blocked or malfunctioning (such as signal interruption) to ensure the effectiveness of the voice acquisition hardware; Real-time vehicle speed: Voice interaction needs are higher at low speeds (such as when parked or in congestion), and wind / tire noise is greater at high speeds, so the noise reduction strategy needs to be adjusted in a timely manner according to the real-time vehicle speed; Vehicle operation status: By observing whether the vehicle is turning (steering wheel angle sensor) or shifting gears (transmission signal), it is determined whether the driver is in a "busy state" to avoid accidental triggering of voice interaction;

[0055] Number and distribution of occupants inside the cabin: Determined by seat pressure sensors, infrared sensors, or cameras (e.g., whether there is a driver / passenger in the front row, or whether there are people in the back row) to avoid accidental interaction triggered by the voices of non-target personnel (e.g., conversations in the back row); Operating status of cabin equipment reflects the status of electrical equipment in the vehicle, and the resulting operating noise may interfere with the recognition of voice signals; Intensity of light inside the cabin: For example, whether it is in a backlit state (direct sunlight on the driver's face) to avoid strong light causing facial shadows to obscure the lips and affect the analysis of lip movement trajectory; Camera field of view: Determined whether the camera is obstructed (e.g., the driver's hands, decorations) to ensure that the lip recognition area is unobstructed and to ensure the effectiveness of video information;

[0056] Environmental perception parameters ultimately provide a comprehensive environmental assessment basis for the voice triggering unit, ensuring that the voice interaction mode is triggered at the appropriate time (such as when the ambient noise is low and the driver's intention is clear), thereby improving the accuracy of the interaction.

[0057] Methods for acquiring voice interaction input signals include:

[0058] By collecting ambient audio and driver facial video information through microphone arrays and cameras deployed in the cockpit, and combining this with environmental perception parameters inside the vehicle for comprehensive evaluation, a judgment is made on whether to enter voice interaction mode. In the ambient audio, it detects whether there are audio segments whose matching degree with the preset voice interaction wake-up words exceeds a preset matching degree threshold. Audio segments exceeding the preset matching degree threshold are extracted and further generated into an audio trigger candidate set. The preset matching degree threshold can be set by staff based on historical data analysis results or expert experience.

[0059] It should be noted that the preset voice interaction wake-up words are derived from a pre-stored set of voice interaction wake-up words used to trigger in-vehicle voice interaction modes. These words are characterized by high distinguishability (unique pronunciation, not easily confused with ordinary speech), strong relevance to vehicle operation, and international definition capabilities. They can include fixed phrases with triggering attributes such as "Hello Xiaozhi," "Open Voice," "Start Command," or "Hi Drive." The setting of preset voice interaction wake-up words can be based on the driver's language habits, human-machine interaction design specifications, or industry standards, ensuring high recognizability and trigger stability in driving scenarios. The matching degree of preset voice interaction wake-up words can be determined by the probability of candidate voice interaction wake-up words output by a neural network model, or by calculating the cosine similarity between an audio segment and the preset voice interaction wake-up words.

[0060] Extract the start and end times of all audio segments in the audio trigger candidate set as the audio trigger candidate time window; in the video information, select a time period with the same start and end time range as the audio trigger candidate time window as the video trigger candidate time window, and extract video frames within the video trigger candidate time window in which the driver faces the camera and has a lip opening and closing trajectory as the video trigger candidate set.

[0061] Using environmental perception parameters as constraints, the audio trigger candidate set and the video trigger candidate set are verified synchronously over a time period to determine whether the audio trigger and the video trigger are valid within the same time period. When the audio trigger candidate set and the video trigger candidate set are valid within the same time period, it is determined that the voice interaction mode is entered and a voice interaction input signal adapted to the current environment is generated.

[0062] It should be noted that after determining that the voice interaction mode has been entered, the audio trigger candidate set and video trigger candidate set that are simultaneously established within the same time period are aligned in the time domain and frame level, noise interference segments and non-speaking segments are removed, and a voice interaction input signal is generated. Specifically, the voice interaction input signal includes the effective speech segment after noise suppression and time domain segmentation, the driver's lip movement video segment corresponding to the timestamp of the effective speech segment, and the corresponding environmental perception parameters (such as noise, light intensity, driving status, etc.).

[0063] Methods for extracting acoustic features from voice interaction input signals include:

[0064] The voice interaction input signal is divided into frames according to a preset frame length and frame shift, and a window function is applied to each frame. The windowed voice interaction input signal is then subjected to a fast Fourier transform to obtain the spectral representation of the voice interaction input signal. Acoustic features are then calculated based on the spectral representation.

[0065] It should be noted that in speech signal processing, frame segmentation, applying window functions, and fast Fourier transform are the core steps in converting time-domain speech signals into frequency-domain features. While the speech interaction input signal is a time-varying, non-stationary signal, it can be approximated as a stationary signal within a short time frame (e.g., 20-30ms). Therefore, frame segmentation is necessary to divide the continuous speech interaction input signal into short segments according to a preset frame length and frame shift (the preset frame length and frame shift can be set using expert experience). Within each short segment, the speech interaction input signal can be approximated as having short-term stationarity, thus ensuring that subsequent time-frequency domain feature extraction has a valid physical basis and numerical stability.

[0066] However, frame-segmentation can cause abrupt changes in the voice interaction input signal at frame edges. This can be mitigated by applying a window function to smooth the edges and reduce spectral leakage in the subsequent Fast Fourier Transform (FFT) (i.e., the phenomenon of signal energy spreading to non-real frequencies). Commonly used window functions include Hanning windows, Hamming windows, and rectangular windows, which can be configured based on task requirements and computational costs. The framed voice interaction input signal is multiplied point by point with the window function to obtain the windowed voice interaction input signal. Subsequently, a Fast Fourier Transform is performed on the windowed voice interaction input signal to map it from the time domain to the frequency domain, thereby obtaining the spectral representation of the voice interaction input signal.

[0067] Statistical features are extracted from the spectral representation, such as spectral energy, zero-crossing rate, spectral centroid, spectral bandwidth, spectral flatness, and spectral roll-off point. Mel frequency cepstral coefficients are extracted through spectral transformation. The statistical features and Mel frequency cepstral coefficients are combined to obtain acoustic features.

[0068] Methods for obtaining joint parsed features include:

[0069] Face detection is performed on the collected video information to obtain the driver's face bounding box. For example, the driver's face bounding box can be obtained by using the MTCNN model for face boundary recognition. Within the driver's face bounding box, the positions of key points in the lip region are located based on facial key points. The positions of key points in the lip region include the upper and lower corners of the mouth, the upper and lower lip edges, the upper middle lip point, and the lower middle lip point. The lip region boundary is constructed based on the two-dimensional coordinates of the key point positions in the lip region (which can be determined by the circumscribed rectangle or the convex hull polygon of the key points). The lip region boundary positions in continuous video frames in the video information are tracked, and the tracked lip region boundaries are geometrically normalized to convert them into a continuous lip image time series. The lip image time series is then used to extract lip motion trajectory features (such as lip height, lip width, opening degree, and changes in lip area over time) to form a lip motion trajectory sequence.

[0070] The acquired acoustic features are aligned with the execution time of the lip movement trajectory sequence to establish the correspondence between speech phonemes and lip movements, forming joint analytical features.

[0071] Methods for constructing multimodal fragment sequences include:

[0072] The joint analytical features are segmented and divided into different multimodal segments based on continuous time windows. The joint analytical features within each multimodal segment are then aggregated to obtain a multimodal segment sequence.

[0073] It should be noted that, in order to group joint analytic features of different modalities that are synchronous in time into the same multimodal segment, segmentation processing based on continuous time windows is performed on the joint analytic features. This can be achieved by presetting the time window length and the overlap ratio between adjacent time windows, and by dividing the joint analytic features along the time axis according to timestamps. This ensures that joint analytic features falling within the same time window are grouped into the same multimodal segment. The preset time window length can be preset based on expert experience and adjusted later according to technical requirements.

[0074] Methods for obtaining the candidate intent set include:

[0075] The multimodal segment sequence is time-aligned on a global time scale (e.g., using timestamp-based interpolation resampling) to further eliminate time offsets caused by inconsistent connections between multimodal segments, ensuring that the entire multimodal segment sequence has consistent time reference globally. A predefined intent space is a set of semantic categories (e.g., "adjust air conditioning temperature", "play music", "answer phone call"), with each intent corresponding to a multimodal feature distribution region. The time-aligned multimodal segment sequence is projected into the predefined intent space to obtain the candidate intent and initial confidence for each multimodal segment. All projection results are collected to generate a candidate intent set containing different candidate intents.

[0076] It should be noted that the predefined intent space refers to the set of semantic categories formed by prior modeling of the types of intents that can be understood and executed by the system in the in-cabin human-computer interaction scenario. The prior modeling of the semantic category set can be based on functional requirement modeling, automatic extraction from corpus statistics, and filtering by system policy constraints. It is determined by a combination of functional support range, data statistics, and security policy constraints. After projecting the time-aligned multimodal segment sequence onto the predefined intent space, initial confidence scores can be assigned to each multimodal segment using metrics such as KL divergence or Mahalanobis distance to indicate the degree of proximity between the current multimodal segment and the intent.

[0077] Functional requirement-based modeling refers to: based on the actual functions supported by the intelligent cockpit multimodal voice interaction system, manually enumerating intent categories in advance, such as air conditioning control (heating, cooling, switching modes, etc.), entertainment control (play, pause, skip songs, adjust volume, etc.), communication (answer, reject, dial), navigation (set destination, cancel navigation, etc.).

[0078] Automatic extraction based on corpus statistics refers to: clustering or extracting topics from labeled voice commands, lip movements, gesture trajectories and trigger intention results in historical voice interaction logs, automatically learning the distribution clusters of intentions that occur frequently, and forming a set of candidate intention categories;

[0079] System policy constraint-based filtering refers to: performing policy filtering on the candidate intent set (such as tailoring according to safety rules, driving status, and regulatory constraints) to remove intent categories that are not allowed or unnecessary to be executed in the current environment, thus forming an effective intent space.

[0080] Methods for determining the final instruction to be executed include:

[0081] For the candidate intent set, calculate the overall confidence score for each candidate intent; the overall confidence score is: ;in, Indicates candidate intent The overall confidence level; Indicates candidate intent The corresponding speech modality confidence; Indicates candidate intent The corresponding video modal confidence; The dynamic weights representing the confidence level of the speech modality; The dynamic weights representing the video modality confidence; An index representing the candidate intent;

[0082] It should be noted that: comparing candidate intents The similarity between the corresponding acoustic features and the preset speech feature template is calculated and normalized to obtain the candidate intent. Corresponding speech modal confidence; comparing candidate intents The similarity between the corresponding lip movement trajectory and the preset lip movement reference template is calculated and normalized to obtain candidate intents. The corresponding video modal confidence;

[0083] After candidate intents are generated, a conflict resolution function is introduced to adjust the overall confidence of candidate intents that may have conflicts, so that the overall confidence of candidate intents with high conflict intensity decreases.

[0084] It should be noted that potential conflicts between candidate intents may include: semantic overlap or operational exclusivity: the operations corresponding to different candidate intents may be functionally mutually exclusive, such as "start navigation" and "cancel navigation"; and inconsistent modal information: the speech modality may be biased towards the intent. The video modality is biased towards intent. This leads to multiple candidate intentions having similar overall confidence levels, resulting in potential selection conflicts; safety constraint conflicts: some candidate intentions may be unsafe or unexecutable in the current driving environment, such as entertainment operations while driving at high speeds.

[0085] Conflict intensity is used to characterize the degree to which two or more candidate intentions cannot be simultaneously established or are logically mutually exclusive within the same time period, while confidence decrease is a quantitative feedback of conflict intensity, used to suppress the probabilistic output of unreasonable candidate intentions and ensure that the output candidate intentions are executable.

[0086] Conflict intensity can be calculated based on one or a combination of the following indicators: semantic mutual exclusion (e.g., "turn on the air conditioner" and "turn off the air conditioner" are strictly mutually exclusive, with a conflict intensity of 1, indicating a strong conflict; "adjust the temperature" and "play music" can be performed in parallel, with a conflict intensity of 0, indicating no conflict), the degree of competition for operational resources (e.g., if two intentions require the same interaction channel, such as the navigation interface occupying the screen, there is a resource conflict, and the intensity can be quantified by sharing rate or exclusivity), and the harm of behavior execution (e.g., when driving at high speed, the simultaneous occurrence of "opening a large number of pop-up windows to display settings" and "adjusting the navigation route" is considered a high-risk conflict, and the conflict intensity is amplified). Conflict intensity can be formalized as a conflict coefficient with a value range between [0,1]. The larger the conflict coefficient, the higher the degree of mutual exclusion.

[0087] When there is a conflict among multiple candidate intentions, the monotonicity principle can be used to penalize and decay the overall confidence based on the conflict coefficient. The larger the conflict coefficient, the greater the decay, which in turn reduces the overall confidence of the candidate intention with strong conflict.

[0088] The conflict resolution function is: ;in, Indicates the adjusted candidate intent The overall confidence level; Representing candidate intent in speech modality The intensity of conflict with other candidate intentions; Indicates candidate intent in lip shape modality The intensity of conflict with other candidate intentions;

[0089] Based on security constraint rules, it is determined whether the candidate intent is allowed to be executed under the current environmental perception parameters; the candidate intents that are allowed to be executed are screened layer by layer, and the candidate intent with the highest comprehensive confidence and that meets the security constraint rules is selected as the final execution instruction.

[0090] The safety constraint rules are as follows: ;in, Indicates the final operation to be performed; This represents the candidate intent with the highest overall confidence level; This indicates that execution is prohibited and will not trigger any actions within the cockpit. Represents the safety constraint function (Boolean function); Indicates candidate intent Allowed to execute; Indicates candidate intent Enforcement is prohibited;

[0091] It should be noted that the essential meaning of security constraint rules is: first find the candidate intent with the highest overall confidence. Then judge In the current environment, is execution allowed? If so, what are the security constraint functions? This indicates that the intention is to be executed as the final instruction; if not, the security constraint function... This indicates that the intention is directly prohibited;

[0092] For example: Safety suppression in high-speed driving scenarios: Current vehicle state parameters: driving on a highway, vehicle speed = 110km / h, only the driver is in the cabin; the candidate intent with the highest overall confidence is: turn on the central control video to play entertainment programs; at this time, the safety constraint function determines: entertainment operations that affect the driver's attention are prohibited in the current high-speed driving state, so the output result is not to execute the candidate intent and not to trigger any execution actions in the cabin.

[0093] Methods for generating and outputting multimodal cooperative response signals include:

[0094] Based on the final execution command, control signals are sent to the devices inside the smart cockpit, causing the corresponding devices to perform corresponding operations, such as navigation, playing music, adjusting the air conditioning, and adjusting the seats. The operation results of the corresponding devices are recorded in real time, and corresponding multimodal collaborative response signals are generated and output through the interactive interface of the smart cockpit. The multimodal collaborative response signals include at least one of voice feedback, visual feedback, and motion feedback to achieve synchronous output of feedback on the final execution command results and user perception.

[0095] This embodiment combines environmental audio, video, and vehicle interior environmental perception parameters to form a triple judgment mechanism, which can effectively avoid accidental triggering of voice interaction mode, while ensuring accurate response in complex environments. It compensates for the shortcomings of single voice recognition, and assists in correcting recognition results and reducing recognition errors in situations such as excessive noise, unclear pronunciation, or heavy dialect. The precise triggering logic avoids unnecessary interaction wake-ups, reducing driver distraction due to irrelevant prompts. At the same time, it quickly locks into valid interaction intentions through multimodal information, shortens system response preparation time, and allows drivers to quickly enter the process when interaction is needed, reducing the time spent with hands off the steering wheel or eyes off the road, and reducing driving risks.

[0096] By adjusting the overall confidence of candidate intentions through conflict resolution functions, potential conflicting intentions can be actively suppressed, making the final executed instructions more reasonable and avoiding erroneous operations caused by semantic or operational mutual exclusion. By associating real-time environmental perception parameters with candidate intentions through safety constraint rules, only candidate intentions that meet the safety constraint rules can be executed, fundamentally avoiding improper operations in dangerous driving scenarios, improving the safety and controllability of the system, and realizing intelligent, highly reliable and safe multimodal voice interaction.

[0097] Example 2

[0098] Please refer to Figure 2. For details not described in this embodiment, please refer to the description in Embodiment 1. A multimodal voice interaction method for an intelligent cockpit is provided, including:

[0099] S1. Collect ambient audio and video information in the cabin, combine it with the environmental perception parameters inside the vehicle, and determine whether to enter the voice interaction mode. When the voice interaction trigger condition is met, generate a voice interaction input signal that is adapted to the current environment.

[0100] S2. Extract acoustic features from the voice interaction input signal, simultaneously analyze the driver's lip movement trajectory in the video information, establish the correspondence between speech phonemes and mouth movements, and form joint analytical features;

[0101] S3. Segment and align the joint parsed features to construct a continuous multimodal segment sequence; synchronize the multimodal segment sequence in time and project it onto a predefined intent space to obtain a set of candidate intents containing different candidate intents;

[0102] S4. Based on the candidate intent set, execute the predefined dynamic rule arbitration logic to filter candidate intents layer by layer and determine the final execution instruction;

[0103] S5. Based on the final execution command, control the internal devices of the intelligent cockpit to perform corresponding operations, update the cockpit dynamic context information, generate multimodal collaborative response signals and output them.

[0104] Since the electronic device described in this embodiment is the electronic device used to implement the intelligent cockpit multimodal voice interaction system and method described in this application embodiment, those skilled in the art can understand the specific implementation method and various variations of the electronic device in this embodiment based on the intelligent cockpit multimodal voice interaction system and method described in this application embodiment. Therefore, how the electronic device implements the method in this application embodiment will not be described in detail here. As long as those skilled in the art implement the electronic device used in the intelligent cockpit multimodal voice interaction system and method described in this application embodiment, it falls within the protection scope of this application.

[0105] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0106] The above description is merely a preferred embodiment of the present invention, and the scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for users of ordinary technical skills, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal voice interaction system for an intelligent cockpit, characterized in that, include: The voice triggering unit collects ambient audio and video information in the cabin and combines it with environmental perception parameters inside the vehicle to determine whether to enter the voice interaction mode. When the voice interaction triggering condition is met, it generates a voice interaction input signal that is adapted to the current environment. The lip movement analysis unit extracts acoustic features from the voice interaction input signal and simultaneously analyzes the driver's lip movement trajectory in the video information to establish the correspondence between speech phonemes and lip movements, forming a joint analysis feature. A candidate generation unit segments and aligns the joint parsed features to construct a continuous multimodal segment sequence. It then synchronizes the multimodal segment sequence in time and projects it onto a predefined intent space to obtain a candidate intent set containing different candidate intents. The method for obtaining the candidate intent set includes: aligning the multimodal segment sequence in time on a global time scale; defining the intent space as a set of semantic categories, with each intent corresponding to a multimodal feature distribution region; projecting the time-aligned multimodal segment sequence into the predefined intent space to obtain the candidate intent and initial confidence score for each multimodal segment; collecting all projection results to generate a candidate intent set containing different candidate intents. An intent arbitration unit, based on the candidate intent set, executes predefined dynamic rule arbitration logic, filtering candidate intents layer by layer and determining the final execution instruction. The method for determining the final execution instruction includes: calculating the comprehensive confidence score for each candidate intent in the candidate intent set; the comprehensive confidence score is: ;in, Indicates candidate intent The overall confidence level; Indicates candidate intent The corresponding speech modality confidence; Indicates candidate intent The corresponding video modal confidence; The dynamic weights representing the confidence level of the speech modality; The dynamic weights representing the video modality confidence; This represents the index of the candidate intent; after the candidate intents are generated, a conflict resolution function is introduced to adjust the overall confidence of candidate intents that may have conflicts, so that the overall confidence of candidate intents with high conflict intensity decreases; the conflict resolution function is: ;in, Indicates the adjusted candidate intent The overall confidence level; Representing candidate intent in speech modality The intensity of conflict with other candidate intentions; Indicates candidate intent in lip shape modality The conflict strength with other candidate intents; determining whether the candidate intent is allowed to be executed under the current environment perception parameters through security constraint rules; filtering the allowed candidate intents layer by layer, and selecting the candidate intent with the highest comprehensive confidence and satisfying the security constraint rules as the final execution instruction; the security constraint rules are: ;in, Indicates the final operation to be performed; This represents the candidate intent with the highest overall confidence level; This indicates that execution is prohibited and will not trigger any actions within the cockpit. Represents the safety constraint function (Boolean function); Indicates candidate intent Allowed to execute; Indicates candidate intent The execution is prohibited; the collaborative response unit, according to the final execution command, controls the internal equipment of the intelligent cockpit to perform corresponding operations, updates the cockpit dynamic context information, generates multimodal collaborative response signals and outputs them.

2. The intelligent cockpit multimodal voice interaction system according to claim 1, characterized in that, The environmental perception parameters include: background noise decibel value, microphone array status, real-time vehicle speed, vehicle operation status, number and distribution of people inside the cabin, operating status of cabin equipment, light intensity inside the cabin, and camera field of view.

3. The intelligent cockpit multimodal voice interaction system according to claim 2, characterized in that, The method for acquiring the voice interaction input signal includes: collecting ambient audio and video information of the driver's face, comprehensively evaluating them in conjunction with environmental perception parameters inside the vehicle, and determining whether to enter the voice interaction mode; in the ambient audio, detecting whether there are audio segments whose matching degree with the preset voice interaction wake-up words exceeds a preset matching degree threshold, extracting the audio segments exceeding the preset matching degree threshold, and further generating an audio trigger candidate set; extracting the start and end times of all audio segments in the audio trigger candidate set as the audio trigger candidate time window; in the video information, selecting a time period with the same start and end time range as the audio trigger candidate time window as the video trigger candidate time window, and extracting video frames within the video trigger candidate time window in which the driver faces the camera and has a lip opening and closing trajectory as the video trigger candidate set; using environmental perception parameters as constraints, performing synchronous time period verification on the audio trigger candidate set and the video trigger candidate set to determine whether the audio trigger and video trigger are valid within the same time period; when the audio trigger candidate set and the video trigger candidate set are simultaneously valid within the same time period, it is determined that the voice interaction mode has been entered, and a voice interaction input signal adapted to the current environment is generated.

4. The intelligent cockpit multimodal voice interaction system according to claim 3, characterized in that, The method for extracting acoustic features from voice interaction input signals includes: dividing the voice interaction input signal into frames according to a preset frame length and frame shift, and applying a window function to each frame; performing a fast Fourier transform on the windowed voice interaction input signal to obtain a spectral representation of the voice interaction input signal, and calculating acoustic features based on the spectral representation.

5. The intelligent cockpit multimodal voice interaction system according to claim 4, characterized in that, The method for obtaining the joint analytical features includes: performing face detection on the collected video information to obtain the driver's face bounding box; locating the key points of the lip region based on facial key points within the driver's face bounding box; constructing the lip region boundary based on the two-dimensional coordinates of the key points of the lip region; tracking the lip region boundary position of consecutive video frames in the video information; performing geometric normalization on the tracked lip region boundary to convert it into a continuous lip image time series; extracting lip movement trajectory features from the lip image time series to form a lip movement trajectory sequence; aligning the acquired acoustic features with the lip movement trajectory sequence in time to establish the correspondence between speech phonemes and lip movements, thus forming joint analytical features.

6. The intelligent cockpit multimodal voice interaction system according to claim 5, characterized in that, The method for constructing the multimodal fragment sequence includes: segmenting the joint analytical features, dividing the joint analytical features into different multimodal fragments according to a continuous time window; aggregating the joint analytical features within each multimodal fragment to obtain the multimodal fragment sequence.

7. The intelligent cockpit multimodal voice interaction system according to claim 6, characterized in that, The method for generating and outputting multimodal cooperative response signals includes: sending control signals to devices inside the smart cockpit according to the final execution instruction, causing the corresponding devices to perform corresponding operations, recording the operation results of the corresponding devices in real time, generating corresponding multimodal cooperative response signals, and outputting them through the interactive interface of the smart cockpit.

8. A multimodal voice interaction method for an intelligent cockpit, implemented using a multimodal voice interaction system for an intelligent cockpit as described in any one of claims 1 to 7, characterized in that, include: S1. Collect ambient audio and video information within the cabin, and combine it with environmental perception parameters inside the vehicle to determine whether to enter voice interaction mode. When the voice interaction trigger condition is met, generate a voice interaction input signal adapted to the current environment. S2. Extract acoustic features from the voice interaction input signal, simultaneously analyze the driver's lip movement trajectory in the video information, establish the correspondence between speech phonemes and lip movements, and form joint analytical features. S3. Segment and align the joint analytical features to construct a continuous multimodal segment sequence. By synchronizing the multimodal segment sequence in time and projecting it onto a predefined intent space, a candidate intent set containing different candidate intents is obtained. S4. Based on the candidate intent set, execute predefined dynamic rule arbitration logic, filter candidate intents layer by layer, and determine the final execution instruction. S5. Based on the final execution command, control the internal devices of the intelligent cockpit to perform corresponding operations, update the cockpit dynamic context information, generate multimodal collaborative response signals and output them.

Citation Information

Patent Citations

  • Method and equipment for triggering voice interaction response

    CN111028842A

  • Pedestrian intention prediction method based on Transform multi-modal fusion strategy

    CN118781575A

  • Vehicle control method, device and equipment based on voice recognition and vehicle

    CN119541480A

  • Vehicle control method, server and computer readable storage medium

    CN120071927A