Vehicle-mounted voice noise reduction multi-mode interaction method and system
By integrating a MEMS microphone array and a thermal imaging camera, combined with V2X data and a 3D context-aware model, the problem of decreased recognition rate and unstable noise reduction effect of in-vehicle voice interaction system in complex environments has been solved, achieving efficient voice recognition and dynamic noise reduction.
Patent Information
- Application Number
- CN202511733218.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-13
AI Technical Summary
Existing in-vehicle voice interaction systems suffer from reduced recognition rates when faced with sudden noise, lack contextual understanding and proactive service capabilities, and the noise reduction mode is not linked to the user's driving habits, making it difficult to dynamically balance the noise reduction effect and voice wake-up sensitivity.
It integrates a MEMS microphone array and a thermal imaging camera to track noise sources in real time, predicts noise types through V2X data, dynamically adjusts noise reduction algorithm parameters, builds a three-dimensional context-aware model, and integrates lip movement recognition, eye tracking, and speech recognition to learn user driving habits and optimize noise reduction modes.
It significantly improves the voice recognition rate, enhances the efficiency and user experience of in-vehicle voice interaction, effectively reduces noise in complex environments, avoids interference in multi-speaker scenarios, and enables dynamic interactive networks.
Smart Images

Figure CN121528231A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of automotive intelligent cockpit technology, in particular to a vehicle-mounted voice noise reduction multi-modal interaction method and system. BACKGROUND
[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Vehicle-mounted voice interaction is a technology that allows drivers or passengers to control various functions of the vehicle through voice commands, without the need for manual operation. It combines speech recognition, natural language processing (NLP) and artificial intelligence (AI) technologies to improve driving safety, convenience and user experience.
[0004] The existing vehicle-mounted voice interaction system has technical bottlenecks such as static noise reduction technology, surface multi-modal fusion, and lack of personalized configuration. On the one hand, it relies too much on pre-set noise models, and has insufficient adaptability to sudden noises (such as sudden braking and horn sounds), resulting in a sharp drop in instantaneous speech recognition rate. On the other hand, multi-modal interaction is mainly based on simple command superposition, lacking context understanding and active service capabilities, such as the inability to combine lip movements and voice for complex command analysis, and the lack of linkage between noise reduction mode and user driving habits, resulting in difficulty in dynamically balancing noise reduction effect and voice wake-up sensitivity. SUMMARY
[0005] To solve the above problems, the present disclosure proposes a vehicle-mounted voice noise reduction multi-modal interaction method and system, which integrates MEMS microphone arrays and thermal imaging cameras to track noise sources in real time, dynamically adjust noise reduction algorithm parameters, predict noise types through V2X data, and switch noise reduction modes in advance; a three-dimensional context perception model is constructed, combining lip movement recognition, gaze tracking and speech recognition, learning user driving habits, and automatically optimizing noise reduction mode.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions: A vehicle-mounted voice noise reduction multi-modal interaction method, comprising: real-time acquisition of external environmental noise data, prediction of noise type through V2X data, and switching of noise reduction mode; constructing a multi-modal input model, inputting external environmental noise data and V2X data into the multi-modal input model, analyzing the spectral characteristics of external environmental noise data, outputting a frequency domain mask matrix to dynamically suppress the energy of a specific frequency band, and performing automatic noise reduction of the external environment; The automatic noise reduction is performed while real-time acquisition of user lip action, line of sight coordinate and voice data, input of the lip action, line of sight coordinate and voice data into a three-dimensional context perception model, formation of a dynamic interaction network through a multi-modal complementary and collaborative mechanism, output of an analysis result of a composite instruction, and realization of automatic noise reduction and instruction analysis of the vehicle-mounted voice interaction.
[0007] According to some embodiments, the present disclosure adopts the technical scheme as follows: A vehicle-mounted voice noise reduction multi-modal interaction system comprises: A data acquisition module is configured to acquire external environment noise data in real time, to pre-judge a noise type through V2X data, and to switch a noise reduction mode. An automatic noise reduction module is configured to construct a multi-modal input model, to input the external environment noise data and the V2X data into the multi-modal input model, to analyze spectral characteristics of the external environment noise data, to output a frequency domain mask matrix to dynamically suppress energy of a specific frequency band, and to perform automatic noise reduction of the external environment. A noise reduction interaction module is configured to perform the automatic noise reduction while real-time acquisition of user lip action, line of sight coordinate and voice data, input of the lip action, line of sight coordinate and voice data into a three-dimensional context perception model, formation of a dynamic interaction network through a multi-modal complementary and collaborative mechanism, output of an analysis result of a composite instruction, and realization of automatic noise reduction and instruction analysis of the vehicle-mounted voice interaction.
[0008] According to some embodiments, the present disclosure adopts the technical scheme as follows: A computer program product comprises a computer program, which, when executed by a processor, implements the vehicle-mounted voice noise reduction multi-modal interaction method.
[0009] According to some embodiments, the present disclosure adopts the technical scheme as follows: A non-transitory computer readable storage medium is configured to store computer instructions, which, when executed by a processor, implement the vehicle-mounted voice noise reduction multi-modal interaction method.
[0010] According to some embodiments, the present disclosure adopts the technical scheme as follows: An electronic device comprises a processor, a memory and a computer program, wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the vehicle-mounted voice noise reduction multi-modal interaction method.
[0011] Compared with the prior art, the present disclosure has the beneficial effects as follows: The vehicle-mounted voice noise reduction multi-modal interaction method of the present disclosure integrates a MEMS microphone array and a thermal imaging camera to track noise sources in real time, dynamically adjust noise reduction algorithm parameters, predict noise types through V2X data, and switch noise reduction modes in advance. Lip movement recognition, gaze tracking, and voice recognition are fused to build a three-dimensional context perception model. The driving habits of users (such as common routes and music preferences) are learned to automatically optimize the noise reduction mode.
[0012] The vehicle-mounted voice noise reduction multi-modal interaction method of the present disclosure integrates lip movement recognition, gaze tracking, and voice recognition to build a three-dimensional context perception model. The three form a dynamic interaction network through multi-modal complementation and synergy mechanism. Lip movement information makes up for the lack of voice signals in a noisy environment, and voice signals provide acoustic features to verify the accuracy of lip movement recognition. Gaze tracking determines the speaker or interactive object that the user is focusing on through eye movement, and lip movement recognition extracts lip movement features for the target area to avoid interference in a multi-speaker scenario. Through innovative dynamic environment modeling and multi-modal context perception technology, the failure problem of traditional vehicle-mounted voice interaction systems in complex environments is solved, and the interaction efficiency and user experience are significantly improved. The voice recognition rate is higher than that of traditional systems. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which form a part of the present disclosure, are intended to provide further understanding of the present disclosure, and the illustrative embodiments of the present disclosure and their descriptions serve to explain the present disclosure, and do not constitute improper limitations on the present disclosure.
[0014] Figure 1 A vehicle-mounted voice noise reduction multi-modal interaction method architecture diagram of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0015] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0016] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.
[0017] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present disclosure. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and furthermore, it should be understood that when the terms "comprise" and / or "include" are used in the specification, they refer to the presence of a feature, step, operation, device, component, and / or combination thereof.
[0018] Embodiment 1 In an embodiment of the present disclosure, a vehicle-mounted voice noise reduction multi-modal interaction method is provided, and the method steps include: Step 1: Real-time acquisition of external environmental noise data, prediction of noise type through V2X data, and switching of noise reduction mode; Step 2: Construct a multi-modal input model, input external environmental noise data and V2X data into the multi-modal input model, analyze the spectral characteristics of external environmental noise data, output a frequency domain mask matrix to dynamically suppress the energy of a specific frequency band, and automatically reduce the noise of the external environment; Step 3: While automatically reducing noise, real-time acquisition of user lip movement, gaze coordinates, and voice data, input of the lip movement, gaze coordinates, and voice data into a three-dimensional context perception model, formation of a dynamic interaction network through multi-modal complementation and cooperation mechanism, output of analysis results of composite instructions, and realization of automatic noise reduction and instruction analysis of vehicle-mounted voice interaction.
[0019] As an embodiment, the vehicle-mounted voice noise reduction multi-modal interaction method of the present disclosure integrates a MEMS microphone array and a thermal imaging camera to real-time track noise sources, dynamically adjust noise reduction algorithm parameters, predict noise type through V2X data, and switch noise reduction mode in advance. Then, lip movement recognition, gaze tracking, and voice recognition are fused to construct a three-dimensional context perception model. And the user's driving habits (such as common routes and music preferences) are learned to automatically optimize the noise reduction mode. The specific implementation process is as follows: Step 1: Real-time acquisition of external environmental noise data, prediction of noise type through V2X data, and switching of noise reduction mode; construct a multi-modal input model, input external environmental noise data and V2X data into the multi-modal input model, analyze the spectral characteristics of external environmental noise data, output a frequency domain mask matrix to dynamically suppress the energy of a specific frequency band, and automatically reduce the noise of the external environment; Specifically, by integrating a MEMS microphone array (supporting four sound zone positioning) and a thermal imaging camera (integrated in a roof control module, achieving low power consumption (<3W) synchronous work through a gallium nitride chip), real-time tracking of noise sources (such as road construction ahead), prediction of noise type (such as traffic warning sound) through V2X data, and switching of noise reduction mode in advance.
[0020] Dynamically adjust noise reduction algorithm parameters by real-time analysis of environmental noise spectral characteristics (such as frequency band distribution and energy intensity), adjust filter order and algorithm complexity, and classify noise spectral characteristics using an AI model (such as CNN). In a dynamic environment (such as noise fluctuation caused by changes in vehicle speed), update the filter weights according to the error signal through the LMS algorithm: Noise reduction intensity = α · Noise source distance + β · Noise type weight + Gamma · User preference coefficient Wherein, a, β, γ are dynamic weight coefficients (updated in real time through reinforcement learning).
[0021] A multi-modal input model (acoustic + thermal imaging + V2X) is constructed to output a frequency domain mask matrix to dynamically suppress the energy of a specific frequency band.
[0022] Step 2: While automatically reducing noise, real-time acquisition of user lip movement, gaze coordinates, and speech data is performed. The lip movement, gaze coordinates, and speech data are input into a three-dimensional context perception model. Through a multi-modal complementary and collaborative mechanism, a dynamic interaction network is formed to output the analysis result of the composite instruction, thereby achieving automatic noise reduction and instruction analysis for vehicle-mounted voice interaction.
[0023] Specifically, a three-dimensional context perception model is constructed by fusing lip-reading, gaze tracking, and speech recognition. Through a multi-modal complementary and collaborative mechanism, a dynamic interaction network is formed. Lip movement information compensates for the lack of speech signals in a noisy environment, and speech signals provide acoustic features to verify the accuracy of lip-reading. Gaze tracking determines the speaker or interactive object that the user is focusing on through eye movement, and lip-reading extracts lip movement features for the target area to avoid interference in a multi-speaker scenario. Speech recognition analyzes semantic content, and gaze tracking reflects user attention distribution through gaze duration and trajectory to jointly infer user intent.
[0024] A cross-modal collaborative framework is constructed through multi-modal data fusion, temporal and spatial alignment, and dynamic reasoning. A multi-camera array captures lip movement, an eye tracking camera (such as an RGB-D sensor) acquires gaze direction, and a microphone array collects speech signals. A hardware synchronization module (such as an FPGA clock) ensures that the timestamps of multi-modal data are aligned, with an error of less than 10ms.
[0025] The present disclosure uses a 3D convolutional neural network (such as a LipNet variant) to extract spatiotemporal features of lip movement, outputting parameters such as lip opening degree and movement frequency. Based on an improved VGG16 network, the pupil position and iris texture in the eye image are extracted, and the LSTM is used to capture the temporal dependence of the gaze trajectory. A dual-channel input of Mel spectrogram and MFCC is used, and a Conformer model is used to extract noise-robust acoustic features. Through a cross-modal attention mechanism (such as the cross-attention layer in the Transformer), weights are dynamically allocated to integrate the recognition results of each modality. The gaze tracking locks the speaker region, reducing the matching range of lip movement and speech. A three-dimensional latent space is constructed to project multi-modal data into the same space, and a graph neural network (GNN) is used to model the spatiotemporal relationship between lip movement, speech, and gaze.
[0026] The lip movement sequence, gaze coordinates, and speech waveform are input into the multi-modal instruction analysis model, and the composite instruction analysis result (such as "play the recommended song list in the gaze positioning area") is output. Further, an AR-HUD noise reduction mode visualization adjustment interface is developed, and a user can adjust noise reduction parameters through gesture dragging. Based on camera and three-dimensional modeling to identify sound source direction, directional noise reduction is achieved by adjusting the pickup coverage angle of the microphone array. The user can adjust the target noise reduction area through gesture dragging the virtual sound source direction arrow (such as 0°-180° horizontal range), for example, expanding the noise reduction focus from 90° straight ahead to 30° on both sides to suppress lateral wind noise when changing lanes. The mixing ratio of environmental sound and noise reduction (such as 0-100% slide bar) can also be adjusted through gestures to balance the noise reduction effect and safety warning needs. Learn the driving habits of users (such as common routes, music preferences), automatically optimize the noise reduction mode, and iteratively optimize the noise reduction strategy through the historical noise characteristic library. Gesture selection of noise mode learning intensity (such as weak / medium / strong) controls the adaptive suppression ability of the system to repeated noise (such as air conditioner noise).
[0027] Embodiment 2 In an embodiment of the present disclosure, a vehicle-mounted voice noise reduction multi-modal interaction system is provided, comprising: A data acquisition module is configured to acquire external environmental noise data in real time, predict the type of noise through V2X data, and switch the noise reduction mode. An automatic noise reduction module is configured to construct a multi-modal input model, input the external environmental noise data and V2X data into the multi-modal input model, analyze the spectral characteristics of the external environmental noise data, output a frequency domain mask matrix to dynamically suppress the energy of a specific frequency band, and perform automatic noise reduction of the external environment. A noise reduction interaction module is configured to acquire user lip movements, gaze coordinates, and voice data in real time while automatically reducing noise, input the lip movements, gaze coordinates, and voice data into a three-dimensional context perception model, form a dynamic interaction network through a multi-modal complementary and collaborative mechanism, output the analysis result of the composite instruction, and realize automatic noise reduction and instruction analysis of vehicle-mounted voice interaction.
[0028] Embodiment 3 In an embodiment of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the vehicle-mounted voice noise reduction multi-modal interaction method.
[0029] Embodiment 4 In an embodiment of the present disclosure, a non-transitory computer readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the vehicle-mounted voice noise reduction multi-modal interaction method.
[0030] Embodiment 5 An embodiment of the present disclosure provides an electronic device, comprising: a processor, a memory and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the vehicle-mounted voice noise reduction multi-modal interaction method. The present disclosure is described with reference to the flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks. Figure 1 The functions specified in one flow or multiple flows and / or blocks.
[0031] These computer program instructions can also be loaded into a computer or other programmable data processing device to cause a series of operation steps to be executed on the computer or other programmable data processing device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing device provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks. Figure 1 The functions specified in one flow or multiple flows and / or blocks.
[0032] Although the specific embodiments of the present disclosure are described above with reference to the accompanying drawings, the present disclosure is not limited to the above embodiments, and various modifications or changes can be made to the embodiments without departing from the scope of the present disclosure.
Claims
1. A method for in-vehicle voice noise reduction and multimodal interaction, characterized in that, include: Real-time acquisition of external environmental noise data; prediction of noise type based on V2X data; and switching of noise reduction mode accordingly. A multimodal input model is constructed, and external environmental noise data and V2X data are input into the multimodal input model. The spectral characteristics of the external environmental noise data are analyzed, and a frequency domain mask matrix is output to dynamically suppress energy in specific frequency bands, thereby performing automatic noise reduction of the external environment. While automatically reducing noise, the system acquires the user's lip movements, gaze coordinates, and voice data in real time. The lip movements, gaze coordinates, and voice data are then input into a three-dimensional context-aware model. Through multimodal complementarity and collaboration mechanisms, a dynamic interactive network is formed, and the parsing results of composite commands are output, thus achieving automatic noise reduction and command parsing for in-vehicle voice interaction.
2. The in-vehicle voice noise reduction multimodal interaction method as described in claim 1, characterized in that, The real-time acquisition of external environmental noise data, the prediction of noise type based on V2X data, and the switching of noise reduction modes include: By integrating a MEMS microphone array and a thermal imaging camera, noise sources can be tracked and acquired in real time. The noise type can be predicted using V2X data, and the noise reduction mode can be switched in advance.
3. The in-vehicle voice noise reduction multimodal interaction method as described in claim 1, characterized in that, A multimodal input model is constructed, and external environmental noise data and V2X data are input into the multimodal input model. The spectral characteristics of the external environmental noise data are analyzed, and a frequency domain mask matrix is output to dynamically suppress energy in specific frequency bands, thereby performing automatic noise reduction of the external environment, including: The noise reduction algorithm parameters are dynamically adjusted. By analyzing the environmental noise spectrum characteristics in real time, the filter order and algorithm complexity are adjusted. The noise spectrum characteristics are classified using an AI model. In a dynamic environment, the LMS algorithm is used to dynamically update the filter weights based on the error signal. A multimodal input model is constructed. External environmental noise data and V2X data are input into the multimodal input model, and the frequency domain mask matrix is output to dynamically suppress specific frequency band energy. The noise spectrum characteristics include frequency band distribution and energy intensity.
4. The in-vehicle voice noise reduction multimodal interaction method as described in claim 1, characterized in that, Real-time acquisition of user lip movements, gaze coordinates, and speech data; input of lip movements, gaze coordinates, and speech data into a 3D context-aware model; formation of a dynamic interaction network through multimodal complementarity and collaboration mechanisms, including: User lip movements compensate for the lack of speech signals in noisy environments, eye tracking determines the speaker or interactive object that the user is paying attention to through eye movements, extracts lip movement features for the target area, speech recognition parses semantic content, and eye tracking reflects the distribution of user attention through gaze duration and trajectory, jointly inferring user intent. A cross-modal dynamic interaction network is constructed through multimodal data fusion, spatiotemporal alignment, and dynamic reasoning. Among them, a multi-camera array captures lip movements, an eye-tracking camera obtains the coordinates of the gaze, a microphone array collects voice signals, and a hardware synchronization module ensures that the timestamps of the multimodal data are aligned.
5. The in-vehicle voice noise reduction multimodal interaction method as described in claim 4, characterized in that, A 3D convolutional neural network is used to extract spatiotemporal features of lip movements, outputting lip opening and closing degree and movement frequency parameters. Based on an improved VGG16 network, pupil position and iris texture in eye images are extracted. LSTM is combined to capture the temporal dependence of gaze trajectory. Mel spectrogram and MFCC dual-channel input are used. Conformer model is combined to extract noise robustness acoustic features. Weights are dynamically allocated through cross-modal attention mechanism, and recognition results of each modality are integrated to construct a three-dimensional latent space. Multimodal data are projected into the same space, and the spatiotemporal relationship between lip movement, speech and gaze is modeled through graph neural network.
6. The in-vehicle voice noise reduction multimodal interaction method as described in claim 1, characterized in that, Based on the parsing results of the compound commands, automatic noise reduction and command parsing for in-vehicle voice interaction are implemented, including: The AR-HUD noise reduction mode features a visual adjustment interface. Users can adjust noise reduction parameters by dragging and dropping gestures. The system identifies the direction of the sound source based on the camera and 3D modeling, and achieves directional noise reduction by adjusting the microphone array's pickup coverage angle. Users can also drag virtual sound source direction arrows to adjust the target noise reduction area and execute voice-defined interactive intentions.
7. A vehicle-mounted voice noise reduction multimodal interaction system, characterized in that, include: The data acquisition module is used to acquire external environmental noise data in real time, predict the noise type through V2X data, and switch the noise reduction mode accordingly. The automatic noise reduction module is used to construct a multimodal input model. It inputs external environmental noise data and V2X data into the multimodal input model, analyzes the spectral characteristics of the external environmental noise data, and outputs a frequency domain mask matrix to dynamically suppress energy in specific frequency bands, thereby performing automatic noise reduction of the external environment. The noise reduction and interaction module is used to automatically reduce noise while acquiring the user's lip movements, gaze coordinates, and voice data in real time. The lip movements, gaze coordinates, and voice data are input into a three-dimensional context-aware model. Through multimodal complementarity and collaboration mechanisms, a dynamic interaction network is formed, and the parsing results of composite commands are output, realizing automatic noise reduction and command parsing for in-vehicle voice interaction.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the in-vehicle voice noise reduction multimodal interaction method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a vehicle-mounted voice noise reduction multimodal interaction method as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform a vehicle-mounted voice noise reduction multimodal interaction method as described in any one of claims 1-6.