Multi-modal in-vehicle noise reduction method, controller and computer readable storage medium
By using a multimodal fusion microphone array and camera device, combining visual and audio data, and employing a Kalman filter for sound source location fusion, the noise reduction problem under multiple interfering sound sources inside the vehicle was solved, achieving accurate noise reduction effect and a stable user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 贵州华鑫信息技术有限公司
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to effectively handle multiple interfering sound sources in complex in-vehicle acoustic environments, resulting in poor noise suppression and system vulnerability to information uncertainty and dynamic changes.
A multimodal fusion method is adopted, which combines a microphone array and a camera device. Visual data and audio data are processed collaboratively, and a Kalman filter is used to fuse the sound source locations, construct a list of interfering sound sources, and perform precise noise reduction based on a preset noise reduction algorithm.
It achieves precise localization and noise reduction of target sound sources in complex in-vehicle environments, improving the accuracy and targeting of noise reduction and ensuring the stability and continuity of user experience.
Smart Images

Figure CN122493873A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of noise reduction technology, and in particular to a multimodal in-vehicle noise reduction method, controller, and computer-readable storage medium. Background Technology
[0002] With the increasing popularity of smart cockpits, in-vehicle voice calls, voice assistants, and in-car conferencing systems have become core human-machine interaction methods. Users expect a clear voice experience similar to "face-to-face conversation" in these scenarios. However, the in-vehicle environment is complex: there is continuous road noise, wind noise, air conditioning noise, and transient engine / motor noise, along with the "cocktail party effect" of multiple conversations and severe acoustic reverberation. These factors pose significant challenges to voice communication and quality.
[0003] The relevant technologies use active noise cancellation to process interfering sound sources. However, the active noise cancellation technology provided by the relevant technologies mainly targets a single noise source and is difficult to deal with the complex scenario where multiple interfering sound sources coexist in the vehicle. It is easy to have problems such as the target sound source being falsely suppressed and the interfering sound source being not completely suppressed. Therefore, the noise suppression effect of the relevant technologies is not good. Summary of the Invention
[0004] One objective of this application is to provide a multimodal in-vehicle noise reduction method, controller, and computer-readable storage medium to improve the poor noise reduction effect of related technologies.
[0005] In a first aspect, embodiments of this application provide a multimodal in-vehicle noise reduction method. The vehicle includes a microphone array and a camera device. The method includes: acquiring audio data collected by the microphone array in a target space and visual data captured by the camera device in the target space, wherein the target space includes a target sound source (a target person) and interfering sound sources; determining the visual position and visual covariance matrix of the target sound source based on the visual data; determining the audio position and audio covariance matrix of the target sound source based on the audio data; and applying a preset Kalman filter to the visual position and audio covariance matrix. The variance matrix, the audio location, and the audio covariance matrix are fused to obtain the sound source location of the target sound source; an interference source list is obtained, which includes the sound source locations of one or more interference sources, and a confidence level is matched for each interference source; based on the interference source list, interference sources that meet the preset confidence level conditions are identified as sound sources to be suppressed; based on a preset noise reduction algorithm, the audio data is denoised using the sound source locations of the target sound source and the sound source locations of the sound sources to be suppressed, so that the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced.
[0006] Optionally, the step of using a preset noise reduction algorithm to perform noise reduction processing on the audio data using the sound source location of the target sound source and the sound source location of the sound source to be suppressed, so as to suppress the audio of the sound source to be suppressed and enhance the audio of the target sound source, includes: filtering out multiple invalid audio segments in the audio data based on the visual data, the audio data including multiple time-domain audio segments, the invalid audio segments being time-domain audio segments that do not belong to the time when the target person is speaking; generating a data covariance matrix of noise space based on the multiple invalid audio segments; performing array manifold processing on the sound source location of the target sound source and the sound source location of the sound source to be suppressed based on the noise reduction algorithm to obtain a dynamic constraint matrix; performing weight calculation processing on the data covariance matrix and a preset target response vector based on the noise reduction algorithm to obtain a target weight matrix; and performing suppression processing on the audio of the sound source to be suppressed and enhancement processing on the audio of the target sound source based on the target weight matrix and the audio data.
[0007] Optionally, the step of filtering out multiple invalid audio segments from the audio data based on the visual data includes: performing frame segmentation processing on the visual data and the audio data respectively to obtain multiple visual segments and multiple temporal audio segments, aligning one visual segment with one temporal audio segment; and filtering out multiple invalid audio segments among the multiple temporal audio segments based on the alignment relationship between the visual segments and the temporal audio segments.
[0008] Optionally, the step of filtering out multiple invalid audio segments among multiple time-domain audio segments based on the alignment relationship between the visual segments and the time-domain audio segments includes: performing a person recognition operation on each visual segment to obtain the identity information of the visual segment; in response to the identity information of the visual segment being the identity information of the target person, performing a lip movement detection operation on the visual segment to obtain the lip movement detection information of the visual segment; in response to the lip movement detection information of the visual segment being speech activity information, setting the audio segment aligned with the visual segment as a valid audio segment; in response to all visual segments completing the person recognition operation and lip movement detection operation, removing all valid audio segments from all audio segments to obtain a set of remaining segments; and setting each audio segment in the set of remaining segments as an invalid audio segment.
[0009] Optionally, generating the data covariance matrix of the noise space based on multiple invalid audio segments includes: performing short-time Fourier transform processing on the invalid audio segments to obtain invalid frequency domain segments; calculating a periodogram based on the invalid frequency domain segments; and performing arithmetic mean processing on the periodograms of all invalid frequency domain segments to obtain the data covariance matrix of the noise space where the microphone array is located.
[0010] Optionally, the microphone array includes multiple array elements, and each interfering sound source in the list of interfering sound sources is configured with a confidence level. The step of performing array manifold processing on the sound source positions of the target sound source and the sound source positions of the sound source to be suppressed based on the noise reduction algorithm to obtain a dynamic constraint matrix includes: determining the target steering vector from the target sound source to all array elements based on the positions of the array elements and the sound source positions of the target sound source; determining the interference steering vector from the sound source to be suppressed to all array elements based on the positions of the array elements and the sound source positions of the sound source to be suppressed; and determining the dynamic constraint matrix based on the target steering vector and all interference steering vectors.
[0011] Optionally, the step of performing weight calculation processing on the data covariance matrix and the preset target response vector based on the noise reduction algorithm to obtain the target weight matrix includes: constructing an optimization problem based on the weight matrix to be solved and the data covariance matrix; constructing constraints based on the weight matrix to be solved and the preset target response vector; and solving the optimization problem based on the constraints to obtain the target weight matrix.
[0012] Optionally, the step of performing suppression processing on the audio of the sound source to be suppressed and enhancement processing on the audio of the target sound source based on the target weight matrix and the audio data includes: performing short-time Fourier transform processing on the time-domain audio segment to obtain a frequency-domain audio segment; determining a denoised frequency-domain audio segment based on the target weight matrix and the frequency-domain audio segment; performing inverse short-time Fourier transform on the denoised frequency-domain audio segment to obtain a denoised time-domain audio segment; and combining all denoised time-domain audio segments to obtain a denoised speech signal.
[0013] Optionally, obtaining the list of interfering sound sources includes: obtaining various types of sound source description data; determining the sound source location of the interfering sound sources based on the sound source description data; recording the sound source location of all interfering sound sources to obtain a list of interfering sound sources.
[0014] Optionally, the various sound source description data include the visual data, the audio data, vehicle operation data, and inherent knowledge source data. The interfering sound sources include actively predicted noise, passively perceived noise, vehicle noise, and fixed noise. Determining the sound source location of the interfering sound source based on the sound source description data includes: performing lip movement detection and position calculation operations based on the visual data to obtain the visual reference position of the non-target person, and setting the visual reference position of the non-target person as the sound source location of the actively predicted noise; determining the audio reference position of the non-target person based on the audio data, and setting the audio reference position of the non-target person as the sound source location of the passively perceived noise; determining the vehicle state based on the vehicle operation data, generating an acoustic context description vector based on the vehicle state, and determining the sound source location of the vehicle noise based on the acoustic context description vector; and determining the sound source location of the fixed noise based on the inherent knowledge source data.
[0015] Optionally, each of the interfering sound sources corresponds to an information source attribute, the information source attribute being used to indicate the determined source of the interfering sound source. The step of determining interfering sound sources satisfying preset confidence conditions as sources to be suppressed based on the list of interfering sound sources includes: acquiring the information source attribute of the interfering sound source; determining the basic weight of the interfering sound source based on the information source attribute; determining a weight correction coefficient based on the sound source description data of the interfering sound source; determining the confidence level of the interfering sound source based on the basic weight and the weight correction coefficient; and determining interfering sound sources satisfying preset confidence conditions as sources to be suppressed based on the confidence levels of each interfering sound source.
[0016] Optionally, determining the interference source that satisfies the preset confidence condition as the source to be suppressed based on the confidence of each interference source includes: determining that the interference source satisfies the preset confidence condition in response to the confidence of the interference source being greater than the preset confidence threshold; and setting the interference source as the source to be suppressed.
[0017] In a second aspect, embodiments of this application provide a controller, including a memory and a processor, wherein the memory is connected to the processor, and the processor is configured to execute one or more computer programs stored in the memory, wherein when the processor executes the one or more computer programs, the controller enables the controller to implement the above-described multimodal in-vehicle noise reduction method.
[0018] In a third aspect, embodiments of this application provide a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the aforementioned multimodal in-vehicle noise reduction method.
[0019] The embodiments of this application can achieve the following technical effects: Firstly, these embodiments introduce the collaborative work of visual and audio data, utilizing both to capture the location of the target sound source, achieving precise localization and providing effective data support for subsequent precise noise reduction. Secondly, by constructing an interference source list and recording the location and confidence level of all interfering sound sources, these embodiments provide accurate interference information support for subsequent targeted noise reduction. This makes noise reduction no longer blind but can specifically focus on interfering sound sources, achieving the core objective of "suppressing interference and protecting the target," significantly improving the accuracy, targeting, and effectiveness of noise reduction. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This application provides a schematic diagram of the system architecture of a multimodal noise reduction system. Figure 2 A flowchart illustrating a multimodal in-vehicle noise reduction method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a controller provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0023] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0024] The inventors also discovered two noise reduction schemes, the details of which are as follows: 1) Option 1: Beamforming technology based on pure audio. Scheme 1 widely employs Generalized Cross-Correlation-Phase Transform (GCC-PHAT) or Controlled Response Power-Phase Transform (SRP-PHAT) algorithms for sound source localization, thereby driving the microphone array to form a directional beam. Specifically, the technical principle of Scheme 1 is as follows: by calculating the time difference of sound arrival between microphone array elements, the direction of the sound source is estimated, and the sound in the target direction is enhanced through spatial filtering. However, Scheme 1 is highly sensitive to noise and reverberation when estimating the generalized cross-correlation function; in-vehicle reverberation can cause multiple peaks in the correlation function, resulting in localization ambiguity.
[0025] Solution 1 is essentially an "open-loop search" problem in a huge space. In environments with strong noise and reverberation, the time delay estimation will be severely distorted, leading to ambiguous positioning, incorrect beam pointing, and a sharp decline in performance. As a result, the noise reduction results are easily affected by noise and reverberation, resulting in inaccurate positioning and an inability to solve interference in the same direction (as the co-pilot speaks).
[0026] 2) Solution 2: Visually guided audio technology based on simple logic.
[0027] Option 2: Use the in-vehicle camera to identify the speaker's location (such as head or mouth coordinates), and then directly transmit the speaker's location to the beamforming algorithm to control the beam to point at that location.
[0028] This approach is an "if-then" type of static, open-loop switching method, which has the following bottlenecks; ① Visual single point of failure: Completely dependent on vision. If the camera is obstructed (passenger hands something to the driver, driver wears sunglasses, sudden change in lighting), the performance of the entire system drops drastically.
[0029] ② Lack of error correction capability: When the visual algorithm gives an incorrect visual position, the audio system will "blindly follow" that visual position without realizing it.
[0030] ③ Ignoring audio value: Even with accurate visual positioning, audio itself contains valuable spatial information, but this solution completely discards it.
[0031] Overall, both Scheme 1 and Scheme 2 are "functional superpositions" rather than "intelligent fusions". Their core flaw lies in the lack of an "intelligent brain" that can assess the credibility of multi-sensor information and make optimal decisions in uncertain environments.
[0032] Existing in-vehicle voice noise reduction and enhancement technologies have fundamental shortcomings in complex acoustic environments, which can be summarized in the following three aspects: 1) Vulnerability of system architecture: Open-loop operation, lack of feedback and fault tolerance.
[0033] Existing solutions, whether pure audio beamforming technology or simple visual guidance technology, are all unidirectional, open-loop systems. After the system executes instructions, it cannot verify the reliability and effectiveness of the output results. When strong noise or reverberation causes inaccurate sound source localization, or when occlusion or changes in lighting cause visual information to fail, the system lacks self-perception and error correction capabilities, thus triggering a chain reaction of performance failures.
[0034] 2) Rigidity of decision-making logic: binary switching, lack of adaptive fusion.
[0035] The existing technologies employ a static, either-or decision-making mechanism, relying entirely on audio or blindly following vision, failing to address the uncertain "grey areas" of sensor information in reality. The system lacks an intelligent core capable of dynamically assessing the credibility of both visual and audio information and performing smooth, weighted fusion. This results in poor performance in complex scenarios such as slight visual biases or audio frequency band interference.
[0036] 3) Limited state perception: lack of in-depth utilization of the dynamic characteristics of sound sources.
[0037] The relevant technologies have limitations at the perception level: they either view the sound source location statically (such as simple visual guidance schemes that treat the mouth coordinates as a fixed point) or only perform instantaneous, memoryless localization of the sound source (such as pure audio schemes that search independently in each frame). This approach ignores the valuable information inherent in the sound source's motion itself. Specifically, the relevant technologies fail to effectively model and utilize the sound source's motion state (such as the speed and trajectory of head / mouth movement), resulting in the following two technical shortcomings: 3.1) Lack of predictive ability: When the sensor signal is briefly interrupted (such as when the speaker turns their head quickly, causing a brief loss of vision, or when the speaker is briefly overwhelmed by noise), the system cannot predict the possible location of the current sound source based on the previous movement trend, which manifests as discontinuous tracking or sudden loss of tracking.
[0038] 3.2) Insufficient filtering capability: For instantaneous jumps or outliers in sensor data (such as sudden jitter of the visual detection box, or positioning jumps caused by noise interference in audio), the system lacks a kinematic model-based "smoothing" mechanism to identify and filter out these unreliable interferences.
[0039] In view of this, the embodiments of this application aim to fundamentally solve the problem of the sharp decline in performance and overall vulnerability of existing voice technologies in complex and dynamic vehicle acoustic environments due to open-loop system architecture, rigid decision-making mechanisms and isolated environmental perception.
[0040] This application's embodiments construct an "intelligent closed-loop system based on uncertainty self-perception," fundamentally solving the core pain point of how to continuously and stably lock onto and enhance target sound sources in a dynamic and noisy in-vehicle environment. Specifically, this is reflected in the following five value dimensions: 1) Vehicle noise sensor and multimodal data fusion: This innovative approach transforms the vehicle itself from a passive acoustic environment into an active noise prediction sensor. By parsing CAN bus data and incorporating a built-in "state-noise" mapping model, the system can proactively predict the emergence and changes in noise, achieving a leap from passive noise reduction to active acoustic environment management and providing the entire system with forward-looking context awareness capabilities.
[0041] 2) Uncertainty-driven dynamic fusion mechanism: The system can quantify the credibility of visual and acoustic signals in real time and intelligently determine when to "emphasize vision" or "rely on acoustics". This solves the dilemma of related technologies being at a loss when sensor information conflicts, and achieves a fundamental leap in robustness.
[0042] 3) Visual Anchoring Serial Processing Architecture: Through a unique process of "visual positioning first, followed by acoustic correction," a stable spatial anchor point is provided for the interference-prone acoustic system. Even in extreme cases of temporary acoustic signal failure, the system can still maintain reliable output, ensuring absolute consistency and stability of the user experience.
[0043] 4) Dynamic noise reduction driven by fusion results: The high-confidence fusion positioning results are converted into precise spatial filtering instructions in real time, thus upgrading beamforming from static "directional reception" to dynamic "active tracking". The noise reduction beam can closely follow the speaker's movement, achieving a qualitative leap in effect.
[0044] 5) Continuous tracking model with predictive capabilities: By introducing motion states (such as speed), the system evolves from "discrete search frame by frame" to "continuous trajectory tracking". This not only brings unprecedented smooth sound source trajectory and completely eliminates positioning jumps and jitters, but also gives the system short-term predictive capabilities, laying a solid foundation for high-end voice interaction.
[0045] The following embodiments of this application provide a multimodal noise reduction system. Please refer to... Figure 1 The multimodal noise reduction system 100 includes a microphone array 11, a camera device 12, and a controller 13.
[0046] Microphone array 11 is installed inside vehicle 14, for example, in the driver's cabin, for collecting audio data in a target space, where the target space is the vehicle interior. Microphone array 11 includes multiple microphones; in some embodiments, the microphones are spaced apart on the same circuit board, while in other embodiments, the microphones are positioned at different locations within the target space to collect audio data from all directions.
[0047] The camera device 12 is installed inside the vehicle 14, for example, on one side of the vehicle's windshield, for collecting visual data in the target space. The camera device 12 may consist of one or more camera units, which are respectively set at different locations in the target space to collect visual data from all directions.
[0048] The controller 13 is communicatively connected to the microphone array 11 and the camera device 12, respectively, and is used to process and analyze various business logics.
[0049] The following embodiments of this application provide a multimodal in-vehicle noise reduction method. Please refer to... Figure 2 In this embodiment of the application, in-vehicle noise reduction is performed through steps S21 to S27, as detailed below: Step S21: Acquire audio data collected by the microphone array in the target space and visual data captured by the camera device in the target space.
[0050] The target space includes the target sound source of the target person and interfering sound sources. The target person is either someone whose audio output needs to be enhanced or someone whose audio output does not need to be suppressed. For example, the target person can be a driver or a passenger. When the target person is a driver, this embodiment of the application enhances the driver's audio and suppresses the audio of surrounding interfering sound sources, thereby achieving noise reduction. Interfering sound sources are sound sources that do not belong to the target person. For example, interfering sound sources can be sound sources of non-target personnel, inherent noise sources of the vehicle, etc.
[0051] The audio data is data collected by a microphone array in the target space, and the audio data can be composed of multi-channel sub-audio data. Embodiments of this application can preprocess the audio data to obtain cleaner audio data. Preprocessing includes echo cancellation and gain control, etc.
[0052] Visual data refers to data collected by the camera device in the target space. Visual data includes multiple frames of captured images, which include images of local areas corresponding to the target person and images of local areas corresponding to non-target persons.
[0053] Step S22: Determine the visual location and visual covariance matrix of the target sound source based on visual data.
[0054] Visual position is the estimated location of a target person based on visual data. This application embodiment performs target detection processing, facial landmark processing, and position estimation processing on the visual data to obtain the visual position of the target person. In some embodiments, this application embodiment inputs the visual data into a pre-trained neural network model, enabling the neural network model to perform position estimation operations on the target person based on the visual data, thereby obtaining the visual position of the target person. Specifically, the neural network model includes an object detection network and a depth estimation network. This application embodiment processes the visual data based on the object detection network to obtain the 2D position of the target person, processes the visual data based on the depth estimation network to determine the depth information of the target person, and determines the visual position of the target person based on the 2D position and depth information. This visual position is the 3D position of the target person in a three-dimensional coordinate system. .in, Let x be the x-coordinate of the target person in a three-dimensional coordinate system. Let y be the coordinate of the target person on the 3D coordinate system. Let Z be the z-coordinate of the target person in the three-dimensional coordinate system. For visual location. The object detection network can be YOLOv8 or Faster R-CNN, and the depth estimation network can be Monodepth2 or DPT monocular depth estimation network, etc.
[0055] The visual covariance matrix is the observation noise covariance matrix of the Kalman filter. It integrates multiple visual noise penalty parameters, quantifying the uncertainty of visual position from multiple dimensions. The visual covariance matrix provides the basis for the observation weights in the Kalman filter.
[0056] When the quality of visual data is good (e.g., small visual covariance matrix, low uncertainty), the Kalman filter assigns higher weights to visual positions, making the final sound source location more consistent with the visual localization result. When the quality of visual data is poor (large visual covariance matrix, high uncertainty), the Kalman filter automatically reduces the weights of visual positions to prevent visual noise from causing drift or deviation in sound source localization.
[0057] Determining the visual covariance matrix based on visual data includes the following steps: determining multiple types of visual noise penalty parameters based on visual data, wherein the visual noise penalty parameters are used to characterize the degree of uncertainty of visual position, and generating the visual covariance matrix based on the multiple types of visual noise penalty parameters.
[0058] The visual noise penalty parameter is used to characterize the uncertainty of visual position. The magnitude of the visual noise penalty parameter is positively correlated with the uncertainty of visual position; the larger the visual noise penalty parameter, the greater the uncertainty of visual position, and the less reliable the visual position. Conversely, the smaller the visual noise penalty parameter, the smaller the uncertainty of visual position, and the more reliable the visual position. This application's embodiments integrate multiple types of visual noise penalty parameters to quantify the reliability of the estimated visual position, providing a solid data foundation for obtaining more accurate and reliable visual positions.
[0059] In some embodiments, multiple visual noise penalty parameters include occlusion rate, which represents the degree to which the facial region of the target person is occluded by external factors. When the target person is occluded by car seats, other passengers, car windows, etc., image features of key parts such as the head will be missing (e.g., only the body can be seen, but the head cannot). For example, occlusion can cause the outline, texture, key points, and other features of the target person to be incomplete or interfered with, and the visual algorithm cannot capture the core features corresponding to the sound source location, resulting in the inability to accurately calculate the visual position.
[0060] Determining multiple types of visual noise penalty parameters based on visual data includes the following steps: determining the occlusion rate based on visual data and a preset occlusion penalty function.
[0061] The occlusion penalty function is used to obtain the occlusion rate. It dynamically quantifies the decrease in visual localization reliability caused by facial occlusion and transforms it into a penalty value that can be used for mathematical fusion. The occlusion rate is output by the occlusion penalty function. The higher the occlusion rate, the less reliable the visual localization becomes due to occlusion. The smaller the value, the more reliable the visual positioning is due to reduced occlusion.
[0062] The visual data includes multiple frames of captured images. Determining the occlusion rate based on the visual data and a preset occlusion penalty function includes the following steps: extracting the face region of the target person and the occlusion area of the face region from the captured images; determining the face area of the face region and the occlusion area of the occlusion area; calculating a first ratio between the occlusion area and the face area; and multiplying the first ratio by a preset occlusion penalty coefficient based on the occlusion penalty function to obtain the occlusion rate.
[0063] Both the face region and the occlusion region are dynamic parameters calculated in real time based on each frame of captured images, while the occlusion penalty coefficient is a parameter calibrated in advance by the designer based on engineering experience. Extracting the face region of a target person from a captured image includes the following steps: performing face extraction processing on the captured image based on a pre-trained face detection model to obtain the face region of the target person.
[0064] Extracting the occluded region of a face from a captured image involves the following steps: Using a pre-trained face segmentation model, the captured image is processed to extract occluded regions. These occluded regions include objects such as hands, mobile phones, clothing, and chairs.
[0065] For example, the occlusion penalty function is: ,in, For occlusion rate, The occlusion penalty coefficient, The area covered by the shading is the area of the shading region. The face area is the region of the face. This is the first ratio.
[0066] The occlusion rate is a normalized value, ranging from [0,1]. Higher occlusion levels result in a higher occlusion rate, and lower occlusion levels result in a lower occlusion rate. As can be seen from the expression of the occlusion penalty function, the occlusion rate proposed in this embodiment is a normalized index, independent of the absolute number of pixels. It can map the severity of occlusion to the contribution to the uncertainty of visual positioning through a linear occlusion penalty coefficient. This allows the Kalman filter to adaptively adjust the weights of the visual position, taking into account the noise of occlusion, which is beneficial for obtaining more reliable and accurate sound source locations.
[0067] This application embodiment constructs an occlusion rate to accurately quantify the impact of target person being occluded on visual positioning, providing an uncertainty parameter of "occlusion dimension" for the subsequent visual covariance matrix, enabling the Kalman filter to automatically adjust the weight of visual position and reduce positioning errors caused by occlusion.
[0068] In some embodiments, the multiple visual noise penalty parameters include illumination deviation rate, which is used to represent the degree to which the target person is affected by the lighting inside the vehicle.
[0069] Abnormal lighting can impair the clarity of features of a target person in visual data, causing visual algorithms to fail to accurately extract the key features required for localization. Specifically, low light blurs key features related to the sound source in the target person, while backlighting and reflections can lead to feature loss or overexposure, resulting in feature extraction bias and a mismatch between the visual position and the actual position of the person. At the same time, dynamic changes in lighting reduce the stability of visual data and increase localization uncertainty.
[0070] Determining multiple types of visual noise penalty parameters based on visual data includes the following steps: determining the illumination deviation rate based on visual data and a preset illumination penalty function.
[0071] The illumination penalty function is used to obtain the illumination deviation rate. It dynamically quantifies the attenuation of visual positioning accuracy caused by excessively dim or bright ambient light, converting it into a calculable penalty value to adjust the weights of visual position in the Kalman filtering process. The illumination penalty function outputs the illumination deviation rate. The larger the value, the less reliable the visual positioning is due to poor lighting; the higher the lighting deviation rate. The smaller the value, the more reliable the visual positioning is due to better lighting.
[0072] Determining the illumination deviation rate based on visual data and a preset illumination penalty function includes the following steps: determining the actual illumination intensity of the captured image, calculating the absolute difference between the preset optimal illumination value and the actual illumination intensity, calculating the second ratio of the absolute difference to the optimal illumination value, and multiplying the second ratio by a preset illumination penalty coefficient based on the illumination penalty function to obtain the illumination deviation rate.
[0073] In some embodiments, determining the actual illumination intensity of a captured image includes the following steps: extracting the face region of the target person from the captured image, obtaining the brightness values of all pixels in the face region, calculating the average of the brightness values of all pixels, and obtaining the actual illumination intensity.
[0074] In other embodiments, a light sensor is installed inside the vehicle to collect the ambient light intensity and obtain the actual light intensity.
[0075] It is understandable that actual light intensity is a highly dynamic value, and it changes rapidly when a vehicle enters a tunnel, passes through a shady area, or drives at night.
[0076] The optimal illumination value is the best illumination intensity that allows for accurate and reliable visual positioning. It measures how much the actual illumination intensity deviates from the optimal illumination intensity. The optimal illumination value can be customized by the designer based on engineering experience.
[0077] The illumination penalty coefficient is a parameter calibrated in advance by the designer based on engineering experience, and it expresses the sensitivity of the system to changes in illumination.
[0078] In some embodiments, the illumination penalty function is: ,in, This refers to the illumination deviation rate. This is the light penalty factor. For optimal illumination, The absolute difference The second ratio, This represents the actual light intensity.
[0079] The illumination deviation rate is a normalized value, ranging from [0,1]. Higher illumination deviation results in a higher illumination deviation rate, and lower illumination deviation results in a lower illumination deviation rate. As can be seen from the expression of the illumination penalty function, the illumination deviation rate proposed in this embodiment is a normalized index. It can map the degree of illumination deviation to a contribution to the uncertainty of visual positioning through a linear illumination penalty coefficient. This allows the Kalman filter to adaptively adjust the weight of the visual position, taking into account the noise of illumination deviation. This is beneficial for obtaining a more reliable and accurate sound source location and solves the problem of decreased visual positioning accuracy caused by complex in-vehicle lighting (such as backlight, low light at night, and glare from the central control screen).
[0080] This application embodiment constructs an illumination deviation rate to accurately quantify the impact of illumination on visual position, providing an uncertainty parameter of "illumination dimension" for the subsequent visual covariance matrix, enabling the Kalman filter to automatically adjust the weight of visual position and reduce positioning errors caused by illumination.
[0081] In some embodiments, the multiple visual noise penalty parameters include a posture deviation rate, which represents the degree of head deflection of the target person.
[0082] Head rotation can obscure or render key parts associated with the sound source invisible. In this application, the visual position is determined by locating the mouth of the target person (i.e., the sound source location). When the head rotates (e.g., turning left or right, looking down, or looking up), key parts strongly related to the sound source, such as the mouth and the center point of the head, may be obscured by facial structures (e.g., cheeks, forehead) or exceed the camera's field of view. This prevents the visual algorithm from capturing the core features corresponding to the sound source location, thus hindering the accurate determination of the visual position.
[0083] Visual algorithms (such as object detection and keypoint recognition networks) estimate head position based on features of a standard frontal or frontal head, with pre-defined head center point and keypoint extraction logic adapted to the frontal head shape. When the head turns, the relative positions of the head's contours, textures, and keypoints (such as the corners of the eyes and mouth) change. The visual algorithm may misjudge the head center point position (e.g., misjudge the center point of the side profile as the center point of the frontal head), resulting in a deviation between the output visual position and the sound source position.
[0084] Determining multiple types of visual noise penalty parameters based on visual data includes the following steps: determining the posture deviation rate based on visual data and a preset posture penalty function.
[0085] The posture penalty function is used to obtain the posture deviation rate. Specifically, it dynamically quantifies the localization error caused by visual feature deformation or occlusion of the mouth due to head rotation, and converts it into a calculable penalty value to adjust the weights of the visual position in the Kalman filtering process. The posture penalty function outputs the posture deviation rate. The larger the value, the less reliable the visual positioning becomes due to head rotation; this indicates a higher posture deviation rate. The smaller the value, the more reliable the visual positioning is due to minimal or no head rotation.
[0086] Determining the posture deviation rate based on visual data and a preset posture penalty function includes the following steps: extracting the face region of the target person from the captured image; performing posture estimation on the face region based on a preset head posture estimation algorithm to obtain the yaw angle about the face region; calculating the absolute value of the cosine function value of the yaw angle; calculating the difference between the natural number 1 and the absolute value of the cosine function value; and multiplying the difference with a preset posture penalty coefficient based on the posture penalty function to obtain the posture deviation rate.
[0087] The process of estimating the pose of a face region based on a pre-defined head pose estimation algorithm to obtain the yaw angle of the face region includes the following steps: inputting the image of the face region into a pre-trained head pose estimation network to obtain the yaw angle of the face region. The head pose estimation network can be a 3DDFA model or a MediaPipe Face Mesh model.
[0088] Yaw angle is used to indicate the degree of head deflection of a target person. The range of head deflection angle is [-90°, 90°], and the corresponding yaw angle range is [-90°, 90°]. When the yaw angle is 0°, the target person's head has not deflected.
[0089] To obtain a normalized value for the yaw angle, this embodiment of the application uses a cosine function to normalize the yaw angle. The effect of the user's head turning to the left on visual positioning is the same as the effect of the user's head turning to the right on visual positioning. Therefore, this embodiment of the application performs absolute processing on the cosine function value of the yaw angle to obtain the absolute value of the cosine function value.
[0090] To express the positive correlation between the degree of head deflection and the posture deviation rate, this embodiment of the application subtracts the absolute value of the natural number 1 from the value of the cosine function. The resulting difference positively reflects the degree of head deflection. Since the degree of head deflection is positively correlated with the uncertainty of visual position, this difference is also positively correlated with the uncertainty of visual position.
[0091] The posture penalty coefficient is a parameter calibrated in advance by the designer based on engineering experience, which expresses the system's sensitivity to head deflection.
[0092] In some embodiments, the attitude penalty function is: ,in, This refers to the attitude deviation rate. The attitude penalty coefficient is... Yaw angle The difference is the sum of its parts.
[0093] The posture deviation rate is a normalized value, ranging from [0,1]. The higher the degree of head deflection, the higher the posture deviation rate; the lower the degree of head deflection, the lower the posture deviation rate.
[0094] The head posture of occupants inside a vehicle is dynamically changing (e.g., talking, looking in the rearview mirror, looking at the central control screen). Changes in the head deflection angle can cause instability in the head features extracted by the visual algorithm, leading to fluctuations in the visual position estimation results. As can be seen from the expression of the posture penalty function, the posture deviation rate proposed in this application is a normalization index. It can map the degree of head deflection to the contribution to the uncertainty of visual positioning through a linear posture penalty coefficient. This allows the Kalman filter to adaptively adjust the weight of visual position considering the noise of head deflection, which is beneficial for obtaining more reliable and accurate sound source positions and solves the problem of decreased visual positioning accuracy caused by severe head deflection of the target person inside the vehicle.
[0095] This application embodiment constructs an attitude deviation rate to accurately quantify the impact of head deflection on visual position, providing an uncertainty parameter of "head attitude transmission change" for the subsequent visual covariance matrix, enabling the Kalman filter to automatically adjust the weight of visual position and reduce the positioning error caused by head deflection.
[0096] Overall, the occlusion rate, illumination deviation rate, and attitude deviation rate correspond to the three core interference factors (occlusion, illumination, and attitude) of in-vehicle visual positioning, comprehensively covering the noise sources of visual data and avoiding incomplete quantification of uncertainties caused by a single parameter.
[0097] Occlusion rate, illumination deviation rate, and posture deviation rate together provide multi-dimensional inputs for the generation of the visual covariance matrix, enabling the visual covariance matrix to accurately reflect the overall uncertainty of visual positioning, providing a reliable basis for the adaptive fusion of Kalman filtering, and ultimately improving the accuracy and robustness of in-vehicle sound source localization.
[0098] The steps for generating the visual covariance matrix of the target person based on multiple types of visual noise penalty parameters are as follows: weighted summation of multiple types of visual noise penalty parameters to obtain the visual penalty result, and generating the visual covariance matrix of the target person based on the visual penalty result.
[0099] The weighted summation of multiple visual noise penalty parameters to obtain the visual penalty result includes the following steps: calculating the first product of the preset first weight coefficient and the occlusion rate, calculating the second product of the preset second weight coefficient and the illumination deviation rate, calculating the third product of the preset third weight coefficient and the posture deviation rate, and adding the first product, the second product and the third product to obtain the visual penalty result.
[0100] For example, the expression for the visual penalty result is as follows:
[0101] in, As a result of visual punishment, As the first weighting coefficient, This is the second weighting coefficient. This is the third weighting coefficient.
[0102] This application implements a "comprehensive quantification" of three types of visual noise penalty parameters, avoiding the problem that a single parameter cannot fully reflect the overall uncertainty of visual positioning. It integrates the three interference factors of occlusion, illumination, and posture into a unified quantitative index. In addition, this application implements different weights for different types of visual noise penalty parameters, which can specifically highlight the differences in the impact of different noises on visual positioning (for example, the interference of occlusion on sound source positioning inside the vehicle is usually greater than that of illumination interference, so it can be assigned a higher weight), making the visual penalty results more consistent with the actual scene inside the vehicle, which is conducive to obtaining accurate and reliable sound source locations.
[0103] First weighting coefficient Second weighting coefficient Third weighting coefficient This is used to balance the relative contributions of the three penalty terms—occlusion, illumination, and pose—to the total uncertainty. First weighting coefficient. Second weighting coefficient Third weighting coefficient Its functions are as follows: 1) Balancing the contributions of multiple sensor modes: This is responsible for answering a key question: "In the current in-vehicle environment, which factor has the greatest negative impact on positioning reliability: occlusion, illumination, or attitude?"
[0104] 2) Achieving cross-scenario adaptation: By using fixed weights, the system achieves a static optimum. For example, the first weight coefficient... Second weighting coefficient Third weighting coefficient This determines the system's priority for various factors in different typical scenarios, such as "driving on highways during the day" (where lighting changes rapidly but head movement is minimal) and "conversing in urban areas at night" (where lighting is dim and head movement is frequent).
[0105] 3) Optimizing the final fusion performance: The ultimate goal of weighting is not to provide an accurate uncertainty estimate, but to ensure that the Kalman filter, which depends on the visual covariance matrix Pv, outputs an accurate sound source location. The three weight coefficients indirectly control whether the Kalman filter trusts the observed visual location or the optimal estimated location more by influencing the visual covariance matrix, thereby optimizing the sound source localization accuracy of the entire system.
[0106] First weighting coefficient Second weighting coefficient Third weighting coefficient The source is a systematic, data-driven engineering process, and the process is determined as follows: A1: Construct a comprehensive dataset. Collect a large-scale, full-coverage dataset. The dataset needs to include various combinations of occlusion, lighting, and pose, and must be acquired synchronously: video streams (used to calculate visual position (Pv) and occlusion rate). Illumination deviation rate Attitude deviation rate The mouth is provided with true 3D coordinates by a high-precision truth system (Vicon motion capture).
[0107] A2: Define the optimization objective. The ultimate goal is not to control the visual position Pv to be as close as possible to the true 3D coordinates Pv_true, but rather to make the final sound source position Pf, after Kalman filtering fusion, as close as possible to the true 3D coordinates Pv_true. Therefore, the optimization objective function is: to minimize the average localization error between the sound source position Pf and the true 3D coordinates Pv_true.
[0108] A3: Perform regression analysis. Treat the entire Kalman fusion system (including the visual uncertainty model) as a whole; the design variables are... , , Using automated hyperparameter tuning algorithms such as Bayesian optimization or grid search, we search within the large dataset for the set of parameters that minimizes the final localization error. , , ]combination.
[0109] , , Once confirmed through the aforementioned offline process, it is burned into the system. , , This represents a statistical understanding of the patterns in the in-vehicle environment. , , It is obtained by optimizing on massive amounts of data, representing the optimal decision-making strategy in the sense of statistical average, and ensuring the overall optimal performance of the system in most scenarios.
[0110] This application provides a complete set of parameters for each coefficient in the model. , , , , , A complete experimental methodology for calibration and optimization was developed. Although regression analysis was used, a controlled and systematic data acquisition and fitting process was designed for in-vehicle visual positioning to determine parameters for a specific model.
[0111] The visual covariance matrix is a diagonal covariance matrix; for example, it is a 3x3 diagonal matrix. The visual covariance matrix includes first visual diagonal elements, second visual diagonal elements, and third visual diagonal elements. The target space is configured with a three-dimensional coordinate system, and each visual diagonal element corresponds to a coordinate axis of the three-dimensional coordinate system.
[0112] For example, the expression for the visual covariance matrix is as follows:
[0113] in, The visual covariance matrix, The first visual diagonal element corresponding to the x-axis of the three-dimensional coordinate system xyz. The second visual diagonal element corresponding to the y-axis of the three-dimensional coordinate system xyz. The third visual diagonal element corresponding to the z-axis of the three-dimensional coordinate system xyz.
[0114] The steps for generating the visual covariance matrix of the target person based on the visual penalty result are as follows: obtaining the first visual basic calibration error corresponding to the x-axis direction of the three-dimensional coordinate system, the second visual basic calibration error corresponding to the y-axis direction of the three-dimensional coordinate system, and the third visual basic calibration error corresponding to the z-axis direction of the three-dimensional coordinate system; determining the first visual diagonal element based on the first visual basic calibration error and the visual penalty result; determining the second visual diagonal element based on the second visual basic calibration error and the visual penalty result; determining the third visual diagonal element based on the third visual basic calibration error and the visual penalty result; and generating the visual covariance matrix of the target person based on the first visual diagonal element, the second visual diagonal element, and the third visual diagonal element.
[0115] In this embodiment, the first visual diagonal element, the second visual diagonal element, and the third visual diagonal element are calculated according to the following formulas, as shown below:
[0116] The first visual baseline calibration error corresponds to the x-axis direction of the three-dimensional coordinate system. The second visual baseline calibration error corresponds to the y-axis direction of the three-dimensional coordinate system. The third visual basis calibration error corresponds to the z-axis direction of the three-dimensional coordinate system.
[0117] The basic calibration errors of the first vision, second vision, and third vision all represent the inherent errors of the vision system under ideal conditions (frontal view, good lighting, no occlusion), and include the following: (1) Camera intrinsic parameter calibration error: residual errors in the calibration process of focal length, principal point coordinates, distortion coefficient, etc.
[0118] (2) Camera external parameter calibration error: Measurement error of the installation position and angle of the camera device inside the vehicle.
[0119] (3) Depth estimation model error: The minimum error inherent in ranging, whether it is a binocular vision camera device or a TOF camera device.
[0120] In this embodiment, the visual penalty result is substituted into a preset covariance matrix generation rule, and the visual penalty result is mapped to the variance element of the visual covariance matrix. This enables the visual covariance matrix to accurately reflect the overall uncertainty of visual positioning. When the visual penalty result is larger (the more severe the noise interference), the larger the variance of the visual covariance matrix, the Kalman filter will automatically reduce the weight of the visual position to avoid sound source positioning drift caused by visual noise.
[0121] Step S23: Determine the audio location and audio covariance matrix of the target sound source based on the audio data.
[0122] The audio location is estimated based on audio data to pinpoint the location of the target person. This application's embodiments estimate the audio location of the target person in a preset three-dimensional coordinate system based on either a time delay estimation (TDOA / GCC-PHAT) algorithm or a spatial spectrum (SRP-PHAT) algorithm. .in, Let x be the coordinates of the target person along the x-axis in a three-dimensional coordinate system. Let y be the coordinates of the target person in the 3D coordinate system along the y-axis. Let Z be the coordinates of the target person along the z-axis in a three-dimensional coordinate system. This indicates the audio location.
[0123] Determining the audio covariance matrix based on audio data includes the following steps: determining multiple types of audio noise penalty parameters based on the audio data, and generating the audio covariance matrix based on the multiple types of audio noise penalty parameters.
[0124] Audio noise penalty parameters are used to characterize the uncertainty of audio location. The magnitude of the audio noise penalty parameter is positively correlated with the uncertainty of audio location; the larger the audio noise penalty parameter, the greater the uncertainty of audio location, and the less reliable the audio location. Conversely, the smaller the audio noise penalty parameter, the smaller the uncertainty of audio location, and the more reliable the audio location. This application's embodiments integrate multiple types of audio noise penalty parameters to quantify the reliability of the estimated audio location, providing a solid data foundation for obtaining more accurate and reliable audio locations.
[0125] The audio covariance matrix is the observation noise covariance matrix of the Kalman filter. It integrates multiple types of audio noise penalty parameters, quantifying the uncertainty of audio location from multiple dimensions. The audio covariance matrix provides the basis for the observation weights in the Kalman filter.
[0126] When the audio data quality is good (e.g., low covariance, low uncertainty), the Kalman filter assigns higher weights to the audio location, making the final sound source location more consistent with the audio localization result. When the audio data quality is poor (high covariance, high uncertainty), the Kalman filter automatically reduces the weights of the audio location to prevent audio noise from causing drift or deviation in sound source localization.
[0127] In some embodiments, multiple audio noise penalty parameters include a noise level ratio, which represents the proportion of noise intensity in the in-vehicle environment. High noise intensity can easily lead to positioning ambiguity.
[0128] The stronger the noise level inside the vehicle and the more chaotic the environment, the worse the quality of the underlying signal for acoustic positioning, and the greater the uncertainty in positioning. Conversely, the weaker the noise level inside the vehicle and the quieter the environment, the better the quality of the underlying signal for acoustic positioning, and the lower the uncertainty in positioning.
[0129] Determining multiple types of audio noise penalty parameters based on audio data includes the following steps: determining the noise level rate based on audio data and a preset noise level diagnostic function.
[0130] The noise level diagnostic function is a function used to obtain the noise level rate, specifically used to diagnose the signal-to-noise ratio of the in-vehicle environment. The noise level diagnostic function outputs the noise level rate. The larger the value, the less reliable the audio localization becomes due to higher ambient noise levels; noise level ratio. The smaller the value, the more reliable the audio localization is due to lower ambient noise levels.
[0131] The noise level diagnostic function is an exponential function with base e. Determining the noise level rate based on audio data and a preset signal-to-noise penalty coefficient includes the following steps: calculating the signal-to-noise ratio based on the audio data, obtaining the preset signal-to-noise penalty coefficient and the preset attenuation coefficient, multiplying the inverse of the attenuation coefficient by the signal-to-noise ratio, using the result as the superscript of the base e of the noise level diagnostic function to obtain the exponential value, and multiplying the signal-to-noise penalty coefficient by the exponential value to obtain the noise level rate.
[0132] Signal-to-noise ratio (SNR) is used to reflect the audio quality of the in-vehicle environment. The higher the SNR, the weaker the noise in the in-vehicle environment and the higher the audio quality. Conversely, the lower the SNR, the stronger the noise in the in-vehicle environment and the worse the audio quality.
[0133] The signal-to-noise ratio (SNR) is inversely proportional to the noise level. The higher the SNR, the lower the noise level, and vice versa.
[0134] Calculating the signal-to-noise ratio (SNR) based on audio data includes the following steps: dividing the audio data into multiple audio segments according to a preset step size; detecting whether the audio segments contain the target person's speech based on a speech activity detection algorithm; when the audio segment does not contain the target person's speech, setting the audio segment as a background noise segment; performing power spectral density estimation on the background noise segment to obtain the noise power, which is used to represent the energy distribution of noise at different frequencies; when the audio segment contains the target person's speech, performing total power calculation on the audio segment to obtain the total signal power; and calculating the SNR based on the total signal power and the noise power.
[0135] The signal-to-noise ratio is calculated according to the following formula in the embodiments of this application, as shown below:
[0136] in, The total signal power, For noise power, This refers to the signal-to-noise ratio.
[0137] The signal-to-noise penalty coefficient is used to calibrate the contribution of the signal-to-noise ratio to the uncertainty of the audio position. The signal-to-noise penalty coefficient is directly proportional to the noise level rate. When the signal-to-noise penalty coefficient is larger, the noise level rate is larger, and when the signal-to-noise penalty coefficient is smaller, the noise level rate is smaller.
[0138] The attenuation factor is used to control the degree of uncertainty in audio location as the signal-to-noise ratio deteriorates.
[0139] The signal-to-noise penalty coefficient and attenuation coefficient are determined in advance by the designer. Specifically, the process for determining the signal-to-noise penalty coefficient and attenuation coefficient is as follows: 1) Data collection: 1.1) Create an environment in an anechoic chamber or a real vehicle where the signal-to-noise ratio (SNR) changes continuously.
[0140] 1.2) Use a fixed sound source (such as an artificial mouth) to play the test signal and ensure its first true position. It is precisely recorded by high-precision equipment (laser tracker).
[0141] 1.3) Synchronously record the audio data of the microphone array.
[0142] 2) Data preprocessing and labeling: 2.1) For each audio data segment, calculate its SNR value.
[0143] 2.2) Run the SRP-PHAT localization algorithm to obtain the location of the first audio sample. .
[0144] 2.3) Calculate the squared error of this positioning: Thus, we have a dataset where each sample is the squared localization error.
[0145] 3) Nonlinear curve fitting: 3.1) Model Construction: Using the model Describe the square of the positioning error The relationship between SNR and SNR.
[0146] 3.2) Optimization objective: Based on the least squares method, determine a set of ( , This makes the predicted values calculated by the above model consistent with the actual measurements in the dataset. The overall difference between the values is minimal. Among them, The signal-to-noise penalty coefficient is... This is the attenuation coefficient.
[0147] 3.3) Performing the fitting: Input the dataset into the programming library (Python's scipy.optimize.curve_fit), perform nonlinear least squares fitting, the algorithm iterates automatically, and finally outputs the optimal result. and The value of .
[0148] It is worth noting that the signal-to-noise penalty coefficient With attenuation coefficient These are a pair of twin parameters, simultaneously determined in the same nonlinear regression process, and the fitted signal-to-noise penalty coefficients are... With attenuation coefficient This represents the optimal quantification of the microphone array's "sensitivity" to noise in a specific vehicle environment.
[0149] The noise level diagnostic function is: ,in, Noise level rate, The signal-to-noise penalty coefficient is... The attenuation coefficient is... For signal-to-noise ratio, This is the exponential value.
[0150] The noise level rate is a normalized value, ranging from [0,1]. The noise level diagnostic function provided in this application uses an exponential decay model to describe the relationship between acoustic positioning performance and SNR. When the SNR is high, the in-vehicle environment is relatively quiet with weak noise, and the noise level rate output by the exponential decay model is low. When the SNR is low, the in-vehicle environment is relatively noisy with strong noise, and the noise level rate output by the exponential decay model increases sharply, indicating that the underlying signal of acoustic positioning has been severely contaminated, and the reliability of the results drops precipitously. Therefore, the exponential decay model perfectly fits this "breakdown point" effect, reflecting the real physical phenomenon better than the linear model and enhancing sensitivity to noisy environments.
[0151] The noise level rate proposed in this application is a normalized index. It can map the severity of noise to the contribution to the uncertainty of audio localization through an exponential noise level diagnostic function. This allows the Kalman filter to adaptively adjust the weight of the audio location taking into account the noise situation, which is beneficial to obtaining a more reliable and accurate sound source location.
[0152] In some embodiments, multiple audio noise penalty parameters include sound field distortion, which is used to diagnose the coherence between signals acquired by each element in the microphone array.
[0153] Coherence is used to characterize the similarity and synchronization of signals acquired by each element of a microphone array. Its core measure is whether the signals acquired by each element originate from the same target person's sound source and whether the phase and amplitude changes of the signals are consistent. High coherence is defined as the signals acquired by each element primarily being the direct sound of the target person, with similar waveforms, synchronized phases, and consistent amplitude trends. Low coherence is defined as the signals are affected by factors such as in-vehicle reverberation, diffused noise (e.g., wind noise, air conditioning noise), and the target person's head movement or obstruction, causing each element to simultaneously receive direct sound, multipath reflections, and irregular noise, resulting in chaotic signal waveforms, disordered phases, and uncorrelated amplitudes.
[0154] It should be noted that the sound source localization algorithms (such as GCC-PHAT and SRP-PHAT) used in the embodiments of this application are based on the premise that the signals collected by each array element have high coherence. Low coherence will destroy this premise, leading to problems such as localization ambiguity and drift.
[0155] Determining multiple types of audio noise penalty parameters based on audio data includes the following steps: determining the sound field distortion based on audio data and a preset sound field health diagnostic function.
[0156] The sound field health diagnostic function is used to obtain the sound field distortion degree, specifically to assess the distortion of the acoustic environment. The acoustic environment is primarily affected by reverberation, which causes sound to propagate along different paths, disrupting the ideal wavefront structure of the signal between microphone pairs and rendering sound source localization algorithms ineffective. Specifically, the lower the coherence, the more severe the reverberation, and the less reliable the localization.
[0157] Determining the sound field distortion degree based on audio data and a preset sound field health diagnosis function includes the following steps: calculating the average coherence value based on the audio data, subtracting the natural number 1 from the average coherence value to obtain the coherence penalty factor, obtaining the preset sound field distortion coefficient, and multiplying the sound field distortion coefficient and the coherence penalty factor based on the preset sound field health diagnosis function to obtain the sound field distortion degree.
[0158] The average coherence value is used to quantify the coherence between any two elements in a microphone array, and its value ranges from 0 to 1. The average coherence value is positively correlated with the coherence level; the higher the coherence level, the closer the average coherence value is to 0, and the lower the coherence level, the closer the average coherence value is to 1.
[0159] Calculating the average coherence value based on audio data includes the following steps: 1. Audio data preprocessing: The audio data acquired by the microphone array is segmented and windowed to obtain audio segments. Then, a fast Fourier transform is performed on each frame of audio segment to convert the time-domain audio signal into a frequency-domain signal, obtaining the frequency domain amplitude and phase information of each frame of audio for each array element.
[0160] 2. Calculation of coherence coefficient between pairs of array elements: Select any two array elements in the microphone array as a pair, and calculate the cross power spectrum and self power spectrum between the two array elements based on the frequency domain signal; where the cross power spectrum characterizes the correlation between the signals of the two array elements at each frequency point, and the self power spectrum characterizes the energy distribution of the self signal of each array element at each frequency point; the coherence coefficient is obtained by dividing the magnitude of the cross power spectrum by the product of the square roots of the self power spectra of the two channels.
[0161] 3. Coherence coefficient fusion at multiple frequency points: Since there are frequency differences in the in-vehicle audio signal, the coherence coefficient at a single frequency point cannot fully reflect the overall coherence. Therefore, the coherence coefficients at all frequency points in the target audio segment are processed by arithmetic average or weighted average to obtain the overall coherence coefficient between the two channels.
[0162] 4. Calculation of average coherence value: Repeat steps 2-3 to calculate the overall coherence coefficient of all paired elements in the microphone array, and then average the coherence coefficients of all paired elements to obtain the average coherence value of the entire microphone array.
[0163] The average coherence value is a normalized value. In this embodiment, the natural number 1 is subtracted from the average coherence value. The resulting coherence penalty factor is still a normalized value. The lower the average coherence value (the weaker the coherence), the closer the coherence penalty factor is to 1, and the better it can characterize the degree of sound field distortion caused by in-vehicle reverberation and diffused noise.
[0164] The sound field distortion coefficient is an empirical or calibration value adapted to the vehicle scenario. It can be flexibly adjusted according to different vehicle models and microphone array layouts. It is used to calibrate the overall magnitude of the sound field distortion to ensure that it matches the magnitude of the noise level rate and the location reliability.
[0165] The process for determining the sound field distortion coefficient is as follows: 1) Data Acquisition: Construct a “coherence-error” dataset; this step needs to be performed in a controlled acoustic environment (such as a reverberation chamber or a real vehicle).
[0166] 1.1) Creating different reverberation conditions: This is the main means of altering coherence. For example: Place / remove sound-absorbing materials inside the vehicle.
[0167] Open / close the windows and sunroof to change the reflection of the interior.
[0168] Use professional equipment to play noise signals with different reverberation times.
[0169] 1.2) Synchronous Data Recording: Under each different reverberation condition, the following data is recorded synchronously: Audio stream: Audio captured by the microphone array, used for subsequent calculation of the average coherence level value. Second audio sample position .
[0170] True coordinates: The second true position of the speaker's mouth is recorded using a high-precision positioning system (such as an ultra-wideband UWB positioning system or an optical motion capture system). .
[0171] 2.) Data Analysis: Process each frame of data collected: Calculate coherence: For each frame of audio, calculate the coherence of the average squared amplitude of the signal between microphone pairs, and obtain... .
[0172] Calculate the coherence penalty factor: Calculate (1- The coherence penalty factor ranges from 0 (ideal coherence) to 1 (complete incoherence).
[0173] Calculate acoustic localization error: For the same frame of audio, run the SRP-PHAT algorithm to obtain the audio location. Calculate the squared error of this positioning: Now, we have a dataset consisting of many data points (x, y), where x is the coherence penalty factor (1- ). ), where y is the square of the acoustic positioning error. .
[0174] 3.) Curve fitting: Determining the sound field distortion coefficient .
[0175] To plot a scatter plot: combine all data points ((1- ), Mark it on the coordinate system.
[0176] Model establishment: Assume a linear relationship exists between the two, i.e. This is a reasonable preliminary assumption, indicating that the error variance increases linearly with the deterioration of coherence.
[0177] Perform linear regression: Use the least squares method to fit a straight line to the data points. The goal of linear regression is to find a straight line y = k*x that minimizes the sum of the squares of the perpendicular distances from all data points to this line. The slope k obtained from the fit is the sound field distortion coefficient. .
[0178] The sound field health diagnosis function is: ,in, For sound field distortion, The sound field distortion coefficient is... The average coherence value, (1- ) is the coherence penalty factor.
[0179] The higher the coherence, the closer the sound field distortion is to 0; the lower the coherence, the closer the sound field distortion is to 1. The embodiments of this application achieve adaptive penalty for positioning observations in low coherence scenarios by adjusting the sound field distortion, thereby ensuring positioning accuracy and stability.
[0180] As can be seen from the expression of the sound field health diagnostic function, the sound field distortion proposed in this application embodiment is a normalized index. It can map the sound field distortion to the contribution to the uncertainty of audio localization through a linear sound field health diagnostic function. This allows the Kalman filter to adaptively adjust the weight of the audio position by taking into account the noise of the sound field distortion, which is beneficial to obtaining a more reliable and accurate sound source position and effectively suppressing the interference of in-vehicle reverberation and diffused noise on the calculation of coherence and sound field health.
[0181] This application embodiment constructs a sound field distortion degree to accurately quantify the reverberation effect on the audio position, providing an uncertainty parameter of "sound field distortion" for the subsequent audio covariance matrix, enabling the Kalman filter to automatically adjust the weight of the audio position and reduce the positioning error caused by sound field distortion.
[0182] Traditional sound source localization algorithms (such as the SRP-PHAT algorithm) rely on the peak characteristics of the spatial spectrum for reliable localization. The core principle of the SRP-PHAT algorithm is to obtain the controllable response power (SRP) of each spatial location by summing the GCC-PHAT generalized cross-correlation function of multi-channel audio data in a spatial grid, thus forming a spatial spectrum curve. The algorithm assumes that "the target sound source location corresponds to a unique maximum peak in the spatial spectrum". That is, when the spatial spectrum shows "a sharp single peak and weak surrounding pseudo-peaks", the preliminary localization result (the location corresponding to the maximum peak) is accurate and reliable. However, when there is strong reverberation, multipath reflection, or diffused noise in the vehicle, the spatial spectrum output by the SRP-PHAT algorithm will show phenomena such as multiple peaks, enhanced pseudo-peaks, and flattened peaks. At this time, the maximum peak may correspond to a pseudo-peak (not the real target location), resulting in a large deviation in the preliminary localization result and a significant decrease in reliability.
[0183] In some embodiments, the multiple audio noise penalty parameters include location confidence, which represents the confidence level in the location of the audio source. Location confidence answers questions such as: How certain is the audio location? Are there other possible locations of the sound source? Determining multiple types of audio noise penalty parameters based on audio data includes the following steps: determining location confidence based on audio data and a preset confidence diagnostic function.
[0184] The confidence diagnostic function is a function used to obtain location confidence, specifically to diagnose whether an audio location is reliable. The confidence diagnostic function outputs the location confidence. The higher the value, the less reliable the audio localization is due to more environmental noise sources; the lower the localization reliability. The smaller the value, the more reliable the audio localization is due to fewer sources of environmental noise.
[0185] Determining location reliability based on audio data and a preset confidence diagnostic function includes the following steps: determining the peak-to-surge ratio coefficient based on the audio data, obtaining a preset peak-to-surge ratio penalty coefficient, and dividing the peak-to-surge ratio penalty coefficient by the peak-to-surge ratio coefficient based on the confidence diagnostic function to obtain the location reliability.
[0186] The peak-to-peak ratio coefficient is used to quantify the reliability of spatial spectral peaks. Determining the peak-to-peak ratio coefficient based on audio data includes the following steps: processing audio data based on the SRP-PHAT algorithm to obtain the controllable response power of each spatial grid point. The controllable response power of each spatial grid point forms a three-dimensional spatial power spectrum. The maximum controllable response power is found in the spatial power spectrum. The second largest controllable response power in the surrounding spatial range of the maximum controllable response power is extracted. The ratio of the maximum controllable response power to the second largest controllable response power is calculated to obtain the peak-to-peak ratio coefficient.
[0187] Understandably, the larger the peak-to-peak ratio, the sharper the maximum peak of the spatial spectrum output by the SRP-PHAT algorithm and the weaker the false peaks, resulting in a stronger uniqueness and higher reliability of the preliminary audio localization result. Conversely, the smaller the peak-to-peak ratio, the more likely the spatial spectrum output by the SRP-PHAT algorithm has multiple peaks and obvious false peaks, and is severely affected by reverberation and noise interference, making the preliminary audio localization result less reliable.
[0188] The peak ratio penalty coefficient is used to quantify the impact of the peak ratio on uncertainty. This coefficient is an empirically calibrated value adapted to automotive scenarios and the characteristics of the SRP-PHAT algorithm. It can be flexibly adjusted according to the reverberation characteristics and microphone array layout of different vehicle models to regulate the overall magnitude of location reliability. Specifically, the peak ratio penalty coefficient is a scaling factor that maps the inverse of the dimensionless, relative peak ratio coefficient to a specific contribution value to the audio covariance matrix.
[0189] The process for determining the peak ratio penalty coefficient is as follows: 1) Data Acquisition: Acquire a large amount of audio data, which needs to cover various location-based reliability scenarios, including: 1.1) High confidence scenario: speaking alone in a quiet environment.
[0190] 1.2) Low confidence scenarios: multiple people talking at the same time, strong reverberation environment.
[0191] 1.3) The third real location of the sound source needs to be recorded synchronously. .
[0192] 2) Data preprocessing and labeling: 2.1) For each frame of data, run the SRP-PHAT algorithm to obtain the position of the third audio sample. Peak ratio coefficient .
[0193] 2.2) Calculate the square of the positioning error: .
[0194] 3) Model building and fitting: 3.1) Assume the square of the positioning error is equal to... Proportional: .
[0195] 3.2) Transfer the data points ( , Perform linear regression (usually with a forced intercept of 0).
[0196] 3.3) The slope obtained from linear regression is the peak-to-slope penalty coefficient. .
[0197] The expression for the confidence diagnostic function is: .in, For location reliability.
[0198] As can be seen from the expression of the confidence diagnostic function, when the peak-to-peak ratio coefficient is small (the spatial spectrum output by the SRP-PHAT algorithm has obvious multi-peaks and spurious peaks, and the initial positioning is unreliable), the positioning confidence increases, and the subsequent penalty for the audio position output by the SRP-PHAT algorithm is strengthened accordingly; when the peak-to-peak ratio coefficient is large (the spatial spectrum output by the SRP-PHAT algorithm has a sharp single peak, and the initial positioning is reliable), the positioning confidence decreases, and the penalty is weakened accordingly, realizing adaptive adjustment of the positioning reliability of the SRP-PHAT algorithm and improving its positioning stability in vehicle-mounted strong reverberation and multi-noise scenarios.
[0199] The location reliability proposed in this application is a normalized index, which can map the multi-peak interference of the spatial spectrum into a contribution to the uncertainty of audio localization through linear location reliability. This allows the Kalman filter to adaptively adjust the weight of the audio location considering multi-source noise, which is beneficial to obtaining a more reliable and accurate sound source location.
[0200] Overall, noise level rate, sound field health, and location reliability comprehensively quantify the uncertainty of audio observation from three dimensions: noise intensity, sound field quality, and location reliability. This provides accurate quantitative basis for the subsequent construction of adaptive audio covariance matrix and Kalman filtering, ensuring the accuracy and stability of vehicle sound source localization.
[0201] In some embodiments, generating the audio covariance matrix of the target person based on multiple types of audio noise penalty parameters includes the following steps: summing the multiple types of audio noise penalty parameters to obtain the audio penalty result, and generating the audio covariance matrix of the target person based on the audio penalty result.
[0202] In some embodiments, the present application can directly sum up multiple types of audio noise penalty parameters to obtain the audio penalty result.
[0203] For example, the expression for the audio penalty result is as follows:
[0204] in, This is the result of the audio penalty.
[0205] In other embodiments, the embodiments of this application can perform weighted summation processing on multiple types of audio noise penalty parameters. Specifically, the embodiments of this application calculate the first product of a preset third weighting coefficient and the noise level rate, calculate the second product of a preset fourth weighting coefficient and the sound field distortion, calculate the third product of a preset fifth weighting coefficient and the location confidence, and add the first, second, and third product results to obtain the audio penalty result.
[0206] For example, the expression for the audio penalty result is as follows:
[0207] in, For audio penalty results, It is the fourth weighting coefficient. This is the fifth weighting coefficient. It is the sixth weighting coefficient.
[0208] This application implements a "comprehensive quantification" of three types of audio noise penalty parameters, avoiding the problem that a single parameter cannot fully reflect the overall uncertainty of audio localization. It integrates noise, sound field distortion, and multi-peak interference into a unified quantitative index. Furthermore, by configuring different weights for different types of audio noise penalty parameters, this application can specifically highlight the differences in the impact of different noises on audio localization, making the audio penalty results more closely resemble the actual in-vehicle scenario and facilitating the acquisition of accurate and reliable sound source locations.
[0209] The audio covariance matrix is a diagonal covariance matrix; for example, it is a 3x3 diagonal matrix. The audio covariance matrix includes first, second, and third audio diagonal elements. The target space is configured with a three-dimensional coordinate system, and each audio diagonal element corresponds to a coordinate axis of the three-dimensional coordinate system.
[0210] For example, the expression for the audio covariance matrix is as follows:
[0211] in, The audio covariance matrix, The first audio diagonal element corresponding to the x-axis of the three-dimensional coordinate system xyz. The second audio diagonal element corresponding to the y-axis of the three-dimensional coordinate system xyz. This is the third audio diagonal element corresponding to the z-axis of the three-dimensional coordinate system xyz.
[0212] The steps for generating the audio covariance matrix of the target person based on the audio penalty result are as follows: the audio penalty result is set as the first audio diagonal element, the second audio diagonal element, and the third audio diagonal element, respectively, and the audio covariance matrix of the target person is generated based on the first audio diagonal element, the second audio diagonal element, and the third audio diagonal element.
[0213] In this embodiment, the first audio diagonal element, the second audio diagonal element, and the third audio diagonal element are calculated according to the following formulas, as shown below:
[0214] In this embodiment, the audio penalty result is substituted into a preset covariance matrix generation rule, and the audio penalty result is mapped to the variance element of the audio covariance matrix. This enables the audio covariance matrix to accurately reflect the overall uncertainty of audio localization. When the audio penalty result is larger (the more severe the noise interference), the larger the variance of the audio covariance matrix, the Kalman filter will automatically reduce the weight of the audio position to avoid sound source localization drift caused by audio noise.
[0215] Step S24: Based on a preset Kalman filter, the visual position, visual covariance matrix, audio position, and audio covariance matrix are fused to obtain the sound source position of the target sound source.
[0216] This application embodiment uses a preset Kalman filter to perform Kalman filtering on the visual covariance matrix and visual position to obtain candidate sound source positions. Based on the audio covariance matrix, it generates credibility information of the audio position. Based on the credibility information, it performs fusion processing on the candidate sound source positions, audio positions, and audio covariance matrix to obtain the sound source position of the target sound source.
[0217] This application's embodiments fundamentally improve the state-space model of the Kalman filter, constructing a physically meaningful uniform kinematic model for dynamic tracking of in-vehicle sound sources. This state-space model is not only a container for data fusion but also an "intelligent engine" that endows the system with predictive and smoothing capabilities. Its innovation lies in the following three closely related and mutually supportive elements: 1) Extended definition of state prediction values: from static positioning to dynamic tracking.
[0218] Traditional sound source localization schemes typically limit state predictions to location information. This is essentially a static positioning model. The embodiments of this application extend the state prediction value to... This means that it simultaneously includes the three-dimensional position and three-dimensional velocity of the target sound source, which upgrades the system's task objective from answering "Where is the sound source?" to answering "Where is the sound source and how is it moving?", laying the data structure foundation for subsequent prediction and smoothing.
[0219] 2) The physical meaning of the state transition model: from memory to prediction.
[0220] Based on the extended state prediction values, the state transition model of this application embodiment It is endowed with a clear physical kinematic meaning. The specific form of the state transition matrix F precisely characterizes the physical law of uniform motion: the position update equation. Velocity-based extrapolation prediction was implemented. Velocity update equation. This embodies the uniform velocity assumption, transforming the Kalman filter from a smoother with only "memory function" into a tracker with "short-term prediction capability." It can predict the next position of the target sound source based on the motion trend, thereby significantly improving the continuity and smoothness of tracking.
[0221] 3) Refined modeling of process noise covariance: intelligent compromise on model limitations.
[0222] In this embodiment, the process noise covariance matrix Q is designed as follows: , Let Q be the energy spectral density of the accelerated white noise. The expression for the process noise covariance matrix Q reflects a profound understanding and sophisticated trade-offs in the physical model. The expression for the process noise covariance matrix Q has been shown to be: 3.1) Position uncertainty stems from velocity: The change in position is entirely caused by velocity, therefore, no process noise is directly introduced into the position component.
[0223] 3.2) Velocity is the main source of uncertainty: The "uniform velocity model" is an ideal assumption. The actual sound source motion has unknown acceleration. Therefore, all process noise is concentrated and assigned to the velocity component. Its variance q*dt cleverly simulates the cumulative effect of random acceleration over time.
[0224] The design of the process noise covariance matrix Q ensures that the model achieves an optimal balance between maintaining predictive smoothness (through a uniform velocity model) and adapting to motion variability (through velocity noise).
[0225] 4) As mentioned above, the expression for the state prediction value is: The state prediction value includes a position part and a velocity part. The position part is: Let the coordinates of the target sound source be in the target space; velocity part: The velocity of the target sound source is defined by the three coordinate axes of the three-dimensional coordinate system in the target space.
[0226] This application embodiment introduces a velocity component into the state prediction value. It can not only estimate the location of the target sound source, but also model the dynamic characteristics of the sound source, enabling the system to do the following two things: 4.1) Predictive tracking: Predict the possible location of the target sound source at the next moment based on the motion trend, making the tracking smoother and more resistant to instantaneous jitter of sensor data.
[0227] 4.2) Provide dynamic context: Speed itself is an important state information that can be used to determine whether the target task is turning its head rapidly, etc.
[0228] This application embodiment focuses on state modeling of the characteristics of in-vehicle sound sources and optimizes the process noise distribution assumptions based on the head movement characteristics of the target task. Traditional solutions only know "where" the sound source is, while this application embodiment enables the system to know "how fast and in which direction" the target sound source is moving, achieving the best balance between computational complexity and tracking accuracy.
[0229] It is worth noting that the velocity section provided in this application embodiment is based on the state prediction value of the Kalman filter. The velocity, obtained through internal maintenance and estimation, refers to the instantaneous velocity of the target sound source in a three-dimensional coordinate system. Velocity is a physical quantity containing both magnitude and direction, specifically decomposed into components in three orthogonal directions in the state prediction value: : The instantaneous velocity of the target sound source along the x-axis in a three-dimensional coordinate system. In this system, the x-axis is the opposite direction of the vehicle's forward movement.
[0230] : The instantaneous velocity of the target sound source along the y-axis in a three-dimensional coordinate system. In this system, the y-axis represents the direction from the left to the right of the vehicle.
[0231] The instantaneous velocity of the target sound source along the z-axis in a three-dimensional coordinate system. In this system, the z-axis runs from the bottom of the vehicle to its roof.
[0232] 5) The calculation process for the velocity part is as follows: 5.1) Obtain the visual position at time k-1. ; 5.2) Obtain the optimal state estimate of the Kalman filter output at time k-1. .
[0233] 5.3) The velocity unit is calculated based on the visual position at time k-1, the optimal state estimate at time k-1, and the preset position observation time, wherein the position observation time is the time required to observe each visual position each time. Specifically, in this embodiment, the visual position at time k-1 is subtracted from the optimal state estimate at time k-1 to obtain the position deviation, and the position deviation is divided by the position observation time to obtain the velocity unit.
[0234] When the visual position at time k-1 is equal to or almost identical to the optimal state estimate at time k-1, it indicates that the target person's head has not moved and is similar to being stationary. The stationary state belongs to the state of zero velocity in the uniform velocity model.
[0235] When the visual position at time k-1 is greater than the optimal state estimate at time k-1, that is, the observed visual position is to the right of the predicted position, the Kalman filter will infer: "It seems that I underestimated your speed, you are actually moving faster." So the Kalman filter will correct the speed estimate upward.
[0236] When the visual position at time k-1 is less than the optimal state estimate at time k-1, that is, the observed visual position is to the left of the predicted position, the Kalman filter will infer: "It seems that I overestimated your speed, or you are decelerating / turning", and then the Kalman filter will correct the speed estimate downward.
[0237] 6) Explanation of the process noise covariance matrix.
[0238] As mentioned earlier, the expression for the process noise covariance matrix is:
[0239] The parameters of the above expressions are described in Table 1: Table 1
[0240] As shown in Table 1, the embodiments of this application assume that the change in position is caused entirely by velocity. Therefore, the position noise is set to (0,0,0) in the process noise covariance matrix.
[0241] Speed noise is Velocity noise is a discretization of a continuous-time white noise acceleration model.
[0242] q represents the energy spectral density of the acceleration white noise, a fixed parameter that needs to be tuned, quantifying the degree of distrust in the uniform velocity model. In this application, different energy spectral densities can be assigned based on the acceleration of the velocity component. The energy spectral density q is positively correlated with the acceleration. A larger q value indicates a greater acceleration (velocity change) of the target person's head, causing the Kalman filter to readily accept new observations, resulting in more agile tracking but potentially more jitter. Conversely, a smaller q value indicates a smaller acceleration (velocity change) of the target person's head, causing the Kalman filter to place more trust in the model's predictions, resulting in smoother tracking but potentially increased response delay.
[0243] 7) The expression for the state transition model is:
[0244] Let k be the state vector at time k. Here is the state transition matrix. Let be the state vector at time k-1. This is process noise.
[0245] State transition matrix Defined by the system model, and calculated based on a preset physical model (such as a uniform velocity model) and the current time interval dt, it is a known and fixed matrix. State transition matrix. Includes position submatrices and velocity submatrices, and state transition matrix. The expression is:
[0246] In the state transition matrix In the diagram, the first three rows are partial matrix positions, as shown below: Line 1: ; Line 2: ; Line 3: ; From the partial matrices in the first three rows, we can obtain:
[0247] The above formula is the standard uniform kinematics formula, used to update the position.
[0248] In the state transition matrix In the diagram, the last three rows are partial velocity submatrices, as shown below: Line 4: ; Line 5: ; Line 6: ; From the partial matrices of the preceding and following rows, we can obtain:
[0249] As can be seen from the above formula, the embodiments of this application assume that the speed remains constant, and the speed is compensated and updated by the process noise w.
[0250] As can be seen from the above formulas, the embodiments of this application configure a position submatrix and a velocity submatrix for the state transition matrix. Based on the state transition matrix configured in this way, a state vector that simultaneously contains position and velocity relationships can be obtained.
[0251] The Kalman filter is configured with a first state prediction model and a first state update model. The first state prediction model is used to estimate the state of the target person's head when it moves at a constant speed in a preset three-dimensional coordinate system. The first state update model is used to output the sound source position of the final output of the Kalman filter.
[0252] Based on a preset Kalman filter, the visual covariance matrix and visual position are processed by Kalman filtering to obtain the candidate sound source positions, including the following steps: Step S241: Determine the predicted state value at the current moment based on the state prediction model.
[0253] Step S242: Determine the first Kalman gain based on the visual covariance matrix.
[0254] Step S243: Determine the candidate sound source location based on the first state update model, the first Kalman gain, and the visual position.
[0255] The current state prediction value is the prior state of the Kalman filter when predicting the head of the target person. Determining the current state prediction value based on the first state prediction model includes the following steps: obtaining the optimal state estimate and state transition matrix of the previous time step, and substituting the optimal state estimate and state transition matrix of the previous time step into the first state prediction model to obtain the current state prediction value.
[0256] The expression for the first-state prediction model is:
[0257] The predicted state value at the current moment. Here is the state transition matrix. This is the optimal state estimate from the previous time step. The predicted state value at the current time step. The state transition matrix is a 6x1 matrix. It is a 6x6 matrix, and the optimal state estimate of the previous time step is a 6x1 matrix.
[0258] Related technologies view the sound source location statically (e.g., simply treating the coordinates of the mouth as a fixed point), or they perform only instantaneous, memoryless localization of the sound source (searching independently in each frame of the captured image). This approach ignores the valuable information inherent in the sound source's motion itself. This approach has at least two drawbacks: (1) Lack of predictive ability: When the camera device fails to capture the image of the target person for a moment (such as when the speaker turns his head quickly and causes a brief loss of vision), the system cannot predict the current position of the target sound source based on the previous movement trend, which manifests as discontinuous tracking or sudden loss.
[0259] (2) Insufficient filtering capability: The system lacks a kinematic model-based "smoothing" mechanism to identify and filter out such unreliable interference noise due to instantaneous jumps or outliers in the visual data (such as sudden jitter of the visual detection box).
[0260] In this embodiment, the first state prediction model provided by this application estimates the state of the target person's head in the three-dimensional coordinate system in a uniform manner. It cleverly combines and utilizes the visual position at different time points contained in all captured images to construct a first state prediction model related to position and velocity. This transforms the Kalman filter from a smoother with only "memory function" into a tracker with "short-term prediction capability". Even if the camera does not capture the position of the target sound source for a moment or a short time, it can still predict the position of the target sound source at the next moment in advance based on the motion trend (i.e., velocity). There will be no break in the prediction of the sound source position, thereby significantly improving the continuity and smoothness of tracking.
[0261] The expression for the current state prediction value is: : .
[0262] As can be seen from the above formula, the state prediction value includes a position component and a velocity component. The Kalman filter estimates the coordinates of the target sound source of the target person in a preset three-dimensional coordinate system based on visual data. (Velocity part) The Kalman filter estimates the velocity of the target sound source along each coordinate axis in a three-dimensional coordinate system based on visual data.
[0263] Compared to traditional technologies that only configure position information in the state prediction value, the embodiments of this application configure both position and velocity components in the state prediction value. By introducing motion state (such as velocity component), the system evolves from "frame-by-frame discrete search" to "continuous trajectory tracking". This not only brings an unprecedentedly smooth sound source trajectory and completely eliminates positioning jumps and jitters, but also endows the system with short-term prediction capabilities, laying a solid foundation for high-end voice interaction.
[0264] The Kalman filter is configured with a first Kalman gain update model, which outputs the first Kalman gain corresponding to the visual position. Determining the first Kalman gain based on the visual covariance matrix includes the following steps: obtaining the first prediction covariance matrix and the first observation matrix at the current time; substituting the first prediction covariance matrix, the first observation matrix, and the visual covariance matrix at the current time into the first Kalman gain update model to obtain the first Kalman gain, wherein the first Kalman gain is negatively correlated with the visual covariance matrix.
[0265] The Kalman filter is configured with a first covariance prediction model, which is used to update the prediction covariance matrix. Obtaining the first prediction covariance matrix at the current time step includes the following steps: obtaining the state transition matrix, the first prediction covariance matrix at the previous time step, and the process noise covariance matrix. The process noise covariance matrix is determined by the velocity noise, and its magnitude is positively correlated with the velocity noise. Substituting the state transition matrix, the first prediction covariance matrix at the previous time step, and the process noise covariance matrix into the first covariance prediction model yields the first prediction covariance matrix at the current time step.
[0266] The expression for the first covariance prediction model is shown below:
[0267] This is the first prediction covariance matrix at the current time, used to represent the degree of uncertainty in the predicted state, and is a 6*6 matrix. Let be the state transition matrix, which is a 6x6 matrix. Let be the first prediction covariance matrix of the previous time step, which represents the uncertainty of the optimal state estimate at the previous time step. It is a 6x6 matrix. This is the transpose of the state transition matrix. In matrix operations, a transpose operation is required to ensure the symmetry and positive definiteness of the covariance matrix. Let be the process noise covariance matrix, which is a 6x6 matrix.
[0268] The Kalman filter is configured with a first observation model, the expression of which is as follows:
[0269] The first observation model establishes a bridge between the camera device's observations and the system's internal state, describing what kind of observations are expected to be obtained from the system state.
[0270] The physical meanings of each parameter in the first observation model are shown in Table 2: Table 2
[0271] The first observation matrix H1 is used to extract the 3-dimensional visual position from the 6-dimensional state vector X (containing position and velocity). In this embodiment, the physical quantity output by the camera device's visual data is the visual position, not the velocity. v_v is the visual observation noise, modeled as a Gaussian white noise with zero mean and a visual covariance matrix of Rv, representing the inherent, unavoidable error of the visual measurement itself. The observation model defines that "if state X is the true value, then the camera device should see H1*X, but due to the noise, it actually sees Pv."
[0272] The expression for the first Kalman gain update model is shown below:
[0273] The physical meaning of each parameter in the first Kalman gain update model is shown in Table 3: Table 3
[0274] As shown in Table 3, the expression for the first Kalman gain update model indicates a negative correlation between the first Kalman gain and the visual covariance matrix. When the visual covariance matrix... When the value is very small, it indicates that the determination of the visual position is relatively reliable. In the expression of the first Kalman gain update model, the denominator becomes smaller, and the first Kalman gain... An increase in the visual covariance matrix indicates that the Kalman filter places more trust in the observed visual location. When the value is very large, it indicates that the determination of the visual position is relatively unreliable. In the expression of the first Kalman gain update model, the denominator becomes larger, and the Kalman gain... The smaller value indicates that the Kalman filter trusts the observed visual location less and trusts the state predictions made by the model more.
[0275] This application's embodiments calculate the visual covariance matrix by integrating multiple types of visual noise penalty parameters. This ensures that the visual covariance matrix incorporates various uncertainties that can affect visual position, allowing all uncertainties to be quantified within the visual covariance matrix. The first Kalman gain is negatively correlated with the visual covariance matrix. The larger the sum of various uncertainties, the larger the magnitude of the visual covariance matrix, and the smaller the first Kalman gain, indicating that the Kalman model trusts the predicted state value more. Conversely, the smaller the sum of various uncertainties, the smaller the magnitude of the visual covariance matrix, and the larger the first Kalman gain, indicating that the Kalman model trusts the observed visual position more. Ultimately, this results in the Kalman filter outputting an accurate and reliable sound source position.
[0276] By combining the expressions of the first covariance prediction model and the first Kalman gain update model, it can be seen that the magnitude of the process noise covariance matrix is positively correlated with the velocity noise, and the velocity noise is positively correlated with the energy spectral density of the acceleration white noise. When the acceleration white noise is larger, the change in the target's head motion state is more drastic; the larger the velocity noise, the larger the magnitude of the process noise covariance matrix, the larger the prediction covariance matrix at the current moment, and the larger the first Kalman gain. This allows the Kalman filter to accept new observations more quickly, resulting in more agile tracking. Conversely, when the acceleration white noise is smaller, the change in the target's head motion state is slower; the smaller the velocity noise, the smaller the magnitude of the process noise covariance matrix, the smaller the prediction covariance matrix at the current moment, and the smaller the first Kalman gain. This allows the Kalman filter to accept the model's output state prediction more quickly.
[0277] This embodiment of the application ensures that the first state prediction model supports measuring the head state of the target person in a uniform manner, that is, without introducing process noise into the position of the state prediction value. However, considering that the actual sound source motion has unknown acceleration, this embodiment of the application designs the process noise covariance matrix as a matrix that is positively correlated with the velocity noise. The velocity noise cleverly simulates the cumulative effect of random acceleration (i.e., acceleration white noise) over time, thereby effectively following the influence of velocity noise on the first Kalman gain, and then reflecting it to the final output of the Kalman filter through the first Kalman gain, that is, reflecting it to the candidate sound source position. Therefore, this embodiment of the application still uses velocity noise to adjust the position without losing the ability of "continuous trajectory tracking", making the final output candidate sound source position more reliable and accurate.
[0278] Determining the candidate sound source location based on the first state update model, the first Kalman gain, and the visual position includes the following steps: obtaining the current state prediction value and the first observation matrix, and substituting the current state prediction value, the first observation matrix, the first Kalman gain, and the visual position into the first state update model to obtain the candidate sound source location.
[0279] The expression for the first-state update model is as follows:
[0280] The physical meaning of each parameter in the first-state update model is shown in Table 4: Table 4
[0281] As shown in Table 4, Called "news" or "observation residuals," it represents the difference between observed and predicted values. Kv* (news) is the weighted correction term. Kalman gain This not only determines the magnitude of the correction but also how to reasonably allocate the 3D positional differences across the 6-dimensional state (including position and velocity) corrections. For example, a persistent positional residual might be interpreted as an error in velocity estimation, thus correcting both velocity and position simultaneously, ultimately adding the weighted correction term to the current state prediction. By doing so, we can obtain the optimal state estimate for the current time step after the visual update. .
[0282] The optimal state estimate at the current moment This includes the position and velocity components at the current moment. The position component at the current moment represents the sound source location of the target sound source, which is ultimately output by the Kalman filter. Therefore, this embodiment estimates the optimal state value at the current moment. Extract the location of the target sound source.
[0283] The Kalman filter is configured with a first covariance update model. The method also includes substituting the first Kalman gain, the first observation matrix, and the first prediction covariance matrix at the current time into the first covariance update model to obtain the updated first prediction covariance matrix.
[0284] The expression for the first covariance update model is as follows:
[0285] The physical meaning of each parameter in the first covariance update model is shown in Table 5: Table 5
[0286] After incorporating new observational information, the uncertainty of the state estimate should decrease. This has helped to reduce uncertainty. After this correction, the optimal state estimate for the current moment is... The new uncertainty covariance matrix Covariance matrix of the uncertainty of the prediction Smaller.
[0287] The reliability information is used to indicate whether the audio location is reliable. This reliability information includes whether the audio location is reliable or unreliable; reliable information indicates that the audio location is reliable, while unreliable information indicates that the audio location is unreliable.
[0288] Generating credibility information about audio location based on the audio covariance matrix includes the following steps: generating credibility values based on the audio covariance matrix; generating reliable audio location information when the credibility value is less than a preset credibility threshold; and generating unreliable audio location information when the credibility value is greater than or equal to the preset credibility threshold.
[0289] The confidence level value is used to represent the reliability of the audio location. The confidence level value is negatively correlated with the reliability of the audio location; the higher the confidence level value, the less reliable the audio location, and vice versa.
[0290] The embodiments of this application can use various methods to generate confidence levels based on the audio covariance matrix.
[0291] In some embodiments, the audio covariance matrix is a diagonal covariance matrix, and the confidence level value includes the variance of the first audio diagonal element, the variance of the second audio diagonal element, and the variance of the first audio diagonal element. In this embodiment, the audio covariance matrix is analyzed, and the variances of the first audio diagonal element, the second audio diagonal element, and the third audio diagonal element are extracted from the audio covariance matrix.
[0292] Correspondingly, the preset reliable thresholds include a first value, a second value, and a third value. When the variance of the first audio diagonal element is less than the first value, the variance of the second audio diagonal element is less than the second value, and the variance of the third audio diagonal element is less than the third value, this reflects that the audio positioning is relatively concentrated. Therefore, the embodiments of this application generate reliable information on the audio position. When the variance of the first audio diagonal element is greater than the first value, or the variance of the second audio diagonal element is greater than the second value, or the variance of the third audio diagonal element is greater than the third value, unreliable information on the audio position is generated.
[0293] In other embodiments, the audio covariance matrix is a diagonal covariance matrix, and generating a confidence value based on the audio covariance matrix includes the following steps: calculating the trace or determinant of the audio covariance matrix, and setting the trace or determinant of the audio covariance matrix as the confidence value.
[0294] The trace of the audio covariance matrix is calculated by the following steps: Based on the audio covariance matrix, obtain the variances of the first audio diagonal elements, the second audio diagonal elements, and the first audio diagonal elements; add the variances of the first audio diagonal elements, the second audio diagonal elements, and the first audio diagonal elements together to obtain the trace of the audio covariance matrix.
[0295] In this embodiment, the determinant of the audio covariance matrix is calculated using the determinant calculation method. The smaller the trace or determinant of the audio covariance matrix, the higher the location confidence of the audio position. Conversely, the larger the trace or determinant of the audio covariance matrix, the lower the location confidence of the audio position.
[0296] In other embodiments, the confidence level value includes Mahalanobis distance. Generating the confidence level value based on the audio covariance matrix includes the following steps: calculating the Mahalanobis distance between the candidate sound source location and the audio location. When the Mahalanobis distance is large, the deviation between the candidate sound source location and the audio location is large; when the Mahalanobis distance is small, the candidate sound source location and the audio location are relatively consistent.
[0297] This application employs an uncertainty-driven dynamic fusion mechanism. It quantifies the reliability of visual and audio data in real time using visual and audio covariance matrices. Utilizing the reliability information of audio location, it intelligently determines when to "emphasize vision" or "rely on acoustics," thus solving the problem of related technologies being unable to determine a reliable location when visual and audio locations conflict, achieving a fundamental leap in robustness. Simultaneously, this application adopts a "visual location first, audio location later for correction" working mechanism. First, the visual location is anchored, providing a stable spatial anchor point for the susceptible acoustic system. Then, a Kalman filter is used to process the visual location to obtain candidate sound source locations. Finally, based on the reliability information of the audio location, the audio location is used to correct the candidate sound source locations, thereby obtaining an accurate and reliable sound source location.
[0298] Based on the credibility information, the candidate sound source location, audio location, and audio covariance matrix are fused to obtain the sound source location of the target sound source. The steps include: in response to the reliable information that the credibility information is the audio location, Kalman filtering is performed on the candidate sound source location, audio location, and audio covariance matrix to obtain the sound source location of the target sound source; in response to the unreliable information that the credibility information is the audio location, the candidate sound source location is set as the sound source location of the target sound source.
[0299] Even in the extreme case of temporary audio data failure, the embodiments of this application can still maintain a reliable sound source position, ensuring an absolutely consistent and stable user experience.
[0300] The Kalman filter is configured with a second observation model, the expression of which is as follows:
[0301] The second observation model establishes a bridge between the microphone array's observations and the system's internal state, describing what kind of observations are expected to be obtained from the system state.
[0302] The physical meanings of each parameter in the second observation model are shown in Table 6: Table 6
[0303] The Kalman filter is configured with a second state update equation. The Kalman filter is used to process the candidate sound source position, audio position and audio covariance matrix to obtain the sound source position of the target sound source. The steps include: determining the second Kalman gain based on the audio covariance matrix, obtaining the second observation matrix, and substituting the second Kalman gain, the second observation matrix, the audio position and the candidate sound source position into the second state update equation to obtain the sound source position of the target sound source.
[0304] The expression for the second-state update equation is:
[0305] The physical meaning of each parameter in the second-state update model is shown in Table 7: Table 7
[0306] As shown in Table 7, This is called "news" or "observation residual," used to represent the difference between the observed and predicted values. Ka* (news) is the weighted acoustic correction term. Finally, in this embodiment, the acoustic correction term is added to the candidate sound source position X_est_v to obtain the optimal state estimate that integrates all information from prediction, vision, and acoustics, i.e., the sound source position X_est of the target sound source.
[0307] As mentioned earlier, the first covariance update model can output the updated first prediction covariance matrix. The Kalman filter is also equipped with a second Kalman gain update model. Determining the second Kalman gain based on the audio covariance matrix includes the following steps: obtaining the updated first prediction covariance matrix and second observation matrix output by the first covariance prediction model; substituting the updated first prediction covariance matrix, second observation matrix, and audio covariance matrix into the second Kalman gain update model to obtain the second Kalman gain, wherein the second Kalman gain is negatively correlated with the audio covariance matrix.
[0308] As can be seen from the expression of the second state update equation, under the premise that the audio position is reliable, the embodiments of this application absorb the candidate sound source position obtained from video data and the audio position obtained from audio data, perform Kalman filtering, and use the second Kalman gain to weight the information between the audio position and the candidate sound source position obtained from visual data, thereby realizing the correction effect of the candidate sound source position by the audio position, thus obtaining a more accurate and reliable sound source position.
[0309] The expression for the second Kalman gain update model is shown below:
[0310] The physical meaning of each parameter in the second Kalman gain update model is shown in Table 8: Table 8
[0311] From the expression of the second Kalman gain update model, it can be seen that the second Kalman gain is negatively correlated with the audio covariance matrix. When the audio covariance matrix... When the value is very small, it indicates that the determination of the audio location is relatively reliable. In the expression for the second Kalman gain update model, the denominator becomes smaller, and the second Kalman gain... An increase in the audio covariance matrix indicates that the Kalman filter places greater trust in the observed audio location. When the value is very large, it indicates that the determination of the audio location is unreliable. In the expression for the second Kalman gain update model, the denominator becomes larger, and the second Kalman gain... The smaller value indicates that the Kalman filter trusts the observed audio location less and trusts the state predictions made by the model more.
[0312] As can be seen from the expression of the second Kalman gain update model, under the premise that the audio location is reliable, the embodiments of this application absorb the updated first prediction covariance matrix used to characterize the uncertainty of the candidate sound source location, and use the updated first prediction covariance matrix and the audio covariance matrix to jointly vote for the second Kalman gain. The second Kalman gain reflects the degree of trust in the audio location under the joint effect of the updated first prediction covariance matrix and the audio covariance matrix. In this way, the audio location can be effectively used to correct the candidate sound source location, thereby obtaining a more reliable and accurate sound source location.
[0313] The Kalman filter is equipped with a second covariance update module. The method also includes: substituting the second Kalman gain, the second observation matrix and the updated first prediction covariance matrix into the second covariance update model to obtain the second prediction covariance matrix.
[0314] The second prediction covariance matrix serves as the input to the first covariance prediction model of the Kalman filter. Simultaneously, the updated first prediction covariance matrix serves as the input to both the second Kalman gain update model and the second covariance update model.
[0315] The expression for the second covariance update model is as follows:
[0316] The physical meaning of each parameter in the second covariance update model is shown in Table 9: Table 9
[0317] After further incorporating audio data, the uncertainty of state estimation is further reduced on top of the visual update. The operator plays a role in reducing uncertainty. From the expression of the second covariance update model, it can be seen that after acoustic correction, the second prediction covariance matrix P_est of the target sound source's location X_est is updated. The second prediction covariance matrix P_est answers the question, "After successively absorbing visual and audio data, how confident is the current estimation result of the source location X_est?" The answer is: maximum confidence, minimum uncertainty. The second prediction covariance matrix P_est will serve as the initial uncertainty for the next iteration.
[0318] Step S25: Obtain a list of interfering sound sources.
[0319] The list of interfering sound sources includes the location of one or more interfering sound sources. Each interfering sound source is matched with a confidence level, which is used to represent the reliability of the location of the interfering sound source. The confidence level is positively correlated with the reliability of the location of the interfering sound source; that is, the higher the confidence level, the stronger the reliability of the location of the interfering sound source, and the lower the confidence level, the weaker the reliability of the location of the interfering sound source.
[0320] Obtaining a list of interfering sound sources involves the following steps: acquiring various types of sound source description data, determining the location of interfering sound sources based on the sound source description data, recording the location of all interfering sound sources, and obtaining a list of interfering sound sources.
[0321] Sound source description data is used to describe the situation of interfering sound sources. It is understandable that both visual and audio data are types of sound source description data. Various types of sound source description data include visual data, audio data, vehicle operation data, and inherent knowledge source data.
[0322] Vehicle operation data is used to describe the vehicle's status, including vehicle speed, window status, air conditioning status, etc. This application embodiment can acquire CAN bus data transmitted by the vehicle based on the CAN bus, and extract vehicle operation data from the CAN bus data. Vehicle operation data includes vehicle speed, engine speed, air conditioning operating level, window open / closed status, windshield wiper operating status, etc.
[0323] Inherent knowledge source data is used to describe the noise generated at fixed locations inside the vehicle. For example, inherent knowledge source data includes the location of air conditioning vents, the location of fans, or the location of objects that are prone to generating noise, etc.
[0324] Interference sources include actively predicted noise, passively sensed noise, vehicle noise, and stationary noise. Actively predicted noise is speech emitted by non-target individuals detected based on visual data. Passively sensed noise is speech emitted by non-target individuals detected based on audio data. Vehicle noise is noise generated by driving behavior on the vehicle. Stationary noise is noise inherent to the vehicle itself (such as air conditioning, car audio, and air vents) or noise generated by objects inside the vehicle (such as a hanging bell).
[0325] In some embodiments, determining the location of the interfering sound source based on the sound source description data includes the following steps: performing lip movement detection and location calculation operations based on visual data to obtain the visual reference location of the non-target person, and setting the visual reference location of the non-target person as the sound source location of the actively predicted noise.
[0326] Specifically, the visual data collected by the camera device is processed frame by frame, dividing the visual data into multiple visual segments. A target detection algorithm (such as YOLO) is used to perform person recognition on these visual segments, identifying target and non-target individuals. For example, the target individual is a pre-defined passenger whose speech needs enhancement, such as the driver, while the non-target individual is other unrelated passengers. Then, lip movement detection is performed on the visual segments corresponding to the non-target individuals to obtain lip movement features (such as the amplitude of lip opening and closing). Based on these features, it is determined whether the non-target individual is speaking. If so, a position calculation is performed on the speaking non-target individual to obtain their visual reference position. This visual reference position is used as the source location of the actively predicted noise. For example, if an unrelated passenger in the back seat is speaking, their position is set as the source location of the actively predicted noise.
[0327] In some embodiments, determining the location of the interfering sound source based on sound source description data includes the following steps: determining the audio reference location of a non-target person based on audio data, and setting the audio reference location of the non-target person as the sound source location of the passively perceived noise.
[0328] Specifically, the multi-channel audio data acquired by the microphone array is framed to obtain multiple audio segments. Fourier transform is performed on each audio segment to obtain frequency domain segments. Audio features are extracted based on these frequency domain segments. Based on these features, it is determined whether the audio segment belongs to an audio segment generated by a non-target person. If so, an audio localization algorithm consistent with the embodiments of this application is used. Based on the spatial features of the non-target person's audio segment (phase difference and amplitude difference of each microphone channel), the audio reference position of the non-target person is calculated. This non-target person's audio reference position is the sound source position of the passively perceived noise. For example, the whisper of a rear-seat passenger may not be captured by visual lip movement detection, but the rear-seat passenger is perceived to be speaking through audio data. The position of the rear-seat passenger is calculated based on the audio data and set as the sound source position of the passively perceived noise.
[0329] In some embodiments, determining the location of an interfering sound source based on sound source description data includes the following steps: determining the vehicle state based on vehicle operation data, generating an acoustic context description vector based on the vehicle state, and determining the location of the vehicle noise source based on the acoustic context description vector.
[0330] Specifically, the process begins by collecting vehicle operating data (such as vehicle speed, engine speed, air conditioning setting, and window open / closed status) to determine the vehicle's current state (such as high-speed driving, idling, air conditioning on, and windows open). Then, the vehicle operating data is normalized to generate an acoustic context description vector. This vector contains acoustic feature association information corresponding to the vehicle state (such as the relationship between vehicle speed and the intensity and location of tire noise and wind noise at high speeds). Finally, based on the acoustic context description vector and the positional parameters of various vehicle components, the location of the vehicle noise source is determined.
[0331] Please refer to Table 10: Table 10
[0332] As shown in Table 10, when the vehicle speed is greater than 100km / h, the vehicle operation data shows that the windows are open. The generated acoustic context description vector is associated with the wind noise and the window position, thereby determining the location of the vehicle noise source, which mainly includes the position of the front of the vehicle and the side windows. When the vehicle engine is idling, the acoustic context description vector is associated with the engine noise, and the engine position is determined as the location of the vehicle noise source.
[0333] In some embodiments, determining the location of an interfering noise source based on sound source description data includes the following steps: determining the location of a fixed noise source based on inherent knowledge source data.
[0334] Specifically, the inherent knowledge source data pre-stores the installation locations of all fixed devices inside the vehicle, such as the installation location of the car audio system. This application embodiment uses the inherent knowledge source data to obtain the installation locations of the fixed devices, which are the sound source locations of fixed noise. For example, the fixed noise generated by the car audio system is located at the installation location of the car audio system.
[0335] In this embodiment, the location of all interfering sound sources is recorded to obtain a list of interfering sound sources.
[0336] Please refer to Table 11: Table 11
[0337] As shown in Table 11, each interfering sound source corresponds to an information source attribute, which indicates the definite origin of the interfering sound source. Each interfering sound source also corresponds to a confidence level. It is understandable that in some application scenarios, there are no passengers inside the vehicle, only vehicle noise and stationary noise. Therefore, the list of interfering sound sources for some application scenarios does not include actively predicted noise and passively sensed noise. In some application scenarios, there are passengers inside the vehicle, but the vehicle speed is low or the air conditioning is not running; therefore, the list of interfering sound sources for some application scenarios does not include vehicle noise.
[0338] Step S26: Based on the list of interference sources, determine the interference sources that meet the preset confidence conditions as the sources to be suppressed.
[0339] The pre-set confidence conditions are customized by the designer based on engineering experience. The noise reduction algorithm provided in this application is the LCMV algorithm. In the LCMV framework, the constraints are mandatory, and the confidence level is a quantitative representation of the "constraint force". When the interfering sound source is added to the dynamic constraint matrix and the target response vector is set (usually set to 0 for the interfering sound source, i.e., "null"), the LCMV algorithm will, at any cost, generate a deep null in the direction that the main beam points to the audio signal of the target person. This approach may consume the system's degrees of freedom and may cause distortion of the target task's audio signal due to model mismatch (such as direction estimation error). Therefore, the confidence level directly determines whether to impose this "mandatory constraint" on a perceived interfering sound source, and the strictness of the constraint.
[0340] Understandably, a typical system would not simply set all the interference sources in the interference source list to forced null, nor would it only process the highest priority one. Instead, it would implement a hierarchical, hybrid constraint and optimization objective setting based on the confidence level.
[0341] This application constructs a dynamic constraint matrix and adds the directions of "highest confidence" and "high confidence" interference sources to the LCMV's dynamic constraint matrix to impose forced nulls on these interference sources. For interference sources of "medium confidence" and "basic confidence," their locations may not be precise enough, or their characteristics may partially resemble the target speaker's speech. If "medium confidence" and "basic confidence" interference sources are set as forced constraints, the forced nulls may be applied incorrectly if the direction estimation is flawed, potentially distorting or suppressing the target speech. The system does not directly add the directions of these medium / low confidence interference sources to the dynamic constraint matrix. Instead, it incorporates these sources into the calculation of the data covariance matrix to influence the core cost function of LCMV. The aim is to ensure that the LCMV algorithm, when minimizing output power, "sees" strong "interference" in these artificially weighted directions, thus adaptively, rather than forcibly, aligning the beammap nulls with these directions.
[0342] The system's degrees of freedom are limited. When there are many high-priority interference sources, there may not be enough array element degrees of freedom to form good adaptive nulls for all medium / low confidence interference sources. In this case, the confidence level of each interference source in the "interference source list" will guide the focus of the data covariance matrix correction, ensuring that system resources are prioritized to deal with the most threatening interference sources at the top of the list.
[0343] In extreme cases (such as when the number of interfering sources is close to or exceeds the system's degrees of freedom), the system may retain only 1-2 interfering sources with the highest confidence in the dynamic constraint matrix, while using all other interfering sources as input to the data covariance matrix for adaptive processing, in order to avoid system crashes or severe performance degradation.
[0344] The steps for determining which interference sources meet the preset confidence conditions as sources to be suppressed based on the list of interference sources include: obtaining the information source attributes of the interference sources; determining the basic weights of the interference sources based on the information source attributes; determining the weight correction coefficients based on the source description data of the interference sources; determining the confidence level of the interference sources based on the basic weights and the weight correction coefficients; and determining which interference sources meet the preset confidence conditions as sources to be suppressed based on the confidence levels of each interference source.
[0345] This application embodiment configures corresponding information source attributes for each interfering sound source based on the determined source of the interfering sound source. These information source attributes include visual attributes, audio attributes, vehicle status attributes, and known noise attributes. Visual attributes indicate that the interfering sound source originates from video data; that is, the information source attribute for actively predicted noise is a visual attribute. Audio attributes indicate that the interfering sound source originates from audio data; that is, the information source attribute for passively perceived noise is an audio attribute. Vehicle status attributes indicate that the interfering sound source originates from vehicle operation data; that is, the information source attribute for vehicle noise is a vehicle status attribute. Known noise attributes indicate that the interfering sound source originates from inherent knowledge source data; that is, the information source attribute for fixed noise is a known noise attribute.
[0346] Determining the basic weight of an interfering sound source based on its information source attributes includes the following steps: Responding to the interfering sound source's information source attribute being a visual attribute, a first weight is determined as the basic weight for actively predicting noise; responding to the interfering sound source's information source attribute being an audio attribute, a second weight is determined as the basic weight for passively sensing noise; responding to the interfering sound source's information source attribute being a vehicle state attribute, a third weight is determined as the basic weight for vehicle noise; and responding to the interfering sound source's information source attribute being a known noise attribute, a fourth weight is determined as the basic weight for fixed noise.
[0347] The first, second, third, and fourth weights can be customized by the designer based on engineering experience or business needs.
[0348] The method of "detecting the visual reference position of a non-target person based on visual data" has high reliability. Therefore, in this embodiment of the application, a first weight is assigned to the visual attribute. For example, the first weight is 0.9.
[0349] The method of "detecting the audio reference position of non-target persons based on audio data" is susceptible to interference from other noise. Compared with the visual method, the reliability of the method is moderate, and the interference of non-target persons on the speech of target persons is relatively serious. Therefore, the value of the second weight needs to be set higher. Thus, in this embodiment, the second weight is assigned to the audio attribute. For example, the second weight is 0.8.
[0350] The impact of "vehicle noise" on the target person's voice is minimal, or some features of the "vehicle noise" are similar to some features of the target person's voice. Therefore, there is no need to excessively suppress this type of noise to avoid affecting the suppression of the target person's voice. Furthermore, vehicle noise typically occurs for a short period in practical applications, and it is desirable to prioritize processing the speech of non-target persons. Therefore, compared to the second weight, this embodiment sets the third weight to be smaller than the second weight; for example, the third weight is 0.75. This embodiment assigns the third weight to the vehicle state attribute.
[0351] The inherent knowledge source data consists of preset fixed parameters, and fixed noise has the highest reliability. In this embodiment, a fourth weight is assigned to the known noise attribute. For example, the fourth weight is 0.7.
[0352] In some embodiments, the weight correction coefficient includes a first correction coefficient. Determining the weight correction coefficient based on the source description data of the interfering sound source includes the following steps: In response to the information source attribute of the interfering sound source being a visual attribute, the visual data is processed by frame segmentation to obtain multiple visual segments. Among the multiple visual segments, the visual segment containing the non-target person is identified as the non-target visual segment. The lip movement features of the non-target person are extracted from the non-target visual segment. The lip movement amplitude is determined based on the lip movement features. In response to the lip movement amplitude being greater than a preset amplitude threshold, it is determined that the non-target person is in a speaking state. The total speaking time of the non-target person within the sampling period is recorded. The total speaking time of the non-target person is normalized to obtain the first correction coefficient. The speaking time of the non-target person is positively correlated with the first correction coefficient. The longer the speaking time of the non-target person, the closer the first correction coefficient is to the natural number 1.
[0353] In some embodiments, the weight correction coefficient includes a second correction coefficient. Determining the weight correction coefficient based on the source description data of the interfering sound source includes the following steps: in response to the information source attribute of the interfering sound source being an audio attribute, the audio data is processed by frame segmentation to obtain multiple audio segments; among the multiple audio segments, audio segments containing non-target individuals are identified as non-target audio segments; the signal-to-noise ratio (SNR) of the non-target audio segments is calculated; the SNR of all non-target audio segments within the sampling period is calculated to obtain the total SNR; the total SNR is normalized to obtain the second correction coefficient, wherein the total SNR and the second correction coefficient are positively correlated. The larger the total SNR, the closer the second correction coefficient is to the natural number 1.
[0354] In some embodiments, the weight correction coefficient includes a third correction coefficient. Determining the weight correction coefficient based on the source description data of the interfering sound source includes the following steps: Responding to the information source attribute of the interfering sound source being a vehicle state attribute, determining the stable operating coefficient of the vehicle within the sampling period based on vehicle operation data, and normalizing the stable operating coefficient to obtain the third correction coefficient. The stable operating coefficient and the third correction coefficient are positively correlated. The larger the stable operating coefficient, the closer the third correction coefficient is to the natural number 1.
[0355] Determining the stable operating coefficient of a vehicle within a sampling period based on vehicle operation data includes the following steps: determining the vehicle speed stability coefficient, air conditioning gear stability coefficient, and window status stability coefficient based on vehicle operation data; weighting the vehicle speed stability coefficient, air conditioning gear stability coefficient, and window status stability coefficient to obtain the vehicle's stable operating coefficient.
[0356] Determining the vehicle speed stability coefficient based on vehicle operation data includes the following steps: extracting the vehicle speed at each moment within the sampling period from the vehicle operation data; determining the vehicle speed standard deviation and average vehicle speed based on the vehicle speed at all moments; calculating the ratio of the vehicle speed standard deviation to the average vehicle speed; and subtracting this ratio from the natural number 1 to obtain the vehicle speed stability coefficient.
[0357] Determining the air conditioning speed stability coefficient based on vehicle operation data includes the following steps: determining the number of times the user switches the air conditioning speed within the sampling period; calculating the product of the number of speed switches and a first value to obtain the first product result; and subtracting the first product result from the natural number 1 to obtain the air conditioning speed stability coefficient. For example, the first value is 0.2. When the first product result is greater than 1, the air conditioning speed stability coefficient is set to 0.
[0358] Determining the window stability coefficient based on vehicle operation data includes the following steps: determining the number of times the user changes the window status within the sampling period; calculating the product of the number of status changes and a second value to obtain the second product result; and subtracting the second product result from the natural number 1 to obtain the window stability coefficient. For example, the second value is 0.3. When the second product result is greater than 1, the window stability coefficient is set to 0.
[0359] It is understandable that when the information source attribute of the interfering sound source is a known noise attribute, the weight correction coefficient of the fixed noise is the third correction coefficient, that is, the weight correction coefficient of the fixed noise is the same as the weight correction coefficient of the vehicle noise.
[0360] In some embodiments, the confidence level of the interfering sound source includes a first confidence level. Determining the confidence level of the interfering sound source based on the base weight and the weight correction coefficient includes the following steps: in response to the information source attribute of the interfering sound source being a visual attribute, the base weight of the actively predicted noise is multiplied by the first correction coefficient to obtain the first confidence level.
[0361] In some embodiments, the confidence level of the interfering sound source includes a second confidence level. Determining the confidence level of the interfering sound source based on the base weight and the weight correction coefficient includes the following steps: in response to the information source attribute of the interfering sound source being an audio attribute, the base weight of the passively perceived noise is multiplied by the second correction coefficient to obtain the second confidence level.
[0362] In some embodiments, the confidence level of the interfering sound source includes a third confidence level. Determining the confidence level of the interfering sound source based on the base weight and the weight correction coefficient includes the following steps: in response to the information source attribute of the interfering sound source being a vehicle state attribute, the base weight of the vehicle noise is multiplied by the third correction coefficient to obtain the third confidence level.
[0363] In some embodiments, the confidence level of the interfering sound source includes a fourth confidence level. Determining the confidence level of the interfering sound source based on the base weight and the weight correction coefficient includes the following steps: in response to the information source attribute of the interfering sound source being a known noise attribute, the base weight of the fixed noise is multiplied by the third correction coefficient to obtain the fourth confidence level.
[0364] For example, embodiments of this application obtain the following weight correction coefficients and basic weights based on various sound source description data, as shown in Table 12: Table 12
[0365] As shown in Table 12, based on the above-mentioned approach, the confidence levels of various interference noise sources can be obtained in the embodiments of this application. Among them, the first confidence level of actively predicted noise is 0.81, the second confidence level of passively perceived noise is 0.72, the third confidence level of vehicle noise is 0.525, and the fourth confidence level of stationary noise is 0.56.
[0366] Determining the interference sources that meet the preset confidence conditions as sources to be suppressed based on the confidence levels of each interference source includes the following steps: in response to the interference source having a confidence level greater than the preset confidence threshold, determining that the interference source meets the preset confidence conditions and setting the interference source as a source to be suppressed; in response to the interference source having a confidence level less than or equal to the preset confidence threshold, not setting the interference source as a source to be suppressed.
[0367] For example, the preset confidence threshold is 0.6. As shown in Table 12, the first confidence level of actively predicted noise and the second confidence level of passively sensed noise are both greater than 0.6. Therefore, in this embodiment, actively predicted noise and passively sensed noise are set as sound sources to be suppressed. The third confidence level of vehicle noise and the fourth confidence level of stationary noise are both less than the preset confidence threshold. Therefore, in this embodiment, vehicle noise and stationary noise are not set as sound sources to be suppressed. However, in this embodiment, vehicle noise and stationary noise are included in the calculation process of the data covariance matrix, thereby indirectly suppressing vehicle noise and stationary noise.
[0368] This application's embodiments combine the determination of the source of interference sound (information source attributes) and its own characteristics, significantly improving the accuracy and specificity of confidence. At the same time, by screening the sound sources to be suppressed based on confidence conditions, it effectively avoids the problems of false suppression and missed suppression, further optimizing the effect of directional noise reduction. This ensures that while suppressing interference sound sources, the integrity of the target sound source is protected to the greatest extent, adapting to diverse and dynamically changing interference scenarios inside the vehicle, and improving the practicality and reliability of the entire noise reduction method.
[0369] Step S27: Based on a preset noise reduction algorithm, noise reduction processing is performed on the audio data using the sound source location of the target sound source and the sound source location of the sound source to be suppressed, so that the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced.
[0370] The noise reduction algorithm provided in this application embodiment can be the LCMV algorithm or other types of noise reduction algorithms.
[0371] This application's embodiments introduce the collaborative operation of visual and audio data, utilizing both to capture the location of the target sound source, achieving precise localization and providing effective data support for subsequent precise noise reduction. Secondly, this application's embodiments construct an interference source list, recording the location and confidence level of all interfering sound sources, providing precise interference information support for subsequent targeted noise reduction. This ensures that noise reduction is no longer blind but can specifically focus on interfering sound sources, achieving the core objective of "suppressing interference and protecting the target," significantly improving the accuracy, targeting, and effectiveness of noise reduction.
[0372] Based on a preset noise reduction algorithm, the audio data is denoised using the location of the target sound source and the location of the sound source to be suppressed, so that the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced. This includes the following steps: Step S271: Based on visual data, filter out multiple invalid audio segments from the audio data.
[0373] The audio data includes multiple temporal audio segments. Invalid audio segments are those that do not occur during the target person's speech. For example, invalid audio segments may include: lip movements not performed by the target person, non-vocal lip movements of the target person (chewing / coughing), and any non-vocal lip movements. Valid audio segments are those that occur during the target person's speech. For example, valid audio segments may include: the target person opening their mouth to speak, and the target person's vocal lip movements.
[0374] The process of filtering out multiple invalid audio segments from audio data based on visual data includes the following steps: performing frame segmentation on the visual data and audio data respectively to obtain multiple visual segments and multiple temporal audio segments; aligning one visual segment with one temporal audio segment; and filtering out multiple invalid audio segments from among the multiple temporal audio segments based on the alignment relationship between the visual segments and the temporal audio segments.
[0375] For example, after visual data is framed, the following visual segments are obtained: V1 (0-33ms), V2 (33-66ms), V3 (66-99ms)... After audio data is framed, the following temporal audio segments are obtained: A1 (0-33ms), A2 (33-66ms), A3 (66-99ms)..., where V1 is aligned with A1, V2 with A2, and so on. This embodiment utilizes the alignment relationship between visual segments and temporal audio segments to analyze the state of the target person in the visual segment (whether they are speaking) to determine whether the corresponding temporal audio segment is invalid: if the visual segment shows that the target person is not speaking, then the corresponding aligned temporal audio segment is invalid; if the visual segment shows that the target person is speaking, then the corresponding aligned temporal audio segment is valid.
[0376] Based on the alignment relationship between visual segments and temporal audio segments, the process of filtering out multiple invalid audio segments from multiple temporal audio segments includes the following steps: performing a person recognition operation on each visual segment to obtain the identity information of the visual segment; in response to the identity information of the visual segment being the identity information of the target person, performing a lip movement detection operation on the visual segment to obtain the lip movement detection information of the visual segment; in response to the lip movement detection information of the visual segment being the speech activity information; setting the audio segment aligned with the visual segment as a valid audio segment; in response to the completion of person recognition and lip movement detection operations on all visual segments, removing all valid audio segments from all visual segments to obtain a set of remaining segments; and setting each visual segment in the set of remaining segments as an invalid audio segment.
[0377] This application embodiment identifies a person region in a visual segment based on an object detection algorithm, processes the person region based on a face recognition algorithm, and obtains the identity information corresponding to the visual segment. The object detection algorithm can be the YOLO algorithm, etc.
[0378] Lip movement detection information includes speech activity information and non-speech activity information. Performing lip movement detection on a visual segment to obtain lip movement detection information for the visual segment includes the following steps: extracting the lip movement features of the target person from the visual segment, determining the lip movement amplitude based on the lip movement features, determining that the target person is in a speaking state when the lip movement amplitude is greater than a preset amplitude threshold, and generating speech activity information; and generating non-speech activity information when the lip movement amplitude is less than or equal to the preset amplitude threshold.
[0379] This application's embodiments accurately filter out invalid audio segments, ensuring that the invalid audio segments are pure noise segments, avoiding the mixing of the target person's voice signal, and providing a guarantee for the accuracy of the subsequent data covariance matrix.
[0380] Step S272: Generate a data covariance matrix for the noise space based on multiple invalid audio segments.
[0381] In this embodiment, invalid audio segments are processed by short-time Fourier transform to obtain invalid frequency domain segments. A periodogram is calculated based on the invalid frequency domain segments, and an arithmetic mean is performed on the periodograms of all invalid frequency domain segments to obtain the data covariance matrix of the noise space where the microphone array is located.
[0382] Invalid audio segments are time-domain signals and need to be converted into invalid frequency-domain segments X(f, t) using a short-time Fourier transform. These invalid frequency-domain segments contain the frequency distribution information of the noise, where t represents the time frame index. Essentially, the invalid frequency-domain segment X(f, t) corresponds to the frequency domain segment of the speech signal of the non-target person. This ensures that the invalid audio segments do not contain the speech signal of the target person, thus preventing speech signal contamination at the source.
[0383] In this embodiment of the application, for all invalid frequency domain segments (assuming a total of N frames), at each frequency point f, the following calculation is performed: ① For each frame of invalid frequency domain segment Calculate the periodogram, which is the outer product of the invalid frequency domain segment of the frame and its own conjugate transpose: , For the outer product, Invalid frequency domain segment The conjugate transpose of .
[0384] ② The periodograms of all invalid frequency domain segments are arithmetically averaged to obtain the data covariance matrix:
[0385] in, This is the data covariance matrix.
[0386] Data covariance matrix The data covariance matrix quantitatively describes the statistical properties in the noise space. The diagonal elements represent the noise power of each element in the microphone array, while the off-diagonal elements represent the spatial correlation and direction-of-arrival information of the noise between different elements. The beamformer of the LCMV algorithm utilizes the data covariance matrix. The optimal weights that can most effectively suppress noise with this spatial characteristic are calculated.
[0387] Related technologies rely on acoustic VADs (such as those based on energy or spectral entropy) to distinguish between valid and invalid audio segments. However, these methods are prone to misjudgment in low signal-to-noise ratio or non-stationary noise environments, leading to leakage of the target speaker's voice signal into the noise estimation (i.e., "voice smearing"). This causes beamformers to inadvertently suppress parts of the target speaker's voice while enhancing it. This application's embodiments introduce visual VADs as the basis for noise segment selection, offering the following advantages: 1) High reliability: Visual data is unaffected by the acoustic environment, ensuring highly reliable determination of whether a speaker is vocalizing. 2) Clean noise estimation: Ensuring the accuracy of noise estimation for calculating the data covariance matrix. Invalid audio segments do not contain the target person's speech signal, thus obtaining a cleaner and more accurate noise spatial characteristic estimate. 3) Performance improvement: The beamforming weights calculated based on this clean estimate can more effectively suppress real noise, while avoiding distortion of the target person's speech, significantly improving the signal-to-noise ratio and fidelity of the output speech.
[0388] Step S273: Based on the noise reduction algorithm, the sound source positions of the target sound source and the sound source positions of the sound source to be suppressed are processed by array manifold to obtain a dynamic constraint matrix.
[0389] The microphone array includes multiple array elements. Each interfering sound source in the list of interfering sound sources is configured with a confidence level. The array manifold processing based on the noise reduction algorithm to obtain the dynamic constraint matrix includes the following steps: determining the target steering vector from the target sound source to all array elements based on the positions of the array elements and the sound source positions of the sound sources to be suppressed; determining the interference steering vector from the sound source to be suppressed to all array elements based on the positions of the array elements and the sound source positions of the sound sources to be suppressed; and determining the dynamic constraint matrix based on the target steering vector and all interference steering vectors.
[0390] The target steering vector is used to characterize the signal propagation characteristics from the target sound source to each element in the microphone array. In this embodiment, based on the physical model of sound wave propagation, the theoretical sound wave transfer function from the target position Pf to each element of the microphone array is calculated, and the theoretical sound wave transfer function is normalized to obtain the target steering vector.
[0391] Specifically, in this embodiment, based on the location of the target sound source and the location of each array element, the propagation distance from the target sound source to each array element is calculated. Then, based on the propagation distance, sound speed, and frequency of the frequency domain audio segment, the target steering vector from the target sound source to each array element is obtained, as shown below:
[0392]
[0393] in, Location of the sound source Let be the position of the i-th array element, c be the speed of sound, and f be the frequency. The target guiding vector, Let j be the propagation distance from the target sound source to the i-th array element, j be the imaginary unit, and N be the total number of array elements.
[0394] Similarly, in this embodiment, based on the location of the sound source to be suppressed and the location of each array element, the propagation distance from the sound source to be suppressed to each array element is calculated. Then, based on the propagation distance, sound speed, and frequency of the frequency domain audio segment, the interference steering vector from the sound source to be suppressed to each array element is obtained, as shown below:
[0395]
[0396] in, Let K be the location of the k-th sound source to be suppressed. Let be the position of the i-th array element, c be the speed of sound, and f be the frequency. For the k-th interference steering vector, Let j be the propagation distance from the k-th sound source to the i-th array element, where j is the imaginary unit and N is the total number of array elements.
[0397] The target steering vector describes the relative phase difference and amplitude attenuation at each array element caused by the difference in propagation path when the sound wave travels from the source location of the target sound source to each array element. The interference steering vector describes the relative phase difference and amplitude attenuation at each array element caused by the difference in propagation path when the sound wave travels from the source location of the sound source to be suppressed to each array element.
[0398] This application embodiment determines the dynamic constraint matrix based on the target guidance vector and all interfering guidance vectors. Specifically, the dynamic constraint matrix... Where M is the total number of sound sources to be suppressed. This is the dynamic constraint matrix.
[0399] Unlike traditional LCMV which uses constraints in a fixed direction, the dynamic constraint matrix in this embodiment... It is dynamically updated in real time, constraining the dynamic constraint matrix. The factors depend on the real-time output of the target sound source location from the EKF fusion module and the sound source location of the sound source to be suppressed provided by the vision system. This means that the following two points can be achieved: (1) Beam main lobe dynamic tracking: When the head of the target person moves, the "auditory focus" of the beamformer can be automatically adjusted to always maintain the best receiving direction.
[0400] (2) Dynamic tracking of interference nulls: When the interference source (such as other passengers) moves, the system can dynamically generate suppression nulls in the new interference direction to achieve "precision strike".
[0401] This mechanism fundamentally solves the problem of the sharp performance degradation of traditional beamforming in dynamic scenarios, and realizes the leap from "static filtering" to "dynamic tracking filtering".
[0402] Step S274: Based on the noise reduction algorithm, the data covariance matrix and the preset target response vector are weighted to obtain the target weight matrix.
[0403] The target response vector is customized by the designer according to business requirements. It is used to specify the expected gain in each constraint direction. For example, the gain is set to 1 in the direction of the target sound source's location, and the gain is set to 0 in the direction of the sound source's location to be suppressed. Therefore, the target response vector is set to... , Let M be the target response vector, and let M be the dimension of (M+1)*1.
[0404] In the target response vector, the first element "1" corresponds to the dynamic constraint matrix. The first column, i.e., the target guidance vector This indicates that the system's gain in the direction of the target sound source must be 1, meaning it must receive the target person's audio without distortion. All subsequent elements "0" correspond to the dynamic constraint matrix. Each column of interference steering vector This means that the system must have a gain of 0 for all directions of interfering sound sources, that is, completely suppress audio from all directions of interfering sound sources.
[0405] The process of calculating the weights of the data covariance matrix and the preset target response vector based on the noise reduction algorithm to obtain the target weight matrix includes the following steps: constructing an optimization problem based on the weight matrix to be solved and the data covariance matrix; constructing constraints based on the weight matrix to be solved and the preset target response vector; and solving the optimization problem based on the constraints to obtain the target weight matrix.
[0406] The expression for the optimization problem is: .
[0407] The expression for the constraint is: .
[0408] in, Let be the weight matrix to be solved. Let be the conjugate transpose of the weight matrix to be solved. For the data covariance matrix, For dynamic constraint matrices, It is the conjugate transpose of the dynamic constraint matrix. This is the target response vector.
[0409] The embodiments of this application involve a simultaneous optimization problem and constraints, including:
[0410] The above problem is a quadratic optimization problem with linear constraints, and its closed-form solution can be obtained using the Lagrange multiplier method. The objective weight matrix is... The calculation formula is:
[0411] Target weight matrix It can achieve the dual requirements of "minimizing noise power and meeting the constraints of target enhancement and interference suppression", providing core parameters for subsequent audio noise reduction and enhancement processing.
[0412] Step S275: Based on the target weight matrix and audio data, perform suppression processing on the audio of the sound source to be suppressed and enhancement processing on the audio of the target sound source.
[0413] In this embodiment, a short-time Fourier transform is performed on a time-domain audio segment to obtain a frequency-domain audio segment. Based on the target weight matrix and the frequency-domain audio segment, a denoised frequency-domain audio segment is determined. An inverse short-time Fourier transform is performed on the denoised frequency-domain audio segment to obtain a denoised time-domain audio segment. All denoised time-domain audio segments are combined to obtain a denoised speech signal.
[0414] At each frequency f, embodiments of this application will target the weight matrix. conjugate transpose With frequency domain audio segments Perform weighted operations to obtain the denoised frequency domain audio segment. The details are as follows: This results in the signals from the direction of the target sound source being superimposed in phase and thus enhanced, while the signals from the direction of the interfering sound source being canceled out in phase and thus suppressed.
[0415] To obtain a playable and post-processable time-domain speech signal, the denoised frequency-domain audio segment... Perform an inverse short-time Fourier transform to obtain the denoised time-domain audio segment.
[0416] In this embodiment, all denoised time-domain audio segments are sequentially combined in chronological order to form a complete time-domain speech signal, which is the final denoised speech signal. In the final denoised speech signal, the audio of the target sound source is effectively enhanced, the audio of the sound source to be suppressed is significantly suppressed, and there is no obvious distortion, which can meet the needs of in-vehicle target speech recognition, communication, etc.
[0417] The spatial noise reduction module provided in this application embodiment is not an independent traditional beamformer, but a module deeply coupled with the upstream visual positioning and EKF dynamic fusion module. It makes full use of the advantages of multimodal information and realizes continuous, accurate and high-fidelity extraction of the target person's audio in a complex and dynamic in-vehicle environment.
[0418] In summary, the embodiments of this application can bring the following quantifiable technical effects: 1. Improved positioning accuracy: Through EKF fusion, the root mean square error of the target sound source is expected to be reduced by 30%-50% compared to the visual or audio position alone, with the advantage being more obvious in noisy and reverberant environments.
[0419] 2. Enhanced System Robustness: The system no longer has a single point of failure. When vision is briefly obstructed, the system automatically increases the visual covariance matrix, becoming more dependent on audio location; conversely, in noisy environments, the system automatically increases the audio covariance matrix, becoming more dependent on visual location, achieving a fault-tolerant effect of "1+1>2".
[0420] 3. Significantly improved voice quality: Based on high-precision target sound source location and LCMV beamforming, the signal-to-noise ratio of the target person's audio can be improved by more than 15dB, and the audio clarity is significantly improved. In multi-person conversation scenarios, it can effectively suppress noise greater than 10dB.
[0421] 4. Enhanced user experience: Achieves an intelligent experience that is "ready to use and always clear," solving the pain points of voice communication for users in complex scenarios such as high speed, open windows, and conversations among rear passengers.
[0422] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0423] See Figure 3 , Figure 3 This is a schematic diagram of a controller provided in an embodiment of this application. The controller 13 includes one or more processors 131 and a memory 132. The memory 132 is connected to one or more processors 131, for example, via a bus.
[0424] Processor 131 is configured to support the controller in performing the corresponding functions in the methods described in the above method embodiments. The processor may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0425] Memory 132 is used to store program code, etc. Memory may include volatile memory (VM), such as random access memory (RAM); memory may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory may also include combinations of the above types of memory.
[0426] The memory 132 can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the multimodal in-vehicle noise reduction method in the embodiments of this application. The processor executes various functional applications and data processing of the multimodal in-vehicle noise reduction method by running the non-volatile software programs, instructions, and modules stored in the memory, thereby realizing the functions of each module or unit of the multimodal in-vehicle noise reduction method provided in the above method embodiments.
[0427] The memory 132 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function. The data storage area may store data created based on the use of the multimodal in-vehicle noise reduction method, etc.
[0428] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the multimodal in-vehicle noise reduction method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.
[0429] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a controller, cause the controller to perform the method described in the foregoing embodiments.
[0430] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0431] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A multi-modal in-vehicle noise reduction method, characterized by, The vehicle includes a microphone array and a camera device, and the method includes: The audio data collected by the microphone array in the target space and the visual data captured by the camera device in the target space are obtained, wherein the target space includes the target sound source of the target person and the interference sound source; The visual location and visual covariance matrix of the target sound source are determined based on the visual data. The audio location and audio covariance matrix of the target sound source are determined based on the audio data. Based on a preset Kalman filter, the visual position, the visual covariance matrix, the audio position, and the audio covariance matrix are fused to obtain the sound source position of the target sound source. Obtain a list of interfering sound sources, the list of interfering sound sources includes the sound source locations of one or more interfering sound sources, and a confidence level is matched for each interfering sound source; Based on the list of interference sources, interference sources that meet the preset confidence conditions are identified as sources to be suppressed. Based on a preset noise reduction algorithm, the audio data is noise-reduced using the location of the target sound source and the location of the sound source to be suppressed, so that the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced.
2. The method of claim 1, wherein, The preset noise reduction algorithm uses the sound source locations of the target sound source and the sound source locations of the sound source to be suppressed to perform noise reduction processing on the audio data, so that the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced, including: Based on the visual data, multiple invalid audio segments are filtered out from the audio data, which includes multiple temporal audio segments. The invalid audio segments are temporal audio segments that do not belong to the time when the target person is speaking. Generate a data covariance matrix for the noise space based on multiple invalid audio segments; Based on the noise reduction algorithm, array manifold processing is performed on the sound source positions of the target sound source and the sound source to be suppressed to obtain a dynamic constraint matrix; Based on the noise reduction algorithm, the data covariance matrix and the preset target response vector are weighted to obtain the target weight matrix; Based on the target weight matrix and the audio data, the audio of the sound source to be suppressed is suppressed and the audio of the target sound source is enhanced.
3. The method of claim 2, wherein, The step of filtering out multiple invalid audio segments from the audio data based on the visual data includes: The visual data and the audio data are processed into frames to obtain multiple visual segments and multiple temporal audio segments, and one visual segment is aligned with one temporal audio segment. Based on the alignment relationship between the visual segment and the temporal audio segment, multiple invalid audio segments are filtered out from among the multiple temporal audio segments.
4. The method of claim 3, wherein, The step of filtering out multiple invalid audio segments among multiple temporal audio segments based on the alignment relationship between the visual segment and the temporal audio segment includes: Perform a person recognition operation on each of the visual segments to obtain the identity information of the visual segments; In response to the fact that the identity information of the visual segment is the identity information of the target person, a lip movement detection operation is performed on the visual segment to obtain the lip movement detection information of the visual segment; In response to the lip movement detection information of the visual segment being speech activity information, an audio segment aligned with the visual segment is set as a valid audio segment; In response to all visual segments, the person recognition and lip movement detection operations are completed. All valid audio segments are removed from all audio segments to obtain the set of remaining segments. Set each audio segment in the remaining segment set to an invalid audio segment.
5. The method of claim 2, wherein, The generation of the data covariance matrix for the noise space based on multiple invalid audio segments includes: The invalid audio segment is processed by a short-time Fourier transform to obtain an invalid frequency domain segment; Calculate the periodogram based on the invalid frequency domain segment; The periodograms of all invalid frequency domain segments are arithmetically averaged to obtain the data covariance matrix of the noise space in which the microphone array is located.
6. The method of claim 2, wherein, The microphone array includes multiple array elements. Each interfering sound source in the list of interfering sound sources is configured with a confidence level. The array manifold processing is performed on the sound source positions of the target sound source and the sound source positions of the sound sources to be suppressed based on the noise reduction algorithm to obtain a dynamic constraint matrix, including: Based on the positions of the array elements and the positions of the target sound source, the target steering vector from the target sound source to all array elements is determined; The interference steering vector from the sound source to all array elements is determined based on the position of the array elements and the position of the sound source to be suppressed. The dynamic constraint matrix is determined based on the target guidance vector and all interference guidance vectors.
7. The method of claim 2, wherein, The step of performing weight calculation processing on the data covariance matrix and the preset target response vector based on the noise reduction algorithm to obtain the target weight matrix includes: An optimization problem is constructed based on the weight matrix to be solved and the data covariance matrix; Constraints are constructed based on the weight matrix to be solved and the preset target response vector; The optimization problem is solved based on the constraints to obtain the target weight matrix.
8. The method of claim 2, wherein, The step of performing suppression processing on the audio of the sound source to be suppressed and enhancement processing on the audio of the target sound source based on the target weight matrix and the audio data includes: The time-domain audio segment is subjected to a short-time Fourier transform to obtain a frequency-domain audio segment; Based on the target weight matrix and the frequency domain audio segment, the denoised frequency domain audio segment is determined; Perform an inverse short-time Fourier transform on the denoised frequency domain audio segment to obtain the denoised time domain audio segment. Combine all the denoised time-domain audio segments to obtain the denoised speech signal.
9. The method according to any one of claims 1 to 8, characterized in that, The list of interference sources includes: Acquire various sound source description data; The location of the interfering sound source is determined based on the sound source description data; Record the location of all interfering sound sources to obtain a list of interfering sound sources.
10. The method of claim 9, wherein, The various sound source description data include the visual data, the audio data, vehicle operation data, and inherent knowledge source data. The interfering sound sources include actively predicted noise, passively sensed noise, vehicle noise, and stationary noise. Determining the sound source location of the interfering sound source based on the sound source description data includes: Based on the visual data, lip movement detection and position calculation operations are performed to obtain the visual reference position of the non-target person, and the visual reference position of the non-target person is set as the sound source position of the actively predicted noise. Based on the audio data, determine the audio reference position of the non-target person, and set the audio reference position of the non-target person as the sound source position of the passively perceived noise. The vehicle status is determined based on the vehicle operation data, an acoustic context description vector is generated based on the vehicle status, and the sound source location of the vehicle noise is determined based on the acoustic context description vector. The location of the fixed noise source is determined based on the inherent knowledge source data.
11. The method according to any one of claims 1 to 8, characterized in that, Each of the aforementioned interfering sound sources corresponds to an information source attribute, the information source attribute being used to indicate the determined source of the interfering sound source, and the step of determining interfering sound sources that meet preset confidence conditions as sound sources to be suppressed based on the list of interfering sound sources includes: Obtain the information source attributes of the interference sound source; The basic weight of the interference sound source is determined based on the information source attributes of the interference sound source. The weight correction coefficient is determined based on the sound source description data of the interference sound source; The confidence level of the interfering sound source is determined based on the basic weights and the weight correction coefficients. Based on the confidence level of each of the aforementioned interference sources, interference sources that meet the preset confidence level conditions are identified as sources to be suppressed.
12. The method of claim 11, wherein, The step of determining the interference sources that meet the preset confidence conditions as the sources to be suppressed based on the confidence levels of each interference source includes: In response to the fact that the confidence level of the interfering sound source is greater than a preset confidence threshold, it is determined that the interfering sound source meets the preset confidence condition; Set the interference sound source as the sound source to be suppressed.
13. A controller characterized by comprising: The system includes a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, and the processor, when executing the one or more computer programs, causing the controller to implement the multimodal in-vehicle noise reduction method as described in any one of claims 1-12.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the multimodal in-vehicle noise reduction method as described in any one of claims 1-12.