User input audio low distortion processing method in multi-source noise environment
By calculating the sound source steering vector and frequency domain observation column vector of the microphone array, updating the noise covariance matrix, and adopting the minimum variance distortionless response criterion, the problem of target echo being misjudged as interference noise in a strong reverberation environment in an in-vehicle intelligent cockpit is solved, and low-distortion audio processing is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI SHENGWANG TECH CO LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-01
AI Technical Summary
In the voice interaction system of in-vehicle intelligent cockpit, existing technologies are prone to misidentifying the target echo as interference noise in strong reverberation environments, resulting in audio distortion. Existing algorithms such as MVDR and GSC perform poorly in strong reverberation environments and cannot effectively suppress interference and maintain the quality of the target voice.
By acquiring the sound source steering vector and frequency domain observation column vector of the microphone array, the power spectra of the superimposed reference signal and the differential reference signal are calculated. Combined with the preset acoustic transmission upper limit energy ratio and interference presence probability weight, the noise covariance matrix is updated. Beamforming is performed using the minimum variance distortionless response criterion to accurately distinguish target echoes from external interference.
It effectively suppresses external interference, reduces audio distortion, improves speech enhancement, and ensures low-distortion audio processing capabilities in multi-source noise environments.
Smart Images

Figure CN121617411B_ABST
Abstract
Description
Low-distortion processing method for user input audio in multi-source noise environments Technical Field
[0001] This invention relates to the field of audio processing technology, and more specifically to a method for low-distortion processing of user input audio in multi-source noise environments. Background Technology
[0002] The voice interaction system of in-vehicle intelligent cockpit faces complex acoustic challenges in practical applications. Among them, the most difficult problem occurs in scenarios with two speakers and strong reverberation. For example, when the driver is giving a voice command, the front passenger speaks at the same time, and the enclosed space of the car causes the driver's voice to reflect through the windows and dashboard, forming high-energy reverberation, which causes audio distortion.
[0003] In existing technologies, adaptive beamforming algorithms, such as Minimum Variance Distortionless Response (MVDR) or Generalized Sidelobe Canceller (GSC), are used to suppress interference. However, in strong reverberation environments, the reflected sound energy of the driver may be higher than the direct sound of the passenger, causing the algorithm to mistakenly identify the target echo as interference noise. Furthermore, once the noise covariance matrix incorrectly absorbs the components of the target echo, the beamformer will form nulls in the target direction, resulting in the target speech being incorrectly suppressed and the audio distortion being greater. Summary of the Invention
[0004] To address the technical problem of misidentifying target echoes as interference noise, resulting in greater audio distortion, this invention aims to provide a low-distortion audio processing method for user input in multi-source noise environments. The specific technical solution adopted is as follows:
[0005] This invention proposes a method for low-distortion processing of user input audio in multi-source noise environments, the method comprising:
[0006] Obtain the sound source steering vector of the microphone array pointing to the target sound source at each frequency, and the frequency domain observation column vector of each frame at each time point;
[0007] Based on the sound source steering vector at each frequency and the frequency observation column vector at each time frame, the superimposed reference signal and differential reference signal at each frequency and time frame are obtained. For any frequency, based on the energy characteristics of the superimposed reference signal and differential reference signal at different time frames, the signal power spectrum of the initial time frame, and the preset acoustic transmission upper limit energy ratio, the superimposed reference signal power spectrum and differential reference signal power spectrum at each time frame are obtained, and the joint interference decision degree at each time frame is obtained.
[0008] Based on the joint interference decision degree distribution of each time frame, the interference presence probability weight of each time frame is obtained; based on the interference presence probability weight of different time frames, the noise covariance matrix of the initial time frame, and the frequency domain observation column vector, the noise covariance matrix of the current time frame is obtained.
[0009] Based on the noise covariance matrix, sound source steering vector, and frequency domain observation column vector of the current frame at each frequency, the target direction enhancement output signal of the current frame at each frequency is obtained.
[0010] Furthermore, the method for obtaining the superimposed reference signal and the differential reference signal includes:
[0011] The product of the conjugate transpose of the sound source steering vector at each frequency and the elements in the same order in the frequency domain observation column vector of each time frame is superimposed and divided by the number of microphones to serve as the superposition reference signal for each time frame at each frequency.
[0012] The product of the conjugate of each element in the frequency domain observation column vector and the corresponding element in the sound source steering vector at each time frame at each frequency is obtained as the first product; the root mean square value of half the difference between the first products between different adjacent elements is obtained as the differential reference signal.
[0013] Furthermore, the method for obtaining the superimposed reference signal power spectrum and the differential reference signal power spectrum includes:
[0014] For any superimposed reference signal or differential reference signal at any frequency, the signal power spectrum of the next time frame is obtained based on the preset smoothing factor, the signal power spectrum of the initial time frame, and the signal amplitude of the next time frame. The smoothing process is continuously iterated to obtain the signal power spectrum of each time frame.
[0015] Furthermore, the method for obtaining the signal power spectrum of each time frame includes:
[0016] The difference between the positive integer 1 and the preset smoothing factor is used as the observation weight;
[0017] The product of the preset smoothing factor and the signal power spectrum of the initial time frame is obtained as the first product; the product of the observation weight and the square of the signal amplitude of the next time frame is obtained as the second product; the sum of the first product and the second product is obtained as the signal power spectrum of the next time frame.
[0018] The next time frame is used as the new initial time frame, and the signal power spectrum of the next time frame is calculated to obtain the signal power spectrum of each time frame.
[0019] Furthermore, the method for obtaining the degree of joint interference decision includes:
[0020] For any frequency, the over-limit reflection energy value is obtained based on the superimposed reference signal power spectrum and the differential reference signal power spectrum of each time frame, as well as the preset acoustic transmission upper limit energy ratio.
[0021] The relative mutation rate of the differential signal at each time frame is obtained by the difference in signal amplitude between the superimposed reference signal and the differential reference signal at each time frame and the previous time frame.
[0022] The joint interference decision level for each time frame is obtained based on the over-limit reflection energy value, the differential reference signal power spectrum, and the differential signal relative mutation rate. The over-limit reflection energy value and the differential signal relative mutation rate are positively correlated with the joint interference decision level, while the differential reference signal power spectrum is negatively correlated with the joint interference decision level.
[0023] Furthermore, the method for obtaining the over-limit reflection energy value includes:
[0024] For each frame at any frequency, the product of the preset acoustic transmission upper limit energy ratio and the superimposed reference signal power spectrum is obtained. The difference between the differential reference signal power spectrum and the product result is obtained. If the difference result is greater than the preset difference threshold, the corresponding difference result is taken as the over-limit reflection energy value; otherwise, the over-limit reflection energy value is set to 0.
[0025] Furthermore, the method for obtaining the relative mutation rate of the differential signal includes:
[0026] For any superimposed reference signal or differential reference signal at any frequency, obtain the difference in signal amplitude between each time frame and the previous time frame. If the difference result is greater than the preset difference threshold, the difference result is used as the positive amplitude increment; otherwise, the positive amplitude increment is set to 0.
[0027] Obtain the first sum of the positive amplitude increment of the superimposed reference signal and the preset adjustment coefficient, and calculate the ratio of the positive amplitude increment of the differential reference signal to the first sum as the relative mutation rate of the differential signal.
[0028] Furthermore, the method for obtaining the probability weight of the interference includes:
[0029] For any frequency, obtain the first difference between the joint interference decision degree and the preset decision threshold for each time frame, calculate the product of the first difference and the preset mapping slope, and normalize it as the interference presence probability weight.
[0030] Furthermore, the method for obtaining the noise covariance matrix includes:
[0031] For any frequency, the product of the preset update step size and the interference existence probability weight of each time frame is obtained as the first weight; the difference between the positive integer 1 and the first weight is obtained as the second weight.
[0032] The product of the noise covariance matrix of the initial frame and the second weight of the next frame is obtained as the first product matrix; the product of the frequency domain observation column vector and the conjugate transpose of the frequency domain observation column vector of the next frame is obtained as the spatial covariance matrix; the product of the first weight and the spatial covariance matrix is obtained as the second product matrix; the sum of the first product matrix and the second product matrix is obtained as the noise covariance matrix of the next frame.
[0033] Using the next time frame as the new initial time frame, the noise covariance matrix of the next time frame is iteratively calculated to obtain the noise covariance matrix of the current time frame.
[0034] Furthermore, the method for acquiring the target direction enhanced output signal includes:
[0035] Diagonally load the noise covariance matrix of the current frame at each frequency to obtain the diagonally loaded noise matrix;
[0036] Based on the inverse matrix of the diagonally loaded noise matrix and the sound source steering vector of the current frame at each frequency, the beamforming weight vector is obtained by adopting the minimum variance distortionless response criterion.
[0037] The product of the conjugate transpose of the ratio vector and the frequency domain observation column vector is calculated and used as the target direction enhancement output signal.
[0038] The present invention has the following beneficial effects:
[0039] This invention obtains the superimposed reference signal and differential reference signal for each time frame at each frequency based on the source steering vector and the frequency observation column vector for each time frame. This helps characterize the dominant target component and the interference and reverberation components. For any frequency, based on the energy characteristics of the superimposed reference signal and the differential reference signal at different time frames, the signal power spectrum of the initial time frame, and the preset acoustic transmission upper limit energy ratio, the interference presence probability weight for each time frame is obtained, reflecting the possibility of independent interference in each time frame. Based on the interference presence probability weights for different time frames, the noise covariance matrix of the initial time frame, and the frequency domain observation column vector, the noise covariance matrix of the current time frame is obtained, reflecting a statistical matrix that only includes spatial information of external interference, excluding target echo contamination. Based on the noise covariance matrix, source steering vector, and frequency domain observation column vector of the current time frame at each frequency, the target direction enhancement output signal for the current time frame at each frequency is obtained. This invention achieves low-distortion speech enhancement by accurately obtaining the interference presence probability weight and precisely distinguishing between target echo and independent external interference. Attached Figure Description
[0040] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 is a flowchart of a method for low-distortion processing of user input audio in a multi-source noise environment according to an embodiment of the present invention.
[0042] Figure 2 is a flowchart of a method for obtaining the degree of joint interference decision provided in an embodiment of the present invention. Detailed Implementation
[0043] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a user input audio low-distortion processing method for multi-source noise environments proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0044] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0045] The following description, in conjunction with the accompanying drawings, details a specific scheme for a low-distortion processing method for user input audio in multi-source noise environments provided by the present invention.
[0046] Please refer to Figure 1, which shows a flowchart of a method for low-distortion processing of user input audio in a multi-source noise environment according to an embodiment of the present invention. The specific method includes:
[0047] Step S1: Obtain the sound source steering vector of the microphone array pointing to the target sound source at each frequency, and the frequency domain observation column vector of each frame at each time point.
[0048] In embodiments of the present invention, considering that in a strong reverberation environment, the reflected sound energy of the driver may be higher than the direct sound of the passenger, causing the algorithm to mistakenly identify the target echo as interference noise, resulting in significant speech distortion; therefore, the spatial characteristics of the target sound source and the physical causality of acoustic propagation are analyzed; the sound source steering vector is a The column vector is represented as ;in, Indicates frequency Lower sound source steering vector; Indicates frequency The first sound source guiding vector The element corresponding to each microphone; M is the number of microphones in the microphone array; Indicates the transpose symbol; the elements in the sound source steering vector from which each microphone points to the target sound source, expressed by the formula: ;in, Indicates frequency The first sound source guiding vector The element corresponding to each microphone; Indicates the distance from the target sound source to the th The propagation time of sound waves from each microphone; Represents the imaginary unit; Indicates frequency; This represents the natural constant. It should be noted that, in the embodiments of the present invention, the sound wave propagation time from the target sound source to each microphone is obtained in advance based on actual vehicle calibration data.
[0049] Audio signals from each microphone are acquired using an onboard microphone array. The signals are then framed, windowed, and subjected to Fast Fourier Transform (FFT) to obtain a frequency domain observation column vector consisting of the complex spectral coefficients of each microphone at each time frame. This vector is represented as follows: ;in, Indicates frequency Time frame The frequency domain observation column vector; Indicates frequency Time frame The frequency domain observation column vector below The element corresponding to each microphone; This represents the transpose symbol. It should be noted that the specific methods used are well-known to those skilled in the art and will not be elaborated upon here.
[0050] Step S2: Based on the sound source steering vector at each frequency and the frequency observation column vector of each time frame, obtain the superimposed reference signal and differential reference signal for each time frame at each frequency; for any frequency, based on the energy characteristics of the superimposed reference signal and differential reference signal at different time frames, the signal power spectrum of the initial time frame, and the preset acoustic transmission upper limit energy ratio, obtain the superimposed reference signal power spectrum and differential reference signal power spectrum for each time frame, and obtain the joint interference decision degree for each time frame.
[0051] To obtain the reference signal containing the highest direct sound energy from the target, the signals from each microphone are superimposed after relative alignment. To obtain a reference signal that theoretically does not contain the direct sound from the target, a spatial null is forcibly formed in the target direction, eliminating possible local noise singularities in a single channel and improving the ability to characterize the spatial distribution of the interference sound field. Based on the source steering vector at each frequency and the frequency observation column vector at each time frame, the superimposed reference signal and differential reference signal at each frequency and each time frame are obtained.
[0052] Preferably, in one embodiment of the present invention, the method for obtaining the superimposed reference signal and the differential reference signal includes:
[0053] The product of the conjugate transpose of the sound source steering vector at each frequency and the elements of the same order in the frequency domain observation column vector at each time frame is superimposed and divided by the number of microphones to obtain the superposition reference signal for each time frame at each frequency; the formula is expressed as: ; Indicates frequency Time frame The superimposed reference signal; Indicates frequency The conjugate transpose of the lower sound source steering vector; Indicates frequency Time frame The frequency domain observation column vector; This indicates the number of microphones in the microphone array.
[0054] The product of the conjugate of each element in the frequency domain observation column vector at each time frame at each frequency and the corresponding element in the source steering vector is obtained as the first product; the difference between the first products of adjacent elements is obtained, and the root mean square value of half the difference between all adjacent elements is calculated as the differential reference signal; the formula is expressed as:
[0055] ;in, Indicates frequency Time frame Differential reference signal below; Indicates frequency Time frame The frequency domain observation column vector below The element corresponding to each microphone; Indicates frequency The first sound source guiding vector The conjugate element of each microphone element, i.e. ; Indicates frequency Time frame The frequency domain observation column vector below The element corresponding to each microphone; Indicates frequency The first sound source guiding vector The conjugate element of each microphone element; Indicates the number of microphones in the microphone array; This represents half of the difference result, to avoid the magnitude doubling after subtraction.
[0056] It should be noted that the original elements in the sound source steering vector represent the phase lag caused by the sound wave traveling from the sound source to each microphone, and the conjugate elements represent the phase compensation required to align the signals. The product of the conjugate results of each element in the frequency domain observation column vector and the corresponding elements in the sound source steering vector is obtained. In other words, in order to restore the phase of the target signal of each microphone to zero phase and thus achieve alignment, the difference between the first products of adjacent sequences can be calculated to eliminate the same spatial components canceled by the difference.
[0057] Considering that the larger the energy characteristics of the differential signal relative to the superimposed signal, the greater the degree of violation of the physical laws of reflection, and the more likely external interference will occur; the power spectrum of the signal in the initial frame reflects the power spectrum of the system at the start-up time, which helps to provide a reference for subsequent events; the preset acoustic transmission upper limit energy ratio reflects the maximum energy gain of the target direct sound after reflection and diffraction into the differential channel, and the larger the preset acoustic transmission upper limit energy ratio, the greater the differential reflection energy; for any frequency, based on the energy characteristics of the superimposed reference signal and the differential reference signal in different frames, the power spectrum of the signal in the initial frame, and the preset acoustic transmission upper limit energy ratio, the power spectrum of the superimposed reference signal and the power spectrum of the differential reference signal in each frame are obtained, and the joint interference decision degree of each frame is obtained.
[0058] Preferably, in one embodiment of the present invention, the method for obtaining the superimposed reference signal power spectrum and the differential reference signal power spectrum includes:
[0059] For any superimposed reference signal or differential reference signal at any frequency, the signal power spectrum of the next time frame is obtained based on the preset smoothing factor, the signal power spectrum of the initial time frame, and the signal amplitude of the next time frame. The smoothing process is continuously iterated to obtain the signal power spectrum of each time frame.
[0060] Based on this, the signal power spectrum reflects the energy distribution intensity of the signal at each frequency component, providing stable input data for subsequent energy comparison.
[0061] Preferably, in one embodiment of the present invention, the method for obtaining the signal power spectrum of each time frame includes:
[0062] The difference between the positive integer 1 and the preset smoothing factor is used as the observation weight;
[0063] The product of the preset smoothing factor and the signal power spectrum of the initial time frame is obtained as the first product; the product of the observation weight and the square of the signal amplitude of the next time frame is obtained as the second product; the sum of the first product and the second product is obtained as the signal power spectrum of the next time frame.
[0064] Using the next time frame as the new initial time frame, the signal power spectrum of the next time frame is calculated to obtain the signal power spectrum of each time frame; the formula is expressed as: ;in, Indicates frequency The next frame The signal power spectrum below; Indicates the preset smoothing factor; Indicates frequency The corresponding previous frame The signal power spectrum below; Indicates frequency The next frame The signal amplitude below.
[0065] It should be noted that the signal power spectrum of the initial frame... The calculation begins by obtaining the signal power spectrum of the next time frame according to the formula. The signal power spectrum of the next time frame is then used as the signal power spectrum of the new initial time frame. The calculation of the signal power spectrum of the next time frame continues, and the process is repeated through recursive smoothing iterations to obtain the signal power spectrum of each time frame.
[0066] It should be noted that a larger smoothing factor results in a larger proportion of historical time frames, stronger noise resistance, and less susceptibility to instantaneous spikes. In one embodiment of the invention, the smoothing factor can be preset to any value within the range of 0.8-0.95 based on relevant historical experience, such as 0.9. The signal power spectrum of the initial time frame is set to a non-zero minimum value, such as... This avoids the error of the denominator being zero in subsequent energy comparison calculations and prevents numerical instability in the initial time frame. In other embodiments of the present invention, the signal power spectrum and the magnitude of the preset smoothing factor of the initial time frame can be set according to specific circumstances, and are not limited or elaborated here.
[0067] The proportion of energy leaked from the target direct sound into the differential signal has an upper limit determined by the physical environment and varies with frequency. If the energy in the differential signal exceeds the upper limit, the excess energy must come from external independent interference that is not constrained by the reflection path, and the greater the degree of joint interference, the better.
[0068] Preferably, in one embodiment of the present invention, the method for obtaining the degree of joint interference decision is shown in Figure 2, which illustrates a flowchart of a method for obtaining the degree of joint interference decision, including:
[0069] Step S201: For any frequency, obtain the over-limit reflection energy value based on the superimposed reference signal power spectrum and the differential reference signal power spectrum of each time frame, as well as the preset acoustic transmission upper limit energy ratio.
[0070] Preferably, in one embodiment of the present invention, the method for obtaining the over-limit reflection energy value includes:
[0071] For each frame at any frequency, the product of the preset acoustic transmission upper limit energy ratio and the superimposed reference signal power spectrum is obtained. The difference between the differential reference signal power spectrum and the product result is obtained. If the difference result is greater than the preset difference threshold, the corresponding difference result is taken as the over-limit reflection energy value; otherwise, the over-limit reflection energy value is set to 0.
[0072] It should be noted that the preset acoustic transmission upper limit energy ratio represents the physical limit of the ratio between the power spectrum of the differential reference signal and the power spectrum of the superimposed reference signal. It reflects the maximum energy gain of the direct sound entering the differential channel after reflection and diffraction in the in-vehicle acoustic environment. In-vehicle sound-absorbing materials have a higher absorption rate for high-frequency signals, and the value of this vector decreases with increasing frequency. In embodiments of this invention, the preset acoustic transmission upper limit energy ratio can be obtained through actual vehicle calibration: in a quiet in-vehicle environment, a white noise test signal is played from the driver's position, and the maximum value of the ratio between the power spectrum of the differential reference signal and the power spectrum of the superimposed reference signal at each frequency is calculated, with a certain safety margin added, such as... This yields the preset upper limit of acoustic transmission energy ratio at each frequency.
[0073] It should be noted that the product of the preset acoustic transmission upper limit energy ratio and the superimposed reference signal power spectrum reflects the maximum ideal energy of the target direct sound entering the differential channel. If the differential reference signal power spectrum is greater than the product result, it indicates that the actual differential signal energy is too strong, exceeding the limit that the target echo can reach. The larger the over-limit reflection energy value, the more it is affected by external interference. In the embodiments of the present invention, the preset difference threshold is set to 0.
[0074] Step S202: Based on the difference in signal amplitude between the superimposed reference signal and the differential reference signal at each time frame and the previous time frame, obtain the relative change rate of the differential signal at each time frame.
[0075] Preferably, according to the acoustic causality law, an echo cannot precede the direct sound, nor can it be stronger than the direct sound at the instant of abrupt change. Therefore, under normal sound source conditions, the amplitude of change of the differential signal is smaller than the amplitude of change of the superimposed signal, and the smaller the relative abrupt change rate of the differential signal; in one embodiment of the present invention, the method for obtaining the relative abrupt change rate of the differential signal includes:
[0076] For any superimposed reference signal or differential reference signal at any frequency, obtain the difference in signal amplitude between each time frame and the previous time frame. If the difference result is greater than the preset difference threshold, the difference result is used as the positive amplitude increment; otherwise, the positive amplitude increment is set to 0.
[0077] Obtain the first sum of the positive amplitude increment of the superimposed reference signal and the preset adjustment coefficient, and calculate the ratio of the positive amplitude increment of the differential reference signal to the first sum as the relative mutation rate of the differential signal.
[0078] It should be noted that the greater the difference in signal amplitude between each frame and the previous frame, the greater the upward trend of the signal energy, and the more important it is to capture the energy injection characteristics at the moment the sound source begins to emit sound. In the embodiments of the present invention, the preset difference threshold is set to 0.
[0079] It should be noted that, in order to avoid the formula being meaningless when the denominator is 0 when calculating the ratio, a very small positive number that is not zero is added as a preset adjustment coefficient. Its value can be set according to the range of the denominator, such as 0.01. The larger the positive amplitude increment of the differential reference signal is than the positive amplitude increment of the superimposed reference signal, the stronger the independent abrupt change of the differential signal is compared with the superimposed signal, the more likely there is external independent interference, and the greater the relative abrupt change rate of the differential signal. The smaller the positive amplitude increment of the differential reference signal is than the positive amplitude increment of the superimposed reference signal, the more it conforms to the physical characteristics of echo energy, which is usually less than that of direct sound, and the smaller the relative abrupt change rate of the differential signal.
[0080] Step S203: Based on the over-limit reflection energy value, differential reference signal power spectrum, and differential signal relative mutation rate of each time frame, obtain the joint interference decision degree of each time frame. The over-limit reflection energy value and differential signal relative mutation rate are positively correlated with the joint interference decision degree, while the differential reference signal power spectrum is negatively correlated with the joint interference decision degree.
[0081] It should be noted that the over-limit reflection energy value quantifies the degree of anomalous energy in the differential signal that cannot be explained by the target reflection model. The larger the over-limit reflection energy value, the more anomalous energy there is, the more susceptible it is to external interference, and the greater the degree of joint interference judgment. The relative mutation rate of the differential signal is used to measure whether the mutation degree of the differential signal is significantly higher than that of the superimposed signal, that is, the degree of incoherent energy surge of the differential signal relative to the superimposed signal. The larger the relative mutation rate of the differential signal, the greater the mutation degree of the differential signal is compared with the superimposed signal, the more susceptible it is to external independent interference, and the greater the degree of joint interference judgment. The power spectrum of the differential reference signal reflects the total energy of the differential signal. The larger the power spectrum of the differential reference signal, the smaller the over-limit reflection energy value is relative to the total energy, the smaller the relative influence, and the smaller the degree of joint interference judgment. Therefore, the over-limit reflection energy value and the relative mutation rate of the differential signal are both positively correlated with the degree of joint interference judgment, while the power spectrum of the differential reference signal is negatively correlated with the degree of joint interference judgment.
[0082] In one embodiment of the present invention, for each time frame of any frequency, the sum of a preset adjustment coefficient and the power spectrum of the differential reference signal is obtained as the differential reference signal adjusted power spectrum; the ratio of the over-limit reflection energy value to the differential reference signal adjusted power spectrum is obtained as the first interference coefficient.
[0083] The product of the preset transient enhancement factor and the relative mutation rate of the differential signal is obtained as the interference adjustment parameter; the product of the positive integer 1 and the interference adjustment parameter is obtained as the interference weight; the product of the interference weight and the first interference coefficient is obtained and normalized as the joint interference decision degree.
[0084] The formula is expressed as: ,in, Indicates frequency Time frame The degree of joint interference in the judgment; Indicates frequency Time frame The excess reflection energy value below; Indicates frequency Time frame The power spectrum of the differential reference signal below; This indicates the preset adjustment coefficient; This indicates the preset transient enhancement factor; Indicates frequency Time frame The relative abrupt change rate of the differential signal; This represents the logistic function, which maps the product of the interference weights and the first interference coefficient to... Normalize.
[0085] It should be noted that, in order to avoid the formula being meaningless when the denominator is 0 when calculating the ratio, a very small positive number that is not 0 is added as a preset adjustment coefficient. Its value can be set according to the specific range of the denominator, such as 0.01. In order to amplify the sensitivity of transient characteristics and ensure that the system can respond to sudden interference in the order of seconds, the transient enhancement factor is preset to 2 based on relevant historical experience.
[0086] Step S3: Obtain the interference presence probability weight of each time frame based on the joint interference decision degree distribution of each time frame; obtain the noise covariance matrix of the current time frame based on the interference presence probability weight of different time frames, the noise covariance matrix of the initial time frame, the frequency domain observation column vector, and the preset update step size.
[0087] The degree of joint interference decision reflects the comprehensive degree of violation of acoustic physics laws. The greater the degree of joint interference decision, the less it violates acoustic physics laws, and the more likely there is external interference. Based on the distribution of the degree of joint interference decision for each time frame, the probability weight of interference presence for each time frame is obtained.
[0088] Preferably, in one embodiment of the present invention, the method for obtaining the interference presence probability weight includes:
[0089] For any frequency, obtain the first difference between the joint interference decision degree and the preset decision threshold for each time frame, calculate the product of the first difference and the preset mapping slope, and normalize it as the interference presence probability weight.
[0090] It should be noted that, in one embodiment of the present invention, the preset judgment threshold reflects the sensitivity threshold of the system in determining the presence of interference. The larger the preset judgment threshold, the higher the threshold, and the greater the degree of joint interference judgment, the more likely interference is to exist. To avoid missing some interference due to excessively high interference and acquiring unnecessary interference due to excessively low interference, the preset judgment threshold is set to 0.5 based on relevant historical experience. The preset mapping slope reflects the hardness or softness of the control decision. The larger the preset mapping slope, the harder the control decision, ensuring the binarization characteristic of the decision, making the update logic of the noise matrix very crisp and reducing intermediate ambiguity states. Based on relevant historical experience, the preset mapping slope is set to 10. In other embodiments of the present invention, the preset judgment threshold and the preset mapping slope can be set according to specific circumstances, and are not limited or elaborated here.
[0091] It should be noted that, in one embodiment of the present invention, a sigmoid function is used to map the product of the first difference and the preset mapping slope to... Normalization is performed so that the probability weight of interference when the threshold exceeds the preset decision threshold is closer to 1, and the probability weight of interference when the threshold is less than the preset decision threshold is closer to 0, so that the feature when interference occurs is more significant; the specific means are well known to those skilled in the art and will not be described in detail here.
[0092] The interference probability weight reflects the proportion of external interference present in the audio at each time frame. The higher the interference probability weight, the more external interference noise there is. The noise covariance matrix of the initial time frame ensures that the system has an all-pass response characteristic at the initial moment, and can immediately output the unfiltered raw signal, avoiding mute or abnormal howling at the moment of startup. The frequency domain observation column vector reflects the mixed sound field information received by the microphone array, which includes the target speech, in-vehicle reverberation, external interference, and background noise. The noise covariance matrix of the current time frame is obtained based on the interference probability weight of different time frames, the noise covariance matrix of the initial time frame, and the frequency domain observation column vector.
[0093] Preferably, in one embodiment of the present invention, the method for obtaining the noise covariance matrix includes:
[0094] For any frequency, the product of the preset update step size and the interference existence probability weight of each time frame is obtained as the first weight; the difference between the positive integer 1 and the first weight is obtained as the second weight.
[0095] The product of the noise covariance matrix of the initial frame and the second weight of the next frame is obtained as the first product matrix; the product of the frequency domain observation column vector and the conjugate transpose of the frequency domain observation column vector of the next frame is obtained as the spatial covariance matrix; the product of the first weight and the spatial covariance matrix is obtained as the second product matrix; the sum of the first product matrix and the second product matrix is obtained as the noise covariance matrix of the next frame.
[0096] Using the next time frame as the new initial time frame, the noise covariance matrix of the next time frame is iteratively calculated to obtain the noise covariance matrix of the current time frame.
[0097] The formula is expressed as: ;in, Indicates frequency The next frame The noise covariance matrix; Indicates frequency The corresponding previous frame The noise covariance matrix; Indicates the preset update step size; Indicates frequency The next frame The interference has probability weights; Indicates frequency The next frame The frequency domain observation column vector; Indicates frequency The next frame The conjugate transpose of the frequency domain observation column vector.
[0098] Based on this Reflected in frequency Time frame The correlation between signals received by different microphones forms a corresponding matrix. The greater the correlation, the greater the probability weight of interference, and the greater the interference. The noise covariance matrix absorbs the spatial information of the frequency domain observation vector and locks the spatial direction of the interference source. The smaller the probability weight of interference, the more likely it is to be an echo, unaffected by external interference, and maintains the noise state of the previous frame.
[0099] It should be noted that, in one embodiment of the present invention, the preset update step size determines the algorithm's tracking response speed to changes in background noise. To avoid the matrix fluctuating too drastically with the current frame when the step size is too large, or to avoid being too sluggish in responding to new interference when the step size is too small, the preset update step size is set to a range based on relevant historical experience. The noise covariance matrix of the initial frame is defined as the identity matrix. In other embodiments of the present invention, the preset update step size is a technical means well known to those skilled in the art, and is not limited or described here.
[0100] It should be noted that the noise covariance matrix of the initial frame... The calculation begins by obtaining the noise covariance matrix of the next time frame according to the formula. The noise covariance matrix of the next time frame is then used as the noise covariance matrix of the new initial time frame. The calculation of the noise covariance matrix of the next time frame continues, and the calculation is iterated until the noise covariance matrix of the current time frame is obtained.
[0101] Step S4: Based on the noise covariance matrix, sound source steering vector, and frequency domain observation column vector of the current frame at each frequency, obtain the target direction enhancement output signal of the current frame at each frequency.
[0102] The noise covariance matrix reflects the energy distribution of the interfering noise, and the sound source steering vector reflects the location of the target sound source. The frequency domain observation column vector reflects the mixed information of the target speech, in-vehicle reverberation, external interference, and background noise. By combining the analysis, the target sound source pointed to by the sound source steering vector is extracted, and the interference and environmental noise of the passenger seat are deeply filtered out.
[0103] Preferably, in one embodiment of the present invention, the method for obtaining the target direction enhanced output signal includes:
[0104] For each frequency, the noise covariance matrix of the current frame is diagonally loaded to obtain a diagonally loaded noise matrix. It should be noted that the diagonal loading process is performed to balance the matrix's adaptability and robustness. This includes adding a small regularization term to the diagonal of the noise covariance matrix of the current frame, and taking the trace of the identity matrix based on relevant historical experience. to Any value in the range.
[0105] Based on the inverse matrix of the diagonally loaded noise matrix and the sound source steering vector of the current frame at each frequency, the beamforming weight vector is obtained by adopting the minimum variance distortionless response criterion.
[0106] The product of the conjugate transpose of the ratio vector and the frequency domain observation column vector is calculated and used as the target direction enhancement output signal.
[0107] It should be noted that the method for obtaining the minimum variance distortionless response criterion is as follows: obtain the product of the inverse matrix of the diagonally loaded noise matrix and the sound source steering vector of the current frame at each frequency, as the product column vector; obtain the product of the conjugate transpose of the sound source steering vector, the inverse matrix of the diagonally loaded noise matrix, and the sound source steering vector of the current frame at each frequency, as the product scalar value; obtain the ratio of the product column vector and the product scalar value as the beamforming weight vector; the specific means are well known to those skilled in the art and will not be elaborated here.
[0108] Based on this, the target direction enhancement output signal achieves deep filtering of co-pilot interference and environmental noise in the frequency domain. At the same time, because the noise matrix is not contaminated by the target echo and has anti-mismatch capability, the driver's direct sound and natural reverberation are completely and clearly preserved.
[0109] In summary, this invention obtains the superimposed reference signal and differential reference signal for each time frame at each frequency based on the source steering vector at each frequency and the frequency observation column vector for each time frame. For any frequency, based on the energy characteristics of the superimposed reference signal and differential reference signal at different time frames, the signal power spectrum of the initial time frame, and the preset acoustic transmission upper limit energy ratio, the interference presence probability weight for each time frame is obtained. Combining the noise covariance matrix of the initial time frame, the source steering vector, and the frequency domain observation column vector, the noise covariance matrix of the current time frame is obtained, and the target direction enhancement output signal for the current time frame at each frequency is obtained. This invention achieves low-distortion speech enhancement by accurately obtaining the interference presence probability weight and precisely distinguishing between target echoes and external independent interference.
[0110] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0111] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for low-distortion processing of user input audio in multi-source noise environments, characterized in that, The method includes: acquiring the sound source steering vector of the microphone array pointing to the target sound source at each frequency, and the frequency domain observation column vector of each time frame; obtaining the superimposed reference signal and differential reference signal of each time frame at each frequency based on the sound source steering vector at each frequency and the frequency observation column vector of each time frame; for any frequency, obtaining the superimposed reference signal power spectrum and differential reference signal power spectrum of each time frame based on the energy characteristics of the superimposed reference signal and differential reference signal at different time frames, the signal power spectrum of the initial time frame, and the preset acoustic transmission upper limit energy ratio, and obtaining the joint interference decision degree of each time frame; and obtaining the interference storage of each time frame based on the distribution of the joint interference decision degree of each time frame. Based on the probability weights, the noise covariance matrix of the initial frame, and the frequency domain observation column vector, the noise covariance matrix of the current frame is obtained. Based on the noise covariance matrix, sound source steering vector, and frequency domain observation column vector of the current frame at each frequency, the target direction enhancement output signal of the current frame at each frequency is obtained. The method for obtaining the superimposed reference signal and the differential reference signal includes: multiplying the conjugate transpose of the sound source steering vector at each frequency with the same-order elements of the frequency domain observation column vector of each frame, and dividing by the number of microphones to obtain the superimposed reference signal for each frame at each frequency; obtaining the frequency domain observation column vector of each frame at each frequency... The product of each element in the vector and the conjugate of the corresponding element in the sound source steering vector is used as the first product; the difference between the first products of adjacent elements is obtained, and the root mean square value of half the difference between all adjacent elements is calculated as the differential reference signal; the method for obtaining the degree of joint interference decision includes: for any frequency, obtaining the over-limit reflection energy value based on the superimposed reference signal power spectrum and the differential reference signal power spectrum of each time frame, as well as the preset acoustic transmission upper limit energy ratio; obtaining the relative change rate of the differential signal in each time frame based on the difference in signal amplitude between the superimposed reference signal and the differential reference signal between each time frame and the previous time frame; and obtaining the relative change rate of the differential signal in each time frame based on the over-limit reflection energy value and the differential reference signal in each time frame. The power spectrum and the relative mutation rate of the differential signal are used to obtain the joint interference decision level for each time frame. The over-limit reflection energy value and the relative mutation rate of the differential signal are positively correlated with the joint interference decision level, while the power spectrum of the differential reference signal is negatively correlated with the joint interference decision level. The method for obtaining the target direction enhancement output signal includes: performing diagonal loading processing on the noise covariance matrix of the current time frame at each frequency to obtain a diagonally loaded noise matrix; based on the inverse matrix of the diagonally loaded noise matrix of the current time frame at each frequency and the sound source steering vector, using the minimum variance distortionless response criterion, to obtain the beamforming weight vector; calculating the product of the conjugate transpose of the ratio vector and the frequency domain observation column vector as the target direction enhancement output signal.
2. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 1, characterized in that, The method for obtaining the superimposed reference signal power spectrum and the differential reference signal power spectrum includes: for any superimposed reference signal or differential reference signal at any frequency, obtaining the signal power spectrum of the next time frame based on a preset smoothing factor, the signal power spectrum of the initial time frame, and the signal amplitude of the next time frame, and continuously recursively smoothing and iterating to obtain the signal power spectrum of each time frame.
3. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 2, characterized in that, The method for obtaining the signal power spectrum of each time frame includes: obtaining the difference between a positive integer 1 and a preset smoothing factor as the observation weight; obtaining the product of the preset smoothing factor and the signal power spectrum of the initial time frame as the first product; obtaining the product of the observation weight and the square of the signal amplitude of the next time frame as the second product; obtaining the sum of the first product and the second product as the signal power spectrum of the next time frame; using the next time frame as the new initial time frame, calculating the signal power spectrum of the next time frame to obtain the signal power spectrum of each time frame.
4. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 1, characterized in that, The method for obtaining the over-limit reflection energy value includes: for each time frame of any frequency, obtaining the product of the preset acoustic transmission upper limit energy ratio and the superimposed reference signal power spectrum, obtaining the difference between the differential reference signal power spectrum and the product result, and if the difference result is greater than the preset difference threshold, the corresponding difference result is taken as the over-limit reflection energy value; otherwise, the over-limit reflection energy value is set to 0.
5. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 1, characterized in that, The method for obtaining the relative mutation rate of the differential signal includes: for any frequency superimposed reference signal or differential reference signal, obtaining the difference in signal amplitude between each time frame and the previous time frame; if the difference result is greater than a preset difference threshold, the difference result is used as the positive amplitude increment; otherwise, the positive amplitude increment is set to 0; obtaining the first sum of the positive amplitude increment of the superimposed reference signal and a preset adjustment coefficient, and calculating the ratio of the positive amplitude increment of the differential reference signal to the first sum as the relative mutation rate of the differential signal.
6. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 1, characterized in that, The method for obtaining the interference presence probability weight includes: for any frequency, obtaining the first difference between the joint interference decision degree and the preset decision threshold for each time frame, calculating the product of the first difference and the preset mapping slope, and normalizing it as the interference presence probability weight.
7. The method for low-distortion processing of user input audio in multi-source noise environments according to claim 1, characterized in that, The method for obtaining the noise covariance matrix includes: for any frequency, obtaining the product of a preset update step size and the interference presence probability weight of each time frame as a first weight; obtaining the difference between the positive integer 1 and the first weight as a second weight; obtaining the product of the noise covariance matrix of the initial time frame and the second weight of the next time frame as a first product matrix; obtaining the product of the frequency domain observation column vector and the conjugate transpose of the frequency domain observation column vector of the next time frame as a spatial covariance matrix; obtaining the product of the first weight and the spatial covariance matrix as a second product matrix; obtaining the sum of the first product matrix and the second product matrix as the noise covariance matrix of the next time frame; taking the next time frame as a new initial time frame, iteratively calculating the noise covariance matrix of the next time frame to obtain the noise covariance matrix of the current time frame.
Citation Information
Patent Citations
Speech recognition method suitable for noise environment
CN110148420A
Sound signal processing method, device, equipment, medium and chip
CN117037833A