Sound source positioning method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202511912023.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-12-17
AI Technical Summary
[0003]本发明提供一种声源定位方法、装置、电子设备和存储介质,用以解决现有的声源定位方法难以在复杂应用场景下提升声源定位的准确性的缺陷
[0015]本发明提供的声源定位方法、装置、电子设备和存储介质,通过对待定位的音频信号进行初步分区,得到初步声源分区结果;基于初步声源分区结果,引导自适应波束形成器对音频信号进行处理,生成空间指向性更明确的增强信号和阻塞信号;将音频信号的特征向量、增强信号的特征向量和阻塞信号的特征向量输入至语音增强模型,由语音增强模型基于多维度特征进行声源分区,输出音频信号的实际声源分区结果,极大地提高了声源定位的准确性,适用于复杂应用场景。
Smart Images

Figure CN121838810B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, and in particular to a sound source localization method, apparatus, electronic device, and storage medium. Background Technology
[0002] In practical applications such as voice communication, intelligent interaction, and intelligent conferencing, accurately distinguishing sound sources in different spatial locations is a crucial prerequisite for achieving efficient audio processing. Currently, existing technologies mostly employ deep clustering algorithms or neural network models for sound source localization; however, these methods are only suitable for conventional application scenarios with good signal-to-noise ratios. In complex real-world scenarios, especially in environments with low to medium signal-to-noise ratios, environmental noise severely contaminates audio signals. This causes phase difference features, which serve as key spatial cues, to be severely disrupted by noise, significantly reducing the spatial distinguishability of sound sources. This makes deep learning models that rely on these features prone to misjudgment, leading to significant errors in sound source localization and partitioning results. Summary of the Invention
[0003] This invention provides a sound source localization method, apparatus, electronic device, and storage medium to address the shortcomings of existing sound source localization methods in improving the accuracy of sound source localization in complex application scenarios.
[0004] In a first aspect, the present invention provides a method for locating a sound source, comprising the following steps: The audio signal to be located is initially partitioned to obtain preliminary sound source partitioning results; The audio signal is processed using an adaptive beamformer to generate enhanced and blocked signals; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results. The feature vectors of the audio signal, the enhanced signal, and the blocking signal are input into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model. The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, the feature vectors of the blocked signal samples, and the actual sound source partitioning labels of the audio signal samples.
[0005] In some embodiments, the preliminary partitioning of the audio signal to be located to obtain preliminary sound source partitioning results includes: Feature extraction is performed on the audio signal to obtain the feature vector of the audio signal; The feature vector of the audio signal is input into the speech partitioning model to obtain the preliminary sound source partitioning result output by the speech partitioning model; the speech partitioning model is trained based on the feature vector of the audio signal sample and the label of the preliminary sound source partitioning result of the audio signal sample.
[0006] In some embodiments, updating the parameters of the initial adaptive beamformer includes: Based on the preliminary sound source partitioning results, speech activity detection is performed on different sound source regions to obtain the speech detection results for each sound source region. Based on the speech detection results of each sound source region, the parameters of the initial adaptive beamformer corresponding to each sound source region are updated to obtain the adaptive beamformer corresponding to each sound source region.
[0007] In some embodiments, the preliminary sound source partitioning result includes multiple sound source regions, and the processing of the audio signal based on an adaptive beamformer to generate enhanced and blocking signals includes: Based on the preliminary sound source partitioning results, sub-audio signals for each sound source region are determined from the audio signals; Based on the adaptive beamformer corresponding to each sound source region, the sub-audio signal of each sound source region is processed to generate the enhanced signal and the blocking signal of each sound source region.
[0008] In some embodiments, the speech enhancement model includes a feature fusion layer and a prediction layer; The feature fusion layer is used to fuse the feature vectors of the audio signal, the feature vector of the enhanced signal, and the feature vector of the jamming signal to obtain a fused feature vector. The prediction layer is used to predict the sound source region of the audio signal based on the fused feature vector, so as to obtain the actual sound source partitioning result of the audio signal.
[0009] In some embodiments, the speech partitioning model is trained based on the following steps: Acquire audio signal samples and determine the preliminary sound source partitioning results labels of the audio signal samples; Feature extraction is performed on the audio signal sample to obtain the feature vector of the audio signal sample; The feature vector of the audio signal sample is input into the initial speech partitioning model to obtain the preliminary sound source partitioning prediction result output by the initial speech partitioning model; Based on the preliminary sound source partitioning prediction results and the labels of the preliminary sound source partitioning results, the total loss function value is calculated. Based on the total loss function value, the parameters of the initial speech partitioning model are iteratively optimized to obtain the speech partitioning model. The total loss function value is calculated based on the mean squared error loss function and the scale-invariant signal-to-noise ratio function.
[0010] In some embodiments, the speech enhancement model is trained based on the following steps: Acquire audio signal samples and determine the actual sound source partitioning label of the audio signal samples; The audio signal samples are initially partitioned to obtain preliminary sound source partitioning results samples; The audio signal samples are processed based on the adaptive beamformer samples to generate enhanced signal samples and jammed signal samples; the adaptive beamformer samples are obtained by updating the parameters of the initial adaptive beamformer samples based on the preliminary sound source partitioning results samples. Extract the feature vectors of the audio signal samples, extract the feature vectors of the enhanced signal samples, and extract the feature vectors of the blocked signal samples; The initial speech enhancement model is trained using the feature vectors of the audio signal sample, the feature vector of the enhanced signal sample, and the feature vector of the blocked signal sample, and the actual sound source partitioning result label of the audio signal sample is used as the sample label. After training, the speech enhancement model is obtained.
[0011] In some embodiments, updating the parameters of the initial adaptive beamformer sample based on the preliminary sound source partitioning result sample includes: Based on the preliminary sound source partitioning results sample, speech activity detection is performed to obtain the initial speech detection result sample; The initial speech detection result sample is probabilistically inverted to obtain a speech detection result sample containing simulated errors; Based on the speech detection result sample, the parameters of the initial adaptive beamformer sample are updated to obtain the adaptive beamformer sample.
[0012] In a second aspect, the present invention also provides a sound source localization device, comprising: The preliminary partitioning unit is used to perform preliminary partitioning of the audio signal to be located, and obtain preliminary sound source partitioning results; The processing unit is used to process the audio signal based on the adaptive beamformer to generate an enhanced signal and a blocking signal; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results; The prediction unit is used to input the feature vector of the audio signal, the feature vector of the enhanced signal and the feature vector of the blocking signal into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model; The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, the feature vectors of the blocked signal samples, and the actual sound source partitioning labels of the audio signal samples.
[0013] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the sound source localization methods described above.
[0014] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the sound source localization method as described above.
[0015] The sound source localization method, apparatus, electronic device, and storage medium provided by this invention obtain preliminary sound source localization results by performing preliminary partitioning on the audio signal to be localized; based on the preliminary sound source localization results, an adaptive beamformer is guided to process the audio signal to generate enhanced and blocking signals with more defined spatial directivity; the feature vectors of the audio signal, the enhanced signal, and the blocking signal are input into a speech enhancement model, which performs sound source localization based on multi-dimensional features and outputs the actual sound source localization results of the audio signal, which greatly improves the accuracy of sound source localization and is suitable for complex application scenarios. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a schematic flowchart of the sound source localization method provided in an embodiment of the present invention.
[0018] Figure 2 This is a flowchart illustrating the training process of the speech partitioning model provided in this embodiment of the invention.
[0019] Figure 3 This is a flowchart illustrating the training process of the speech enhancement model provided in this embodiment of the invention.
[0020] Figure 4 This is a schematic diagram of the sound source localization device provided in an embodiment of the present invention.
[0021] Figure 5 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, in this invention, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0024] Figure 1 This is a flowchart illustrating the sound source localization method provided in an embodiment of the present invention. Figure 1 As shown, a sound source localization method is provided, including the following steps: step 110, step 120, and step 130. This method's steps are merely one possible implementation of the present invention.
[0025] Step 110: Perform preliminary partitioning of the audio signal to be located to obtain preliminary sound source partitioning results.
[0026] Optionally, a multi-channel audio signal is acquired through a microphone array, such as headphones, hearing aids, or conference audio acquisition equipment. This audio signal is a mixed signal that includes all sound components, such as one or more target sound sources from different locations in the space, ambient noise, and reverberation. The audio signal includes multiple time-frequency units.
[0027] Optionally, preliminary sound source partitioning is performed based on the amplitude and phase information of the audio signal to obtain preliminary sound source partitioning results; the preliminary sound source partitioning results can be a set of time-frequency masks. Assuming there are two sound sources, two masks are generated.
[0028] Specifically, a short-time Fourier transform is performed on the audio signal to obtain the complex spectrum of each channel; features that can characterize the spatial location and spectral characteristics of the sound source are extracted from the spectrum; the features of each time-frequency unit are input into a pre-trained deep neural network to obtain the embedding vector of each time-frequency unit; the embedding vectors of all time-frequency units are collected, and these vectors are grouped using a clustering algorithm; based on the grouping results, a preliminary sound source partitioning result is generated, i.e., a set of binary time-frequency masks.
[0029] The preliminary sound source partitioning results include the Ideal Ratio Mask (IRM) values for different sound source regions. The IRM value represents the proportion of energy occupied by the target speech at a specific time and frequency point. If a time-frequency unit is completely dominated by the target speech, the IRM value is close to 1; if a time-frequency unit is completely dominated by noise, the IRM value is close to 0; if a time-frequency unit is in a mixed state, the IRM value is between 0 and 1.
[0030] Step 120: Process the audio signal based on the adaptive beamformer to generate enhanced and blocked signals; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results.
[0031] The adaptive beamformer assigns specific complex weights to the sub-audio signals received by each microphone in the microphone array, and then sums the weighted sub-audio signals to enhance the sound signal coming from a specific direction while suppressing interference and noise from other directions. The initial parameters of the adaptive beamformer can be default values. Based on the preliminary sound source partitioning results, the parameters of the initial adaptive beamformer are updated to obtain the final parameters of the adaptive beamformer.
[0032] Optionally, based on the preliminary sound source zoning results, the signal covariance matrix and noise / interference covariance matrix of the target sound source are estimated; according to optimization criteria, such as the minimum variance distortionless response criterion, a new set of optimized weight parameters is calculated to replace the initial parameters.
[0033] The enhanced signal primarily contains the target sound source component, and has a higher signal-to-noise ratio compared to the original signal. The blocking signal retains only all other components of the audio signal except the target sound source, namely interference sources, background noise, and ambient reverberation.
[0034] Step 130: Input the feature vectors of the audio signal, the enhanced signal, and the blocking signal into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model.
[0035] The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, the feature vectors of the blocked signal samples, and the actual sound source partitioning labels of the audio signal samples.
[0036] Optionally, the main structure of the speech enhancement model is a feed-forward sequential memory network (FSMN), which is obtained by adding memory modules to the feed-forward neural network; the FSMN uses memory modules to capture contextual information.
[0037] Optionally, feature extraction is performed on the audio signal to obtain a feature vector; feature extraction is performed on the enhanced signal to obtain a feature vector; and feature extraction is performed on the slugging signal to obtain a feature vector. The feature vector includes at least the amplitude spectrum and the logarithmic amplitude spectrum.
[0038] Optionally, the actual sound source partitioning result is a high-precision time-frequency mask; compared with the preliminary sound source partitioning result, its accuracy and noise resistance are greatly improved, and it can accurately identify the sound source to which each time-frequency unit belongs.
[0039] In some embodiments, the speech enhancement model includes a feature fusion layer and a prediction layer; The feature fusion layer is used to fuse the feature vectors of the audio signal, the enhanced signal, and the jamming signal to obtain a fused feature vector. The prediction layer is used to predict the sound source region of the audio signal based on the fused feature vector, and obtain the actual sound source partitioning result of the audio signal.
[0040] Specifically, the feature fusion layer is used to: determine the weights of the audio signal, the augmented signal, and the blocking signal; and based on the weights of the audio signal, the augmented signal, and the blocking signal, fuse the feature vectors of the audio signal, the augmented signal, and the blocking signal to obtain a fused feature vector.
[0041] It should be noted that since the input of the speech enhancement model includes the features of the original audio signal and the beam features after preliminary segmentation, if there are errors in the beam features after preliminary segmentation, the misjudged speech in the preliminary segmentation can be corrected by combining the features of the original audio signal, thereby improving the accuracy of sound source segmentation.
[0042] In this embodiment of the invention, a preliminary sound source partitioning result is obtained by performing preliminary partitioning on the audio signal to be located. Based on the preliminary sound source partitioning result, an adaptive beamformer is guided to process the audio signal to generate an enhanced signal and a blocking signal with more specific spatial directivity. The feature vectors of the audio signal, the enhanced signal, and the blocking signal are input into the speech enhancement model, which performs sound source partitioning based on multi-dimensional features and outputs the actual sound source partitioning result of the audio signal. This greatly improves the accuracy of sound source localization and is suitable for complex application scenarios.
[0043] In some embodiments, step 110, performing preliminary partitioning of the audio signal to be located to obtain preliminary sound source partitioning results, includes: Step 111: Extract features from the audio signal to obtain the feature vector of the audio signal.
[0044] Step 112: Input the feature vector of the audio signal into the speech partitioning model to obtain the preliminary sound source partitioning result output by the speech partitioning model; the speech partitioning model is trained based on the feature vector of the audio signal sample and the label of the preliminary sound source partitioning result of the audio signal sample.
[0045] Among them, the feature vector of an audio signal is a set of values extracted from the audio signal to summarize and represent the key acoustic characteristics of the audio signal at a specific time-frequency unit; the feature vector of an audio signal includes at least spectral features and spatial features; for example, the logarithmic amplitude spectrum describes the distribution of sound energy at different frequencies; another example is to reflect the direction of arrival of sound waves by calculating the phase difference of signals received by different microphones.
[0046] Optionally, the main structure of the speech partitioning model is FSMN, which is obtained by improving the traditional deep neural network. A memory module is added to the hidden layer to store the node information of past moments, which can effectively utilize the relevant information of the time series. Compared with recurrent neural networks, it does not need to set up feedback connections to maintain the causality of the network.
[0047] In this embodiment of the invention, by inputting the feature vector of the audio signal into the speech partitioning model, the preliminary sound source partitioning result output by the speech partitioning model is obtained, which improves the real-time performance and accuracy of the preliminary partitioning.
[0048] In some embodiments, updating the parameters of the initial adaptive beamformer includes: Based on the preliminary sound source zoning results, speech activity detection is performed in different sound source regions to obtain the speech detection results for each sound source region. Based on the speech detection results of each sound source region, the parameters of the initial adaptive beamformer corresponding to each sound source region are updated to obtain the adaptive beamformer corresponding to each sound source region.
[0049] Among them, the sound source region is a set of time-frequency units of a specific sound source, and the preliminary sound source partitioning results include multiple sound source regions.
[0050] Voice Activity Detection (VAD) is a key tool for distinguishing between valid information and noise. Its function is to determine whether human speech exists in a given audio segment. In a given time frame, if speech is detected, "1" is output; if there is only background noise or silence, "0" is output; if the IRM value of each frequency point in the current frame is greater than a preset threshold, then speech activity is determined to exist; otherwise, no speech activity is determined to exist.
[0051] Optionally, the parameters of the initial adaptive beamformer can be updated using a generalized sidelobe elimination algorithm, a minimum variance distortionless response algorithm, or a linearly constrained minimum variance algorithm.
[0052] For example, using a generalized sidelobe cancellation algorithm, assuming there are 4 sound source regions, in the first frame, 4 sets of adaptive beam parameters are initialized; then, based on the 4 IRMs of different sound source regions, it is determined whether there is speech activity in each sound source region in the current frame, and the beam of the corresponding sound source region is updated frame by frame accordingly.
[0053] In this embodiment of the invention, independent speech activity detection is first performed on each sound source region based on the preliminary partitioning results. Then, the initial adaptive beamformer parameters corresponding to different sound source regions can be updated in a targeted manner, thereby significantly improving the beamformer's tracking accuracy of the target sound source.
[0054] In some embodiments, the initial sound source partitioning results include multiple sound source regions. The audio signal is processed based on an adaptive beamformer to generate enhanced and blocking signals, including: Based on the preliminary sound source zoning results, sub-audio signals for each sound source region are determined from the audio signals; Based on the adaptive beamformer corresponding to each sound source region, the sub-audio signal of each sound source region is processed to generate the enhanced signal and the blocking signal of each sound source region.
[0055] Among them, the sub-audio signal refers to the signal extracted from the audio signal based on the preliminary sound source partitioning results and associated with a specific sound source region.
[0056] In this embodiment of the invention, the global problem is decomposed into multiple local problems. Each adaptive beamformer only needs to focus on a specific sound source region. The generation processes of the enhanced signal and the blocking signal in each sound source region are independent of each other. The sub-audio signals of multiple sound source regions can be processed in parallel, thereby improving the overall computational efficiency and meeting the requirements of real-time processing.
[0057] Figure 2 This is a flowchart illustrating the training process of the speech partitioning model provided in an embodiment of the present invention. Figure 2 As shown, in some embodiments, the speech partitioning model is trained based on the following steps: Step 210: Obtain audio signal samples and determine the preliminary sound source partitioning results labels of the audio signal samples.
[0058] Optionally, the microphone array of the target device is used to record audio, collecting a large amount of clean speech and noise, and synthesizing audio signal samples.
[0059] Step 220: Extract features from the audio signal samples to obtain the feature vectors of the audio signal samples.
[0060] Optionally, the feature vector of the audio signal sample includes at least spectral features and spatial features.
[0061] Step 230: Input the feature vector of the audio signal sample into the initial speech partitioning model to obtain the preliminary sound source partitioning prediction results output by the initial speech partitioning model.
[0062] Optionally, the main structure of the initial speech partitioning model is FSMN.
[0063] Step 240: Based on the preliminary sound source partitioning prediction results and the preliminary sound source partitioning result labels, calculate the total loss function value. Based on the total loss function value, iteratively optimize the parameters of the initial speech partitioning model to obtain the speech partitioning model. The total loss function value is calculated based on the mean square error loss function and the scale-invariant signal-to-noise ratio function.
[0064] Taking headphones as an example, if each earphone has two microphones, a 4-channel microphone array is formed. When recording with this array, for the sound source locations to be partitioned, such as the front, back, left, and right regions, the room impact response of the target is simulated using the mirror source method, and convolved with clean speech to obtain the array's received signal. Then, scattered noise and point source noise are superimposed and mixed according to different signal-to-noise ratios to generate the original noisy array training data. The input features of the initial partitioning model are the logarithmic amplitude spectrum of the four microphones and the phase difference information of different microphone pairs, and the learning target is the IRM value of different sound source regions.
[0065] In this embodiment of the invention, the initial speech partitioning model is trained using a mean squared error loss function, enabling the speech partitioning model to generate preliminary sound source partitioning results with high accuracy. At the same time, the initial speech partitioning model is trained using a scale-invariant signal-to-noise ratio function, enabling the speech partitioning model to generate preliminary sound source partitioning results with high fidelity. This dual-objective driven training strategy improves the performance of the speech partitioning model and ensures that the preliminary sound source partitioning results generated by the model have higher accuracy and fidelity.
[0066] Figure 3 This is a flowchart illustrating the training process of the speech enhancement model provided in an embodiment of the present invention. Figure 3 As shown, in some embodiments, the speech enhancement model is trained based on the following steps: Step 310: Obtain audio signal samples and determine the actual sound source partitioning result label of the audio signal samples.
[0067] Optionally, the microphone array of the target device is used to record audio, collecting a large amount of clean speech and noise, and synthesizing audio signal samples.
[0068] Step 320: Perform preliminary partitioning of the audio signal samples to obtain preliminary sound source partitioning results samples.
[0069] Optionally, feature extraction is performed on the audio signal samples to obtain feature vectors of the audio signal samples; the feature vectors of the audio signal samples are then input into the speech partitioning model to obtain preliminary sound source partitioning results samples output by the speech partitioning model.
[0070] Step 330: Process the audio signal samples based on the adaptive beamformer samples to generate enhanced signal samples and blocking signal samples; the adaptive beamformer samples are obtained by updating the parameters of the initial adaptive beamformer samples based on the preliminary sound source partitioning results samples.
[0071] In some embodiments, the parameters of the initial adaptive beamformer sample are updated based on the preliminary sound source partitioning results sample, including: Based on the preliminary sound source partitioning results sample, speech activity detection is performed to obtain the initial speech detection result sample; The initial speech detection result sample is probabilistically inverted to obtain a speech detection result sample containing simulated errors. Based on the speech detection results samples, the parameters of the initial adaptive beamformer samples are updated to obtain the adaptive beamformer samples.
[0072] It should be noted that in low to medium signal-to-noise ratio environments, segmentation based solely on the original features may lead to errors. Therefore, training data for situations where the speech segmentation model makes mistakes can be supplemented into the speech enhancement model. Specifically, the ideal VAD value can be probabilistically inverted.
[0073] For example, when the signal-to-noise ratio (SNR) is less than 5dB, the ideal VAD value is inverted with a probability of 2.5%; when the SNR is between 5 and 10dB, the inversion probability is 2%; when the SNR is between 10 and 15dB, the inversion probability is 1.5%; and when the SNR is not less than 15dB, the inversion probability is 1.5%.
[0074] Step 340: Extract the feature vectors of the audio signal samples, extract the feature vectors of the enhanced signal samples, and extract the feature vectors of the blocked signal samples.
[0075] Optionally, feature extraction is performed on the audio signal sample to obtain the feature vector of the audio signal sample; feature extraction is performed on the enhanced signal sample to obtain the feature vector of the enhanced signal sample; and feature extraction is performed on the blocked signal sample to obtain the feature vector of the blocked signal sample.
[0076] Step 350: Using the feature vectors of the audio signal sample, the feature vector of the enhanced signal sample, and the feature vector of the suffocating signal sample as training samples, and the actual sound source partitioning result labels of the audio signal sample as sample labels, train the initial speech enhancement model. After training, the speech enhancement model is obtained.
[0077] Optionally, an initial speech enhancement model is trained based on the mean squared error loss function and the scale-invariant signal-to-noise ratio function.
[0078] Optionally, the initial speech enhancement model includes an initial feature fusion layer and an initial prediction layer; The initial feature fusion layer is used to fuse the feature vectors of audio signal samples, enhanced signal samples, and blocked signal samples to obtain fused feature vector samples. The initial prediction layer is used to predict the sound source region of the audio signal sample based on the fused feature vector sample, and obtain the actual sound source partition prediction result of the audio signal sample.
[0079] Optionally, after the speech partitioning model and speech enhancement model have been trained, they can be tested.
[0080] In this embodiment of the invention, by introducing speech detection results containing simulated errors to update the beamformer, and using the enhanced signal samples and blocking signal samples generated by it as training data, the robustness and generalization ability of the speech enhancement model are significantly improved. This enables the model to maintain stable and reliable performance even when facing inaccurate VAD in real and complex environments, thereby enhancing the fault tolerance and practicality of the system.
[0081] The sound source localization device provided in the embodiments of the present invention is described below. The sound source localization device described below can be referred to in correspondence with the sound source localization method described above.
[0082] Figure 4 This is a schematic diagram of the sound source localization device provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the sound source locating device 400 includes: The preliminary partitioning unit 410 is used to perform preliminary partitioning of the audio signal to be located, and obtain preliminary sound source partitioning results; The processing unit 420 is used to process the audio signal based on the adaptive beamformer to generate enhanced and blocked signals; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results; The prediction unit 430 is used to input the feature vectors of the audio signal, the feature vectors of the enhanced signal, and the feature vectors of the blocking signal into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model. The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, the feature vectors of the blocked signal samples, and the actual sound source partitioning labels of the audio signal samples.
[0083] Optionally, the audio signal to be located is initially partitioned to obtain preliminary sound source partitioning results, including: Feature extraction is performed on the audio signal to obtain the feature vector of the audio signal; The feature vector of the audio signal is input into the speech partitioning model to obtain the preliminary sound source partitioning result output by the speech partitioning model; the speech partitioning model is trained based on the feature vector of the audio signal sample and the label of the preliminary sound source partitioning result of the audio signal sample.
[0084] Optionally, the parameters of the initial adaptive beamformer are updated, including: Based on the preliminary sound source zoning results, speech activity detection is performed in different sound source regions to obtain the speech detection results for each sound source region. Based on the speech detection results of each sound source region, the parameters of the initial adaptive beamformer corresponding to each sound source region are updated to obtain the adaptive beamformer corresponding to each sound source region.
[0085] Optionally, the preliminary sound source zoning results include multiple sound source regions. The audio signal is processed using an adaptive beamformer to generate enhanced and blocking signals, including: Based on the preliminary sound source zoning results, sub-audio signals for each sound source region are determined from the audio signals; Based on the adaptive beamformer corresponding to each sound source region, the sub-audio signal of each sound source region is processed to generate the enhanced signal and the blocking signal of each sound source region.
[0086] Optionally, the speech enhancement model includes a feature fusion layer and a prediction layer; The feature fusion layer is used to fuse the feature vectors of the audio signal, the enhanced signal, and the jamming signal to obtain a fused feature vector. The prediction layer is used to predict the sound source region of the audio signal based on the fused feature vector, and obtain the actual sound source partitioning result of the audio signal.
[0087] Optionally, the speech segmentation model is trained based on the following steps: Acquire audio signal samples and determine the preliminary sound source partitioning labels for the audio signal samples; Feature extraction is performed on the audio signal samples to obtain the feature vector of the audio signal samples; The feature vectors of the audio signal samples are input into the initial speech partitioning model to obtain the preliminary sound source partitioning prediction results output by the initial speech partitioning model. Based on the preliminary sound source partitioning prediction results and the preliminary sound source partitioning result labels, the total loss function value is calculated. Based on the total loss function value, the parameters of the initial speech partitioning model are iteratively optimized to obtain the speech partitioning model. The total loss function value is calculated based on the mean square error loss function and the scale-invariant signal-to-noise ratio function.
[0088] Optionally, the speech enhancement model is trained based on the following steps: Acquire audio signal samples and determine the actual sound source partitioning labels of the audio signal samples; The audio signal samples are initially partitioned to obtain preliminary sound source partitioning results samples; The audio signal samples are processed based on the adaptive beamformer samples to generate enhanced signal samples and jammed signal samples; the adaptive beamformer samples are obtained by updating the parameters of the initial adaptive beamformer samples based on the preliminary sound source partitioning results samples. Extract the feature vectors of audio signal samples, extract the feature vectors of enhanced signal samples, and extract the feature vectors of blocked signal samples; The initial speech enhancement model is trained using the feature vectors of the audio signal sample, the feature vector of the enhanced signal sample, and the feature vector of the blocked signal sample, and the actual sound source partitioning result label of the audio signal sample is used as the sample label. After training, the speech enhancement model is obtained.
[0089] Optionally, based on the preliminary sound source partitioning results samples, the parameters of the initial adaptive beamformer samples are updated, including: Based on the preliminary sound source partitioning results sample, speech activity detection is performed to obtain the initial speech detection result sample; The initial speech detection result sample is probabilistically inverted to obtain a speech detection result sample containing simulated errors. Based on the speech detection results samples, the parameters of the initial adaptive beamformer samples are updated to obtain the adaptive beamformer samples.
[0090] It should be noted that the sound source localization device provided in this embodiment of the invention can implement all the method steps implemented in the above-described sound source localization method embodiment and can achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0091] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a sound source localization method, which includes: performing preliminary partitioning of the audio signal to be localized to obtain preliminary sound source partitioning results; processing the audio signal based on an adaptive beamformer to generate an enhanced signal and a blocking signal; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results; inputting the feature vectors of the audio signal, the enhanced signal, and the blocking signal into a speech enhancement model to obtain the actual sound source partitioning results of the audio signal output by the speech enhancement model; wherein the speech enhancement model is trained based on the feature vectors of the audio signal samples, the enhanced signal samples, and the blocking signal samples, as well as the labels of the actual sound source partitioning results of the audio signal samples.
[0092] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0093] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the sound source localization method provided by the above methods. The method includes: performing preliminary partitioning of the audio signal to be localized to obtain preliminary sound source partitioning results; processing the audio signal based on an adaptive beamformer to generate an enhanced signal and a blocking signal; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results; inputting the feature vectors of the audio signal, the enhanced signal, and the blocking signal into a speech enhancement model to obtain the actual sound source partitioning results of the audio signal output by the speech enhancement model; wherein the speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, the feature vectors of the blocking signal samples, and the actual sound source partitioning result labels of the audio signal samples.
[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0095] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method of acoustic source localization, the method comprising: include: The audio signal to be located is initially partitioned to obtain preliminary sound source partitioning results, which include the ideal ratio mask IRM for different sound source regions. The audio signal is processed using an adaptive beamformer to generate enhanced and blocked signals; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results. The feature vectors of the audio signal, the enhanced signal, and the blocking signal are input into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model. The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, and the feature vectors of the blocked signal samples, as well as the actual sound source partitioning labels of the audio signal samples. The updating of the parameters of the initial adaptive beamformer includes: Based on the preliminary sound source partitioning results, speech activity detection is performed on different sound source regions to obtain the speech detection results for each sound source region. Based on the speech detection results of each sound source region and the IRM of each sound source region, the parameters of the initial adaptive beamformer corresponding to each sound source region are updated to obtain the adaptive beamformer corresponding to each sound source region.
2. The acoustic source positioning method of claim 1, wherein, The preliminary partitioning of the audio signal to be located, to obtain preliminary sound source partitioning results, includes: Feature extraction is performed on the audio signal to obtain the feature vector of the audio signal; The feature vector of the audio signal is input into the speech partitioning model to obtain the preliminary sound source partitioning result output by the speech partitioning model; the speech partitioning model is trained based on the feature vector of the audio signal sample and the label of the preliminary sound source partitioning result of the audio signal sample.
3. The sound source localization method according to claim 2, characterized in that, The preliminary sound source partitioning results include multiple sound source regions. The processing of the audio signal based on the adaptive beamformer to generate enhanced and blocking signals includes: Based on the preliminary sound source partitioning results, sub-audio signals for each sound source region are determined from the audio signals; Based on the adaptive beamformer corresponding to each sound source region, the sub-audio signal of each sound source region is processed to generate the enhanced signal and the blocking signal of each sound source region.
4. The sound source localization method according to claim 1, characterized in that, The speech enhancement model includes a feature fusion layer and a prediction layer; The feature fusion layer is used to fuse the feature vectors of the audio signal, the feature vector of the enhanced signal, and the feature vector of the jamming signal to obtain a fused feature vector. The prediction layer is used to predict the sound source region of the audio signal based on the fused feature vector, so as to obtain the actual sound source partitioning result of the audio signal.
5. The sound source localization method according to claim 2, characterized in that, The speech partitioning model is trained based on the following steps: Acquire audio signal samples and determine the preliminary sound source partitioning results labels of the audio signal samples; Feature extraction is performed on the audio signal sample to obtain the feature vector of the audio signal sample; The feature vector of the audio signal sample is input into the initial speech partitioning model to obtain the preliminary sound source partitioning prediction result output by the initial speech partitioning model; Based on the preliminary sound source partitioning prediction results and the labels of the preliminary sound source partitioning results, the total loss function value is calculated. Based on the total loss function value, the parameters of the initial speech partitioning model are iteratively optimized to obtain the speech partitioning model. The total loss function value is calculated based on the mean squared error loss function and the scale-invariant signal-to-noise ratio function.
6. The sound source localization method according to claim 1, characterized in that, The speech enhancement model is trained based on the following steps: Acquire audio signal samples and determine the actual sound source partitioning label of the audio signal samples; The audio signal samples are initially partitioned to obtain preliminary sound source partitioning results samples; The audio signal samples are processed based on the adaptive beamformer samples to generate enhanced signal samples and jammed signal samples; the adaptive beamformer samples are obtained by updating the parameters of the initial adaptive beamformer samples based on the preliminary sound source partitioning results samples. Extract the feature vectors of the audio signal samples, extract the feature vectors of the enhanced signal samples, and extract the feature vectors of the blocked signal samples; The initial speech enhancement model is trained using the feature vectors of the audio signal sample, the feature vector of the enhanced signal sample, and the feature vector of the blocked signal sample, and the actual sound source partitioning result label of the audio signal sample is used as the sample label. After training, the speech enhancement model is obtained.
7. The sound source localization method according to claim 6, characterized in that, The step of updating the parameters of the initial adaptive beamformer sample based on the preliminary sound source partitioning results sample includes: Based on the preliminary sound source partitioning results sample, speech activity detection is performed to obtain the initial speech detection result sample; The initial speech detection result sample is probabilistically inverted to obtain a speech detection result sample containing simulated errors; Based on the speech detection result sample, the parameters of the initial adaptive beamformer sample are updated to obtain the adaptive beamformer sample.
8. A sound source localization device, characterized in that, include: The preliminary partitioning unit is used to perform preliminary partitioning of the audio signal to be located, and obtain preliminary sound source partitioning results. The preliminary sound source partitioning results include the ideal ratio mask (IRM) for different sound source regions. The processing unit is used to process the audio signal based on the adaptive beamformer to generate an enhanced signal and a blocking signal; the adaptive beamformer is obtained by updating the parameters of the initial adaptive beamformer based on the preliminary sound source partitioning results; The prediction unit is used to input the feature vector of the audio signal, the feature vector of the enhanced signal and the feature vector of the blocking signal into the speech enhancement model to obtain the actual sound source partitioning result of the audio signal output by the speech enhancement model; The speech enhancement model is trained based on the feature vectors of the audio signal samples, the feature vectors of the enhanced signal samples, and the feature vectors of the blocked signal samples, as well as the actual sound source partitioning labels of the audio signal samples. The updating of the parameters of the initial adaptive beamformer includes: Based on the preliminary sound source partitioning results, speech activity detection is performed on different sound source regions to obtain the speech detection results for each sound source region. Based on the speech detection results of each sound source region and the IRM of each sound source region, the parameters of the initial adaptive beamformer corresponding to each sound source region are updated to obtain the adaptive beamformer corresponding to each sound source region.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the sound source localization method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the sound source localization method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice processing method, device and facility, and storage medium
CN108806707A
Multi-channel adaptive speech signal processing with noise reduction
CN1753084A
Low-latency speech separation
US20200322722A1
Speech enhancement method and apparatus, and device and computer-readable storage medium
WO2022105571A1