Sound source positioning method and device, electronic equipment and storage medium

CN122260236BActive Publication Date: 2026-09-04JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610748243.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-04
Estimated Expiration
2046-05-28

AI Technical Summary

Technical Problem

其主要目的在于解决现有技术受环境噪声及混响干扰导致声源定位准确性低的问题

Benefits of technology

[0023]综上所述,本公开提供的声源定位方法、装置、电子设备及存储介质,该方法包括:从采集的多通道音频信号中提取多通道空间特征;其中,所述多通道空间特征包括信道间相位差特征及对数方向波束信噪比特征;将所述信道间相位差特征及所述对数方向波束信噪比特征进行特征融合,得到融合特征;将所述融合特征输入预先训练的声源定位模型,确定声源方位信息。与相关技术相比,本公开的方案能够通过融合信道间相位差特征及对数方向波束信噪比特征,提高了在噪声和混响环境下声源定位的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122260236B_ABST
    Figure CN122260236B_ABST
Patent Text Reader

Abstract

The present disclosure provides a sound source positioning method and device, electronic equipment and storage medium, which relates to the field of signal processing, and the main technical features include: extracting multi-channel spatial features from the collected multi-channel audio signals; wherein the multi-channel spatial features include inter-channel phase difference features and log directional beam signal-to-noise ratio features; the inter-channel phase difference features and the log directional beam signal-to-noise ratio features are fused to obtain fused features; and the fused features are input into a pre-trained sound source positioning model to determine the sound source direction information. By fusing the inter-channel phase difference features and the log directional beam signal-to-noise ratio features, the accuracy of sound source positioning in a noisy and reverberant environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of signal processing, and in particular to a sound source localization method, apparatus, electronic device, and storage medium. Background Technology

[0002] Sound source localization technology refers to the use of multi-channel audio signals collected by microphone arrays to determine the spatial location information of the sound source through signal processing. It is widely used in fields such as conference systems, smart speakers, and smart security.

[0003] Currently, traditional sound source localization methods are easily affected by environmental noise and reverberation in practical applications, resulting in low localization accuracy. Summary of the Invention

[0004] This disclosure provides a sound source localization method, apparatus, electronic device, and storage medium. Its main purpose is to solve the problem of low sound source localization accuracy caused by environmental noise and reverberation interference in the prior art.

[0005] According to a first aspect of this disclosure, a sound source localization method is provided, comprising:

[0006] Multi-channel spatial features are extracted from the acquired multi-channel audio signals; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features; The inter-channel phase difference features and the logarithmic directional beam signal-to-noise ratio features are fused to obtain fused features; The fused features are input into a pre-trained sound source localization model to determine the sound source's location information.

[0007] In some embodiments, extracting multi-channel spatial features from the acquired multi-channel audio signal includes: The inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features are extracted from the multi-channel audio signals acquired by the microphone array.

[0008] In some embodiments, extracting the inter-channel phase difference features from the multi-channel audio signals acquired from the microphone array includes: The multi-channel audio signal is subjected to time-frequency transformation to obtain the frequency domain representation of each channel at the time-frequency point; Based on the frequency domain representation of any two channels at the same time and frequency point, the phase difference between the two channels at the same time and frequency point is determined, and the phase difference characteristics between the channels are obtained.

[0009] In some embodiments, extracting the logarithmic directional beam signal-to-noise ratio features from the multi-channel audio signals acquired from the microphone array includes: Beamforming filters are designed for multiple preset directions, and the multi-channel audio signals are processed based on the beamforming filters to obtain the filtered signals for the multiple preset directions. The ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions is determined, and the ratio is subjected to logarithmic transformation to obtain the logarithmic directional beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

[0010] In some embodiments, before inputting the fused features into a pre-trained sound source localization model to determine the sound source location information, the method further includes: Obtain a training sample set; each training sample in the training sample set includes a multi-channel audio training signal and its corresponding sound source location training label. The inter-channel phase difference training features and logarithmic directional beam signal-to-noise ratio training features are extracted from the multi-channel audio training signal and fused to obtain the training fusion features. The training fusion features and the sound source location training labels are input into the initial sound source localization model, and the initial sound source localization model is trained based on a preset loss function to obtain the pre-trained sound source localization model; wherein, the preset loss function includes a cross-entropy loss function and a supervised contrastive loss function.

[0011] In some embodiments, obtaining the training sample set includes: Select clean speech signals from the speech library and select noise signals from the noise library; The clean speech signal is mixed with the noise signal based on the random signal-to-noise ratio, and random reverb and / or random channel jitter are added to obtain the multi-channel audio training signal. Set corresponding sound source location training labels for the multi-channel audio training signals to obtain the training sample set.

[0012] In some embodiments, inputting the fused features into a pre-trained sound source localization model to determine the sound source location information includes: The fused features are input into the pre-trained sound source localization model to obtain the sound source prediction probability for each preset direction; The preset direction with the highest predicted probability of the sound source is determined as the sound source location information.

[0013] According to a second aspect of this disclosure, a sound source localization device is provided, comprising: An extraction unit is used to extract multi-channel spatial features from the acquired multi-channel audio signal; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features; The first fusion unit is used to fuse the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features to obtain fused features. The determining unit is used to input the fused features into a pre-trained sound source localization model to determine the sound source location information.

[0014] In some embodiments, the extraction unit includes: The extraction module is used to extract the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features from the multi-channel audio signals acquired by the microphone array.

[0015] In some embodiments, the extraction module includes: The first determining submodule is used to perform time-frequency transformation on the multi-channel audio signal to obtain the frequency domain representation of each channel at the time-frequency point; The second determining submodule is used to determine the phase difference between any two channels at the same time and frequency point based on the frequency domain representation of any two channels at the same time and frequency point, thereby obtaining the inter-channel phase difference characteristics.

[0016] In some embodiments, the extraction module includes: The third determining submodule is used to design beamforming filters for multiple preset directions, and process the multi-channel audio signal based on the beamforming filters to obtain the filtered signals for the multiple preset directions. The fourth determining submodule is used to determine the ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions, and to perform logarithmic transformation on the ratio to obtain the logarithmic directional beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

[0017] In some embodiments, the apparatus further includes: The acquisition unit is used to acquire a training sample set before the determining unit inputs the fused features into the pre-trained sound source localization model to determine the sound source orientation information; each training sample in the training sample set includes a multi-channel audio training signal and its corresponding sound source orientation training label. The second fusion unit is used to extract inter-channel phase difference training features and logarithmic direction beam signal-to-noise ratio training features from the multi-channel audio training signal and fuse them to obtain training fusion features. The training unit is used to input the training fusion features and the sound source location training labels into the initial sound source localization model, and to train the initial sound source localization model based on a preset loss function to obtain the pre-trained sound source localization model; wherein, the preset loss function includes a cross-entropy loss function and a supervised contrastive loss function.

[0018] In some embodiments, the acquiring unit includes: The selection module is used to select clean speech signals from the speech library and noise signals from the noise library; The first determining module is used to mix the clean speech signal with the noise signal based on the random signal-to-noise ratio, and add random reverb and / or random channel jitter to obtain the multi-channel audio training signal; The second determining module is used to set corresponding sound source location training labels for the multi-channel audio training signals to obtain the training sample set.

[0019] In some embodiments, the determining unit includes: The input module is used to input the fused features into the pre-trained sound source localization model to obtain the sound source prediction probability for each preset direction; The third determining module is used to determine the preset direction with the highest predicted probability of the sound source as the sound source azimuth information.

[0020] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0021] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0022] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0023] In summary, the sound source localization method, apparatus, electronic device, and storage medium provided in this disclosure include: extracting multi-channel spatial features from acquired multi-channel audio signals; wherein the multi-channel spatial features include inter-channel phase difference features and logarithmic beam signal-to-noise ratio (SNR) features; fusing the inter-channel phase difference features and the logarithmic beam signal-to-noise ratio (SNR) features to obtain fused features; and inputting the fused features into a pre-trained sound source localization model to determine the sound source's directional information. Compared with related technologies, the solution disclosed in this disclosure can improve the accuracy of sound source localization in noisy and reverberant environments by fusing inter-channel phase difference features and logarithmic beam signal-to-noise ratio (SNR) features.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0025] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic flowchart illustrating a sound source localization method provided in an embodiment of the present disclosure. Figure 2 This is a schematic flowchart of a sound source localization method based on a microphone array provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of a channel phase difference feature extraction process provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram of a logarithmic directional beam signal-to-noise ratio feature extraction process provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of a sound source localization model training process provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of a training sample set construction process provided in an embodiment of the present disclosure; Figure 7 This is a schematic diagram of a sound source location determination process provided in an embodiment of the present disclosure; Figure 8 This is a schematic diagram of a sound source localization algorithm provided in an embodiment of the present disclosure; Figure 9 This is a schematic diagram of the structure of a sound source localization device provided in an embodiment of the present disclosure; Figure 10 This is a schematic diagram of another sound source localization device provided in an embodiment of the present disclosure; Figure 11 A schematic block diagram of an electronic device provided in an embodiment of this disclosure; Figure 12 This is a schematic block diagram of an example electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0026] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0027] This invention has a wide range of applications, and its advantages are particularly evident in products for intelligent voice interaction, intelligent conferencing, and intelligent security monitoring.

[0028] The sound source localization method, apparatus, electronic device, and storage medium of this disclosure are described below with reference to the accompanying drawings.

[0029] Figure 1 This is a schematic flowchart of a sound source localization method provided in an embodiment of the present disclosure.

[0030] like Figure 1 As shown, the method includes the following steps: Step 101: Extract multi-channel spatial features from the acquired multi-channel audio signals; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features.

[0031] In some embodiments, the multi-channel audio signal includes information such as the time delay and amplitude differences in the arrival of the sound source at different microphones. Inter-channel phase difference (IPD) refers to the phase difference between different microphone channels at the same frequency. This feature directly reflects the relative time difference of the sound waves arriving at each microphone, is closely related to the incident angle of the sound source, and has strong robustness to amplitude distortion in reverberant environments. Logarithmic Directional Signal-to-Noise Ratio (LDSNR) is a feature obtained by calculating the ratio of the energy of the target direction signal to the energy of the signal in the maximum interference direction after enhancing the target direction signal and suppressing the interference direction using beamforming technology, and then taking the logarithm. This feature can highlight the signal components in the target direction in noisy environments. By extracting these two features, the spatial location of the sound source can be described from both phase and energy information dimensions.

[0032] Using the above method, key features that reflect the spatial location of the sound source can be obtained from multi-channel audio signals, providing a foundation for subsequent feature fusion and sound source localization.

[0033] Step 102: The inter-channel phase difference feature and the logarithmic direction beam signal-to-noise ratio feature are fused to obtain the fused feature.

[0034] In some embodiments, feature fusion is the process of integrating two different types of features to form a single feature that can comprehensively reflect multiple aspects of information. Feature fusion methods include, but are not limited to, dimensional concatenation, feature concatenation, and weighted summation of features. During the fusion process, it is necessary to maintain the original information of the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features to avoid loss of feature information.

[0035] By way of example and not limitation, this disclosure discloses an embodiment that designs a time-frequency adaptive gating unit to dynamically weight and fuse IPD and LDSNR. The specific process is as follows: 1) Normalize the IPD features and LDSNR features to the same scale; 2) Construct the gating weights: ,in It is the Sigmoid activation function. for Convolution operation is used to perform a linear transformation on the input features along the channel dimension, where t represents the time frame index, f represents the frequency index, G(t,f) represents the gating weight at the time-frequency point (t,f), IPD(t,f) represents the inter-channel phase difference feature value at the time-frequency point (t,f), and LDSNR(t,f) represents the logarithmic directional beam signal-to-noise ratio feature value at the time-frequency point (t,f); 3) Adaptive fusion output: ,in, This indicates an element-wise multiplication operation. Fusion(t,f) represents the fused feature value obtained after adaptive weighted fusion at the time-frequency point (t,f). Through adaptive dynamic weighted fusion, the LDSNR weight is automatically increased when noise is dominant, and the IPD weight is automatically increased when reverberation is dominant, achieving complementary and synergistic enhancement.

[0036] The above method can integrate the information of the two types of features to obtain fused features that can comprehensively reflect the spatial attributes of multi-channel audio signals, thereby improving the feature representation ability.

[0037] Step 103: Input the fused features into the pre-trained sound source localization model to determine the sound source location information.

[0038] In some embodiments, the pre-trained sound source localization model can be a deep learning-based neural network model that has learned the mapping relationship between fused features and sound source location through a large amount of labeled training data. After the fused features are input into the model, the model calculates and outputs sound source location information, such as the azimuth angle of the sound source relative to the microphone array. The sound source location information can be used for various subsequent processing, such as controlling the camera to automatically track the speaker, adjusting the beamforming direction to enhance the sound in the target direction, or determining the user's location in intelligent voice interaction. The trained model can adapt to different acoustic environments and maintain good localization performance under noise and reverberation conditions.

[0039] By using the above method and reasoning about the fused features using a pre-trained model, the location information of the sound source can be accurately determined in complex acoustic environments.

[0040] In summary, the sound source localization method provided by the embodiments of this disclosure can improve the accuracy of sound source localization in noisy and reverberant environments by fusing inter-channel phase difference characteristics and logarithmic beam signal-to-noise ratio characteristics.

[0041] Figure 2 This is a schematic flowchart of a sound source localization method based on a microphone array provided in an embodiment of this disclosure, as shown below. Figure 2 As shown, the method includes steps 201-203.

[0042] Step 201: Extract the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features from the multi-channel audio signals acquired by the microphone array.

[0043] In some embodiments, a microphone array is an audio acquisition device composed of multiple microphones arranged in a preset manner. Each microphone corresponds to an independent acquisition channel, and the multiple microphones synchronously acquire audio signals from the surrounding environment to form a multi-channel audio signal. Each channel of the multi-channel audio signal corresponds to the audio data acquired by one microphone, and the audio data of each channel are synchronous and spatially correlated. Extracting inter-channel phase difference features and logarithmic beamwidth signal-to-noise ratio (SNR) features from the multi-channel audio signal acquired by the microphone array requires ensuring that both types of features can be completely extracted based on the channel characteristics of the multi-channel audio signal. Extraction methods include, but are not limited to, extraction methods based on signal processing algorithms. During the extraction process, it is necessary to ensure that the inter-channel phase difference features can accurately represent the phase difference between different microphone channels, and that the logarithmic beamwidth signal-to-noise ratio (SNR) features can accurately represent the logarithmic form of the SNR of audio signals in different directions. Simultaneously, the accuracy and completeness of the two types of feature data must be ensured to avoid feature information loss or distortion.

[0044] Step 202: The inter-channel phase difference feature and the logarithmic direction beam signal-to-noise ratio feature are fused to obtain the fused feature.

[0045] Step 203: Input the fused features into the pre-trained sound source localization model to determine the sound source location information.

[0046] For explanations of steps 202-203, please refer to [link / reference needed]. Figure 1 The detailed descriptions of the relevant embodiments are not repeated here.

[0047] The above method ensures that the extracted inter-channel phase difference features and logarithmic beam signal-to-noise ratio features match the audio signals acquired by the microphone array, providing accurate and effective feature data for subsequent feature fusion and sound source localization.

[0048] Figure 3 This is a schematic diagram of an inter-channel phase difference feature extraction process provided in an embodiment of this disclosure, based on Figure 2 The illustrated embodiment further explains step 201. Figure 3 This may include the following steps: Step 301: Perform time-frequency transformation on the multi-channel audio signal to obtain the frequency domain representation of each channel at the time-frequency point.

[0049] In some embodiments, the multi-channel audio signals acquired by the microphone array are time-domain signals, which need to be converted to the frequency domain for processing through time-frequency transformation. Time-frequency transformation can be implemented in various ways, including but not limited to Fast Fourier Transform (FFT), Short-Time Fourier Transform (SFT), or wavelet transform. Taking Fast Fourier Transform as an example, the multi-channel audio signal is segmented into frames, dividing the continuous audio signal into short time frames of fixed length. After applying a window function to each frame, a Fast Fourier Transform is performed to obtain the frequency domain representation of each channel at each time frame t and each frequency f. For each channel i, its frequency domain representation can be expressed in complex form, containing the real part. and the virtual part By using Fast Fourier Transform, the time-domain signal is decomposed into different frequency components, providing a basis for subsequent extraction of phase information.

[0050] Step 302: Based on the frequency domain representation of any two channels at the same time and frequency point, determine the phase difference between the two channels at the same time and frequency point to obtain the inter-channel phase difference characteristics.

[0051] In some embodiments, the inter-channel phase difference characteristics reflect the phase difference between different microphone channels at the same time-frequency point. For any two channels i and j, the frequency domain representations at the same time frame t and the same frequency f are respectively... and Each frequency domain representation is a complex number. First, calculate the phase angle of each channel at that time frequency. The phase angle can be obtained from the imaginary and real parts of the frequency domain representation using the arctangent function: the phase angle of channel i is... The phase angle of channel j is Then, the phase angles of the two channels are subtracted to obtain the inter-channel phase difference characteristics at that time-frequency point. The specific calculation formula is as follows: Where t represents the time frame index, f represents the frequency index, and i and j represent different microphone channel indices. and Let represent the real and imaginary parts of the frequency domain representation of the i-th channel at the time-frequency point (t,f), respectively. and Let represent the real and imaginary parts of the frequency domain representation of the j-th channel at the time-frequency point (t, f), respectively, and let arctan denote the arctangent function. This represents the inter-channel phase difference characteristic between channel i and channel j at the time-frequency point (t,f).

[0052] The above calculation process is performed on all time frames t, all frequencies f, and all channel pairs i and j to obtain the inter-channel phase difference characteristics. This characteristic directly reflects the relative time delay of sound waves arriving at different microphones, has a direct physical correspondence with the incident angle of the sound source, and exhibits strong robustness to amplitude distortion in reverberant environments.

[0053] Using the above method, the phase difference features between channels can be extracted from the frequency domain representation of multi-channel audio signals, providing spatial information in the phase domain for subsequent sound source localization.

[0054] Figure 4 This is a schematic diagram of a logarithmic directional beam signal-to-noise ratio feature extraction process provided in an embodiment of this disclosure, based on Figure 2 The illustrated embodiment further explains step 201. Figure 4 This may include the following steps: Step 401: Design beamforming filters for multiple preset directions, and process the multi-channel audio signal based on the beamforming filters to obtain the filtered signals for the multiple preset directions.

[0055] In some embodiments, preset directions are multiple pre-defined spatial orientations used to cover the spatial range where sound sources may exist. The number and specific orientations of the preset directions can be set according to the actual application scenario. Beamforming filters are used to weight multi-channel audio signals to enhance signals from specific directions and suppress signals from other directions. For each preset direction, a corresponding beamforming filter is designed. The design methods for beamforming filters include, but are not limited to, guide vector-based design methods, where the noise-dominant direction... The corresponding beamforming filter weight vector is And must meet ; where argmin is the minimum value operation. The beamforming filter weight vector is to be solved. for The conjugate transpose of . Let be the normalized correlation matrix of the isotropic noise field in the target direction. The direction of noise dominance The steering vector at frequency f This represents the gain control coefficient for scattering noise. for The conjugate transpose of the lower triangular matrix obtained after Cholesky decomposition (i.e., the upper triangular factor), along with three constraints, ensures that the solution is obtained. It can accurately extract signals in the direction where noise dominates. Target direction. The corresponding beamforming filter weight vector Can be adopted with The same solution method can be used, or conventional beamforming filter design methods in this field can be employed. Based on the designed beamforming filter, the multi-channel audio signals are weighted and calculated, resulting in a filtered signal for each preset direction. The filtered signal is in complex form and can reflect the audio signal characteristics of that preset direction. Beamforming technology enhances signals in specific directions and suppresses signals in other directions by spatially filtering multi-channel audio signals.

[0056] Step 402: Determine the ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions, and perform logarithmic transformation on the ratio to obtain the logarithmic directional beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

[0057] In some embodiments, the target direction is any one of a plurality of preset directions, and the filtered signal energy is obtained by taking the square of the modulus of the filtered signal. The ratio of the filtered signal energy in the target direction to the maximum filtered signal energy among the plurality of preset directions is the Directional Signal-to-Noise Ratio (DSNR), which is calculated using the following formula: ,in Indicates the direction of the target The directional signal-to-noise ratio at the time-frequency point (t,f), For the target direction The beamforming filter weight vector at frequency f for The conjugate transpose of . Let f be the multi-channel frequency domain signal matrix of the multi-channel audio signal at the time-frequency point (t,f). The direction dominated by noise. The direction of noise dominance The corresponding beamforming filter weight vector is the conjugate transpose. The logarithmic direction beam SNR characteristic is obtained from the target direction SNR through a logarithmic transformation, calculated using the following formula: ,in Indicates the direction of the target The logarithmic directional beam signal-to-noise ratio characteristic at the time-frequency point (t,f), where log is the logarithmic transform.

[0058] Using the above method, the logarithmic direction beam signal-to-noise ratio characteristics can be extracted from multi-channel audio signals, providing spatial information in the energy domain for subsequent sound source localization.

[0059] Figure 5 This is a schematic diagram of a sound source localization model training process provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, the method includes steps 501-503.

[0060] Step 501: Obtain the training sample set; each training sample in the training sample set includes a multi-channel audio training signal and its corresponding sound source location training label.

[0061] In some embodiments, the training sample set is used to train the sound source localization model, enabling it to learn the mapping relationship between fused features and sound source location. The multi-channel audio training signal can be a signal from an actual microphone array or generated through simulation. To obtain a large amount of labeled training data, simulation generation can be used. During simulation generation, a clean speech signal is first selected from a speech database as the sound source signal; a noise signal is selected from a noise database as environmental interference. The clean speech signal and the noise signal are mixed at a certain signal-to-noise ratio to simulate noisy speech in a real environment. Simultaneously, to simulate the effects of different acoustic environments, reverberation effects can be added to include echo components generated by room reflections in the audio signal; random channel jitter can also be added to simulate the inconsistencies between channels in the microphone array. The mixed signal is the multi-channel audio training signal, and its corresponding sound source location (e.g., the angle of the sound source relative to the microphone array) serves as the sound source location training label for that training sample. Through the above methods, a diverse training sample set containing various acoustic conditions can be obtained, providing rich training data for the sound source localization model.

[0062] Step 502: Extract inter-channel phase difference training features and logarithmic direction beam signal-to-noise ratio training features from the multi-channel audio training signal and fuse them to obtain training fusion features.

[0063] In some embodiments, the process of extracting inter-channel phase difference training features from multi-channel audio training signals is the same as the method of extracting inter-channel phase difference features from actual acquired signals, and can be referred to... Figure 3 and Figure 4 The embodiment shown is described below. After extracting the inter-channel phase difference training features and the logarithmic directional beam signal-to-noise ratio training features, feature fusion is performed in the same way as in step 102, for example, by concatenating the two types of features along the feature dimension to obtain the training fused features.

[0064] Step 503: Input the training fusion features and the sound source location training labels into the initial sound source localization model, and train the initial sound source localization model based on the preset loss function to obtain the pre-trained sound source localization model; wherein, the preset loss function includes the cross-entropy loss function and the supervised contrast loss function.

[0065] In some embodiments, the initial sound source localization model can be a deep learning network model to be trained, whose network structure can be designed according to actual needs (e.g., including convolutional layers, recurrent layers, and fully connected layers). During training, the training fusion features are used as input to the model, and the model outputs the predicted sound source location information. A loss function is constructed based on the difference between the predicted value output by the model and the training labels of the real sound source locations in the training samples. In this embodiment, the preset loss function includes a cross-entropy loss function and a supervised contrastive loss function. The cross-entropy loss function is used to measure the difference between the classification probability predicted by the model and the real label, which can effectively guide the model to improve classification accuracy. The supervised contrastive loss function enhances the model's ability to discriminate features by bringing the feature representations of samples with the same location category closer together and pushing the feature representations of samples with different location categories further apart.

[0066] During model training, Mean Absolute Error (MAE) and Classification Accuracy (ACC) can be used as performance metrics. MAE measures the average deviation between the predicted angle value and the actual angle value. For each training sample, the absolute error between the predicted angle p and the actual angle s needs to consider the cyclic characteristics of angles, i.e., the shortest distance between the two angles on a circle. The MAE is calculated as follows: ,in, Let N be the mean absolute error, N be the total number of samples in the training set, and t be the sample index. Let be the true sound source orientation angle of the t-th sample. Let be the prediction angle of the model for the t-th sample. Classification accuracy measures the proportion of samples correctly predicted by the model, and it is calculated as follows: Here, ACC represents classification accuracy, C represents the number of samples correctly predicted by the model (i.e., the number of samples where the difference between the predicted angle p and the true angle s is less than a preset threshold), and N represents the total number of samples in the training set. These evaluation metrics can be used to monitor the model's training performance and guide the optimization of model parameters.

[0067] By employing the backpropagation algorithm, the classification accuracy is improved based on the cross-entropy loss function, and the mean absolute error is reduced based on the supervised contrastive loss function, thereby continuously improving the model's predictive performance. After training, a pre-trained sound source localization model is obtained, which can be used for sound source location reasoning in subsequent practical applications.

[0068] By using the above method and training data that incorporates the fusion of two types of features, combined with cross-entropy loss and supervised contrastive loss, a sound source localization model with high localization accuracy in noisy and reverberant environments can be trained.

[0069] Figure 6 This is a schematic diagram of a training sample set construction process provided in an embodiment of the present disclosure, based on Figure 5 The illustrated embodiment further explains step 501. Figure 6 This may include the following steps: Step 601: Select a clean speech signal from the speech library and a noise signal from the noise library.

[0070] In some embodiments, the speech library is a pre-built database containing a large amount of clean speech data, where clean speech signals refer to original speech signals that do not contain background noise, reverberation, or other interference. The noise library is a pre-built database containing various types of noise signals, including but not limited to point source noise (such as television noise, music noise, and percussion noise) and scattered noise (such as wind noise and ambient noise from public transportation). When generating training data, one or more clean speech signals are randomly selected from the speech library as sound source signals, and one or more noise signals are randomly selected from the noise library as environmental interference signals. The selection process can be randomized to increase the diversity of training samples.

[0071] Step 602: Mix the clean speech signal with the noise signal based on the random signal-to-noise ratio, and add random reverb and / or random channel jitter to obtain the multi-channel audio training signal.

[0072] In some embodiments, signal-to-noise ratio (SNR) refers to the energy ratio of signal to noise, used to control the relative intensity of clean speech signals and noise signals in the mixed signal. Random SNR refers to an SNR value randomly generated within a preset range (e.g., randomly selected within the range of 0dB to 20dB). The clean speech signal and noise signal are mixed according to the random SNR to obtain a noisy mixed signal. Reverberation refers to the echo effect formed by sound reflecting multiple times through walls, objects, etc., in an enclosed space. Adding random reverberation involves applying a room impulse response to the mixed signal to simulate different reverberation environments (e.g., setting the reverberation time to a random value within the range of 0.3 seconds to 1.2 seconds), so that the training signal contains different degrees of reverberation components. Random channel jitter refers to adding random, small time offsets or amplitude differences between different channels of a multi-channel audio signal to simulate the inconsistencies that may exist between channels in an actual microphone array (e.g., randomly setting jitter at 1 to 2 sampling points). Through the above processing, a noisy multi-channel audio training signal simulating a real acoustic environment is obtained. It should be noted that the order of mixing, adding reverb, and adding channel jitter can be adjusted according to actual needs, and this embodiment does not limit this.

[0073] Step 603: Set corresponding sound source location training labels for the multi-channel audio training signals to obtain the training sample set.

[0074] In some embodiments, the sound source orientation training label refers to the actual sound source direction information corresponding to the multi-channel audio training signal, such as the azimuth angle of the sound source relative to the microphone array. During the simulation generation process, the sound source orientation is a known value set manually when generating the signal; therefore, this azimuth angle can be directly used as the label for the training sample. Each multi-channel audio training signal and its corresponding sound source orientation training label together constitute a training sample. Multiple training samples generated in the above manner are aggregated to form a training sample set. For example, 1.02 million noisy multi-channel audio signals can be generated, of which 1 million are used for model training, 10,000 for validation, and 10,000 for testing.

[0075] The above methods can generate a diverse training sample set containing various acoustic conditions, providing rich training data for the sound source localization model, enabling the model to adapt to different noise and reverberation environments, and improving the model's generalization ability and robustness.

[0076] Figure 7 This is a schematic diagram of a sound source location determination process provided in an embodiment of this disclosure, based on Figure 5 The illustrated embodiment further explains step 103. Figure 7 This may include the following steps: Step 701: Input the fused features into the pre-trained sound source localization model to obtain the sound source prediction probability for each preset direction.

[0077] In some embodiments, the pre-trained sound source localization model has undergone parameter optimization and can output corresponding sound source location prediction results based on the input fusion features. The fusion features extracted and fused from the actual application scenario are input into the model, which performs forward computation and outputs a probability vector. Each element in this probability vector corresponds to a probability value in a preset direction, representing the confidence that the sound source is located in that direction. The preset directions can be multiple pre-defined discrete directions; for example, dividing a 360-degree space into K directions at certain intervals, the model output is the probability distribution of these K directions.

[0078] The probability values ​​output by the model satisfy the condition that they are non-negative and sum to 1. The larger the probability value, the higher the probability that the sound source is located in that direction.

[0079] Step 702: Determine the preset direction with the highest predicted probability of the sound source as the sound source azimuth information.

[0080] In some embodiments, based on the probability distribution of each preset direction output by the model, the preset direction with the highest probability value is selected as the final sound source location information. Specifically, the probability vector output by the model is indexed by its maximum value to find the preset direction corresponding to the maximum probability value, which is the location of the sound source determined by the model. In practical applications, if the probability values ​​of all directions are low or the maximum value is below a preset threshold, it can be determined that the system is silent or there is no effective sound source.

[0081] By using the above method and reasoning about the fused features using a pre-trained sound source localization model, the directional information of the sound source can be accurately determined.

[0082] To verify the performance of the sound source localization method provided in this disclosure, a comparative experiment was conducted. The experiment used 1.02 million noisy multi-channel audio signals as test data, including point source noise (such as television noise, music noise, and percussion noise) and scattered noise (such as wind noise and ambient noise from public transportation). The signal-to-noise ratio ranged from 0 to 20 dB, the reverberation time T60 ranged from 0.3 to 1.2 s, and the microphone spacing was 3.7 cm. Sound source localization tests were performed using the Steered Response Power with Phase Transform (SRP-PHAT) algorithm, a deep learning method based on FFT complex features, and the sound source localization method provided in this disclosure, respectively. Experimental results show that the method provided in this disclosure has a mean absolute error (MAE) of 9.37 degrees and a localization accuracy (ACC) of 82.99%, which is significantly better than the SRP-PHAT method (MAE of 20.15 degrees and ACC of 42.92%) and the deep learning method based on FFT complex features (MAE of 17.91 degrees and ACC of 69.77%). This demonstrates that the method provided in this disclosure has higher localization accuracy and robustness in noisy and reverberant environments.

[0083] In some possible ways, Figure 8 This is a schematic diagram of a sound source localization algorithm provided in an embodiment of this disclosure, as shown below. Figure 8 As shown, features are first extracted and fused from multi-channel audio signals to obtain inter-channel phase difference features and logarithmic beam signal-to-noise ratio features (i.e., fused features of IPD+LDSNR). The fused features are then processed sequentially through three layers of a Convolutional Neural Network (CNN), one layer of a Gated Recurrent Unit (GRU), and one fully connected layer to output the corresponding sound source location information. Specifically, the three CNN layers are used for local feature extraction and abstraction of the fused features, capturing spatial correlation information within the features; the GRU is used for temporal modeling of the features output by the CNN layers, mining the temporal dependencies within the features; and the fully connected layer maps the high-dimensional features output by the GRU to an output dimension consistent with the number of preset directions, obtaining the predicted probability of the sound source in each preset direction, thereby determining the sound source location information.

[0084] In another possible implementation, the sound source localization method provided in this embodiment adopts a dual-branch feature extraction + high-level fusion structure. The specific process is as follows: 1) Spatial phase branch: Input IPD features and extract the spatial phase structure by 2-layer CNN; 2) Directional energy branch: Input LDSNR features and extract the directional energy distribution by 2-layer CNN; 3) High-level fusion: Concatenate the two outputs in the channel dimension and send them to 1-layer GRU to model the temporal dependency; 4) Classification output: After passing through a fully connected layer and Softmax, output the predicted probabilities of K directions.

[0085] Furthermore, a triple composite loss is employed, trained using an angle-aware composite loss. Specifically: 1) Classification cross-entropy loss. 1) Constraint direction classification accuracy; 2) Angle cycle loss 3) Considering 360° periodicity, correct for angle jump error; 4) Neighborhood direction smoothing loss. : Punish mutations in adjacent directions to improve stability.

[0086] Total loss The calculation formula is: ,in , , These are learnable weights.

[0087] Corresponding to the sound source localization method described above, this invention also proposes a sound source localization device. Since the device embodiments of this invention correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to in the method embodiments described above, and will not be repeated here.

[0088] Figure 9 This is a schematic diagram of the structure of a sound source localization device provided in an embodiment of this disclosure, as shown below. Figure 9 As shown, it includes: Extraction unit 81 is used to extract multi-channel spatial features from the acquired multi-channel audio signal; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features; The first fusion unit 82 is used to fuse the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features to obtain fused features. The determining unit 83 is used to input the fused features into a pre-trained sound source localization model to determine the sound source location information.

[0089] The sound source localization device provided in this disclosure can improve the accuracy of sound source localization in noisy and reverberant environments by fusing inter-channel phase difference characteristics and logarithmic direction beam signal-to-noise ratio characteristics.

[0090] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the extraction unit 81 includes: Extraction module 811 is used to extract the inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features from the multi-channel audio signals acquired by the microphone array.

[0091] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the extraction module 811 includes: The first determining submodule 8111 is used to perform time-frequency transformation on the multi-channel audio signal to obtain the frequency domain representation of each channel at the time-frequency point; The second determining submodule 8112 is used to determine the phase difference between any two channels at the same time and frequency point based on the frequency domain representation of any two channels at the same time and frequency point, thereby obtaining the inter-channel phase difference characteristics.

[0092] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the extraction module 811 includes: The third determining submodule 8113 is used to design beamforming filters for multiple preset directions, and process the multi-channel audio signal based on the beamforming filters to obtain the filtered signals for the multiple preset directions. The fourth determining submodule 8114 is used to determine the ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions, and to perform logarithmic transformation on the ratio to obtain the logarithmic direction beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

[0093] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the device further includes: The acquisition unit 84 is used to acquire a training sample set before the determination unit 83 inputs the fused features into the pre-trained sound source localization model to determine the sound source orientation information; each training sample in the training sample set includes a multi-channel audio training signal and its corresponding sound source orientation training label. The second fusion unit 85 is used to extract inter-channel phase difference training features and logarithmic direction beam signal-to-noise ratio training features from the multi-channel audio training signal and fuse them to obtain training fusion features. Training unit 86 is used to input the training fusion features and the sound source location training labels into the initial sound source localization model, and train the initial sound source localization model based on a preset loss function to obtain the pre-trained sound source localization model; wherein, the preset loss function includes a cross-entropy loss function and a supervised contrast loss function.

[0094] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the acquisition unit 84 includes: The selection module 841 is used to select clean speech signals from the speech library and noise signals from the noise library; The first determining module 842 is used to mix the clean speech signal with the noise signal based on the random signal-to-noise ratio, and add random reverberation and / or random channel jitter to obtain the multi-channel audio training signal. The second determining module 843 is used to set corresponding sound source location training labels for the multi-channel audio training signals to obtain the training sample set.

[0095] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 10 As shown, the determining unit 83 includes: The input module 831 is used to input the fused features into the pre-trained sound source localization model to obtain the sound source prediction probability in each preset direction; The third determining module 832 is used to determine the preset direction with the highest prediction probability of the sound source as the sound source azimuth information.

[0096] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of the embodiments of this disclosure, and the principle is the same. Therefore, the embodiments of this disclosure are not limited thereto.

[0097] This invention also provides an electronic device for real-time directional positioning of sounds near the electronic device. For example... Figure 11 As shown, it includes an audio acquisition circuit 001, a processor 002, and a storage device 003. The audio acquisition circuit is responsible for acquiring the raw microphone signal and storing it in the storage device 003. The processor 002 reads the raw microphone signal data in the storage device 003 frame by frame and processes it to implement the various methods and processes described above, such as the sound source localization method.

[0098] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0099] Figure 12A schematic block diagram of an example electronic device 1200 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory 1202 or a computer program loaded from a storage unit 1208 into a random access memory 1203. The random access memory 1203 may also store various programs and data required for the operation of the electronic device 1200. The computing unit 1201, the read-only memory 1202, and the random access memory 1203 are interconnected via a bus 1204. An input / output interface 1205 is also connected to the bus 1204.

[0101] Multiple components in electronic device 1200 are connected to input / output interface 1205, including: input unit 1206, such as keyboard, mouse, etc.; output unit 1207, such as various types of monitors, speakers, etc.; storage unit 1208, such as disk, optical disk, etc.; and communication unit 1209, such as network card, modem, wireless transceiver, etc. Communication unit 1209 allows electronic device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0102] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as sound source localization methods. For example, in some embodiments, the sound source localization method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1200 via read-only memory 1202 and / or communication unit 1209. When the computer program is loaded into random access memory 1203 and executed by the computing unit 1201, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform the aforementioned sound source localization method by any other suitable means (e.g., by means of firmware).

[0103] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0105] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0107] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0108] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0109] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0110] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0111] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for locating a sound source, characterized in that, include: Multi-channel spatial features are extracted from the acquired multi-channel audio signals; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features; The inter-channel phase difference features and the logarithmic beam signal-to-noise ratio features are fused to obtain fused features; wherein, the feature fusion method includes at least dynamically weighting and fusing the inter-channel phase difference features and the logarithmic beam signal-to-noise ratio features through a time-frequency adaptive gating unit; The fused features are input into a pre-trained sound source localization model to determine the sound source location information; The extraction of the inter-channel phase difference features includes: The multi-channel audio signal is subjected to time-frequency transformation to obtain the frequency domain representation of each channel at the time-frequency point; Based on the frequency domain representation of any two channels at the same time and frequency point, determine the phase difference between the two channels at the same time and frequency point to obtain the inter-channel phase difference characteristics; The extraction of the logarithmic direction beam signal-to-noise ratio features includes: Beamforming filters are designed for multiple preset directions, and the multi-channel audio signals are processed based on the beamforming filters to obtain the filtered signals for the multiple preset directions. The ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions is determined, and the ratio is subjected to logarithmic transformation to obtain the logarithmic directional beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

2. The method according to claim 1, characterized in that, The extraction of multi-channel spatial features from the acquired multi-channel audio signals includes: The inter-channel phase difference features and the logarithmic direction beam signal-to-noise ratio features are extracted from the multi-channel audio signals acquired by the microphone array.

3. The method according to claim 1, characterized in that, Before inputting the fused features into a pre-trained sound source localization model to determine the sound source location information, the method further includes: Obtain a training sample set; each training sample in the training sample set includes a multi-channel audio training signal and its corresponding sound source location training label. The inter-channel phase difference training features and logarithmic directional beam signal-to-noise ratio training features are extracted from the multi-channel audio training signal and fused to obtain the training fusion features. The training fusion features and the sound source location training labels are input into the initial sound source localization model, and the initial sound source localization model is trained based on a preset loss function to obtain the pre-trained sound source localization model; wherein, the preset loss function includes a cross-entropy loss function and a supervised contrastive loss function.

4. The method according to claim 3, characterized in that, The acquisition of the training sample set includes: Select clean speech signals from the speech library and select noise signals from the noise library; The clean speech signal is mixed with the noise signal based on the random signal-to-noise ratio, and random reverb and / or random channel jitter are added to obtain the multi-channel audio training signal. Set corresponding sound source location training labels for the multi-channel audio training signals to obtain the training sample set.

5. The method according to claim 3, characterized in that, The step of inputting the fused features into a pre-trained sound source localization model to determine the sound source location information includes: The fused features are input into the pre-trained sound source localization model to obtain the sound source prediction probability for each preset direction; The preset direction with the highest predicted probability of the sound source is determined as the sound source location information.

6. A sound source localization device, characterized in that, include: An extraction unit is used to extract multi-channel spatial features from the acquired multi-channel audio signal; wherein, the multi-channel spatial features include inter-channel phase difference features and logarithmic direction beam signal-to-noise ratio features; The first fusion unit is used to fuse the inter-channel phase difference features and the logarithmic beam signal-to-noise ratio features to obtain fused features; wherein, the feature fusion method includes at least dynamically weighting the inter-channel phase difference features and the logarithmic beam signal-to-noise ratio features through a time-frequency adaptive gating unit. The determining unit is used to input the fused features into a pre-trained sound source localization model to determine the sound source location information; The extraction unit includes: The first determining submodule is used to perform time-frequency transformation on the multi-channel audio signal to obtain the frequency domain representation of each channel at the time-frequency point; The second determining submodule is used to determine the phase difference between any two channels at the same time and frequency point based on the frequency domain representation of any two channels at the same time and frequency point, thereby obtaining the inter-channel phase difference characteristics; The third determining submodule is used to design beamforming filters for multiple preset directions, and process the multi-channel audio signal based on the beamforming filters to obtain the filtered signals for the multiple preset directions. The fourth determining submodule is used to determine the ratio of the filtered signal energy in the target direction to the maximum filtered signal energy in the plurality of preset directions, and to perform logarithmic transformation on the ratio to obtain the logarithmic directional beam signal-to-noise ratio characteristics; wherein, the filtered signal energy is the square of the magnitude of the filtered signal.

7. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Sound source positioning method and device and air conditioner

    CN107271963A

  • Method for sound source direction estimation based on time frequency masking and deep neural network

    CN109839612A