Sound source positioning method and device, equipment and medium
By processing multi-channel microphone signals and utilizing the natural properties of noise and reverberation, the accuracy of sound source localization is improved, the problem of direct path signals being affected by noise and reverberation is solved, and more accurate sound source localization and sound separation are achieved.
Patent Information
- Application Number
- CN202410525855.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-10-28
AI Technical Summary
In the prior art, sound source localization is performed because direct path signals are susceptible to noise and reverberation contamination, resulting in inaccurate feature estimation, which in turn affects the accuracy of sound source localization.
By processing multi-channel microphone signals, the time difference of arrival (TDOA) of the direct path propagation between the sound source and different microphones is determined. Combined with predictive localization features and reference localization features, the accuracy of sound source localization is improved.
It effectively suppresses the spatial diffusion of noise and reverberation, improves the accuracy of sound source localization, and supports real-time tracking and sound separation.
Smart Images

Figure CN120857033A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of electronic equipment technology, and in particular to a sound source localization method, apparatus, device, and medium. Background Technology
[0002] Sound source localization refers to using multi-channel microphone signals to estimate the location of a sound source. It is widely used in scenarios such as video conferencing, robot auditions, and live video streaming. For example, during live video streaming, it can automatically locate the speaker's position and voice, and can be combined with other functions to improve the performance of speech enhancement and separation.
[0003] In related technologies, sound source localization is performed by estimating the characteristics of the direct path signal and based on the estimated characteristics.
[0004] In this approach, the direct path signal is easily contaminated by noise and reverberation in the real environment, resulting in inaccurate feature estimation and thus inaccurate sound source localization. Summary of the Invention
[0005] This disclosure aims to at least partially address one of the technical problems in the related art.
[0006] Therefore, this disclosure provides a sound source localization method, apparatus, electronic device, computer-readable storage medium, and computer program product to improve the accuracy of sound source localization.
[0007] To achieve the above objectives, a first aspect of this disclosure provides a sound source localization method, comprising: acquiring sound from a sound source to obtain multi-channel microphone signals; determining predictive localization features based on the multi-channel microphone signals, wherein the predictive localization features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source and different microphones; determining reference localization features corresponding to each candidate direction; and determining a target predictive feature from multiple reference localization features based on the predictive localization features, and using the candidate direction corresponding to the target predictive feature as the localization result of the sound source.
[0008] To achieve the above objectives, a second aspect of this disclosure provides a sound source localization device, comprising: a data acquisition module for acquiring sound from a sound source under test to obtain multi-channel microphone signals; a first determination module for determining a predicted localization feature based on the multi-channel microphone signals, wherein the predicted localization feature represents the time difference of arrival (TDOA) of the direct path propagation between the sound source under test and different microphones; a second determination module for determining a reference localization feature corresponding to each candidate direction; and a localization module for determining a target predicted feature from multiple reference localization features based on the predicted localization feature, and using the candidate direction corresponding to the target predicted feature as the localization result of the sound source under test.
[0009] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the sound source localization method proposed in the first aspect of this disclosure.
[0010] To achieve the above objectives, a fourth aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the sound source localization method proposed in the first aspect of this disclosure.
[0011] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the sound source localization method proposed in the first aspect of this disclosure.
[0012] The sound source localization method, apparatus, electronic device, computer-readable storage medium, and computer program product disclosed herein acquire sound from a sound source under test to obtain multi-channel microphone signals. Based on the multi-channel microphone signals, predictive localization features are determined, wherein the predictive localization features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source under test and different microphones. Reference localization features corresponding to each candidate direction are determined, and a target predictive feature is determined from multiple reference localization features based on the predictive localization features. The candidate direction corresponding to the target predictive feature is used as the localization result of the sound source under test. Because the multi-channel microphone signals are directly processed to estimate the predictive localization features, direct processing of the microphone signals can more effectively utilize the natural properties of noise and reverberation to suppress the spatial diffusion of noise and post-reverberation, thereby estimating more accurate predictive localization features. When using these predictive localization features to locate the sound source under test, the accuracy of sound source localization can be effectively improved.
[0013] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, wherein:
[0015] Figure 1 This is a schematic flowchart illustrating a sound source localization method provided in an embodiment of the present disclosure.
[0016] Figure 2 This is a schematic flowchart illustrating another sound source localization method provided in an embodiment of this disclosure;
[0017] Figure 3 This is a schematic diagram of the structure of the localization feature prediction model in this embodiment of the present disclosure;
[0018] Figure 4 This is a schematic diagram of the structure of a sound source localization device provided in an embodiment of the present disclosure;
[0019] Figure 5 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0020] Some embodiments of this disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.
[0021] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0022] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device. Figure 1 This is a flowchart illustrating a sound source localization method provided in an embodiment of the present disclosure.
[0023] This disclosure provides a sound source localization method to improve the accuracy of sound source localization.
[0024] The application scenarios of the sound source localization method provided in the embodiments of this disclosure are illustrated below with examples:
[0025] (1) In mobile devices, such as during a call or video call, the location of the other party can be determined by the sound source localization method, thereby improving call quality and user experience.
[0026] (2) In a meeting setting, the speaker's voice can be automatically captured, such as by locating the speaker's position, combining speech recognition and enhancement technologies, or by directly separating the speaker's voice to improve the speaker's voice quality and the meeting experience.
[0027] (3) In the live streaming scenario, the voice and location information of the host can be automatically located and the voice of the host can be separated.
[0028] (4) In video recording scenarios, speech enhancement can be performed after the speaker's voice is located, or other sound effects processing methods can be used for subsequent processing, such as increasing the speaker's volume.
[0029] Figure 1 This is a flowchart illustrating a sound source localization method provided in an embodiment of the present disclosure.
[0030] This embodiment illustrates the use of a sound source localization method configured in a sound source localization device. In this embodiment, the sound source localization method can be configured in a sound source localization device, which can be located in a server or an electronic device, without limitation.
[0031] This embodiment uses the example of a sound source localization method configured in an electronic device. The electronic device includes hardware devices with various operating systems, such as smartphones, in-vehicle systems, tablets, personal digital assistants, and e-readers.
[0032] It should be noted that the execution entity of the embodiments disclosed herein may be, in hardware, a central processing unit (CPU) in a server or electronic device, and in software, a related background service in a server or electronic device, without limitation.
[0033] like Figure 1 As shown, the sound source localization method includes the following steps:
[0034] Step S101: Collect sound from the sound source under test to obtain multi-channel microphone signals.
[0035] The sound source to be located can be referred to as the sound source being tested. The sound source being tested can be, for example, a streamer in a live broadcast scenario; there are no restrictions on this.
[0036] In some embodiments, multi-channel microphone signals can be acquired, such as using an electronic device with multiple microphones to capture the voice of a broadcaster. The acquired microphone signals can be referred to as multi-channel microphone signals.
[0037] Step S102: Determine the predicted localization features based on the multi-channel microphone signals, wherein the predicted localization features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source being measured and different microphones.
[0038] Among them, the features used to predict the localization results of sound sources can be called predictive localization features.
[0039] The predictive localization feature used in this embodiment can be obtained from the analysis of multi-channel microphone signals. This predictive localization feature represents the time difference of arrival (TDOA) of the direct path propagation between the sound source being measured and different microphones. This predictive localization feature can also be referred to as the inter-channel phase difference (IPD).
[0040] Because it directly processes multi-channel microphone signals to estimate predictive localization features, direct processing of microphone signals can more effectively utilize the natural properties of noise and reverberation to suppress the spatial diffusion of noise and post-reverberation, thereby estimating more accurate predictive localization features. When using these predictive localization features to locate the sound source under test, the accuracy of sound source localization can be effectively improved.
[0041] When the sound source localization method provided in this embodiment is applied to video conferencing or live streaming scenarios, it can identify and track the speaker's audio source. It can track the speaker's audio smoothly in real time and supports real-time sound separation.
[0042] In some embodiments, when determining the predicted localization features based on multi-channel microphone signals, signal analysis can be performed on the multi-channel microphone signals, and the time difference of arrival (TDOA) of the direct path propagation between the sound source under test and different microphones can be determined based on the signal analysis results; or acoustic modeling can be used to process the multi-channel microphone signals to determine the predicted localization features; or any other possible method can be used to determine the predicted localization features based on multi-channel microphone signals, without limitation.
[0043] Step S103: Determine the reference positioning features corresponding to each candidate direction.
[0044] Candidate directions can be determined by sampling in space. For example, candidate directions can be sampled in the entire positioning space beforehand to obtain multiple candidate directions, and reference positioning features can be calculated for each candidate direction.
[0045] The reference localization feature is the time difference of arrival (TDOA) of the direct path propagation between the tested sound source and different microphones when it is in the candidate direction. This reference localization feature can be pre-calibrated based on experiments. It can be used to compare with the predicted localization feature to determine the direction of the tested sound source, which is then used as the localization result of the tested sound source.
[0046] Step S104: Determine the target prediction feature from multiple reference positioning features based on the predicted positioning features, and use the candidate direction corresponding to the target prediction feature as the positioning result of the sound source being measured.
[0047] In some embodiments, when determining the predicted localization features and the reference localization features in each candidate direction, the target predicted feature can be determined from multiple reference localization features based on the predicted localization features, and the candidate direction corresponding to the target predicted feature can be used as the localization result of the sound source being measured.
[0048] For example, a reference localization feature that matches the predicted localization feature can be selected from multiple reference localization features, and the candidate direction corresponding to the matched reference localization feature can be used as the direction of the sound source being measured. This candidate direction can be directly used as the localization result of the sound source being measured.
[0049] In some embodiments, the process of determining the target predicted feature from multiple reference localization features based on the predicted localization feature may involve determining the similarity between the predicted localization feature and each reference localization feature, determining the maximum similarity among the multiple similarities, and using the reference localization feature corresponding to the maximum similarity as the target predicted feature. This enables rapid and efficient identification of the target predicted feature from multiple reference localization features, thereby supporting improvements in the efficiency of determining the localization results of the measured sound source.
[0050] For example, the inner product of the predicted IPD vector (an optional example of the predicted localization feature) and the IPD vector of the candidate direction (an optional example of the reference localization feature) can be calculated, and the candidate direction with the largest inner product (the candidate direction corresponding to the target predicted feature) can be used as the localization result of the sound source being measured.
[0051] In this embodiment, sound is collected from the sound source under test to obtain multi-channel microphone signals. Based on these signals, predictive localization features are determined. These predictive localization features represent the Time Difference of Arrival (TDOA) of the direct path propagation between the sound source and different microphones. Reference localization features are determined for each candidate direction, and a target predictive feature is selected from these reference features based on the predictive localization features. The candidate direction corresponding to the target predictive feature is then used as the localization result of the sound source under test. Since the multi-channel microphone signals are processed directly to estimate the predictive localization features, this direct processing of the microphone signals can more effectively utilize the natural properties of noise and reverberation to suppress spatial diffusion and post-reverberation, thereby estimating more accurate predictive localization features. When these predictive localization features are used to locate the sound source under test, the accuracy of sound source localization is effectively improved.
[0052] The sound source localization method provided in this embodiment can, after determining the localization result of the sound source under test, further extract the target audio signal corresponding to the localization result of the sound source under test from the multi-channel microphone signal. This supports improving the accuracy of sound separation of the sound source under test.
[0053] For example, the target audio signal can be separated from the multi-channel microphone signal by referring to the localization results of the sound source being tested and using any possible sound separation algorithm.
[0054] For example, the localization results of the sound source under test can be referenced to identify the sound signal characteristics in the direction of the sound source under test from the multi-channel microphone signal, and the sound signal characteristics can be processed to obtain the target audio signal.
[0055] Figure 2 This is a schematic flowchart illustrating another sound source localization method provided in an embodiment of this disclosure.
[0056] like Figure 2 As shown, the sound source localization method includes the following steps:
[0057] Step S201: Collect sound from the sound source under test to obtain multi-channel microphone signals.
[0058] For a detailed description of step S201, please refer to the above embodiments, which will not be repeated here.
[0059] Step S202: Perform frequency domain transformation processing on the multi-channel microphone signal to obtain a multi-channel frequency domain signal.
[0060] In some embodiments, the multi-channel microphone signal can be processed by Short-Time Fourier Transform (STFT) to obtain a multi-channel frequency domain signal in the STFT domain.
[0061] For example, suppose a sound source moving along the θ direction is observed in an environment using two microphones. The two-channel microphone signal can be defined in the STFT domain as follows:
[0062] N mics (t,d)=Fun(t,d) θ • Speech + Noise;
[0063] Where mics represents the number of microphones, t represents the time dimension index, and d represents the frequency dimension index. N, Speech, and Noise are the STFT coefficients of the multi-channel frequency domain signal, source signal, and noise signal, respectively. θ This is the room transfer function, which is the Fourier transform of the room impulse response. The room transfer function can be time-varying or time-dependent. The room transfer function can be decomposed into:
[0064] Fun(t,d) θ =F room +F reverb ;
[0065] Among them, F room and F reverb These represent the direct response and partial reverberation response of the multi-channel microphone signal, respectively.
[0066] Step S203: Input the multi-channel frequency domain signal into the positioning feature prediction model and obtain the predicted positioning features output by the positioning feature prediction model, wherein the predicted positioning features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source being measured and different microphones.
[0067] After obtaining the multi-channel frequency domain signal, the multi-channel frequency domain signal can be input into the pre-trained positioning feature prediction model, and the positioning feature prediction model can be used to model, analyze and predict positioning features.
[0068] The location feature prediction model is used to predict location features based on multi-channel frequency domain signals. This model can be any type of artificial intelligence, such as a neural network model or a machine learning model; there are no restrictions on its application.
[0069] In some embodiments, the location feature prediction model can be pre-trained, and the model has already modeled and analyzed the mapping relationship between multi-channel frequency domain signals and predicted location features. Therefore, this location feature prediction model can support the prediction of more accurate location features.
[0070] In some embodiments, when inputting multi-channel frequency domain signals into a positioning feature prediction model, in order to effectively enrich the frequency domain features, a first frequency domain portion and a second frequency domain portion of the multi-channel frequency domain signal can be determined, wherein the first frequency domain portion represents the real part of the frequency domain coefficients of the multi-channel microphone signal, and the second frequency domain portion represents the imaginary part of the frequency domain coefficients of the multi-channel microphone signal, and the first frequency domain portion and the second frequency domain portion are input into the positioning feature prediction model.
[0071] For example, the real part (an optional example of the first frequency domain part) and the imaginary part (an optional example of the second frequency domain part) of the STFT coefficients of the multi-channel frequency domain signal can be used as inputs to the localization feature prediction model. Therefore, the number of input channels C of the localization feature prediction model is twice the number of microphones.
[0072] In some embodiments, the localization feature prediction model includes a first feature extraction path and a second feature extraction path; wherein, in the first feature extraction path, there are multiple first full-band sub-models and multiple first narrow-band sub-models, which are alternately connected, and the first sub-model in the first feature extraction path is a first full-band sub-model; in the second feature extraction path, there are multiple second full-band sub-models and multiple second narrow-band models, which are alternately connected, and the first sub-model in the second feature extraction path is a second narrow-band model; wherein, the first full-band sub-model and the second full-band sub-model are respectively based on extracting full-band correlation features related to the predicted localization features at different granularities, and the first narrow-band sub-model and the second narrow-band sub-model are respectively based on extracting narrow-band features related to the predicted localization features at different granularities.
[0073] The first full-band sub-model is used to extract full-band correlation features related to the predicted localization features based on one granularity. The second full-band model is used to extract full-band correlation features related to the predicted localization features based on another granularity. The different full-band models use different granularities.
[0074] The first narrowband sub-model is used to extract narrowband features related to the predicted localization features based on one granularity. The second narrowband sub-model is used to extract narrowband features related to the predicted localization features based on another granularity. The different narrowband sub-models use different granularities.
[0075] In other words, embodiments of this disclosure support learning full-band and narrow-band features related to IPD by using alternating full-band and narrow-band sub-models, respectively. This enables more accurate IPD predictions.
[0076] like Figure 3 As shown, Figure 3This is a schematic diagram of the structure of the localization feature prediction model in this embodiment of the present disclosure, including a first feature extraction path 31 and a second feature extraction path 32. The first feature extraction path 31 includes a first full-band sub-model 311 and a first narrow-band model 312, and the second feature extraction path 32 includes a second narrow-band model 321 and a second full-band model 322. There is also an average pooling sub-model 33, a fully connected sub-model 34, and a LayerNorm module 35. The LayerNorm module 35 is used to normalize the input in each layer of the localization feature prediction model.
[0077] exist Figure 3 In this model, the input is processed by alternating full-band GRU (an optional example of a full-band sub-model) and narrow-band GRU (an optional example of a narrow-band sub-model), with parallel full-band and narrow-band processing enabling the extraction of features at different granularities. More modules can be added by repeating the second block (in...). Figure 3 (omitted). The outputs of the full-band and narrow-band sub-models are passed to the average pooling sub-model to compress the frame rate, and then fed into a fully connected (FC) layer (an optional example of the fully connected sub-model) to transform it into the desired output dimension. For the full-band GRU, the full-band GRU processes time frames independently, and all time frames share the same network parameters. The multi-channel frequency domain signal is first input to the first full-band sub-model, while for other full-band sub-models, its input includes the output of the previous narrow-band sub-model. The full-band GRU focuses on learning the inter-frequency dependencies of spatial / localization cues (an optional example of predicting localization features). IPDs at the same frequency are highly correlated because they all originate from the same TDOA. Furthermore, with the help of other frequencies, spatial cues with low direct path energy (an optional example of predicting localization features) can be predicted more accurately. The full-band sub-model does not learn any temporal information; the narrow-band sub-model learns the temporal information. For the narrow-band GRU, the narrow-band GRU processes frequencies independently, and all frequencies share the same network parameters. The narrow-band GRU focuses on utilizing narrow-band inter-channel information. Furthermore, IPD is time-varying for moving sound sources, and the narrowband sub-model also learns the temporal evolution of IPD. To ensure the effective transmission of full-band and narrowband information, a skip connection is established, and a LayerNorm module 35 is added for smoothing. The full-band and narrowband sub-models focus on their respective information, and adding skip connections effectively avoids information loss. The output of the second narrowband sub-model can be used together with the output of the first full-band sub-model as the input of the first narrowband sub-model. Then, the output of the previous full-band sub-model (or narrowband model) is added to the input of the next full-band model (or narrowband model).
[0078] In some embodiments, it can be as follows Figure 3As shown, the input to the first sub-model is a multi-channel frequency domain signal; specifically, the input to the first full-band sub-model includes the output of the preceding first narrow-band sub-model and the output of the second full-band sub-model; the input to the second full-band sub-model includes the output of the preceding second narrow-band sub-model and the output of the first full-band sub-model; the input to the first narrow-band sub-model includes the output of the preceding first full-band sub-model and the output of the second narrow-band sub-model; and the input to the second narrow-band sub-model includes the output of the preceding second full-band sub-model and the output of the first narrow-band sub-model. This effectively avoids information loss and supports improving the prediction accuracy of full-band correlation features and narrow-band features.
[0079] In some embodiments, the localization feature prediction model further includes: an average pooling sub-model connected to the first feature extraction path and the second feature extraction path respectively, and a fully connected sub-model connected to the average pooling sub-module. The average pooling sub-model is used to process the target narrowband features output by the first feature extraction path and the target fullband correlation features output by the second feature extraction path to obtain fused features. The fully connected sub-model is used to process the fused features to obtain predicted localization features.
[0080] In this process, the last sub-model in the first feature extraction path can be the first narrowband sub-model, and the last sub-model in the second feature extraction path can be the second full-band sub-model. Therefore, the features output by the first feature extraction path can be the target narrowband features, and the features output by the second feature extraction path can be the target full-band features.
[0081] In some embodiments, the target narrowband features of the output of the first feature extraction path and the target full-band correlation features of the output of the second feature extraction path can be processed by the average pooling sub-model. The resulting features can be called fused features, which can be used to estimate the IPD of the multi-channel microphone signal.
[0082] In some embodiments, the fusion features output by the average pooling sub-model can be provided to the fully connected sub-model, which processes the fusion features to obtain the IPD of the multi-channel microphone signal.
[0083] Step S204: Determine the reference positioning features corresponding to each candidate direction.
[0084] Step S205: Determine the target prediction feature from multiple reference positioning features based on the predicted positioning features, and use the candidate direction corresponding to the target prediction feature as the positioning result of the sound source being measured.
[0085] For a detailed description of steps S204-S205, please refer to the above embodiments, which will not be repeated here.
[0086] In this embodiment, sound is collected from the sound source under test to obtain multi-channel microphone signals. Based on these signals, predictive localization features are determined. These predictive localization features represent the Time Difference of Arrival (TDOA) of the direct path propagation between the sound source and different microphones. Reference localization features are determined for each candidate direction, and a target predictive feature is determined from these reference features. The candidate direction corresponding to the target predictive feature is then used as the localization result of the sound source under test. Since the multi-channel microphone signals are directly processed to estimate the predictive localization features, this direct processing allows for more effective utilization of the natural properties of noise and reverberation to suppress spatial diffusion and post-reverberation, resulting in more accurate predictive localization features. When these predictive localization features are used to locate the sound source under test, the accuracy of sound source localization is significantly improved.
[0087] The sound source localization method provided in this disclosure proposes a fusion technique based on full-band and narrow-band signals. It establishes a fusion mechanism for different feature bands (full-band and narrow-band) through a neural network to estimate the inter-channel phase difference (IPD) of the direct path of multi-channel microphone signals. Alternating full-band and narrow-band sub-models are used to learn full-band correlation features and narrow-band features related to IPD, respectively. This method fully utilizes both narrow-band and full-band information from multi-channel microphone signals. Multi-channel microphone signals (such as those from a three- or four-microphone mobile phone) can be used as input to the model to achieve real-time frame-by-frame prediction of IPD. The predicted IPD can be used to locate the current speaker's position. This method offers high accuracy and efficiency in localization.
[0088] The full-band and narrow-band fusion network (an optional example of a localization feature prediction model) proposed in this disclosure uses dedicated gated recurrent units (GRUs) for full-band and narrow-band processing, respectively. Full-band correlation features and narrow-band features are processed separately for each time frame, and the same network parameters can be shared across all time frames. Full-band correlation features and narrow-band features of IPD are learned by cascading and alternating full-band and narrow-band sub-models, respectively. A full-band GRU layer (an optional example of a full-band sub-model) and a narrow-band GRU layer (an optional example of a narrow-band sub-model) are cascaded to predict accurate IPD.
[0089] In this embodiment of the disclosure, an IPD can be used for sound source localization. The multi-channel frequency domain signal N corresponding to the dual-channel microphone signal can be... mics (t,d) is taken as input, and the predicted IPD is output.
[0090] The data simulation and basic configuration of the sound source localization method provided in the embodiments of this disclosure can be illustrated by the following examples:
[0091] Consider using two microphones to determine the Direction of Arrival (DOA) at a 180° azimuth angle. The two-channel microphone signals are obtained by convolving the Room Impulse Response (RIR) with the source speech signal. In this embodiment, data from an open-source database (e.g., the LirbriSpeech corpus) can be used to train, develop, and test the localization feature prediction model, with the room reverberation time (RT60) randomly set within the range of [0.1 sec, 1.5 sec]. The room dimensions are randomly set within the range of 8 × 8 × 3 m to 12 × 10 × 6 m. The movement trajectory of the speech source is randomly generated, with each trajectory having a fixed height. Two microphones are randomly placed 10 cm apart in the room, at the same horizontal plane as the sound source. A generated diffused noise signal is added to the clean sensor signal based on a signal-to-noise ratio randomly selected from -5 dB to 15 dB. All data are sampled at a rate of 16 kHz (kilohertz). The STFT window length is 512 samples (20ms), the frame shift is 256 samples (10ms), and a Hanning window is used. The audio length used for training is 4 seconds. The real and imaginary parts of the STFT coefficients are concatenated as the network input, and the number of input channels is C = 2M, where M = 2 or 3, representing the number of microphones. The output dimension of each GRU layer is set to 128. A total of 8 layers are set for the full-band and narrow-band sub-models, each consisting of a single-layer 1D convolution + GRU + a single-layer 1D convolution structure. During training, the batch size is set to 32. The initial learning rate is set to 0.001, and the learning rate decays exponentially with a decay factor of 0.95. The model is trained for 50 epochs (50 epochs refers to 50 forward and backward propagations of the entire training dataset through the network during neural network training. Each epoch includes feeding the entire training dataset into the neural network once and updating the parameters based on the difference between the network output and the true label).
[0092] Figure 4 This is a schematic diagram of the structure of a sound source localization device provided in an embodiment of this disclosure.
[0093] like Figure 4 As shown, the sound source locating device 40 includes:
[0094] The acquisition module 401 is used to acquire sound from the sound source under test and obtain multi-channel microphone signals.
[0095] The first determining module 402 is used to determine the predicted localization features based on the multi-channel microphone signals, wherein the predicted localization features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source being measured and different microphones.
[0096] The second determining module 403 is used to determine the reference positioning features corresponding to each candidate direction.
[0097] The localization module 404 is used to determine the target prediction feature from multiple reference localization features based on the predicted localization features, and to use the candidate direction corresponding to the target prediction feature as the localization result of the sound source being measured.
[0098] It should be noted that the foregoing explanation of the sound source localization method also applies to the sound source localization device of this embodiment, and will not be repeated here.
[0099] In this embodiment, sound is collected from the sound source under test to obtain multi-channel microphone signals. Based on these signals, predictive localization features are determined. These predictive localization features represent the Time Difference of Arrival (TDOA) of the direct path propagation between the sound source and different microphones. Reference localization features are determined for each candidate direction, and a target predictive feature is selected from these reference features based on the predictive localization features. The candidate direction corresponding to the target predictive feature is then used as the localization result of the sound source under test. Since the multi-channel microphone signals are processed directly to estimate the predictive localization features, this direct processing of the microphone signals can more effectively utilize the natural properties of noise and reverberation to suppress spatial diffusion and post-reverberation, thereby estimating more accurate predictive localization features. When these predictive localization features are used to locate the sound source under test, the accuracy of sound source localization is effectively improved.
[0100] Figure 5 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown. Figure 5 The electronic device 12 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0101] like Figure 5 As shown, the electronic device 12 is represented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, memory 28, and bus 18 connecting different system components (including memory 28 and processing unit 16).
[0102] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0103] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.
[0104] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 5 Not shown; usually referred to as a "hard drive".
[0105] although Figure 5 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.
[0106] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.
[0107] Electronic device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable human interaction with electronic device 12, and / or with any device that enables electronic device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, electronic device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of electronic device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0108] The processing unit 16 executes various functional applications and data processing by running programs stored in the memory 28, such as implementing the sound source localization method mentioned in the foregoing embodiments.
[0109] To implement the above embodiments, this disclosure also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0110] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0111] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0112] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this disclosure can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented in hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this disclosure.
[0113] Furthermore, the term “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term “exemplary” is intended to present the concept in a concrete manner. As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clear from the context, “X applies A or B” is intended to mean any of the natural inclusive arrangements. That is, “X applies A or B” satisfies any of the foregoing instances if X applies A; X applies B; or both X applies A and B. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles “a” and “an” as used in this application and the appended claims are generally understood to mean “one or more.”
[0114] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term “including.”
[0115] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0116] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for locating a sound source, characterized in that, The method includes the following steps: Sound is collected from the sound source under test to obtain multi-channel microphone signals; Based on the multi-channel microphone signals, predictive localization features are determined, wherein the predictive localization features represent the time difference of arrival (TDOA) of the direct path propagation between the sound source under test and different microphones; Determine the reference localization features corresponding to each candidate orientation; and Based on the predicted localization features, a target predicted feature is determined from a plurality of reference localization features, and the candidate direction corresponding to the target predicted feature is used as the localization result of the sound source being measured.
2. The method according to claim 1, characterized in that, The step of determining the target prediction feature from a plurality of reference positioning features based on the predicted positioning feature includes: Determine the similarity between the predicted localization feature and each of the reference localization features; Determine the maximum similarity among the multiple similarities; The reference positioning feature corresponding to the maximum similarity is used as the target prediction feature.
3. The method according to claim 1, characterized in that, The step of determining the predicted localization features based on the multi-channel microphone signals includes: The multi-channel microphone signal is subjected to frequency domain transformation processing to obtain a multi-channel frequency domain signal; The multi-channel frequency domain signal is input into the positioning feature prediction model, and the predicted positioning feature output by the positioning feature prediction model is obtained. The positioning feature prediction model has modeled and analyzed the mapping relationship between the multi-channel frequency domain signal and the predicted positioning feature.
4. The method according to claim 3, characterized in that, The step of inputting the multi-channel frequency domain signal into the positioning feature prediction model includes: A first frequency domain portion and a second frequency domain portion of the multi-channel frequency domain signal are determined, wherein the first frequency domain portion represents the real part of the frequency domain coefficients of the multi-channel microphone signal, and the second frequency domain portion represents the imaginary part of the frequency domain coefficients of the multi-channel microphone signal; The first frequency domain portion and the second frequency domain portion are input into the localization feature prediction model.
5. The method according to claim 3, characterized in that, The localization feature prediction model includes a first feature extraction path and a second feature extraction path; wherein... In the first feature extraction path, there are multiple first full-band sub-models and multiple first narrow-band sub-models, and the first full-band sub-models and the first narrow-band sub-models are connected alternately. The first sub-model in the first feature extraction path is the first full-band model. In the second feature extraction path, there are multiple second full-band sub-models and multiple second narrow-band sub-models, and the second narrow-band sub-models and the second full-band models are connected alternately. The first sub-model in the second feature extraction path is the second narrow-band model. Specifically, the first full-band sub-model and the second full-band sub-model extract full-band correlation features related to the predicted localization features based on different granularities, and the first narrow-band sub-model and the second narrow-band sub-model extract narrow-band features related to the predicted localization features based on different granularities.
6. The method according to claim 5, characterized in that, The input to the first sub-model is a multi-channel frequency domain signal; where, The inputs of the first full-band sub-model include: the output of the previous first narrow-band sub-model and the output of the second full-band sub-model; The input to the second full-band sub-model includes: the output of the previous second narrow-band sub-model and the output of the first full-band sub-model; The inputs to the first narrowband sub-model include: the output of the previous first full-band sub-model and the output of the second narrowband sub-model; The inputs to the second narrowband sub-model include the output of the previous second full-band sub-model and the output of the first narrowband sub-model.
7. The method according to claim 5, characterized in that, The localization feature prediction model further includes: an average pooling sub-model connected to the first feature extraction path and the second feature extraction path respectively, and a fully connected sub-model connected to the average pooling sub-model. The average pooling sub-model is used to process the target narrowband features output by the first feature extraction path and the target fullband correlation features output by the second feature extraction path to obtain fused features. The fully connected sub-model is used to process the fused features to obtain the predicted localization features.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: Extract the target audio signal corresponding to the localization result of the sound source under test from the multi-channel microphone signal.
9. A sound source localization device, characterized in that, The device includes: The acquisition module is used to acquire sound from the sound source under test and obtain multi-channel microphone signals; The first determining module is used to determine the predicted localization feature based on the multi-channel microphone signal, wherein the predicted localization feature represents the time difference of arrival (TDOA) of the direct path propagation between the sound source under test and different microphones; The second determining module is used to determine the reference positioning features corresponding to each candidate direction; and The positioning module is used to determine the target prediction feature from a plurality of reference positioning features based on the predicted positioning feature, and to take the candidate direction corresponding to the target prediction feature as the positioning result of the sound source being measured.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-8.