Audio source positioning method and device, electronic equipment and storage medium
By extracting the characteristics of audio signals and motion sensor data, and training the positioning model, the problem of low accuracy in traditional audio source positioning methods in complex environments is solved, and more accurate audio source positioning is achieved.
Patent Information
- Application Number
- CN202510196865.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional audio source positioning methods are difficult to accurately capture the position information of audio sources in complex environments, and the positioning accuracy is low.
By extracting audio features, such as the time difference, frequency and amplitude of sound arrival from the audio signal, and combining the motion characteristics acquired by the motion sensor, the positioning model is trained to achieve accurate positioning of the audio source.
It improves the accuracy of audio source positioning, and can adjust the positioning results in real time in complex scenarios, making up for the deviations of traditional methods in dynamic scenarios.
Smart Images

Figure CN120085255A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of signal processing, and more specifically, relates to an audio source localization method and apparatus, an electronic device, and a storage medium. Background Art
[0002] Audio source localization technology has wide applications in the fields of communication, virtual reality, and smart home. For example, in video conferencing and multimedia application scenarios, audio source localization can accurately focus on the position of the speaker, thereby achieving better sound pickup and directional audio transmission, improving the quality of voice communication, and reducing the interference of ambient noise; in the smart home field, by locating audio sources such as the cry of a baby or the sound of a smoke alarm, the home status can be better monitored and automatically controlled, providing users with a more intelligent and convenient living experience.
[0003] Traditional audio source localization methods, such as Time Difference of Arrival (TDOA) and beamforming methods, have many limitations and are difficult to accurately capture the position information of the audio source in a complex environment, resulting in low localization accuracy. Summary of the Invention
[0004] The purpose of the present disclosure is to provide an audio source localization method and apparatus, an electronic device, and a storage medium to improve the localization accuracy of the audio source.
[0005] In a first aspect of an embodiment of the present disclosure, an audio source localization method is provided, including: Extracting audio features from the audio signals sent by the audio source localization device; the audio features include the time difference of arrival of sound, frequency, and amplitude; the audio source localization device includes an array microphone, and the time difference of arrival of sound is the time difference of the audio source signal arriving at each microphone in the array microphone; Extracting the motion features of the audio source localization device from the operation data sent by the motion sensor; the motion features include speed, acceleration, and direction; Training a localization model based on the audio features and the motion features to perform the localization of the audio source based on the trained localization model.
[0006] In a second aspect of an embodiment of the present disclosure, an audio source localization apparatus is provided, including: A first feature extraction module for extracting audio features from the audio signals sent by the audio source localization device; the audio features include the time difference of arrival of sound, frequency, and amplitude; the audio source localization device includes an array microphone, and the time difference of arrival of sound is the time difference of the audio source signal arriving at each microphone in the array microphone; The second feature extraction module is configured to extract the motion features of the audio source localization device from the operation data sent by the motion sensor; the motion features include speed, acceleration, and direction; The audio source localization module is configured to train a localization model based on the audio features and the motion features, and perform audio source localization based on the trained localization model.
[0007] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above audio source localization method are implemented.
[0008] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above audio source localization method are implemented.
[0009] The beneficial effects of the audio source localization method, device, electronic device, and storage medium provided by the embodiments of the present disclosure are as follows: In the embodiments of the present disclosure, the audio source localization device can collect audio signals from multiple angles. By processing the audio signals, audio features can be obtained, and based on the audio features, the distance and direction between the audio source and the audio source localization device can be obtained; the motion sensor can sense the motion state of the audio source localization device itself. By comprehensively considering the audio features and the motion features, the audio source position information included in the two types of features can be fully utilized. Compared with the traditional localization method that only considers audio features, adding motion features can adjust the localization result of the audio source in real time, so as to achieve more accurate audio source localization in various complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of an audio source localization method provided by an embodiment of the present disclosure; Figure 2 It is a structural block diagram of an audio source localization device provided by an embodiment of the present disclosure; Figure 3 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0012] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0013] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.
[0014] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an audio source localization method provided for an embodiment of the present disclosure. The method includes: S101: Extract audio features from the audio signal sent by the audio source localization device; the audio features include time difference of arrival, frequency, and amplitude; the audio source localization device includes an array microphone, and the time difference of arrival is the time difference of the audio source signal arriving at each microphone in the array microphone.
[0015] In this embodiment, the array microphone in the audio source localization device is composed of multiple microphone units. When the audio source emits sound, due to the different distances and angles between different microphones and the audio source, the time for the sound to reach each microphone will be different. By calculating the time difference of arrival (TDOA) of the sound reaching different microphones, the direction of the audio source can be initially determined, just like a person has two ears and judges the general direction of the sound through the time difference of the sound reaching the two ears.
[0016] At the same time, the frequency feature can reflect the sound generation characteristics of the audio source itself and the changes caused by various factors (such as the Doppler effect, etc.) during the propagation of the audio signal. Different audio sources often have different frequency features or frequency change laws, and based on this, different sounding objects can be distinguished and located. In addition, the closer the distance between the audio source and the microphone, the greater the received sound amplitude usually is. Therefore, the amplitude information can assist in judging the distance of the audio source. Taking the time difference of arrival, frequency, and amplitude of the sound as audio features together helps to more accurately locate the position of the audio source.
[0017] Specifically, the time difference of arrival can be obtained by the cross-correlation function method. Taking the signals received by any two microphones , as an example, the cross-correlation function of the two signals is:
[0018] The cross-correlation function The time corresponding to the peak position is the time difference for the sound to reach the two microphones. 。
[0019] The frequency characteristics can be obtained by performing a short-time Fourier transform (STFT) on the audio signal. Assuming the audio signal is ,First, the audio signal is divided into short time segments by a window function, such as ,where is the moving step size of the window.
[0020] Then, perform a Fourier transform on to obtain ,which represents the distribution of the signal at different times and frequencies. By comparing the spectra at different times, the frequency changes can be obtained. For example, when the Doppler effect is generated by the movement of the audio source, the spectrum of the received signal will shift, and this frequency change can be detected by the short-time Fourier transform.
[0021] The amplitude characteristics can be obtained by calculating the root mean square (RMS) amplitude of the audio signal. By comparing the RMS amplitudes at different times, the amplitude changes can be obtained. In the presence of noise interference, calculating the RMS can better extract the amplitude change characteristics of the audio source.
[0022] S102: Extract the motion characteristics of the audio source localization device from the operation data sent by the motion sensor; the motion characteristics include speed, acceleration, and direction.
[0023] In this embodiment, the motion sensor can include an accelerometer and a gyroscope, etc. When the audio source localization device itself moves or there is relative motion between the audio source and the localization device, motion sensors such as accelerometers and gyroscopes can sense the motion state of the device, including speed, acceleration, direction changes, etc.
[0024] Among them, the accelerometer can be used to measure acceleration. By integrating the acceleration over time, the speed can be obtained, and by integrating the speed again, the displacement can be obtained; the gyroscope can accurately measure the change in the rotation angle of the audio source localization device to determine the change in direction.
[0025] Combining the motion characteristics and the audio characteristics can dynamically determine the position of the audio source, compensating for the possible deviations in localization in dynamic scenarios relying solely on audio characteristics. For example, if the localization device moves forward a certain distance, then the actual spatial position corresponding to the previously calculated relative direction of the audio source needs to be corrected according to the displacement of the device; if the device rotates, the representation of the relative direction of the audio source in the global coordinate system also needs to be adjusted accordingly.
[0026] S103: Train a localization model based on the audio characteristics and motion characteristics, and perform the localization of the audio source based on the trained localization model.
[0027] In this embodiment, the combinations of audio features and motion features generated by different audio sources in different motion states often have certain regularities. Therefore, based on a neural network, a large amount of data with annotations (known audio source positions) can be used to train a positioning model. By learning the complex mapping relationships between audio features such as time difference of arrival, frequency, amplitude, etc., and motion features such as speed, acceleration, direction, etc., and the audio source position, in practical applications, the model can predict the position of the audio source according to the input features, including direction and distance. For example, through the multiple hidden layers of a deep neural network, higher-level feature combinations can be automatically extracted, thus enabling more accurate position prediction.
[0028] Among them, existing neural networks can adopt multi-layer perceptrons (MLP), convolutional neural networks (CNN), recurrent neural networks (RNN), etc.
[0029] It can be concluded from the above that in this embodiment, the audio source positioning device can collect audio signals from multiple angles. By processing the audio signals, audio features can be obtained, and based on the audio features, the distance and direction between the audio source and the audio source positioning device can be obtained; the motion sensor can sense the motion state of the audio source positioning device itself. By comprehensively considering the audio features and motion features, the audio source position information contained in both types of features can be fully utilized. Compared with the traditional positioning method that only considers audio features, adding motion features can adjust the positioning result of the audio source in real time, so as to achieve more accurate audio source positioning in various complex scenarios.
[0030] In an embodiment of the present disclosure, training a positioning model based on audio features and motion features includes: Calculating the correlation coefficients between every two features in the audio features and motion features to obtain a plurality of first correlation coefficients; Screening second correlation coefficients from the plurality of first correlation coefficients based on the magnitudes of the first correlation coefficients; Determining the second correlation coefficients as interaction features, and training a positioning model based on the audio features, motion features, and interaction features.
[0031] In this embodiment, considering that there may be an inherent connection between different audio features and motion features. For example, the moving speed of an audio source may be related to the time difference of arrival of sound. When the audio source moves towards the microphone array, the change in speed may cause a corresponding change in the time difference of arrival of sound. Another example is that the moving direction may be related to the frequency change of sound (due to the Doppler effect, etc.). By calculating the correlation coefficient between features, the relationship hidden between features can be understood. Using the relationship between features in the training of the positioning model can help the model deeply understand the mutual influence mechanism of various factors in the audio source positioning process. Among them, the calculation of the correlation coefficient can be implemented using the existing Pearson correlation coefficient.
[0032] Meanwhile, in order to avoid introducing too much irrelevant or interfering information during the training of the model, so that the model can focus more on learning the key feature interaction patterns for positioning, this embodiment screens based on the magnitude of the first correlation coefficient, selects the coefficients corresponding to the feature combinations with stronger correlations, that is, the second correlation coefficient, and uses the second correlation coefficient as the interaction feature, together with the audio features and motion features, for the training of the positioning model. This can remove those feature relationships with weaker correlations and relatively smaller contributions to positioning, so as to focus on those truly valuable and closely related feature connections, which helps to improve the efficiency of model training and the accuracy of positioning.
[0033] Among them, when screening the second correlation coefficient, a certain threshold can be set. The first correlation coefficient (second correlation coefficient) greater than this threshold indicates that there is a relatively close association or a relatively significant mutual influence between the corresponding two features. Therefore, this second correlation coefficient is introduced as an interaction feature into the training of the positioning model.
[0034] From the above, it can be concluded that in this embodiment, by exploring the correlation between audio features and motion features, valuable interaction features are incorporated into the model training, enabling the model to more comprehensively and deeply understand the cooperative effect and mutual influence of different features in the audio source positioning process. When facing complex actual scenarios (such as when the audio source and the positioning device are both in a complex motion state), it can make a more accurate judgment of the audio source position, thereby effectively improving the positioning accuracy.
[0035] In an embodiment of the present disclosure, calculating the correlation coefficient between each two of the audio features and motion features includes: For any two features, Respectively obtain N sample data of any two features in the first time period to obtain the first sample data and the second sample data; Calculate the probability distribution of the first sample data, the probability distribution of the second sample data, and the joint probability respectively; the joint probability is the joint probability of the first sample data and the second sample data; Calculate the mutual information value between any two features based on the probability distribution of the first sample data, the probability distribution of the second sample data, and the joint probability; Determine the mutual information value as the correlation coefficient between any two features.
[0036] In this embodiment, the method of calculating the mutual information value can be used to obtain the correlation coefficient between features to characterize the mutual dependence between two variables. Taking the time difference of arrival of sound and frequency features as an example, the calculation process of the mutual information value between the above two features can include: First, divide the first sample data corresponding to the time difference of arrival of sound and the second sample data corresponding to the frequency feature into several intervals respectively. For example, for the time difference of arrival of sound, according to its possible value range (assuming between -10 ms and 10 ms), it is divided into equally spaced intervals, such as each 1 ms as an interval; for the frequency feature, according to the frequency change range of the audio signal (assuming -100 Hz to 100 Hz), it is divided into each 10 Hz as an interval.
[0037] Count the number of sample data (including the first sample data and the second sample data) of the time difference of arrival of sound and the frequency feature falling into each interval. For example, represents the number of the time difference of arrival of sound falling into interval i, represents the number of the frequency feature falling into interval j, represents the number of sample data where the time difference of arrival of sound falls into interval i and the frequency feature falls into interval j.
[0038] On this basis, the probability distribution of the first sample data can be obtained 、the probability distribution of the first sample data 、and the joint probability .
[0039] Finally, use the first formula to calculate the mutual information value between the time difference of arrival of sound and the frequency feature. The first formula is specifically:
[0040] The larger the mutual information value, the higher the degree of dependence between the two features. For complex audio source localization scenarios, there may be various non-linear interactions between audio features and motion features. The mutual information method can effectively capture the dependence relationship between audio features and motion features, providing more comprehensive feature relationship information for subsequent model training.
[0041] In one embodiment of the present disclosure, training a localization model based on audio features and motion features includes: In response to the loss function of the localization model being greater than a first threshold, calculate the gradients of the loss function of the localization model with respect to each feature, and adjust the weights of each feature based on the gradients of the loss function with respect to each feature; In response to the loss function of the localization model being less than or equal to the first threshold, determine a first adjustment parameter based on the historical gradients of each feature, and adjust the weights of each feature based on the first adjustment parameter and the gradients of the loss function with respect to each feature; the historical gradients of each feature are the historical data of the gradients of the loss function with respect to each feature.
[0042] In this embodiment, when the loss function of the localization model is greater than the first threshold, it indicates that the prediction effect of the model is far from the expected result. At this time, updating the weights in the opposite direction of the gradient can make the loss function quickly change in the decreasing direction, so that the model can quickly learn the basic relationship pattern between the audio features, motion features and the audio source position in the initial stage of training, and quickly converge to the region with a smaller loss function. Among them, the first threshold is a preset constant, and those skilled in the art can design the specific value of the first threshold according to actual needs.
[0043] Specifically, the weights of each feature can be adjusted according to the following second formula: ; where represents the weights of each feature, represents the loss function, is a preset proportionality coefficient.
[0044] When the loss function of the localization model is less than or equal to the first threshold, it indicates that the model is close to the optimal solution or near the local optimal solution. At this time, if the weights of each feature are still adjusted based on the gradients of the loss function with respect to each feature, for the feature with a relatively large corresponding gradient, it may oscillate back and forth near the optimal solution due to too large an adjustment amount and cannot converge stably. For the feature with a relatively small corresponding gradient, it may converge too slowly due to too small an adjustment amount.
[0045] To solve the above problems, in this embodiment, a first adjustment parameter is determined based on the historical gradients of each feature, the gradients corresponding to each feature are adjusted based on the first adjustment parameter, and the weights of each feature are adjusted based on the adjusted gradients.
[0046] Specifically, determining the first adjustment parameter based on the historical gradients of each feature can be elaborated as: Calculate the square root of the sum of the squares of the historical gradients of each feature to obtain multiple gradient eigenvalues; Determine the first adjustment parameter based on the negative correlation between each gradient eigenvalue and the first adjustment parameter.
[0047] In this embodiment, by calculating the square root of the sum of the historical squared gradients of each feature, multiple gradient eigenvalues are obtained, which can unify the positive and negative information of the gradients and avoid the change of the first adjustment parameter caused by the alternating positive and negative gradients. At the same time, the summation and square root operations comprehensively consider the gradient change of the historical gradient over a period of time, making the first adjustment parameter more stable, and further making the adjustment amount of the weight more stable.
[0048] The negative correlation between each gradient eigenvalue and the first adjustment parameter can be calculated using the following third formula: ; where represents the first adjustment parameter, represents the gradient eigenvalue corresponding to each feature.
[0049] Based on the obtained first adjustment parameter, the following fourth formula can be used to calculate and adjust the weights of each feature: .
[0050] It can be concluded from the above that in this embodiment, when the loss function is large, the gradient-based weight adjustment method can quickly guide the model to move in the direction of reducing the loss function. When the loss function is small, the first adjustment parameter is determined based on the historical gradients of each feature, and the gradients of each feature are adjusted based on the first adjustment parameter. For features with large gradients, the gradients of these features can be adjusted smaller through the first adjustment parameter to avoid over-adjustment; for features with small gradient changes, the gradients of these features can be adjusted larger through the first adjustment parameter, enabling the model to converge to the optimal solution more stably.
[0051] In an embodiment of the present disclosure, extracting audio features from the audio signal sent by the audio source localization device includes: Performing spectral analysis on the audio signal based on the Fourier transform to obtain the power spectrum of the audio signal; Calculating the noise power of the audio signal based on the power spectral density of the audio signal; In response to the noise power being less than or equal to the second threshold, performing echo cancellation processing and noise cancellation processing on the audio signal in sequence; In response to the noise power being greater than the second threshold, performing noise cancellation processing and echo cancellation processing on the audio signal in sequence.
[0052] In this embodiment, considering that in the actual environment, the audio signal is easily affected by background noise, echoes, etc., which affects the accuracy of features such as the time difference of arrival, frequency, and amplitude of the sound. Therefore, before extracting the features of the audio signal, it is necessary to perform echo cancellation processing and noise cancellation processing on the audio signal.
[0053] Furthermore, considering that in the case of relatively low noise, the impact of echo on audio features is more significant. If the noise is processed first, it may change the structure of the audio signal, thereby affecting the effect of echo cancellation. For example, the echo may have complex interweaving relationships with the original audio signal in both the time domain and the frequency domain. Removing the echo first when the noise is low can better retain the useful information of the original audio signal and provide more favorable conditions for subsequent noise cancellation.
[0054] Correspondingly, when the noise is high, the noise may mask the characteristics of the echo, making it difficult to effectively perform echo cancellation. For example, in a noisy environment, strong background noise may interfere with the detection and estimation of the echo. Reducing the noise first can make the characteristics of the echo more obvious, thereby improving the accuracy of echo cancellation.
[0055] Therefore, in this embodiment, first, by performing spectral analysis on the audio signal to obtain the power spectrum of the audio signal, and integrating the power spectrum over the entire frequency range or a set frequency interval, the noise power can be obtained. The noise power can be used to characterize the noise intensity; then, the noise power is compared with a second threshold, and based on the comparison result, the order of echo cancellation processing and noise cancellation processing is determined. Here, the second threshold is a preset constant, and those skilled in the art can design the specific value of the second threshold according to actual needs.
[0056] From the above, it can be concluded that in this embodiment, in different noise environments, based on the noise power, the order of echo cancellation processing and noise cancellation processing is flexibly adjusted, which can more effectively restore the original features of the audio signal, thereby improving the accuracy of audio feature extraction and providing more reliable data support for subsequent audio source localization.
[0057] In an embodiment of the present disclosure, the echo cancellation processing process includes: Estimating the echo delay time and echo amplitude based on the autocorrelation function of the audio signal; Constructing an echo filter based on the echo delay time and echo amplitude; Filtering the audio signal based on the echo filter.
[0058] In this embodiment, when performing echo cancellation processing, first, the echo delay time and echo amplitude are estimated based on the autocorrelation function of the audio signal. Specifically, the calculation formula of the autocorrelation function is:
[0059] Where represents the estimated value of the autocorrelation function of the audio signal, represents the discrete audio signal, represents the delay time.
[0060] In the calculation result of the autocorrelation function, the main peak usually corresponds to the correlation of the signal itself ( ), and the echo will appear after a certain delay, manifested as a secondary peak after the main peak. For example, if there is an echo, the autocorrelation function will have a distinct peak at a delay time greater than 0 . This delay time may be the delay time, and this distinct peak is the echo amplitude.
[0061] Based on obtaining the echo delay time and echo amplitude, an echo filter can be constructed. For example, the following fifth formula can be used to calculate and construct the echo filter:
[0062] Where represents the output result of the echo filter, and is a preset proportionality coefficient.
[0063] Inputting the audio signal into the constructed echo filter can effectively reduce or remove the echo component in the audio signal.
[0064] It can be concluded from the above that in this embodiment, by accurately estimating the echo delay time and echo amplitude and constructing a suitable echo filter for filtering, the echo component in the audio signal can be effectively removed.
[0065] In an embodiment of the present disclosure, the noise cancellation process includes: Removing the noise power spectrum from the power spectrum of the audio signal to obtain a first power spectrum; Performing an inverse Fourier transform on the first power spectrum to obtain the denoised audio signal.
[0066] In this embodiment, the noise can be detected according to a set time period. For example, when each time period arrives, audio signals of multiple time periods are collected. When the root mean square value of the audio signal in a certain time period is less than the third threshold, it indicates that the audio signal in this time period does not contain the signal of the audio source and only contains noise signals. At this time, by performing spectral analysis on the audio signal in this time period, the noise power spectrum can be obtained.
[0067] On this basis, removing the noise power spectrum from the power spectrum of the audio signal can obtain the power spectrum of the audio source, that is, the first power spectrum. Performing an inverse Fourier transform on the first power spectrum can obtain the denoised audio signal.
[0068] Wherein, the third threshold can be obtained by multiplying the average value of the root mean square values of the audio signals of multiple time periods by a preset proportionality coefficient. The preset proportionality coefficient is a coefficient between 0 and 1, such as 0.1, 0.2, etc.
[0069] It can be concluded from the above that the noise cancellation method based on power spectrum subtraction and inverse Fourier transform is adopted in this embodiment, which can be applied to various types of audio signals and noise scenarios, effectively reducing the noise components in the audio signal and making the audio signal purer.
[0070] Corresponding to the audio source localization method in the above embodiment, Figure 2 is a structural block diagram of an audio source localization device provided by an embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 2 The audio source localization device 20 includes: a first feature extraction module 21, a second feature extraction module 22, and an audio source localization module 23.
[0071] Among them, the first feature extraction module 21 is used to extract audio features from the audio signal sent by the audio source localization device; the audio features include time difference of arrival, frequency, and amplitude; the audio source localization device includes an array microphone, and the time difference of arrival is the time difference of the audio source signal arriving at each microphone in the array microphone; The second feature extraction module 22 is used to extract the motion features of the audio source localization device from the operation data sent by the motion sensor; the motion features include speed, acceleration, and direction; The audio source localization module 23 is used to train a localization model based on the audio features and motion features, and perform audio source localization based on the trained localization model.
[0072] In an embodiment of the present disclosure, the audio source localization module 23 is specifically used for: Calculating the correlation coefficients between every two features in the audio features and motion features to obtain a plurality of first correlation coefficients; Screening second correlation coefficients from the plurality of first correlation coefficients based on the magnitudes of the first correlation coefficients; Determining the second correlation coefficients as interaction features, and training a localization model based on the audio features, motion features, and interaction features.
[0073] In an embodiment of the present disclosure, the audio source localization module 23 is specifically further used for: For any two features, Respectively obtaining N sample data of any two features within the first time period to obtain first sample data and second sample data; Respectively calculating the probability distribution of the first sample data, the probability distribution of the second sample data, and the joint probability; the joint probability is the joint probability of the first sample data and the second sample data; Calculating the mutual information value between any two features based on the probability distribution of the first sample data, the probability distribution of the second sample data, and the joint probability; Determine the mutual information value as the correlation coefficient between any two features.
[0074] In an embodiment of the present disclosure, the audio source localization module 23 is specifically configured to: In response to the loss function of the localization model being greater than a first threshold, calculate the gradient of the loss function of the localization model with respect to each feature, and adjust the weight of each feature based on the gradient of the loss function with respect to each feature; In response to the loss function of the localization model being less than or equal to the first threshold, determine a first adjustment parameter based on the historical gradient of each feature, and adjust the weight of each feature based on the first adjustment parameter and the gradient of the loss function with respect to each feature; the historical gradient of each feature is the historical data of the gradient of the loss function with respect to each feature.
[0075] In an embodiment of the present disclosure, the first feature extraction module 21 is specifically configured to: Perform spectral analysis on the audio signal based on the Fourier transform to obtain the power spectrum of the audio signal; Calculate the noise power of the audio signal based on the power spectral density of the audio signal; In response to the noise power being less than or equal to a second threshold, perform echo cancellation processing and noise cancellation processing on the audio signal in sequence; In response to the noise power being greater than the second threshold, perform noise cancellation processing and echo cancellation processing on the audio signal in sequence.
[0076] In an embodiment of the present disclosure, the first feature extraction module 21 is further specifically configured to: Estimate the echo delay time and echo amplitude based on the autocorrelation function of the audio signal; Construct an echo filter based on the echo delay time and echo amplitude; Perform filtering processing on the audio signal based on the echo filter.
[0077] In an embodiment of the present disclosure, the first feature extraction module 21 is further specifically configured to: Remove the noise power spectrum from the power spectrum of the audio signal to obtain a first power spectrum; Perform an inverse Fourier transform on the first power spectrum to obtain the denoised audio signal.
[0078] See Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 3The electronic device 300 in the present embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above device embodiments, for example Figure 2 the functions of the modules 21 to 23 shown.
[0079] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0080] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0081] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0082] In specific implementation, the processors 301, input devices 302, and output devices 303 described in the embodiments of the present disclosure may implement the implementation manners described in the first embodiment and the second embodiment of the audio source localization method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated here.
[0083] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0084] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0085] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0086] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.
[0087] In several embodiments provided by the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, or can also be an electrical, mechanical or other form of connection.
[0088] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0089] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0090] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for localizing an audio source, characterized in that: include: Extracting audio features from the audio signal sent by the audio source positioning device; the audio features include sound arrival time difference, frequency and amplitude; The audio source positioning device includes an array microphone, and the sound arrival time difference is the time difference between the audio source signal arriving at each microphone in the array microphone; Extracting motion characteristics of the audio source positioning device from the operation data sent by the motion sensor; the motion characteristics include speed, acceleration and direction; A positioning model is trained based on the audio features and the motion features, so as to locate the audio source based on the trained positioning model.
2. The audio source localization method according to claim 1, characterized in that: The training of the positioning model based on the audio feature and the motion feature includes: Calculating the correlation coefficient between every two features of the audio features and the motion features to obtain a plurality of first correlation coefficients; Selecting a second correlation coefficient from among a plurality of first correlation coefficients based on the magnitude of the first correlation coefficient; The second correlation coefficient is determined as an interaction feature, and a positioning model is trained based on the audio feature, the motion feature and the interaction feature.
3. The audio source localization method according to claim 2, characterized in that: The calculating the correlation coefficient between every two features of the audio features and the motion features includes: For any two features, Respectively obtain N sample data of the arbitrary two features in the first time period to obtain first sample data and second sample data; Calculate the probability distribution of the first sample data, the probability distribution of the second sample data, and the joint probability respectively; the joint probability is the joint probability of the first sample data and the second sample data; Calculate the mutual information value between any two features based on the probability distribution of the first sample data, the probability distribution of the second sample data and the joint probability; The mutual information value is determined as a correlation coefficient between the arbitrary two features.
4. The audio source localization method according to claim 1 or 2, characterized in that: The training of the positioning model based on the audio feature and the motion feature includes: In response to the loss function of the positioning model being greater than a first threshold, calculating the gradient of the loss function of the positioning model with respect to each feature, and adjusting the weight of each feature based on the gradient of the loss function with respect to each feature; In response to the loss function of the positioning model being less than or equal to a first threshold, a first adjustment parameter is determined based on the historical gradient of each feature, and the weight of each feature is adjusted based on the first adjustment parameter and the gradient of the loss function for each feature; the historical gradient of each feature is the historical data of the gradient of the loss function for each feature.
5. The audio source localization method according to claim 1, characterized in that: The step of extracting audio features from an audio signal sent by an audio source positioning device comprises: Performing spectrum analysis on the audio signal based on Fourier transform to obtain a power spectrum of the audio signal; Calculating the noise power of the audio signal based on the power spectral density of the audio signal; In response to the noise power being less than or equal to a second threshold, sequentially performing echo cancellation processing and noise cancellation processing on the audio signal; In response to the noise power being greater than a second threshold, the audio signal is sequentially subjected to noise cancellation processing and echo cancellation processing.
6. The audio source localization method according to claim 5, characterized in that: The echo cancellation process includes: Estimating the echo delay time and echo amplitude based on the autocorrelation function of the audio signal; constructing an echo filter based on the echo delay time and the echo amplitude; The audio signal is filtered based on the echo filter.
7. The audio source localization method according to claim 5, characterized in that: The noise cancellation process includes: Removing a noise power spectrum from the power spectrum of the audio signal to obtain a first power spectrum; Perform inverse Fourier transform on the first power spectrum to obtain a denoised audio signal.
8. An audio source localization device, characterized in that: include: A first feature extraction module is used to extract audio features from an audio signal sent by an audio source positioning device; the audio features include a sound arrival time difference, a frequency, and an amplitude; the audio source positioning device includes an array microphone, and the sound arrival time difference is a time difference between an audio source signal arriving at each microphone in the array microphone; A second feature extraction module is used to extract the motion features of the audio source positioning device from the operation data sent by the motion sensor; the motion features include speed, acceleration and direction; The audio source localization module is used to train a localization model based on the audio features and the motion features, so as to localize the audio source based on the trained localization model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.