A method and system for estimating the number of sound sources
By performing short-time Fourier transform and mono sound source DOA estimation after receiving the signal in the microphone array, combined with frequency statistics and peak point analysis, the problem of large error in the number of sound sources when the ambient reverberation is strong or the noise does not meet the Gaussian white noise requirements is solved, and fast and accurate estimation of the number of sound sources and improving the voice signal processing effect is achieved.
Patent Information
- Application Number
- CN202210445311.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-26
AI Technical Summary
When the prior art reverberation is relatively strong in the environment or the background noise does not meet the Gaussian white noise requirements, the estimation error of the number of sound sources is large, which affects the subsequent voice signal processing effect.
The signal is received through the microphone array and short-time Fourier transform is performed, and the mono sound source DOA estimation is performed on multiple frequency points of the transform domain signal, the frequencies obtained by the estimated directions of different sound source are counted, the peak points are found to determine the number of sound sources, and the number of active sound sources is updated in real time.
It realizes rapid and accurate estimation of the number of sound sources in complex acoustic environments, reduces the error in estimating the number of sound sources, and improves the effect of voice signal processing.
Smart Images

Figure CN114724585B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of signal processing, and in particular, to a method and system for estimating the number of sound sources. Background Art
[0002] The problem of estimating the number of sound sources has always been a hot topic in the field of speech signal processing. Many classic signal processing methods require knowing the number of sound sources. For example, in the multiple signal classification (MUSIC) method for direction of arrival (DOA) estimation, its working premise is that the number of sound sources is known. Then, the MUSIC method divides the signal subspace and the noise subspace according to the number of sound sources to calculate the spatial spectrum, and finally determines the azimuths of multiple sound sources by finding the positions of the extreme points on the spatial spectrum. However, if there is an error between the estimated number of sound sources and the actual number of sound sources, the DOA estimation performance of the MUSIC method will decrease significantly. Therefore, how to quickly and accurately estimate the number of sound sources has important research value in the field of speech signal processing.
[0003] However, the limitation of the existing methods is that only Gaussian white noise and non-reverberant conditions are considered in the signal model. When the reverberation in the environment is strong or the background noise does not meet the requirements of Gaussian white noise, there is a large estimation error in the number of sound sources, which in turn affects the subsequent speech signal processing effect. Summary of the Invention
[0004] The object of the present invention is to provide a method and system for estimating the number of sound sources to solve the technical problem in the prior art that when the reverberation in the environment is strong or the background noise does not meet the requirements of Gaussian white noise, there is a large estimation error in the number of sound sources, which in turn affects the subsequent speech signal processing effect.
[0005] In view of the above problems, the present invention provides a method and system for estimating the number of sound sources.
[0006] In a first aspect, the present invention provides a method for estimating the number of sound sources, which is implemented by a system for estimating the number of sound sources. Wherein, the method includes: receiving a first signal y m (t) through a microphone array; performing short-time Fourier transform on the first signal to obtain a transformed-domain signal Y m (k, l); performing single sound source DOA estimation on multiple frequency points within the transformed-domain signal to obtain an estimation result; within the estimation result, statistically obtaining the frequencies estimated for different sound source azimuths within a preset time period to obtain a statistical frequency result Q(θ); finding peak points within the statistical frequency result Q(θ) to obtain multiple sound source azimuths for which the estimated frequency is greater than a threshold τ1, obtaining a sound source number estimation result; and updating the number of currently active sound sources based on the sound source number estimation result.
[0007] On the other hand, the present invention also provides an estimation system for the number of sound sources, which is used to execute an estimation method for the number of sound sources as described in the first aspect. Wherein, the system includes: a first acquisition unit: the first acquisition unit is used to receive and acquire a first signal y through a microphone array m (t); a second acquisition unit: the second acquisition unit is used to perform a short-time Fourier transform on the first signal to obtain a transformed-domain signal Y m (k, l); a third acquisition unit: the third acquisition unit is used to perform single sound source DOA estimation on multiple frequency points within the transformed-domain signal to obtain an estimation result; a fourth acquisition unit: the fourth acquisition unit is used to statistically obtain the frequencies obtained by estimating different sound source azimuths within the preset time period in the estimation result to obtain a statistical frequency result Q(θ); a fifth acquisition unit: the fifth acquisition unit is used to find the peak points within the statistical frequency result Q(θ) to obtain multiple sound source azimuths where the estimated frequency is greater than a threshold value τ1 to obtain a sound source number estimation result; a first update unit: the first update unit is used to update the number of currently active sound sources based on the sound source number estimation result.
[0008] In a third aspect, an electronic device includes a processor and a memory;
[0009] The memory is used for storage;
[0010] The processor is used to execute the method described in any one of the above first aspects by calling.
[0011] In a fourth aspect, a computer program product includes a computer program and / or instructions, and when the computer program and / or instructions are executed by a processor, the steps of the method described in any one of the above first aspects are implemented.
[0012] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0013] 1. Receive the first signal y m (t) through a microphone array and perform a short-time Fourier transform to obtain a transformed-domain signal Y m(k, l), and then perform single-source DOA estimation at multiple frequency points in the transformed-domain signal respectively, and obtain the corresponding estimation results. Further, count the frequencies estimated for different sound source orientations within a preset time period to obtain the corresponding statistical frequency result Q(θ). Finally, screen out multiple sound source orientations with peak point frequencies greater than the threshold τ1 in the statistical frequency result Q(θ) to determine the sound source number estimation result. In addition, the number of active sound sources in the sound source number estimation result is updated in real time. By performing single-source DOA estimation on each frequency point in the transformed-domain signal after short-time Fourier transform in sequence, and then estimating the number of sound sources through statistical analysis, the technical effect of quickly and accurately estimating the number of sound sources in a complex acoustic environment and reducing the sound source number estimation error is achieved.
[0014] 2. By performing single-source DOA estimation on each frequency point in the transformed-domain signal after short-time Fourier transform in sequence, and then estimating the number of sound sources through statistical analysis, the eigenvalue decomposition of the covariance matrix is avoided, and the technical goal of reducing the computational complexity and improving the sound source number estimation speed is achieved.
[0015] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the embodiments of the present invention. Brief Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0017] Figure 1 It is a schematic flowchart of a method for estimating the number of sound sources according to the present invention;
[0018] Figure 2 It is a schematic flowchart of the activation, maintenance, and termination strategies in a method for estimating the number of sound sources according to the present invention;
[0019] Figure 3 It is a schematic structural diagram of a system for estimating the number of sound sources according to the present invention;
[0020] Figure 4 It is a schematic structural diagram of an exemplary electronic device according to the present invention;
[0021] Description of the Reference Numerals:
[0022] The first acquisition unit 11, the second acquisition unit 12, the third acquisition unit 13, the fourth acquisition unit 14, the fifth acquisition unit 15, the first update unit 16, the bus 300, the receiver 301, the processor 302, the transmitter 303, the memory 304, the bus interface 305. Detailed implementation manners
[0023] By providing a method and a system for estimating the number of sound sources, the present invention solves the technical problem in the prior art that when the reverberation in the environment is strong or the background noise does not meet the requirements of Gaussian white noise, the estimation error of the number of sound sources is large, which further affects the subsequent speech signal processing effect. By performing single sound source DOA estimation on each frequency point in the transformed domain signal after short-time Fourier transform, and then statistically analyzing to estimate the number of sound sources, the technical effect of quickly and accurately estimating the number of sound sources in a complex acoustic environment and reducing the estimation error of the number of sound sources is achieved.
[0024] In the technical solution of the present invention, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0025] Next, the technical solutions in the present invention will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the example embodiments described herein. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In addition, it should be noted that for the sake of description, only the parts related to the present invention are shown in the accompanying drawings rather than all of them.
[0026] The present invention provides a method for estimating the number of sound sources, and the method is applied to a system for estimating the number of sound sources. Wherein, the method includes: receiving and acquiring a first signal y m (t) through a microphone array; performing short-time Fourier transform on the first signal to obtain a transformed domain signal Y m (k, l); performing single sound source DOA estimation on multiple frequency points in the transformed domain signal to obtain an estimation result; within the estimation result, statistically obtaining the frequencies estimated for different sound source orientations within a preset time period to obtain a statistical frequency result Q(θ); finding the peak points in the statistical frequency result Q(θ) to obtain multiple sound source orientations for which the estimated frequency is greater than a threshold value τ1, to obtain a sound source number estimation result; and updating the number of currently active sound sources based on the sound source number estimation result.
[0027] After introducing the basic principle of the present invention, the various non-limiting implementation manners of the present invention will be specifically introduced below with reference to the accompanying drawings of the specification.
[0028] Embodiment 1
[0029] Please refer to the appendix Figure 1 The present invention provides a method for estimating the number of sound sources. Specifically, the method is applied to a system for estimating the number of sound sources, and the method specifically includes the following steps:
[0030] Step S100: Receive and obtain the first signal y m (t);
[0031] Specifically, the method for estimating the number of sound sources is applied to the system for estimating the number of sound sources. By performing single-source DOA estimation on each frequency point in the transformed-domain signal after short-time Fourier transform, and then through statistical analysis, the accurate and reliable number of sound sources can be quickly estimated.
[0032] Estimating the number of sound sources is the basis of speech signal processing. Before calculating and estimating the number of sound sources, the microphone array is first used to receive the sound source signal, thereby obtaining the corresponding sound source signal, that is, the first signal y m (t). By receiving the sound source signal, the technical effect of providing a data basis for subsequent statistical analysis and then determining the number of sound sources is achieved.
[0033] Step S200: Perform short-time Fourier transform on the first signal to obtain the transformed-domain signal Y m (k, l);
[0034] Furthermore, the short-time Fourier transform is calculated by the following formula:
[0035]
[0036] where m is the m-th microphone in the microphone array, k is the frequency index, l is the frame index, w(n) is a window function, N is the window length, and L is the hop length.
[0037] The result calculated using the formula is the transformed-domain signal Y m (k, l). Among them, the short-time Fourier transform refers to determining the frequency of the sine wave in the local area by selecting any window function with local time-frequency variation in the first signal. By determining the transformed-domain signal through short-time Fourier transform, the technical effect of providing a basis for subsequent estimating the direction of arrival of waves at each frequency point is achieved.
[0038] Step S300: Perform single-source DOA estimation on multiple frequency points in the transformed-domain signal to obtain the estimation results;
[0039] Specifically, according to the short-time Fourier transform, the transformed-domain signal Y m(k, l), calculate the manifold vector of the microphone array, then determine the number of microphones in the microphone array, and finally obtain the direction-of-arrival estimation results of a single sound source at each frequency point in sequence, which is the estimation result. Among them, the direction of arrival refers to the arrival direction of a spatial signal, that is, the single sound source signal. By performing the direction-of-arrival estimation of a single sound source on each frequency point in the transformed-domain signal in sequence, the technical effect of providing a data basis for subsequent estimation of the direction of arrival of all single sound sources and further determining the frequency is achieved.
[0040] Step S400: In the estimation result, statistically obtain the frequencies at which different sound source azimuths are estimated within a preset time period, and obtain the statistical frequency result Q(θ);
[0041] Further, step S400 of the present invention further includes:
[0042] Step S410: According to the estimation result, obtain all the direction-of-arrival estimation results of all the single sound sources at multiple frequency points within all time;
[0043] Step S420: Based on all the direction-of-arrival estimation results of all the single sound sources, statistically obtain the frequencies Q(θ) at which all the direction-of-arrival of the single sound sources are estimated, where θ ∈ {θ1, θ2, …, θ D}.
[0044] Specifically, the estimation result includes all the direction-of-arrival estimations of the signals received by the microphone array within all time. Further, within the estimation result, statistically obtain the frequencies at which different sound source azimuths are estimated within a preset time period. Among them, the preset time period is a time range comprehensively determined based on the actual direction-of-arrival estimation result of a single sound source, the actual sound source environment, etc. Exemplarily, according to the above estimation result, statistically obtain the frequencies at which all the direction-of-arrival of all single sound sources at all frequency points within the past T frames of time are estimated, and this frequency is the statistical frequency result Q(θ). The T-frame time is the preset time period, and the value of T can be set according to the requirements of sound source number estimation, for example, it can be all time.
[0045] First, according to the estimation result, determine all the direction-of-arrival estimation results of all the single sound sources at multiple frequency points within all time, and then based on all the direction-of-arrival estimation results of all the single sound sources, statistically obtain the frequencies Q(θ) at which all the direction-of-arrival of the single sound sources are estimated, where θ ∈ {θ1, θ2, …, θ D}. By obtaining the statistical frequency result, the technical goal of providing a range for subsequent determination of the peak point is achieved, and the technical effect of providing a basis for improving the accuracy of sound source number estimation is achieved.
[0046] Step S500: Search for and obtain the peak points within the statistical frequency result Q(θ), acquire multiple sound source azimuths with the estimated frequency greater than the threshold τ1, and obtain the estimated result of the number of sound sources;
[0047] Further, step S500 of the present invention further includes:
[0048] Step S510: Search for and obtain all the peak points within the statistical frequency result Q(θ);
[0049] Step S520: Select all the peak points greater than the threshold τ1, denoted as θ i1 , θ i2 , …, θ iI ;
[0050] Step S530: Construct and obtain the first vector θ cand , θ cand with a length of D, as follows:
[0051]
[0052] Step S540: Use the first vector as the estimated result of the number of sound sources.
[0053] Specifically, according to the statistically determined statistical frequency result Q(θ), successively compare and analyze to determine its peak points, record the multiple sound source azimuths with a frequency greater than the threshold τ1, and form the estimated result of the number of sound sources with the multiple sound source azimuths greater than the threshold τ1.
[0054] First, successively analyze to obtain all the peak points within the statistical frequency result Q(θ), and then compare and screen all the peak points to obtain all the peak points greater than the threshold τ1, denoted as θ i1 , θ i2 , …, θ iI . Further, use each determined peak point, that is, the θ i1 , θ i2 , …, θ iI as the data basis to construct the first vector θ cand , and use the first vector θ cand as the estimated result of the number of sound sources. Among them, the threshold τ1 refers to being set in advance through comprehensive analysis based on the analysis and processing experience of historical sound source signals. Among them, the constructed first vector θ cand is:
[0055]
[0056] By constructing the first vector θ cand, achieving the technical effect of quickly calculating the number of sound sources.
[0057] Step S600: Based on the sound source quantity estimation result, update the number of currently active sound sources.
[0058] Furthermore, step S600 of the present invention further includes:
[0059] Step S610: Update the number of currently active sound sources by adopting an activation, retention, and termination strategy.
[0060] Furthermore, step S610 of the present invention further includes:
[0061] Step S611: Set the input parameters to θ active and θ cand , where the initial value of θ active is a zero vector and has the same length as the first vector;
[0062] Step S612: Subtract 1 from the values of all elements in θ active , and assign 0 to the elements with values less than 0;
[0063] Step S613: Add θ active and θ cand together, and use the result as the updated θ active ;
[0064] Step S614: Assign the elements in the updated θ active with values greater than the threshold τ max to τ max ;
[0065] Step S615: Count the number S of elements in the updated θ active whose values exceed the threshold τ2;
[0066] Step S616: Use the updated θ active and S as the output result, and S is the number of currently active sound sources.
[0067] Specifically, after estimating the number of sound sources, that is, obtaining the sound source quantity estimation result, an activation, retention, and termination strategy is adopted to perform real-time update of the number of currently active sound sources, that is, to determine the number of currently active sound sources.
[0068] As Figure 2 shown, Figure 2 shows a schematic flowchart of a possible activation, retention, and termination strategy according to an embodiment of the present application. Specifically, first, set the input parameters to θ active and θ cand , where the θ activeThe initial value is a zero vector, and its length is the same as that of the first vector θ cand Then subtract 1 from the values of all elements in θ active and assign all the θ active elements after the value subtraction to 0. Further, add θ active and θ cand and use the resulting sum as the updated θ active , and for the updated θ active , assign all elements with values greater than the threshold τ max to τ max . Finally, count the number S of elements in the updated θ active whose values exceed the threshold τ2, and use the updated θ active and S as the output results. S is the number of currently active sound sources. Among them, the settings of the threshold τ2 and the threshold τ max are also preset based on the analysis and processing experience of historical sound source signals to implement this strategy. Through the activation, maintenance, and termination strategies, the technical goal of real-time updating of the number of currently active sound sources is achieved.
[0069] Furthermore, step S300 of the present invention further includes:
[0070] Step S310: Calculate: P k,l (θ) = d H (θ)y(k, l)y(k, l) H d(θ), where d(θ) is the steering vector of the microphone array, θ ∈ {θ1, θ2, …, θ D}, and θ is the azimuth where the sound source may appear;
[0071] Step S320: Calculate: y(k, l) = [Y1(k, l) Y2(k, l) … Y M (k, l)] T , where M is the number of microphones in the microphone array;
[0072] Step S330: Find the angle at which the maximum peak point appears, as follows:
[0073]
[0074] Step S340: is the single sound source arrival azimuth estimation result at the l-th frame and the k-th frequency point;
[0075] Step S350: Calculate and obtain all the single sound source arrival azimuth estimation results at multiple said frequency points over all time to obtain the said estimation results.
[0076] Specifically, when performing single-source direction-of-arrival (DOA) estimation on each frequency point in the transform-domain signal in sequence, it is first calculated based on the following formula:
[0077] P k,l (θ) = d H (θ)y(k, l)y(k, l) H d(θ)
[0078] y(k, l) = [Y1(k, l) Y2(k, l) … Y M (k, l)] T
[0079] where d(θ) is the steering vector of the microphone array, θ ∈ {θ1, θ2, …, θ D}, θ is the azimuth where the sound source may appear, and M is the number of microphones in the microphone array.
[0080] Then, the angle of the maximum peak point is determined using the following formula:
[0081]
[0082] where is the single-source DOA estimation result at the l-th frame and the k-th frequency point. Finally, the single-source DOA estimation results at all frequencies and all times form the estimation result. By calculating the single-source DOA estimation result, the technical effect of providing a data basis for subsequent determination of the number of sound sources is achieved.
[0083] In summary, the method for estimating the number of sound sources provided by the present invention has the following technical effects:
[0084] 1. By receiving the first signal y m (t) through the microphone array and performing short-time Fourier transform to obtain the transform-domain signal Y m (k, l), then performing single-source DOA estimation on multiple frequency points in the transform-domain signal and obtaining the corresponding estimation results. Further, the frequencies estimated for different sound source azimuths within a preset time period are statistically analyzed to obtain the corresponding statistical frequency result Q(θ). Finally, multiple sound source azimuths with peak point frequencies greater than the threshold τ1 in the statistical frequency result Q(θ) are selected to determine the sound source number estimation result. In addition, the number of active sound sources in the sound source number estimation result is updated in real time. By performing single-source DOA estimation on each frequency point in the transform-domain signal after short-time Fourier transform and then statistically analyzing to estimate the number of sound sources, the technical effects of quickly and accurately estimating the number of sound sources in a complex acoustic environment and reducing the sound source number estimation error are achieved.
[0085] 2. By performing single-source DOA estimation on each frequency point in the transformed-domain signal after short-time Fourier transform in sequence, and then estimating the number of sound sources through statistical analysis, eigenvalue decomposition of the covariance matrix is avoided, achieving the technical objectives of reducing the computational complexity and improving the estimation speed of the number of sound sources.
[0086] Embodiment 2
[0087] Based on the method for estimating the number of sound sources in the foregoing embodiment and with the same inventive concept, the present invention also provides a system for estimating the number of sound sources. Please refer to the attached Figure 3 , the system includes:
[0088] The first acquisition unit 11 is configured to receive and acquire a first signal y m (t) through a microphone array;
[0089] The second acquisition unit 12 is configured to perform short-time Fourier transform on the first signal to acquire a transformed-domain signal Y m (k, l);
[0090] The third acquisition unit 13 is configured to perform single-source DOA estimation on multiple frequency points in the transformed-domain signal respectively to acquire an estimation result;
[0091] The fourth acquisition unit 14 is configured to statistically acquire the frequencies of different sound source azimuths estimated within a preset time period in the estimation result to acquire a statistical frequency result Q(θ);
[0092] The fifth acquisition unit 15 is configured to find peak points in the statistical frequency result Q(θ), obtain multiple sound source azimuths with the estimated frequencies greater than a threshold τ1, and acquire a sound source number estimation result;
[0093] The first update unit 16 is configured to update the number of currently active sound sources based on the sound source number estimation result.
[0094] Further, the system further includes:
[0095] The first calculation unit is configured to calculate:
[0096] P k,l (θ) = d H (θ)y(k, l)y(k, l) H d(θ), where d(θ) is the manifold vector of the microphone array, θ ∈ {θ1, θ2,..., θ D} and θ is the azimuth where the sound source may appear;
[0097] A second calculation unit, which is used to calculate:
[0098] y(k, l) = [Y1(k, l) Y2(k, l) … Y M (k, l)] T , where M is the number of microphones in the microphone array;
[0099] A second execution unit, which is used to find the angle where the maximum peak point appears, as follows:
[0100]
[0101] A second setting unit, which is used to That is, the single-source wave arrival azimuth estimation result at the l-th frame and the k-th frequency point;
[0102] A sixth acquisition unit, which is used to calculate and obtain all single-source wave arrival azimuth estimation results at multiple frequency points during all time to obtain the estimation results.
[0103] Furthermore, the system further includes:
[0104] A seventh acquisition unit, which is used to obtain all single-source wave arrival azimuth estimation results at multiple frequency points during all time according to the estimation results;
[0105] An eighth acquisition unit, which is used to statistically obtain the frequency Q(θ) at which all single-source wave arrival azimuths are estimated, where θ ∈ {θ1, θ2, …, θ D}.
[0106] Furthermore, the system further includes:
[0107] A third execution unit, which is used to find and obtain all peak points in the statistical frequency result Q(θ);
[0108] A fourth execution unit, which is used to select all peak points greater than the threshold τ1, denoted as θ i1 , θ i2 , …, θ iI ;
[0109] A first construction unit, which is used to construct and obtain a first vector θ cand , θ cand with a length of D, as follows:
[0110]
[0111] A third setting unit, which is configured to use the first vector as the sound source number estimation result.
[0112] Furthermore, the system further includes:
[0113] A second update unit, which is configured to perform updates using activation, retention, and suspension strategies to update the number of currently active sound sources.
[0114] Furthermore, the system further includes:
[0115] A fourth setting unit, which is configured to set the input parameter to θ active and θ cand , where the initial value of θ active is a zero vector and has the same length as the first vector;
[0116] A fifth execution unit, which is configured to subtract 1 from the values of all elements in θ active and assign the elements with values less than 0 to 0;
[0117] A fifth setting unit, which is configured to add θ active and θ cand and use the obtained result as the updated θ active ;
[0118] A sixth execution unit, which is configured to assign the elements with values greater than the threshold τ active in the updated θ max to τ max ;
[0119] A seventh execution unit, which is configured to count the number S of elements with values exceeding the threshold τ2 in the updated θ active ;
[0120] A sixth setting unit, which is configured to use the updated θ active and S as the output result, and S is the number of currently active sound sources.
[0121] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The foregoing Figure 1The method for estimating the number of sound sources and the specific example in the first embodiment are equally applicable to the system for estimating the number of sound sources in this embodiment. Through the foregoing detailed description of the method for estimating the number of sound sources, those skilled in the art can clearly know the system for estimating the number of sound sources in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method section.
[0122] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0123] Exemplary electronic device
[0124] Reference is made below Figure 4 to describe the electronic device of the present invention.
[0125] Figure 4 FIG. shows a schematic structural diagram of an electronic device according to the present invention.
[0126] Based on the inventive concept of the method for estimating the number of sound sources in the foregoing embodiment, the present invention further provides a system for estimating the number of sound sources, on which a computer program is stored, and when the program is executed by a processor, the steps of any of the methods for estimating the number of sound sources described above are implemented.
[0127] Among them, in Figure 4 , the bus architecture (represented by bus 300), bus 300 may include any number of interconnected buses and bridges, and bus 300 links various circuits including one or more processors represented by processor 302 and a memory represented by memory 304 together. Bus 300 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, will not be further described herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, which provides a unit for communicating with various other devices on a transmission medium.
[0128] Processor 302 is responsible for managing bus 300 and general processing, while memory 304 may be used to store data used by processor 302 when performing operations.
[0129] The present invention provides a method for estimating the number of sound sources, which is applied to a system for estimating the number of sound sources. Wherein, the method includes: receiving a first signal y m through a microphone array; performing a short-time Fourier transform on the first signal to obtain a transformed-domain signal Y m (k, l); performing single sound source DOA estimation on multiple frequency points in the transformed-domain signal to obtain an estimation result; within the estimation result, statistically obtaining the frequencies at which different sound source directions are estimated within a preset time period to obtain a statistical frequency result Q(θ); searching for peak points in the statistical frequency result Q(θ), and obtaining multiple sound source directions at which the estimated frequencies are greater than a threshold value τ1 to obtain a sound source number estimation result; based on the sound source number estimation result, updating the number of currently active sound sources. It solves the technical problem that the limitation of the existing method is that only Gaussian white noise and no reverberation are considered in the signal model, and when the reverberation in the environment is relatively strong or the background noise does not meet the requirements of Gaussian white noise, there is a large estimation error in the number of sound sources, which in turn affects the subsequent speech signal processing effect. By sequentially performing single sound source DOA estimation on each frequency point in the transformed-domain signal after short-time Fourier transform, and then statistically analyzing to estimate the number of sound sources, the technical effect of quickly and accurately estimating the number of sound sources in a complex acoustic environment and reducing the sound source number estimation error is achieved.
[0130] The present invention also provides an electronic device, which includes a processor and a memory;
[0131] The memory is used for storage;
[0132] The processor is used to execute the method described in any one of the above-mentioned Embodiment 1 by calling.
[0133] The present invention also provides a computer program product, including a computer program and / or instructions, and when the computer program and / or instructions are executed by a processor, the steps of the method described in any one of the above-mentioned Embodiment 1 are implemented.
[0134] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete software embodiment, a complete hardware embodiment, or an embodiment combining software and hardware aspects. In addition, the present invention is in the form of a computer program product that can be implemented on one or more computer-usable storage media containing computer-usable program code. The computer-usable storage media include, but are not limited to, various media that can store program code, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disk storage, compact disc read-only memory (CD-ROM), optical storage, etc.
[0135] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction system that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts.
[0138] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for estimating the number of sound sources, characterized in that, The method includes: Receive the first signal y m (t); Perform a short-time Fourier transform on the first signal to obtain a transformed-domain signal Y m (k, l); Performing single-source DOA estimation on multiple frequency points within the transform-domain signal to obtain an estimation result; Within the estimation result, statistically obtaining the frequencies estimated for different sound source azimuths within a preset time period to obtain a statistical frequency result Q(θ); Searching for peak points within the statistical frequency result Q(θ) to obtain multiple sound source azimuths for which the estimated frequency is greater than a threshold value τ1, thereby obtaining an estimation result of the number of sound sources; Based on the estimation result of the number of sound sources, updating the currently active number of sound sources; The single-source DOA estimation includes: Calculation: P k,l (θ) = d H (θ)y(k, l)y(k, l) H d(θ), where d(θ) is the steering vector of the microphone array, θ ∈ {θ1, θ2, …, θ D}, θ is the azimuth where the sound source may appear; Calculation: y(k, l) = [Y1(k, l) Y2(k, l) … Y M (k, l)] T , where M is the number of microphones in the microphone array; Searching for the angle at which the maximum peak point appears, as follows: That is the arrival direction estimation result of a single sound source at the l-th frame and the k-th frequency point; Calculating to obtain all single-source wave arrival azimuth estimation results at multiple frequency points over all time to obtain the estimation result.
2. The method according to claim 1, wherein The short-time Fourier transform is calculated by the following formula: where m is the m-th microphone within the microphone array, k is the frequency index, l is the frame index, w(n) is a window function, N is the window length, and L is the hop length.
3. The method according to claim 1, characterized in that, The statistically obtaining the frequencies estimated for different sound source azimuths within a preset time period includes: Based on the estimation result, obtaining all single-source wave arrival azimuth estimation results at multiple frequency points over all time; Based on all the single-source direction-of-arrival estimation results, the frequency Q(θ) at which all the single-source directions-of-arrival are estimated is statistically obtained, where θ ∈ {θ1, θ2, …, θ D}.
4. The method according to claim 1, characterized in that, The searching for peak points within the statistical frequency result Q(θ) to obtain multiple sound source azimuths for which the estimated frequency is greater than a threshold value τ1 includes: Searching to obtain all peak points within the statistical frequency result Q(θ); Select all the peak points greater than the threshold value τ1, denoted as θ i1 , θ i2 , …, θ iI ; Construct to obtain the first vector θ cand , θ cand with a length of D, as follows: Taking the first vector as the estimation result of the number of sound sources.
5. The method according to claim 4, characterized in that, The updating the currently active number of sound sources includes: Using activation, retention, and termination strategies for updating to update the currently active number of sound sources.
6. The method according to claim 5, characterized in that, The using activation, retention, and termination strategies for updating includes: Set the input parameter to θ active and θ cand , where the initial value of θ active is a zero vector and has the same length as the length of the first vector; Subtract 1 from the value of all elements in θ active and assign 0 to the elements whose values are less than 0; Add θ active and θ cand together, and use the result as the updated θ active ; Assign the updated θ active where all values greater than the threshold τ max to τ max ; Statistically updated θ active The number S of elements in active whose all values exceed the threshold τ2; The updated θ active and S as the output result, where S is the number of currently active sound sources.
7. An apparatus for estimating the number of sound sources, characterized in that, The apparatus is applied to the method according to any one of claims 1 to 6, and the apparatus includes: First acquisition unit: The first acquisition unit is configured to receive and acquire a first signal y m (t); Second acquisition unit: The second acquisition unit is configured to perform short-time Fourier transform on the first signal to obtain a transform-domain signal Y m (k, l); A third obtaining unit: The third obtaining unit is configured to perform single-source DOA estimation on multiple frequency points within the transform-domain signal to obtain an estimation result; A fourth obtaining unit: The fourth obtaining unit is configured to statistically obtain the frequencies estimated for different sound source azimuths within a preset time period within the estimation result to obtain a statistical frequency result Q(θ); A fifth obtaining unit: The fifth obtaining unit is configured to search for peak points within the statistical frequency result Q(θ) to obtain multiple sound source azimuths for which the estimated frequency is greater than a threshold value τ1, thereby obtaining an estimation result of the number of sound sources; A first updating unit: The first updating unit is configured to update the currently active number of sound sources based on the estimation result of the number of sound sources.
8. An electronic device, characterized in that, Including a processor and a memory; The memory is used for storage; The processor is configured to execute the method according to any one of claims 1 to 6 by calling.
9. A computer program product, comprising a computer program and / or instructions, characterized in that, When the computer program and / or instruction is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.