A method and apparatus for simultaneous counting and localization of multiple sound sources

By using frequency domain processing of microphone arrays and peak iterative search algorithms, the problems of low accuracy and high computational complexity in multi-source localization are solved, achieving efficient multi-source counting and localization. In particular, it can still accurately estimate the number and direction of sound sources even in the absence of prior knowledge and data.

CN120669197BActive Publication Date: 2025-11-14HANGZHOU AIHUA INSTR
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510891204.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-14
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy and high computational complexity in multi-source localization, and it is difficult to balance the estimation of the number of sound sources with localization, especially when there is a lack of prior knowledge and sufficient data.

Method used

By acquiring the time-domain data of the microphone array, preprocessing it into the frequency domain in frames, selecting the frequency band of interest, using phase information to initially estimate the direction of the sound source, and estimating the number of sound sources through a peak iterative search algorithm, the direction of each sound source is finally determined, thus realizing the counting and localization of multiple sound sources.

Benefits of technology

It achieves multi-source counting and localization with low computational cost and high accuracy, avoiding complex eigenvalue decomposition and large amount of data training, and improving the accuracy of sound source number estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669197B_ABST
    Figure CN120669197B_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for simultaneous counting and localization of multiple sound sources, relating to the field of sound source counting and localization technology. It solves the problems of low sound source localization accuracy, high computational complexity, and difficulty in simultaneously estimating the number of sound sources and localization in the prior art. The method includes: collecting data from multiple sound sources based on a microphone array; firstly, performing frame-based preprocessing on the time-domain data; selecting the processing frequency band of interest and selecting frequencies with higher energy within the frequency band; preliminarily estimating the direction of the sound sources using phase information; then summarizing the sound source direction estimation results of multiple time frames into a histogram; and finally using a peak iterative search algorithm to extract the number of sound sources and their corresponding directions from the histogram. This achieves simultaneous counting and localization of multiple sound sources, ensuring not only high counting and localization accuracy but also reducing computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound source counting and localization technology, and in particular to a method and apparatus for simultaneous counting and localization of multiple sound sources. Background Technology

[0002] Sound source localization technology uses microphone arrays to measure sound signals and processes these signals using algorithms to determine the direction of arrival of the sound source. Multi-source localization involves locating multiple sound sources simultaneously; however, because multiple sound sources often have different sound pressure levels, stronger sources can mask weaker ones, making localization difficult, and the number of sound sources is often unknown. In recent years, various technologies for multi-source localization have been developed, such as sound source cancellation and deep learning, but these technologies still have the following problems:

[0003] 1. Low accuracy of sound source localization estimation: Some existing technologies require sufficient prior knowledge and data for estimation. When these conditions are insufficient, the estimation accuracy drops sharply.

[0004] 2. High computational complexity: For example, deep learning algorithms require a large amount of various sound source data for training, which results in excessively high computational costs and makes them difficult to implement.

[0005] 3. Difficulty in simultaneously estimating and locating the number of sound sources: As mentioned earlier, the sound source elimination technology can locate multiple sound sources, but it cannot estimate the number of sound sources. Summary of the Invention

[0006] The purpose of this application is to overcome the problems of low sound source localization accuracy, high computational complexity, and difficulty in simultaneously estimating the number of sound sources and localizing them in the existing technology, and to provide a method and apparatus for simultaneously counting and localizing multiple sound sources.

[0007] Firstly, a method for simultaneously counting and locating multiple sound sources is provided, including:

[0008] Acquire the time-domain data collected by the microphone array, and preprocess the time-domain data into the frequency domain by dividing it into frames;

[0009] Select the processing frequency band of interest, and select a calculation frequency from the processing frequency band based on the signal energy;

[0010] The direction of the sound source is initially estimated using the phase information of the selected calculation frequency;

[0011] The sound source direction estimation results for each time frame are summarized and the number of sound sources is estimated based on the peak iterative search algorithm;

[0012] Determine the direction of each sound source to achieve counting and localization of multiple sound sources.

[0013] In some possible implementations, the signal time-domain model of the time-domain data is as follows:

[0014] ;

[0015] Where, x m (t) represents the time-domain data of the signal received by each microphone at time t, where m = 1, 2, 3, ..., M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, n is the additive noise, P is the number of sound sources, and θ is the incident angle of the sound source.

[0016] In some possible implementations, the time-domain data is preprocessed into the frequency domain by frame division, including:

[0017] The time-domain data is divided into L frames according to the time length, and each frame processes x data. l Where l = 1, 2, ..., L;

[0018] The processed data for each frame is converted to the frequency domain using Fourier transform:

[0019] ;

[0020] Among them, X ml (f) k ) is x m (t) The converted frequency domain data, where N is additive noise, k=1,2,…K, and K is the number of frequency points within the frequency band.

[0021] In some possible implementations, the calculation frequency is selected from the processing bandwidth based on the signal energy, including:

[0022] Calculate the signal energy at each frequency point within the processing frequency band, wherein the signal energy E at each frequency point within the frequency band is... k The calculation formula is:

[0023] ;

[0024] All frequencies whose signal energy is greater than a preset energy threshold are selected as the calculation frequencies.

[0025] In some possible implementations, the direction of the sound source is initially estimated using phase information from the selected calculation frequency, including:

[0026] Calculate the cross-power spectrum phase between microphone pairs in the microphone array. :

[0027] ;

[0028] in, The cross-power spectrum is shown, and * represents complex conjugate.

[0029] Calculate the phase rotation factor :

[0030] ;

[0031] in, Let φ be the relative delay difference between microphone pair {1,2} and microphone pair {m,m+1}, where φ∈[0,2).

[0032] Constructing spectral functions :

[0033] ;

[0034] By searching for spectral peaks and maximizing the spectral function Obtain sound source direction estimation :

[0035] ;

[0036] Here, argmax represents finding the independent variable that maximizes the function. This is a modulo operation.

[0037] In some possible implementations, the source direction estimation results for each time frame are summarized and the number of sound sources is estimated based on a peak iterative search algorithm, including:

[0038] The sound source direction estimation results for each time frame The data is summarized and plotted as a histogram;

[0039] Set a sound source counter =1 indicates the first sound source that has been identified so far;

[0040] Initialize the first peak position :

[0041] ;

[0042] Where argmax represents finding the independent variable that maximizes the function;

[0043] Set dynamic threshold , Thresholds set by the user This indicates taking the maximum value. For in position The peak value of the histogram;

[0044] Search for the next peak in the histogram. ;

[0045] like ,and Then the peak value is considered to be As a new source of sound, among which, In position The peak value of the histogram. For the identified peak values, Let B be the minimum angular interval between sound sources, and let B be the total number of bars in the histogram. This is a union operation, where the subscript j=1 indicates the starting index of the union operation, and the superscript c... s Indicates the terminating index of the union operation;

[0046] If a new sound source is discovered, update. , ;

[0047] If no new sound source is found or the preset maximum number of sound sources is reached, the search stops, and the number of sound sources is obtained. .

[0048] In some possible implementations, the direction corresponding to each sound source is determined, including: finding the estimated direction value corresponding to each sound source based on the estimated number of sound sources, so as to realize the counting and localization of multiple sound sources.

[0049] Secondly, a device for simultaneously counting and locating multiple sound sources is provided, comprising:

[0050] The data acquisition and preprocessing module is used to acquire time-domain data collected by the microphone array and preprocess the time-domain data into the frequency domain by frame.

[0051] The selection module is used to select the processing frequency band of interest and select a calculation frequency from the processing frequency band based on the signal energy.

[0052] The sound source direction estimation module is used to make a preliminary estimate of the sound source direction using the phase information of the selected calculation frequency;

[0053] The sound source number estimation module is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on the peak iterative search algorithm;

[0054] The matching module is used to determine the direction of each sound source in order to count and locate multiple sound sources.

[0055] Thirdly, a computer-readable storage medium is provided that stores program code for execution by a device, the program code including steps for performing a method as described in any of the implementations of the first aspect above.

[0056] Fourthly, an electronic device is provided, the electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the method as described in any of the implementations of the first aspect above.

[0057] This application has the following advantages: Compared with other subspace-based technologies in the prior art, the multi-source counting and localization method of this application does not require complex operations such as eigenvalue decomposition, and also ensures high counting and localization accuracy. Secondly, by summarizing multiple results of preliminary source estimation, and then using peak iterative search to extract the number of sources, it makes full use of source data information, without requiring a large amount of data training or relying on various prior knowledge, and has obvious advantages in terms of computational cost and estimation accuracy, while realizing the counting and localization of multiple sources. Attached Figure Description

[0058] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an undue limitation of this application.

[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1 This is a flowchart of the method for simultaneous counting and localization of multiple sound sources according to Embodiment 1 of this application;

[0061] Figure 2 This is a histogram of the sound source direction estimation results for each time frame in the method for simultaneous counting and localization of multiple sound sources in Embodiment 1 of this application.

[0062] Figure 3 This is a cumulative distribution diagram of sound source angle error in the method for simultaneous counting and localization of multiple sound sources according to Embodiment 1 of this application;

[0063] Figure 4 This is a structural block diagram of the device for simultaneous counting and locating of multiple sound sources according to Embodiment 2 of this application;

[0064] Figure 5 This is a schematic diagram of the internal structure of the electronic device according to Embodiment 4 of this application.

[0065] Figure label:

[0066] 100. Data acquisition and preprocessing module; 200. Selection module; 300. Sound source direction estimation module; 400. Sound source number estimation module; 500. Matching module. Detailed Implementation

[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] Example 1

[0069] like Figure 1 As shown, the method for simultaneous counting and localization of multiple sound sources according to Embodiment 1 of this application includes:

[0070] S100: Acquire the time-domain data collected by the microphone array, and preprocess the time-domain data into the frequency domain by dividing it into frames.

[0071] Assuming there are P broadband sound sources in space incident from different directional angles θ, the time-domain model of the signal received by each microphone at time t is as follows:

[0072] ;

[0073] Where, x m (t) represents the time-domain data of the signal received by each microphone at time t, where m = 1, 2, 3, ..., M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, and n is the additive noise. It should be noted that the parameters with subscripts in the above signal time-domain model have the same definition as those without subscripts, for example: τ is the signal delay, τ... m This represents the signal delay of the m-th microphone, where m has already been defined.

[0074] Next, the time-domain data needs to be preprocessed into the frequency domain by dividing it into frames. First, the time-domain data is divided into L frames according to the time length, and each frame processes x data. l , where l=1,2,…L.

[0075] Using Fourier transform, each frame of processed data is converted to the frequency domain. The above equation can be written in the frequency domain as:

[0076] ;

[0077] Among them, X ml (f) k ) is x m(t) The transformed frequency domain data, k=1,2,…K, where K is the number of frequency points within the frequency band. It should be noted that when transforming the time domain model to the frequency domain, according to convention, the parameters in the time domain model are lowercase letters, while those in the frequency domain are uppercase letters (such as xX, nN), both representing the same variable. To avoid redundancy, there is no need to redefine them. For example, in the time domain model, n is additive noise, while in the frequency domain, N is additive noise.

[0078] S200. Select the processing frequency band of interest, and select the calculation frequency from the processing frequency band according to the signal energy.

[0079] Selecting the processing frequency band of interest, the signal energy at each frequency point within the processing frequency band can be expressed as:

[0080] ;

[0081] Among them, E k Let X be the signal energy at the k-th frequency point within the frequency band. ml (f) k ) is x m (t) The converted frequency domain data.

[0082] To select a suitable calculation frequency, an energy threshold is introduced. Select a frequency band that satisfies E k > All frequencies are used as the frequencies for subsequent positioning calculations. The energy threshold is set manually based on the overall energy distribution of each calculated frequency point, and is not a fixed value.

[0083] S300. Use the phase information of the selected calculation frequency to make a preliminary estimate of the direction of the sound source.

[0084] It should be noted that, to maintain the simplicity of the formula, all parameters in the following text are defined in each time frame of data, and therefore no subscript 'l' is added. Based on this, the preliminary estimation of the sound source direction includes the following steps:

[0085] Calculate the cross-power spectrum phase between microphone pairs in the microphone array. :

[0086] ;

[0087] in, is the cross-power spectrum, and * represents complex conjugate.

[0088] Calculate the phase rotation factor :

[0089] ;

[0090] in, Let φ be the relative delay difference between microphone pair {1,2} and microphone pair {m,m+1}, where φ∈[0,2].

[0091] Constructing spectral functions :

[0092] .

[0093] By searching for spectral peaks and maximizing the spectral function Obtain sound source direction estimation :

[0094] ;

[0095] Here, argmax represents finding the independent variable that maximizes the function. This is a modulo operation.

[0096] The method used in this embodiment simplifies the sound source direction estimation into a nonlinear optimization problem, which significantly improves computational efficiency and also has high estimation accuracy in complex environments.

[0097] S400. Summarize the sound source direction estimation results for each time frame and estimate the number of sound sources based on the peak iterative search algorithm. This includes the following steps:

[0098] S401. Estimate the sound source direction for each time frame. The data is summarized and plotted as a histogram. Assuming there are three broadband sound sources with incident directions of 60°, 120°, and 160°, the method in this embodiment is used to first perform a preliminary estimation of the sound source directions, and the histogram obtained from the summarized results is then analyzed. like Figure 2 As shown, the reason for the two peaks at 120° is that the algorithm divides the sound source point equally when calculating it. Therefore, around 120°, due to the sound source estimation by the algorithm, there may be sound source directions such as 119°, 120° and 121°.

[0099] S402. Estimating the number of sound sources based on the peak iterative search algorithm, specifically including the following steps:

[0100] S4021, Initialization:

[0101] Set a sound source counter =1 indicates the first sound source that has been identified so far;

[0102] Initialize the first peak position :

[0103] ;

[0104] Where argmax represents finding the independent variable that maximizes the function;

[0105] Set dynamic threshold , Thresholds set by the user These are empirical values, and their values ​​are set based on the specific distribution results of the histogram. This indicates taking the maximum value. For in position The peak value of the histogram.

[0106] S4022, Iterative search for the next peak:

[0107] First, search for the next peak in the histogram. ;

[0108] If the search found peak value ,in, The dynamic threshold set in step S4021, and The purpose is to avoid being too close to already identified peak values, hence this minimum angular interval condition is set. If both of the above conditions are met simultaneously, the peak value is considered to be... As a new source of sound, among which, In position The peak value of the histogram. For the identified peak values, Let B be the minimum angular interval between sound sources, and let B be the total number of bars in the histogram. This is the union operation, used to combine elements from multiple sets. The subscript j=1 indicates the starting index of the union operation, and the superscript c... s The terminating index (usually a positive integer) represents the union operation.

[0109] If a new sound source is discovered during the above steps, then... and renew: , ;

[0110] If no new sound source is found or the preset maximum number of sound sources is reached during the above steps, the search stops, and the number of sound sources is obtained. It can be seen that Figure 3 To estimate the probability density function of the histogram results, three distinct peaks were extracted, corresponding to angles of 60°, 120°, and 160°, respectively. This enabled the estimation and localization of the number of three sound sources, verifying the effectiveness of the method of this invention.

[0111] The sound source number estimation method used in this embodiment is different from the traditional method. It directly extracts the sound source number from multiple estimation results through peak search and combines the two judgment conditions of peak value and minimum angular interval between sound sources. This not only reduces the computational complexity, but also effectively improves the estimation accuracy of the sound source number.

[0112] S500: Determine the direction of each sound source to achieve counting and localization of multiple sound sources.

[0113] Specifically, in step S300, the initial estimation of the sound source direction indicates that the sound source direction has been located. However, due to the large number of results, the exact number of sound sources and their corresponding directions are unknown. Step S400 first estimates the exact number of sound sources and then finds the estimated direction value for each sound source, thus completing the simultaneous counting and localization of multiple sound sources.

[0114] For example: Suppose there are 3 sound sources with incident angles of 30°, 60°, and 90°. The preliminary estimate in step 3 may have many angles, including the directional results of the 3 sound sources, but the actual number of sound sources and their corresponding directions are unknown. Therefore, step S400 is needed to estimate the number of sound sources and find their corresponding angle values, thus achieving simultaneous counting and localization.

[0115] In this embodiment, no complex operations such as eigenvalue decomposition are required, and high counting and positioning accuracy are guaranteed. By summarizing multiple results of preliminary sound source estimation, and then using peak iterative search to extract the number of sound sources, the sound source data information is fully utilized. It does not require a large amount of data training or rely on various prior knowledge, and has obvious advantages in terms of computational cost and estimation accuracy. It has also been verified by simulation and experimental data.

[0116] Example 2

[0117] like Figure 4 As shown, the device for simultaneous counting and locating of multiple sound sources according to Embodiment 2 of this application includes:

[0118] The data acquisition and preprocessing module 100 is used to acquire time-domain data collected by the microphone array and preprocess the time-domain data into the frequency domain in frames.

[0119] Selection module 200 is used to select the processing frequency band of interest and select a calculation frequency from the processing frequency band according to the signal energy.

[0120] The sound source direction estimation module 300 is used to initially estimate the sound source direction using the phase information of the selected calculation frequency;

[0121] The sound source number estimation module 400 is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on the peak iterative search algorithm;

[0122] The matching module 500 is used to determine the direction of each sound source in order to realize the counting and localization of multiple sound sources.

[0123] It should be noted that other specific implementations of the device for simultaneous counting and locating of multiple sound sources in this embodiment can be found in the specific implementations of the method for simultaneous counting and locating of multiple sound sources described above. To avoid redundancy, they will not be repeated here.

[0124] Example 3

[0125] This application relates to a computer-readable storage medium in embodiment 3, which stores program code for execution by a device, the program code including steps for performing the method as described in any implementation of embodiment 1 of this application;

[0126] The computer-readable storage medium may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM); the computer-readable storage medium may store program code, and when the program stored in the computer-readable storage medium is executed by a processor, the processor is used to perform the steps of the method in any of the implementations of Embodiment 1 of this application.

[0127] Example 4

[0128] like Figure 5 As shown, an electronic device according to Embodiment 4 of this application includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the method in any of the implementations in Embodiment 1 of this application.

[0129] The processor can be a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute related programs to implement the method in any of the implementations of Embodiment 1 of this application.

[0130] The processor can also be an integrated circuit electronic device with signal processing capabilities. In implementation, each step of the method in any of the implementations of Embodiment 1 of this application can be completed by the integrated logic circuitry in the processor's hardware or by software instructions.

[0131] The aforementioned processor can also be a general-purpose processor, digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the functions required by the units included in the data processing apparatus of the embodiments of this application, or executes the methods in any implementation of Embodiment 1 of this application.

[0132] The above are merely preferred embodiments of this application; however, the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and its improved concept, should be covered within the scope of protection of this application.

Claims

1. A method for simultaneously counting and locating multiple sound sources, characterized in that, include: Acquire the time-domain data collected by the microphone array, and preprocess the time-domain data into the frequency domain by dividing it into frames; Select the processing frequency band of interest, and select a calculation frequency from the processing frequency band based on the signal energy; The direction of the sound source is initially estimated using the phase information of the selected calculation frequency; The sound source direction estimation results for each time frame are summarized and the number of sound sources is estimated based on the peak iterative search algorithm; Determine the direction of each sound source to achieve counting and localization of multiple sound sources; The process involves summarizing the source direction estimation results for each time frame and estimating the number of sound sources based on a peak iterative search algorithm, including: The sound source direction estimation results for each time frame The data is summarized and plotted as a histogram; Set a sound source counter =1 indicates the first sound source that has been identified so far; Initialize the first peak position : ; Where argmax represents finding the independent variable that maximizes the function; Set dynamic threshold , Thresholds set by the user This indicates taking the maximum value. For in position The peak value of the histogram; Search for the next peak in the histogram. ; like ,and Then the peak value is considered to be As a new source of sound, among which, In position The peak value of the histogram. For the identified peak values, Let B be the minimum angular interval between sound sources, and let B be the total number of bars in the histogram. This is a union operation, where the subscript j=1 indicates the starting index of the union operation, and the superscript cs indicates the ending index of the union operation. If a new sound source is discovered, update. , ; If no new sound source is found or the preset maximum number of sound sources is reached, the search stops, and the number of sound sources is obtained. .

2. The method for simultaneous counting and locating of multiple sound sources according to claim 1, characterized in that, The signal time-domain model of the time-domain data is: ; Where, x m (t) represents the time-domain data of the signal received by each microphone at time t, where m = 1, 2, 3, ..., M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, n is the additive noise, P is the number of sound sources, and θ is the incident angle of the sound source.

3. The method for simultaneous counting and locating of multiple sound sources according to claim 2, characterized in that, The time-domain data is preprocessed into frames and then transferred to the frequency domain, including: The time-domain data is divided into L frames according to the time length, and each frame processes x data. l Where l = 1, 2, ..., L; The processed data for each frame is converted to the frequency domain using Fourier transform: ; Among them, X ml (f) k ) is x m (t) The converted frequency domain data, where N is additive noise, k=1,2,…K, and K is the number of frequency points within the frequency band.

4. The method for simultaneous counting and locating of multiple sound sources according to claim 3, characterized in that, Selecting a calculation frequency from the processing bandwidth based on signal energy includes: Calculate the signal energy at each frequency point within the processing frequency band, wherein the signal energy E at each frequency point within the frequency band is... k The calculation formula is: ; All frequencies whose signal energy is greater than a preset energy threshold are selected as the calculation frequencies.

5. The method for simultaneous counting and locating of multiple sound sources according to claim 4, characterized in that, The direction of the sound source is initially estimated using the phase information of the selected calculation frequency, including: Calculate the cross-power spectrum phase between microphone pairs in the microphone array. : ; in, The cross-power spectrum is shown, and * represents complex conjugate. Calculate the phase rotation factor : ; in, Let φ be the relative delay difference between microphone pair {1,2} and microphone pair {m,m+1}, where φ∈[0,2). Constructing spectral functions : ; By searching for spectral peaks and maximizing the spectral function Obtain sound source direction estimation : ; Here, argmax represents finding the independent variable that maximizes the function. This is a modulo operation.

6. The method for simultaneous counting and locating of multiple sound sources according to any one of claims 1-5, characterized in that, Determine the direction corresponding to each sound source, including: based on the estimated number of sound sources, find the estimated direction value corresponding to each sound source, so as to realize the counting and localization of multiple sound sources.

7. A device for simultaneous counting and locating of multiple sound sources, characterized in that, include: The data acquisition and preprocessing module is used to acquire time-domain data collected by the microphone array and preprocess the time-domain data into the frequency domain by frame. The selection module is used to select the processing frequency band of interest and select a calculation frequency from the processing frequency band based on the signal energy. The sound source direction estimation module is used to make a preliminary estimate of the sound source direction using the phase information of the selected calculation frequency; The sound source number estimation module is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on the peak iterative search algorithm; The matching module is used to determine the direction of each sound source in order to count and locate multiple sound sources; The process involves summarizing the source direction estimation results for each time frame and estimating the number of sound sources based on a peak iterative search algorithm, including: The sound source direction estimation results for each time frame The data is summarized and plotted as a histogram; Set a sound source counter =1 indicates the first sound source that has been identified so far; Initialize the first peak position : ; Where argmax represents finding the independent variable that maximizes the function; Set dynamic threshold , Thresholds set by the user This indicates taking the maximum value. For in position The peak value of the histogram; Search for the next peak in the histogram. ; like ,and Then the peak value is considered to be As a new source of sound, among which, In position The peak value of the histogram. For the identified peak values, Let B be the minimum angular interval between sound sources, and let B be the total number of bars in the histogram. This is a union operation, where the subscript j=1 indicates the starting index of the union operation, and the superscript cs indicates the ending index of the union operation. If a new sound source is discovered, update. , ; If no new sound source is found or the preset maximum number of sound sources is reached, the search stops, and the number of sound sources is obtained. .

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code for execution by the device, the program code including steps for performing the method as described in any one of claims 1-6.

9. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Indoor multi-sound-source positioning method based on DOA estimation and DOA association

    CN120214697A