Method and device for simultaneously counting and positioning multiple sound sources

Through time domain data processing and frequency domain analysis of the microphone array, combined with a peak iterative search algorithm, high-precision counting and positioning of multiple sound sources is achieved, solving the problems of low sound source positioning accuracy and high computational complexity in the existing technology and simplifying the sound source number estimation process.

CN120669197AActive Publication Date: 2025-09-19HANGZHOU AIHUA INSTR
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510891204.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-19
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing technology of multi-sound source localization has the problems of low sound source localization accuracy, high computational complexity, and difficulty in balancing the estimation and localization of the number of sound sources, especially when strong sound sources mask weak sound sources and the number of sound sources is unknown.

Method used

By acquiring the time domain data of the microphone array, pre-processing it into the frequency domain by dividing the frames, using the signal energy selection calculation, using the phase information to preliminarily estimate the direction of the sound source, and combining the peak iterative search algorithm to estimate the number of sound sources, the number of each sound source is determined, and the counting and positioning of multiple sound sources is achieved.

Benefits of technology

The method for counting and locating multiple sound sources is implemented, which solves the problems of low sound source positioning accuracy, high computational complexity, and difficulty in balancing sound source number estimation and positioning in the existing technology, and achieves high-precision counting and positioning of multiple sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669197A_ABST
    Figure CN120669197A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for simultaneously counting and positioning multiple sound sources, relates to the technical field of sound source counting and positioning, and solves the problems of low sound source positioning precision, high calculation complexity, difficulty in giving consideration to sound source number estimation and positioning and the like in the prior art. The method comprises the following steps: collecting data of a plurality of sound sources based on a microphone array, firstly, carrying out framing preprocessing on time domain data, selecting an interested processing frequency band, selecting a frequency with relatively high energy in the frequency band, preliminarily estimating a sound source direction by utilizing phase information, summarizing sound source direction estimation results of a plurality of time frames into a histogram, and carrying out direction estimation on the sound sources according to the histogram; and finally, the number of sound sources and corresponding directions are extracted from the histogram by using a peak iterative search algorithm, so that simultaneous counting and positioning of multiple sound sources are realized, relatively high counting and positioning precision is ensured, and the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of sound source counting and positioning, and in particular to a method and device for simultaneously counting and positioning multiple sound sources. Background Art

[0002] Sound source localization technology uses a microphone array to measure sound signals and processes the measured sound signals through algorithms to obtain the direction of arrival of the sound source. Multi-source localization is the simultaneous localization of multiple sound sources. However, because multiple sound sources often have different sound pressure levels, strong sound sources can mask weak sound sources, making localization difficult. In addition, the number of sound sources is often unknown. In recent years, various technologies have been developed for multi-source localization, such as sound source elimination technology and deep learning. However, these technologies still have the following problems:

[0003] 1. Low sound source localization estimation accuracy: Some existing technologies require sufficient prior knowledge and data volume for estimation. When these conditions are insufficient, the estimation accuracy drops sharply.

[0004] 2. High computational complexity: Deep learning algorithms, for example, require a large amount of sound source data for training, which is computationally expensive and difficult to implement.

[0005] 3. It is difficult to estimate and locate the number of sound sources at the same time: For example, the sound source elimination technology mentioned above can locate multiple sound sources, but it cannot estimate the number of sound sources. Summary of the Invention

[0006] The purpose of this application is to overcome the problems existing in the prior art, such as low sound source localization accuracy, high computational complexity, and difficulty in balancing sound source number estimation and localization, and to provide a method and device for simultaneously counting and localizing multiple sound sources.

[0007] In a first aspect, a method for simultaneously counting and locating multiple sound sources is provided, comprising:

[0008] Acquire time domain data collected by the microphone array, and pre-process the time domain data into frequency domain by framing;

[0009] selecting a processing frequency band of interest, and selecting a calculation frequency from the processing frequency band based on signal energy;

[0010] The sound source direction is preliminarily estimated using the phase information of the selected calculation frequency;

[0011] The sound source direction estimation results of each time frame are summarized and the number of sound sources is estimated based on the peak iterative search algorithm;

[0012] Determine the direction corresponding to each sound source to achieve the counting and positioning of multiple sound sources.

[0013] In some possible implementations, the signal time domain model of the time domain data is:

[0014] ;

[0015] Among them, x m (t) is the time domain data of the signal received by each microphone at time t, m=1, 2, 3,…M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, n is the additive noise, P is the number of sound sources, and θ is the incident angle of the sound source.

[0016] In some possible implementations, pre-processing the time domain data into frequency domain frames includes:

[0017] The time domain data is divided into L frames according to the time length, and the processing data of each frame is x l , where l=1,2,…L;

[0018] Use Fourier transform to convert each frame of processed data into the frequency domain:

[0019] ;

[0020] Among them, X ml (f k ) is x m (t) The converted frequency domain data, N is the additive noise, k = 1, 2, …, K, K is the number of frequency points in the frequency band.

[0021] In some possible implementations, selecting a calculation frequency from a processing frequency band based on signal energy includes:

[0022] Calculate the signal energy of each frequency point in the processing frequency band, where the signal energy E of each frequency point in the frequency band is k The calculation formula is:

[0023] ;

[0024] All frequencies whose signal energy is greater than a preset energy threshold are selected as calculation frequencies.

[0025] In some possible implementations, preliminarily estimating the direction of the sound source using the phase information of the selected calculation frequency includes:

[0026] Calculate the cross-power spectrum phase between microphone pairs in the microphone array :

[0027] ;

[0028] in, is the cross power spectrum, * is the complex conjugate;

[0029] Calculate the phase rotation factor :

[0030] ;

[0031] in, is the relative delay difference between the microphone pair {1,2} and the microphone pair {m,m+1}, φ∈[0,2);

[0032] Constructing spectral functions :

[0033] ;

[0034] Search by spectral peak and maximize the spectral function Get the sound source direction estimate :

[0035] ;

[0036] Among them, argmax means finding the independent variable that maximizes the function. It is a modulo operation.

[0037] In some possible implementations, the sound source direction estimation results of each time frame are aggregated and the number of sound sources is estimated based on a peak iterative search algorithm, including:

[0038] The sound source direction estimation results of each time frame Summarize and make statistics as histogram;

[0039] Setting the sound source counter =1, indicating the first sound source currently identified;

[0040] Initialize the first peak position :

[0041] ;

[0042] Among them, argmax means finding the independent variable that maximizes the function;

[0043] Setting dynamic thresholds , For user-set thresholds, Indicates taking the maximum value, For the location The peak of the histogram;

[0044] Search for the next peak in the histogram ;

[0045] like ,and , then the peak is a new sound source, among which, In position The peak of the histogram, is the identified peak, is the minimum angular interval between sound sources, B is the total number of columns in the histogram, It is a union operation, where the subscript j=1 represents the starting index of the union operation and the superscript c s Indicates the ending index of the union operation;

[0046] If a new sound source is found, update , ;

[0047] If no new sound source is found or the preset maximum number of sound sources is reached, the search stops and the number of sound sources is obtained. .

[0048] In some possible implementations, determining the direction corresponding to each sound source includes: finding a direction estimation value corresponding to each sound source based on the estimated number of sound sources, so as to achieve counting and positioning of multiple sound sources.

[0049] In a second aspect, a device for simultaneously counting and locating multiple sound sources is provided, comprising:

[0050] A data acquisition and preprocessing module is used to acquire time domain data collected by the microphone array and preprocess the time domain data into frequency domain frames;

[0051] a selection module for selecting a processing frequency band of interest and selecting a calculation frequency from the processing frequency band according to signal energy;

[0052] A sound source direction estimation module is used to preliminarily estimate the sound source direction using the phase information of the selected calculation frequency;

[0053] A sound source number estimation module is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on the peak iterative search algorithm;

[0054] The matching module is used to determine the direction corresponding to each sound source to achieve the counting and positioning of multiple sound sources.

[0055] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable medium stores program code for execution by a device, the program code including steps for executing the method in any one of the implementations of the first aspect.

[0056] In a fourth aspect, an electronic device is provided, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements a method as in any one of the implementations of the first aspect described above.

[0057] The present application has the following beneficial effects: compared with other subspace technologies in the prior art, the multi-sound source counting and positioning method of the present application does not require complex operations such as eigenvalue decomposition, and also ensures high counting and positioning accuracy. Secondly, by summarizing multiple results of preliminary sound source estimation and then using peak iterative search to extract the number of sound sources, the sound source data information is fully utilized, without the need for large-scale data training and without relying on various types of prior knowledge. It has obvious advantages in terms of computational cost and estimation accuracy, and at the same time realizes the counting and positioning of multiple sound sources. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The drawings that constitute a part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application.

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0060] Figure 1 This is a flow chart of the method for simultaneously counting and locating multiple sound sources in Example 1 of the present application;

[0061] Figure 2 is a histogram of the sound source direction estimation results for each time frame in the method for simultaneous counting and positioning of multiple sound sources in Example 1 of the present application;

[0062] Figure 3 This is a cumulative distribution diagram of sound source angle errors in the method for simultaneously counting and locating multiple sound sources in Example 1 of the present application;

[0063] Figure 4 This is a structural block diagram of a device for simultaneously counting and locating multiple sound sources according to Example 2 of the present application;

[0064] Figure 5 This is a schematic diagram of the internal structure of the electronic device of Example 4 of the present application.

[0065] Reference numerals:

[0066] 100, data acquisition and preprocessing module; 200, selection module; 300, sound source direction estimation module; 400, sound source number estimation module; 500, matching module. DETAILED DESCRIPTION

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0068] Example 1

[0069] like Figure 1 As shown, a method for simultaneously counting and locating multiple sound sources involved in Example 1 of the present application includes:

[0070] S100: Acquire time domain data collected by a microphone array, and pre-process the time domain data into frequency domain by dividing the frames into frames.

[0071] Assuming that there are P broadband sound sources in space incident from different direction angles θ, the time domain model of the signal received by each microphone at time t is as follows:

[0072] ;

[0073] Among them, x m (t) is the time domain data of the signal received by each microphone at time t, m=1, 2, 3, ... M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, and n is the additive noise. It should be noted that the parameters with subscripts in the above signal time domain model have the same definition as the above parameters without subscripts, for example: τ is the signal delay, τ m is the signal delay of the mth microphone, where m has been defined.

[0074] Then we need to pre-process the time domain data into frequency domain by dividing it into frames. First, we divide the time domain data into L frames according to the time length. The processing data of each frame is x. l , where l=1,2,…L.

[0075] Use Fourier transform to convert each frame of processed data into the frequency domain. The above formula can be written in the frequency domain as:

[0076] ;

[0077] Among them, X ml (f k ) is x m(t) Converted frequency domain data, k = 1, 2, ... K, K is the number of frequency points in the frequency band. It should be noted that when transforming the time domain model to the frequency domain, according to the conventional writing method, the parameters in the time domain model are lowercase letters, and those in the frequency domain are uppercase letters (such as xX, nN), both representing the same variable. To avoid redundancy, there is no need to repeat the definition. For example: in the time domain model, n is additive noise, while in the frequency domain, N is additive noise.

[0078] S200 , selecting a processing frequency band of interest, and selecting a calculation frequency from the processing frequency band according to signal energy.

[0079] Select the processing frequency band of interest. The signal energy of each frequency point in the processing frequency band can be expressed as:

[0080] ;

[0081] Among them, E k is the signal energy at the kth frequency point in the frequency band, X ml (f k ) is x m (t) Frequency domain data after transformation.

[0082] In order to select the appropriate calculation frequency, the energy threshold is introduced , select the frequency band that satisfies E k > All frequencies are used as subsequent positioning calculation frequencies, wherein the energy threshold is manually set according to the overall energy distribution of each frequency point calculated, and is not a fixed value.

[0083] S300: Preliminarily estimate the direction of the sound source using the phase information of the selected calculation frequency.

[0084] It should be noted that, in order to keep the formula simple, all parameters in the following text are defined in each time frame data, so there is no subscript l. Based on this, the preliminary estimation of the sound source direction includes the following steps:

[0085] Calculate the cross-power spectrum phase between microphone pairs in the microphone array :

[0086] ;

[0087] in, is the cross power spectrum, and * is the complex conjugate.

[0088] Calculate the phase rotation factor :

[0089] ;

[0090] in, is the relative delay difference between the microphone pair {1, 2} and the microphone pair {m, m+1}, φ∈[0, 2).

[0091] Constructing spectral functions :

[0092] .

[0093] Search by spectral peak and maximize the spectral function Get the sound source direction estimate :

[0094] ;

[0095] Among them, argmax means finding the independent variable that maximizes the function. It is a modulo operation.

[0096] The method used in this embodiment simplifies the sound source direction estimation into a nonlinear optimization problem, which significantly improves the computational efficiency and also has high estimation accuracy in complex environments.

[0097] S400, summarizing the sound source direction estimation results of each time frame and estimating the number of sound sources based on a peak iterative search algorithm, specifically comprising the following steps:

[0098] S401, the sound source direction estimation result of each time frame Summarize and calculate the statistics as a histogram Assuming there are three broadband sound sources with incident directions of 60°, 120°, and 160° respectively, the method of this embodiment is used to perform a preliminary estimation of the sound source direction and the histogram obtained by summarizing the results is like Figure 2 As shown, the reason why two peaks are displayed at 120° is that the algorithm divides the sound source point into equal parts when calculating it. Therefore, around 120°, due to the sound source estimation performed by the algorithm, there may be sound source directions such as 119°, 120° and 121°.

[0099] S402, estimating the number of sound sources based on a peak iterative search algorithm, specifically comprising the following steps:

[0100] S4021, Initialization:

[0101] Setting the sound source counter =1, indicating the first sound source currently identified;

[0102] Initialize the first peak position :

[0103] ;

[0104] Among them, argmax means finding the independent variable that maximizes the function;

[0105] Setting dynamic thresholds , For user-set thresholds, It is an empirical value, and its size is set according to the specific distribution results of the histogram. Indicates taking the maximum value, For the location The peak of the histogram.

[0106] S4022, iterative search for the next peak:

[0107] First search for the next peak in the histogram ;

[0108] If the peak value found ,in, is the dynamic threshold set in step S4021, and The purpose is to avoid being too close to the identified peak, so this minimum angle interval condition is set. If the above two conditions are met at the same time, the peak is considered is a new sound source, among which, In position The peak of the histogram, is the identified peak, is the minimum angular interval between sound sources, B is the total number of columns in the histogram, that is, the total number of columns in the histogram, It is a union operation used to merge elements of multiple sets, where the subscript j=1 represents the starting index of the union operation and the superscript c s The ending index of the union operation (usually a positive integer).

[0109] If a new sound source is found in the above steps, and renew: , ;

[0110] If no new sound source is found in the above steps or the preset maximum number of sound sources is reached, the search is stopped and the number of sound sources is obtained. , it can be seen that Figure 3 In order to estimate the probability density function of the histogram result, three obvious peaks were extracted, and the corresponding angles were 60°, 120°, and 160°, respectively. The number estimation and positioning of three sound sources were realized, verifying the effectiveness of the method of the present invention.

[0111] The sound source number estimation method used in this embodiment is different from the traditional method. It directly extracts the number of sound sources from multiple estimation results through the peak search method, and combines the two judgment conditions of peak value and minimum angular interval between sound sources. This not only reduces the computational complexity, but also effectively improves the estimation accuracy of the number of sound sources.

[0112] S500: Determine the direction corresponding to each sound source to achieve counting and positioning of multiple sound sources.

[0113] Specifically, the sound source directions in the preliminary sound source direction estimation in step S300 have already been located, but due to the large number of possible results, the specific number of sound sources and their corresponding directions are unknown. In step S400, the specific number of sound sources is estimated, and then the direction estimate corresponding to each sound source is found, thus completing the simultaneous counting and localization of multiple sound sources.

[0114] For example, suppose there are three sound sources with incident angles of 30°, 60°, and 90°, respectively. The initial estimate in step 3 may yield many angles, including the directions of the three sound sources. However, the actual number of sound sources and their corresponding directions are unknown. Therefore, step S400 is required to estimate the number of sound sources and then find the corresponding angle values ​​for each, thus achieving simultaneous counting and localization.

[0115] In this embodiment, there is no need for complex operations such as eigenvalue decomposition, and high counting and positioning accuracy is guaranteed. By summarizing multiple results of preliminary sound source estimation and then using peak iterative search to extract the number of sound sources, the sound source data information is fully utilized. Without the need for large-scale data training and without relying on various types of prior knowledge, the method has obvious advantages in terms of computational cost and estimation accuracy, and has been verified by simulation and experimental data.

[0116] Example 2

[0117] like Figure 4 As shown, a device for simultaneously counting and locating multiple sound sources according to embodiment 2 of the present application includes:

[0118] The data acquisition and preprocessing module 100 is used to acquire the time domain data collected by the microphone array and preprocess the time domain data into frequency domain by framing;

[0119] A selection module 200 for selecting a processing frequency band of interest and selecting a calculation frequency from the processing frequency band according to signal energy;

[0120] A sound source direction estimation module 300 is used to preliminarily estimate the sound source direction using the phase information of the selected calculation frequency;

[0121] A sound source number estimation module 400 is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on a peak iterative search algorithm;

[0122] The matching module 500 is used to determine the direction corresponding to each sound source to achieve the counting and positioning of multiple sound sources.

[0123] It should be noted that, for other specific implementations of the device for simultaneously counting and locating multiple sound sources in this embodiment, reference can be made to the specific implementations of the method for simultaneously counting and locating multiple sound sources described above, which will not be described again here to avoid redundancy.

[0124] Example 3

[0125] A computer-readable storage medium according to embodiment 3 of the present application, wherein the computer-readable storage medium stores program code for execution by a device, the program code including steps for executing the method in any one of the implementations in embodiment 1 of the present application;

[0126] Among them, the computer-readable storage medium can be a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM); the computer-readable storage medium can store program code, and when the program stored in the computer-readable storage medium is executed by the processor, the processor is used to execute the steps of the method in any one of the implementation methods in Example 1 of the present application.

[0127] Example 4

[0128] like Figure 5 As shown, an electronic device involved in Example 4 of the present application includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the method in any one of the implementations in Example 1 of the present application;

[0129] Among them, the processor can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the method in any one of the implementation methods in Example 1 of the present application.

[0130] The processor may also be an integrated circuit electronic device with signal processing capabilities. In the implementation process, each step of the method in any one of the implementations in Example 1 of the present application may be completed by hardware integrated logic circuits in the processor or software instructions.

[0131] The above-mentioned processor can also be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in combination with its hardware, completes the functions required to be executed by the units included in the data processing device of the embodiment of the present application, or executes the method in any one of the implementation modes in Example 1 of the present application.

[0132] The above are only preferred specific implementations of this application; however, the scope of protection of this application is not limited thereto. Any person skilled in the art who, within the technical scope disclosed in this application, makes equivalent substitutions or modifications based on the technical solutions and improved concepts of this application shall be covered by the scope of protection of this application.

Claims

1. A method for simultaneously counting and locating multiple sound sources, characterized in that: include: Acquire time domain data collected by the microphone array, and pre-process the time domain data into frequency domain by framing; selecting a processing frequency band of interest, and selecting a calculation frequency from the processing frequency band based on signal energy; The sound source direction is preliminarily estimated using the phase information of the selected calculation frequency; The sound source direction estimation results of each time frame are summarized and the number of sound sources is estimated based on the peak iterative search algorithm; Determine the direction corresponding to each sound source to achieve the counting and positioning of multiple sound sources.

2. The method for simultaneously counting and locating multiple sound sources according to claim 1, wherein: The signal time domain model of the time domain data is: ; Among them, x m (t) is the time domain data of the signal received by each microphone at time t, m=1, 2, 3,…M, M is the number of microphones, S is the signal emitted by the sound source, τ is the signal delay, n is the additive noise, P is the number of sound sources, and θ is the incident angle of the sound source.

3. The method for simultaneously counting and locating multiple sound sources according to claim 2, characterized in that: The time domain data is frame-preprocessed into the frequency domain, including: The time domain data is divided into L frames according to the time length, and the processing data of each frame is x l , where l=1,2,…L; Use Fourier transform to convert each frame of processed data into the frequency domain: ; Among them, X ml (f k ) is x m (t) The converted frequency domain data, N is the additive noise, k = 1, 2, …, K, K is the number of frequency points in the frequency band.

4. The method for simultaneously counting and locating multiple sound sources according to claim 3, wherein: The calculation frequency is selected from the processing frequency band based on the signal energy, including: Calculate the signal energy of each frequency point in the processing frequency band, where the signal energy E of each frequency point in the frequency band is k The calculation formula is: ; All frequencies whose signal energy is greater than a preset energy threshold are selected as calculation frequencies.

5. The method for simultaneously counting and locating multiple sound sources according to claim 4, characterized in that: The phase information of the selected calculation frequency is used to preliminarily estimate the direction of the sound source, including: Calculate the cross-power spectrum phase between microphone pairs in the microphone array : ; in, is the cross power spectrum, * is the complex conjugate; Calculate the phase rotation factor : ; in, is the relative delay difference between the microphone pair {1,2} and the microphone pair {m,m+1}, φ∈[0,2); Constructing spectral functions : ; Search by spectral peak and maximize the spectral function Get the sound source direction estimate : ; Among them, argmax means finding the independent variable that maximizes the function. It is a modulo operation.

6. The method for simultaneously counting and locating multiple sound sources according to claim 5, characterized in that: The sound source direction estimation results of each time frame are summarized and the number of sound sources is estimated based on the peak iterative search algorithm, including: The sound source direction estimation results of each time frame Summarize and make statistics as histogram; Setting the sound source counter =1, indicating the first sound source currently identified; Initialize the first peak position : ; Among them, argmax means finding the independent variable that maximizes the function; Setting dynamic thresholds , For user-set thresholds, Indicates taking the maximum value, For the location The peak of the histogram; Search for the next peak in the histogram ; like ,and , then the peak is a new sound source, among which, In position The peak of the histogram, is the identified peak, is the minimum angular interval between sound sources, B is the total number of columns in the histogram, It is a union operation, where the subscript j=1 represents the starting index of the union operation and the superscript c s Indicates the ending index of the union operation; If a new sound source is found, update , ; If no new sound source is found or the preset maximum number of sound sources is reached, the search stops and the number of sound sources is obtained. .

7. The method for simultaneously counting and locating multiple sound sources according to any one of claims 1 to 6, characterized in that: Determining the direction corresponding to each sound source includes: finding the direction estimation value corresponding to each sound source based on the estimated number of sound sources, so as to achieve counting and positioning of multiple sound sources.

8. A device for simultaneously counting and locating multiple sound sources, characterized in that: include: A data acquisition and preprocessing module is used to acquire time domain data collected by the microphone array and preprocess the time domain data into frequency domain frames; a selection module for selecting a processing frequency band of interest and selecting a calculation frequency from the processing frequency band according to signal energy; A sound source direction estimation module is used to preliminarily estimate the sound source direction using the phase information of the selected calculation frequency; A sound source number estimation module is used to summarize the sound source direction estimation results of each time frame and estimate the number of sound sources based on the peak iterative search algorithm; The matching module is used to determine the direction corresponding to each sound source to achieve the counting and positioning of multiple sound sources.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program codes for execution by a device, wherein the program codes include steps for executing the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sound source positioning method and system

    CN105204001A

  • Sound source positioning method and device, equipment and storage medium

    CN117724045A

  • Real-time thunder and lightning positioning method based on FPGA signal acquisition and preprocessing

    CN120142771A

  • Indoor multi-sound-source positioning method based on DOA estimation and DOA association

    CN120214697A

  • Microphone array position estimation device, microphone array position estimation method, and program

    US20200275224A1