A method, device, equipment and storage medium for sound source localization
By calculating the delay value of the microphone array signal and the historical sound source position, the search range of sound source positioning is narrowed, and the existing algorithm has solved the problem of high computing volume in scenarios with high real-time requirements, and efficient sound source positioning is achieved.
Patent Information
- Application Number
- CN202210248170.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-03-14
AI Technical Summary
The existing SRP-PHAT-based sound source positioning algorithm has too much computing capacity in scenarios with high real-time requirements and cannot meet the needs of fast positioning.
By calculating the signal delay value between the microphone arrays, the initial position of the real-time microphone signal is determined, and local spatial search is performed with this center, and weighted calculation is performed in combination with the historical sound source position to narrow the search range and reduce the positioning calculation amount.
It greatly narrows the spatial search range of sound source positioning, reduces the amount of computing, and is suitable for scenarios with high real-time requirements.
Smart Images

Figure CN114646920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio signal processing, and specifically, to a sound source localization method, device, equipment, and storage medium. Background Art
[0002] In fields such as video conferencing, smart home, robotics, and fault detection and localization, the localization of sound sources is of great significance. Existing technologies usually adopt sound source localization technologies based on microphone arrays to perform sound source localization. For example, a sound source localization algorithm based on Steered Response Power and Phase Transform (SRP-PHAT) is used. However, the computational complexity of this method during sound source localization is relatively large, and it has high requirements for the computational power of the device. Therefore, on the basis of this method, in the existing technology, based on the SRP-PHAT sound source localization algorithm, the cross-power spectral components that do not contribute to the phase accumulation sum are removed, and the full search for all frequency bands in the original method is changed to a coarse-to-fine search by frequency band. Although the computational complexity of the improved SRP-PHAT sound source localization algorithm during sound source localization has decreased compared with the original solution, there is still a problem of large computational complexity and it is not suitable for scenarios with high real-time requirements. Summary of the Invention
[0003] Based on this, the present invention provides a sound source localization method, device, equipment, and storage medium, which can perform preliminary localization by using the time delay between microphone arrays, greatly reducing the initial range of spatial search, reducing the localization computational complexity, and being suitable for scenarios with high real-time requirements.
[0004] To achieve the above object, an embodiment of the present invention provides a sound source localization method, including:
[0005] Calculating the signal time delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signals; wherein, the real-time microphone signals include the real-time signals of the relative microphones and the real-time signals of the reference microphone, and the number of the relative microphones is at least 3;
[0006] Calculating the initial position of the real-time microphone signals according to the spatial positions of all the relative microphones acquired, the spatial position of the reference microphone, the preset speed of sound, and all the signal time delay values;
[0007] Searching the local space to obtain the position with the maximum local response power, and taking the position with the maximum local response power as the sound source position of the real-time microphone signals; wherein, the center of the local space is the initial position, and the radius of the local space is the preset real-time search radius.
[0008] As an improvement of the above solution, the real-time search radius is obtained by the following method:
[0009] Search the global space to obtain the maximum global response power and the global position corresponding to the maximum global response power, calculate the distance error between the global position and the initial position, and set a power threshold according to the maximum global response power;
[0010] Set a search radius according to the distance error;
[0011] Use the search radius as the real-time search radius until the maximum local response power obtained in real-time is less than the power threshold, and then return to the step of searching the global space.
[0012] As an improvement to the above solution, after taking the position of the maximum local response power as the sound source position of the real-time microphone signal, the following steps are further included:
[0013] Perform weighted calculation based on the sound source positions of several recent historical microphone signals and the sound source position of the real-time microphone signal to obtain a final position, and update the sound source position of the real-time microphone signal according to the final position.
[0014] As an improvement to the above solution, the final position is calculated in the following way:
[0015]
[0016] IF(||r s (t)-r s (t - 1)|| 2 >μ);
[0017]
[0018]
[0019] Among them, r 终 (t) represents the final position, ω(t - i) represents the weight coefficient of the sound source position at the previous i moments, r s (t) represents the sound source position of the real-time microphone signal, r s (t - 1) represents the sound source position at the previous moment, and μ represents a preset distance threshold.
[0020] As an improvement to the above solution, before calculating the signal time delay value between each relative microphone and the reference microphone according to the obtained real-time microphone signal, the following steps are further included:
[0021] Normalize the real-time microphone signal;
[0022] Perform effective frame detection on the real-time microphone signal to delete invalid frames.
[0023] As an improvement of the above solution, calculating the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal specifically includes:
[0024] Based on the GCC-PHAT algorithm, calculating the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal.
[0025] As an improvement of the above solution, calculating the initial position of the real-time microphone signal according to the spatial positions of each acquired relative microphone, the preset speed of sound, and all the signal delay values specifically includes:
[0026] Obtaining the reference position of the reference microphone, the relative positions of all the relative microphones, and setting the sound source position of the real-time microphone signal as an unknown position;
[0027] Listing a system of equations for solving the spatial delay value between each relative microphone and the reference microphone according to each reference position, the reference position, the unknown position, and the acquired speed of sound;
[0028] Establishing a sound source position calculation system of equations with the signal delay value and the system of equations for solving the spatial delay value:
[0029] Based on the least mean square error method, solving the sound source position calculation system of equations to obtain the initial position of the real-time microphone signal.
[0030] To achieve the above object, an embodiment of the present invention further provides a sound source localization device, including:
[0031] A signal delay value calculation module, configured to calculate the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal; wherein, the real-time microphone signal includes the real-time signal of the relative microphone and the real-time signal of the reference microphone, and the number of the relative microphones is at least 3;
[0032] An initial position calculation module, configured to calculate the initial position of the real-time microphone signal according to the spatial positions of all the acquired relative microphones, the spatial position of the reference microphone, the preset speed of sound, and all the signal delay values;
[0033] A sound source position calculation module, configured to search a local space to obtain the position with the maximum local response power, and use the position with the maximum local response power as the sound source position of the real-time microphone signal; wherein, the center of the local space is the initial position, and the radius of the local space is a preset real-time search radius.
[0034] To achieve the above object, an embodiment of the present invention further provides a sound source localization device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the sound source localization method described in any of the above embodiments is implemented.
[0035] To achieve the above object, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. When the computer program runs, the device where the computer-readable storage medium is located is controlled to execute the sound source localization method described in any of the above embodiments.
[0036] Compared with the prior art, the sound source localization method, device, equipment, and storage medium disclosed in the embodiments of the present invention calculate the signal delay value between each relative microphone and the reference microphone by obtaining the real-time microphone signal, and then calculate the initial position of the real-time microphone according to the signal delay value, the spatial positions of all the relative microphones and the reference microphone obtained, and the preset sound speed. Taking the initial position as the center of the local space, a local space is searched with a preset real-time search radius, so as to obtain the position with the maximum local response power and use it as the sound source position of the real-time microphone signal. It can be seen that the embodiments of the present invention calculate the signal delay value between each relative microphone and the reference microphone through the real-time microphone signal, combine the spatial positions of each microphone and the sound speed to determine the initial position of the real-time microphone signal, and search with the initial position as the search center to obtain the sound source position of the real-time microphone, greatly reducing the initial range of spatial search and reducing the positioning calculation amount, which is applicable to scenarios with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] To more clearly illustrate the technical solutions of the present invention, the drawings required for the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 is a flowchart of a sound source localization method provided by an embodiment of the present invention;
[0039] Figure 2 is a schematic diagram of a microphone array model provided by an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of the positional relationship between a microphone pair and a sound source provided by an embodiment of the present invention;
[0041] Figure 4It is a schematic structural diagram of a sound source localization device provided by an embodiment of the present invention;
[0042] Figure 5 It is a schematic structural diagram of a sound source localization device provided by an embodiment of the present invention. Specific embodiments
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] See Figure 1 , which is a schematic flowchart of a sound source localization method provided by an embodiment of the present invention. The sound source localization method can be executed by a computing terminal, and the computing terminal is a device such as a computer or a tablet computer, which is not limited herein.
[0045] Specifically, the sound source localization method includes the following steps:
[0046] S1. Calculate the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signals; wherein, the real-time microphone signals include the real-time signals of the relative microphones and the real-time signals of the reference microphone, and the number of the relative microphones is at least 3;
[0047] S2. Calculate the initial position of the real-time microphone signals according to the spatial positions of all the relative microphones, the spatial position of the reference microphone, the preset sound speed, and all the signal delay values;
[0048] S3. Search the local space to obtain the position with the maximum local response power, and use the position with the maximum local response power as the sound source position of the real-time microphone signals; wherein, the center of the local space is the initial position, and the radius of the local space is the preset real-time search radius.
[0049] Specifically, by way of example, the microphone array is composed of 4 microphones. One of the microphones is selected as the reference microphone, and the other 3 are used as relative microphones. The microphone array performs real-time acquisition of the audio signal to obtain the real-time microphone signals, and calculates the signal delay value between each relative microphone and the reference microphone according to the real-time microphone signals, that is, calculates the time delay value of the signals received between the microphones by using the signals; see Figure 2The shown microphone array model, where M represents the positions of microphones (including 4 microphones in the figure). Assume that the distance from the sound source to the center of the array element is D, the azimuth angle is α, and the elevation angle is β. Since the time delay values of the signals received between microphones are related to the distances between each microphone and the sound source, as well as the propagation speed, therefore, using the spatial position relationship between each microphone and the preset sound speed (the sound speed in the medium where the audio signal propagates), combined with the signal time delay values, the initial position of the real-time microphone signal is calculated. Taking the initial position as the center of the local space and the preset real-time search radius as the search radius of the local space, search in the space near the initial position to find the local maximum response power point and take this point as the sound source position of the real-time microphone signal. Thus, it can be seen that the embodiment of the present invention calculates the signal time delay values between each relative microphone and the reference microphone through the real-time microphone signal, combines the spatial positions and sound speed of each microphone to determine the initial position of the real-time microphone signal, and searches with the initial position as the search center to obtain the sound source position of the real-time microphone, greatly reducing the initial range of the spatial search and reducing the positioning calculation amount, which is applicable to scenarios with high real-time requirements.
[0050] In one implementation, the real-time search radius is obtained in the following manner:
[0051] Search the global space to obtain the maximum global response power and the global position corresponding to the maximum global response power, calculate the distance error between the global position and the initial position, and set a power threshold according to the maximum global response power;
[0052] Set the search radius according to the distance error;
[0053] Take the search radius as the real-time search radius until the maximum local response power calculated in real time is less than the power threshold, and return to the step of searching the global space.
[0054] Specifically, in this embodiment, the SRG-PHAT algorithm is improved. When the SRP-PHAT algorithm is calculated for the first time, search the global space U = {(d, α, β)|d represents the distance to the center of the array element, α represents the azimuth angle, β represents the elevation angle} to obtain the maximum global response power P max, record the position (global position) where the maximum global response power is located (represented in the form of a vector), calculate the distance error between the global position and the initial position, and set a power threshold according to the maximum global response power (the maximum global power is positively correlated with the power threshold). Optionally, multiply the distance error by a preset radius coefficient (constant) to obtain a search radius, which is used as the real-time search radius for the subsequently obtained real-time microphone signals. As time goes by, if the position of the sound source moves significantly and moves away from the original local space, in this case, the maximum local response power obtained by searching in the original local space will also decrease significantly. Therefore, if the maximum local response power is less than the power threshold during subsequent calculations, it is necessary to re-search the global space to re-determine the new local space.
[0055] In one implementation, after taking the position of the maximum local response power as the sound source position of the real-time microphone signal in step S13, it further includes:
[0056] Perform weighted calculation based on the sound source positions of several recent historical microphone signals and the sound source position of the real-time microphone signal to obtain a final position, and update the sound source position of the real-time microphone signal according to the final position.
[0057] Specifically, after searching for the position of the maximum local response power, consider the sound source positions at several recent historical moments, perform weighted processing on them, perform smooth positioning, calculate the final position, and update the sound source position of the real-time microphone signal according to the final position.
[0058] In one implementation, the final position is calculated by the following method:
[0059]
[0060] IF(||r s (t)-r s (t - 1)|| 2 >μ);
[0061]
[0062]
[0063] where, r 终 (t) represents the final position, ω(t - i) represents the weight coefficient of the sound source position at the previous i moments, r s (t) represents the sound source position of the real-time microphone signal, r s (t - 1) represents the sound source position at the previous moment, μ represents a preset distance threshold, and p is greater than or equal to 1.
[0064] Specifically, the final position is obtained by weighted calculation of the sound source position of the real-time microphone signal and the sound source positions of the most recent p - 1 times. When the square of the difference between the sound source position of the current real-time microphone signal and the sound source position of the previous time is greater than a preset distance threshold, it indicates that there is a large change in the position of the sound source at the current moment and the previous moment, that is, the sound source moves relatively significantly. At this time, the weight coefficient is calculated using the first coefficient calculation method. When the square of the difference between the sound source position of the current real-time microphone signal and the sound source position of the previous time is not greater than the preset distance threshold, the idea of averaging (1 / p) is used to set the coefficient. It should be noted that the specific value of μ is set according to the experience of engineering applications.
[0065] In one implementation, before calculating the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal, it further includes:
[0066] Normalize the real-time microphone signal;
[0067] Perform effective frame detection on the real-time microphone signal to delete invalid frames.
[0068] Specifically, the original audio signal is acquired, and the original audio signal is translated and scaled in amplitude to obtain a standard audio signal (normalized signal) with an average value of 0 and a maximum amplitude of 1. The normalized signal is preprocessed, windowed (optional sine window, Hanning window), and framed to obtain a series of audio frames (one audio frame represents a real-time microphone signal). Each frame is processed in turn, and effective frame detection is performed on this frame (using the VAD algorithm to detect as a speech frame or a noise frame). If it is a non-speech frame (noise frame), it is skipped. If it is an effective frame, it is retained for subsequent algorithm processing.
[0069] In one implementation, in step S1, calculating the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal specifically includes:
[0070] Based on the GCC-PHAT algorithm, calculate the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signal.
[0071] Specifically, the GCC-PHAT algorithm is:
[0072]
[0073]
[0074]
[0075] Wherein, R(τ) is the cross-correlation function, A(w) is the weighted value of the reciprocal of the cross-power spectrum of X and Y, nFFt is the Fourier transform length, Fs is the sampling rate, X(w) and Y(w) are the Fourier transforms of the microphone signals x(t) and y(t). In this embodiment, x(t) is the real-time microphone signal of one of the relative microphones, and y(t) is the real-time microphone signal of the reference microphone. The above specific GCC-PHAT algorithm is used to calculate the signal delay value between each relative microphone and the reference microphone.
[0076] Further, taking the valid frame as the object, perform GCC-PHAT (Generalized Cross-Correlation) calculation, calculate the tde (signal delay value) between each relative microphone and the reference microphone, and perform delay value screening. If the screening fails, jump back to the step of valid frame detection and reselect the valid frame.
[0077] Specifically, the method of delay value screening is:
[0078] tde <= d / c;
[0079] That is, tde should be less than the maximum delay between the corresponding relative microphone and the reference microphone. d is the distance between the corresponding microphone and the reference microphone, and c is the preset speed of sound.
[0080] In one implementation manner, the step of calculating the initial position of the real-time microphone signal according to the spatial positions of all the relative microphones, the spatial position of the reference microphone, the preset speed of sound, and all the signal delay values in step S2 specifically includes:
[0081] Obtain the reference position of the reference microphone and the relative positions of all the relative microphones, and set the sound source position of the real-time microphone signal as an unknown position;
[0082] List a system of equations for solving the spatial delay value between each relative microphone and the reference microphone according to the relative position, the reference position, the unknown position, and the obtained speed of sound;
[0083] Establish a system of equations for calculating the sound source position by combining the signal delay value and the system of equations for solving the spatial delay value;
[0084] Based on the least mean square method, solve the system of equations for calculating the sound source position to obtain the initial position of the real-time microphone signal.
[0085] Specifically, referring to Figure 3 the schematic diagram of the position relationship between the microphone pair and the sound source, list the system of equations for solving the spatial delay value:
[0086]
[0087] Among them, c represents the preset sound speed, and τ ij represents the spatial time delay value between the i-th relative microphone and the reference microphone. (x Mi , y Mi , z Mi ) represents the spatial position of the i-th relative microphone, and (x Mj , y Mj , z Mj ) represents the spatial position of the reference microphone, and (x S , y S , z S ) represents the sound source position (unknown position) of the real-time microphone signal;
[0088] It can be seen from the system of equations for solving the spatial time delay value that the sound source is located on a hyperboloid with the relative microphone M i and the reference microphone M j as the foci. When the microphone array has multiple array elements (more than three), each pair of microphones and its spatial time delay value can determine a hyperboloid. Due to the existence of time delay estimation errors, these hyperboloids cannot intersect at an absolute point. However, generally speaking, the overlapping area of the hyperboloids is concentrated around the sound source.
[0089] Figure 3 In j , M j represents the spatial position (reference position) of the reference microphone, Mi represents the spatial position of the i-th relative microphone, S represents the sound source position (unknown position) of the real-time microphone signal, Rs is the distance from the reference microphone M j to the sound source S, Ri is the distance from the relative microphone Mi to the sound source S, and di
[0090]
[0091]
[0092]
[0093]
[0094]
[0095] After simplification, it is obtained:
[0096] R i 2 -d ij 2 -2d ij R s -2r i T rs = 0
[0097] Since d ij is obtained by estimating the time delay between microphones, d ij = cτ ij , there must be an error compared with the actual value. Therefore, the above equation is not 0. Assume its error is:
[0098] ε = R i 2 - d ij 2 - 2d ij R s - 2r i T r s
[0099] Assume there are M microphones in the sound source localization system, numbered from 0 to M - 1. Taking the position of microphone M0 as the reference point, a coordinate system is established with it as the origin. It can be obtained that:
[0100] ε = δ - 2R s d - 2Sr s
[0101]
[0102] The purpose of the minimum mean square error method is to minimize the mean variance of the above equation, and at this time the localization result is the most accurate. After derivation, in order to achieve the minimum mean square error, the position of the sound source is:
[0103]
[0104]
[0105] Specifically, the specific formula of SRP - PHAT is as follows:
[0106]
[0107]
[0108] Compared with the prior art, the sound source localization method disclosed in the embodiments of the present invention calculates the signal delay value between each relative microphone and the reference microphone by acquiring real-time microphone signals, and then calculates the initial position of the real-time microphone according to the signal delay value, the spatial positions of all the relative microphones and the reference microphone obtained, and the preset sound speed. Taking the initial position as the center of the local space, a local space is searched with a preset real-time search radius, so as to obtain the position with the maximum local response power and take it as the sound source position of the real-time microphone signal. It can be seen that the embodiments of the present invention calculate the signal delay value between each relative microphone and the reference microphone through real-time microphone signals, combine the spatial positions of each microphone and the sound speed to determine the initial position of the real-time microphone signal, and search with the initial position as the search center to obtain the sound source position of the real-time microphone, greatly reducing the initial range of spatial search and reducing the positioning calculation amount, which is applicable to scenarios with high real-time requirements.
[0109] See Figure 4 , which is a schematic structural diagram of a sound source localization device provided by an embodiment of the present invention. The sound source localization device includes:
[0110] A signal delay value calculation module 11, configured to calculate the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signals; wherein, the real-time microphone signals include the real-time signals of the relative microphones and the real-time signals of the reference microphone, and the number of the relative microphones is at least 3;
[0111] An initial position calculation module 12, configured to calculate the initial position of the real-time microphone signal according to the spatial positions of all the relative microphones obtained, the spatial position of the reference microphone, the preset sound speed, and all the signal delay values;
[0112] A sound source position calculation module 13, configured to search the local space to obtain the position with the maximum local response power, and take the position with the maximum local response power as the sound source position of the real-time microphone signal; wherein, the center of the local space is the initial position, and the radius of the local space is the preset real-time search radius.
[0113] It should be noted that the specific working process of the sound source localization device may refer to the working process of the sound source localization method in the above embodiments, and will not be elaborated here.
[0114] Compared with the prior art, the sound source localization device disclosed in the embodiments of the present invention calculates the signal delay value between each relative microphone and the reference microphone by acquiring real-time microphone signals, and then calculates the initial position of the real-time microphone based on the signal delay value, the spatial positions of all the relative microphones and the reference microphone obtained, and a preset sound speed. Taking the initial position as the center of the local space, a local space is searched with a preset real-time search radius, so as to obtain the position with the maximum local response power and use it as the sound source position of the real-time microphone signal. It can be seen that the device in the embodiments of the present invention calculates the signal delay value between each relative microphone and the reference microphone through real-time microphone signals, combines the spatial positions of each microphone and the sound speed to determine the initial position of the real-time microphone signal, and searches with the initial position as the search center to obtain the sound source position of the real-time microphone, greatly reducing the initial range of spatial search and reducing the positioning calculation amount, and being applicable to scenarios with high real-time requirements.
[0115] See Figure 5 , which is a schematic structural diagram of a sound source localization device provided by an embodiment of the present invention. The sound source localization device includes a processor 21, a memory 22, and a computer program stored in the memory 22 and configured to be executed by the processor 21. When the processor 21 executes the computer program, it implements the steps in the embodiment of the above sound source localization method, such as Figure 1 the steps S1 to S3 described in
[0116] Exemplarily, the computer program may be divided into one or more modules. The one or more modules are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the sound source localization device. For example, the computer program may be divided into a signal delay value calculation module 11, an initial position calculation module 12, and a sound source position calculation module 13. The specific functions of each module are as follows:
[0117] The signal delay value calculation module 11 is configured to calculate the signal delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signals; wherein, the real-time microphone signals include the real-time signals of the relative microphones and the real-time signals of the reference microphone, and the number of the relative microphones is at least 3;
[0118] An initial position calculation module 12 is configured to calculate the initial position of the real-time microphone signal based on all the obtained spatial positions relative to the microphone, the spatial position of the reference microphone, a preset sound speed, and all the signal delay values.
[0119] A sound source position calculation module 13 is configured to search a local space to obtain the position with the maximum local response power, and use the position with the maximum local response power as the sound source position of the real-time microphone signal; wherein, the center of the local space is the initial position, and the radius of the local space is a preset real-time search radius.
[0120] The specific working processes of each module can refer to the working process of the sound source localization device described in the above embodiments, and will not be elaborated here.
[0121] The sound source localization device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The sound source localization device may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the sound source localization device, and does not constitute a limitation on the sound source localization device. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the sound source localization device may further include an input / output device, a network access device, a bus, etc.
[0122] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor 21 is the control center of the sound source localization device, and connects various parts of the entire sound source localization device through various interfaces and lines.
[0123] The memory 22 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 22, and invoking the data stored in the memory 22, the processor 21 realizes various functions of the sound source localization device. The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0124] Among them, if the modules integrated in the sound source localization device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0125] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A sound source localization method, characterized in that, Comprising: Calculating the signal time delay value of each relative microphone and the reference microphone according to the acquired real-time microphone signal; wherein, the real-time microphone signal includes the real-time signal of the relative microphone and the real-time signal of the reference microphone, and the number of the relative microphones is at least 3; Calculating the initial position of the real-time microphone signal according to the spatial positions of all the acquired relative microphones, the spatial position of the reference microphone, the preset speed of sound, and all the signal time delay values; Taking the initial position as the center of the local space and the preset real-time search radius as the radius of the local space, searching the local space to obtain the position with the maximum local response power, and taking the position with the maximum local response power as the sound source position of the real-time microphone signal.
2. The sound source localization method according to claim 1, wherein The real-time search radius is obtained by the following method: Searching the global space to obtain the maximum global response power, the global position corresponding to the maximum global response power, calculating the distance error between the global position and the initial position, and setting a power threshold according to the maximum global response power; Setting the search radius according to the distance error; Taking the search radius as the real-time search radius until the maximum local response power calculated in real time is less than the power threshold, and returning to the step of searching the global space.
3. The sound source localization method according to claim 1, characterized in that, After taking the position with the maximum local response power as the sound source position of the real-time microphone signal, it further includes: Performing weighted calculation according to the sound source positions of several recent historical microphone signals and the sound source position of the real-time microphone signal to obtain the final position, and updating the sound source position of the real-time microphone signal according to the final position.
4. The sound source localization method according to claim 3, characterized in that, The final position is calculated by the following method: ; ; ; ; Among them, represents the final position, represents the weight coefficient of the sound source position at the previous i-th moment, represents the sound source position of the real-time microphone signal, represents the sound source position at the previous moment, represents a preset distance threshold; is greater than or equal to 1.
5. The sound source localization method according to claim 1, wherein Before calculating the signal time delay value of each relative microphone and the reference microphone according to the acquired real-time microphone signal, it further includes: Performing normalization processing on the real-time microphone signal; Performing effective frame detection on the real-time microphone signal to delete invalid frames.
6. The sound source localization method according to claim 1, wherein Calculating the signal time delay value of each relative microphone and the reference microphone according to the acquired real-time microphone signal specifically includes: Based on the GCC-PHAT algorithm, calculating the signal time delay value of each relative microphone and the reference microphone according to the acquired real-time microphone signal.
7. The sound source localization method according to claim 1, characterized in that Calculating the initial position of the real-time microphone signal according to the spatial positions of all the acquired relative microphones, the spatial position of the reference microphone, the preset speed of sound, and all the signal time delay values specifically includes: Obtaining the reference position of the reference microphone and the relative positions of all the relative microphones, and setting the sound source position of the real-time microphone signal as an unknown position; Listing a system of equations for solving the spatial time delay value of each relative microphone and the reference microphone according to the relative position, the reference position, the unknown position, and the acquired speed of sound; Establishing a system of equations for calculating the sound source position by combining the signal time delay value and the system of equations for solving the spatial time delay value; Based on the least mean square method, solving the system of equations for calculating the sound source position to obtain the initial position of the real-time microphone signal.
8. A sound source localization device, characterized in that, Comprising: A signal time delay value calculation module, configured to calculate the signal time delay value between each relative microphone and the reference microphone according to the acquired real-time microphone signals; wherein, the real-time microphone signals include the real-time signals of the relative microphones and the real-time signals of the reference microphone, and the number of the relative microphones is at least 3; An initial position calculation module, configured to calculate the initial position of the real-time microphone signals according to the spatial positions of all the relative microphones, the spatial position of the reference microphone, a preset speed of sound, and all the signal time delay values; A sound source position calculation module, configured to search a local space with the initial position as the center of the local space and a preset real-time search radius as the radius of the local space, so as to obtain the position with the maximum local response power, and use the position with the maximum local response power as the sound source position of the real-time microphone signals.
9. A sound source localization device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the sound source localization method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the sound source localization method according to any one of claims 1 to 7.