A sound source positioning method, device, equipment and medium

By obtaining the direct sound direction and the stability judgment of the candidate sound source direction of the historical frame speech signal, the target direction of the current frame speech signal is determined, which solves the false positioning problem of traditional sound source localization algorithm in reflection environment and achieves higher positioning accuracy and stability.

CN122632191APending Publication Date: 2026-08-25HUNAN GOKE MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610775236.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Traditional sound source localization algorithms suffer from false localization in reflective environments because the reflected peak value is greater than the direct sound peak value.

Method used

By acquiring the direct sound direction and candidate sound source direction determined by several historical frame speech signals, and combining the stability judgment results of the candidate sound source direction, the target direction of the current frame speech signal is determined. The direct sound direction information of the historical frame speech signals is introduced to determine the candidate sound source direction, and the final target direction is determined by combining the stability judgment results of several candidate sound source directions corresponding to several historical frame speech signals.

Benefits of technology

It effectively solves the problem of false localization caused by the reflection peak being greater than the direct sound peak in a reflective environment, thus improving the accuracy and stability of sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122632191A_ABST
    Figure CN122632191A_ABST
Patent Text Reader

Abstract

The application discloses a sound source positioning method and device, equipment and medium, and relates to the technical field of sound source positioning. The method comprises the following steps: acquiring direct sound directions determined by a plurality of historical frame voice signals and a plurality of candidate sound source directions corresponding to the plurality of historical frame voice signals; acquiring a current frame voice signal, and determining the candidate sound source direction of the current frame voice signal according to the direct sound direction; and determining the target direction of the current frame voice signal according to the candidate sound source direction and the stability judgment result of the plurality of candidate sound source directions. It can be seen that the application effectively solves the problem that the traditional sound source positioning algorithm misjudges the reflection direction as the sound source direction due to the reflection peak being greater than the direct sound peak in the reflection environment, thereby causing false positioning, and improves the accuracy of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, and in particular to a sound source localization method, apparatus, equipment and medium. Background Technology

[0002] Sound source localization technology is widely used in intelligent conferencing systems, voice interaction devices, and other fields. Traditional sound source localization algorithms (such as SRP-PHAT, MUSIC, and beamforming algorithms) receive signals through microphone arrays and calculate the energy spectrum of the signals in various directions to find the direction with the highest energy as the direct sound estimation direction of the sound source. These methods have good localization performance in free-field or weak reflection environments. However, in indoor applications, sound wave propagation is often affected by reflective surfaces such as walls, ceilings, and furniture, generating a large amount of reflected sound. When the reflected sound energy is strong, a peak will appear on the energy spectrum corresponding to the reflection direction. If this reflection peak is greater than the direct sound peak, the localization algorithm based on the maximum energy principle will misjudge the reflection direction as the sound source direction, leading to false localization. Therefore, the above problems urgently need to be solved. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a sound source localization method, apparatus, device, and medium, which effectively solves the problem that traditional sound source localization algorithms misjudge the reflection direction as the sound source direction in reflective environments because the reflection peak value is greater than the direct sound peak value, thus leading to false localization, thereby improving the accuracy of sound source localization. The specific solution is as follows: Firstly, this application discloses a sound source localization method, including: Obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals. Acquire the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location; The target location of the current frame speech signal is determined based on the location of the candidate sound sources and the stability judgment results of the locations of the candidate sound sources.

[0004] Optionally, the process of determining the location of a direct sound from several historical frames of speech signals includes: The direction of the direct sound is determined based on the stability judgment results of the directions of the candidate sound sources.

[0005] Optionally, the process of determining the location of candidate sound sources corresponding to any frame of speech signal includes: Obtain the energy value of any frame of speech signal at each sound source location; The location of the sound source corresponding to the maximum energy value is determined as the corresponding candidate sound source location.

[0006] Optionally, the direction of the direct sound is determined based on the stability judgment results of the directions of the candidate sound sources, including: The effective sound source among the several candidate sound source locations is determined based on the noise threshold. Remove the sound sources with the maximum and minimum energy values ​​from the effective sound sources to obtain the target effective sound sources; If the number of the target effective sound sources is greater than a preset value, and the azimuth difference between any two of the target effective sound sources meets a preset condition, then the azimuth of the direct sound is determined based on the azimuth of the target effective sound sources. If the number of target effective sound sources is less than or equal to a preset value, or if the azimuth difference between any two target effective sound sources does not meet the preset condition, then the direct sound azimuth or initial azimuth of the sound source in the previous frame will be used as the direct sound azimuth.

[0007] Optionally, determining the direct sound direction based on the location of the target effective sound source includes: The average value of the azimuth of the target effective sound source is taken as the azimuth of the direct sound.

[0008] Optionally, determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: If the current frame speech signal has an energy peak within a preset range of the direct sound location, then the direct sound location is taken as the candidate sound source location of the current frame speech signal. If the current frame speech signal does not have an energy peak within a preset range of the direct sound location, then the sound source location corresponding to the maximum energy value of the current frame speech signal is taken as the candidate sound source location of the current frame speech signal.

[0009] Optionally, determining the target location of the current frame speech signal based on the location of the candidate sound sources and the stability judgment result of the locations of the plurality of candidate sound sources includes: The candidate sound source locations are recombined with the several candidate sound source locations to obtain several new candidate sound source locations, and the sound source localization method described above is executed to determine the new direct sound location. The new direct sound location is used as the target location of the current frame speech signal.

[0010] Optionally, the step of recombining the candidate sound source locations with the plurality of candidate sound source locations to obtain a new plurality of candidate sound source locations includes: If the candidate sound source location is the direct sound location, and the number of the candidate sound source locations exceeds the preset number, then delete the candidate sound source location corresponding to the first frame of speech signal among the candidate sound source locations, and add the candidate sound source location to the end of the remaining candidate sound source locations after deletion to obtain the new candidate sound source locations. If the candidate sound source location is the direct sound location, and the number of the plurality of candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the plurality of candidate sound source locations to obtain the new plurality of candidate sound source locations. If the candidate sound source location is not the direct sound location, and the number of the candidate sound source locations exceeds the preset number, then delete the candidate sound source location corresponding to the first frame of speech signal among the candidate sound source locations, and add the candidate sound source location to the end of the remaining candidate sound source locations after deletion to obtain the new candidate sound source locations. If the candidate sound source location is not the direct sound location, and the number of the candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the candidate sound source locations to obtain the new candidate sound source locations.

[0011] Optionally, after determining the target location of the current frame speech signal, the method further includes: Enable the suppress reflection command; Accordingly, determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: If the suppression of reflections command is enabled, the candidate sound source location of the current frame speech signal is determined based on the direct sound location.

[0012] Optionally, the enable / suppress reflection command includes:

[0013] The location of the sound source whose energy peak in the time domain of the current frame speech signal arrives later than the location of the direct sound is determined as the location of the reflected sound, and the energy peak corresponding to the location of the reflected sound is removed.

[0014] Secondly, this application discloses a sound source localization device, comprising: The historical location determination module is used to obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals. The current candidate location determination module is used to acquire the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location; The target orientation determination module is used to determine the target orientation of the current frame speech signal based on the orientation of the candidate sound sources and the stability judgment results of the orientations of the candidate sound sources.

[0015] Thirdly, this application discloses an electronic device, including: Memory, used to store computer programs; A processor is used to execute the computer program to implement the aforementioned sound source localization method.

[0016] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned sound source localization method.

[0017] As can be seen, this application proposes a sound source localization method, including: acquiring the direct sound direction determined by several historical frame speech signals and several candidate sound source directions corresponding to the several historical frame speech signals; acquiring the current frame speech signal and determining the candidate sound source directions of the current frame speech signal based on the direct sound direction; and determining the target direction of the current frame speech signal based on the stability judgment results of the candidate sound source directions and the several candidate sound source directions. In summary, this application determines the candidate sound source directions of the current frame speech signal based on the direct sound direction determined by several historical frame speech signals, and then determines the target direction of the current frame speech signal based on the stability judgment results of the candidate sound source directions and the several candidate sound source directions corresponding to the several historical frame speech signals. In this way, this application no longer relies on the maximum energy value of a single frame for sound source localization, but instead introduces the direct sound direction information of historical frame speech signals to determine the candidate sound source directions of the current frame speech signal, and simultaneously combines the stability judgment results of the several candidate sound source directions corresponding to several historical frame speech signals to determine the final target direction. Therefore, this application can effectively solve the problem that traditional sound source localization algorithms misjudge the direction of reflection as the direction of sound source in a reflective environment because the reflection peak is greater than the direct sound peak, thus leading to false localization. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a flowchart of a sound source localization method disclosed in this application; Figure 2 This is a flowchart of a specific sound source localization method disclosed in this application; Figure 3 This is a flowchart of a specific sound source localization method disclosed in this application; Figure 4 This is a schematic diagram of the structure of a sound source localization device disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] In indoor applications, sound wave propagation is often affected by reflective surfaces such as walls, ceilings, and furniture, generating a large amount of reflected sound. When the reflected sound energy is strong, a peak corresponding to the reflection direction will appear on the energy spectrum. If this reflection peak is greater than the direct sound peak, the localization algorithm based on the maximum energy principle will misjudge the reflection direction as the sound source direction, leading to false localization. These problems urgently need to be addressed.

[0022] To address this issue, this application proposes a sound source localization scheme that effectively solves the problem of traditional sound source localization algorithms misinterpreting the reflection direction as the sound source direction in reflective environments because the reflection peak value is greater than the direct sound peak value, thus leading to false localization.

[0023] This application discloses a sound source localization method. See also Figure 1 As shown, the method includes: Step S11: Obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals.

[0024] In this embodiment, the method for determining the location of a direct sound from several historical frame speech signals includes: determining the location of the direct sound based on the stability judgment result of the locations of the several candidate sound sources.

[0025] The process of determining the candidate sound source location for any frame of speech signal includes: acquiring the energy value of the speech signal in each frame at each sound source location; and determining the sound source location corresponding to the maximum energy value as the corresponding candidate sound source location. That is, first, a microphone array is used to collect the speech signal; the speech signal collected by the m-th microphone is... Where n is the discrete-time index, the signal preprocessing module divides the speech signal into frames, applies windows, and performs FFT (Fast Fourier Transform) to obtain the frequency domain signal. Here, k represents the frequency point position, l represents the frame index position, and M represents the number of microphones. The energy value of the speech signal in each frame is calculated using beamforming via the energy spectrum calculation module at each sound source location. Then, the sound source location corresponding to the maximum energy value is identified and recorded as a candidate sound source location. This candidate sound source location is stored in a buffer of length N, thus obtaining several candidate sound source locations corresponding to several historical frames of speech signals. In this way, storing the candidate sound source locations of multiple consecutive frames in a buffer of length N is equivalent to constructing a time-series data pool, providing a foundation for subsequently determining the target location using the stability judgment results of several candidate sound source locations corresponding to several historical frames of speech signals.

[0026] The determination of the direct sound direction based on the stability judgment results of several candidate sound source directions includes: determining the effective sound source among the sound sources corresponding to the directions of several candidate sound sources based on a noise threshold; removing the sound sources with the maximum and minimum energy values ​​corresponding to the effective sound sources to obtain the target effective sound source; if the number of target effective sound sources is greater than a preset value, and the azimuth difference between any two sound sources among the target effective sound sources meets a preset condition, then the direct sound direction is determined based on the direction of the target effective sound source; if the number of target effective sound sources is less than or equal to a preset value, or the azimuth difference between any two sound sources among the target effective sound sources does not meet the preset condition, then the direct sound direction or initial direction of the sound source in the previous frame is taken as the direct sound direction. It should be noted that VAD (Voice Activity Detection) information is introduced here. VAD is a detection technology used to distinguish between speech signals and noise signals. It calculates the cepstral energy of speech data every 20ms and compares this cepstral energy with a preset threshold (i.e., the noise threshold mentioned above) to determine whether the current data is a valid sound source. A VAD flag of 0 indicates that the current 20ms data is noise, and a VAD flag of 1 indicates that the current 20ms data is a valid sound source. Based on this, the sound sources with the maximum and minimum energy values ​​are removed from the valid sound sources to obtain the target valid sound sources. It is then determined whether the number of the target valid sound sources is greater than a preset value, and whether the azimuth difference between any two of the target valid sound sources meets a preset condition. If both of the above conditions are met, the azimuth of the direct sound is determined based on the azimuth of the target valid sound sources. For example, the preset value here is M (M≤N). If the number of target effective sound sources is greater than M, and the azimuth difference between any two target effective sound sources is less than or equal to 10 degrees, then the direct sound azimuth is determined according to the azimuth of the target effective sound sources; otherwise, the direct sound azimuth or initial azimuth of the sound source in the previous frame is used as the direct sound azimuth.

[0027] It should be noted that the reason for first determining whether the number of effective sound sources is greater than M before making subsequent judgments in the above process is that all of the N candidate sound source angles in the buffer may come from noise frames (VAD=0). Only when the number of effective speech frames corresponding to VAD=1 reaches the preset value M or more can it be said that there are enough effective speech angles in the buffer. Only then will the subsequent stability judgment based on these angles be meaningful, thereby reducing the false localization under noise dominance and ensuring the reliability of the direct sound direction judgment.

[0028] It should be noted that when the number of effective target sound sources is greater than M, another condition is that the azimuth difference between any two effective target sound sources meets a preset condition. In this case, the azimuth of the direct sound is determined based on the azimuth of the effective target sound sources. This is because direct sound occurs continuously in time and has a stable direction. If the azimuth difference between any two effective target sound sources is small (e.g., no more than 10 degrees), it indicates that the azimuths of these sound sources remain highly consistent over multiple consecutive moments, conforming to the physical law that the direction of direct sound is continuously stable and does not frequently change. Conversely, if the azimuth difference is too large, it indicates that the currently detected sound source azimuth may contain reflected sound or noise interference, lacking stability and unsuitable as a basis for determining the azimuth of the direct sound. Therefore, by setting a preset condition for the azimuth difference, the azimuths of sound sources that meet the temporal stability characteristics of direct sound can be effectively screened, thereby ensuring that the determined azimuth of the direct sound is accurate and reliable.

[0029] It should also be noted that the direct sound location of the sound source in the previous frame refers to the effective sound source location obtained by combining the locations of N candidate sound sources in the buffer with the VAD stability determination during the processing of the previous frame. In this way, the localization result is prevented from jumping due to noise or interference, ensuring the continuity and reliability of the output.

[0030] Furthermore, determining the direct sound direction based on the location of the target effective sound source includes: using the average value of the locations of the target effective sound sources as the direct sound direction. This helps to offset the random fluctuations of a single target effective sound source, making the final determined direct sound direction more accurate.

[0031] Step S12: Obtain the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location.

[0032] In this embodiment, determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: if the current frame speech signal has an energy peak within a preset range of the direct sound location, then the direct sound location is used as the candidate sound source location of the current frame speech signal; if the current frame speech signal does not have an energy peak within the preset range of the direct sound location, then the sound source location corresponding to the maximum energy value of the current frame speech signal is used as the candidate sound source location of the current frame speech signal. Specifically, the sound source localization algorithm traverses each direction within the range of 0° to 360° and calculates the energy value of each direction. For example, if the traversal step size is set to 5°, then 360° is divided into 72 directions, each direction corresponding to an index (1 to 72). When the direct sound location is output based on the stability judgment in the historical buffer, the system records the index of the direction corresponding to the direct sound location. Assuming the index corresponding to the direct sound location is 19, the preset range of the direct sound location can be set to the range corresponding to index 19. For a newly received current frame speech signal, first detect whether there is an energy peak at index 19. If there is, the direct sound direction (the direction corresponding to index 19) is taken as the candidate sound source direction of the current frame speech signal; if there is no direct sound direction, the direction corresponding to the maximum energy value in the current frame speech signal is taken as the candidate sound source direction of the current frame speech signal.

[0033] The reason for the above design is as follows: According to the laws of acoustic propagation, direct sound arrives at the microphone array before reflected sound. Therefore, the peak of the direct sound wave appears at the beginning of the time axis, while the peak of the reflected sound wave appears at the end of the time axis. As long as the peak of the direct sound wave still exists, even if its energy has been surpassed by the subsequent peak of the reflected sound wave, it indicates that the current stage is still dominated by the direct sound, and the true direction of the sound source has not changed. Therefore, the system first checks whether there is a peak in the direction of the direct sound in the current frame of the speech signal: if it exists, the direction of the direct sound is continued to be used as the candidate sound source direction of the current frame of the speech signal, thereby avoiding being misled by the higher-energy reflected peak; if it does not exist, it means that the direct sound may have disappeared or has not yet appeared, and the direction corresponding to the maximum energy value in the current frame of the speech signal is taken as the candidate sound source direction of the current frame of the speech signal.

[0034] Step S13: Determine the target location of the current frame speech signal based on the location of the candidate sound source and the stability judgment result of the location of the candidate sound source.

[0035] In this embodiment, determining the target location of the current frame speech signal based on the candidate sound source locations and the stability judgment results of the candidate sound source locations includes: recombining the candidate sound source locations with the candidate sound source locations to obtain new candidate sound source locations, and executing the aforementioned sound source localization method to determine the new direct sound location; and using the new direct sound location as the target location of the current frame speech signal. Specifically, the candidate sound source locations corresponding to the current frame speech signal are recombined with several historical candidate sound source locations in the buffer to obtain several new candidate sound source locations containing current frame information; the aforementioned multi-frame stability judgment process combined with VAD is executed on the new candidate sound source locations to determine the updated direct sound location; and finally, the new direct sound location is used as the target location of the current frame speech signal. In this way, this application no longer relies on the maximum energy value of a single frame for sound source localization, but introduces the direct sound location information of historical frame speech signals to determine the candidate sound source locations of the current frame speech signal, and at the same time combines the stability judgment results of several candidate sound source locations corresponding to several historical frame speech signals to determine the final target location. The following is a detailed description of how to reorganize: It should be noted that in this embodiment, the buffer length is N. After processing each frame of speech signal, a buffer update operation is performed, which involves shifting the candidate sound source positions of frames 2 to N in the buffer forward by one position and adding a new candidate sound source position to the end of the buffer. This embodiment provides several implementation methods for whether and how to perform the update operation: I. Real-time updates In the first embodiment, if the candidate sound source location is the direct sound location and the number of the plurality of candidate sound source locations exceeds a preset number, then the candidate sound source location corresponding to the first frame of speech signal among the plurality of candidate sound source locations is deleted, and the candidate sound source location is added to the end of the remaining candidate sound source locations after deletion to obtain the new plurality of candidate sound source locations.

[0036] In the second embodiment, if the candidate sound source location is the direct sound location and the number of the plurality of candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the plurality of candidate sound source locations to obtain the new plurality of candidate sound source locations.

[0037] In the third implementation, if the candidate sound source location is not the direct sound location and the number of the plurality of candidate sound source locations exceeds a preset number, then the candidate sound source location corresponding to the first frame of speech signal among the plurality of candidate sound source locations is deleted, and the candidate sound source location is added to the end of the remaining candidate sound source locations after deletion to obtain the new plurality of candidate sound source locations.

[0038] In the fourth embodiment, if the candidate sound source location is not the direct sound location and the number of the plurality of candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the plurality of candidate sound source locations to obtain the new plurality of candidate sound source locations.

[0039] It should be noted that the preset number is generally taken as the buffer length N. The advantage of the above real-time update scheme is that the buffer always maintains the candidate sound source location information of the most recent N frames, which can track the changes in sound source location in real time, avoid outputting historical information due to long-term lack of updates, and improve the continuous tracking capability of localization.

[0040] II. Simplified Update

[0041] In this scheme, if the candidate sound source location is a direct sound location, then the aforementioned candidate sound source locations become the new candidate sound source locations. That is, when the candidate sound source location is a direct sound location, no update operation is performed on the buffer; the candidate sound source locations remain unchanged, and the direct sound location is directly used as the target location of the current frame's speech signal. In this case, there is no need to perform the addition or deletion operations described in the first and second embodiments. When the candidate sound source location is not a direct sound location (i.e., the current frame's speech signal does not have an energy peak in the direct sound location, and the candidate sound source location is the direction corresponding to the maximum energy value of the current frame), then a buffer update operation is performed, specifically using the third or fourth embodiment described above. The advantage of the simplified update scheme is that it reduces the frequency of buffer update operations and is easier to implement.

[0042] It is understandable that the aforementioned real-time update scheme and simplified update scheme each have their advantages, and those skilled in the art can choose one of them according to actual application needs. Specifically: if the application scenario has high requirements for real-time tracking capability of sound source location, such as intelligent conference systems where speakers move frequently or robot hearing scenarios, then the real-time update scheme is preferred to keep the buffer always containing the effective angle information of the most recent N frames, ensuring continuous tracking capability of positioning; if the application scenario has high requirements for computational overhead and ease of implementation, such as resource-constrained embedded devices or voice interaction systems with extremely high real-time requirements, then the simplified update scheme is preferred to reduce the frequency of buffer operations and reduce computational burden.

[0043] As can be seen, this application proposes a sound source localization method, including: acquiring the direct sound direction determined by several historical frame speech signals and several candidate sound source directions corresponding to the several historical frame speech signals; acquiring the current frame speech signal and determining the candidate sound source directions of the current frame speech signal based on the direct sound direction; and determining the target direction of the current frame speech signal based on the stability judgment results of the candidate sound source directions and the several candidate sound source directions. In summary, this application determines the candidate sound source directions of the current frame speech signal based on the direct sound direction determined by several historical frame speech signals, and then determines the target direction of the current frame speech signal based on the stability judgment results of the candidate sound source directions and the several candidate sound source directions corresponding to the several historical frame speech signals. In this way, this application no longer relies on the maximum energy value of a single frame for sound source localization, but instead introduces the direct sound direction information of historical frame speech signals to determine the candidate sound source directions of the current frame speech signal, and simultaneously combines the stability judgment results of the several candidate sound source directions corresponding to several historical frame speech signals to determine the final target direction. Therefore, this application can effectively solve the problem that traditional sound source localization algorithms misjudge the direction of reflection as the direction of sound source in a reflective environment because the reflection peak is greater than the direct sound peak, thus leading to false localization.

[0044] This application discloses a specific sound source localization method. Compared to the previous embodiment, this embodiment further explains and optimizes the technical solution, specifically by adding a reflection suppression processing step. In actual sound field environments, there are often sound reflections from walls and obstacles, which can interfere with the actual direct sound signal and easily cause errors in sound source location determination. Therefore, this embodiment adds a reflection suppression processing step to further improve the accuracy of sound source localization. See also... Figure 2 As shown, it specifically includes: Step S21: Obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals.

[0045] Step S22: Obtain the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location.

[0046] Step S23: Determine the target location of the current frame speech signal based on the location of the candidate sound source and the stability judgment result of the location of the candidate sound source.

[0047] For more detailed information on steps S21, S22, and S23, please refer to the aforementioned embodiments, which will not be repeated here.

[0048] Step S24: Enable the suppress reflection command.

[0049] In this embodiment, after determining the target location of the current frame speech signal based on the location of the candidate sound sources and the stability judgment results of the locations of the candidate sound sources, a reflection suppression command is enabled. Specifically, enabling the reflection suppression command involves: determining the location of the sound source whose energy peak in the time domain of the current frame speech signal arrives later than the location of the direct sound as the location of the reflected sound, and removing the energy peak corresponding to the location of the reflected sound. After enabling the reflection suppression command, determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: if the reflection suppression command is enabled, determining the candidate sound source location of the current frame speech signal based on the direct sound location; determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: if the current frame speech signal has an energy peak within a preset range of the direct sound location, then the direct sound location is taken as the candidate sound source location of the current frame speech signal; if the current frame speech signal does not have an energy peak within the preset range of the direct sound location, then the sound source location corresponding to the maximum energy value of the current frame speech signal is taken as the candidate sound source location of the current frame speech signal.

[0050] See Figure 3 As shown, enabling the reflection suppression command enables the reflection suppression module. After the reflection suppression module is enabled, the reflection suppression process begins: After calculating the energy spectrum of each direction of the subsequent frame speech signal, it enters the judgment branch of "whether the reflection suppression module is enabled". If the judgment is "yes", it enters the peak judgment process to detect whether there is a direct sound peak in the current frame speech signal within the preset range of the historical direct sound direction. If there is a direct sound peak, the direct sound direction remains unchanged and is directly used as the candidate sound source direction. If there is no direct sound peak, the angle update operation is performed to redetermine the candidate sound source direction of the current frame speech signal (that is, the sound source direction corresponding to the maximum energy value of the current frame speech signal is used as the candidate sound source direction of the current frame speech signal). Then, the candidate sound source direction is updated and cached in the buffer to reorganize the candidate sound source direction with the several candidate sound source directions to obtain several new candidate sound source directions, and the subsequent stability judgment process continues to be executed. It should be noted that, according to the laws of acoustic propagation, direct sound arrives at the microphone array before reflected sound. The peak of the direct sound wave appears at the beginning of the time axis, while the peak of the reflected sound wave appears at the end of the time axis. Furthermore, the energy of the reflected sound is often higher than that of the direct sound, easily forming a more prominent peak in the energy spectrum. If these reflected sound peaks are not removed, subsequent buffer angle updates will be misled, thus affecting the accuracy of stability judgment and sound source localization. Therefore, this application determines the sound source location in the current frame's speech signal where the energy peak in the time domain arrives later than the location of the direct sound as the location of the reflected sound, and removes its corresponding energy peak. This fundamentally avoids interference from reflected sound and ensures the reliability of subsequent update and judgment processes.

[0051] It should be noted that the reflection suppression module, once enabled after obtaining a reliable direct sound location, can continuously eliminate reflected sound in subsequent frames, ensuring that the angle information updated to the buffer does not contain reflection interference. However, before the reflection suppression module is enabled, the system cannot suppress reflected sound. To further improve the accuracy and stability of sound source localization, this scheme can: in the preprocessing stage or system initialization stage, buffer a segment of speech signal, perform VAD discrimination and stability analysis, output an initial stable direct sound location, and enable the reflection suppression module accordingly. Afterward, the system enters the actual operation stage, using the enabled reflection suppression module to continuously eliminate reflection peaks in subsequent frames and update the buffer with the suppressed angle information, thereby ensuring that the angle within the buffer in the entire closed-loop system always remains the direction of the direct sound without reflection interference. This design solves the problem of reflections not being suppressed before the suppression module is activated, and also ensures the continuous accuracy of subsequent localization.

[0052] See further Figure 3 As shown, the following specific example illustrates this solution: For the first frame of the speech signal, signal preprocessing and energy spectrum calculation in each direction are performed. The location of the sound source corresponding to the maximum energy value of the speech signal is determined as the candidate sound source location of the speech signal and then stored in the first position of the buffer. Similarly, for the second frame, the third frame, and the Nth frame of the speech signal, the above steps are performed in sequence to obtain N frames of angle information, where the length of the buffer is N. Furthermore, by combining the angle information of N frames, a direct sound azimuth (also called direct sound angle) is obtained. If the direct sound angle is proven stable after the aforementioned stability judgment, the reflection suppression module is enabled. At this time, for the N+1th frame speech signal, the N+2th frame speech signal, and so on, for the real-time processing flow, they all belong to the current frame speech signal. For these speech signals, after real-time acquisition, signal preprocessing, and energy spectrum calculation in each direction, it is determined whether there is a peak in the energy spectrum of these speech signals for the direct sound azimuth. If a peak exists, the direct sound azimuth is maintained and used as the candidate sound source azimuth of the current frame speech signal. Then, this azimuth is directly used in the subsequent target azimuth determination process. If no peak exists, the sound source azimuth corresponding to the maximum value of the energy spectrum in each direction of the current frame speech signal is recalculated, and this sound source azimuth is used as the candidate sound source azimuth of the current frame speech signal. Then, the N frame angle information in the buffer is updated, and the subsequent stability judgment process is entered. It should be noted that the reflection suppression module remains enabled after the first activation.

[0053] In summary, the main problem and objective of this solution is that sound source localization technology faces the challenge of reflection environments: the reflected sound energy may be greater than the direct sound energy, causing traditional algorithms to incorrectly identify the reflection direction as the true direction of the sound source, resulting in false localization. Specifically, traditional methods struggle to identify the direct sound direction when multiple energy peaks exist, leading to insufficient localization accuracy and robustness in reflection scenarios. To address these shortcomings, this solution aims to provide a method to improve the accuracy of sound source localization and solve the problem of algorithms locating false directions in reflection scenarios. This includes: 1) not relying on the assumption that reflected sound energy is necessarily less than direct sound, thus achieving accurate localization even when reflected sound energy is greater than direct sound; 2) utilizing the temporal characteristic that direct sound arrives earlier than reflected sound, identifying and recording the direct sound direction during the arrival phase, and verifying the continuity of the direct sound direction during the arrival phase of reflected sound; 3) combining VAD and cached historical angle information to improve the reliability of localization; 4) implementing a low-complexity, low-latency reflection suppression algorithm requiring only simple peak detection, statistical analysis, and logical judgment, suitable for embedded device implementation.

[0054] In summary, the technical steps of this solution are as follows: Based on the physical law that direct sound arrives earlier than reflected sound, and the stability and persistence characteristics of the direction of direct sound in time sequence, a historical angle buffer mechanism is used to save the temporal information, and stable direct sound directions are extracted from the buffer by combining VAD information. The specific process includes: dividing the speech signal into frames, windowing and performing FFT to obtain the frequency domain signal, calculating the energy spectrum of each direction in the current frame and finding the direction corresponding to the maximum value as the candidate sound source direction, and then storing the estimated sound source direction in a buffer of length N; next, combining VAD and buffered historical angle information to determine and output a stable angle direction, which is recorded as the direct sound azimuth; finally, by detecting whether there is a peak in the current frame speech signal at the direct sound azimuth, if there is, the current azimuth is determined to be the direct sound azimuth, and peaks in other azimuths are discarded. The key to this scheme lies in the continuous verification of the direct sound direction and the decision mechanism for suppressing reflections: After confirming the direction of the direct sound, peak detection is performed on the energy spectrum of the current frame to determine whether the current frame contains a peak of the direct sound (allowing an error of 1-2 positions). If it persists, the direction of the direct sound is output and other reflection peaks are suppressed; if it does not exist, the angle is updated to the buffer. This scheme does not rely on traditional energy criteria, but fully utilizes the characteristic that the direct sound peak still exists even when the reflected sound energy is greater, which can effectively identify and suppress reflection peaks, improving the accuracy of localization in reflective environments. In terms of algorithm complexity, this scheme does not require complex matrix operations and iterative optimization, resulting in low computational complexity and low processing latency, meeting the real-time application requirements of intelligent conferencing, security monitoring, and voice interaction. At the same time, the historical angle buffering mechanism fully utilizes the direct sound direction information in earlier frames, avoiding it from being covered by subsequent reflected sound, thereby improving the stability of sound source localization.

[0055] Accordingly, this application also discloses a sound source localization device, see [link to relevant documentation]. Figure 4 As shown, the device includes: The historical location determination module 11 is used to obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals. The current candidate location determination module 12 is used to acquire the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location; The target orientation determination module 13 is used to determine the target orientation of the current frame speech signal based on the orientation of the candidate sound sources and the stability judgment results of the orientations of the candidate sound sources.

[0056] For more detailed information on the working process of each of the above modules, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0057] Furthermore, embodiments of this application also provide an electronic device. Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0058] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the sound source localization method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0059] In this embodiment, the power supply 26 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 24 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0060] Furthermore, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon may include computer programs 221, and the storage method may be temporary storage or permanent storage. In addition to including computer programs capable of performing the sound source localization method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 221 may further include computer programs capable of performing other specific tasks.

[0061] Furthermore, embodiments of this application also disclose a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned sound source localization method.

[0062] For the specific steps of this method, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.

[0063] The various embodiments in this application are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts between the various embodiments, refer to each other. As for the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section.

[0064] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0065] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0066] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0067] The above provides a detailed description of a sound source localization method, apparatus, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for locating a sound source, characterized in that, include: Obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals. Acquire the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location; The target location of the current frame speech signal is determined based on the location of the candidate sound sources and the stability judgment results of the locations of the candidate sound sources.

2. The sound source localization method according to claim 1, characterized in that, The process of determining the location of a direct sound from several historical frames of speech signals includes: The direction of the direct sound is determined based on the stability judgment results of the directions of the candidate sound sources.

3. The sound source localization method according to claim 2, characterized in that, The process of determining the location of candidate sound sources corresponding to any frame of speech signal includes: Obtain the energy value of any frame of speech signal at each sound source location; The location of the sound source corresponding to the maximum energy value is determined as the corresponding candidate sound source location.

4. The sound source localization method according to claim 2, characterized in that, The direction of the direct sound is determined based on the stability assessment results of the several candidate sound source directions, including: The effective sound source among the several candidate sound source locations is determined based on the noise threshold. Remove the sound sources with the maximum and minimum energy values ​​from the effective sound sources to obtain the target effective sound sources; If the number of the target effective sound sources is greater than a preset value, and the azimuth difference between any two of the target effective sound sources meets a preset condition, then the azimuth of the direct sound is determined based on the azimuth of the target effective sound sources. If the number of target effective sound sources is less than or equal to a preset value, or if the azimuth difference between any two target effective sound sources does not meet the preset condition, then the direct sound azimuth or initial azimuth of the sound source in the previous frame will be used as the direct sound azimuth.

5. The sound source localization method according to claim 4, characterized in that, Determining the direct sound direction based on the direction of the target effective sound source includes: The average value of the azimuth of the target effective sound source is taken as the azimuth of the direct sound.

6. The sound source localization method according to claim 1, characterized in that, Determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: If the current frame speech signal has an energy peak within a preset range of the direct sound location, then the direct sound location is taken as the candidate sound source location of the current frame speech signal. If the current frame speech signal does not have an energy peak within a preset range of the direct sound location, then the sound source location corresponding to the maximum energy value of the current frame speech signal is taken as the candidate sound source location of the current frame speech signal.

7. The sound source localization method according to claim 6, characterized in that, The step of determining the target location of the current frame speech signal based on the location of the candidate sound sources and the stability judgment results of the locations of the candidate sound sources includes: The candidate sound source locations are recombined with the plurality of candidate sound source locations to obtain a plurality of new candidate sound source locations, and the sound source localization method as described in any one of claims 4 and 5 is performed to determine the new direct sound location; The new direct sound location is used as the target location of the current frame speech signal.

8. The sound source localization method according to claim 7, characterized in that, The step of recombining the candidate sound source locations with the plurality of candidate sound source locations to obtain a plurality of new candidate sound source locations includes: If the candidate sound source location is the direct sound location, and the number of the candidate sound source locations exceeds the preset number, then delete the candidate sound source location corresponding to the first frame of speech signal among the candidate sound source locations, and add the candidate sound source location to the end of the remaining candidate sound source locations after deletion to obtain the new candidate sound source locations. If the candidate sound source location is the direct sound location, and the number of the plurality of candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the plurality of candidate sound source locations to obtain the new plurality of candidate sound source locations. If the candidate sound source location is not the direct sound location, and the number of the candidate sound source locations exceeds the preset number, then delete the candidate sound source location corresponding to the first frame of speech signal among the candidate sound source locations, and add the candidate sound source location to the end of the remaining candidate sound source locations after deletion to obtain the new candidate sound source locations. If the candidate sound source location is not the direct sound location, and the number of the candidate sound source locations does not exceed the preset number, then the candidate sound source location is added to the end of the candidate sound source locations to obtain the new candidate sound source locations.

9. The sound source localization method according to claim 1, characterized in that, After determining the target location of the current frame speech signal, the method further includes: Enable the suppress reflection command; Accordingly, determining the candidate sound source location of the current frame speech signal based on the direct sound location includes: If the suppression of reflections command is enabled, the candidate sound source location of the current frame speech signal is determined based on the direct sound location.

10. The sound source localization method according to claim 9, characterized in that, The enable / suppress reflection command includes: The location of the sound source whose energy peak in the time domain of the current frame speech signal arrives later than the location of the direct sound is determined as the location of the reflected sound, and the energy peak corresponding to the location of the reflected sound is removed.

11. A sound source localization device, characterized in that, include: The historical location determination module is used to obtain the direct sound location determined by several historical frame speech signals and the locations of several candidate sound sources corresponding to the several historical frame speech signals. The current candidate location determination module is used to acquire the current frame speech signal and determine the candidate sound source location of the current frame speech signal based on the direct sound location; The target orientation determination module is used to determine the target orientation of the current frame speech signal based on the orientation of the candidate sound sources and the stability judgment results of the orientations of the candidate sound sources.

12. An electronic device, characterized in that, include: Memory is used to store computer programs; A processor for executing the computer program to implement the sound source localization method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the sound source localization method as described in any one of claims 1 to 10.