A high-precision sound source positioning method and system based on a two-dimensional microphone array

By using a grouping and iterative optimization method for two-dimensional microphone arrays, the problem of low sound source localization accuracy in near-field environments is solved, achieving high-precision sound source localization, which is applicable to fields such as intelligent security, robot navigation, and video conferencing.

CN120522639BActive Publication Date: 2025-11-07CHINA SOUTHERN POWER GRID INTERNET SERVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511028959.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-07
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

Existing sound source localization technologies have low localization accuracy in near-field environments and rely on external ranging equipment, resulting in large errors.

Method used

A two-dimensional microphone array is used, which is divided into two groups of microphone arrays. The azimuth angle of the sound source is initially estimated using a far-field model. Combined with geometric relationships and near-field model, iterative optimization is performed. Finally, the accuracy is improved by a full-array near-field localization optimization algorithm.

Benefits of technology

No additional ranging equipment is required, which significantly improves the accuracy and practicality of sound source localization in near-field environments. It is simple and efficient, and suitable for fields such as intelligent security, robot navigation and video conferencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120522639B_ABST
    Figure CN120522639B_ABST
Patent Text Reader

Abstract

The application discloses a high-precision sound source positioning method and system based on a two-dimensional microphone array, aiming at solving the problem of low sound source positioning accuracy in a near-field environment. The microphone array is grouped, and the sound source azimuth is preliminarily estimated by using a far-field model; the preliminary sound source position is calculated through geometric relationship; the sound source position is iteratively optimized in combination with a near-field model; and finally, the final sound source position is obtained through a full-array near-field positioning optimization algorithm. The high-precision sound source positioning system disclosed by the application does not need external ranging equipment, is suitable for a near-field scene, and improves the practicability and robustness of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, and in particular to a high-precision sound source positioning method and system based on a two-dimensional microphone array. BACKGROUND

[0002] Sound source positioning technology has a wide and important application in many fields. In the field of intelligent security, by accurately determining the position of the sound source, the source of abnormal sound can be found in time to ensure the safety of the place; in robot navigation, sound source positioning can help the robot to identify the position of the sound source in the surrounding environment, to realize more intelligent interaction and action; in the video conference system, the accurate positioning of the speaker's sound source position can optimize the camera tracking effect and improve the conference experience.

[0003] However, the existing sound source positioning technology has many problems in the near-field environment. The traditional sound source positioning method is mostly based on a far-field model to construct a steering vector, and the positioning accuracy of this model will decrease significantly in the actual near-field environment where the distance between the sound source and the microphone is close. Moreover, many existing technologies rely on external ranging devices (such as laser ranging) to determine the distance of the sound source, but this way often leads to large errors in sound source positioning in complex scenes. SUMMARY

[0004] The purpose of the present application is to provide a high-precision sound source positioning method and system based on a two-dimensional microphone array, which can improve the above problems.

[0005] In a first aspect, the present application provides a high-precision sound source positioning method based on a two-dimensional microphone array, which includes steps S1 to S4, wherein S1, S2, etc. are only step identifiers, and the execution order of the method does not necessarily follow the order from small to large, such as executing step S2 first and then executing step S1, which is not limited by the present application.

[0006] S1: dividing the two-dimensional microphone array into two groups of microphone arrays, in the far-field model, the first sound source azimuth is preliminarily estimated by the sound information collected by the first microphone array, and the second sound source azimuth is preliminarily estimated by the sound information collected by the second microphone array.

[0007] S2: calculating the intersection position of the first sound source azimuth and the second sound source azimuth as the preliminary sound source estimation position according to the geometric relationship.

[0008] S3: bringing the preliminary sound source estimation position into the near-field model to iteratively optimize the near-field sound source estimation position.

[0009] S4: re-optimizing the near-field sound source estimation position by a full-array near-field positioning optimization algorithm to obtain the final sound source estimation position.

[0010] This application proposes a high-precision sound source localization method based on a two-dimensional microphone array, comprising: grouping the microphone array and using a far-field model to initially estimate the azimuth angle of the sound source; calculating the preliminary sound source position through geometric relationships; iteratively optimizing the sound source position by combining a near-field model; and finally obtaining the final sound source position through a full-array near-field localization optimization algorithm. This application eliminates the need for additional ranging equipment, avoiding the inconsistencies in direction or the inability to reflect light that may arise from laser ranging, thus reducing costs. It is particularly suitable for near-field environments, improving the accuracy and practicality of sound source localization. Through iterative optimization, the localization process is simple, efficient, and highly accurate. Furthermore, this method has good scalability and can be applied to two-dimensional / three-dimensional array structures and complex actual sound fields, providing a more accurate and reliable sound source localization solution for multiple fields such as intelligent security, robot navigation, and video conferencing.

[0011] In an optional embodiment of this application, step S1 includes: dividing the microphones in the two-dimensional microphone array into two groups of microphone arrays, using beamforming estimation or generalized cross-correlation delay estimation, to initially estimate the first sound source azimuth based on the far-field steering vector of the sound information collected by the first microphone array, and to initially estimate the second sound source azimuth based on the far-field steering vector of the sound information collected by the second microphone array.

[0012] In an optional embodiment of this application, the first sound source azimuth angle is initially estimated based on the far-field steering vector of the sound information collected by the first microphone array using beamforming estimation, and the second sound source azimuth angle is initially estimated based on the far-field steering vector of the sound information collected by the second microphone array, including:

[0013] The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated using the following formula:

[0014] ;

[0015] in, The frequency at which the amplitude is maximum is obtained after performing Fourier decomposition on the sound signal generated by the sound source. Represents the speed of sound. Represents the first microphone array Each microphone location This represents the total number of microphones in the first microphone array. The unit vector representing the direction of the sound source, i.e. , Represents the angle of the sound source;

[0016] The beam output power of the sound information collected by the first microphone array is calculated using the following formula. :

[0017] ;

[0018] wherein, represents the frequency domain form of the beam collected by the first microphone array, i.e.

[0019] ,

[0020] represents the Fourier transform form of the sound information collected by the i-th microphone in the first microphone array;

[0021] The first sound source azimuth angle corresponding to the maximum beam output power is found according to the following formula: ; wherein, represents the first sound source azimuth angle;

[0022] The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated by the following formula:

[0023] ;

[0024] wherein, represents the position of the i-th microphone in the second microphone array, represents the total number of microphones in the second microphone array;

[0025] The beam output power of the sound information collected by the second microphone array is calculated by the following formula:

[0026] ;

[0027] wherein, represents the frequency domain form of the beam collected by the second microphone array, i.e.

[0028] ,

[0029] represents the Fourier transform form of the sound information collected by the i-th microphone in the second microphone array;

[0030] The second sound source azimuth angle corresponding to the maximum beam output power is found according to the following formula:

[0031] ; wherein, represents the second sound source azimuth angle.

[0032] ​​​​In an optional embodiment of the present application, S2 comprises the following steps S21-S23, wherein S21, S22, etc. are only step identifiers, and the execution order of the method does not necessarily follow the order from small to large numbers, for example, step S22 can be executed first and then step S21, which is not limited in the present application.

[0033] S21: Calculate the distance between the geometric center of the first microphone array and the preliminary sound source estimation position according to the following formula:

[0034] ;

[0035] wherein, represents the distance between the geometric center of the first microphone array and the preliminary sound source estimation position, represents the distance between the preliminary sound source estimation position and the plane where the two-dimensional microphone array is located, represents the distance between the geometric center of the first microphone array and the geometric center of the second microphone array, represents the first sound source azimuth angle, represents the second sound source azimuth angle;

[0036] S22: Calculate the distance between the geometric center of the second microphone array and the preliminary sound source estimation position according to the following formula:

[0037] ;

[0038] represents the distance between the geometric center of the second microphone array and the preliminary sound source estimation position;

[0039] S23: Calculate the intersection position of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimation position according to the following formula:

[0040] ;

[0041] wherein, represents the preliminary sound source estimation position, represents the geometric center position of the first microphone array, represents the geometric center position of the second microphone array;

[0042] represents the directional vector of the first sound source azimuth angle, i.e. ,

[0043] represents the directional vector of the second sound source azimuth angle, i.e. .

[0044] In an optional embodiment of this application, step S3 includes: calculating a near-field steering vector based on the previously calculated sound source estimation position; iteratively calculating the first sound source azimuth angle and the second sound source azimuth angle using the near-field steering vector; calculating the position of the intersection point of the line between the iterated first sound source azimuth angle and the second sound source azimuth angle based on geometric relationships, until the distance between the currently calculated line intersection point and the previously calculated line intersection point is less than a preset threshold; when the distance between the currently calculated line intersection point and the previously calculated line intersection point is less than the preset threshold, the currently calculated line intersection point is used as the near-field sound source estimation position; when initially calculating the near-field steering vector, calculating the near-field steering vector based on the preliminary sound source estimation position.

[0045] The purpose of step S3 is to improve the accuracy of sound source localization through iterative optimization. It first calculates the near-field steering vector based on the previously calculated estimated sound source location. Then, it iteratively calculates the sound source azimuth using this vector and determines the intersection point of the azimuths based on geometric relationships. This process iterates until the distance between the intersection points of two consecutive calculations is less than a preset threshold, thus determining the estimated near-field sound source location. This process effectively corrects the initial estimation error, improving the accuracy and robustness of sound source localization, and is particularly suitable for high-precision sound source localization requirements in near-field environments.

[0046] In an optional embodiment of this application, the step of calculating the near-field steering vector based on the previously calculated sound source estimated position includes:

[0047] According to the following formula, the previously calculated sound source estimation location is substituted into the near-field steering vector of the sound information collected by the first microphone array in the near-field model:

[0048] ;

[0049] in, The frequency at which the amplitude is maximum is obtained after performing Fourier decomposition on the sound signal generated by the sound source. Represents the speed of sound. Represents the first microphone array Each microphone location This represents the total number of microphones in the first microphone array;

[0050] According to the following formula, the estimated location of the sound source calculated in the previous calculation is substituted into the near-field steering vector of the sound information collected by the second microphone array in the near-field model:

[0051] ;

[0052] in, Represents the second microphone array a microphone position, representing the total number of microphones in the second microphone array.

[0053] In an optional embodiment of the present application, the first sound source azimuth angle and the second sound source azimuth angle are iteratively calculated using the near-field steering vector, including: calculating the first sound source azimuth angle according to the near-field steering vector of the sound information collected by the first microphone array, and calculating the second sound source azimuth angle according to the near-field steering vector of the sound information collected by the second microphone array, by a beamforming estimation method or a generalized cross-correlation delay estimation method.

[0054] In an optional embodiment of the present application, the step S4 includes: re-optimizing the near-field sound source estimated position according to the following cost function to obtain the final sound source estimated position:

[0055] ;

[0056] wherein, representing the frequency at the maximum amplitude obtained after Fourier decomposition of the sound signal generated by the sound source, representing the sound speed, representing the position of the i-th microphone in the two-dimensional microphone array, a microphone position, representing the total number of microphones in the two-dimensional microphone array, representing the near-field sound source estimated position, representing the final sound source estimated position.

[0057] It can be understood that the step S4 optimizes the near-field sound source estimated position by a full-array near-field positioning optimization algorithm, aiming to further improve the accuracy of sound source positioning. This algorithm comprehensively considers the sound information collected by all microphones in the two-dimensional microphone array, and finely adjusts the preliminary estimated near-field sound source position through complex mathematical models and calculations. After this step of optimization, the final sound source estimated position obtained will be closer to the real sound source position, effectively reducing errors and providing more reliable sound source positioning data support for intelligent security, robot navigation and other fields.

[0058] In a second aspect, the present application discloses a high-precision sound source positioning system based on a two-dimensional microphone array, comprising a two-dimensional microphone array and a processor, the two-dimensional microphone array is used to collect sound information in the environment, and the processor is used to execute the method according to any one of the first aspect.

[0059] It can be understood that the high-precision sound source positioning system based on the two-dimensional microphone array provided in the application does not need to rely on external devices such as laser ranging, avoids direction deviation and reflection problems, and reduces costs; it is particularly suitable for near-field environments, significantly improves the accuracy and practicality of sound source positioning through near-field model iterative optimization; combined with the full-array near-field positioning optimization algorithm, the accurate estimation of the sound source position is realized; the system also has good scalability and can be flexibly applied to two-dimensional or three-dimensional array structures and complex sound field environments, providing an efficient and reliable sound source positioning solution for intelligent security, robot navigation and other fields.

[0060] In an optional embodiment of the application, the arrangement form of the two-dimensional microphone array can be at least one of the following: a linear array; a circular array; a multi-spiral arm array. Beneficial effects

[0061] The application provides a high-precision sound source positioning method and system based on a two-dimensional microphone array, aiming to solve the problem of low sound source positioning accuracy in a near-field environment. The microphone array is grouped, the sound source azimuth is preliminarily estimated using a far-field model, the preliminary sound source position is calculated through geometric relationships, the sound source position is iteratively optimized in combination with a near-field model, and finally the final sound source position is obtained through a full-array near-field positioning optimization algorithm. The application does not need additional ranging devices, avoids the problem of inconsistent directions or inability to reflect caused by laser ranging, reduces costs, is particularly suitable for near-field environments, improves the accuracy and practicality of sound source positioning, and through iterative optimization, the positioning process is simple and efficient, and the positioning accuracy is high. The method has good scalability and can be applied to two-dimensional / three-dimensional array structures and actual complex sound fields, providing a more accurate and reliable sound source positioning solution for intelligent security, robot navigation, video conferencing and other fields.

[0062] The application has been verified by MATLAB simulation and actual measurement data, and in a typical experimental scenario (the distance between the sound source and the array center is 0.5-2 meters), the error is reduced by more than 60% compared with the traditional far-field model positioning method. In the environment without laser assistance, the sound source position can still be stably output, and the application has engineering feasibility and practical application potential.

[0063] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the following optional embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the application, and therefore should not be regarded as a limitation on the scope. Other related drawings can also be obtained by those skilled in the art without creating any creative labor.

[0065] Figure 1 is a schematic diagram provided by the present application for dividing a two-dimensional microphone array into two groups of microphone arrays;

[0066] Figure 2 is a schematic diagram provided by the present application for calculating the intersection position of the first sound source azimuth and the second sound source azimuth as a preliminary sound source estimation position according to the geometric relationship;

[0067] Figure 3 is a schematic diagram of the connection relationship of the high-precision sound source positioning system based on the two-dimensional microphone array provided by the present application. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0069] In a first aspect, the present application provides a high-precision sound source positioning method based on a two-dimensional microphone array, which includes steps S1 to S4. S1, S2, etc. are only step identifiers, and the execution order of the method does not necessarily follow the order from small to large, such as executing step S2 first and then executing step S1. The present application does not make any limitation.

[0070] S1: Divide the two-dimensional microphone array into two groups of microphone arrays. In the far-field model, the first sound source azimuth is preliminarily estimated by the sound information collected by the first microphone array, and the second sound source azimuth is preliminarily estimated by the sound information collected by the second microphone array.

[0071] The arrangement mode of the two-dimensional microphone array can include linear array, circular array, and multi-spiral arm array, etc. Different arrangement modes have their own advantages: linear array has simple structure and is easy to implement, and is suitable for scenes of preliminary estimation of sound source azimuth; circular array can provide more uniform azimuth coverage, which is beneficial to omnidirectional sound source positioning; multi-spiral arm array combines the advantages of the former two, and can realize high-precision positioning in more complex environments.

[0072] The division of the microphone is mainly to divide the microphone array into two groups, and the specific division manner is not limited, which can be divided into two groups of microphone arrays with different numbers, or can be divided into two groups of microphone arrays with different numbers. Each group of microphones can work independently to reduce mutual interference, and through comparison and fusion of the data of the two groups, the accuracy and robustness of sound source positioning can be improved. Especially in a near-field environment, this division method combined with near-field model iterative optimization can significantly improve the accuracy of sound source positioning.

[0073] As shown in Figure 1 , the microphone array arranged in a multi-helix arm array is divided into a first microphone array and a second microphone array with different numbers of microphones.

[0074] S2: According to the geometric relationship, the intersection position of the first sound source azimuth and the second sound source azimuth is calculated as the preliminary sound source estimation position.

[0075] S3: The preliminary sound source estimation position is brought into the near-field model to iteratively optimize the near-field sound source estimation position.

[0076] S4: The near-field sound source estimation position is further optimized by a full-array near-field positioning optimization algorithm to obtain the final sound source estimation position.

[0077] It can be understood that the present application proposes a high-precision sound source positioning method based on a two-dimensional microphone array, which includes: grouping the microphone array, using a far-field model to preliminarily estimate the sound source azimuth; calculating the preliminary sound source position through geometric relationship; iteratively optimizing the sound source position combined with the near-field model; and finally obtaining the final sound source position through a full-array near-field positioning optimization algorithm. The present application does not need additional ranging equipment, avoids the problem of inconsistent direction or inability to reflect caused by laser ranging, and reduces the cost; it is especially suitable for near-field environment, improves the accuracy and practicality of sound source positioning; through iterative optimization method, the positioning process is simple and efficient, and the positioning accuracy is high; and the method has good scalability, can be applied to two-dimensional / three-dimensional array structure and actual complex sound field, and provides more accurate and reliable sound source positioning solutions for intelligent security, robot navigation, video conference and other fields.

[0078] In an optional embodiment of the present application, the above step S1 includes: dividing the microphones in the two-dimensional microphone array into two groups of microphone arrays, preliminarily estimating the first sound source azimuth according to the far-field steering vector of the sound information collected by the first microphone array through beam forming estimation method or generalized cross-correlation delay estimation method, and preliminarily estimating the second sound source azimuth according to the far-field steering vector of the sound information collected by the second microphone array.

[0079] Beamforming estimation method adjusts the weighting coefficients of each microphone in the microphone array, so that the array enhances the sound signal in a specific direction and suppresses the sound signal in other directions, thereby achieving preliminary estimation of the sound source azimuth. This method uses the far-field steering vector, combined with the calculation of beam output power, to find the direction that maximizes the beam output power, which is the preliminary estimated sound source azimuth. This method can effectively improve the signal-to-noise ratio of sound source positioning and enhance the sensitivity to sound sources in a specific direction.

[0080] Generalized cross-correlation phase transform (GCC-PHAT) estimates the time difference of sound signals arriving at different microphones by calculating the cross-correlation function between two microphone signals and performing weighted processing, and then calculates the sound source azimuth. This method performs well in handling multipath effects and noise interference, and can provide more accurate time delay estimation, thereby improving the accuracy of sound source positioning. GCC-PHAT method emphasizes the peak value of cross-correlation function through phase transform (PHAT), making time delay estimation more accurate and reliable.

[0081] In an optional embodiment of the present application, the first sound source azimuth is preliminarily estimated according to the far-field steering vector of the sound information collected by the first microphone array, and the second sound source azimuth is preliminarily estimated according to the far-field steering vector of the sound information collected by the second microphone array, including:

[0082] The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated by the following formula:

[0083] ;

[0084] wherein, represents the frequency at which the amplitude is maximum after Fourier decomposition of the sound signal generated by the sound source, represents the speed of sound, represents the position of the i-th microphone in the first microphone array, represents the total number of microphones in the first microphone array, represents the sound source direction unit vector, i.e. , represents the sound source angle;

[0085] The beam output power of the sound information collected by the first microphone array is calculated by the following formula:

[0086] ;

[0087] wherein, represents the frequency domain form of the beam collected by the first microphone array, i.e. ​​

[0088] ,

[0089] represents the Fourier transform of the sound information collected by the i th microphone in the first microphone array;

[0090] The first sound source azimuth angle corresponding to the maximum beam output power is found according to the following formula:

[0091] ; wherein, represents the first sound source azimuth angle;

[0092] The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated by the following formula:

[0093] ;

[0094] wherein, represents the position of the i th microphone in the second microphone array, represents the total number of microphones in the second microphone array; The beam output power of the sound information collected by the second microphone array is calculated by the following formula:

[0095]

[0096] ;

[0097] wherein, represents the frequency domain form of the beam, that is,

[0098] ,

[0099] represents the Fourier transform of the sound information collected by the i th microphone in the second microphone array;

[0100] The second sound source azimuth angle corresponding to the maximum beam output power is found according to the following formula: ; wherein, represents the second sound source azimuth angle.

[0101] In an optional embodiment of the present application, the above S2 includes steps S21 to S23, wherein S21, S22, etc. are only step identifiers, and the execution order of the method does not necessarily follow the order from small to large, such as executing step S22 first and then executing step S21. The present application does not make any limitation.

[0102] ​​​​S21: Calculate the distance between the geometric center of the first microphone array and the preliminary sound source estimation position according to the following formula:

[0103] ;

[0104] As shown in Figure 2 , represents the distance between the geometric center of the first microphone array and the preliminary sound source estimation position, represents the distance between the preliminary sound source estimation position and the plane where the two-dimensional microphone array is located, represents the distance between the geometric center of the first microphone array and the geometric center of the second microphone array , represents the first sound source azimuth angle, represents the second sound source azimuth angle.

[0105] S22: Calculate the distance between the geometric center of the second microphone array and the preliminary sound source estimation position according to the following formula:

[0106] .

[0107] As shown in Figure 2 , represents the distance between the geometric center of the second microphone array and the preliminary sound source estimation position.

[0108] S23: Calculate the intersection position of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimation position according to the following formula:

[0109] .

[0110] Where, represents the preliminary sound source estimation position, represents the geometric center position of the first microphone array, represents the geometric center position of the second microphone array;

[0111] represents the directional vector of the first sound source azimuth angle, i.e. ,

[0112] represents the directional vector of the second sound source azimuth angle, i.e. .

[0113] In an optional embodiment of this application, step S3 includes: calculating a near-field steering vector based on the previously calculated sound source estimation position; iteratively calculating the first sound source azimuth angle and the second sound source azimuth angle using the near-field steering vector; calculating the position of the intersection of the first sound source azimuth angle and the second sound source azimuth angle after iteration based on geometric relationships, until the distance between the currently calculated intersection position and the previously calculated intersection position is less than a preset threshold; when the distance between the currently calculated intersection position and the previously calculated intersection position is less than the preset threshold, the currently calculated intersection position is taken as the near-field sound source estimation position; when calculating the near-field steering vector for the first time, the near-field steering vector is calculated based on the preliminary sound source estimation position.

[0114] The purpose of step S3 is to improve the accuracy of sound source localization through iterative optimization. It first calculates the near-field steering vector based on the previously calculated estimated sound source location. Then, it iteratively calculates the sound source azimuth using this vector and determines the intersection point of the azimuths based on geometric relationships. This process iterates until the distance between the intersection points of two consecutive calculations is less than a preset threshold, thus determining the estimated near-field sound source location. This process effectively corrects the initial estimation error, improving the accuracy and robustness of sound source localization, and is particularly suitable for high-precision sound source localization requirements in near-field environments.

[0115] In an optional embodiment of this application, calculating the near-field steering vector based on the previously calculated sound source estimated location includes:

[0116] According to the following formula, the previously calculated sound source estimation location is substituted into the near-field steering vector of the sound information collected by the first microphone array in the near-field model:

[0117] ;

[0118] in, The frequency at which the amplitude is maximum is obtained after performing Fourier decomposition on the sound signal generated by the sound source. Represents the speed of sound. Represents the first microphone array One microphone location, This represents the total number of microphones in the first microphone array;

[0119] According to the following formula, the estimated location of the sound source calculated in the previous calculation is substituted into the near-field steering vector of the sound information collected by the second microphone array in the near-field model:

[0120] ;

[0121] in, Represents the second microphone array One microphone location, This represents the total number of microphones in the second microphone array.

[0122] In an optional embodiment of this application, the near-field steering vector is used to iteratively calculate the first sound source azimuth angle and the second sound source azimuth angle, including: calculating the first sound source azimuth angle based on the near-field steering vector of the sound information collected by the first microphone array using beamforming estimation or generalized cross-correlation delay estimation, and calculating the second sound source azimuth angle based on the near-field steering vector of the sound information collected by the second microphone array.

[0123] In an optional embodiment of this application, step S4 includes: further optimizing the near-field sound source estimation location according to the following cost function to obtain the final sound source estimation location:

[0124] ;

[0125] in, The frequency at which the amplitude is maximum is obtained after performing Fourier decomposition on the sound signal generated by the sound source. Represents the speed of sound. Represents the second microphone in a two-dimensional microphone array One microphone location, This represents the total number of microphones in a two-dimensional microphone array. This represents the estimated location of the near-field sound source. This represents the final estimated location of the sound source.

[0126] It is understandable that step S4, through a full-array near-field localization optimization algorithm, deeply optimizes the estimated location of the near-field sound source, aiming to further improve the accuracy of sound source localization. This algorithm comprehensively considers the sound information collected by all microphones in the two-dimensional microphone array, and through complex mathematical models and calculations, finely adjusts the initially estimated near-field sound source location. After this optimization step, the final estimated sound source location will be closer to the actual sound source location, effectively reducing errors and providing more reliable sound source localization data support for fields such as intelligent security and robot navigation.

[0127] Secondly, such as Figure 3 As shown, this application discloses a high-precision sound source localization system based on a two-dimensional microphone array, including a two-dimensional microphone array and a processor. The two-dimensional microphone array is used to collect sound information in the environment, and the processor is used to execute the method as described in any of the first aspects.

[0128] It can be understood that the high-precision sound source positioning system based on the two-dimensional microphone array provided in the application does not need to rely on external devices such as laser ranging, avoids directional deviation and reflection problems, and reduces costs; it is particularly suitable for near-field environments, and through near-field model iterative optimization, the accuracy and practicability of sound source positioning are significantly improved; combined with the full-array near-field positioning optimization algorithm, the accurate estimation of the sound source position is realized; the system also has good scalability and can be flexibly applied to two-dimensional or three-dimensional array structures and complex sound field environments, providing an efficient and reliable sound source positioning solution for intelligent security, robot navigation and other fields.

[0129] In the optional embodiment of the application, the arrangement form of the two-dimensional microphone array can be at least one of the following: a linear array; a circular array; a multi-spiral arm array.

[0130] The expressions "first", "second", "the first" or "the second" used in various embodiments of the present disclosure can modify various components regardless of order and / or importance, but these expressions do not limit the corresponding components. The above expressions are only configured for the purpose of distinguishing elements from other elements. For example, the first user equipment and the second user equipment represent different user equipment, although both are user equipment. For example, without departing from the scope of the present disclosure, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element.

[0131] When an element (for example, a first element) is referred to as being "(operatively or communicatively) coupled with" or "(operatively or communicatively) coupled to" another element (for example, a second element) or "connected to" another element (for example, a second element), it should be understood that the one element is directly connected to the other element or the one element is indirectly connected to the other element via another element (for example, a third element). In contrast, it can be understood that when an element (for example, a first element) is referred to as being "directly connected" or "directly coupled" to another element (a second element), then no element (for example, a third element) is inserted between the two.

[0132] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a", "comprising", or "comprises" does not, without further qualification, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element. The singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. The terms "connected" and "coupled" are used broadly and encompass both direct and indirect connections and couplings.

[0133] The above description is only optional embodiments of the present application and the explanation of the technical principles applied. Those skilled in the art should understand that the inventive scope of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by the above features and the similar technical features disclosed in the present application (but not limited to) are replaced with each other.

[0134] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."

[0135] The above description is only optional embodiments of the present application and the explanation of the technical principles applied. Those skilled in the art should understand that the inventive scope of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by the above features and the similar technical features disclosed in the present application (but not limited to) are replaced with each other.

[0136] The above description is only optional embodiments of the present application and the explanation of the technical principles applied. Those skilled in the art should understand that the inventive scope of the present application is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by the above features and the similar technical features disclosed in the present application (but not limited to) are replaced with each other.

Claims

1. A high-precision sound source positioning method based on a two-dimensional microphone array, characterized in that, The method comprises the following steps: dividing a two-dimensional microphone array into two groups of microphone arrays, preliminarily estimating a first sound source azimuth angle in a far-field model through sound information collected by a first microphone array, and preliminarily estimating a second sound source azimuth angle through sound information collected by a second microphone array; calculating a straight line intersection position of the first sound source azimuth angle and the second sound source azimuth angle as a preliminary sound source estimation position according to a geometric relationship; bringing the preliminary sound source estimation position into a near-field model to iteratively optimize a near-field sound source estimation position; optimizing the near-field sound source estimation position again through a full-array near-field positioning optimization algorithm to obtain a final sound source estimation position; the step of calculating the straight line intersection position of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimation position according to the geometric relationship comprises the following steps: calculating a distance between a geometric center of the first microphone array and the preliminary sound source estimation position according to the following formula: ; wherein, a distance between a geometric center of the first microphone array and the preliminary sound source estimate position, a distance between the preliminary sound source estimate position and a plane in which the two-dimensional microphone array is located, a distance between a geometric center of the first microphone array and a geometric center of the second microphone array, a first sound source azimuth angle, a second sound source azimuth angle; calculating a distance between a geometric center of the second microphone array and the preliminary sound source estimation position according to the following formula: ; wherein, represents a distance between a geometric center of the second microphone array and the preliminary sound source estimate location; calculating the straight line intersection position of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimation position according to the following formula: ; wherein, represents the preliminary sound source estimation position, represents a geometric center position of the first microphone array, represents a geometric center position of the second microphone array; The orientation vector representing the azimuth angle of the first sound source, i.e. , The orientation vector representing the azimuth angle of the second sound source, i.e. ; the step of bringing the preliminary sound source estimation position into the near-field model to iteratively optimize the near-field sound source estimation position comprises the following steps: calculating a near-field steering vector according to a sound source estimation position calculated in a previous iteration, iteratively calculating the first sound source azimuth angle and the second sound source azimuth angle by using the near-field steering vector, and calculating a straight line intersection position of the first sound source azimuth angle and the second sound source azimuth angle after the iteration according to a geometric relationship until a distance between the straight line intersection position calculated in the current iteration and the straight line intersection position calculated in the previous iteration is less than a preset threshold value; when the distance between the straight line intersection position calculated in the current iteration and the straight line intersection position calculated in the previous iteration is less than the preset threshold value, taking the straight line intersection position calculated in the current iteration as the near-field sound source estimation position; when the near-field steering vector is calculated for the first time, calculating the near-field steering vector according to the preliminary sound source estimation position.

2. The high-precision sound source positioning method based on a two-dimensional microphone array according to claim 1, wherein the step of preliminarily estimating the first sound source azimuth angle in the far-field model through the sound information collected by the first microphone array and preliminarily estimating the second sound source azimuth angle through the sound information collected by the second microphone array comprises the following steps: preliminarily estimating the first sound source azimuth angle according to a far-field steering vector of the sound information collected by the first microphone array and preliminarily estimating the second sound source azimuth angle according to a far-field steering vector of the sound information collected by the second microphone array by using a beamforming estimation method or a generalized cross-correlation delay estimation method.

3. The high-precision sound source positioning method based on a two-dimensional microphone array according to claim 2, wherein the step of preliminarily estimating the first sound source azimuth angle according to the far-field steering vector of the sound information collected by the first microphone array and preliminarily estimating the second sound source azimuth angle according to the far-field steering vector of the sound information collected by the second microphone array by using the beamforming estimation method comprises the following steps: ​ ​ The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated by the following formula: ; wherein, a frequency of a maximum amplitude obtained after Fourier decomposition of a sound signal representing a sound source, represents a sound velocity, represents a position of a th microphone in the first microphone array, represents a total number of microphones in the first microphone array, represents any positive integer in a range of 1 to , represents a sound source direction unit vector, i.e. , represents a sound source angle; Beam output power of sound information captured by the first microphone array is calculated by the following formula : ; wherein, represent a frequency domain form of the beams collected by the first microphone array, i.e. , a Fourier transform representative of sound information collected by a first microphone of said first microphone array; and a Fourier transform representative of sound information collected by a second microphone of said first microphone array. The first sound source azimuth corresponding to the maximum beam output power is found according to the following formula: ; wherein represents the first sound source azimuth angle; The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated by the following formula: ; wherein represents a position of a th microphone in the second microphone array, represents a total number of microphones in the second microphone array, represents any positive integer in the range of 1 to . Beam output power of sound information captured by the second microphone array is calculated by the following formula : ; wherein, represent a frequency domain form of the beams collected by the second microphone array, i.e. , a Fourier transform representative of sound information collected by a first microphone of said first microphone array; a Fourier transform representative of sound information collected by a first microphone of said first microphone array; The second sound source azimuth corresponding to the maximum beam output power is found according to the following formula: ; wherein, represents the second sound source azimuth angle.

4. The high-precision sound source positioning method based on a two-dimensional microphone array according to claim 1, wherein the near-field steering vector is calculated according to the last calculated sound source estimated position, and the near-field steering vector of the sound information collected by the first microphone array in the near-field model is calculated by the following formula: The near-field steering vector of the sound information collected by the second microphone array in the near-field model is calculated by the following formula:

5. The high-precision sound source positioning method based on a two-dimensional microphone array according to claim 4, wherein the first sound source azimuth and the second sound source azimuth are iteratively calculated using the near-field steering vector, and the first sound source azimuth is calculated according to the near-field steering vector of the sound information collected by the first microphone array by a beamforming estimation method or a generalized cross-correlation delay estimation method, and the second sound source azimuth is calculated according to the near-field steering vector of the sound information collected by the second microphone array by the beamforming estimation method or the generalized cross-correlation delay estimation method. ; wherein, a frequency of a maximum amplitude obtained after Fourier decomposition of a sound signal representing a sound source, represents a sound velocity, represents a position of a th microphone in the first microphone array, represents a total number of microphones in the first microphone array, represents any positive integer in the range of 1 to ​ 6. The high-precision sound source positioning method based on a two-dimensional microphone array according to claim 1, wherein the near-field sound source estimated position is further optimized by a full-array near-field positioning optimization algorithm to obtain a final sound source estimated position, and the near-field sound source estimated position is further optimized by the following cost function optimization to obtain the final sound source estimated position: ; wherein represents a position of a microphone in the second microphone array, represents the total number of microphones in the second microphone array, represents any positive integer in the range of 1 to .​ 7. A high-precision sound source positioning system based on a two-dimensional microphone array, comprising a two-dimensional microphone array and a processor, wherein the two-dimensional microphone array is used to collect sound information in an environment, and the processor is used to execute the method according to any one of claims 1 to 6.

8. The high-precision sound source positioning system based on a two-dimensional microphone array according to claim 7, wherein the two-dimensional microphone array has at least one of the following arrangement forms: a linear array; a circular array; a multi-spiral arm array. ​ ; wherein, a frequency of a maximum amplitude obtained after Fourier decomposition of a sound signal representing a sound source, representing a sound velocity, representing a position of a th microphone in the two-dimensional microphone array, representing a total number of microphones in the two-dimensional microphone array, representing any positive integer in a range of 1 to representing any positive integer in a range of 1 to representing an estimated position of the near-field sound source, representing a final estimated position of the sound source. ​ ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Video and audio conference equipment, terminal equipment, sound source positioning method and medium

    CN114095687A

  • Sound source localization method and device based on microphone linear double arrays and medium

    CN115128544A

  • Sound source localization method and system based on improved adaptive beam forming

    CN119511200A