High-precision sound source positioning method and system based on two-dimensional microphone array

Through two-dimensional microphone array grouping and iterative optimization algorithms, the problem of low sound source positioning accuracy in near-field environments is solved, and high-precision and efficient sound source positioning is achieved. It is suitable for intelligent security, robot navigation and video conferencing and other fields.

CN120522639AActive Publication Date: 2025-08-22CHINA SOUTHERN POWER GRID INTERNET SERVICE CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511028959.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-08-22
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

The existing sound source positioning technology has low positioning accuracy in near-field environments and has large errors relying on external ranging equipment.

Method used

Two-dimensional microphone array grouping is used, and the far-field model is used to initially estimate the azimuth angle of the sound source, combined with geometric relationships and iterative optimization of the near-field model, the final sound source position is obtained through the full array near-field positioning optimization algorithm, avoiding the directional deviation of laser ranging.

Benefits of technology

It improves the accuracy and practicality of sound source positioning, especially suitable for near-field environments, is simple and efficient, and reduces costs. It is suitable for intelligent security, robot navigation and video conferencing and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120522639A_ABST
    Figure CN120522639A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision sound source positioning method and system based on a two-dimensional microphone array, and aims to solve the problem of low sound source positioning precision in a near-field environment. Grouping microphone arrays, and preliminarily estimating a sound source azimuth angle by using a far-field model; calculating an initial sound source position through the geometrical relationship; iteratively optimizing the sound source position in combination with the near-field model; and finally, obtaining a final sound source position through a full-array near-field positioning optimization algorithm. The high-precision sound source positioning system disclosed by the invention does not need external distance measuring equipment, is suitable for a near-field scene, and improves the practicability and robustness of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of sound source localization, and in particular to a high-precision sound source localization method and system based on a two-dimensional microphone array. Background Art

[0002] Sound source localization technology has broad and important applications in numerous fields. In intelligent security, accurately determining the sound source can promptly identify the source of unusual sounds and ensure safety. In robot navigation, sound source localization helps robots identify the source of voice commands in their surroundings, enabling more intelligent interactions and actions. In video conferencing systems, precisely locating the speaker's sound source optimizes camera tracking and enhances the conferencing experience.

[0003] However, existing sound source localization technologies face numerous challenges in near-field environments. Traditional sound source localization methods mostly construct steering vectors based on far-field models. This model significantly reduces localization accuracy in near-field environments, where the distance between the sound source and the microphone is relatively close. Furthermore, many existing technologies rely on external ranging devices (such as laser rangefinders) to determine the distance to the sound source. This approach often results in significant errors in sound source localization in complex scenarios. Summary of the Invention

[0004] The purpose of this application is to provide a high-precision sound source localization method and system based on a two-dimensional microphone array, which can improve the above-mentioned problems.

[0005] First, the present application provides a high-precision sound source localization method based on a two-dimensional microphone array, which includes steps S1 to S4, where S1, S2, etc. are only step identifiers, and the execution order of the method is not necessarily in ascending order of numbers. For example, step S2 may be executed first and then step S1. This application does not impose any restrictions.

[0006] S1: Divide the two-dimensional microphone array into two groups of microphone arrays. In the far-field model, the azimuth of the first sound source is preliminarily estimated through the sound information collected by the first microphone array, and the azimuth of the second sound source is preliminarily estimated through the sound information collected by the second microphone array.

[0007] S2: Calculate the intersection of the first sound source azimuth angle and the second sound source azimuth angle according to the geometric relationship as a preliminary sound source estimated position.

[0008] S3: Bringing the preliminary sound source estimated position into the near-field model to iteratively optimize the near-field sound source estimated position.

[0009] S4: Optimizing the near-field sound source estimated position again using a full-array near-field positioning optimization algorithm to obtain a final sound source estimated position.

[0010] It can be understood that this application proposes a high-precision sound source localization method based on a two-dimensional microphone array, including: grouping the microphone array and using the far-field model to preliminarily estimate the sound source azimuth; calculating the preliminary sound source position through geometric relationships; iteratively optimizing the sound source position in combination with the near-field model; and finally obtaining the final sound source position through a full-array near-field positioning optimization algorithm. This application does not require additional ranging equipment, avoids the direction inconsistency or lack of reflection problems that may be caused by laser ranging, and reduces costs; is particularly suitable for near-field environments, and improves the accuracy and practicality of sound source localization; through iterative optimization methods, the positioning process is simple and efficient, with high positioning accuracy; and the method has good scalability and can be applied to two-dimensional / three-dimensional array structures and actual complex sound fields, providing more accurate and reliable sound source localization solutions for multiple fields such as intelligent security, robot navigation, and video conferencing.

[0011] In an optional embodiment of the present application, the above-mentioned step S1 includes: dividing the microphones in the two-dimensional microphone array into two groups of microphone arrays, and preliminarily estimating the azimuth angle of the first sound source based on the far-field steering vector of the sound information collected by the first microphone array through a beamforming estimation method or a generalized cross-correlation delay estimation method, and preliminarily estimating the azimuth angle of the second sound source based on the far-field steering vector of the sound information collected by the second microphone array.

[0012] In an optional embodiment of the present application, preliminarily estimating the azimuth of the first sound source based on the far-field steering vector of the sound information collected by the first microphone array by a beamforming estimation method, and preliminarily estimating the azimuth of the second sound source based on the far-field steering vector of the sound information collected by the second microphone array, includes: The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated using the following formula: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array, Represents the unit vector of the sound source direction, that is , Represents the angle of the sound source; The beam output power of the sound information collected by the first microphone array is calculated by the following formula: : ; in, Represents the frequency domain form of the beam collected by the first microphone array, that is, , Represents the first microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, the azimuth angle of the first sound source corresponding to the maximum output power of the beam is found: ;in, Represents the azimuth of the first sound source; The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated using the following formula: ; in, Represents the second microphone array microphone positions, represents the total number of microphones in the second microphone array; The beam output power of the sound information collected by the second microphone array is calculated by the following formula: : ; in, Represents the frequency domain form of the beam collected by the second microphone array, that is, , Represents the second microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, the azimuth angle of the second sound source corresponding to the maximum output power of the beam is found: ;in, Represents the azimuth of the second sound source.

[0013] In an optional embodiment of the present application, the above-mentioned S2 includes the following steps S21 to S23, wherein S21, S22, etc. are merely step identifiers, and the execution order of the method is not necessarily in ascending order of numbers. For example, step S22 may be executed first and then step S21. This application does not impose any restrictions.

[0014] S21: Calculate the distance between the geometric center of the first microphone array and the preliminary sound source estimated position according to the following formula: ; in, represents the distance between the geometric center of the first microphone array and the estimated position of the preliminary sound source, represents the distance between the preliminary sound source estimation position and the plane where the two-dimensional microphone array is located, represents the distance between the geometric center of the first microphone array and the geometric center of the second microphone array, represents the azimuth of the first sound source, represents the azimuth of the second sound source; S22: Calculate the distance between the geometric center of the second microphone array and the preliminary sound source estimated position according to the following formula: ; represents the distance between the geometric center of the second microphone array and the estimated position of the preliminary sound source; S23: Calculate the intersection of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimated position according to the following formula: ; in, represents the preliminary estimated location of the sound source, represents the geometric center position of the first microphone array, represents the geometric center position of the second microphone array; The directional vector representing the azimuth of the first sound source, i.e. , Represents the directional vector of the second sound source azimuth, that is .

[0015] In an optional embodiment of the present application, the above-mentioned step S3 includes: calculating the near-field steering vector based on the sound source estimated position calculated last time, iteratively calculating the first sound source azimuth and the second sound source azimuth using the near-field steering vector, and calculating the straight line intersection position of the iterative first sound source azimuth and the second sound source azimuth according to the geometric relationship, until the distance between the straight line intersection position calculated this time and the straight line intersection position calculated last time is less than a preset threshold; when the distance between the straight line intersection position calculated this time and the straight line intersection position calculated last time is less than a preset threshold, the straight line intersection position calculated this time is used as the near-field sound source estimated position; when the near-field steering vector is calculated for the first time, the near-field steering vector is calculated based on the preliminary sound source estimated position.

[0016] As you can understand, step S3 aims to improve the accuracy of sound source localization through iterative optimization. It first calculates the near-field steering vector based on the previously calculated estimated sound source position. This vector is then used to iteratively calculate the sound source azimuth, and the intersection of these azimuths is determined based on geometric relationships. This process continues until the distance between the intersections of two consecutive calculations is less than a preset threshold, thereby determining the estimated near-field sound source position. This process effectively corrects errors in the initial estimate, improving the accuracy and robustness of sound source localization, making it particularly suitable for high-precision sound source localization in near-field environments.

[0017] In an optional embodiment of the present application, the calculating of the near-field steering vector according to the last calculated estimated sound source position includes: According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the first microphone array in the near-field model: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array; According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the second microphone array in the near-field model: ; in, Represents the second microphone array microphone positions, represents the total number of microphones in the second microphone array.

[0018] In an optional embodiment of the present application, the first sound source azimuth and the second sound source azimuth are iteratively calculated using the near-field steering vector, including: calculating the first sound source azimuth according to the near-field steering vector of the sound information collected by the first microphone array through a beamforming estimation method or a generalized cross-correlation delay estimation method, and calculating the second sound source azimuth according to the near-field steering vector of the sound information collected by the second microphone array.

[0019] In an optional embodiment of the present application, the above step S4 includes: re-optimizing the near-field sound source estimated position according to the following cost function optimization to obtain a final sound source estimated position: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone positions, represents the total number of microphones in the two-dimensional microphone array, represents the estimated position of the near-field sound source, Represents the final estimated sound source position.

[0020] As you can understand, step S4 uses a full-array near-field localization optimization algorithm to deeply optimize the estimated near-field sound source position, aiming to further improve the accuracy of sound source localization. This algorithm comprehensively considers the sound information collected by all microphones in the two-dimensional microphone array and, through complex mathematical models and calculations, fine-tunes the initial estimated near-field sound source position. After this optimization step, the final sound source estimate is closer to the actual sound source location, effectively reducing errors and providing more reliable sound source localization data support for fields such as intelligent security and robotic navigation.

[0021] In a second aspect, the present application discloses a high-precision sound source localization system based on a two-dimensional microphone array, comprising a two-dimensional microphone array and a processor, wherein the two-dimensional microphone array is used to collect sound information in the environment, and the processor is used to execute the method described in any one of the first aspects.

[0022] It can be understood that the high-precision sound source localization system based on a two-dimensional microphone array provided by this application does not need to rely on external equipment such as laser ranging, avoids directional deviation and reflection problems, and reduces costs; it is particularly suitable for near-field environments, and through iterative optimization of the near-field model, it significantly improves the accuracy and practicality of sound source localization; combined with the full-array near-field positioning optimization algorithm, it achieves accurate estimation of the sound source position; the system also has good scalability and can be flexibly applied to two-dimensional or three-dimensional array structures, as well as complex sound field environments, providing efficient and reliable sound source localization solutions for intelligent security, robot navigation and other fields.

[0023] In an optional embodiment of the present application, the arrangement form of the two-dimensional microphone array can be at least one of the following: a straight array; a circular array; a multi-spiral arm array. Beneficial effects

[0024] This application proposes a high-precision sound source localization method and system based on a two-dimensional microphone array, aiming to solve the problem of low sound source localization accuracy in near-field environments. The microphone array is grouped, and the far-field model is used to preliminarily estimate the azimuth of the sound source; the preliminary sound source position is calculated through geometric relationships; the sound source position is iteratively optimized in combination with the near-field model; and finally, the final sound source position is obtained through a full-array near-field positioning optimization algorithm. This application does not require additional ranging equipment, avoids the direction inconsistency or lack of reflection problems that may be caused by laser ranging, and reduces costs; it is particularly suitable for near-field environments, and improves the accuracy and practicality of sound source localization; through iterative optimization methods, the positioning process is simple and efficient, and the positioning accuracy is high; and the method has good scalability and can be applied to two-dimensional / three-dimensional array structures and actual complex sound fields, providing more accurate and reliable sound source localization solutions for intelligent security, robot navigation, video conferencing and other fields.

[0025] This application has been validated through MATLAB simulations and field measurements. In typical experimental scenarios (with a sound source and array center distance of 0.5–2 meters), the error reduction compared to traditional far-field model positioning methods is over 60%. The sound source location can be stably output even without laser assistance, demonstrating engineering feasibility and potential for practical application.

[0026] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, optional embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0028] Figure 1 This is a schematic diagram of dividing a two-dimensional microphone array into two groups of microphone arrays provided by the present application; Figure 2 This is a schematic diagram provided by the present application of calculating the intersection of the first sound source azimuth angle and the second sound source azimuth angle as the preliminary sound source estimation position based on the geometric relationship; Figure 3 This is a schematic diagram of the connection relationship of the high-precision sound source localization system based on a two-dimensional microphone array provided in this application. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] First, the present application provides a high-precision sound source localization method based on a two-dimensional microphone array, which includes steps S1 to S4, where S1, S2, etc. are only step identifiers, and the execution order of the method is not necessarily in ascending order of numbers. For example, step S2 may be executed first and then step S1. This application does not impose any restrictions.

[0031] S1: Divide the two-dimensional microphone array into two groups of microphone arrays. In the far-field model, the azimuth of the first sound source is preliminarily estimated through the sound information collected by the first microphone array, and the azimuth of the second sound source is preliminarily estimated through the sound information collected by the second microphone array.

[0032] Two-dimensional microphone arrays can be arranged in a linear array, a circular array, or a multi-helical array. Each arrangement has its own advantages: a linear array is simple and easy to implement, suitable for initial estimation of the azimuth angle of a sound source; a circular array provides more uniform azimuth coverage, facilitating omnidirectional sound source localization; and a multi-helical array combines the advantages of the first two, enabling high-precision positioning in more complex environments.

[0033] The microphone division method primarily divides the microphone array into two groups. The specific division method is not limited and can be split 50 / 50 or divided into two groups of unequal numbers. Each microphone group can operate independently, reducing mutual interference. At the same time, the comparison and fusion of the two sets of data can improve the accuracy and robustness of sound source localization. Especially in near-field environments, this division method, combined with iterative optimization of the near-field model, can significantly improve the accuracy of sound source localization.

[0034] like Figure 1 As shown, the microphone array arranged in a multi-spiral arm array is divided into a first microphone array and a second microphone array with different numbers of microphones.

[0035] S2: Calculate the intersection of the first sound source azimuth and the second sound source azimuth based on the geometric relationship as the preliminary sound source estimation position.

[0036] S3: Bring the preliminary sound source estimated position into the near-field model to iteratively optimize the near-field sound source estimated position.

[0037] S4: The near-field sound source estimated position is further optimized using the full-array near-field positioning optimization algorithm to obtain the final sound source estimated position.

[0038] It can be understood that this application proposes a high-precision sound source localization method based on a two-dimensional microphone array, including: grouping the microphone array and using the far-field model to preliminarily estimate the sound source azimuth; calculating the preliminary sound source position through geometric relationships; iteratively optimizing the sound source position in combination with the near-field model; and finally obtaining the final sound source position through a full-array near-field positioning optimization algorithm. This application does not require additional ranging equipment, avoids the direction inconsistency or lack of reflection problems that may be caused by laser ranging, and reduces costs; is particularly suitable for near-field environments, and improves the accuracy and practicality of sound source localization; through iterative optimization methods, the positioning process is simple and efficient, with high positioning accuracy; and the method has good scalability and can be applied to two-dimensional / three-dimensional array structures and actual complex sound fields, providing more accurate and reliable sound source localization solutions for multiple fields such as intelligent security, robot navigation, and video conferencing.

[0039] In an optional embodiment of the present application, the above-mentioned step S1 includes: dividing the microphones in the two-dimensional microphone array into two groups of microphone arrays, and preliminarily estimating the azimuth angle of the first sound source based on the far-field steering vector of the sound information collected by the first microphone array through a beamforming estimation method or a generalized cross-correlation delay estimation method, and preliminarily estimating the azimuth angle of the second sound source based on the far-field steering vector of the sound information collected by the second microphone array.

[0040] Beamforming estimation methods adjust the weighting coefficients of each microphone in a microphone array so that the array enhances sound signals from specific directions while suppressing those from other directions, thereby providing a preliminary estimate of the sound source's azimuth. This method utilizes far-field steering vectors, combined with beam output power calculations, to find the direction that maximizes beam output power, which serves as the initial estimate of the sound source's azimuth. This method effectively improves the signal-to-noise ratio for sound source localization and enhances sensitivity to sound sources in specific directions.

[0041] The generalized cross-correlation delay estimation method (GCC-PHAT) calculates the cross-correlation function between two microphone signals and performs a weighted process to estimate the time difference between the sound signal reaching the different microphones, thereby inferring the azimuth of the sound source. This method excels in dealing with multipath effects and noise interference, providing more accurate time delay estimates and thus improving the accuracy of sound source localization. The GCC-PHAT method uses phase shifting (PHAT) to emphasize the peak of the cross-correlation function, making the time delay estimation more accurate and reliable.

[0042] In an optional embodiment of the present application, preliminarily estimating the azimuth of the first sound source based on the far-field steering vector of the sound information collected by the first microphone array by a beamforming estimation method, and preliminarily estimating the azimuth of the second sound source based on the far-field steering vector of the sound information collected by the second microphone array, includes: The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated using the following formula: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array, Represents the unit vector of the sound source direction, that is , Represents the angle of the sound source; The beam output power of the sound information collected by the first microphone array is calculated by the following formula: : ; in, Represents the frequency domain form of the beam collected by the first microphone array, that is, , Represents the first microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, find the azimuth angle of the first sound source corresponding to the maximum beam output power: ;in, Represents the azimuth of the first sound source; The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated using the following formula: ; in, Represents the second microphone array microphone positions, represents the total number of microphones in the second microphone array; The beam output power of the sound information collected by the second microphone array is calculated by the following formula: : ; in, Represents the frequency domain form of the beam, that is , Represents the second microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, find the azimuth angle of the second sound source corresponding to the maximum beam output power: ;in, Represents the azimuth of the second sound source.

[0043] In an optional embodiment of the present application, the above-mentioned S2 includes the following steps S21 to S23, wherein S21, S22, etc. are merely step identifiers, and the execution order of the method is not necessarily in ascending order of numbers. For example, step S22 may be executed first and then step S21. This application does not impose any restrictions.

[0044] S21: Calculate the distance between the geometric center of the first microphone array and the preliminary sound source estimated position according to the following formula: ; like Figure 2 As shown, Represents the geometric center of the first microphone array The distance from the initial sound source estimate, Represents the distance between the estimated position of the preliminary sound source and the plane where the two-dimensional microphone array is located, Represents the geometric center of the first microphone array and the geometric center of the second microphone array distance, represents the azimuth of the first sound source, Represents the azimuth of the second sound source.

[0045] S22: Calculate the distance between the geometric center of the second microphone array and the preliminary sound source estimated position according to the following formula: .

[0046] like Figure 2 As shown, Represents the geometric center of the second microphone array The distance to the preliminary sound source estimate.

[0047] S23: Calculate the intersection of the first sound source azimuth and the second sound source azimuth according to the following formula as the preliminary sound source estimated position: .

[0048] in, Represents the preliminary estimated location of the sound source, represents the geometric center position of the first microphone array, represents the geometric center position of the second microphone array; The directional vector representing the azimuth of the first sound source, i.e. , The directional vector representing the azimuth of the second sound source, i.e. .

[0049] In an optional embodiment of the present application, the above-mentioned step S3 includes: calculating the near-field steering vector based on the sound source estimated position calculated last time, iteratively calculating the first sound source azimuth and the second sound source azimuth using the near-field steering vector, and calculating the straight line intersection position of the iterative first sound source azimuth and the second sound source azimuth according to the geometric relationship, until the distance between the straight line intersection position calculated this time and the straight line intersection position calculated last time is less than a preset threshold; when the distance between the straight line intersection position calculated this time and the straight line intersection position calculated last time is less than a preset threshold, the straight line intersection position calculated this time is used as the near-field sound source estimated position; when the near-field steering vector is calculated for the first time, the near-field steering vector is calculated based on the preliminary sound source estimated position.

[0050] As you can understand, step S3 aims to improve the accuracy of sound source localization through iterative optimization. It first calculates the near-field steering vector based on the previously calculated estimated sound source position. This vector is then used to iteratively calculate the sound source azimuth, and the intersection of these azimuths is determined based on geometric relationships. This process continues until the distance between the intersections of two consecutive calculations is less than a preset threshold, thereby determining the estimated near-field sound source position. This process effectively corrects errors in the initial estimate, improving the accuracy and robustness of sound source localization, making it particularly suitable for high-precision sound source localization in near-field environments.

[0051] In an optional embodiment of the present application, calculating the near-field steering vector according to the last calculated estimated sound source position includes: According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the first microphone array in the near-field model: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array; According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the second microphone array in the near-field model: ;

[0052] in, Represents the second microphone array microphone positions, Represents the total number of microphones in the second microphone array.

[0053] In an optional embodiment of the present application, the azimuth angle of the first sound source and the azimuth angle of the second sound source are iteratively calculated using a near-field steering vector, including: calculating the azimuth angle of the first sound source based on the near-field steering vector of the sound information collected by the first microphone array through a beamforming estimation method or a generalized cross-correlation delay estimation method, and calculating the azimuth angle of the second sound source based on the near-field steering vector of the sound information collected by the second microphone array.

[0054] In an optional embodiment of the present application, the above step S4 includes: re-optimizing the near-field sound source estimated position according to the following cost function optimization to obtain a final sound source estimated position: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone positions, represents the total number of microphones in the two-dimensional microphone array, represents the estimated position of the near-field sound source, Represents the final estimated sound source position.

[0055] As you can understand, step S4 uses a full-array near-field localization optimization algorithm to deeply optimize the estimated near-field sound source position, aiming to further improve the accuracy of sound source localization. This algorithm comprehensively considers the sound information collected by all microphones in the two-dimensional microphone array and, through complex mathematical models and calculations, fine-tunes the initial estimated near-field sound source position. After this optimization step, the final sound source estimate is closer to the actual sound source location, effectively reducing errors and providing more reliable sound source localization data support for fields such as intelligent security and robotic navigation.

[0056] Second, as Figure 3 As shown, the present application discloses a high-precision sound source localization system based on a two-dimensional microphone array, comprising a two-dimensional microphone array and a processor, wherein the two-dimensional microphone array is used to collect sound information in the environment, and the processor is used to execute any method as in the first aspect.

[0057] It can be understood that the high-precision sound source localization system based on a two-dimensional microphone array provided by this application does not need to rely on external equipment such as laser ranging, avoids directional deviation and reflection problems, and reduces costs; it is particularly suitable for near-field environments, and through iterative optimization of the near-field model, it significantly improves the accuracy and practicality of sound source localization; combined with the full-array near-field positioning optimization algorithm, it achieves accurate estimation of the sound source position; the system also has good scalability and can be flexibly applied to two-dimensional or three-dimensional array structures, as well as complex sound field environments, providing efficient and reliable sound source localization solutions for intelligent security, robot navigation and other fields.

[0058] In an optional embodiment of the present application, the arrangement of the two-dimensional microphone array may be at least one of the following: a linear array; a circular array; a multi-spiral arm array.

[0059] The terms "first," "second," "the first," or "the second" used in various embodiments of the present disclosure may modify various components regardless of order and / or importance, but these terms do not limit the corresponding components. The above terms are configured solely for the purpose of distinguishing an element from other elements. For example, a first user device and a second user device represent different user devices, even though both are user devices. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the present disclosure.

[0060] When one element (for example, a first element) is referred to as being “(operably or communicably) coupled” or “(operably or communicably) coupled to” or “connected to” another element (for example, a second element), it should be understood that the one element is directly connected to the other element or that the one element is indirectly connected to the other element via yet another element (for example, a third element). Conversely, it should be understood that when an element (for example, a first element) is referred to as being “directly connected” or “directly coupled” to another element (the second element), there is no element (for example, a third element) interposed therebetween.

[0061] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.

[0062] The above description is merely an optional embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

[0063] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0064] The above description is merely an optional embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.

[0065] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A high-precision sound source localization method based on a two-dimensional microphone array, characterized in that: include: The two-dimensional microphone array is divided into two groups of microphone arrays. In a far-field model, the azimuth of the first sound source is preliminarily estimated using the sound information collected by the first microphone array, and the azimuth of the second sound source is preliminarily estimated using the sound information collected by the second microphone array. Calculate the intersection of the first sound source azimuth angle and the second sound source azimuth angle according to the geometric relationship as a preliminary sound source estimated position; Bringing the preliminary sound source estimated position into the near-field model for iterative optimization to obtain the near-field sound source estimated position; The near-field sound source estimated position is further optimized using a full-array near-field positioning optimization algorithm to obtain a final sound source estimated position.

2. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 1, characterized in that: In the far-field model, preliminarily estimating the azimuth of the first sound source through the sound information collected by the first microphone array, and preliminarily estimating the azimuth of the second sound source through the sound information collected by the second microphone array, includes: Through the beamforming estimation method or the generalized cross-correlation delay estimation method, the azimuth angle of the first sound source is preliminarily estimated based on the far-field steering vector of the sound information collected by the first microphone array, and the azimuth angle of the second sound source is preliminarily estimated based on the far-field steering vector of the sound information collected by the second microphone array.

3. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 2, characterized in that: The method of preliminarily estimating the azimuth of the first sound source based on the far-field steering vector of the sound information collected by the first microphone array and preliminarily estimating the azimuth of the second sound source based on the far-field steering vector of the sound information collected by the second microphone array by using the beamforming estimation method includes: The far-field steering vector of the sound information collected by the first microphone array in the far-field model is calculated using the following formula: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array, Represents the unit vector of the sound source direction, that is , Represents the angle of the sound source; The beam output power of the sound information collected by the first microphone array is calculated by the following formula: : ; in, represents the frequency domain form of the beam collected by the first microphone array, that is, , Represents the first microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, the azimuth angle of the first sound source corresponding to the maximum output power of the beam is found: ;in, Represents the azimuth of the first sound source; The far-field steering vector of the sound information collected by the second microphone array in the far-field model is calculated using the following formula: ; in, Represents the second microphone array microphone positions, represents the total number of microphones in the second microphone array; The beam output power of the sound information collected by the second microphone array is calculated by the following formula: : ; in, represents the frequency domain form of the beam collected by the second microphone array, that is, , Represents the second microphone array The Fourier transform of the sound information collected by the microphone; According to the following formula, the azimuth angle of the second sound source corresponding to the maximum output power of the beam is found: ;in, Represents the azimuth of the second sound source.

4. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 1, characterized in that: The step of calculating the intersection of the first sound source azimuth angle and the second sound source azimuth angle as a preliminary sound source estimated position according to the geometric relationship includes: The distance between the geometric center of the first microphone array and the estimated position of the preliminary sound source is calculated according to the following formula: ; in, represents the distance between the geometric center of the first microphone array and the estimated position of the preliminary sound source, represents the distance between the preliminary sound source estimation position and the plane where the two-dimensional microphone array is located, represents the distance between the geometric center of the first microphone array and the geometric center of the second microphone array, represents the azimuth of the first sound source, represents the azimuth of the second sound source; The distance between the geometric center of the second microphone array and the preliminary sound source estimated position is calculated according to the following formula: ; in, represents the distance between the geometric center of the second microphone array and the estimated position of the preliminary sound source; The intersection of the first sound source azimuth angle and the second sound source azimuth angle is calculated as the preliminary sound source estimated position according to the following formula: ; in, represents the preliminary estimated location of the sound source, represents the geometric center position of the first microphone array, represents the geometric center position of the second microphone array; The directional vector representing the azimuth of the first sound source, i.e. , Represents the directional vector of the second sound source azimuth, that is .

5. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 4, characterized in that: The step of bringing the preliminary sound source estimated position into a near-field model for iterative optimization to obtain the near-field sound source estimated position includes: Calculating a near-field steering vector based on a previously calculated estimated sound source position, iteratively calculating the first sound source azimuth and the second sound source azimuth using the near-field steering vector, and calculating the position of a straight line intersection of the first sound source azimuth and the second sound source azimuth after iteration based on a geometric relationship, until a distance between the currently calculated straight line intersection position and the previously calculated straight line intersection position is less than a preset threshold; When the distance between the intersection of the straight lines calculated this time and the intersection of the straight lines calculated last time is less than a preset threshold, the intersection of the straight lines calculated this time is used as the estimated position of the near-field sound source; When the near-field steering vector is calculated for the first time, the near-field steering vector is calculated according to the preliminary sound source estimated position.

6. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 5, characterized in that: The calculating of the near-field steering vector according to the sound source estimated position calculated last time includes: According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the first microphone array in the near-field model: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone array microphone positions, represents the total number of microphones in the first microphone array; According to the following formula, the estimated sound source position calculated last time is substituted into the near-field steering vector of the sound information collected by the second microphone array in the near-field model: ; in, Represents the second microphone array microphone positions, represents the total number of microphones in the second microphone array.

7. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 6, characterized in that: Iteratively calculating the first sound source azimuth angle and the second sound source azimuth angle using the near-field steering vector includes: The first sound source azimuth is calculated based on the near-field steering vector of the sound information collected by the first microphone array through the beamforming estimation method or the generalized cross-correlation delay estimation method, and the second sound source azimuth is calculated based on the near-field steering vector of the sound information collected by the second microphone array.

8. The high-precision sound source localization method based on a two-dimensional microphone array according to claim 1, characterized in that: The near-field sound source estimated position is further optimized by using a full-array near-field positioning optimization algorithm to obtain a final sound source estimated position, including: The near-field sound source estimated position is further optimized according to the following cost function optimization to obtain the final sound source estimated position: ; in, The frequency at which the amplitude of the sound signal generated by the sound source is maximum after Fourier decomposition. represents the speed of sound, Represents the first microphone positions, represents the total number of microphones in the two-dimensional microphone array, represents the estimated position of the near-field sound source, Represents the final estimated sound source position.

9. A high-precision sound source localization system based on a two-dimensional microphone array, characterized in that: The method comprises a two-dimensional microphone array and a processor, wherein the two-dimensional microphone array is used to collect sound information in an environment, and the processor is used to execute the method according to any one of claims 1 to 8.

10. The high-precision sound source localization system based on a two-dimensional microphone array according to claim 9, characterized in that: The arrangement of the two-dimensional microphone array is at least one of the following: Linear array; circular array; Multi-helical arm array.

Citation Information

Patent Citations

  • Parallelized sound source positioning system based on embedded GPU system and method

    CN104535965A

  • Microphone array self-calibration sound source positioning system based on iterative optimization algorithm

    CN104898091A

  • Multi-robot cooperative 3D sound source identification and positioning method

    CN112379330A

  • Long-distance sound source positioning method and system based on microphone array

    CN112394324A

  • Video and audio conference equipment, terminal equipment, sound source positioning method and medium

    CN114095687A