Multi-target positioning and speech enhancement method based on microphone array
The proposed method using a four-element cross microphone array and neural networks effectively addresses multi-source localization and enhancement challenges by accurately determining source positions and enhancing audio signals in complex environments.
Patent Information
- Application Number
- CN202510398322.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The existing microphone array technology is difficult to achieve accurate positioning and effective voice enhancement in multi-sound source scenarios. Especially in multi-sound source environments, the positioning accuracy and stability of the existing technology are affected, and the signal processing efficiency is low, so it is impossible to make full use of multi-sound source information.
The quaternary cross microphone array is combined with a uniform linear array, and the relative delay time and arrival angle of the sound source are calculated through a generalized cross-correlation-weighting function and a broadband MUSIC algorithm. The sound source position is determined by combining the feedforward neural network model, and speech enhancement is performed through a fast independent component analysis method, and signal alignment is performed considering the sound source transmission delay.
It realizes that multiple sound source locations are accurately positioned and false targets are eliminated without prior information, improving the clarity and recognizability of speech signals, and improving signal processing efficiency in multi-sound source environments.
Smart Images

Figure CN120314871A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech signal processing, and particularly relates to a multi-target localization and speech enhancement method based on a microphone array. Background Art
[0002] A microphone array is a broadband acoustic signal processing device. With its characteristics of small volume, easy to carry and easy to hide, it is widely used in sound source localization, tracking and target acoustic signal enhancement in special environments. By processing acoustic signals through a microphone array, key information can be effectively captured. Existing microphone array localization technologies mainly focus on the localization of a single sound source, while the research on multi-source scenarios is relatively insufficient. Compared with single-source scenarios, the complexity of multi-source scenarios increases significantly, which is not only reflected in the increase in the number of sound sources, but also includes problems such as mutual interference between sound sources, complex sound field distribution, and signal overlap. Therefore, existing single-source localization technologies are difficult to be directly applied to multi-source environments, and their localization accuracy and stability are greatly affected in multi-source scenarios. At the same time, in multi-source environments, the difficulty of obtaining the acoustic signal information of interest also increases significantly, because background noise, reverberation effects, and interference from other sound sources will have a significant impact on the extraction of target signals. In addition, in existing microphone array signal processing technologies, a single signal model is usually adopted, and this model often seems powerless when processing multi-source signals, resulting in low signal processing efficiency and unable to fully explore and utilize the information potential in multi-source environments. Therefore, it is urgent to develop more advanced multi-source localization technologies and signal processing algorithms to cope with the challenges in complex multi-source scenarios.
[0003] In existing microphone array speech enhancement technologies, delay-and-sum beamforming and nulling techniques are implemented based on the spatial positions of known signal sources, which can effectively suppress environmental noise and interference signals, thereby improving the clarity and recognizability of target speech. However, when the noise is replaced with the same type of speech clutter, these technologies have poor effects. In addition, in multi-source environments, due to the large number of sound sources and signal overlap and other interferences, it is particularly crucial to accurately obtain the speech information of a specific person. Therefore, for such special environments, it is urgent to develop more advanced speech enhancement algorithms to improve the extraction and enhancement effects of specific person speech in multi-source scenarios and meet the actual application requirements in complex environments. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a multi-target localization and speech enhancement method based on a microphone array.
[0005] The technical solution of the present invention is as follows:
[0006] A multi-target localization method based on a microphone array, used for multi-source coordinate localization in a two-dimensional space, includes the following steps:
[0007] S1. Simultaneously receive acoustic signals using a quaternion cross microphone array and a uniform linear array, and perform sampling and discretization processing to obtain the first received signal and the second received signal of multiple sound sources respectively; the quaternion cross microphone array refers to arranging four microphone elements at the following four positions in a two-dimensional coordinate system: (-d, 0), (0, d), (d, 0), and (0, -d), where d is the first element spacing; the uniform linear array includes multiple microphone elements, and the first element is located at the coordinate origin, and the other elements are arranged in sequence along one direction of the coordinate axis with a second element spacing L;
[0008] S2. Select one element from the quaternion cross microphone array as the reference element, calculate the generalized cross-correlation-weighted function of other elements relative to this element according to the first received signal, and then obtain the relative delay time of different sound source signals relative to the reference element from the function curve;
[0009] Calculate the arrival angles of different sound source signals using the broadband MUSIC algorithm according to the second received signal;
[0010] S3. Classify the obtained different relative delay times to obtain all relative delay time combinations. The specific method is as follows: Define the number of sound sources as J. Then, any one generalized cross-correlation-weighted function curve obtained through S2 contains J peak points, and the positions of the peak points are the relative delay times of each sound source. Select 1 peak point from the J peak points of each generalized cross-correlation-weighted function curve for sequential combination to obtain 1 relative delay time combination. After all J combinations, all relative delay time combinations are obtained; Pass each relative delay time combination through a feedforward neural network model to obtain all possible positions of the sound sources, and calculate the spatial angles corresponding to all positions. The feedforward neural network model is trained by the true positions of the sound sources and the true relative delay time combinations;
[0011] In this step, after obtaining three generalized cross-correlation-weighted function curves of other elements relative to the reference element based on step S2, when the number of sound sources is J, any one generalized cross-correlation-weighted function curve contains J peak points, and the positions of the peak points are the relative delay times of each sound source. To obtain the position of a certain sound source, it is necessary to select the relative delay time corresponding to this sound source from the first to the last generalized cross-correlation-weighted function curves for sequential combination and then calculate. However, at this time, it is impossible to determine which sound source the relative delay time specifically belongs to. To cover all possible positions of the sound sources, it is necessary to select one relative delay time from the first to the last generalized cross-correlation-weighted function curves for sequential combination to obtain all relative delay time combinations;
[0012] S4. Match the spatial angle obtained in S3 with the arrival angle obtained in S2, so as to obtain the spatial position of the sound source target according to the matching result.
[0013] Further, the quaternion cross microphone array uses a real number signal model, and the uniform linear array uses a complex number signal model.
[0014] A speech enhancement method based on a microphone array and a multi-target localization method based on a microphone array. After obtaining the spatial position of the sound source target, calculate the transmission delays of different sound sources. Based on the transmission delays, align the signals received by the quaternion cross array to obtain an enhanced speech signal.
[0015] The beneficial effects of the present invention are as follows: The method of the present invention does not require prior information and can achieve the exclusion of false targets and multi-target accurate positioning only by relying on two microphone arrays with different structures to work simultaneously to process the received sound signal information. Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the composite microphone array structure used in the present invention.
[0017] Figure 2 It is a flowchart of multi-target localization of the microphone array.
[0018] Figure 3 It is a schematic diagram of the multi-source signal model received by the microphone array in the embodiment of the present invention.
[0019] Figure 4 It is a schematic diagram of the generalized cross-correlation-weighted curve calculated in the embodiment of the present invention.
[0020] Figure 5 It is a schematic diagram of all classification combinations of relative delay times in the embodiment of the present invention.
[0021] Figure 6 It is a schematic diagram of the feedforward neural network model.
[0022] Figure 7 It is a schematic diagram of the specific architecture of the feedforward neural network used in the embodiment of the present invention.
[0023] Figure 8 It is a schematic diagram of the training situation of the feedforward neural network in the embodiment of the present invention.
[0024] Figure 9 It is a schematic diagram comparing the positioning accuracy of the feedforward neural network of the present invention with the accuracy of the traditional geometric positioning algorithm.
[0025] Figure 10 It is a schematic diagram of the spatial spectrum curve calculated by the broadband MUSIC algorithm in the embodiment of the present invention.
[0026] Figure 11This is a schematic diagram of the working principle of the multi-channel speech enhancement algorithm proposed by the present invention.
[0027] Figure 12 This is a schematic diagram of the working performance curve of the fast independent component analysis method in the embodiment of the present invention. Detailed implementation manners
[0028] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0029] The present invention provides a composite microphone array structure in which a four-element cross microphone array is embedded in a uniform linear microphone array. As Figure 1 shown, the four microphone elements in the four-element cross microphone array are respectively located at the following positions in the two-dimensional coordinate system: (-d, 0), (0, d), (d, 0), and (0, -d), where d is a preset value of the element spacing; the uniform linear array includes a plurality of microphone elements, the first element is located at the coordinate origin, and the other elements are arranged in sequence with a spacing of L. The element spacing d of the four-element cross array, the element spacing L of the uniform linear array, and the number of elements in the uniform linear array can be reasonably selected according to the actual problem.
[0030] The four-element cross microphone array and the uniform linear array receive acoustic signals simultaneously. The four-element cross microphone array uses a real signal model, and the uniform linear array uses a complex signal model. The reason for using different signal models for the two different arrays is that different arrays need to achieve different functions respectively, but the real signal model and the complex signal model can be essentially converted to each other. Suppose there are M sound sources and N microphone elements, and let y n (t) be the acoustic signal received by the nth microphone element, a nm be the amplitude attenuation of the signal emitted by the mth sound source reaching the nth microphone element, s m (t) be the initial signal of the mth sound source, t nm be the transmission delay of the signal emitted by the mth sound source reaching the nth microphone element, be the noise. Then, for the real signal model, there is: For the complex signal model, there is: In the complex signal model, for the microphone broadband speech signal, f m represents all frequency components of the mth sound source. At this time, s m (t) also contains a complex signal with multiple frequency components. Actually, it means that the sum is calculated after separately calculating according to the relationship of this formula for all different frequency components. The acoustic signals received by different elements are used for subsequent calculations after sampling and discretization processing. The sampling rate can be reasonably set according to the actual problem requirements.
[0031] As a theoretical premise, the sound source signals need to meet the following conditions: when the number of sound source signals is 1, the sound source is uncorrelated with the external noise, and the noises received by different array elements are also uncorrelated with each other; when the number of sound source signals is greater than 1, any two sound sources are uncorrelated with each other, any sound source is uncorrelated with any noise, and at the same time, the noises received by different array elements are also uncorrelated with each other.
[0032] After the four-element cross array receives the acoustic signal and performs sampling and discretization processing, a certain array element is selected as the reference array element, and the Generalized Cross-Correlation with Phase Transform (GCC-PHAT) function of other array elements relative to this array element is calculated. The relative delay time of different sound source signals relative to the reference array element is obtained from the function curve, and the reference array element can be arbitrarily selected according to the actual problem; after the uniform linear array receives the acoustic signal and performs sampling and discretization processing, the broadband MUSIC algorithm is used to calculate the arrival angles of different sound source signals; the acoustic signals received by the four-element cross array and the acoustic signals received by the uniform linear array are not filtered or truncated to avoid losing important information.
[0033] After obtaining the relative delay time, it is actually impossible to directly determine the sound source attribution corresponding to different relative delay times. Therefore, the relative delay times from different function curves are classified, and the classification results consider all possible situations. All the classified relative delay time combinations are sent into the feedforward neural network model to obtain all possible positions of the sound sources, and the spatial angles corresponding to these positions are calculated and compared with the angle measurement results of the uniform linear array. At this time, the accurate spatial positions of multiple sound source targets can be obtained.
[0034] The present invention also provides a multi-channel speech enhancement method. Since the distances from different sound sources to different array elements are inconsistent, there is a time delay in the signals received by each array element. Based on the accurate multi-source localization results, the signal from a certain sound source in the signals received by the four-element cross array is aligned. The aligned received data matrix uses the Fast Independent Component Analysis (FASTICA) blind source separation algorithm to achieve speech separation and enhancement. Compared with the existing theory, the innovation of the present invention lies in considering the sound source transmission delay condition and performing alignment processing.
[0035] Embodiment
[0036] As Figure 2 shown, the signals emitted by the sound source will be received by Figure 1The composite microphone array shown is used for reception. After sampling and discretizing the received continuous analog signal, it is used for subsequent calculation processes. The signal processing module is the core of the multi-target localization method for the microphone array. It can be divided into two parts: the signal processing module for the four-element cross array to receive signals and the signal processing module for the uniform linear array to receive signals. These two signal processing modules work simultaneously and obtain conclusions, and finally obtain the accurate spatial positions of multiple sound sources. As an example, the sampling rate is set to 48KHZ, the two-dimensional spatial positions of the two sound sources are set to (5.49m, 5.58m) and (-2.66m, 0.56m) respectively, and Figure 2 The process shown is used for experimental verification. The corresponding scenario of the microphone array receiving signals in this example is as Figure 3 shown. In this example, the spacing d of the four-element cross array is set to 1m, the spacing L of the uniform linear array is set to 0.08m, the number of array elements of the uniform linear array is set to 4, and environmental noise is not considered at the same time.
[0037] For the signal processing module of the four-element cross array to receive signals, when each element of the four-element cross array receives the sound signal and samples and discretizes it, select Figure 3 the element (1) shown as the reference element, and calculate the generalized cross-correlation-weighted function between the signals received by other elements and the signal received by element (1). For the above example, three groups of generalized cross-correlation-weighted curves as Figure 4 shown can be obtained.
[0038] Refer to Figure 4 , different generalized cross-correlation function curves characterize the signal reception situations of different elements relative to the reference element. The position of the peak point of the generalized cross-correlation function curve is the relative delay time of different sound sources arriving at this element relative to arriving at the reference element. However, in different curves, the sound source attribution of the peak point cannot be uniquely determined. Therefore, the peak points from different curves are classified, as Figure 5 shown.
[0039] As Figure 5 shown, to calculate the spatial position coordinates of the signal source, it is necessary to obtain and classify the relative delay data from the three generalized cross-correlation function curves. The classification result is as Figure 5 , that is, select a peak point from each of the three curves for combination. At this time, it is necessary to calculate the possible positions of the sound sources according to the possible combinations. The present invention uses a feedforward neural network to calculate the possible positions of the sound sources. 15,000 groups of combinations of the true positions of the sound sources and the true relative delays are generated before for training the feedforward neural network. The structure model of the feedforward neural network, the specific parameters of the feedforward neural network used in the present invention, the training situation of the feedforward neural network of the present invention, and the accuracy comparison between the feedforward neural network and the traditional TDOA geometric positioning algorithm are respectively as Figure 6 , Figure 7 , Figure 8 , Figure 9As shown. So far, the four-element cross-array received signal module has completed the calculation of all possible positions of the sound sources. At this time, there are false targets, and it cannot be completed only by the four-element cross-array received signal module.
[0040] Refer back to Figure 3 , for the uniform linear array received signal processing module, first, the sound signal is received through the uniform linear array. It should be understood that the sound signal received by the uniform linear array and the sound signal received by the four-element cross-array work synchronously. Then, the sound signal received by the uniform linear array is sampled and discretized. For the signals received by each array element, finally, the broadband MUSIC algorithm is used to solve the spatial spectrum function curve. As Figure 10 shown, the position of the peak point of the spatial spectrum curve is the true spatial angle of each sound source. The spatial angle of the sound source corresponds to each different sound source without ambiguity, and the false targets can be excluded by using the spatial angle of the sound source as verification information.
[0041] Refer back to Figure 5 and Figure 10 , use Figure 10 the spatial angle information of the sound source shown as the verification condition and compare it with the positions of each different sound source shown in Figure 5 . Finally, it is found that the positions (5.08m, 5.23m) and (-2.72m, 0.5m) match the angle verification information. At this time, the multi-target localization module has achieved the exclusion of false targets and the determination of the spatial positions of the signal sources.
[0042] The present invention also proposes a multi-channel speech enhancement method for a microphone array. This method is based on the blind source separation theory of fast independent component analysis. Compared with the existing methods using blind source separation algorithms to achieve microphone array speech enhancement, the method proposed by the present invention fully considers the energy loss and transmission delay problems in the sound signal transmission process. Compared with the existing applications, the scope of investigation of the present invention is wider. The process of using the fast independent component analysis method to achieve speech enhancement is as Figure 11 .
[0043] Refer to Figure 11 . The fast independent component analysis algorithm currently has a complete theoretical system. Therefore, in this process, the core is to calculate the transmission delay and align the signals. For the four-element cross-array received signal, it can be expressed as: When the positions of different sound sources are obtained through the multi-target localization method, the transmission delay t nm can be obtained. At this time, although there is an error from the true delay, the estimated delay can still be calculated. Considering aligning the first sound source, after the alignment process, the four-element cross-array received signal can be expressed as: At this time, the aligned signal compensates for the loss of independence caused by transmission delay to a certain extent. It is possible to consider aligning multiple sound sources separately and sending the aligned received signal matrices into the fast independent component analysis algorithm for speech separation and enhancement respectively.
[0044] In this example, the true positions of the two sound sources in the two-dimensional space are set as (5.49m, 5.58m) and (-2.66m, 0.56m) respectively. The positions obtained by the positioning method are (5.08m, 5.23m) and (-2.72m, 0.5m) respectively. This example is consistent with the positioning example. After solving the estimated delay and performing alignment processing on the received signals of the four-element cross array in this example according to the above description, they are sent into the fast independent component analysis algorithm, and the algorithm performance graph is as Figure 12 , STOI is a parameter used to analyze the performance of speech enhancement. It analyzes the performance of the speech enhancement method by solving the frequency-domain correlation between the enhanced speech and the original speech. Therefore, the higher the STOI value, the better the speech enhancement performance. And from Figure 12 it can be seen that the performance of the method proposed in the present invention is relatively excellent.
Claims
1. A multi-target localization method based on a microphone array for multi-sound source coordinate localization in a two-dimensional space, characterized in that Including the following steps: S1. Use a four - element cross - microphone array and a uniform linear array to simultaneously receive acoustic signals and perform sampling and discretization processing, respectively obtaining the first received signal and the second received signal of multiple sound sources; the four - element cross - microphone array refers to setting four microphone elements at the following four positions in a two - dimensional coordinate system: (-d, 0), (0, d), (d, 0), and (0, -d), where d is the first element spacing; the uniform linear array includes multiple microphone elements, and the first element is located at the coordinate origin, and other elements are arranged in sequence along one direction of the coordinate axis with a second element spacing L; S2. Select one element from the four - element cross - microphone array as the reference element, calculate the generalized cross - correlation - weighted function of other elements relative to this element according to the first received signal, and then obtain the relative delay time of different sound source signals relative to the reference element from the function curve; According to the second received signal, use the broadband MUSIC algorithm to calculate the arrival angles of different sound source signals; S3. Classify the obtained different relative delay times to obtain all relative delay time combinations. The specific method is as follows: Define the number of sound sources as J. Then, any generalized cross - correlation - weighted function curve obtained through S2 contains J peak points, and the peak point positions are the relative delay times of each sound source. Select 1 peak point from the J peak points of each generalized cross - correlation - weighted function curve for sequential combination, so as to obtain 1 relative delay time combination. After all J times of combination, all relative delay time combinations are obtained; Pass each relative delay time combination through a feed - forward neural network model to obtain all possible positions of the sound sources and calculate the spatial angles corresponding to all positions. The feed - forward neural network model is trained by the true positions of the sound sources and the true relative delay time combination; S4. Match the spatial angles obtained in S3 with the arrival angles obtained in S2, so as to obtain the target spatial position of the sound source according to the matching result.
2. The multi-target positioning method based on a microphone array according to claim 1, wherein The four - element cross - microphone array uses a real - number signal model, and the uniform linear array uses a complex - number signal model.
3. A voice enhancement method based on a microphone array, based on the multi-target localization method based on a microphone array described in claim 1, characterized in that, After obtaining the target spatial position of the sound source, calculate the transmission delays of different sound sources. Based on the transmission delays, align the received signals of the four - element cross - array to obtain enhanced speech signals.
Citation Information
Cited By
Control system of controllable anti-recording device
CN120582744A
A control system for a controllable anti-recording device
CN120582744B