Sound source positioning and identification method based on four-channel audio signals of vector microphone
By using a sound source localization and recognition method based on four-channel audio signals from a vector microphone, combined with an improved CRNN network, the problems of large localization error and high computational cost of traditional sound source localization technology in complex environments are solved, and high-precision sound source localization and recognition in three-dimensional environments are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
AI Technical Summary
Existing sound source localization technologies have large positioning angle errors when faced with complex speech signals, making it difficult to meet practical needs. Furthermore, traditional microphone arrays have high hardware complexity, large data processing volume, high computational cost, and limited adaptability to various scenarios.
A sound source localization and recognition method based on four-channel audio signals from a vector microphone is adopted. By simulating the four-channel audio signals output by the vector microphone and combining them with an improved CRNN network, short-time Fourier transform and normalization processing are performed to construct a sample dataset and train a sound source localization/recognition model to achieve high-precision sound source localization and recognition in complex three-dimensional environments.
It reduces hardware layout complexity and computational costs, and achieves high-precision sound source localization of single-frequency signals and single-voice signals and recognition of multiple voice signals in complex three-dimensional environments, making it suitable for real-world scenarios.
Smart Images

Figure CN122372895A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization technology, and in particular to a sound source localization and identification method based on four-channel audio signals from a vector microphone. Background Technology
[0002] Currently, sound source localization technology is relatively mature. Among them, the MUSIC algorithm and beamforming algorithm have high localization accuracy for single-frequency signals and pulse signals, and can also show high stability in noisy environments. Neural network models, as a popular model in machine learning, are also applied to sound source localization technology. They mainly focus on high-precision localization of single-frequency and pulse signals in more complex environments, or two-dimensional localization of single speech signals, that is, outputting only the azimuth angle result, which can have high accuracy and strong stability.
[0003] However, current technologies still have unresolved issues: when faced with more complex speech signals, the positioning angle error is significant, failing to meet practical needs. Speech localization and even multi-speech recognition in complex three-dimensional environments remain unsolved problems in this research field. Furthermore, existing traditional algorithms and neural network models process data collected by microphone arrays. While these arrays can collect sound signals from multiple locations and obtain substantial sound information, they require processing large amounts of data, incur high computational costs, and are bulky and difficult to assemble, limiting their applicability in practical applications. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a sound source localization and recognition method based on four-channel audio signals from a vector microphone. By simulating the four-channel audio output of a vector microphone to obtain raw audio data, the invention reduces computational costs and hardware layout complexity, while achieving high-precision sound source localization and recognition of single-frequency signals and single-voice signals in complex three-dimensional environments, as well as recognition of multiple voice signals.
[0005] Technical Solution: To achieve the above objectives, the sound source localization and identification method based on four-channel audio signals from a vector microphone, as described in this invention, includes the following steps:
[0006] S1. Generate the original audio signal and add reverb, noise and four-channel encapsulation processing to generate a four-channel audio signal output by an analog vector microphone. The original audio signal is a single-frequency signal, a single speech signal or a multi-speech signal.
[0007] S2. Perform Short Time Fourier Transform (STFT) and normalization on all frame signals of all channels in the four-channel audio signal to obtain a three-dimensional feature array. Labels are added to construct a sample dataset, where a three-dimensional feature array... It is generated based on a single-frequency signal or a single speech signal. The tag includes the state information of whether the target sound source exists in each frame, as well as the azimuth and elevation angles of the target sound source. If the three-dimensional feature array... It is generated based on multiple speech signals, and the tag is the status information of whether each target sound source exists in each frame;
[0008] S3. Construct a sound source localization / identification model based on an improved CRNN network and train it using a sample dataset;
[0009] S4. Based on the real four-channel audio signal collected and output by the vector microphone, after processing by S2, it is input into the trained sound source localization / recognition model for prediction: if the vector microphone collects a single-frequency signal or a single speech signal, the prediction result is the state information, azimuth angle and elevation angle of the target sound source; if the vector microphone collects multiple speech signals, the prediction result is the state information of each target sound source.
[0010] Preferably, the original audio signal generated in S1 , =1, 2, 3, including:
[0011] when When =1, Indicates a single-frequency signal: , The frequency of the signal;
[0012] when When =2, Indicates a single speech signal: , This indicates the total number of frames in the single audio signal. Indicates the first The amplitude value of the frame;
[0013] when When =3, Indicates multiple speech signals: , This indicates the number of individual speech signals that make up a multi-speech signal. Indicates an index. Indicates the first A single voice signal.
[0014] Preferably, the original audio signal after adding reverb in S1 is represented as follows:
[0015] ,
[0016] Among them, when When =1, for Sampling is performed first, then reverberation is added; The room impulse response is represented as:
[0017] ,
[0018] In the formula, They represent respectively for The order of the coordinate axis symmetric treatment, and , Indicates the maximum number of sound reflections. , For room volume, This represents the total surface area of the room. Speed of sound; For reverberation time, , The average sound absorption coefficient of the wall surface; The distance from the virtual sound source to the microphone; and Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of the two walls along the axial direction; , They are respectively , of Power of 1 They are respectively of Power of 1 They are respectively of Power Item is time The Dirac function indicates that the signal from the virtual sound source only changes during the delay time. It was collected.
[0019] Preferably, the four-channel audio signal in S1 is represented as follows:
[0020] ,
[0021] In the formula, Indicates noise. , The signal after reverberation processing power, Signal-to-noise ratio; , , , These represent the signals of the four channels respectively. Indicates the frame index. ;
[0022] The four-channel audio signal is simplified as follows: ,in, , Corresponding to , , , The signal of the channel.
[0023] Preferably, in S2, a Short-Time Fourier Transform (STFT) is performed on all frame signals of all channels in the four-channel audio signal to obtain a three-dimensional feature array P, represented as:
[0024] ,
[0025] In the formula, Indicates channel signal The Frame, First The feature vector list on each channel is obtained by first processing the signal of the c-th channel in the t-th frame. STFT processing is performed to obtain complex spectrum values. Further extraction of complex spectrum values amplitude spectrum Phase spectrum After that, with the first Frame, First The amplitude and phase spectra of all frequency points on each channel are arranged in sequence. The amplitude and phase spectra of all frequency points are then vertically stitched together to obtain the following:
[0026] ,
[0027]
[0028] ,
[0029] In the formula, For frequency index, =0~F-1, where F is the total number of frequency points; Indicates frame shift, For the length of the window, This is the index for the sample offset within the window. For the signal at the sample point The amplitude value at that point, The term is a window function;
[0030] Feature vector list Represented as:
[0031] Abbreviated as: , , represents the index of a feature vector within the feature vector list.
[0032] Preferably, the three-dimensional feature array P is obtained by normalizing the three-dimensional feature array. , represented as:
[0033] ,
[0034] To The result is obtained by normalizing all elements of all feature vectors in the dataset, and is represented as: ;
[0035] in, ,
[0036] In the formula, For list The Middle Frame, First Normalized feature vectors of each channel, , , where represent the mean vector and standard deviation vector of the eigenvectors, respectively.
[0037] Preferably, a three-dimensional feature array is added in S2. The labeling method is as follows: first, process the three-dimensional feature array... All feature vectors from the same channel are concatenated in time series order to obtain four forms. The four two-dimensional arrays are then concatenated along the channel dimension to obtain the form: A three-dimensional array, denoted as ;
[0038] If a three-dimensional array If the signal is generated based on a single-frequency signal or a single-voice signal, the added tag will be in the form of: A three-dimensional array, wherein, in the form of The two-dimensional array represents the state information of the audio signal. If the sound source does not emit sound in a certain time frame of a single-frequency signal or single-speech signal, the state label of that time frame is set to 0; otherwise, it is set to 1. The format is... The two-dimensional array represents the angular information of the audio signal, including the azimuth angle of the sound source relative to the microphone for each time frame. and pitch angle ;
[0039] If a three-dimensional array Based on multiple speech signals, the tagging format is as follows: A two-dimensional array represents the state information of the audio signal. This is obtained by acquiring the state labels of all individual speech signals that make up the multi-speech signal. The form is The two-dimensional array is obtained by concatenating it along the second dimension.
[0040] Preferably, the sound source localization / recognition model includes a CNN module, an RNN module, and an FC module connected in sequence, wherein the number of convolutional kernels in the CNN module is set to be [number missing]. The pooling factor list is as follows The RNN module sets a list of bidirectional gated recurrent unit neurons. The list of fully connected units is set in the FC module. It includes a Time Detection Component (SED) and a Direction of Arrival (DOA). If the sound source localization / recognition model performs a single-frequency signal or single-speech signal localization and recognition task, the RNN module output result is simultaneously input to the fully connected layer's Time Detection Component (SED) and Direction of Arrival (DOA), and the final output format is... A three-dimensional array, where the form is A two-dimensional array represents the state information of the audio signal, in the form of... The two-dimensional array represents the angle information of the audio signal; if performing a multi-speech signal recognition task, the RNN module outputs the result input to the time detection part (SED), and the final output format is... A two-dimensional array, representing the state information of the audio signal.
[0041] Preferably, in the DOA section, the output of the RNN module is first input into a first fully connected layer with a time distribution for loop processing. After looping, the result is input into another second fully connected layer, and the output is specified to be in the form of a two-dimensional array. , in the form of The two-dimensional array is input into the hyperbolic tangent function tanh for activation, and the DOA output is further transformed into the following form: Transformation of a two-dimensional array into The transformation is based on the unit distance assumption, treating the output vector as three-dimensional coordinates, and obtaining the azimuth angle θ and elevation angle φ by solving the following system of equations:
[0042] ;
[0043] In the SED section, the output of the RNN module is first fed into the first fully connected layer of the time distribution for iterative processing. After the loop, the result is fed into another second fully connected layer, and the output is specified to be in the form of a two-dimensional array. , in the form of The two-dimensional array is input into the sigmoid function for activation, and the final output two-dimensional array is the result of SED.
[0044] Preferably, the sound source localization / recognition model training method is as follows: A loss function and optimization strategy are selected based on the specific task: In localization and recognition tasks for single-frequency signals and single-speech signals, joint loss is used for optimization. Specifically, the DOA output is for a regression task, using mean absolute error or mean squared error as the loss function to measure the deviation between the predicted sound source angle and the true label; the SED output is for a binary classification task, using binary cross-entropy as the loss function to measure the deviation between the sound source state and the true label. The total loss function is a weighted sum of the two loss functions according to set weights, calculating the weighted error between the model output and the corresponding label. In the recognition task for multiple speech signals, binary cross-entropy is used as the loss function to calculate the error between the model output and the label. After completing one round of forward propagation, backpropagation is performed based on the calculated error value, and all trainable parameters in the network are updated through the optimizer until a set stopping condition is reached, completing the model training.
[0045] Beneficial effects: The present invention has the following advantages: The present invention uses a four-channel audio signal output from an analog vector microphone, which greatly reduces the complexity and size of the hardware layout compared to traditional microphone arrays. It also reduces the amount of data processing, effectively reduces the computational cost, and is easier to promote and apply in real-world scenarios. By training an improved CRNN network model in combination with a three-dimensional feature array obtained from four-channel audio signal processing, it can not only achieve high-precision sound source localization and recognition of single-frequency signals and single speech signals in complex three-dimensional environments, but also effectively recognize multiple speech signals. Attached Figure Description
[0046] Figure 1 This is a flowchart of the method of the present invention;
[0047] Figure 2 Flowchart for generating a four-channel audio signal from an analog vector microphone output;
[0048] Figure 3 This is a schematic diagram showing the relationship between the vector microphone and the sound source.
[0049] Figure 4 This is a schematic diagram of a vector microphone structure;
[0050] Figure 5 Flowchart for constructing the sample dataset;
[0051] Figure 6 This is a diagram of the sound source localization / identification model architecture. Detailed Implementation
[0052] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.
[0053] like Figure 1As shown, the sound source localization and identification method based on four-channel audio signals from a vector microphone according to the present invention includes the following:
[0054] S1. Generate the original audio signal The signal is then processed by adding reverb, noise, and four-channel encapsulation to generate a four-channel audio signal that mimics the output of a vector microphone, such as... Figure 2 As shown.
[0055] S101, Original Audio Signal , =1, 2, 3, when When =1, Indicates a single-frequency signal, when When =2, Represents a single speech signal, when When =3, It indicates multiple speech signals.
[0056] A single-frequency signal is represented as: , The frequency of the signal is a sine wave. The signal will be sampled to obtain a discrete time series, which will be stored in array form.
[0057] A single speech signal is represented as: A single speech signal is obtained by parsing clean speech, resulting in a sequence with 1 row and the number of columns equal to the length of the time series. Array; It also indicates the total number of frames in the single voice signal. Indicates the first The amplitude value of the frame.
[0058] Multiple speech signals are represented as follows: Multi-speech signals are composed of the linear superposition of multiple sets of single-speech signals. This indicates the number of individual speech signals that make up a multi-speech signal. Indicates an index. Indicates the first A single voice signal.
[0059] S102. To more realistically simulate audio signals collected in actual environments (especially for enclosed spaces such as conference rooms and living rooms), it is necessary to process the original audio signals. Perform reverberation processing.
[0060] In reverberant scenarios, sound reflects multiple times through walls, ceilings, and floors before reaching the microphone. The "Mirror Source Method" (ISM), which uses virtual sound sources to equivalently replace these reflected sounds, effectively simulates this situation. A spatial Cartesian coordinate system is established using a corner of the room as an example. By aligning the real sound source with the length, width, and height of each wall in the room and performing symmetrical processing with respect to each wall, we obtain first-order virtual sources. These first-order virtual sources are then symmetrically processed with respect to other walls to obtain second-order virtual sources, i.e., secondary reflections. This process can be repeated to generate an infinite number of virtual sources, corresponding to an infinite number of reflections.
[0061] Voice signal and The array has the same format and dimensions, and is a single-frequency signal. The initial sine wave form is transformed into a form similar to the one obtained by setting the sampling rate. and The array format and dimensions are consistent, therefore single-frequency signals and voice signals and The above methods can be used for reverberation processing in all cases.
[0062] The reverberated audio signal is represented as follows: ,
[0063] In the formula, For room impulse response, :
[0064]
[0065] In the formula, They represent respectively for The order of the coordinate axis symmetric treatment, and To constrain The value is used to truncate infinite terms. Indicates the maximum number of sound reflections. , For reverberation time, For room volume, This represents the total surface area of the room. The speed of sound is [value missing]. ;
[0066] in addition, , The average sound absorption coefficient of the wall surface, i.e., the wall surface sound absorption coefficient. The surface-weighted average is expressed as:
[0067]
[0068] In the formula, For room number The area of each wall surface. For the first The sound absorption coefficient of each wall surface;
[0069] The order of axisymmetry for each set of walls in three orthogonal directions This represents a virtual sound source. This represents the distance from the virtual sound source to the microphone;
[0070] and Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of both walls along the axial direction are related to the wall material, and the first... The reflectance value of the wall surface is ;
[0071] , They are respectively , of Power of 1 They are respectively of Power of 1 They are respectively of Power
[0072] Item is time The Dirac function has the following specific form:
[0073]
[0074] This function indicates that the signal from the virtual sound source only occurs during the delay time. The pulse is collected, that is, a pulse is generated over the delay time.
[0075] During reverberation processing, the room size, sound source location, and microphone location can be freely set. Considering practical application scenarios, the room size can be set to... The microphone is located in the room. The sound source location is randomly generated within a certain range to improve the algorithm's generalization ability under different reverberation environments, while the sound source location is randomly generated in space.
[0076] In practical applications, according to the formula That is, the arc length is the product of the central angle corresponding to the arc and the radius. Therefore, as... Figure 3As shown, when the angular error between the line connecting the actual sound source location and the microphone and the line connecting the predicted sound source location and the microphone is constant, the greater the distance between the sound source location and the microphone, the greater the positional deviation between the actual sound source and the predicted sound source. Therefore, the sound source localization effect is best when the microphone is located in the center of space.
[0077] S103, Using the formula The amplitude of the Gaussian white noise is obtained, where The signal after reverberation processing power, The signal-to-noise ratio is used to generate a signal-to-noise ratio. Gaussian white noise with the same time series length is also in the form of a one-dimensional array.
[0078] S104, Signal Noise and four-channel encapsulation processing are added to generate a four-channel audio signal for analog vector microphone output.
[0079] like Figure 4 As shown, a vector microphone is an integrated device based on a thermal convection acoustic particle velocity sensor and a sound pressure sensor. It can collect acoustic velocity information in three orthogonal directions (x, y, z directions) and sound pressure information at its location in three-dimensional space, forming a four-channel audio signal (x, y, z acoustic velocity information and total sound pressure information). This invention simulates, processes, and analyzes this four-channel information to achieve high-precision sound source localization and speech recognition in more complex environments. Furthermore, vector microphones are small in size, easy to package, and more conducive to practical deployment.
[0080] When a vector microphone collects signals, the relationship between its four channel signals and the sound source location can be categorized as a steering vector.
[0081]
[0082] In the formula, These are the azimuth angles of the sound source relative to the vector microphone. Pitch angle, in generating the four-channel audio signal from the analog vector microphone output, These are the azimuth and elevation angles of the virtual signal source relative to the virtual acquisition point, respectively, used to simulate the spatial position of the sound source relative to the vector microphone.
[0083] The generated four-channel audio signal is represented as follows:
[0084]
[0085] The four-channel audio signal has a line number of A two-dimensional array with 4 columns It can be represented as , Corresponding to , , , The signals of the channel are all from Composed of frame audio signals A 3D time series array, Indicates the frame index.
[0086] The present invention utilizes the above method to directly generate batches of four-channel audio signals for use in simulating signals collected by a vector microphone.
[0087] S2. Perform Short Time Fourier Transform (STFT) and normalization on all frame signals of all channels in the four-channel audio signal to obtain a three-dimensional feature array. And add labels to construct a sample dataset, such as Figure 5 As shown.
[0088] S201. Perform a short-time Fourier transform (STFT) on each frame of each channel of the four-channel audio signal to obtain the amplitude spectrum and phase spectrum of the signal.
[0089] For the c-th channel t-th frame signal , The STFT processing method is as follows:
[0090]
[0091] in, Indicates the first frame( ), frequency index is ( =0~F-1, where F is the total number of frequency points), the... The complex spectral values of each channel are in the form of a two-dimensional array; This represents the frame shift, which is the number of samples between the start points of two adjacent frames; The window length is the number of sampling points contained in each selected frame. This is the in-window sample offset index, used to traverse every sampling point within the current frame; For the signal at the sample point The amplitude value at that point; The term is a window function, which weights the signal of each frame; it is in the form of a one-dimensional array. The expression is:
[0092] .
[0093] Further extraction amplitude spectrum Phase spectrum :
[0094]
[0095] With the first Frame, First The amplitude and phase spectra of all frequency points on each channel are sequentially arranged, and the amplitude and phase spectra of all frequency points are vertically spliced to obtain the channel signal. The Frame, First Feature vectors on each channel , represented as:
[0096] .
[0097] Let F be a one-dimensional array with 1 row and 2F columns. For ease of explanation, let F be a one-dimensional array with 1 row and 2F columns. Rewritten as , ;
[0098] The first All time frame signals of the first channel are processed as described above to obtain the second channel. List of feature vectors for each channel: [ , ,... ,..... ]
[0099] By performing the above processing on all time frame signals of the four channels respectively, a four-channel audio signal can be obtained. The three-dimensional feature array P is composed of The feature vectors form a row with 4 rows and a column with 12 columns. A three-dimensional array, specifically represented as follows:
[0100]
[0101] S202. Normalize array P by first calculating the mean vector and standard deviation vector of all eigenvectors in array P:
[0102]
[0103]
[0104] in, The mean vector is the first One element, The standard deviation vector is the first One element, Indicates the first Frame, First Feature vectors of each channel The Middle Each element.
[0105] For arrays Normalize all elements in all feature vectors to obtain the normalized result. ,in:
[0106]
[0107] In the formula, For array The Middle Frame, First Normalized eigenvectors of each channel.
[0108] Normalized Represented as:
[0109] ;
[0110] Normalized 3D feature array Represented as:
[0111] .
[0112] S203, Regarding the three-dimensional feature array Reorganize and add a three-dimensional feature array. Labels are used to construct a sample dataset.
[0113] First, concatenate all feature vectors of the same channel in time series order to obtain the form: A two-dimensional array, when subjected to the same operation on other channels, yields four forms. The four two-dimensional arrays are then concatenated along the channel dimension to obtain the final form. A three-dimensional array, denoted as .
[0114] If a three-dimensional array If the audio data is generated based on a single-frequency signal or a single-voice signal, the methods for adding tags include: first, adding the azimuth angle when generating the audio data based on S103. With pitch angle Extending this to the time series, we obtain the form: Angle labels of a two-dimensional array, and The value is the same in different time frames; if ( If, in a certain time frame (=1, 2), the sound source does not emit sound (meaning the microphone did not collect a signal in that time frame), then the status label for that time frame is set to 0; otherwise, it is set to 1 (meaning the sound source emitted sound in that time frame). The format is as follows: A two-dimensional array of state labels. The labels for each time frame are sorted by state and azimuth. and pitch angle Arranged in order, the result is in the form of A three-dimensional array, generated based on a single-frequency signal or a single speech signal. The tag.
[0115] If a three-dimensional array Based on multiple speech signals, the methods for adding tags include: In the process, acquire all single speech signals. , ... Status labels, the status label format for a single voice signal is as follows: A two-dimensional array, then a total of The form is A two-dimensional array, for The form is If the two-dimensional array is concatenated along its second dimension, the result is of the form: The state labels of the two-dimensional array are used as the basis for generating multiple speech signals. The tag.
[0116] Using the above method, this invention can construct a sample dataset including several labeled audio features and divide it into a training set, a validation set, and a test set.
[0117] S3. Construct a sound source localization / recognition model based on an improved CRNN network, including: a CNN module (Convolutional Neural Network), an RNN module (Recurrent Neural Network), and an FC module (Fully Connected Layer), such as... Figure 6 As shown.
[0118] (1) First, for the three-dimensional array Reorganize by swapping the order of the channel and time dimensions, transforming it into the form of A three-dimensional array, represented as Input the CNN module.
[0119] The number of convolutional kernels in a CNN module is The pooling factor list is as follows 3D array After being input into the CNN module, it first undergoes convolution processing through convolutional blocks, resulting in the form of... The three-dimensional array is then processed sequentially through batch normalization and non-linear activation before entering the max pooling layer, ultimately yielding the form of... A three-dimensional array, that is .
[0120] During the pooling layer processing, the pooling factor is selected from the pooling list based on the number of loops. Taking the first loop as an example... The data format after pooling is as follows: A three-dimensional matrix, in which for Gaussian floor function:
[0121]
[0122] Subsequently, the data is randomly deactivated to prevent overfitting during training, and the data format remains unchanged during this process.
[0123] The entire process from the convolutional block to random deactivation described above constitutes one loop. The output of the first loop is then used as the input for the second loop, and the CNN module is processed again. This process is repeated three times. During these three loops, the pooling layer is processed according to a list of pooling factors. The pooling factors are selected in the correct order, so the results of the three iterations are as follows: 3D array, 3D array, A three-dimensional array.
[0124] in, for The Gaussian floor function, and for Gaussian floor function:
[0125] .
[0126] (2) Set the list of bidirectional gated recurrent unit (GRU) neurons in the RNN module. The elements in the list represent the number of neurons in the bidirectional GRU during the three RNN loops. The RNN module receives the three-dimensional array output by the CNN module. Next, the first two dimensions of the array are arranged along the same dimension for merging, resulting in the form: Given a two-dimensional array, swap the two dimensions of the array to obtain the form: A two-dimensional array. (The second-dimensional array is...) The input is fed into a bidirectional GRU, and the output of the previous iteration is used as the input for the next iteration. This process is repeated three times. The number of neurons in the bidirectional GRU that process the array each time is adjusted from the list. The array is obtained based on the number of loops, i.e., the arrays after three bidirectional GRU processing steps. , , A two-dimensional array; the last output is in the form of The two-dimensional array is the output of the RNN module, represented as... .
[0127] (3) Set the list of fully connected units in the FC module The FC module structure is selected based on the task being performed.
[0128] (3.1) If performing the localization and recognition task of a single-frequency signal and a single speech signal, then Simultaneously input to the sound timing detection section (SED) and direction of arrival (DOA) section of the FC module, the sound timing detection section (SED) and direction of arrival section respectively... Nonlinear fitting was performed to obtain the results of SED and DOA.
[0129] In the DOA section, first... The input is fed into a fully connected layer (TimeDistributed) with a time distribution. The number of units in TimeDistributed is obtained from list C based on the loop count. Taking the first loop as an example, after passing through TimeDistributed, the output is in the form of... Two-dimensional array Then, it is randomly deactivated, while the data format remains unchanged. This TimeDistributed + random deactivation process is repeated three times, with the output of the previous loop used as the input for the next loop. The results obtained in the three loops will be in the following formats: Two-dimensional arrays Two-dimensional arrays A two-dimensional array. Then, the result obtained after looping is of the form... The two-dimensional array is input into another TimeDistributed array, and the output two-dimensional array is specified to be in the form of... Where 3 represents the output of each time frame, including three numbers corresponding to the x, y, and z coordinates. Finally, this form is... The two-dimensional array is input to the hyperbolic tangent function (tanh) for activation. In this layer, the array form remains unchanged, and the final output two-dimensional array is the result of DOA.
[0130] In the SED section, its structure is similar to DOA. The specific structural change compared to the DOA section is that the output format of the second TimeDistributed value in the DOA section is changed to... The final activation function is changed from tanh to sigmoid, while the rest of the structure remains unchanged. The final output is in the form of... The two-dimensional array is the output of SED.
[0131] The output format of DOA is as follows The two-dimensional array is processed, which represents the three-dimensional coordinates of each channel positioning in each time frame, assuming a distance of 1 unit length. The value can be obtained from the following system of equations. and :
[0132]
[0133] Then transform the DOA output into A two-dimensional array, corresponding to the label.
[0134] The output format of SED is: A two-dimensional array represents the signal state (0 or 1) in each time frame. Finally, the output of SED is concatenated with the output of DOA conversion in a third dimension to obtain the form: The three-dimensional array represents the final output of the sound source localization / recognition model in the task of locating and recognizing single-frequency and single-speech signals, and is expressed as follows: Generate based on single-frequency signals or single-voice signals The corresponding tag.
[0135] (3.2) If performing a multi-speech signal recognition task, then... Input only into the SED section of FC, and modify the output format of the second TimeDistributed in the SED section to... (The remaining structure of the SED part is the same as that of the SED in the localization and recognition tasks of single-frequency signals and single-speech signals.) After processing by the SED part, the result is in the form of... The two-dimensional array represents the final output of the sound source localization / recognition model in the multi-speech signal recognition task, and is related to the multi-speech signal generation... The corresponding tag is represented as .
[0136] During the training of the sound source localization / recognition model, an appropriate loss function and optimization strategy are selected based on the specific task. In the localization and recognition tasks of single-frequency signals and single-speech signals, the model uses a joint loss for optimization. The DOA output is for a regression task, using mean squared error (MSE) as the loss function to measure the deviation between the predicted angle (or coordinates) and the true label; the SED output is for a binary classification task, using binary cross-entropy as the loss function to measure the accuracy of frame-level sound source state detection. The total loss function is a weighted sum of the two loss functions with a weighted ratio of 100:1, calculating the weighted error between the model output and the corresponding label.
[0137] In the task of recognizing multiple speech signals, the model only outputs the SED part, and it is a multi-label classification task. Therefore, multi-label binary cross-entropy is used as the loss function to directly calculate the error between the model output and the label.
[0138] After completing a round of forward propagation in training, backpropagation is performed based on the calculated error value, and all trainable parameters in the network are updated through the optimizer (Adam), so that the model gradually fits the mapping relationship between the input features and labels in the training data.
[0139] After training is completed, based on the error reduction trend during training, the calculation of positioning accuracy or recognition accuracy, and the performance on the test set, corresponding analysis charts are generated (such as angle error distribution histograms drawn based on loss reduction curves, SED performance curves, and DOA error curves, etc.) to compare and improve model performance.
[0140] S4. Based on the vector microphone, real four-channel audio signals are acquired. After processing in S2, these signals are input into the trained sound source localization / recognition model. For single-frequency signal and single-speech signal localization and recognition tasks, the model output format is as follows: The prediction results, among which, This represents the number of frames in the actual four-channel audio signal. This indicates whether the target signal (0 or 1) exists in each time frame. This represents the azimuth and elevation angles of the sound source relative to the vector microphone at each time frame. For multi-speech signal recognition tasks, the model output format is as follows: The prediction results, among which, This represents the number of frames in the actual four-channel audio signal. This indicates the number of individual speech particles that make up a multi-speech signal.
[0141] This embodiment further provides a comparative example of using the method of the present invention and existing positioning and recognition algorithms to predict the location of single-frequency signals. Specifically, the generated original single-frequency signals are preprocessed and then input into the sound source localization / recognition model of the present invention and the existing localization and recognition model, respectively. The results show that the method of the present invention is comparable to the existing method in terms of positioning accuracy for single-frequency signals; while for single-speech audio with both noise and reverberation interference, the positioning error of the method of the present invention is reduced to 11°. In addition, the present invention achieves a speech recognition accuracy of 99.89% for single-audio audio, and can accurately detect the time frame of the sound source; in the case of three-source superimposed audio, the speech recognition accuracy can still reach 87%, and can effectively identify the number of sound sources and the sound sources in each time frame.
[0142] The above results demonstrate that the feature engineering processing method and network architecture proposed in this invention have higher accuracy advantages compared to existing localization algorithms, and can more effectively extract sound feature information to achieve high-precision localization. Unlike most current studies that target different types of sound detection, the audio identified in this invention is a scene of multiple voices superimposed. In this scene, the differences between different sound sources are small, and the recognition is difficult. Existing recognition algorithms are limited in such complex acoustic environments and cannot achieve accurate multi-source detection and localization.
Claims
1. A sound source localization and identification method based on four-channel audio signals from a vector microphone, characterized in that, Includes the following steps: S1. Generate the original audio signal and add reverb, noise and four-channel encapsulation processing to generate a four-channel audio signal output by an analog vector microphone. The original audio signal is a single-frequency signal, a single speech signal or a multi-speech signal. S2. Perform Short Time Fourier Transform (STFT) and normalization on all frame signals of all channels in the four-channel audio signal to obtain a three-dimensional feature array. Labels are added to construct a sample dataset, where a three-dimensional feature array... It is generated based on a single-frequency signal or a single speech signal. The tag includes the state information of whether the target sound source exists in each frame, as well as the azimuth and elevation angles of the target sound source. If the three-dimensional feature array... It is generated based on multiple speech signals, and the tag is the status information of whether each target sound source exists in each frame; S3. Construct a sound source localization / identification model based on an improved CRNN network and train it using a sample dataset; S4. Based on the real four-channel audio signal collected and output by the vector microphone, after processing by S2, it is input into the trained sound source localization / recognition model for prediction: if the vector microphone collects a single-frequency signal or a single speech signal, the prediction result is the state information, azimuth angle and elevation angle of the target sound source; if the vector microphone collects multiple speech signals, the prediction result is the state information of each target sound source.
2. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 1, characterized in that, The original audio signal generated in S1 , =1, 2, 3, including: when When =1, Indicates a single-frequency signal: , The frequency of the signal; when When =2, Indicates a single speech signal: , This indicates the total number of frames in the single audio signal. Indicates the first The amplitude value of the frame; when When =3, Indicates multiple speech signals: , This indicates the number of individual speech signals that make up a multi-speech signal. Indicates an index. Indicates the first A single voice signal.
3. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 2, characterized in that, The original audio signal after adding reverb in S1 is represented as follows: , Among them, when When =1, for Sampling is performed first, then reverberation is added; The room impulse response is represented as: , In the formula, They represent respectively for The order of the coordinate axis symmetric treatment, and , Indicates the maximum number of sound reflections. , For room volume, This represents the total surface area of the room. Speed of sound; For reverberation time, , The average sound absorption coefficient of the wall surface; The distance from the virtual sound source to the microphone; and Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of the two walls along the axial direction. Represents respectively in The reflection coefficients of the two walls along the axial direction; , They are respectively , of Power of 1 They are respectively of Power of 1 They are respectively of Power Item is time The Dirac function indicates that the signal from the virtual sound source only changes during the delay time. It was collected.
4. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 3, characterized in that, The four-channel audio signal described in S1 is represented as follows: , In the formula, Indicates noise. , The signal after reverberation processing power, Signal-to-noise ratio; , , , These represent the signals of the four channels respectively. Indicates the frame index. ; The four-channel audio signal is simplified as follows: ,in, , Corresponding to , , , The signal of the channel.
5. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 4, characterized in that, In S2, a Short-Time Fourier Transform (STFT) is performed on all frame signals of all channels in the four-channel audio signal to obtain a three-dimensional feature array P, represented as: , In the formula, Indicates channel signal The Frame, First The feature vector list on each channel is obtained by first processing the signal of the c-th channel in the t-th frame. Perform STFT processing to obtain complex spectrum values. Further extraction of complex spectrum values amplitude spectrum Phase spectrum After that, with the first Frame, First The amplitude and phase spectra of all frequency points on each channel are arranged in sequence, and the results are obtained by vertically stitching the amplitude and phase spectra of all frequency points, where: , , , In the formula, For frequency index, =0~F-1, where F is the total number of frequency points; Indicates frame shift, For the length of the window, This is the index for the sample offset within the window. For the signal at the sample point The amplitude value at that point, The term is a window function; Feature vector list Represented as: Abbreviated as: , , represents the index of a feature vector within the feature vector list.
6. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 5, characterized in that, The three-dimensional feature array P is obtained by normalizing the three-dimensional feature array. , is represented as: , To The result is obtained by normalizing all elements of all feature vectors in the dataset, and is represented as: ; in, , In the formula, For list The Middle Frame, First Normalized feature vectors of each channel, , , where represent the mean vector and standard deviation vector of the eigenvectors, respectively.
7. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 6, characterized in that, Add a three-dimensional feature array in S2 The labeling method is as follows: first, process the three-dimensional feature array... All feature vectors from the same channel are concatenated in time series order to obtain four forms. The four two-dimensional arrays are then concatenated along the channel dimension to obtain the form: A three-dimensional array, denoted as ; If a three-dimensional array If the data is generated based on a single-frequency signal or a single speech signal, the added tag will be in the form of: A three-dimensional array, where, in the form of The two-dimensional array represents the state information of the audio signal. If the sound source does not emit sound in a certain time frame of a single-frequency signal or a single-voice signal, the state label of that time frame is set to 0; otherwise, it is set to 1. The format is... The two-dimensional array represents the angular information of the audio signal, including the azimuth angle of the sound source relative to the microphone for each time frame. and pitch angle ; If a three-dimensional array Based on multiple speech signals, the tagging format is as follows: A two-dimensional array represents the state information of the audio signal. This is obtained by acquiring the state labels of all individual speech signals that make up the multi-speech signal. The form is The two-dimensional array is obtained by concatenating it along the second dimension.
8. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 1, characterized in that, The sound source localization / recognition model includes a CNN module, an RNN module, and an FC module connected in sequence. The number of convolutional kernels in the CNN module is set to be... The pooling factor list is as follows The RNN module sets a list of bidirectional gated recurrent unit neurons. The list of fully connected units is set in the FC module. It includes the Time Detection (SED) section and the Direction of Arrival (DOA) section. If the sound source localization / recognition model performs the task of locating and recognizing a single-frequency signal or a single speech signal, the RNN module output is simultaneously input to the Time Detection Part (SED) and Direction of Arrival (DOA) of the fully connected layer, and the final output format is as follows: A three-dimensional array, where the form is A two-dimensional array represents the state information of the audio signal, in the form of... The two-dimensional array represents the angle information of the audio signal; If performing a multi-speech signal recognition task, the RNN module outputs the input time detection part (SED), and the final output format is as follows: A two-dimensional array, representing the state information of the audio signal.
9. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 8, characterized in that, In the DOA section, the output of the RNN module is first fed into a first fully connected layer with a time distribution for iterative processing. After the iteration, the result is fed into another second fully connected layer, and the output is specified to be in the form of a two-dimensional array. , in the form of The two-dimensional array is input into the hyperbolic tangent function tanh for activation, and the DOA output is further transformed into the following form: Transformation of a two-dimensional array into The transformation is based on the unit distance assumption, treating the output vector as three-dimensional coordinates, and obtaining the azimuth angle θ and elevation angle φ by solving the following system of equations: ; In the SED section, the output of the RNN module is first fed into the first fully connected layer of the time distribution for iterative processing. After the loop, the result is fed into another second fully connected layer, and the output is specified to be in the form of a two-dimensional array. , in the form of The two-dimensional array is input into the sigmoid function for activation, and the final output two-dimensional array is the result of SED.
10. The sound source localization and identification method based on four-channel audio signals from a vector microphone according to claim 8, characterized in that, The training method for the sound source localization / recognition model is as follows: Loss functions and optimization strategies are selected based on the specific task. In single-frequency signal and single-speech signal localization and recognition tasks, joint loss is used for optimization. Specifically, DOA output is for regression tasks, using mean absolute error or mean squared error as the loss function to measure the deviation between the predicted sound source angle and the true label; SED output is for binary classification tasks, using binary cross-entropy as the loss function to measure the deviation between the sound source state and the true label. The total loss function is a weighted sum of the two loss functions according to set weights, calculating the weighted error between the model output and the corresponding label. In multi-speech signal recognition tasks, binary cross-entropy is used as the loss function to calculate the error between the model output and the label. After completing one round of forward propagation, backpropagation is performed based on the calculated error values, and all trainable parameters in the network are updated through the optimizer until the set stopping condition is reached, completing the model training.