A sound source positioning method based on a neural network model

By using a neural network model-based sound source localization method, ADMM-Net is employed to achieve integrated solution for the separation and localization of direct sound and reverberant sound. This solves the problems of accuracy and reliability in sound source localization in enclosed spaces, adapts to varied and complex acoustic conditions, and improves localization accuracy and robustness.

CN122345836APending Publication Date: 2026-07-07HARBIN ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN ENG UNIV
Filing Date
2026-04-08
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing sound source localization technologies suffer from large deviations and poor reliability in enclosed or semi-enclosed spaces due to reverberation. They are also difficult to adapt to varied and complex acoustic conditions. Traditional methods require precise environmental modeling and parameter adjustment, making the process cumbersome and lacking in generalization ability.

Method used

A sound source localization method based on a neural network model is adopted. By introducing a decomposable observation model of direct sound and reverberation, the iterative process is expanded into a deep network using ADMM-Net to achieve integrated solution of direct sound and reverberation separation and sound source localization. Learnable nonlinear proximal mapping and adaptive adjustment are used to reduce dependence on environmental parameters.

Benefits of technology

It significantly suppresses spurious peaks and sidelobe interference, improves positioning accuracy and robustness, is applicable to a variety of complex reverberation scenarios, reduces the engineering burden of parameter adjustment and iterative solution, and has strong practicality and promotion value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122345836A_ABST
    Figure CN122345836A_ABST
Patent Text Reader

Abstract

In order to solve the aforementioned problem that the existing sound source positioning technology cannot adapt to the complex acoustic conditions, the present application provides a sound source positioning method based on a neural network model, comprising the following steps: obtaining sound pressure data of a sound field, performing discrete Fourier transform on the time-domain sound pressure data to map it to a frequency domain to obtain complex-valued sound pressure data; inputting the complex-valued sound pressure data as input data into a neural network model; the neural network model comprises a reconstruction layer, a nonlinear threshold layer and a multiplier update layer; the structure of the reconstruction layer comprises a convolution layer, an activation function, a deconvolution layer and a skip connection; and the neural network model outputs a sound source positioning result. The technical scheme of the present application can be widely applied to the field of sound source positioning technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source localization technology, and in particular to the field of using neural network models to analyze sound field data and locate sound sources. Background Technology

[0002] Existing sound source localization technologies are typically based on the free-field assumption, employing techniques such as sparse reconstruction and beamforming to invert the sound source location using array measurement data. However, in enclosed or semi-enclosed spaces, due to multiple reflections and scatterings of sound waves on walls and obstacles, the sound field is superimposed with a large number of indirectly propagating reverberation components, forming multiple illusory sound sources (pseudo-sources). This severely interferes with the accuracy of sound source localization, resulting in large deviations and poor reliability in the localization results. Existing dereverberation sound source localization methods proposed to reduce the impact of reverberation often rely on the precise measurement and modeling of acoustic environmental parameters (such as room geometry and material absorption coefficients). Furthermore, they require repeated adjustments to regularization weights and threshold parameters for specific application scenarios, making them difficult to adapt to varied and complex acoustic conditions. The process is cumbersome and lacks generalization ability. Summary of the Invention

[0003] To address the aforementioned problem that existing sound source localization technologies struggle to adapt to diverse and complex acoustic conditions, this invention provides a sound source localization method based on a neural network model.

[0004] The technical solution of the present invention is as follows:

[0005] A sound source localization method based on a neural network model includes the following steps:

[0006] S1. Obtain the time-domain sound pressure data of the sound field, and perform a discrete Fourier transform on the time-domain sound pressure data to map it to the frequency domain to obtain complex-valued sound pressure data.

[0007] S2. Input the complex sound pressure data as input data into the neural network model; the neural network model includes a reconstruction layer, a nonlinear threshold layer, and a multiplier update layer; the structure of the reconstruction layer includes: a convolutional layer, an activation function, a deconvolutional layer, and skip connections;

[0008] The activation function expression is:

[0009]

[0010] in, It is the identity matrix. ; For the sound source component, For reverberation components, For the free field Green's function, For plane wave dictionary matrix, , For plane wave dictionary matrix, As an additional variable, Let Lagrange multiplier vectors be used. The penalty parameter is defined as follows: the superscript i represents the i-th iteration, and the superscript i-1 represents the (i-1)-th iteration.

[0011] S3. The neural network model outputs the sound source localization result.

[0012] Optionally, the reconstruction layer includes sub-layers: a free field reconstruction layer X and a reverberation field reconstruction layer U;

[0013] The expression for the free field reconstruction layer X is:

[0014]

[0015] The expression for the reverberation field reconstruction layer U is:

[0016]

[0017] in, As an additional variable, Let Lagrange multiplier vectors be used. This is the penalty parameter.

[0018] Optionally, the structure of the nonlinear threshold layer includes: a piecewise linear function and a convolutional layer; the expression of the piecewise linear function is:

[0019]

[0020] in The input variables represent the piecewise linear function. Indicates input The output result after processing by a piecewise linear function Indicates the first The x-coordinate of each segment node, Each of the segment nodes satisfies a monotonically increasing relationship:

[0021]

[0022] in Let x be the x-coordinate of the leftmost node. The x-coordinate of the rightmost node;

[0023] Represents nodes The corresponding ordinate function value, where the superscript Indicates the first The function value of the next iteration This represents the total number of nodes in a piecewise linear function. This represents a range index used to determine the input. The current segment interval, Indicates that it contains input The x-coordinate of the left endpoint of the current segment interval. This represents the x-coordinate of the right endpoint of the current segment interval. and The left and right endpoints of the current segmented interval are respectively represented as the points in the first, second, and third positions. The function value corresponding to the next iteration.

[0024] Optionally, the nonlinear threshold layer includes sublayers: a free-field nonlinear layer Z and a reverberant field nonlinear layer W;

[0025] The expression for the free-field nonlinear layer Z is:

[0026]

[0027] The expression for the nonlinear layer W of the reverberation field is:

[0028] .

[0029] Optionally, the multiplier update layer includes sub-layers: a free field multiplier update layer N and a reverberant field multiplier update layer M;

[0030] The expression for the free-field multiplier update layer N is:

[0031]

[0032] The expression for the reverberant field multiplier update layer M is:

[0033] .

[0034] The technical effects of this invention are as follows:

[0035] This invention presents a sound source localization method based on a neural network model. It introduces a decomposable observation model for direct sound and reverberation, and expands the ADMM iterative process for solving this model layer by layer into an interpretable deep network, ADMM-Net. While retaining optimized structures such as data consistency updates and proximal constraints, the neural network model extends the key shrinkage operator from a fixed soft threshold to a learnable nonlinear proximal mapping. Furthermore, it performs data-driven adaptive adjustment of the iterative correlation coefficient, enabling the model to automatically match the sparse characteristics of direct sound with the structural characteristics of reverberation components under different reverberation intensities and frequencies. Therefore, this invention achieves integrated solution for the separation of direct sound and reverberation and sound source localization in the multi-channel observation domain without relying on precise room reflection parameters or rigorous environmental modeling. It significantly suppresses spurious peaks and sidelobe interference, improves localization accuracy and robustness, and achieves the objectives of this invention. This method reduces the engineering burden of repeated parameter adjustments and iterative solutions in traditional iterative algorithms, is applicable to various complex reverberation scenarios, and has strong practicality and promotional value.

[0036] The further effects of the above-mentioned alternative methods will be explained in detail below with reference to specific implementation methods. Attached Figure Description

[0037] Figure 1 This is a flowchart of the sound source localization method of the present invention.

[0038] Figure 2 This is a schematic diagram of the structure of a neural network model.

[0039] Figure 3 This is a schematic diagram of the reconstruction layer.

[0040] Figure 4 This is a schematic diagram of the nonlinear threshold layer.

[0041] Figure 5 This is a comparison chart of the first results from the simulation experiment.

[0042] Figure 6 This is the second comparison chart of the simulation experiment results.

[0043] Figure 7 This is a diagram of the experimental setup.

[0044] Figure 8 This is the first comparison chart of the experimental results.

[0045] Figure 9 This is a comparison chart of the second result from the experiment. Detailed Implementation

[0046] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0047] Figure 1 The main steps of the sound source localization method of the present invention are shown, and will be described below.

[0048] Obtain complex data of the sound field

[0049] S1. In this step, sound pressure data of the sound field is obtained using existing technology, and the time-domain sound pressure data is mapped to the frequency domain using a discrete Fourier transform to obtain complex-valued sound pressure data:

[0050]

[0051] in, For complex sound pressure levels in the frequency domain, This is a Discrete Fourier Transform operation. The time-domain signal acquired by the microphone. For time.

[0052] Input neural network model

[0053] S2. In this step, the complex sound pressure data obtained in the previous step is used as input data to input the neural network model. The neural network model of this invention is based on the sound field morphology decomposition and interpretable depth alternating direction multiplier method network (ADMM-Net), that is, the iterative optimization process of alternating direction multiplier method (ADMM) is mapped to a neural network model that can be trained end-to-end.

[0054] The process of obtaining a neural network model is explained below.

[0055] In a reverberant environment, location Sound pressure Satisfies the nonhomogeneous Helmholtz equation:

[0056]

[0057] in It is the Laplace operator. It is the wave number. It refers to the distribution of sound sources.

[0058] Decompose the sound field into:

[0059]

[0060] in For direct sound (particular solution), it can be represented as the convolution of free-field Green's functions:

[0061]

[0062] in The Green's function is a free-field function used to describe the sound wave's movement from... spread to Acoustic propagation characteristics It is the location of the sound source. It refers to the distribution of sound sources.

[0063] The reverberant component can be expanded into a plane wave basis:

[0064]

[0065] in It is a plane wave index. It is the correlation coefficient. The basis function components of a plane wave can be represented as follows: These are plane wave basis functions. and Indicates the first Two-dimensional Cartesian coordinates of a microphone and These represent the sine and cosine functions, respectively.

[0066] This allows us to obtain the following model for the decomposition of the sound field morphology:

[0067]

[0068] The discrete form of the above formula can be obtained:

[0069]

[0070] in It is the sound source component. It is the reverberation component. For the free field Green's function, This is the dictionary matrix for plane waves.

[0071] The model is optimized and sound source components are recovered using the following methods. and reverberation component :

[0072]

[0073] The solution is obtained using the alternating direction multiplier method (ADMM), which decomposes the problem into:

[0074]

[0075] in It is an additional variable. and It is a Lagrange multiplier vector. These are the regularization parameters. These are the penalty parameters.

[0076] The minimum value of the augmented Lagrange function can be decomposed into the following six subproblems:

[0077] renew :

[0078]

[0079] renew :

[0080]

[0081] renew :

[0082]

[0083] renew :

[0084]

[0085] renew :

[0086]

[0087] renew :

[0088]

[0089] The aforementioned iterative process using ADMM is summarized as follows:

[0090] Free field:

[0091]

[0092] Reverberation field:

[0093]

[0094] superscript This indicates the i-th iteration.

[0095] By fixing the aforementioned ADMM solution iterative process into network layers, with each layer corresponding to one iteration of ADMM, the following neural network layers are formed:

[0096] Reconstruction layer: Matrix inversion is replaced by a neural network (implemented through Neumann series approximation + convolutional neural network). The reconstruction layer further has two sub-layers: a free field reconstruction layer (X) and a reverberant field reconstruction layer (U).

[0097] The expression for the free-field reconstruction layer X is:

[0098] The expression for the reverberation field reconstruction layer U is:

[0099] Reconstruction layer Its function is to analyze the sound source signal. Perform estimation and updates. (Based on known...) Layer output In the case of the first The output of the layer is defined as:

[0100]

[0101] in express The learnable penalty parameter for the stage. Solving the above equation involves matrices. The inverse operation of the matrix is ​​computationally expensive. This computation has high time complexity, leading to a significant decrease in model training efficiency; therefore, an alternative method is needed. This paper utilizes Neumann series expansion to determine the inverse matrix and models each expanded term using convolutional and deconvolutional neural networks, effectively solving the matrix inversion problem. Specifically, the inverse matrix is ​​expanded as follows:

[0102]

[0103] Assumption and Representing the convolution and transpose convolution operations respectively, the above formula can be written as:

[0104]

[0105] in Activation functions are used to introduce nonlinear features into the network structure, enabling deep neural networks to learn and fit complex sound field distribution patterns. Commonly used activation functions include ReLU, Leaky-ReLU, and Softplus. Taking a finite number of terms while satisfying accuracy requirements, it can be further expressed as...

[0106]

[0107] in The reconstruction layer consists of convolutional layers, activation functions, deconvolutional layers, and skip connections. Convolutional layers extract spatial correlation information from the input features; activation functions introduce nonlinearity into the network, thereby enhancing the neural network model's fitting performance to complex sound fields; deconvolutional layers are responsible for restoring the spatial resolution of the features and generating reconstruction results; skip connections establish direct mapping paths between input and output to alleviate gradient decay problems in deep structures.

[0108] If all the numerical calculations described above are used, the computational efficiency of the reconstruction layer will become extremely low. To address this issue, this invention uses convolution and deconvolution to replace complex matrix calculations. The overall structure of the reconstruction layer is as follows: Figure 3 As shown. The reconstruction layer first receives input. The signal is processed using complex convolution to simultaneously capture the amplitude and phase information of the sound pressure signal, enabling joint modeling of complex acoustic propagation characteristics. Subsequently, the LeakyReLU activation function is applied to allow small negative gradients to pass through, thus avoiding neuron inactivation and enhancing the neural network model's ability to learn complex sound field features. Next, deconvolution (transposed convolution) is used to achieve upsampling, gradually restoring detailed information about the spatial distribution of the sound source, allowing for accurate reconstruction of the sound source location in the spatial domain.

[0109] Nonlinear Thresholding Layer: The threshold shrinkage operation is learned using a piecewise linear function (PLF). The nonlinear thresholding layer has two sublayers: a free-field nonlinear layer (Z) and a reverberant-field nonlinear layer (W).

[0110] The expression for the free-field nonlinear layer Z is: .

[0111] The expression for the nonlinear layer W of the reverberation field is: .

[0112] To improve the flexibility and adaptability of neural network models when processing complex nonlinear sound field data, a piecewise linear function (PLF) is used instead of the traditional soft thresholding function. The piecewise linear function can employ different linear mappings in different intervals, thus more effectively describing complex input-output relationships and enabling the model to more precisely characterize the nonlinear features in the reverberant sound field. Given... Layer output and its corresponding Lagrange multipliers , No. The output of the layer is

[0113]

[0114] in This represents a piecewise linear function, defined by a set of control points. Decide, This indicates the number of control points. The expression for PLF is:

[0115]

[0116] in The input variables represent the piecewise linear function, which are usually scalar input values ​​to be mapped. Indicates input The output result after processing with a piecewise linear function. Indicates the first The x-coordinate of each segment node, Each segment point satisfies a monotonically increasing relationship:

[0117]

[0118] in Let x be the x-coordinate of the leftmost node. is the x-coordinate of the rightmost node.

[0119] Represents nodes The corresponding ordinate function value, where the superscript Indicates the first The function value of the next iteration. This represents the total number of nodes in a piecewise linear function. This represents a range index used to determine the input. The current segment interval. When the input satisfies:

[0120]

[0121] When, explain Falling in each interval Inside, Indicates that it contains input The x-coordinate of the left endpoint of the current segment interval. This represents the x-coordinate of the right endpoint of the current segment interval. and These represent the left and right endpoints of the current segment interval at the [number]th position, respectively. The function value corresponding to the next iteration.

[0122] Figure 4 The structure of the nonlinear threshold layer is shown. For example... Figure 4 As shown, the threshold function is approximated by a piecewise linear function, enabling the network to adaptively learn the nonlinear characteristics of the acoustic environment in the sound source localization problem. This piecewise linear function-based design allows the network to flexibly approximate any continuous function, effectively addressing the complex nonlinear relationships in the sound source localization process under reverberant environments. Through a hierarchical linear approximation mechanism, the network can still learn patterns from data even when acoustic characteristics are unknown or difficult to determine, achieving higher localization accuracy and robustness.

[0123] Multiplier Update Layer: Implements the same Lagrange multiplier update as ADMM.

[0124] The expression for the free-field multiplier update layer N is:

[0125] The expression for the reverberant field multiplier update layer M is:

[0126] Figure 2 ADMM was shown in the Data nodes and their flow in the next iteration. For example... Figure 2 As shown, the neural network model has 6 layers:

[0127] (1) Corresponding variables Reconstruction layer ;

[0128] (2) Corresponding variables Reconstruction layer ;

[0129] (3) Corresponding variables nonlinear layer ;

[0130] (4) Corresponding variables nonlinear layer ;

[0131] (5) Corresponding variables Update layer ;

[0132] (6) Corresponding variables Update layer .

[0133] The six neural network layers described above correspond to the continuous iterative process of ADMM. By truncating a fixed number of iterations and expanding each of the six neural network layers into a network layer with learnable parameters, the overall framework of the neural network model can be constructed. This method preserves the original optimization structure of ADMM while endowing each layer with learnability, enabling the neural network model to possess both physical interpretability and data-driven adaptive optimization.

[0134] After training, neural network models can... Figure 1 The "Input Neural Network Model" step shown is used. The neural network model is trained using multiple sets of sound source locations, different reverberation times, and frequencies in a closed-space virtual environment. The training loss function is MSE, which simultaneously evaluates localization error and normalized power error.

[0135] Obtain sound source localization results

[0136] S3. In this step, the neural network model outputs the sound source localization result.

[0137] The technical solution of the present invention is verified through simulation experiments and real-world scenario experiments, which are described below.

[0138] Simulation Experiment

[0139] A simulated indoor sound field was built using Pyroomacoustics, generating 1600 sets of data for training and testing at different reverberation times. Comparisons were made with ADMM-Net (the method of this invention), FISTA-Net (a fast iterative thresholding algorithm for network expansion), U-Net (a classic deep learning method), DAMAS+De-reverb (a traditional method of source mapping deconvolution + dereverberation), FISTA+De-reverb (a fast iterative thresholding algorithm + dereverberation), and SF-MCA (sound field morphology decomposition). At high reverberation (420ms), ADMM-Net exhibited the lowest localization error, the smallest normalized power error, and accurately separated direct sound from spurious sources. Figure 5 The squares represent the identified sound sources, while the cyan dots represent the actual sound source locations. When the actual sound source location coincides with a square, the sound source has been located. Figure 5 In (a), the calculation time of the method of the present invention is 1 hour, and in (f), the calculation time is 1 day. Figure 6 The results are from the sound source localization quantization. The bar chart represents the normalized power error (NPE), with smaller values ​​indicating better sound source localization. The curve represents the localization error (LE), with smaller values ​​indicating better results. These two methods are complementary. We compared the method with five other methods to verify the accuracy of this invention.

[0140] Real-world scenario experiment

[0141] Figure 7 The image shows a real-world experimental setup and parameter settings. The enclosed space used in the experiment measures 1.5m x 1.5m x 1.5m.

[0142] The reverberation time T20 of the experimental environment was measured to be 0.12 s using the interrupted sound source and inverse integration method. The sound source localization results at different frequencies were compared with those of ADMM-Net (the method of this invention), FISTA-Net, U-Net, traditional DAMAS+De-reverb, FISTA+De-reverb, and SF-MCA methods as follows: Figure 8 and Figure 9 As shown. Figure 8 and Figure 9 The results show that the method of the present invention yields the optimal result.

[0143] It is worth noting that the above description is only a preferred embodiment of the present invention and does not limit the scope of patent protection of the present invention. The present invention can also be replaced by equivalent technologies. Therefore, all equivalent changes made based on the description and figures of the present invention, or direct or indirect applications to other related technical fields, are included within the scope of the present invention.

Claims

1. A sound source localization method based on a neural network model, characterized in that: Includes the following steps: S1. Obtain the time-domain sound pressure data of the sound field, and perform a discrete Fourier transform on the time-domain sound pressure data to map it to the frequency domain to obtain complex-valued sound pressure data. S2. Input the complex sound pressure data as input data into the neural network model; The neural network model includes a reconstruction layer, a nonlinear threshold layer, and a multiplier update layer; the structure of the reconstruction layer includes: a convolutional layer, an activation function, a deconvolutional layer, and skip connections; The activation function expression is: in, It is the identity matrix. ; For the sound source component, For reverberation components, For the free field Green's function, For plane wave dictionary matrix, , For plane wave dictionary matrix, As an additional variable, Let Lagrange multiplier vectors be used. The penalty parameter is defined as follows: the superscript i represents the i-th iteration, and the superscript i-1 represents the (i-1)-th iteration. S3. The neural network model outputs the sound source localization result.

2. The sound source localization method based on a neural network model according to claim 1, characterized in that: The reconstruction layer includes sub-layers: free field reconstruction layer X and reverberant field reconstruction layer U; The expression for the free field reconstruction layer X is: The expression for the reverberation field reconstruction layer U is: in, As an additional variable, Let Lagrange multiplier vectors be used. This is the penalty parameter.

3. The sound source localization method based on a neural network model according to claim 2, characterized in that: The structure of the nonlinear threshold layer includes: a piecewise linear function and a convolutional layer; the expression of the piecewise linear function is: in The input variables represent the piecewise linear function. Indicates input The output result after processing by a piecewise linear function Indicates the first The x-coordinate of each segment node, Each of the segment nodes satisfies a monotonically increasing relationship: in Let x be the x-coordinate of the leftmost node. The x-coordinate of the rightmost node; Represents nodes The corresponding ordinate function value, where the superscript Indicates the first The function value of the next iteration This represents the total number of nodes in a piecewise linear function. This represents a range index used to determine the input. The current segment interval, Indicates that it contains input The x-coordinate of the left endpoint of the current segment interval. This represents the x-coordinate of the right endpoint of the current segment interval. and The left and right endpoints of the current segmented interval are respectively represented as the points in the first, second, and third positions. The function value corresponding to the next iteration.

4. The sound source localization method based on a neural network model according to claim 3, characterized in that: The nonlinear threshold layer includes sublayers: a free-field nonlinear layer Z and a reverberant-field nonlinear layer W; The expression for the free-field nonlinear layer Z is: The expression for the nonlinear layer W of the reverberation field is: 。 5. The sound source localization method based on a neural network model according to claim 4, characterized in that: The multiplier update layer includes sub-layers: free field multiplier update layer N and reverberant field multiplier update layer M; The expression for the free-field multiplier update layer N is: The expression for the reverberant field multiplier update layer M is: 。