Depth expansion sound source localization method and system based on physical information inspiration
By combining physical models and data-driven learning, a deep unfolding sound source localization method was developed, which solved the problems of accuracy and robustness in sound source localization in complex industrial environments, and achieved high-precision and high-robustness fault sound source localization.
Patent Information
- Application Number
- CN202511007938.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
AI Technical Summary
Existing sound source localization technologies face problems such as inaccurate localization and poor robustness in complex industrial environments, especially in fault diagnosis of converter valves in high-voltage power equipment. Traditional methods fail due to model mismatch, while data-driven methods lack interpretability and have limited generalization ability.
By combining prior knowledge from physical models with data-driven adaptive learning, a physical information-inspired deep unfolding sound source localization method is constructed through sparse feature enhancement, deep acoustic feature extraction, and iterative optimization localization. A hybrid loss function is used for end-to-end training to adaptively handle strong noise and reverberation interference.
It significantly improves the positioning accuracy and robustness in complex industrial environments, overcomes the model mismatch of traditional methods and the lack of interpretability of data-driven methods, and achieves high-precision and high-robust fault sound source localization.
Smart Images

Figure CN120847723A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of acoustic signal processing and artificial intelligence technology. Specifically, it relates to a physical information-inspired deep expansion sound source localization method and system for use in complex industrial environments, particularly suitable for acoustic monitoring and diagnosis of faults in high-voltage power equipment (converter valves). Background Technology
[0002] However, in practical industrial scenarios where sound source localization technology is applied to converter valve fault diagnosis, existing technical solutions face severe challenges, which can be mainly divided into the following two categories: The first category is traditional sound source localization methods based on accurate physical models, such as beamforming methods based on generalized cross-correlation (GCC-PHAT) time-delay estimation (SRP-PHAT) and high-resolution subspace methods such as multiple signal classification (MUSIC). These methods rely on an idealized acoustic propagation model, assuming that the sound source is in a free field and far field. However, in real industrial environments such as converter halls where converter valves are located, this physical model assumption is severely compromised: First, the large internal space and hard surfaces such as walls in converter halls lead to severe acoustic reverberation and multipath effects, resulting in the microphone receiving a superposition of direct sound and a large number of reflected sounds, which seriously interferes with time-delay-based localization algorithms; second, auxiliary equipment such as cooling fans generate strong broadband background noise that may overlap with the fault sound frequency band; and third, strong electromagnetic interference (EMI) also contaminates the electrical signals collected by the microphone. Under the combined effect of these factors, the physical assumptions of traditional model-driven methods fail, leading to a sharp decline in their positioning accuracy, or even complete failure, and poor robustness.
[0003] The second category is data-driven deep learning-based sound source localization methods, such as using convolutional neural networks (CNNs) or convolutional recurrent neural networks (CRNNs) to directly learn the mapping relationship between sound source locations from acoustic features. These methods have shown good performance on specific datasets, but their limitations are also significant: First, they are typically "black box" models, lacking physical interpretability in their decision-making process. In fields requiring high reliability, such as power safety, their diagnostic results are difficult to trust and verify. Second, the performance of deep learning models is highly dependent on large-scale, diverse labeled datasets. However, acquiring and labeling acoustic data covering all fault types and operating conditions in industrial settings is extremely costly, and the scarcity of labeled samples severely restricts their application. Finally, purely data-driven models have limited generalization ability; a model trained in one environment will show a significant performance drop in another environment with slightly different noise and reverberation characteristics.
[0004] In summary, existing sound source localization technologies, whether traditional model-driven or emerging data-driven methods, struggle to effectively address the inaccurate and robust localization of fault sound sources in complex industrial scenarios such as converter valves, where factors like strong noise, strong reverberation, and weak, sparse signals contribute to their inaccurateness. Therefore, there is an urgent need in this field for a novel sound source localization technology that combines the prior knowledge of physical models with the adaptive learning capabilities of data-driven approaches to overcome the shortcomings of existing technologies. Summary of the Invention
[0005] To address the shortcomings mentioned in the background art, the present invention aims to provide a deep unfolding sound source localization method based on physical information inspiration, which aims to deeply integrate the prior knowledge of the physical model with the data-driven adaptive learning capability to achieve high-precision and high-robustness sound source localization.
[0006] In a first aspect, the present invention provides a depth-deployed sound source localization method inspired by physical information, the method comprising the following steps:
[0007] Step 1: Acquisition and time-frequency transformation of multi-channel audio signals; Receive multi-channel time-domain audio signals x(t) = [x1(t), x2(t), ..., x9(t)] acquired by a microphone array containing 9 microphones. T For each channel's signal x m (t) Perform a short-time Fourier transform (STFT) to convert it into a time-frequency domain representation, resulting in a complex time-frequency spectrum:
[0008]
[0009] Where l is the time frame index, f is the frequency point index, w(n) is the window function (using a Hanning window), N is the number of Fourier transform points, and H is the frame shift step size. The time-spectrum graphs of all channels constitute the original input tensor f(X) of the network;
[0010] Step Two: Feature Enhancement Based on Sparse Physical Priors; This step leverages the physical prior that industrial fault acoustic signals (such as partial discharge, electric arcs, etc.) typically exhibit sparse impacts in the time domain. It processes the input time-spectrum graph f(X) to suppress background noise and highlight fault features. This step is accomplished by a learnable sparse feature enhancement module. This module implements an iterative sparse recovery algorithm (i.e., Variational Bayesian Inference, VBI) to model the input signal f(X) as the result of convolving the sparse fault signal f(S) with the system transfer function, aiming to recover f(S). Its core is mapping the iterative update process of the VBI algorithm to a recurrent neural network layer. The update rule of this layer simulates the change in the posterior probability distribution of the sparse signal. The update can be abstracted as:
[0011]
[0012] Among them, f update It is a nonlinear function parameterized by the GRU neural network, θ S These are the learnable weights of the network. This module iterates several times and eventually outputs an enhanced, sparser feature representation.
[0013] Step 3: Deep acoustic feature extraction and initial localization; representing the enhanced sparse features X S The input is fed into a spatiotemporal feature extractor, which consists of a convolutional neural network (CNN). The CNN extracts features from X through multiple convolutional layers, ReLU activation functions, and pooling layers. S Learn a deep acoustic feature map F that is insensitive to changes in translation, scale, etc. and contains rich spatial information:
[0014] F = CNN(X) S ;θ CNN (3)
[0015] Where θ CNN These are the learnable weights of the CNN. Subsequently, the deep acoustic feature map F is input into an initial localization module (composed of a multilayer perceptron MLP) to regress a coarse initial location estimate θ0 of the sound source.
[0016] θ0 = MLP(Flatten(F); θ Init (4)
[0017] Where θ0 is a two-dimensional coordinate vector (azimuth and elevation angles);
[0018] Step 4: Iterative optimization localization through deep network expansion; This step is the core of this invention. It expands the optimization process of the classic iterative localization algorithm (iteration frequency focusing) into a deep network structure with a preset K layers. The network module of the kth layer (1≤k≤K) is called the Iterative Refinement Block (IRB), and its function is to receive the position estimate θ from the previous time step. k-1 And depth acoustic features F, and calculate a more accurate position estimate θ. k The update process in the k-th iteration is parameterized as a learnable function:
[0019] θ k =f IRB (θ k-1 ,F;θ IRB (5)
[0020] Specifically, this update can be decomposed into a residual update form:
[0021] θ k =θ k-1 +Δθ k (6)
[0022] Among them, the position update amount Δθ k The update prediction network (a GRU) within the IRB is dynamically generated. To improve robustness, a gating signal g generated by a context-aware network is introduced into the update process. k This signal is used to adaptively adjust the magnitude of the update:
[0023] Δθ k =g k ·f update_net (θ k-1 ,F;θ update (7)
[0024] g k =σ(f context_net (F;θ context (8)
[0025] Where σ is the ReLU activation function. The entire process starts with the initial position estimation θ0, iterates K times, and then outputs the final localization result θ. K Since all modules are differentiable, the network can be trained in an end-to-end manner.
[0026] Step 5: End-to-end network training based on a hybrid loss function; To collaboratively optimize the two tasks of feature enhancement and localization, this invention employs a hybrid loss function L... total Train the entire network:
[0027] L total =L localization +λL sparsity (9)
[0028] Where λ is a scalar hyperparameter used to balance the importance of the two loss terms. Localization loss L localization Used to measure the difference between the final positioning result and the true position θ gt The difference lies in the network's use of smoothing L1 loss, which is insensitive to outliers, and sparsity regularization loss L... sparsity This is used to guide the module in step two to learn sparse feature representations, by calculating its output X. S It is achieved using the L1 norm.
[0029] Secondly, in order to achieve the above objectives, the present invention provides a physical information-inspired depth-deployed sound source localization system for implementing the above method. This network, as a computer-implemented system or device, has the following specific structure:
[0030] The signal receiving and time-frequency conversion module is configured to receive multi-channel time-domain audio signals acquired by a microphone array containing M microphones. The short-time Fourier transform (STFT) is then performed on each channel signal to convert it into a complex time-spectrum tensor X∈C. M×L×F Where M is the number of channels, L is the number of time frames, and F is the number of frequency points;
[0031] The learnable sparse feature enhancement module, which is key to the physical prior injection in this invention, is configured to receive the original temporal spectrum tensor X and output a feature representation X with enhanced sparsity. S The internal structure of this module is not a traditional fixed filter. Instead, it unfolds the iterative update process into a learnable neural network layer with a preset number of iterations through a variational Bayesian inference (VBI) algorithm based on sparse Bayesian learning. In the i-th internal iteration, the posterior statistic (mean μ) of the estimated sparse signal is... S The update of ) is performed by a learnable weight θ S Parameterized function f VBI-update Replaced by:
[0032]
[0033] In this way, the network can learn the optimal sparse signal extraction strategy under specific noise and signal characteristics through end-to-end training, rather than simply executing a fixed algorithm.
[0034] A deep acoustic feature extraction module is configured to receive the feature representation X output by the sparse feature enhancement module. S It extracts deep, abstract spatial information useful for sound source localization tasks. Its internal structure is a convolutional neural network (CNN), containing multiple alternating stacked convolutional layers, non-linear activation function layers (ReLU), and pooling layers;
[0035] The initial localization and depth unfolding localization module consists of two cascaded sub-modules: the first module is the initial localization module, which internally contains a multilayer perceptron (MLP) configured to receive the feature map F output by the depth acoustic feature extraction module after flattening, and calculate a coarse initial position estimate θ0 of the sound source. The second module is the depth unfolding localization module, the core of the network, which internally consists of K cascaded, structurally identical, and weight-shared iterative optimization modules (IRBs). This module maps the iterative optimization process of traditional optimization algorithms to the forward propagation process of the network. The function of the k-th IRB (1≤k≤K) is to receive the position estimate θ0 from the previous time step. k-1It combines the deep acoustic feature map F and outputs an optimized, more accurate position estimate θ. k Its update rule is determined by a learnable weight θ. IRB The parameterized function is defined as: θ k =f IRB (θ k-1 ,F;θ IRB In one specific embodiment, the update function employs residual learning and introduces an adaptive gating mechanism: θ k =θ k-1 +Δθ k Among them, the update amount Δθ k It is generated by the update prediction network (a gated recurrent unit, GRU) within the IRB, while the adaptive gating signal g... k It is dynamically generated by the context-aware network (a small attention network) inside the IRB based on the input acoustic features F, and is used to automatically adjust the update step size under low signal-to-noise ratio conditions to prevent the model from diverging.
[0036] The storage and processing module, wherein the storage module is a non-volatile computer-readable storage medium, is used to permanently store the complete structural parameters of the network of this invention, as well as the optimal set of weight parameters {θ} obtained after training. S ,θ CNN ,θ Init ,θ IRB The processor is one or more central processing units (CPUs), graphics processing units (GPUs), or dedicated artificial intelligence (AI) chips (ASICs / FPGAs), configured to load the network structure and weights from the storage module and perform computational tasks defined by all the aforementioned modules, including a complete forward propagation process during the inference phase, and gradient calculation and backpropagation based on the hybrid loss function during the training phase to update the network weights.
[0037] The beneficial effects of this invention are:
[0038] This invention proposes a physics-inspired deep unfolded sound source localization system, which deeply integrates a physics-based iterative algorithm framework with a data-driven deep learning method. This effectively addresses the problems of poor robustness due to model mismatch in traditional sound source localization methods in industrial environments, as well as the insufficient interpretability and limited generalization ability of purely data-driven models. By "unfolding" the classic iterative localization algorithm into a network hierarchical structure, it retains the physical logic of the step-by-step solution of the traditional algorithm to improve the interpretability of the decision-making process. Furthermore, end-to-end training enables the network to autonomously learn and iterate key parameters, adaptively handling factors such as strong noise and strong reverberation. This method addresses the model mismatch caused by noise, significantly enhancing the robustness of localization in complex industrial environments. A learnable sparse feature enhancement module is designed, injecting prior physical knowledge of the sparse impact of fault sound signals to suppress background noise and enhance the separation of weak fault features, providing high-quality input to the localization module and fundamentally improving localization accuracy. A hybrid loss function, incorporating localization loss and sparsity regularization loss, is constructed, placing signal enhancement and localization tasks under a unified framework for end-to-end joint training. This avoids the accumulation of suboptimal solutions in traditional multi-stage processing, achieving multi-task collaborative optimization, accelerating convergence speed, and improving localization performance. Experimental results show that this method effectively overcomes strong noise and complex reverberation interference, achieving accurate and robust localization of weak and sparse fault sound sources in industrial scenarios, with significant performance improvements compared to traditional physical model methods and pure data-driven methods. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a schematic diagram of the overall process of a physical information-inspired deep unfolding sound source localization method according to the present invention.
[0041] Figure 2 This is the overall network architecture diagram of the sparse unfolded network inspired by physical information proposed in this invention;
[0042] Figure 3 This is the flowchart of the IRB module undergoing iterative optimization;
[0043] Figure 4 This is a diagram showing the experimental results comparing the positioning error of the method of this invention with other existing methods under different signal-to-noise ratio conditions.
[0044] Figure 5 This is a schematic diagram of the structure of a physical information-inspired depth unfolding sound source localization system according to the present invention. Detailed Implementation
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0046] Example 1:
[0047] like Figure 1 As shown, the depth-deployment sound source localization method based on physical information is characterized by the following steps:
[0048] S101: Receives and preprocesses multi-channel audio signals, injecting physical prior knowledge through the sparse feature enhancement module;
[0049] The method of this invention first receives multi-channel time-domain audio signals x(t) from the industrial equipment under test (high-pressure converter valve) through a microphone array containing nine microphones. These signals not only contain the fault sound signal to be located, but also contain strong background noise and reverberation. Based on the prior knowledge that fault sound signals are generally physically represented as a series of sparse impacts, this invention models the received noisy signal x(t) as follows:
[0050] x(t)=h(t)*s(t)+n(t) (1)
[0051] Where s(t) is the fault source signal with sparse characteristics to be determined, h(t) is the unknown acoustic system transfer function, and n(t) represents noise and interference.
[0052] To recover or enhance a sparse fault signal s(t) from a noisy background, this invention designs a sparse feature enhancement module. This module is not a fixed filter, but rather a variational Bayesian inference (VBI) algorithm based on sparse Bayesian learning, which unfolds the iterative update process into a learnable neural network layer with a preset number of iterations. This layer receives the time-spectrum representation f(X) of the original audio signal and, through several internal iterations simulating the VBI algorithm, enhances the posterior statistic (mean μ) of the sparse signal. s The update process for the i-th iteration can be abstracted as follows:
[0053]
[0054] Among them, f update It is composed of learnable weights θ SA parameterized nonlinear function (a small recurrent neural network). In this way, the network can autonomously learn the optimal denoising and sparse feature enhancement strategies, ultimately outputting an enhanced sparse feature representation X. S .
[0055] S102: Extract depth acoustic features and generate an initial localization estimate;
[0056] The sparse feature representation X output from step S101, enhanced with physical priors, is used to represent... S The input is fed into a spatiotemporal feature extractor. In this embodiment, the extractor is a convolutional neural network (CNN), which extracts features from X through a stack of multiple convolutional layers and non-linear activation functions (ReLU). S Extract the deep acoustic feature map F, which is invariant to changes in translation and scale, and contains rich spatial information of the sound source:
[0057] F = CNN(X) S ;θ CNN (3)
[0058] Where θ CNN These are the learnable weights of this CNN module.
[0059] The extracted depth acoustic feature map F is then input into an initial localization module. This module consists of a multilayer perceptron (MLP) used to perform regression analysis on the feature map to calculate a coarse initial location estimate θ0 of the sound source. This step provides a starting point for subsequent iterative fine-tuning of the localization.
[0060] S103: Iterative optimization and localization are performed through deep unfolding structure to obtain the final position estimate;
[0061] This step is the core of this invention. For example... Figure 2 As shown, this invention expands the physical optimization process of the classic iterative localization algorithm into a deep network structure with K cascaded iterative refinement blocks (IRBs).
[0062] The structure receives the initial position estimate θ0 and depth acoustic feature map F generated in S102, and performs K iterations of optimization. In the k-th iteration (1≤k≤K), the k-th IRB module receives the position estimate θ0 from the previous time step. k-1 The deep features F are used to calculate a position update Δθ through an internal learnable network. k This allows for a more accurate position estimate θ. k The update process employs residual learning:
[0063] θ k=θ k-1 +Δθ k (4)
[0064] To improve robustness in complex environments, this invention introduces an adaptive gating signal g generated by a context-aware network within the IRB. k This network can analyze the input acoustic features F, determine the signal-to-noise ratio or reverberation level of the current environment, and dynamically adjust the update amplitude. After introducing this mechanism, the update formula is optimized as follows:
[0065] θ k =θ k-1 +g k ·Δθ k (5)
[0066] The process starts at θ0, iterates K times, and then the last IRB module outputs θ. K This is the final result of the sound source location estimation.
[0067] S104: End-to-end training is performed using a hybrid loss function, and the total loss is calculated.
[0068] To achieve collaborative optimization between the sparse feature enhancement task in S101 and the localization task in S103, this invention employs a hybrid loss function L during the network training phase. total The loss function consists of two parts, and a balancing coefficient λ is introduced to weigh their importance:
[0069] L total =L localization +λL sparsity (6)
[0070] For the positioning loss L localization A smooth L1 loss function, which is insensitive to outliers, is used to minimize the final predicted position θ. K With the real location label θ gt The distance between them is calculated using the following formula:
[0071] L localization =SmoothL1(θ) K -θ gt (7)
[0072] For the sparsity regularization loss L sparsity The L1 norm is used to measure the feature representation X output by the sparse feature enhancement module in S101. S The sparsity of the loss term. This loss term acts as a physical constraint, guiding the network to learn and preserve the inherent sparse physical properties of the fault signal. Its mathematical expression is as follows:
[0073] Lsparsity =‖‖X S ‖‖1 (8)
[0074] During training, by minimizing the total loss L total Using the backpropagation algorithm, all learnable parameters (θ) of the entire network are processed. S ,θ CNN The weights of all IRB modules are synchronized end-to-end.
[0075] Specifically, the present invention will be further illustrated below through embodiments:
[0076] 3.1 Dataset
[0077] The dataset used for training and verification of the method of the present invention is constructed by combining simulation generation and field measurement to comprehensively cover various complex working conditions of high-pressure converter valve fault diagnosis.
[0078] The simulation dataset utilizes the acoustic simulation software (pyroomacoustics) to construct an acoustic model that matches the actual converter hall (30m × 20m × 15m, wall reverberation time T60 of 1.2 seconds). The sound source signals are actual fault sounds from converter valves, such as partial discharge and arcing, recorded on-site. Background noise is superimposed with on-site recorded cooling fan operating sounds and simulated broadband electromagnetic interference pulses. By combining these signals at different signal-to-noise ratios (-5dB to 20dB) and different sound source locations, a training sample containing 10,000 samples with precise location labels was generated.
[0079] The experimental dataset was constructed in a laboratory environment, using a test platform that included a 9-channel MEMS microphone array and a high-frequency signal generator. By playing simulated broadband fault sound signals at known locations, 500 sets of experimental data were collected for model validation and testing.
[0080] 3.2 Experimental Setup
[0081] The equipment used in the verification experiments of this invention consisted of an Intel Core i9-13900K CPU, an NVIDIA GeForce RTX 4090 24G graphics card, 64GB of RAM, and an Ubuntu 22.04 platform. The deep learning framework used was PyTorch, and the Python version was 3.10. During training, the batch size was set to 16, the number of epochs was set to 100, the AdamW optimizer was used, and the initial learning rate was 0.001. The ratio of the training set, validation set, and test set was 8:1:1.
[0082] Figure 4The figure shows the experimental results comparing the positioning errors of the proposed method with other existing methods under different signal-to-noise ratio (SNR) conditions. As can be seen from the figure, the PISU-Net method proposed in this invention (shown by the solid line) exhibits significantly lower mean absolute positioning error (MAE) than the traditional SRP-PHAT method and the purely data-driven CRNN network under various SNR conditions. Its performance advantage is particularly evident in the low SNR range (e.g., -5dB to 5dB), fully demonstrating the robustness and superiority of the proposed method.
[0083] Example 2: Second aspect, such as Figure 5 As shown, in order to achieve the above objectives, this invention discloses a depth-deployment sound source localization method inspired by physical information, comprising:
[0084] The signal receiving module 11 and the feature enhancement module 12 are used to receive multi-channel audio signals acquired by the microphone array, and based on the sparse physical prior of the fault sound signal, process the audio signals through an expanded, learnable sparse recovery network to suppress noise and output an enhanced sparse feature representation X. S ;
[0085] The acoustic feature extraction module 13 and the iterative depth unrolling localization module 14 are used to receive the sparse feature representation X. S The internal structure of this module is based on the physical model of the classic iterative localization algorithm. It includes an initial localization unit for generating an initial position estimate θ0, and K cascaded IRB modules for iterative optimization, ultimately calculating the final position estimate θ of the sound source. K ;
[0086] Model optimization and storage / processing module 15 is used during the training process of the system to optimize and store the model using a model containing localization loss L. localization and sparsity regularization loss L sparsity The hybrid loss function optimizes all learnable parameters in the system end-to-end, and stores and retrieves all network parameters and model processing results through the storage and processing module 15.
[0087] In conjunction with the second aspect, in some implementations of the second aspect, the system further includes: the signal receiving module 11 and the feature enhancement module 12 parameterize the iterative update process of solving the sparse signal s(t) into a parameter composed of learnable weights θ. S Control function f update :
[0088]
[0089] In the iterative positioning module 14, the position update process of the k-th IRB is described by the following formula: θ k=θ k-1 +g k ·Δθ k Among them, the update amount Δθ k and gating signal g k All of these are estimated by the learnable subnetworks within the IRB, based on the sparse feature representation of the input and the position at the previous time step, θ. k-1 Dynamically generated.
[0090] The model optimization, storage, and processing module 15 employs a hybrid loss function L. total Defined by the following formula:
[0091] L total =SmoothL1(θ) K -θ gt )+λ‖‖X S ‖‖1 (2)
[0092] Where, θ gt For the actual location label, |||X S ‖‖1 is the L1 norm of the sparse feature representation, and λ is the balance coefficient.
[0093] Based on the same inventive concept, this invention also provides a computer device, comprising: one or more processors, and a memory for storing one or more computer programs; the programs include program instructions, and the processor executes the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, used to implement one or more instructions, specifically for loading and executing one or more instructions stored in a computer storage medium to implement the above-described method.
[0094] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, performs the above-described method. This storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0095] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0096] The foregoing has shown and described the basic principles, main features, and advantages of this disclosure. Those skilled in the art should understand that this disclosure is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of this disclosure. Various changes and modifications can be made to this disclosure without departing from its spirit and scope, and all such changes and modifications fall within the scope of this disclosure as claimed.
Claims
1. A depth-expanded sound source localization method inspired by physical information, characterized in that, The method includes the following steps: Receive multi-channel audio signals acquired by a microphone array and input the audio signals into a pre-constructed, physically-inspired depth unfolding sound source localization method; The audio signal first flows through the sparse feature enhancement module within the network. The sparse feature enhancement module processes the audio signal based on the sparse physical prior of the fault sound signal to suppress noise and output an enhanced sparse feature representation. The sparse feature representation is input into the deep unrolling localization module in the network. The sparse feature enhancement module is formed by unrolling the physical model of the classic iterative localization algorithm. Through an iterative optimization process with a preset number of layers, the final location estimate of the sound source is calculated. During the training of the network, end-to-end parameter optimization is performed on the network using a hybrid loss function that includes localization loss and sparsity regularization loss.
2. The depth-unfolding sound source localization method based on physical information as described in claim 1, characterized in that, The sparse feature enhancement module will solve the sparse signal iterative algorithm including: a learned neural network layer; the iterative update process of the learned neural network layer is as follows: the received multi-channel audio signal x(t) is modeled as a convolution of sparse fault signal s(t) and system transfer function h(t) and noise n(t) is superimposed, and the calculation process is x(t) = h(t) * s(t) + n(t); The sparse feature enhancement module simulates the iterative process of solving the sparse fault signal s(t). In the i-th internal iteration, it calculates the posterior statistic of the sparse signal, where the posterior statistic is the mean μ. s The posterior statistic is updated by a learnable weight θ s Parameterized function f update Replaced by: The enhanced sparse feature representation is output through a preset number of iterations.
3. The depth-deployment sound source localization method based on physical information as described in claim 2, characterized in that, The iterative optimization process of the depth unfolding localization module includes: first, calculating the initial position estimate θ0 of the sound source based on the sparse feature representation through an initial localization module; Then, the initial position estimate θ0 and the sparse feature representation are jointly input into a deep unfolding structure consisting of K cascaded iterative optimization modules; the k-th iterative optimization module, 1≤k≤K, receives the position estimate θ0 output by the (k-1)-th module. k-1 And calculate the updated position estimate θ k .
4. The depth-unfolding sound source localization method based on physical information as described in claim 3, characterized in that, The position update process of the kth iterative optimization module is described by the following formula: i k =θ k-1 +Δθ k (1) Among them, the position update amount Δθ k The update prediction network within the IRB is based on θ k-1 The sparse feature representation is dynamically generated; a gating signal g generated by a context-aware network is introduced. k Adaptive adjustments to the update process: i k =θ k-1 +g k ·Dth k (2) Among them, g k The value is determined by the contextual information of the signal-to-noise ratio or reverberation level represented by the sparse features.
5. The depth-unfolding sound source localization method based on physical information as described in claim 1, characterized in that, The mathematical expression for the hybrid loss function is as follows: THE total =L localization +λL sparsity (3) Among them, L localization For the positioning loss, L sparsity Let λ be the sparsity regularization loss, and λ be the weight hyperparameter used to balance the two loss terms.
6. The depth-unfolding sound source localization method based on physical information as described in claim 5, characterized in that, The positioning loss L localization The final position estimate θ is calculated using a smoothed L1 loss function. K With the real location label θ gt Differences between them: L localization =SmoothL1(θ K -θ gt ) (4) The sparsity regularization loss L sparsity The feature representation X output by the sparse feature enhancement module is calculated. S The L1 norm is used to guide the network in learning the sparse structure of the signal: L sparsity =||X S ||1 (5)。 7. A depth-expanded sound source localization method inspired by physical information, characterized in that, include: The signal receiving and feature enhancement module is used to receive multi-channel audio signals acquired by a microphone array, and process the audio signals based on the sparse physical prior of the fault sound signal to suppress noise and output an enhanced sparse feature representation. The iterative localization module, whose internal structure is based on the physical model of the classic iterative localization algorithm, is used to receive the sparse feature representation and calculate the final location estimate of the sound source through a process that includes initial localization and iterative optimization with a preset number of layers. The model training and optimization module is used to perform end-to-end optimization of all learnable parameters in the system during the training process of the system using a hybrid loss function that includes localization loss and sparsity regularization loss.
8. The depth-unfolding sound source localization method based on physical information as described in claim 7, characterized in that: The signal receiving and feature enhancement module parameterizes the iterative update process of solving sparse signals into a function consisting of learnable weights θ. S Control function f update ; The iterative localization module includes an initial localization module and K cascaded iterative optimization modules (IRBs), wherein the position update process of the k-th IRB is derived by the following formula: i k =θ k-1 +g k ·Dth k (6) Where the update amount Δθ k and gating signal g k All are dynamically generated by the learnable subnetworks within the IRB; The hybrid loss function used in the model training and optimization module is defined by the following formula: L total =SmoothL1(θ K -θ gt )+λ||X S ||1 (7)。 9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it performs the physical information-inspired depth unfolding sound source localization method according to any one of claims 1 to 6.