Three-dimensional sound source positioning method and system based on frequency domain saliency adaptation and multi-task decoupling

CN122836664APending Publication Date: 2026-09-29DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611290090.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

若将上述任务直接放入完全共享的网络结构中联合优化,可能导致不同任务之间出现特征竞争或梯度干扰,进而影响模型训练稳定性和最终预测性能

Benefits of technology

1.提高了复杂声学环境下的频域特征利用能力:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122836664A_ABST
    Figure CN122836664A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a three-dimensional sound source positioning method and system based on frequency domain saliency adaptation and multi-task decoupling, which comprises the following steps: S1, adaptively recalibrating original time-frequency acoustic features to obtain enhanced features after recalibration; S2, obtaining high-dimensional acoustic representation based on the enhanced features; S3, obtaining decoupling processing results based on the high-dimensional acoustic representation through the decoupling branch structure; S4, extracting and utilizing the confidence mask to constrain the loss calculation or gradient back propagation range of the sound source distance related parameter estimation branch; S5, inputting the audio features to be processed into the optimized model to obtain corresponding parameters and perform post-processing to generate three-dimensional sound source parameter estimation results containing sound source direction information and distance related information. The application can realize joint estimation of sound event categories, direction parameters and distance related parameters, thereby improving the stability and availability of three-dimensional sound source perception results in complex acoustic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of audio signal processing, deep learning, and spatial acoustic modeling, and particularly to a three-dimensional sound source localization method and system based on frequency domain saliency adaptation and multi-task decoupling. Background Technology

[0002] With the development of technologies such as human-computer interaction, intelligent robots, security monitoring, autonomous driving, and spatial audio, sound source localization and detection technology plays a crucial role in perceiving complex acoustic scenes. This technology not only needs to determine the presence of a specific sound event in the environment but also needs to further estimate the time of the sound event and the spatial location of the sound source. Therefore, accurately acquiring the spatial parameters of the sound source is a vital foundation for achieving sound field understanding, target tracking, acoustic interaction, and intelligent decision-making.

[0003] Existing sound event localization and detection methods typically combine sound event detection with sound source direction estimation. Some methods employ Cartesian coordinate-based direction representations, such as Activity-Coupled Cartesian Direction of Arrival (ACCDOA) and its variants, which simultaneously represent the activity state and direction of arrival of a sound event using a three-dimensional direction vector. These methods have achieved some success in azimuth and elevation angle estimation, and can meet some two-dimensional or spherical direction sensing requirements.

[0004] However, in practical applications, sound source spatial perception often requires not only directional information but also radial distance information between the sound source and the receiving device. Most existing methods mainly focus on the mapping of direction vectors on a unit sphere, with insufficient modeling of the distance dimension. This makes it difficult for the system to directly obtain three-dimensional sound source spatial parameters, including azimuth, pitch, and distance. This limits the system's ability to express real spatial relationships in applications such as robot navigation, target tracking, spatial interaction, and sound field reconstruction.

[0005] Furthermore, reverberation, background noise, and frequency-selective attenuation are common acoustic phenomena in indoor and semi-open environments. The semantic information, spatial positioning information, and distance-related information of sound events contained in different frequency bands are not entirely consistent. If a fixed frequency domain feature representation is used, the model will find it difficult to adaptively adjust the contribution of each frequency band feature according to different acoustic scenarios, which can easily affect the stability of sound source direction estimation and distance-related parameter estimation.

[0006] Meanwhile, sound event detection, direction estimation, and distance estimation are related but not entirely identical tasks. Sound event detection relies more on semantic time-frequency patterns, direction estimation relies more on multi-channel spatial cues, and distance estimation is also affected by factors such as sound source energy attenuation, reverberation intensity, and frequency band response variations. If these tasks are directly put into a completely shared network structure for joint optimization, it may lead to feature competition or gradient interference between different tasks, thereby affecting the model's training stability and final prediction performance. Summary of the Invention

[0007] Based on this, in order to address the shortcomings of existing technologies, a three-dimensional sound source localization method and system based on frequency domain saliency adaptation and multi-task decoupling is proposed. It improves the joint modeling capabilities of sound event detection, direction estimation and distance perception in complex acoustic environments by using frequency domain feature adaptive modeling, multi-task decoupling learning and distance-related parameter estimation.

[0008] To achieve the above design objectives, the technical solution of the present invention is as follows: A three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling includes: S1: Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed, perform statistical compression on the original time-frequency acoustic features in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension, generate frequency domain saliency weights through nonlinear mapping, and use the frequency domain saliency weights to weight the original time-frequency acoustic features to obtain the recalibrated enhanced features. The enhanced features can adaptively recalibrate the original time-frequency acoustic features in the frequency dimension. S2: Based on the enhanced features, perform deep feature extraction to obtain a high-dimensional acoustic representation; S3: Based on the high-dimensional acoustic representation, it is input into the pre-constructed decoupled branch structure for forward propagation calculation to obtain the joint prediction result of acoustic event activity and sound source direction, as well as the predicted value of distance-related parameters; the decoupled branch structure includes a joint estimation branch of acoustic event activity and sound source direction and a sound source distance-related parameter estimation branch; S4: During the model training phase of the decoupled branch structure, the corresponding confidence information is extracted based on the output of the joint estimation branch of sound event activity and sound source direction, and a confidence mask is constructed. The confidence mask is then used to constrain the loss calculation or gradient backpropagation range of the sound source distance related parameter estimation branch. S5: Input the audio features to be processed into the model optimized by S4, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure reasoning; and perform post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.

[0009] Furthermore, S1 specifically includes the following steps: S11. Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed; S12. The original time-frequency acoustic features are statistically compressed in the time dimension and feature channel dimension to obtain a feature description vector in the frequency dimension. The feature channel dimension includes at least an acoustic energy characterization channel and a spatial cue characterization channel. The spatial cue characterization channel is used to characterize spatial information related to the direction of the sound source. S13. Frequency domain saliency weights are generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weights to obtain the recalibrated enhanced features; wherein, the frequency domain saliency weights are generated through a bottleneck mapping structure, which consists of a dimension-reducing fully connected layer, a nonlinear activation layer, and an dimension-increasing fully connected layer; the dimension-reducing fully connected layer is used to compress the frequency dimension of the feature description vector; the nonlinear activation layer is used to perform nonlinear mapping on the compressed features to introduce nonlinear expressions; the dimension-increasing fully connected layer is used to restore the features with introduced nonlinear expressions to the original frequency dimension.

[0010] Furthermore, the specific processing steps for each branch of the decoupled branch structure are as follows: The joint estimation branch of acoustic event activity and acoustic source direction is used to output a direction prediction vector that couples activity state and direction parameters for each time frame and acoustic event category. The magnitude of the direction prediction vector is used to characterize the activity confidence of the corresponding acoustic event category, and the normalized direction prediction vector is used to characterize the acoustic source direction parameters. The sound source distance-related parameter estimation branch is used to output distance parameters related to the radial position of the sound source.

[0011] Furthermore, in S4, constraining the loss calculation or gradient backpropagation range of the sound source distance correlation parameter estimation branch using the confidence mask includes: The confidence mask is used to weight the estimated loss between the predicted value of the distance-related parameters and the true label. When the confidence of the corresponding time frame or sound event category is higher than a preset threshold, the distance estimation loss corresponding to that time frame or sound event category is retained and used in gradient backpropagation to update the model parameters. When the confidence of the corresponding time frame or sound event category is lower than or equal to the preset threshold, the distance estimation loss corresponding to that time frame or sound event category is masked to block its gradient backpropagation path.

[0012] Furthermore, in S5, the post-processing includes using category-related thresholds to determine the activity state for different sound event categories, and performing time-domain sliding window smoothing on the direction parameters and distance-related parameters in consecutive time frames.

[0013] Based on the same inventive essence, this application also proposes a three-dimensional sound source localization system based on frequency domain saliency adaptation and multi-task decoupling, which includes: The frequency domain saliency adaptive module is connected to the audio feature input terminal and is used to obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed. The original time-frequency acoustic features are statistically compressed in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension. Frequency domain saliency weights are generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weights to obtain the recalibrated enhanced features. The deep feature encoding module, which is connected to the frequency domain saliency adaptive module, is used to perform deep feature extraction based on the enhanced features to obtain a high-dimensional acoustic representation. The multi-task decoupling processing module is connected to the deep feature encoding module, and includes a joint estimation branch for acoustic event activity and sound source direction and a sound source distance-related parameter estimation branch set in parallel; the joint estimation branch for acoustic event activity and sound source direction is used to output the direction prediction vector coupled with the activity state and direction parameter, and the sound source distance-related parameter estimation branch is used to output the distance-related parameter. The confidence constraint training module is connected to the multi-task decoupling processing module. It is used to obtain the active confidence based on the magnitude of the direction prediction vector, construct a confidence mask based on the active confidence, and use the confidence mask to constrain the loss calculation or gradient backpropagation range of the sound source distance related parameter estimation branch. The output post-processing module is connected to the confidence constraint training module. It is used to input the audio features to be processed into the model optimized by the confidence constraint training module, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure inference. The module then performs post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.

[0014] Implementing the embodiments of the present invention will have the following beneficial effects: 1. Improved the ability to utilize frequency domain characteristics in complex acoustic environments: This invention introduces a saliency adaptive module for the frequency dimension, which generates learnable weight coefficients for the input acoustic features in the frequency dimension and uses these weight coefficients to recalibrate features in different frequency bands. This approach enables the model to adaptively adjust the participation of each frequency band feature in sound event recognition, direction estimation, and distance-related parameter estimation tasks based on training data, avoiding complete reliance on fixed frequency band representations. This improves the model's adaptability to different sound source types, different spectral distributions, and complex acoustic conditions, providing a more flexible frequency domain representation for subsequent three-dimensional sound source parameter estimation.

[0015] 2. It expands the spatial perception dimension of traditional Sound Event Localization and Detection (SELD) systems: Traditional sound event localization and detection systems typically output sound event category, event time, and direction of arrival. Their spatial information, often expressed as azimuth, elevation, or unit direction vectors, is insufficient to directly describe the radial distance relationship between the sound source and the receiving device. To address this issue, this invention introduces a distance-related parameter estimation branch on top of the original direction estimation task, enabling the system to simultaneously obtain sound event category, direction information, and distance-related information. Thus, this invention extends sound source perception from traditional direction estimation to three-dimensional sound source localization that includes radial information, providing richer sound source state information for applications such as robot hearing, intelligent monitoring, human-computer interaction, and spatial sound field understanding. 3. Reduced mutual interference in multi-task joint modeling: Given the following characteristics: although acoustic event recognition, direction estimation, and distance-related parameter estimation all originate from the same multi-channel acoustic signal, the acoustic cues they rely on are not entirely consistent. Acoustic event recognition focuses more on semantic time-frequency patterns, direction estimation relies more on spatial cues such as inter-channel phase differences and intensity differences, and distance-related parameter estimation is also affected by factors such as sound source energy attenuation, reverberation ratio, and frequency band response changes. Directly using a completely shared network structure for joint optimization may lead to feature competition or gradient interference between different tasks. Therefore, this invention constructs a multi-task decoupled branch structure and combines it with confidence masks, loss weighting, or gradient constraint mechanisms to ensure that the distance-related parameter estimation branch primarily learns within the effective acoustic event region, reducing the impact of silent segments or low-confidence predictions on distance-related parameter regression. This design helps alleviate the optimization conflict between the direction estimation task and the distance-related parameter estimation task, improving the stability of the multi-task training process.

[0016] 4. Improved the continuity and usability of the three-dimensional sound source output results: This invention combines category-related thresholding and temporal smoothing strategies in the output stage to perform consistency processing on the sound event activity state, direction estimation results, and distance-related parameter estimation results. This processing method can reduce short-term false detections, spatial jumps, and random jitter commonly found in frame-by-frame prediction, making the output three-dimensional sound source parameter estimation results more continuous in the temporal dimension and more consistent with the basic laws of real sound source motion in the spatial dimension. Therefore, this invention can improve the interpretability and practical application usability of three-dimensional sound source perception results.

[0017] In summary, the frequency domain saliency adaptive module and multi-task decoupling branch structure adopted in this invention can be embedded as functional modules into existing acoustic event localization and detection frameworks without requiring a complete redesign of the basic encoder structure. Therefore, it can be said that the scheme adopted in this invention can introduce distance correlation sensing capabilities while maintaining the basic structure of the original backbone model, thereby achieving joint estimation of acoustic event categories, direction parameters, and distance correlation parameters. Furthermore, because the added module structure is relatively simple, it is easy to combine with existing multi-channel audio features, deep learning models, and pre-trained acoustic models, thus possessing good engineering scalability and application adaptability. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] in: Figure 1 This is a flowchart of the basic steps corresponding to the solution described in this invention; Figure 2 This is a schematic diagram of the overall system architecture corresponding to the implementation case of the present invention. This diagram is used to show the overall flowchart from audio input, feature extraction, frequency domain saliency adaptation, deep feature encoding, multi-task decoupling prediction to output post-processing. Figure 3 This is the internal logic structure diagram of the frequency domain saliency adaptive module corresponding to the embodiment of the present invention; Figure 4 This is a schematic diagram of the distance-related parameter estimation constraint process based on confidence mask during the training phase of an embodiment of the present invention. Figure 5 This is a schematic diagram of the output post-processing and result smoothing process corresponding to the implementation examples of this invention; Figure 6 This is a schematic diagram comparing the traditional directional positioning results with the three-dimensional sound source positioning results after introducing distance-related parameters in the present invention, according to an embodiment of the present invention. In the diagram: 101 represents the audio input, 102 represents the feature extraction unit, 103 represents the frequency domain explicitness enhancement module, 104 represents the deep feature encoding backbone network, 105 represents the bifurcation point, 106 represents the decoupled multi-task processing region, 107 represents the sound source distance-related parameter estimation branch, 108 represents the distance-related parameter estimation branch, 109 represents the 3D sound source coordinate output, 201 represents the original time-frequency acoustic feature tensor, 202 represents the global temporal average pooling unit, 203 represents the dimensionality reduction fully connected layer, 204 represents the nonlinear activation layer, 205 represents the dimensionality increase fully connected layer, and 206 represents the Sigmoid function. The mapping unit, 207 represents the element-wise multiplication operator, 208 represents the time-frequency acoustic features, 301 represents the confidence information, 302 represents the predicted value output by the distance-related parameter estimation branch, 303 represents the reference distance-related parameters, 304 represents the activity probability discrimination unit, 305 represents the mask generation module, 306 represents the distance calculation unit, 307 represents the original distance-related error, 308 represents the mask weighted multiplication operator, 309 represents the weighted distance-related loss input loss calculation and training control unit, 310 represents the model parameters, 401 represents the original model output sequence, 402 represents the category-related threshold discrimination unit, 403 represents the time-domain sliding window smoothing unit, and 404 represents the three-dimensional sound source parameter estimation result. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. It is understood that the terms “first,” “second,” etc., as used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element without departing from the scope of this application, and similarly, a second element may be referred to as a first element. Both the first element and the second element are elements, but they are not the same element.

[0022] This invention essentially designs a three-dimensional sound source perception and localization method based on frequency domain saliency adaptation and multi-task decoupling, oriented towards sound event detection, sound source direction estimation, and distance-related parameter estimation. Its overall architecture is established through the following core steps: First, we introduce a three-dimensional sound source localization framework that includes distance-related parameters: Unlike traditional sound event localization and detection methods that mainly focus on modeling sound event categories, activity times, and sound source direction information, this invention further introduces a distance-related parameter estimation branch on the basis of sound source direction estimation, enabling the system to simultaneously output the sound event activity status, sound source direction parameters, and distance-related parameters. Through a multi-task decoupling architecture, this invention extends traditional direction perception to three-dimensional sound source localization that includes radial information, which helps to supplement the shortcomings of existing SELD systems in distance-related spatial information modeling. Secondly, a frequency-oriented saliency adaptive recalibration mechanism is proposed. Addressing the issue that different frequency bands in multi-channel acoustic features contribute differently to the joint estimation of acoustic event activity and direction, as well as the estimation of distance-related parameters, this invention proposes a frequency-axis-oriented saliency adaptive module. This module statistically compresses the input time-frequency acoustic features along the time and feature channel dimensions to obtain a feature description vector along the frequency dimension. Subsequently, a bottleneck mapping structure generates frequency domain saliency weights, which are then applied to the original time-frequency acoustic features to achieve adaptive recalibration of features in different frequency bands. Through this method, the model can adaptively adjust the participation level of different frequency band features in subsequent tasks based on training data, improving the flexibility of frequency domain feature representation and scene adaptability. Furthermore, a confidence-mask-based training constraint mechanism for distance-related parameters is proposed. While acoustic event activity detection, source direction estimation, and distance-related parameter estimation are correlated, their respective acoustic cues are not entirely consistent. Direct joint optimization could lead to feature competition or gradient interference between different tasks. Therefore, this invention proposes a confidence-mask-based training constraint mechanism. A mask is constructed using the confidence information output from the joint estimation branch of acoustic event activity and source direction, and this mask is used to constrain the loss calculation of the distance-related parameter estimation branch. In this way, the distance-related parameter estimation branch mainly participates in training within the effective acoustic event region, thereby reducing the impact of silent segments, low-confidence regions, or invalid acoustic segments on the learning of distance-related parameters, and contributing to improving the stability of the multi-task joint training process.

[0023] Finally, by combining category-related thresholds and temporal smoothing to improve the stability of the output results, this invention calculates the activity confidence based on the direction prediction vector output by the joint estimation branch of sound event activity and sound source direction during the inference stage, and uses a preset threshold or category-related threshold to distinguish the activity state of different sound event categories; the direction prediction vector determined to be active is normalized to obtain the sound source direction parameter; simultaneously, the distance correlation parameter output by the distance correlation parameter estimation branch is combined to generate the three-dimensional sound source parameter estimation result. Furthermore, temporal smoothing is performed on the sound event activity state, sound source direction parameter, and distance correlation parameter on continuous time frames to reduce short-term false detections and parameter jumps.

[0024] Based on the above design framework, this embodiment proposes a three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling, such as... Figures 1-2 As shown, the method includes the following steps: S1: Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed, perform statistical compression on the original time-frequency acoustic features in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension, generate frequency domain saliency weights through nonlinear mapping, and use the frequency domain saliency weights to weight the original time-frequency acoustic features to obtain the recalibrated enhanced features. The enhanced features can adaptively recalibrate the original time-frequency acoustic features in the frequency dimension. S2: Based on the enhanced features, perform deep feature extraction to obtain a high-dimensional acoustic representation; S3: Based on the high-dimensional acoustic representation, it is input into the pre-constructed decoupled branch structure for forward propagation calculation to obtain the joint prediction result of acoustic event activity and sound source direction, as well as the predicted value of distance-related parameters; the decoupled branch structure includes a joint estimation branch of acoustic event activity and sound source direction and a sound source distance-related parameter estimation branch; S4: During the model training phase of the decoupled branch structure, the corresponding confidence information is extracted based on the output of the joint estimation branch of sound event activity and sound source direction, and a confidence mask is constructed. The confidence mask is then used to constrain the loss calculation or gradient backpropagation range of the sound source distance related parameter estimation branch. S5: Input the audio features to be processed into the model optimized by S4, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure reasoning; and perform post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.

[0025] In some specific embodiments, such as Figure 3S1 mainly utilizes the Squeeze-and-Excitation mechanism to model the frequency axis. Specifically, feature descriptors in the frequency dimension are extracted through global pooling, and a frequency domain weight vector is calculated through a nonlinear bottleneck mapping structure. This weight vector is then applied to the original acoustic feature tensor to achieve adaptive recalibration of different frequency band features of the input acoustic features. Based on the aforementioned core design principles, step S1 specifically includes the following steps: S11. Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed. The original time-frequency acoustic features have dimensions [B,C,T,F], where B, C, T, and F represent the batch size, number of channels, number of time frames, and frequency dimension, respectively. The multi-channel audio signal to be processed is obtained through the audio input terminal 101. The original time-frequency acoustic features corresponding to the multi-channel audio are obtained through the feature extraction unit 102. Preferably, the time-frequency acoustic features include the time-frequency energy information of the sound event and multi-channel spatial information. For example, the original time-frequency acoustic features may include acoustic features such as the log-mel spectrum and the intensity vector (IV). The log-mel spectrum is used to characterize the energy distribution of the audio signal in the time and frequency dimensions, and the intensity vector is used to provide multi-channel spatial clues. For example, the intensity vector IV is used to characterize the direction of sound field energy propagation and spatial information related to the direction of the sound source, providing spatial clues for sound source direction estimation. S12. The original time-frequency acoustic features are statistically compressed in the time and channel dimensions to obtain a feature description vector in the frequency dimension. The feature channel dimension includes at least an acoustic energy representation channel and a spatial cue representation channel. The spatial cue representation channel is used to represent spatial information related to the sound source direction, such as spatial orientation cues (sound intensity vector IV). The statistical compression operation refers to using a squeeze operation, i.e., performing global time-domain average pooling on the original tensor—the original time-frequency acoustic features—using the global time-domain average pooling unit 202. This unit is used to perform statistical compression in the time dimension T and channel dimension C to obtain the feature description vector V in the frequency dimension. f This step is used to extract global response features of different frequency bands, providing input for subsequent frequency domain weight generation; S13. A frequency domain saliency weight (S) is generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weight to obtain the recalibrated enhanced features. The enhanced features can adaptively recalibrate the original time-frequency acoustic features in the frequency dimension, that is, adaptively recalibrate features in different frequency bands. The frequency domain saliency weight can be applied to the original time-frequency acoustic features to obtain the recalibrated enhanced features (X'=X). S). This step enables the frequency domain saliency adaptive module to adaptively adjust the participation of different frequency band features in subsequent tasks based on the training data, thereby improving the flexibility and scene adaptability of frequency domain feature representation for sound event detection, direction estimation, and distance-related parameter estimation. The process of generating frequency domain saliency weights (S) through nonlinear mapping is the feature transformation process (203-205): the feature description vector V... f The input is a bottleneck mapping structure consisting of two fully connected layers. First, the frequency dimension is compressed to 1 / 4 of the original frequency dimension through a dimension reduction fully connected layer 203; then, nonlinear expression capability is introduced through a nonlinear activation layer 204; finally, the feature is restored to the original frequency dimension F through a dimension increase fully connected layer 205. The process of obtaining the recalibrated enhanced features based on the frequency domain saliency weights is as follows: the features processed by the bottleneck mapping structure are input into a Sigmoid mapping unit 206, generating a frequency domain weight vector with values ​​ranging from 0 to 1, which represents the relative participation of different frequency band features in subsequent tasks; then, using an element-wise multiplication operator 207, the frequency domain weight vector is applied to the original time-frequency acoustic feature tensor 201 through a broadcasting mechanism to obtain the recalibrated time-frequency acoustic features 208.

[0026] In some specific embodiments, in S2, the recalibrated features, i.e. the enhanced features, are subjected to deep acoustic extraction via a backbone network or a shared deep feature encoding backbone network 104 to obtain a high-dimensional acoustic representation. This representation is used to model the time-frequency patterns, inter-channel spatial relationships, and temporal context information in the input features. This step obtains a high-dimensional acoustic representation that simultaneously contains time-frequency information, inter-channel spatial information, and temporal context information, which then serves as the shared input for subsequent multi-task decoupling branches. Preferably, the deep feature encoding backbone network can specifically adopt various neural network architectures or combinations thereof to extract different dimensional features from the input data. Preferably, convolutional networks, such as Convolutional Neural Networks (CNNs) or Residual Networks (ResNets), can be used; temporal modeling networks, such as Recurrent Neural Networks (RNNs) or Long Short-Term Memory Networks (LSTMs), can also be used; attention mechanism networks, such as Transformer network structures, or Hierarchical Token-Semantic Audio Transformers (HTS-AT) based on Transformers, can also be used; and it is not limited to the above single types of networks, but can also adopt a hybrid architecture of the aforementioned network structures, such as Convolutional Recurrent Neural Networks (CRNNs), which achieve joint extraction of spatial and temporal features through the cascaded combination of convolutional layers and recurrent layers. Those skilled in the art can adaptively configure the number of network layers, convolutional kernel size, and number of attention heads as needed.

[0027] In some specific embodiments, in S3, forward propagation calculation of shared deep acoustic features (i.e., high-dimensional acoustic representation) is performed through a decoupled branch structure (which is a network structure of multiple parallel task branches located at the back end of the shared deep feature encoder). The decoupled branch structure is a branch for joint estimation of acoustic event activity and sound source direction and a branch for estimation of sound source distance related parameters set in parallel. The specific processing of each branch is as follows: The joint estimation branch of acoustic event activity and acoustic source direction is used to output a direction prediction vector that couples activity state and direction parameters for each time frame and acoustic event category. The magnitude of the direction prediction vector is used to characterize the activity confidence of the corresponding acoustic event category, and the normalized direction prediction vector is used to characterize the acoustic source direction parameters. The sound source distance-related parameter estimation branch is used to output distance parameters related to the radial position of the sound source.

[0028] Through the above-mentioned parallel decoupled branch structure, the sound event category, spatial direction and distance-related information are independently modeled, thereby effectively reducing the mutual interference between different tasks in the feature learning process and improving the generalization ability of the overall system. The decoupling refers to decomposing the sound source localization task into a joint task of sound event activity detection and sound source direction estimation and an independent task of sound source distance estimation, and realizing the mutual independence of feature learning and loss calculation of each task through the branch network structure.

[0029] Preferably, in this embodiment, the branch refers to a parallel computing path with independent weight parameters constructed based on a neural network. Each branch is structurally decoupled from the others, and the dimensions and physical meanings of the output tensors are different. For example, the parallel branches can introduce shared features into a decoupled multi-task processing region (e.g., reference numeral 106) through a bifurcation point (e.g., reference numeral 105). This region includes at least a sound source distance-related parameter estimation branch 107 and a distance-related parameter estimation branch 108, thereby obtaining the corresponding three-dimensional sound source coordinate output 109. Furthermore, a sound event activity detection branch can be added according to actual needs to flexibly adapt to different application scenarios.

[0030] In some specific embodiments, such as Figure 4 In the model training phase of the decoupled branch structure, confidence information is extracted from the output of the joint estimation branch of acoustic event activity and sound source direction, and a confidence mask is constructed. This confidence mask is then used to constrain the loss calculation or gradient backpropagation range of the sound source distance-related parameter estimation branch. In other words, this invention introduces a training constraint mechanism based on a confidence mask. This mechanism uses the confidence score of the joint estimation branch of acoustic event activity and sound source direction as a control signal to construct a mask corresponding to a time frame or event category. When the confidence score is higher than a preset value, the location (time frame or event category) is considered a valid acoustic event region, allowing the distance-related parameter estimation branch to participate in the loss calculation. When the confidence score is lower than a preset threshold, the distance estimation loss at that location is reduced or masked. This approach reduces the impact of silent segments, low-confidence regions, or invalid acoustic segments on the training of the distance-related parameter estimation branch, helping to alleviate task interference during multi-task joint optimization. The confidence information refers to the vector representation of the joint estimation branch of sound event activity and sound source direction, which uses the coupling of activity state and direction parameters. The magnitude of the direction prediction vector is used to characterize the confidence of sound event activity, and the normalized direction prediction vector is used to characterize the sound source direction parameters.

[0031] The confidence level is a scalar reliability used to characterize whether a valid sound event exists in the corresponding time frame and sound event category. When the activity confidence level is greater than or equal to a preset threshold, the mask value is 1; when the activity confidence level is less than the preset threshold, the mask value is 0. The confidence mask is used to constrain the loss calculation or gradient backpropagation range of the distance-related parameter estimation branch, so that this branch mainly participates in training within the valid sound event region. Preferably, when the sound source direction estimation branch adopts a vector representation that couples the activity state with the direction parameter, the sound source direction estimation branch is for time frames. t Harmony Event Category c Output direction prediction vector

[0032] in, Represents time frame t Harmony Event Category c The corresponding joint prediction vector of activity state and direction parameters; , and These represent the predicted components of the direction prediction vector on the three spatial coordinate axes, respectively.

[0033] The activity confidence level at the corresponding position is obtained based on the magnitude of the predicted direction vector.

[0034] in, Represents time frame t Harmony Event Category c Corresponding activity confidence level; The L2 norm of the vector is used to represent the activity confidence level. The activity confidence level reflects whether a valid sound event exists at the corresponding time frame and sound event category. When a valid sound event is determined to exist, the unit direction vector of the corresponding sound source can be obtained by normalizing the direction prediction vector.

[0035] in, Represents time frame t Harmony Event Category c The corresponding normalized sound source direction vector, This is a preset positive number used to avoid a denominator of zero. Therefore, the magnitude of the direction prediction vector is used to characterize the confidence level of the sound event, while the normalized direction vector is used to characterize the sound source direction parameters. Both originate from the same prediction output but have different physical meanings.

[0036] Furthermore, let For time frames t Harmony Event Category cThe corresponding binary confidence mask. The confidence mask takes the value 0 or 1, and it is not used as the final sound source parameter output, but as the training control signal of the distance-related parameter estimation branch, used to determine whether the distance estimation error at the corresponding position participates in the loss calculation or gradient backpropagation.

[0037] The activity confidence level Compared with the preset confidence threshold Compare the results and generate a confidence mask according to the following formula:

[0038] when Greater than or equal to the preset confidence threshold At that time, the corresponding time frame and sound event category are determined as valid sound event regions, and the corresponding confidence mask is set. A value of 1 enables the distance-related parameter estimation error at that location to be included in loss calculation and model parameter update.

[0039] when Less than the preset confidence threshold When the corresponding location is identified as a silent region, a low-confidence region, or an invalid sound event region, the corresponding confidence mask is set accordingly. The value is set to 0, thereby reducing or masking the distance-related parameter estimation error at that location and minimizing its impact on the training process of the distance-related parameter estimation branch.

[0040] In summary, the confidence mask is used to weight or constrain the loss calculation or gradient backpropagation range of the distance-related parameter estimation branch, so that the distance-related parameter estimation branch mainly participates in training within the effective acoustic event region, and reduces the interference of silent regions, low-confidence regions or invalid acoustic segments on the learning of distance-related parameters, thereby improving the stability of the multi-task joint training process.

[0041] Preferably, in step S4, constraining the loss calculation or gradient backpropagation range of the sound source distance correlation parameter estimation branch using the confidence mask includes: The confidence mask is used to weight the estimated loss between the predicted value of the distance-related parameters and the true label. When the confidence is higher than a preset threshold, the distance estimation loss at that position is retained and used to participate in gradient backpropagation to update the model parameters. When the confidence is lower than the preset threshold, the distance estimation loss at that position is masked to block its gradient backpropagation path.

[0042] Preferably, a confidence mask is constructed based on the confidence information of the branch output jointly estimated by the acoustic event activity and the direction of the sound source. Specifically, when the confidence of the corresponding time or acoustic event category is higher than a preset threshold ( When the confidence level is lower than the preset value, the output is a valid mask, and the location is used as the valid acoustic event region so that the distance-related parameter estimation branch participates in the loss calculation; when the confidence level is lower than the preset value, the output is a valid mask. When a certain condition is met, an invalid mask is output; this mask reduces or blocks the distance estimation loss at that location. This ensures that distance-related parameter estimation primarily participates in training within the effective acoustic event region. In other words, this reduces the impact of silent segments, low-confidence regions, or invalid acoustic segments on the distance estimation branch training, thus improving the stability of multi-task joint training.

[0043] The processing flow based on the above design scheme is as follows: 1. Multiple data inputs (301-303): Obtain the confidence information 301 of the joint estimation branch of sound event activity and sound source direction, and the predicted value 302 of the distance correlation parameter estimation branch; at the same time, obtain the reference distance correlation parameters 303 corresponding to the training samples as a reference for calculating the distance correlation error. 2. Confidence Determination and Mask Generation (304-305): Confidence information 301 is input to the active probability discrimination unit 304. The system performs logical determination based on a preset threshold: when the confidence level is greater than or equal to the preset threshold, the mask generation module 305 outputs a valid mask Mask=1; when the confidence level is lower than the preset threshold, the mask generation module 305 outputs an invalid mask Mask=0. The mask is used to control whether the corresponding time frame or event category participates in the loss calculation of the distance-related parameter estimation branch.

[0044] 3. Raw Error Calculation (306-307): The predicted distance correlation parameter 302 and the reference distance correlation parameter 303 are input into the error calculation unit, which can also be called the distance calculation unit 306, and the raw distance correlation error 307 without masking is calculated. The error calculation unit uses a mean square error algorithm to calculate the error between the predicted value and the reference value; in other embodiments, the mean absolute error or relative error can also be used to calculate this error.

[0045] 4. Mask Weighting: The mask value output by the mask generation module 305 is multiplied element-wise by the original distance correlation error 307 using the multiplication operator 308 to obtain the weighted distance correlation loss. When the mask value is 1, the distance correlation error at the corresponding position participates in the training; when the mask value is 0, the distance correlation error at the corresponding position is reduced or masked. 5. Training and Update: The weighted distance correlation loss is input into the loss calculation and training control unit 309, and used together with other task losses to update the model parameters 310. Through this method, the distance correlation parameter estimation branch mainly participates in training within the effective acoustic event region, thereby helping to alleviate optimization interference between the distance correlation parameter estimation task and the acoustic event detection and direction estimation tasks, and improving the stability of the multi-task joint training process. In summary, this step proposes a training constraint protection strategy based on a confidence mask. It utilizes the confidence information from the joint estimation branch of acoustic event activity and sound source direction to construct a mask, and then uses this mask to weight or constrain the training of the distance correlation parameter estimation branch. This method ensures that the distance correlation parameter estimation branch primarily participates in training within the effective acoustic event region, thereby reducing the impact of silent segments or low-confidence regions on the learning of distance correlation parameters.

[0046] In some specific embodiments, during the inference phase, frame-by-frame neural network predictions may be affected by background noise, sound source overlap, or model uncertainty, often resulting in short-term false detections, discrete jumps, or local fluctuations in the original output sequence. To improve these issues, this invention performs post-processing on the original results output from the decoupled multi-task processing steps (results in sound event activity probability, direction estimation, and distance-related parameter estimation) to improve the temporal continuity and practical usability of the 3D sound source perception results; specifically, such as... Figure 5 In S5, the post-processing steps include category-related threshold discrimination and temporal sliding window smoothing. The category-related threshold discrimination involves: based on the output distribution of different sound event categories, using a preset threshold or a category-related threshold matrix to binarize and discriminate the activity probability of sound events, filtering out low-confidence predictions. For example, based on the category probability output by the sound event activity detection branch, different categories are discriminated using a preset threshold or a category-related threshold. Simultaneously, combining the sound source direction estimation result and the distance-related parameter estimation result, a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information is generated. The temporal sliding window smoothing involves: using a sliding window of preset length to smooth the sound event activity state, sound source direction parameters, and distance-related parameters on consecutive time frames, reducing short-term fluctuations and discrete jumps in frame-by-frame prediction. Through this step, a more temporally continuous and spatially stable three-dimensional sound source perception result is obtained.

[0047] The specific processing procedure is as follows: (1) Acquisition of raw output: Acquire the raw model output sequence 401 output from the decoupled multi-task processing region. The raw model output sequence includes at least one of the following: sound event activity probability, sound source direction parameter, and distance-related parameter. At this time, since frame-by-frame neural network prediction may be affected by background noise, sound source repetition, or model uncertainty, there may be short-term false detections, discrete jumps, or local fluctuations in the raw output sequence. (2) Category-related threshold discrimination: The original output sequence is input into a predefined category-related threshold discrimination unit 402. This unit uses a preset threshold or a category-related threshold matrix to discriminate the activity status of the sound events based on the output distribution of different sound event categories, and obtains the binarized sound event activity results. This step is used to filter out low-confidence predictions and reduce the impact of false detections on the final results.

[0048] (3) Temporal smoothing: The discriminated output is input to a predefined temporal sliding window smoothing unit 403. This unit uses a sliding window of a preset length (e.g., the window length is set to 20 frames) to smooth the acoustic event activity state, sound source direction parameters, and distance-related parameters on consecutive time frames. Through this step, short-term fluctuations and discrete jumps in frame-by-frame prediction can be reduced, making the output results more continuous in the temporal dimension.

[0049] (4) 3D parameter generation: Output the post-processed model results and generate 3D sound source parameter estimation results 404, which include the sound event activity state, orientation information, and distance-related parameters. This result can more stably reflect the spatial change trend of the sound source in continuous time frames and can be directly used for subsequent target tracking, spatial interaction, or sound field analysis.

[0050] Preferably, the post-processing step may further include one or more processing methods such as short-term false detection filtering or trajectory continuity constraints, in order to further suppress outlier jumps and improve the smoothness and stability of the output results.

[0051] In addition, based on the above design scheme, specific simulation experimental cases and results are given, the details of which are as follows: (1) Experimental setup and parameters: This embodiment is based on the public dataset STARSS23 to verify the effectiveness of the three-dimensional sound source localization scheme based on frequency domain saliency adaptation and multi-task decoupling described in this invention.

[0052] Audio feature input (input features): Extract 64-dimensional Log-mel spectral features (n - (mels=64), and combined with the sound intensity vector IV as the model input.

[0053] Frequency domain saliency adaptive module: Set the frequency dimension compression factor reduction=4 to generate frequency domain weights and recalibrate the input features; Training strategy: Employ a training constraint mechanism based on confidence masks and set an activity probability threshold. =0.5, used to constrain the distance-related parameter estimation branch to participate in training mainly within the effective acoustic event region.

[0054] Post-processing parameters: The category-related threshold matrix is ​​used to determine the activity state of sound events, and the time-domain smoothing window length is set to window=20 to reduce short-term fluctuations in continuous frame output.

[0055] (2) Performance evaluation comparison: The performance of this invention is compared with that of current benchmark models. Evaluation metrics include: positioning error (LE). CD (Unit: degrees) and location-related acoustic event detection F-score (F) 20° A true positive is only counted when the predicted event category is correct and the angular error between the predicted direction and the true direction does not exceed 20°. Location recall (LR) CD ), comprehensive SELD index ( SELD This invention also introduces the Event Conditional Relative Distance Error (EC-RDE). EC-RDE measures the relative deviation between the predicted distance-related parameters and the reference distance under valid acoustic event conditions; a smaller value indicates a more accurate distance estimation. Experimental results comparing the performance of the method of this invention with the benchmark model are shown in Table 1.

[0056] Table 1 shows the experimental results comparing the performance of the method of the present invention with that of the benchmark model.

[0057] (3) Experimental conclusions: Table 1 presents the performance comparison results of the proposed method and the benchmark model on the STARSS23 dataset. Experimental results show that the proposed method exhibits better overall performance in acoustic event detection, spatial localization, and distance perception. Compared to PSELENets, the proposed method demonstrates better performance in ER... 20° F 20° LE CD LR CD and SELD Improvements were achieved in all five key indicators. Among them, F 20° Increased to 57.1%, LR CDThe accuracy was increased to 69.8%, and the overall SELD index was reduced to 0.325, indicating that the method of the present invention can effectively improve the accuracy of acoustic event detection and spatial positioning recall capability.

[0058] Compared with Xue et al., the method of this invention has advantages in F 20° and LR CD The optimal results of 57.1% and 69.8% were achieved respectively, indicating that the method of the present invention has stronger ability to identify and locate effective sound events. Although the ER of the method of the present invention is... 20°、 LE CD and SELD Slightly higher than Xue et al., but the overall SELD index difference is only 0.004, indicating that the method of this invention still has strong competitiveness in terms of overall SELD performance and angle positioning accuracy.

[0059] Furthermore, this invention introduces the Event-Conditioned Relative Distance Error (EC-RDE) index to evaluate the estimation error of distance-related parameters under valid acoustic event conditions. Experimental results show that the EC-RDE of this invention is 0.2909. This result indicates that the proposed distance-related parameter estimation branch can further learn acoustic information related to the radial position of the sound source, building upon the completion of acoustic event detection and direction estimation. Since the comparison method does not report the EC-RDE index, a direct numerical comparison on this index is not possible. However, this result reflects, from a supplementary evaluation perspective, that this invention's method possesses distance-related sensing capabilities beyond the traditional SELD method.

[0060] Therefore, the experimental results validate the effectiveness of the method of this invention to a certain extent. This method not only improves the performance of acoustic event detection and localization recall, but also introduces distance-related acoustic cues, enabling the model to obtain richer spatial representation capabilities in complex real-world acoustic scenarios. Considering all indicators, this method achieves a good balance between event detection performance, localization recall capability, and distance-related parameter estimation capability, making it suitable for acoustic event localization and detection tasks in complex acoustic environments.

[0061] Meanwhile, the corresponding simulation results also demonstrate the advantages of this step, such as... Figure 6As shown, this paper compares the traditional direction positioning function with the three-dimensional sound source positioning function of this invention. The left figure (a) is a schematic diagram of the traditional direction positioning result. Traditional sound event localization and detection methods usually mainly output sound source direction information, such as azimuth angle, pitch angle, or direction vector on a unit sphere. When two sound sources are located in close proximity or in the same direction but at different distances, a system that relies solely on direction information has difficulty distinguishing the difference in their radial positions.

[0062] Figure (b) on the right is a schematic diagram of the three-dimensional sound source localization result of the present invention. Based on the output sound source direction parameters, the present invention introduces a distance-related parameter estimation branch, enabling the system to obtain sound source direction information and distance-related information. Based on azimuth, elevation, and distance-related parameters, the system can generate three-dimensional sound source parameter estimation results containing radial information. Compared with traditional methods that only output direction information, the present invention can provide richer information for expressing the spatial relationship of sound sources, helping to improve the three-dimensional spatial perception capability in complex acoustic scenes.

[0063] Based on the same inventive concept, this invention also proposes a three-dimensional sound source localization system based on frequency domain saliency adaptation and multi-task decoupling, comprising: The frequency domain saliency adaptive module is connected to the audio feature input terminal and is used to obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed. The original time-frequency acoustic features are statistically compressed in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension. Frequency domain saliency weights are generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weights to obtain the recalibrated enhanced features. The deep feature encoding module, which is connected to the frequency domain saliency adaptive module, is used to perform deep feature extraction based on the enhanced features to obtain a high-dimensional acoustic representation. The multi-task decoupling processing module is connected to the deep feature encoding module, and includes a joint estimation branch for acoustic event activity and sound source direction and a sound source distance-related parameter estimation branch set in parallel; the joint estimation branch for acoustic event activity and sound source direction is used to output the direction prediction vector coupled with the activity state and direction parameter, and the sound source distance-related parameter estimation branch is used to output the distance-related parameter. The confidence constraint training module is connected to the multi-task decoupling processing module. It is used to obtain the active confidence based on the magnitude of the direction prediction vector, construct a confidence mask based on the active confidence, and use the confidence mask to constrain the loss calculation or gradient backpropagation range of the sound source distance related parameter estimation branch. The output post-processing module is connected to the confidence constraint training module. It is used to input the audio features to be processed into the model optimized by the confidence constraint training module, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure inference. The module then performs post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.

[0064] Based on the same inventive concept, the present invention also proposes a computer-readable storage medium including computer instructions that, when executed on a computer, cause the computer to perform the method described thereon.

[0065] Based on the same inventive concept, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the steps of the method when executing the program.

[0066] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling, characterized in that, include: S1: Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed, perform statistical compression on the original time-frequency acoustic features in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension, generate frequency domain saliency weights through nonlinear mapping, and use the frequency domain saliency weights to weight the original time-frequency acoustic features to obtain the recalibrated enhanced features, which are used to adaptively recalibrate the original time-frequency acoustic features in the frequency dimension; S2: Based on the enhanced features, perform deep feature extraction to obtain a high-dimensional acoustic representation; S3: Based on the high-dimensional acoustic representation, it is input into the pre-constructed decoupled branch structure for forward propagation calculation to obtain the joint prediction results of acoustic event activity and sound source direction, as well as the predicted values ​​of distance-related parameters. The decoupled branch structure includes a joint estimation branch for acoustic event activity and acoustic source direction and a branch for estimation of acoustic source distance-related parameters. S4: During the model training phase of the decoupled branch structure, the corresponding confidence information is extracted based on the output of the joint estimation branch of sound event activity and sound source direction, and a confidence mask is constructed. The confidence mask is then used to constrain the loss calculation or gradient backpropagation range of the sound source distance related parameter estimation branch. S5: Input the audio features to be processed into the model optimized by S4, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure reasoning; and perform post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.

2. The three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling according to claim 1, characterized in that, S1 specifically includes the following steps: S11. Obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed; S12. The original time-frequency acoustic features are statistically compressed in the time dimension and feature channel dimension to obtain a feature description vector in the frequency dimension. The feature channel dimension includes at least an acoustic energy characterization channel and a spatial cue characterization channel. The spatial cue characterization channel is used to characterize spatial information related to the direction of the sound source. S13. Frequency domain saliency weights are generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weights to obtain the recalibrated enhanced features; wherein, the frequency domain saliency weights are generated through a bottleneck mapping structure, which consists of a dimension-reducing fully connected layer, a nonlinear activation layer, and an dimension-increasing fully connected layer; the dimension-reducing fully connected layer is used to compress the frequency dimension of the feature description vector; the nonlinear activation layer is used to perform nonlinear mapping on the compressed features to introduce nonlinear expressions; the dimension-increasing fully connected layer is used to restore the features with introduced nonlinear expressions to the original frequency dimension.

3. The three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling according to claim 1, characterized in that, S3 specifically includes the following steps: The specific processing steps for each branch of the decoupled branch structure are as follows: The joint estimation branch of acoustic event activity and acoustic source direction is used to output a direction prediction vector that couples activity state and direction parameters for each time frame and acoustic event category. The magnitude of the direction prediction vector is used to characterize the activity confidence of the corresponding acoustic event category, and the normalized direction prediction vector is used to characterize the acoustic source direction parameters. The sound source distance-related parameter estimation branch is used to output distance parameters related to the radial position of the sound source.

4. The three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling according to claim 1, characterized in that, In step S4, constraining the loss calculation or gradient backpropagation range of the sound source distance correlation parameter estimation branch using the confidence mask includes: The estimated loss between the predicted values ​​of the distance-related parameters and the true labels is weighted using the confidence mask. When the confidence level of the corresponding time frame or sound event category is higher than the preset threshold, the distance estimation loss corresponding to that time frame or sound event category is retained and used in gradient backpropagation to update the model parameters. When the confidence level of the corresponding time frame or sound event category is lower than or equal to the preset threshold, the distance estimation loss corresponding to that time frame or sound event category is masked to block its gradient backpropagation path.

5. The three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling according to claim 1, characterized in that, In S5, the post-processing includes using category-related thresholds to determine the activity state for different sound event categories, and performing time-domain sliding window smoothing on the direction parameters and distance-related parameters on consecutive time frames.

6. A system for implementing the three-dimensional sound source localization method based on frequency domain saliency adaptation and multi-task decoupling as described in any one of claims 1-5, characterized in that, include: The frequency domain saliency adaptive module is connected to the audio feature input terminal and is used to obtain the original time-frequency acoustic features corresponding to the multi-channel audio signal to be processed. The original time-frequency acoustic features are statistically compressed in the time dimension and channel dimension to obtain the feature description vector in the frequency dimension. Frequency domain saliency weights are generated through nonlinear mapping, and the original time-frequency acoustic features are weighted using the frequency domain saliency weights to obtain the recalibrated enhanced features. A deep feature encoding module, which is connected to the frequency domain saliency adaptive module, is used to perform deep feature extraction based on the enhanced features to obtain a high-dimensional acoustic representation. The multi-task decoupling processing module is connected to the deep feature encoding module, including a joint estimation branch for acoustic event activity and sound source direction and a sound source distance-related parameter estimation branch set in parallel. The joint estimation branch for acoustic event activity and sound source direction is used to output a direction prediction vector coupled with the activity state and direction parameters. The confidence constraint training module is connected to the multi-task decoupling processing module and is used to obtain the activity confidence based on the magnitude of the direction prediction vector, construct a confidence mask based on the activity confidence, and use the confidence mask to constrain the loss calculation or gradient backpropagation range of the sound source distance-related parameter estimation branch. The output post-processing module is connected to the confidence constraint training module. It is used to input the audio features to be processed into the model optimized by the confidence constraint training module, and obtain the corresponding sound event activity state, sound source direction parameters and distance-related parameters through the decoupled branch structure inference. The module then performs post-processing on the above parameters to generate a three-dimensional sound source parameter estimation result containing sound source direction information and distance-related information.