A three-dimensional meshless rotating sound source positioning method and system based on a sparse attention transformer neural network

CN122836662APending Publication Date: 2026-09-29ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610956699.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]然而,现有方法仍存在一定局限性:一方面,多数声源定位方法基于预定义空间网格进行离散化建模,受限于网格划分精度,难以实现更高分辨率的连续空间定位,且通常仅能获得二维平面位置结果,难以直接获取声源的三维空间坐标;另一方面,传统方法在复杂旋转工况下仍面临计算复杂度较高、对噪声敏感以及对强非平稳声场适应性不足等问题

Benefits of technology

本发明提出了一种基于稀疏注意力Transformer神经网络的三维无网格旋转声源定位方法及系统,将传统声学建模方法与深度学习全局特征建模能力相结合,在保证计算效率的同时显著提升了旋转声源三维空间定位的精度与鲁棒性。通过采用基于模态分解的旋转波束形成算法,本发明在模型训练前对旋转声场进行物理约束建模,有效增强了声场数据的空间表达能力,为后续网络学习提供了更加稳定且具有物理意义的输入特征。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122836662A_ABST
    Figure CN122836662A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of rotating sound source identification, and particularly relates to a three-dimensional meshless rotating sound source positioning method and system based on a sparse attention Transformer neural network. The method comprises collecting sound pressure signals of a rotating sound source field to construct a sound pressure cross-spectrum matrix, and generating a three-dimensional rotating sound source distribution map based on a rotating beam forming algorithm of modal decomposition; inputting the three-dimensional rotating sound source distribution map into a SATNN model to obtain predicted three-dimensional coordinates of the sound source; and determining the spatial position of the rotating sound source based on the predicted three-dimensional coordinates of the sound source. Thanks to the meshless strategy, the present application effectively breaks through the resolution limitation caused by the dependence of traditional rotating sound source imaging on a fixed spatial mesh, so that the positioning result is significantly improved in terms of spatial accuracy and continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of rotating sound source recognition technology, specifically relating to a three-dimensional meshless rotating sound source localization method and system based on a sparse attention Transformer neural network. Background Technology

[0002] Rotating machinery is widely used in aerospace, wind power generation, and transportation. The noise generated during its operation significantly impacts equipment performance and environmental comfort, making the analysis and control of its noise characteristics of great engineering significance. Due to the significant Doppler effect and spatial modulation effect accompanying the motion of rotating sound sources, the array-received signal exhibits obvious non-stationary characteristics. Its radiated sound field is characterized by strong time-varying properties, complex spatial distribution, and significant phase modulation, posing a considerable challenge to high-precision sound source localization. Currently, rotating sound source localization methods are mainly based on array signal processing techniques. For example, by constructing a sound pressure cross-spectrum matrix and introducing methods such as rotating reference frame transformation, Doppler effect compensation, and mode decomposition, the rotating sound field is modeled and imaged. Simultaneously, high-resolution deconvolution methods are also used to improve the image resolution of sound sources, thereby addressing the insufficient spatial resolution of traditional beamforming methods.

[0003] However, existing methods still have certain limitations: on the one hand, most sound source localization methods are based on discretized modeling using predefined spatial grids. Limited by the grid division accuracy, they struggle to achieve higher-resolution continuous spatial localization and typically only obtain two-dimensional planar position results, making it difficult to directly acquire the three-dimensional spatial coordinates of the sound source. On the other hand, traditional methods still face problems such as high computational complexity, sensitivity to noise, and insufficient adaptability to strong non-stationary sound fields under complex rotating conditions. In recent years, with the development of deep learning methods, the mapping relationship between sound field features and sound source positions has improved localization efficiency and robustness to some extent. However, their output is still mostly based on gridded representations, making it difficult to break through the constraints of discrete grids and achieve truly gridless continuous three-dimensional spatial localization. Therefore, there is an urgent need to develop a high-precision rotating sound source localization method that can integrate the physical characteristics of rotating sound fields with deep feature modeling capabilities, while simultaneously achieving direct regression of three-dimensional continuous coordinates. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention provides a three-dimensional meshless rotating sound source localization method and system based on a sparse attention Transformer neural network. The aim is to overcome the resolution limitations imposed by traditional rotating sound source imaging, which relies on a fixed spatial grid, and to improve the spatial accuracy and continuity of the localization results.

[0005] To achieve the above objectives, the present invention provides the following solution: A method for localizing a 3D meshless rotating sound source based on a sparse attention Transformer neural network, the method comprising: The sound pressure signals of the rotating sound source field are collected to construct the sound pressure cross spectrum matrix, and a three-dimensional rotating sound source distribution map is generated based on the rotating beamforming algorithm of mode decomposition. The three-dimensional rotating sound source distribution map is input into the SATNN model to obtain the predicted three-dimensional coordinates of the sound sources; Based on the predicted three-dimensional coordinates of the sound source, the spatial location of the rotating sound source is determined.

[0006] Preferably, the method for acquiring sound pressure signals from a rotating sound source field to construct a sound pressure cross-spectrum matrix, and generating a three-dimensional rotating sound source distribution map based on a mode decomposition-based rotating beamforming algorithm includes: Several microphone sensors are arranged in the rotating sound source field to construct a measurement plane that is parallel to the rotating plane of the sound source and coaxial with its center. Within the measurement plane, the sound pressure cross-spectrum matrix is ​​obtained based on the sound pressure signal acquired by the microphone sensor; The rotating beamforming algorithm based on mode decomposition performs motion compensation and mode expansion on the sound pressure cross spectrum matrix to generate a three-dimensional rotating sound source distribution map.

[0007] Preferably, the method for generating a three-dimensional rotating sound source distribution map by performing motion compensation and mode expansion on the sound pressure cross-spectrum matrix using a mode decomposition-based rotating beamforming algorithm includes: Calculate the turning vector of the rotating sound source: ; in, Let Green's function be the acoustic propagation between the s-th grid point and the m-th microphone sensor in a rotating coordinate system in a free field. Its expression is: ; Where i is the imaginary unit, Here, n is the modal factor, and n is the order. It is the rotational angular frequency. The wavenumber after frequency shift caused by rotation effect. As the normalization factor, for Step Legendre functions of the first kind Let be the radial function in the Green's function expansion of the rotating coordinate system. and Coordinates; At grid points The output at this location is: ; Where C is the sound pressure cross spectrum matrix, and the superscript H indicates the conjugate transpose; Output Normalization is performed for acoustic imaging to generate a three-dimensional rotating sound source distribution map.

[0008] Preferably, the method for inputting the three-dimensional rotating sound source distribution map into the SATNN model to obtain the predicted three-dimensional coordinates of the sound source includes: The three-dimensional rotating sound source distribution map is input into the feature embedding module of the SATNN model to obtain a low-level feature representation containing local spatial information. The low-level feature representation is input into the encoder of the SATNN model to extract the deep feature information of the three-dimensional rotating sound source distribution map layer by layer. The deep feature information is input into the decoder of the SATNN model, and layer-by-layer upsampling and fusion are performed to reconstruct the sound source distribution map and obtain the sound source reconstruction features. The reconstructed sound source features are input into the feature compensation and mapping module of the SATNN model for adaptive compensation and mapping to obtain a low-dimensional feature representation. The low-dimensional feature representation is input into the three-dimensional coordinate regression module of the SATNN model to predict the coordinate parameters of the sound source in three-dimensional space, thereby obtaining the predicted three-dimensional coordinates of the sound source.

[0009] Preferably, the encoder of the SATNN model consists of four cascaded encoding modules, each encoding module including several sparse Transformer modules and a downsampling layer; The sparse Transformer module consists of a layer normalization module, a sparse attention module, and a feedforward network module. It is used for global feature modeling and multi-scale feature extraction of the input low-level feature representation. Specifically, the layer normalization module normalizes the input features; the sparse attention module constructs query, key, and value mappings and uses a sparse selection strategy to filter attention weights, achieving efficient modeling of global feature dependencies; and the feedforward network module employs a multi-scale deep convolutional structure to perform non-linear mapping and enhancement of features, achieving the fusion of local and global features. The downsampling layer uses a combination of convolution and pixel rearrangement to achieve spatial downsampling. The pixel rearrangement operation compresses the feature map, reducing spatial resolution while increasing channel dimension, thus obtaining a more hierarchical representation of sound source features.

[0010] Preferably, the sparse attention module includes: a feature mapping submodule, an attention calculation submodule, a sparse selection submodule, a multi-scale fusion submodule, and an output mapping submodule; The feature mapping submodule performs linear mapping on the input features through 1×1 convolution and depthwise separable convolution to generate queries, keys and values, which are used to characterize the spatial and channel information in the three-dimensional rotating sound source distribution map. The attention calculation submodule normalizes the query and key, calculates the similarity matrix, obtains the correlation distribution between global features, and establishes long-distance dependencies. The sparse selection submodule filters the attention matrix based on the Top-K strategy, retaining the most important relevance information under different sparsity ratios, while suppressing irrelevant or preset low contribution regions. The multi-scale fusion submodule performs weighted fusion of attention results obtained under different sparsity ratios and adaptively combines multi-scale attention information through learnable parameters. The output mapping submodule maps the fused features back to the original feature space through linear projection, thereby realizing the integration and reconstruction of feature information.

[0011] Preferably, the decoder of the SATNN model includes several upsampling modules and a feature fusion module; The upsampling module upsamples the input features by pixel rearrangement to obtain upsampled features; The feature fusion module is used to concatenate the upsampled features with the feature map output by the corresponding layer of the encoder, and then use the concatenated multi-scale fused features to perform feature reconstruction through the sparse Transformer module to obtain the sound source reconstruction features.

[0012] The present invention also provides a three-dimensional meshless rotating sound source localization system based on a sparse attention Transformer neural network. The system is used to implement the aforementioned method and includes: an acquisition module, a prediction module, and a localization module. The acquisition module is used to acquire the sound pressure signal of the rotating sound source field to construct the sound pressure cross spectrum matrix, and generate a three-dimensional rotating sound source distribution map based on the rotating beamforming algorithm of mode decomposition. The prediction module is used to input the three-dimensional rotating sound source distribution map into the SATNN model to obtain the predicted three-dimensional coordinates of the sound source; The positioning module is used to determine the spatial position of the rotating sound source based on the predicted three-dimensional coordinates of the sound source.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned method.

[0014] The present invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the aforementioned method.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a three-dimensional meshless rotating sound source localization method and system based on a sparse attention Transformer neural network. It combines traditional acoustic modeling methods with the global feature modeling capabilities of deep learning, significantly improving the accuracy and robustness of three-dimensional spatial localization of rotating sound sources while maintaining computational efficiency. By employing a rotating beamforming algorithm based on mode decomposition, this invention performs physical constraint modeling of the rotating sound field before model training, effectively enhancing the spatial representation capability of the sound field data and providing more stable and physically meaningful input features for subsequent network learning.

[0016] This invention combines a sparse attention Transformer neural network structure with a 3D rotating sound source distribution map as input to construct an encoder-decoder joint feature extraction framework, achieving step-by-step modeling and fusion of deep spatial features of the sound field. During feature extraction, a sparse attention mechanism is introduced to selectively model global dependencies, reducing computational complexity while enhancing the responsiveness of key sound source regions. Combined with a multi-scale feature fusion strategy, this improves the expressive power under complex rotating sound fields. Furthermore, a 3D coordinate regression head is introduced at the network endpoint to achieve a direct mapping from the sound field distribution map to the continuous 3D spatial coordinates of the sound sources, thus overcoming the limitations of traditional gridded localization methods.

[0017] The method of this invention exhibits superior performance in terms of three-dimensional spatial positioning accuracy, noise resistance, and computational efficiency, enabling stable and high-precision localization of rotating sound sources even under conditions of limited microphone array size and low signal-to-noise ratio. Furthermore, this method represents a technological leap from traditional two-dimensional grid imaging to three-dimensional gridless continuous coordinate regression, significantly improving the engineering applicability and robustness of rotating sound source localization, and providing a new technical approach for the identification of sound sources and the reconstruction of three-dimensional sound fields in complex rotating machinery. Attached Figure Description

[0018] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a three-dimensional meshless rotating sound source localization method based on a sparse attention Transformer neural network according to an embodiment of the present invention. Figure 2 This is a framework diagram of a three-dimensional meshless rotating sound source localization method based on a sparse attention Transformer neural network according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the specific structure of the model in an embodiment of the present invention; wherein, (a) is the overall architecture of the SATNN model; (b) is the expert hybrid feature compensator (MEFC), i.e., the feature compensation and mapping module; (c) is the Top-k sparse attention mechanism (TKSA), i.e., the sparse attention module; (d) is the hybrid scale feedforward network (MSFN), i.e., the feedforward network module; and (e) is the three-dimensional coordinate regression head (3D-CRH). Figure 4 The diagram shows the three-dimensional dual-sound source location results according to an embodiment of the present invention; wherein, (a) is the location result diagram at the dual-sound source positions [0,0.25, 0.05], [0, -0.25, 0.05], and f=1600 Hz, (b) is the location result diagram at the dual-sound source positions [0,0.25, 0.05], [0, -0.25, 0.05], and f=2500 Hz, (c) is the location result diagram at the dual-sound source positions [0,0.25, 0.05], [0, -0.25, 0.05], and f=4500 Hz, and (d) is the location result diagram at the dual-sound source positions [0,0.25, 0.05], [0, -0.25, 0.05], and f=6000 Hz. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Example 1 like Figure 1 As shown, this invention provides a three-dimensional meshless rotating sound source localization method based on a sparse attention Transformer neural network, comprising: The sound pressure signals of the rotating sound source field are collected to construct the sound pressure cross spectrum matrix, and a three-dimensional rotating sound source distribution map is generated based on the rotating beamforming algorithm of mode decomposition. The three-dimensional rotating sound source distribution map is input into the SATNN model to obtain the predicted three-dimensional coordinates of the sound sources; Based on the predicted three-dimensional coordinates of the sound source, the spatial location of the rotating sound source is determined.

[0023] like Figures 1-4 As shown, the specific implementation process of the present invention is as follows: Step 1: Collect the sound pressure signal of the rotating sound source field to construct the sound pressure cross spectrum matrix (CSM). Use the rotating beamforming algorithm of mode decomposition to perform motion compensation and mode expansion on the sound pressure cross spectrum matrix to generate a three-dimensional rotating sound source distribution map in the rotating scene.

[0024] Step 1 is as follows: Step 1.1: In a rotating sound source field formed by K sound sources, M microphone sensors are arranged, forming a planar array, defined as the measurement plane. The measurement plane is parallel to the rotating sound source plane, and its center is coaxial with the rotation center of the rotating sound source plane to ensure the uniformity and symmetry of the spatial sampling of the rotating sound source field by the measurement array. The distance between the two is z. The array of microphone sensors is arranged in a ring to acquire spatial sound pressure information.

[0025] Step 1.2: Divide the 3D sound source region into a grid, discretizing the continuous sound source distribution into several discrete sound source units, forming a 3D spatial grid composed of multiple grid points, defined as the 3D focusing domain. The 3D focusing domain contains N grid points, each of which is also a focal point. Each focal point serves as a potential sound source location for sound source energy estimation and reconstruction (see steps 1.3-1.6 for details). By calculating the steering vector corresponding to each focal point and combining it with array measurement data, the sound source energy is estimated, thereby obtaining the spatial energy distribution within the entire focusing domain and generating a 3D rotating sound source distribution map for a rotating scene. The sound source location is determined based on the energy magnitude of each grid point, thus achieving sound source localization.

[0026] Step 1.3: Receive the sound pressure signal X collected by the microphone sensor array. The sound pressure signal X includes the target sound source signal and the background noise signal. The collected time-domain sound pressure signal is segmented and windowed. A fast Fourier transform is used to convert the time-domain signal into a frequency-domain signal. The cross-power spectrum is calculated based on the frequency-domain sound pressure data corresponding to each microphone channel. By statistically averaging multiple time frames, the sound pressure cross-spectrum matrix C is obtained, which is used to characterize the spatial correlation between the sensors.

[0027] ; in, , This represents the sound pressure signal from the m-th microphone. H represents the conjugate transpose of the matrix.

[0028] Step 1.4: Calculate the turning vector of the rotating sound source.

[0029] ; in, Let Green's function be the acoustic propagation between the s-th grid point and the m-th microphone sensor in a rotating coordinate system in a free field. Its expression is: ; Where i is the imaginary unit. The range of the number of terms in an infinite series expansion can be approximated as follows, provided that the truncation error is small: , ,in , , , ,in, Modal factors The lower and upper cutoffs, The upper cutoff is for order n. For wave number, For frequency, For the speed of sound, Indicates the radius of rotation. This indicates rounding down to the nearest integer. It is the rotational angular frequency. and These are the spherical coordinate vectors of grid points in the measurement plane and scanning plane (i.e., the sound source rotation plane) formed by the microphone sensor, respectively. and The coordinates are .

[0030] The expression for the normalization factor is: ; The wavenumber after frequency shift caused by the rotation effect is: ; for Step Legendre functions of the second kind; radial functions in the expansion of Green's functions in a rotating coordinate system. Composed of spherical Bessel and spherical Hankel functions, it is used to describe the radial propagation characteristics of sound waves from the sound source to the microphone: ; ; in, for The first-order Bessel function of the sphere of the first kind yes Rank Ball-like Hankel function.

[0031] Step 1.5: At grid points The output at this location is: ; The superscript H indicates the conjugate transpose.

[0032] Step 1.6: Output Normalization is performed for acoustic imaging to obtain a three-dimensional rotating sound source distribution map.

[0033] Step 2: Input the 3D rotating sound source distribution map into the SATNN model, and obtain the predicted 3D coordinates of the rotating sound sources through network inference. SATNN uses the 3D rotating sound source distribution map as input and the real sound source coordinates as labels for supervised training to obtain model parameters. The SATNN model structure is as follows: Figure 3 As shown.

[0034] Step 2 is as follows: Step 2.1: Input the three-dimensional rotating sound source distribution map into the feature embedding module of the SATNN model. The feature embedding module adopts an overlapping image block embedding structure and performs local feature extraction and channel mapping on the input three-dimensional rotating sound source distribution map through a 3×3 convolution kernel. Initial feature encoding is achieved without changing the spatial size of the feature map, thereby obtaining a low-level feature representation containing local spatial information.

[0035] Step 2.2: Input the low-level feature representation into the encoder of the SATNN model to extract the deep feature information of the three-dimensional rotating sound source distribution map layer by layer. The encoder consists of four cascaded encoding modules, each of which includes multiple sparse Transformer modules and downsampling layers, used to extract the deep feature information of the three-dimensional rotating sound source distribution map layer by layer.

[0036] The feature encoding process of the sparse Transformer module can be represented as: ; ; Where LN represents layer normalization, Indicates the number of layers in the network. and These represent the outputs of the sparse attention module and the feedforward network module, respectively. TKSA and MSFN represent the operations of the sparse attention module and the feedforward network module, respectively.

[0037] The sparse Transformer module consists of a layer normalization module, a sparse attention module, and a feedforward network module. The layer normalization module is used to normalize the input features. The sparse attention module constructs query, key, and value mappings and uses a sparse selection strategy to filter attention weights, thereby achieving efficient modeling of global feature dependencies. The feedforward network module uses a multi-scale deep convolutional structure to perform nonlinear mapping and enhancement of features, thereby achieving the fusion of local and global features and improving the expressive power of multi-scale spatial features.

[0038] The downsampling layer uses a combination of convolution and pixel rearrangement to achieve spatial downsampling. The pixel rearrangement operation compresses the feature map to reduce spatial resolution while increasing channel dimension, thereby expanding the receptive field and reducing computational complexity to obtain a more hierarchical representation of sound source features.

[0039] Specifically, in this embodiment, each encoding module adopts a sparse Transformer module structure, including a layer normalization module, a sparse attention module, and a feedforward network module, which are used to perform global feature modeling and multi-scale feature extraction on the low-level feature representation of the input.

[0040] The sparse attention module constructs query, key, and value mappings and performs sparse selection on the attention matrix. This preserves important feature responses while suppressing redundant information, thereby achieving efficient modeling of long-distance dependencies and enhancing the global perception of the spatial distribution characteristics of rotating sound sources.

[0041] The feedforward network module employs a multi-scale deep convolutional structure for feature enhancement, including deep convolutional operations with different receptive fields, to extract local spatial features and achieve multi-scale information fusion, thereby improving the richness and robustness of feature representation.

[0042] Layer normalization is introduced between each submodule to standardize the feature maps, thereby reducing the differences in feature distribution, improving the stability of network training, and accelerating the convergence speed.

[0043] Specifically, in this embodiment, a sparse attention mechanism is introduced in the sparse Transformer module of each encoding module to adaptively weight features globally, thereby enhancing the expressive power of key sound source features.

[0044] The sparse attention module includes a feature mapping submodule, an attention computation submodule, a sparse selection submodule, a multi-scale fusion submodule, and an output mapping submodule.

[0045] The sparse attention computation process can be represented as: ; in, Scaling factor Represents the Top-k selection operator: in Representation matrix The Middle In the middle of the line Large elements, Represents the attention score matrix S The Middle i Line 1 j Column elements. This method transforms dense attention computation into sparse attention computation, thereby improving the efficiency of feature modeling.

[0046] The feature mapping submodule performs linear mapping on the input features through 1×1 convolution and depthwise separable convolution to generate query, key, and value features, which are used to characterize the spatial and channel information in the three-dimensional rotating sound source distribution map; The attention calculation submodule normalizes the query and key and calculates their similarity matrix to obtain the correlation distribution between global features, thereby establishing long-distance dependencies. The sparse selection submodule filters the attention matrix based on the Top-K strategy, retaining the most important relevance information under different sparsity ratios, while suppressing irrelevant or low-contribution regions, thereby reducing computational complexity and improving the effectiveness of feature selection. The multi-scale fusion submodule performs weighted fusion of attention results obtained under different sparsity ratios, and adaptively combines multi-scale attention information through learnable parameters, thereby taking into account both local details and global structural features. The output mapping submodule maps the fused features back to the original feature space through linear projection, thereby realizing the integration and reconstruction of feature information. By introducing the sparse attention mechanism mentioned above during the encoding process, the model can adaptively focus on the key regions of the rotating sound source in the global scope, effectively modeling long-distance dependencies in complex sound fields, while reducing computational complexity and improving the ability to express the features of the rotating sound source and its noise resistance performance.

[0047] Step 2.3: Input the deep feature information into the decoder of the SATNN model for layer-by-layer upsampling and fusion to reconstruct the sound source distribution map and obtain the sound source reconstruction features. The decoder includes multiple upsampling modules and feature fusion modules, used to restore the spatial resolution of the feature map layer by layer and reconstruct the sound source distribution features.

[0048] The calculation process of the feature fusion module is as follows: For the input feature map ( and (representing the height and width of the feature map, respectively). First, channel averaging is performed to generate a C-dimensional channel descriptor. : in Representation of feature map exist The value of the position. Then, based on the learnable weight matrix... and Assign a coefficient vector to each operator branch, where T is the weight dimension and O is the number of operator branches. To ensure that the input and output dimensions remain unchanged, the input feature map is zero-paddinged before calculating the features in each operator branch. Finally, the... The output of the layer feature fusion module is calculated as follows: in, Indicates the first Layer weight dimensions Represents the ReLU activation function. express convolution, This represents the characteristic operations of each operator branch.

[0049] Each upsampling module upsamples the input features by rearranging pixels to obtain upsampled features, thereby improving spatial resolution. After upsampling, the upsampled features are concatenated with the feature maps output by the corresponding layers in the encoder based on the feature fusion module to form multi-scale fused features. After multi-scale fusion features are dimensionally aligned via channel-adjusted convolution, they are input into a decoding unit composed of multiple sparse Transformer modules for feature reconstruction, obtaining sound source reconstruction features. These features are used to recover high-level semantic information and spatial detail features, achieving a fine representation of the distribution features of rotating sound sources. The sparse Transformer modules here are structurally identical to those in the encoder's encoding module.

[0050] Step 2.4: Input the sound source reconstruction features output from the last layer of the decoder into the feature compensation and mapping module for adaptive compensation and mapping to obtain a low-dimensional feature representation. The feature compensation and mapping module includes a feature compensation submodule and an output mapping submodule.

[0051] The feature compensation submodule adopts a feature enhancement structure based on hybrid operation adaptive weighting to adaptively compensate and optimize the decoded features to obtain high-dimensional feature representation, thereby further improving the expressive power of the features. The output mapping submodule maps the compensated high-dimensional feature representation to a low-dimensional feature representation through convolution operations, providing input features for subsequent sound source 3D coordinate regression.

[0052] Step 2.5: Input the output low-dimensional feature representation into the three-dimensional coordinate regression module (3D-CRH three-dimensional coordinate regression head). The three-dimensional coordinate regression module performs nonlinear transformation on the features through fully connected mapping to directly predict the coordinate parameters of the sound source in three-dimensional space and obtain the predicted three-dimensional coordinates of the sound source.

[0053] The three-dimensional coordinate regression module outputs the spatial coordinate information of the sound source, including the position parameters (x, y, z) of the sound source in three-dimensional space. The coordinates are continuous values ​​and are used to characterize the precise position of the sound source in the actual physical space.

[0054] The 3D coordinate regression module is connected to the decoder output and includes a global average pooling module and a multi-layer fully connected network, which are used to compress the 2D feature map into a 1D feature vector and map the feature vector into a 3D continuous coordinate output, respectively.

[0055] Step 3: Determine the spatial position of the rotating sound source based on the predicted three-dimensional coordinates of the sound source.

[0056] Step 3 specifically involves: Based on the three-dimensional coordinate information of the predicted three-dimensional coordinates of the sound source, the position of the rotating sound source in space can be directly determined without the need for grid-based peak search or coordinate mapping processing, thereby achieving gridless continuous spatial positioning of the rotating sound source.

[0057] This process, based on the end-to-end feature mapping capability of the SATNN model, directly realizes the regression prediction of the three-dimensional spatial coordinates of the sound source, thereby improving the accuracy and robustness of the location of the rotating sound source.

[0058] The three-dimensional dual-source localization results are as follows: Figure 4 As shown.

[0059] In summary, this invention provides a three-dimensional meshless rotating sound source localization method based on a sparse attention Transformer neural network. The method includes: firstly, performing motion compensation and modal expansion on the sound pressure cross-spectrum matrix using a rotating beamforming algorithm based on mode decomposition to generate a three-dimensional sound source distribution map in a rotating scene, and then compressing its channels to reduce data size and improve the efficiency of model training and inference; subsequently, inputting this map into a three-dimensional meshless SATNN model to directly predict the continuous coordinates of the sound source in three-dimensional space. Thanks to the meshless strategy, this method effectively overcomes the resolution limitations imposed by traditional rotating sound source imaging relying on a fixed spatial grid, resulting in a significant improvement in both spatial accuracy and continuity of the localization results.

[0060] Example 2 Based on the same inventive concept, the present invention also provides a three-dimensional meshless rotating sound source localization system based on a sparse attention Transformer neural network, for implementing the method described in the foregoing embodiments. The system includes: an acquisition module, a prediction module, and a localization module. The acquisition module is used to acquire the sound pressure signal of the rotating sound source field to construct the sound pressure cross spectrum matrix, and generate a three-dimensional rotating sound source distribution map based on the rotating beamforming algorithm of mode decomposition. The prediction module is used to input the three-dimensional rotating sound source distribution map into the SATNN model to obtain the predicted three-dimensional coordinates of the sound source; The positioning module is used to determine the spatial position of the rotating sound source based on the predicted three-dimensional coordinates of the sound source.

[0061] Example 3 The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in the foregoing embodiments.

[0062] Example 4 The present invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the methods described in the foregoing embodiments.

[0063] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A three-dimensional meshless rotating sound source localization method based on a sparse attention Transformer neural network, characterized in that, The method includes: The sound pressure signals of the rotating sound source field are collected to construct the sound pressure cross spectrum matrix, and a three-dimensional rotating sound source distribution map is generated based on the rotating beamforming algorithm of mode decomposition. The three-dimensional rotating sound source distribution map is input into the SATNN model to obtain the predicted three-dimensional coordinates of the sound sources; Based on the predicted three-dimensional coordinates of the sound source, the spatial location of the rotating sound source is determined.

2. The method according to claim 1, characterized in that, Methods for acquiring sound pressure signals from a rotating sound source field to construct a sound pressure cross-spectrum matrix, and generating a three-dimensional rotating sound source distribution map based on a mode decomposition-based rotating beamforming algorithm, include: Several microphone sensors are arranged in the rotating sound source field to construct a measurement plane that is parallel to the rotating plane of the sound source and coaxial with its center. Within the measurement plane, the sound pressure cross-spectrum matrix is ​​obtained based on the sound pressure signal acquired by the microphone sensor; The rotating beamforming algorithm based on mode decomposition performs motion compensation and mode expansion on the sound pressure cross spectrum matrix to generate a three-dimensional rotating sound source distribution map.

3. The method according to claim 2, characterized in that, The method for generating a three-dimensional rotating sound source distribution map by performing motion compensation and mode expansion on the sound pressure cross-spectrum matrix based on mode decomposition rotating beamforming algorithm includes: Calculate the turning vector of the rotating sound source: ; in, Let Green's function be the acoustic propagation between the s-th grid point and the m-th microphone sensor in a rotating coordinate system in a free field. Its expression is: ; Where i is the imaginary unit, Here, n is the modal factor, and n is the order. It is the rotational angular frequency. The wavenumber after frequency shift caused by rotation effect. As the normalization factor, for Step Legendre functions of the first kind Let be the radial function in the Green's function expansion of the rotating coordinate system. and Coordinates; At grid points The output at this location is: ; Where C is the sound pressure cross spectrum matrix, and the superscript H indicates the conjugate transpose; Output Normalization is performed for acoustic imaging to generate a three-dimensional rotating sound source distribution map.

4. The method according to claim 1, characterized in that, The method for inputting the three-dimensional rotating sound source distribution map into the SATNN model to obtain the predicted three-dimensional coordinates of the sound sources includes: The three-dimensional rotating sound source distribution map is input into the feature embedding module of the SATNN model to obtain a low-level feature representation containing local spatial information. The low-level feature representation is input into the encoder of the SATNN model to extract the deep feature information of the three-dimensional rotating sound source distribution map layer by layer. The deep feature information is input into the decoder of the SATNN model, and layer-by-layer upsampling and fusion are performed to reconstruct the sound source distribution map and obtain the sound source reconstruction features. The reconstructed sound source features are input into the feature compensation and mapping module of the SATNN model for adaptive compensation and mapping to obtain a low-dimensional feature representation. The low-dimensional feature representation is input into the three-dimensional coordinate regression module of the SATNN model to predict the coordinate parameters of the sound source in three-dimensional space, thereby obtaining the predicted three-dimensional coordinates of the sound source.

5. The method according to claim 4, characterized in that, The encoder of the SATNN model consists of four cascaded encoding modules, each of which includes several sparse Transformer modules and a downsampling layer. The sparse Transformer module consists of a layer normalization module, a sparse attention module, and a feedforward network module. It is used for global feature modeling and multi-scale feature extraction of the input low-level feature representation. Specifically, the layer normalization module normalizes the input features; the sparse attention module constructs query, key, and value mappings and uses a sparse selection strategy to filter attention weights, achieving efficient modeling of global feature dependencies; and the feedforward network module employs a multi-scale deep convolutional structure to perform non-linear mapping and enhancement of features, achieving the fusion of local and global features. The downsampling layer uses a combination of convolution and pixel rearrangement to achieve spatial downsampling. The pixel rearrangement operation compresses the feature map, reducing spatial resolution while increasing channel dimension, thus obtaining a more hierarchical representation of sound source features.

6. The method according to claim 5, characterized in that, The sparse attention module includes: a feature mapping submodule, an attention computation submodule, a sparse selection submodule, a multi-scale fusion submodule, and an output mapping submodule; The feature mapping submodule performs linear mapping on the input features through 1×1 convolution and depthwise separable convolution to generate queries, keys and values, which are used to characterize the spatial and channel information in the three-dimensional rotating sound source distribution map. The attention calculation submodule normalizes the query and key, calculates the similarity matrix, obtains the correlation distribution between global features, and establishes long-distance dependencies. The sparse selection submodule filters the attention matrix based on the Top-K strategy, retaining the most important relevance information under different sparsity ratios, while suppressing irrelevant or preset low contribution regions. The multi-scale fusion submodule performs weighted fusion of attention results obtained under different sparsity ratios and adaptively combines multi-scale attention information through learnable parameters. The output mapping submodule maps the fused features back to the original feature space through linear projection, thereby realizing the integration and reconstruction of feature information.

7. The method according to claim 4, characterized in that, The decoder of the SATNN model includes several upsampling modules and feature fusion modules; The upsampling module upsamples the input features by pixel rearrangement to obtain upsampled features; The feature fusion module is used to concatenate the upsampled features with the feature map output by the corresponding layer of the encoder, and then use the concatenated multi-scale fused features to perform feature reconstruction through the sparse Transformer module to obtain the sound source reconstruction features.

8. A three-dimensional meshless rotating sound source localization system based on a sparse attention Transformer neural network, the system being used to implement the method according to any one of claims 1-7, characterized in that, The system includes: a data acquisition module, a prediction module, and a positioning module; The acquisition module is used to acquire the sound pressure signal of the rotating sound source field to construct the sound pressure cross spectrum matrix, and generate a three-dimensional rotating sound source distribution map based on the rotating beamforming algorithm of mode decomposition. The prediction module is used to input the three-dimensional rotating sound source distribution map into the SATNN model to obtain the predicted three-dimensional coordinates of the sound source; The positioning module is used to determine the spatial position of the rotating sound source based on the predicted three-dimensional coordinates of the sound source.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method described in any one of claims 1-7.