Transformer abnormal sound source positioning method and system based on gru-rnn and spectral enhancement
By using a method based on GRU-RNN and spectral enhancement, sensor array and neural network technology, the accuracy and anti-interference problems of transformer partial discharge sound source positioning were solved, and accurate positioning and real-time response to abnormal sound sources of transformers were achieved.
Patent Information
- Application Number
- CN202511122486.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-12
AI Technical Summary
The existing transformer partial discharge sound source localization technology has problems such as insufficient positioning accuracy, poor anti-interference ability, weak real-time performance and complex operation, making it difficult to accurately determine the specific location of the discharge point.
A method based on GRU-RNN and spectrum enhancement is adopted. Acoustic signals are collected through a sensor array, processed by gated recurrent unit (GRU) to generate a pseudo-covariance matrix, and eigenvalue decomposition is performed to construct a weighted noise subspace matrix. An optimized spectrum is generated through a multi-layer perceptron network and an adaptive resolution enhancement network. Finally, the spatial coordinates of the abnormal sound source are located through spectral peak detection.
The system can accurately locate the abnormal sound source of the transformer, improve the positioning accuracy and anti-interference ability, enhance the real-time performance and ease of operation of the system, and overcome the defects of traditional methods.
Smart Images

Figure CN120629848B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power equipment monitoring, in particular to a transformer abnormal sound source positioning method and system based on GRU-RNN and spectral enhancement. BACKGROUND
[0002] The transformer sound field imaging monitoring system is mainly applied to the electric power industry, especially the key equipment monitoring in the power transmission and distribution system. With the rapid development of smart grid, higher requirements are put forward for the state monitoring and fault diagnosis of power equipment. As one of the core equipment in the power grid, the running state of the transformer directly affects the safe and stable operation of the power grid. Therefore, developing efficient and accurate transformer monitoring technology is of great significance for ensuring power supply and preventing major accidents.
[0003] At present, the electric power industry has made significant progress in transformer monitoring technology. Traditional methods such as ultrasonic detection and infrared thermal imaging have been widely used in transformer fault detection. These methods can identify abnormal states of transformers to some extent. However, in the precise positioning of partial discharge sound sources, existing technologies still have many shortcomings. Ultrasonic detection can detect discharge signals, but the positioning accuracy is limited due to environmental noise, propagation attenuation and other factors; infrared thermal imaging focuses more on temperature anomalies and has weak direct positioning ability for discharge sound sources.
[0004] The main defects of existing transformer partial discharge sound source positioning technology are: insufficient positioning accuracy, which makes it difficult to accurately determine the specific location of the discharge point.
[0005] In addition, there are also: poor anti-interference ability, easy to be affected by environmental noise and other electromagnetic interference; poor real-time performance, unable to quickly respond to sudden failures; complex operation, requiring professional personnel to perform on-site operation and analysis.
[0006] The main reason for these defects is that existing technologies mostly rely on single sensors or simple signal processing methods, making it difficult to effectively extract and analyze complex sound field information. SUMMARY
[0007] In view of one of the defects in the prior art, the purpose of the present application is to provide a transformer abnormal sound source positioning method and system based on GRU-RNN and spectral enhancement.
[0008] The first aspect of the present application provides a transformer abnormal sound source positioning method based on GRU-RNN and spectral enhancement, comprising:
[0009] acquiring sound signals through a sensor array;
[0010] Processing the acoustic signal using a recurrent neural network (RNN) implemented by a gated recurrent unit (GRU) to generate a pseudo covariance matrix;
[0011] Performing eigenvalue decomposition on the pseudo-covariance matrix to obtain eigenvalues;
[0012] Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multi-layer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, which is then input into an adaptive resolution enhancement network to generate an optimized spectrum;
[0013] Performing logarithmic-difference processing on the optimized spectrum and then inputting the result into a multi-layer perceptron to obtain a beam weight vector;
[0014] A beamforming signal is generated according to the beam weight vector, and the spatial coordinates of an abnormal sound source in the transformer are located by spectrum peak detection.
[0015] Optionally, the processing of the acoustic signal by a recursive neural network (RNN) using a gated recurrent unit (GRU) to generate a pseudo-covariance matrix includes:
[0016] Processing the acoustic signal, dividing it into overlapping time frames, performing short-time Fourier transform on the signal of each time frame and converting it into the frequency domain to obtain N×F frequency point signals, where N is the number of time frames and F is the number of frequency points;
[0017] For one of the frequency point signals, construct a matrix obtained by multiplying the frequency point signal and its conjugate transpose;
[0018] Performing vectorization processing on the matrix to obtain an eigenvector including the real part and the imaginary part of the matrix elements;
[0019] For each time frame, the feature vectors of all frequency points are concatenated into the input vector of the frame, and N frames of input vectors constitute an input sequence;
[0020] The input vector of each time frame is used as the current time step, input into the GRU gated recurrent unit, and output the hidden state of the current time step;
[0021] A fully connected layer is used to process the hidden state, mapping it into a real vector, which is then reorganized into a conjugate symmetric matrix as a pseudo-covariance matrix.
[0022] Optionally, the step of constructing a weighted noise subspace matrix based on the eigenvalues through a multilayer perceptron network and introducing probability weights into a MUSCI algorithm to obtain an initial spatial spectrum includes:
[0023] Arrange the eigenvalues in ascending order and obtain corresponding eigenvector matrices;
[0024] Constructing a noise subspace according to the number of signal sources and the eigenvector matrix;
[0025] Inputting the eigenvectors of the noise subspace into a multilayer perceptron network, calculating the probability that each eigenvector belongs to the noise subspace, and constructing a weighted noise subspace matrix based on the probability;
[0026] The orthogonality between the weighted noise subspace matrix and the signal steering vector is utilized and the MUSCI algorithm is adopted to calculate the initial spatial spectrum.
[0027] Optionally, inputting the spatial spectrum into an adaptive resolution enhancement network for enhancement to obtain an optimized spectrum includes:
[0028] Taking the initial values of the spatial spectrum at L discrete angle points as input vectors;
[0029] Inputting the input vector into multiple serially connected weight learning modules to generate attention-enhanced features; wherein each weight learning module includes a one-dimensional convolution unit and an attention unit in sequence, the one-dimensional convolution unit of the first weight learning module extracts features from the input vector, the attention unit learns the importance weights of different angular regions from the features extracted by the one-dimensional convolution unit, and generates attention-enhanced features as the input vector of the one-dimensional convolution unit of the next weight learning module;
[0030] For the attention enhancement feature output by the last weight learning module, a convolution layer with a convolution kernel size of 1 is used to reduce the number of channels to 1, and the enhanced spectral value is obtained through a linear activation function.
[0031] Optionally, performing logarithmic-difference processing on the optimized spectrum and then inputting the processed spectrum into a multilayer perceptron to obtain a beam weight vector includes:
[0032] performing a logarithmic transformation on the optimized spectrum;
[0033] Based on the logarithmic transformation results, central difference calculation and range normalization processing are performed to obtain normalized results;
[0034] uniformly sampling a number of points between the angles [0-2π), extracting a local window of the sampling points, and calculating statistical features of the local window based on the normalized result;
[0035] Constructing an angle feature vector based on the normalization result and the statistical features to form a final angle feature matrix;
[0036] The angle feature matrix is input into a multi-layer perceptron network to obtain beam weights.
[0037] Optionally, inputting the angle feature matrix into a multilayer perceptron network to obtain beam weights includes:
[0038] The linear layer of the multilayer perceptron network outputs a complex weight w beam= w real + j w imag , as the beam value;
[0039] Among them, w real is the real part of the beam weight, given by W real h3+b real get;
[0040] w imag is the imaginary part of the beam weight, given by W imag h3+b imag get;
[0041] h3 is the shared feature vector in the multilayer perceptron network, W real is the weight matrix for calculating the real part of the beam weight, b real is the bias vector for calculating the real part of the beam weight; W imag is the weight matrix for calculating the imaginary part of the beam weight, b imag is the bias vector used to calculate the imaginary part of the beam weights.
[0042] Optionally, generating a beamforming signal according to the beam weight vector and locating the spatial coordinates of an abnormal sound source in the transformer by spectrum peak detection includes:
[0043] forming a beamformed signal based on the beam weight and the acoustic signal;
[0044] generating a time-frequency spectrum matrix by short-time Fourier transform for the beamforming signal;
[0045] Performing Mel spectrum conversion on the time-frequency spectrum data matrix to obtain a Mel spectrum;
[0046] Detecting local maximum points on the Mel spectrum;
[0047] Calculating the arrival time difference based on the local maximum point;
[0048] The spatial coordinates of the abnormal sound source are determined based on the arrival time difference.
[0049] The second aspect of the present application provides a transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement, comprising:
[0050] Signal module: collects acoustic signals through sensor array;
[0051] GRU-RNN module: Processes the acoustic signal using a recurrent neural network (RNN) implemented by a gated recurrent unit (GRU) to generate a pseudo-covariance matrix;
[0052] Eigenvalue decomposition module: performing eigenvalue decomposition on the pseudo covariance matrix to obtain eigenvalues;
[0053] Optimization module: Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multi-layer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, which is then input into an adaptive resolution enhancement network to generate an optimized spectrum;
[0054] Logarithmic difference module: performing logarithmic-difference processing on the optimized spectrum and then inputting the result into a multi-layer perceptron to obtain a beam weight vector;
[0055] Detection module: generates a beamforming signal according to the beam weight vector, and locates the spatial coordinates of the abnormal sound source in the transformer through spectrum peak detection.
[0056] In a third aspect of the present application, a terminal is provided, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it can be used to execute the method for localizing abnormal sound sources of transformers based on GRU-RNN and spectral enhancement, or to run the system for localizing abnormal sound sources of transformers based on GRU-RNN and spectral enhancement.
[0057] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute the transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement, or to run the transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement.
[0058] The transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement provided in this application uses RNN to replace the traditional MUSIC algorithm to estimate the covariance matrix and generate a pseudo-covariance matrix, which solves the sensitivity of the traditional MUSIC algorithm to model mismatch; through the collaboration of adaptive resolution enhancement network and probability weighted subspace, the processing capability and robustness of processing correlated and broadband sources are improved, and the accurate positioning of the transformer abnormal sound source is achieved.
[0059] A spatial spectrum is constructed based on the noise subspace and optimized using an adaptive resolution enhancement network. After logarithmic-difference processing, the optimized spectrum is used to predict the DOA angle and beam weight vector using a multi-layer perceptron. The beam weights are used to generate a beamforming signal, and the coordinates of the transformer's abnormal sound source are located through spectral peak detection. This application overcomes the traditional MUSIC algorithm's sensitivity to model mismatch and its inability to handle coherent and broadband sources, improving its processing capability and robustness for coherent and broadband signals, thereby accurately locating the transformer's abnormal sound source.
[0060] Other technical effects brought about by the additional features will be further explained in the corresponding embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0062] Figure 1 4 is a flowchart of a method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to an exemplary embodiment;
[0063] Figure 2 2 is an overall structural diagram of a transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement according to an exemplary embodiment;
[0064] Figure 3 This is a flowchart of calculating an optimized spectrum according to an exemplary embodiment;
[0065] Figure 4 is a flowchart of calculating a beam weight vector according to an exemplary embodiment;
[0066] Figure 5 The figure is a structural diagram of a transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement according to an exemplary embodiment. DETAILED DESCRIPTION
[0067] The present application is described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that, without departing from the concept of the present application, a number of variations and improvements may be made by those skilled in the art, and these all fall within the scope of protection of the present application. Parts not described in detail in the following examples may be implemented using existing technologies.
[0068] At present, the defects of transformer partial discharge sound source positioning technology are mainly: poor positioning accuracy, and it is difficult to accurately determine the specific position of the discharge point. Based on the above problems, the embodiment of the application provides a transformer abnormal sound source positioning method based on GRU-RNN and spectrum enhancement to solve the above problems.
[0069] Referring to Figure 1 and Figure 2 As shown in the embodiment of the application, a transformer abnormal sound source positioning method based on GRU-RNN and spectrum enhancement includes the following steps:
[0070] S100, acquiring sound signals through a sensor array;
[0071] S200, implementing recursive neural network (RNN) processing of the sound signals through a gated recurrent unit (GRU) to generate a pseudo-covariance matrix;
[0072] S300, performing eigenvalue decomposition on the pseudo-covariance matrix to obtain eigenvalues;
[0073] S400, based on the eigenvalues, constructing a weighted noise subspace matrix through a multi-layer perception network, introducing probability weights into the MUSCI algorithm, obtaining an initial spatial spectrum, and inputting an adaptive resolution enhancement network to generate an optimized spectrum;
[0074] S500, inputting the optimized spectrum into a multi-layer perception after performing log-difference processing to obtain a beam weight vector;
[0075] S600, generating a beam forming signal according to the beam weight vector, and positioning the spatial coordinates of the abnormal sound source in the transformer through spectrum peak detection.
[0076] Specifically, in the signal acquisition link, sound signals are acquired through a sensor array. Compared with a single sensor, the sensor array covers the sound field inside the transformer, can acquire sound signals from multiple positions and angles, and obtain more comprehensive sound field information.
[0077] The sound signals are processed by a recurrent neural network using a gated recurrent unit (GRU) to generate a pseudo-covariance matrix. The GRU can effectively process the time sequence characteristics of the sound signals, can capture the complex change rules of the sound signals in the time sequence, and can mine more valuable information. Compared with traditional simple signal processing methods, the extracted features can better reflect the essential characteristics of the abnormal sound source, and provide high-quality feature data for subsequent analysis.
[0078] The pseudo-covariance matrix is subjected to eigenvalue decomposition, and a weighted noise subspace matrix is constructed using a multi-layer perceptron network, incorporating probability weights into the algorithm. Compared to the traditional MUSIC algorithm, which simply selects the eigenvector corresponding to the minimum eigenvalue to form the noise subspace, this approach is more flexible and intelligent. Especially in low signal-to-noise ratio environments, it can effectively distinguish between signal and noise, accurately extract features related to abnormal sound sources, improve the ability to analyze complex sound fields, and reduce the impact of noise interference on positioning.
[0079] The initial spatial spectrum is calculated by utilizing the orthogonality of the noise subspace and the signal steering vector, and a weighted noise subspace matrix is used. This allows the calculated initial spatial spectrum to more accurately reflect the directional information of the abnormal sound source in complex situations such as low signal-to-noise ratio, and has higher reliability than traditional methods.
[0080] The initial spatial spectrum is input into the adaptive resolution enhancement network to generate an optimized spectrum. This process further optimizes the resolution of the spatial spectrum, which can more clearly present the distribution characteristics of abnormal sound sources in space and provide more detailed spatial information for precise positioning.
[0081] Logarithmic-difference processing is performed on the optimized spectrum to compress the data's dynamic range, enhance low-amplitude features, and extract trend characteristics, making the signal characteristics more suitable for the Multilayer Perceptron (MLP). By learning these processed features, the MLP can accurately obtain the set of Direction of Arrival (DOA) angles and beam weight vectors, achieving precise estimation of the direction of the abnormal sound source.
[0082] A beamforming signal is generated based on the beam weight vector, enhancing the signal in the target direction and suppressing interference from other directions. Spectral peak detection then locates the location with the highest signal energy, thereby determining the spatial coordinates of the abnormal sound source. This seamless process from direction estimation to coordinate determination fully leverages the information obtained from previous steps to ensure accurate positioning.
[0083] The method of the above embodiment of the present application comprehensively improves the ability to collect, process and analyze complex sound field information inside the transformer through collaborative optimization of multiple links, effectively overcomes the defects of traditional methods, and thus realizes the precise positioning of abnormal sound sources.
[0084] Furthermore, the GRU network can learn signal characteristics in complex noisy environments, filter noise through a gating mechanism, and generate a more robust pseudo-covariance matrix. Multi-layer perceptron probability weighting is introduced in the noise subspace construction, enabling accurate separation of signal and noise subspaces even at low signal-to-noise ratios. The adaptive resolution enhancement network learns through training how to enhance valid signals and suppress interference, improving the quality of the spatial spectrum. Logarithmic-difference processing extracts local features, enhancing sensitivity to anomalous sound sources while simultaneously suppressing ambient noise. Beamforming utilizes weight vectors to create gain in the target direction and null in the interference direction, effectively suppressing interference. This effectively improves interference resistance.
[0085] To more comprehensively and accurately acquire acoustic signals from the transformer's internal sound field, in some specific embodiments of the present application, the sensor array employed in S100 covers the transformer's internal sound field. The sensor array employed is an M-element sensor array. The M-element sensor array herein refers to a system that acquires acoustic field signals using M spatially distributed sensors.
[0086] Specifically, the acoustic signals received from the M-element sensor array constitute a signal matrix:
[0087] ;
[0088] Where: x(t) represents the signal vector at time t, the synchronous measurement value of all sensors at time t. t Is the time index, sampling time number ; T: total number of sampling points.
[0089] Among them, each signal vector in the signal matrix is expressed as:
[0090] x(t)=A(θ)s(t)+n(t)
[0091] in, x(t) is the signal vector received by the microphone sensor array, which is an M-dimensional complex column vector; A(θ) is the array manifold matrix, which contains the gain coefficients and time delay information of all microphone units to the sound source; θ is the direction of arrival of the sound source DOA (Direction Of Arrival); s( t ) is the sound source signal vector; n( t ) is the noise vector.
[0092] The array manifold matrix contains the steering vector ,Right now
[0093] ;
[0094] in θ d is the source direction, d represents the sound source index, identifying the dth independent sound source, .
[0095] The above-mentioned embodiments of the present application, by means of a sensor array, collect acoustic signals from multiple spatial locations, and can obtain richer information on the spatial distribution of the sound field, providing sufficient data support for accurately locating abnormal sound sources, thereby effectively improving positioning accuracy. Secondly, the data collected by multiple sensors can reduce the impact of environmental noise and electromagnetic interference on effective signals through mutual verification and interference cancellation, significantly enhancing the system's anti-interference ability. In addition, the design of the M-element sensor array can cover a wider frequency range, and can fully capture the acoustic signal characteristics generated by the transformer under different fault conditions, achieving effective coverage of the transformer's full fault frequency band.
[0096] In order to more effectively process the temporal characteristics of the acoustic signal, a pseudo covariance matrix that can more accurately reflect the signal characteristics is generated. In some specific embodiments of the present application, Figure 3 As shown, in S200, the recursive neural network RNN is implemented by a gated recurrent unit GRU to process the acoustic signal and generate a pseudo covariance matrix. The following steps can be used:
[0097] S201, acoustic signal preprocessing: Split the continuous acoustic signal into overlapping time frames, perform short-time Fourier transform (STFT) on each frame signal to convert it into the frequency domain, and the signal of the fth frequency point in the nth frame is x n,f , 1≤n≤N, 1≤f≤F, N is the number of overlapping time frames; F is the number of STFT frequency points.
[0098] S202, construct GRU input features:
[0099] Define the vector matrix: ;
[0100] Vectorized processing:
[0101] ;
[0102] in, r ij is a matrix The (i,j)th element of .
[0103] Input vector: For each frame n, the feature vectors of all frequency points (or selected F frequency points) are concatenated to form the input sequence of the frame.
[0104] Specifically, the frequency point eigenvector is: ; The input sequence is , where N is the number of frames.
[0105] S203, GRU network structure:
[0106] The input is the current time step input v n and the hidden state at the previous time step , the output is the hidden state h of the current time step n .
[0107] ;
[0108] ;
[0109] ;
[0110] ;
[0111] in:
[0112] σ is the sigmoid activation function.
[0113] ⊙ is element-wise multiplication.
[0114] W z, W r, W h is the weight matrix input to the hidden state.
[0115] U z, U r, U h is the weight matrix from hidden state to hidden state.
[0116] b z, b r, b h is the bias vector.
[0117] Specifically, a recurrent neural network (RNN) is a type of neural network specialized for processing time series data. It retains historical information through recurrent connections in the hidden layer and is suitable for analyzing time-varying sequence signals (such as transformer acoustic signals). The gated recurrent unit (GRU) in this application is an optimized variant of the RNN. By introducing reset and update gates, it addresses the vanishing gradient problem of long sequences in traditional RNNs and can more efficiently capture the temporal dependencies of acoustic signals. Therefore, the above-mentioned embodiments of this application essentially use the GRU as the specific structure of the RNN, leveraging its gating mechanism to accurately model the dynamic characteristics of acoustic signals, providing information support in the temporal dimension for the subsequent generation of the pseudo-covariance matrix.
[0118] S204, generate a pseudo covariance matrix:
[0119] Use a fully connected layer to transform the hidden state h n Mapped to a real vector (through dense layers, network layers
[0120] complex layer), and then reorganize it into a complex conjugate symmetric matrix. Specifically:
[0121] The hidden state h n Mapped to a dimension of 2 through a fully connected layer M 2 A real vector of (M = the number of microphones in the sensor array)
[0122] o n =W o h n +b o
[0123] The output vector o n Convert to a complex matrix:
[0124] Will o n Divided into two parts: front M 2 elements as the real part, and then M 2 elements as the imaginary part.
[0125] Reshape into a complex matrix :
[0126] ;
[0127] Enforce conjugate symmetry (since the covariance matrix must be conjugate symmetric):
[0128] ;
[0129] Ensuring positive definiteness: Positive definiteness is guaranteed by adding a small multiple of the identity matrix:
[0130] ;
[0131] in, is a small positive number.
[0132] The above embodiment of the present application uses GRU-RNN to dynamically generate pseudo-covariance matrix to replace the traditional statistical calculation method, and replaces the traditional statistical average with deep time series modeling. This innovation effectively reduces the demand for the number of samples and improves the real-time performance and anti-interference ability of the system. The matrix conjugate symmetry operation is embedded at the output of the neural network to construct a conjugate symmetry enforcement mechanism, which greatly reduces the model mismatch error and significantly improves the resolution of coherent sources. Design based on matrix norm Dynamically adjusting the strategy to achieve adaptive positivity guarantees avoids the over-disturbance problem that may be caused by manually setting fixed values. This enables the system to effectively reduce the positioning failure rate in extreme scenarios such as sensor failure and sudden changes in oil temperature, further enhancing the stability and reliability of the technology application.
[0133] Furthermore, the pseudo covariance matrix is subjected to eigenvalue decomposition. In some specific embodiments of the present application, in S300, the pseudo covariance matrix is subjected to eigenvalue decomposition to obtain eigenvalues, and the following steps may be employed:
[0134] S301, matrix symmetry processing, eliminate the possible asymmetric error of GRU output, and meet the EVD condition R=R H 。
[0135] Specifically, the pseudo covariance matrix generated by GRU Perform symmetry operation:
[0136] ;
[0137] H : conjugate transpose (complex field) or transpose (real field)
[0138] S302, eigenvalue decomposition, solve the characteristic equation:
[0139] ;
[0140] S303, matrix decomposition form:
[0141] ;
[0142] in:
[0143] , is the eigenvalue diagonal matrix;
[0144] , is the eigenvector matrix.
[0145] The above-described embodiment of the present application utilizes a dual symmetrization mechanism, adding a pre-symmetrization process (S301) prior to standard eigendecomposition. This, combined with the subsequent Riemann optimization (S303), provides dual guarantees, thus overcoming the limitations of traditional processes. Pre-symmetrization eliminates asymmetric errors in the GRU output, thereby ensuring the stability of the eigendecomposition. Riemann optimization provides a secondary guarantee at the eigenvalue level to ensure positive definiteness, addressing the issue of negative eigenvalue drift and ultimately ensuring convergence of the MUSIC algorithm.
[0146] Eigenvalue decomposition is the core method for dividing the signal subspace and the noise subspace. After decomposition, large eigenvalues correspond to the signal subspace (carrying target signal information), and small eigenvalues correspond to the noise subspace (carrying noise information). In order to further determine which features belong to the noise subspace, in some specific embodiments of the present application, S400, based on the eigenvalues, constructs a spatial spectrum through the noise subspace eigenvector matrix, and inputs it into the adaptive resolution enhancement network to generate an optimized spectrum, such as Figure 3 As shown, the following steps can be taken:
[0147] S401, weighted selection of noise subspace eigenvectors.
[0148] Enter the eigenvalues: (Sort by ascending order, i.e. ) and the corresponding eigenvector matrix: ;
[0149] If the number of signal sources D (scope ) is known, the noise subspace consists of the minimum The eigenvalues corresponding to the eigenvectors are composed of: .
[0150] S402, the feature vector The input is fed into the multi-layer perceptron network structure, which learns the probability that each eigenvector belongs to the noise subspace, thereby constructing a weighted noise subspace matrix.
[0151] Specifically, the multi-layer perceptron network structure includes:
[0152] Input layer: M eigenvalues (or consider other features, such as the logarithm of the eigenvalue, the ratio of adjacent eigenvalues, etc.)
[0153] Hidden layer: 2 layers, 64 neurons per layer, using ReLU activation function;
[0154] Output layer: M output nodes, using Sigmoid activation function, output probability qi ∈[0,1]
[0155] ;
[0156] in:
[0157] W 1, b1 is the weight and bias of the first layer;
[0158] W 2, b2 is the weight and bias of the second layer;
[0159] is the probability that each eigenvector is selected as the noise subspace
[0160] Construct the weighted noise subspace matrix:
[0161] ;
[0162] It is worth noting that the traditional MUSIC algorithm adopts a “hard selection” strategy when selecting the noise subspace: only the smallest N The eigenvectors corresponding to the eigenvalues form the noise subspace. The selected eigenvectors have a weight of 1, and the unselected ones have a weight of 0. This is a discrete, either-or selection method. The above embodiment of the present application uses a multi-layer perceptron (MLP) to learn the continuous weight (weight value between 0 and 1) of each eigenvector belonging to the noise subspace. Instead of directly determining whether the eigenvector "belongs" or "does not belong" to the noise subspace, it uses continuous weights to reflect its "degree of belonging to the noise subspace". These weights are then used to weight the eigenvectors to construct a weighted noise subspace matrix.
[0163] In low signal-to-noise ratio scenarios, the eigenvalue distinction between signal and noise is significantly reduced. Traditional MUSIC's "hard selection" can easily lead to incorrect selection of eigenvectors in the noise subspace due to fuzzy eigenvalues, resulting in increased errors in subsequent processing. However, a multilayer perceptron can learn to assign appropriate continuous weights to each eigenvector, even when eigenvalues are difficult to clearly distinguish. This allows for a more flexible and precise characterization of the role of each eigenvector in the noise subspace, making the noise subspace construction more reliable. Therefore, this method is more stable (and more robust) than traditional MUSIC at low signal-to-noise ratios.
[0164] S403, initial spatial spectrum calculation.
[0165] This step mainly uses the orthogonality between the noise subspace and the signal steering vector to calculate the spatial spectrum. Specifically:
[0166] For a given scanning angle φ , calculate the steering vector a( φ) (i.e., array manifold vector). The scanning angle φ is a discretized search variable for the sound source direction angle, which is used to systematically detect potential sound source locations in space.
[0167] Its core function is to discretize the continuous spatial azimuth angle by calculating the spatial spectrum value at each discrete angle. , find the true sound source direction corresponding to the spectrum peak. The specific calculation method is as follows:
[0168] ;
[0169] : Scan angle (sound source direction assumption) ;
[0170] g m : Gain coefficient of the mth sensor;
[0171] λ: Wavelength of sound waves;
[0172] : Time delay from sound source to sensor m
[0173] ;
[0174] r m :sensor m Location coordinates
[0175] : unit vector of the sound source direction
[0176] C: Speed of sound
[0177] The initial spatial spectrum is defined as:
[0178] ;
[0179] It is worth noting that this embodiment uses a weighted noise subspace matrix E instead of the subspace consisting only of noise eigenvectors in traditional MUSIC.
[0180] The above embodiment of the present application is compatible with tradition and innovation: when the weighted vector q is close to the traditional "hard selection" (the weights corresponding to the smallest N eigenvalues are approximately 1, and the rest are approximately 0), it can approximate the traditional MUSIC spectrum, ensuring consistency with the classical method in applicable scenarios.
[0181] The above-mentioned embodiments of the present application have improved robustness and accuracy: due to the use of a weighted noise subspace (rather than the traditional subspace consisting only of noise eigenvectors), the contribution of each eigenvector in the noise subspace can be more flexibly integrated. Combined with the continuous weights learned by the multi-layer perceptron (MLP), in scenarios with low signal-to-noise ratio and low eigenvalue discrimination, the spatial spectrum calculation is more accurate and robust than traditional MUSIC, providing a more reliable foundation for subsequent tasks such as sound source localization.
[0182] S404, adaptive resolution enhancement network.
[0183] The adaptive resolution enhancement network enhances the initial spatial spectrum, improves the angular resolution and suppresses the side lobes. The structure of the adaptive resolution enhancement network is as follows:
[0184] Input: the values of the initial spatial spectrum at multiple angles;
[0185] Network: 1D Convolutional Neural Network (CNN);
[0186] Output: Optimized spatial spectrum values at the same angle.
[0187] Specifically:
[0188] S4041, input representation: The values of the initial spatial spectrum at L discrete angle points are used as input vectors: .
[0189] S4042, input the input vector of S4041 into multiple series-connected weight learning modules, each weight learning module includes a convolution unit and an attention unit.
[0190] Specifically, the convolution layer uses multiple one-dimensional convolution kernels connected in series for feature extraction.
[0191] Exemplary:
[0192] First layer: 32 convolution kernels, size 5, stride 1, ReLU activation
[0193] Second layer: 32 convolution kernels, size 3, stride 1, ReLU activation
[0194] The third layer: 16 convolution kernels, size 3, stride 1, ReLU activation
[0195] The attention mechanism adds an attention unit after each convolution unit to learn the importance weights of different angular regions, allowing the network to focus on the angular regions where signals may appear, improving the sharpness of the main lobe and suppressing the side lobes.
[0196] Specifically, in the attention unit, a global average pooling layer is used, and then a fully connected layer is used to generate the weight of each angle:
[0197] ;
[0198] where f i is the feature of the i-th angle.
[0199] : the first fully connected layer weight
[0200] : the first fully connected layer bias
[0201] ReLU: activation function
[0202] : second fully connected layer weight
[0203] : the second fully connected layer bias
[0204] Sigmoid: activation function
[0205] S4043, output layer: using a convolution layer with a kernel size of 1 (equivalent to a fully connected layer), the channel number of the feature output by the last group of weight learning modules is reduced to 1, and then a linear activation function is used to obtain the enhanced spectral value:
[0206] ; where:
[0207] f : input feature (from the third layer attention output, i.e. the output feature of the third group of weight learning modules)
[0208] Wout: 1x1 convolution kernel weight (channel dimension reduction)
[0209] bout: bias vector (angle dimension independent bias)
[0210] Final output: ;
[0211] ;
[0212] Nattn represents the mapping function of the entire adaptive resolution enhancement network: an end-to-end nonlinear transformation from the original spatial spectrum to the enhanced spatial spectrum.
[0213] In the above-mentioned embodiments of this application, continuous probability modeling of noise subspaces: While traditional methods use hard threshold segmentation, the above-mentioned embodiments use a multi-layer perceptron to learn continuous probabilities, resolving the hard decision flaws in low signal-to-noise ratio scenarios and combating eigenvalue ambiguity. A spatial spectral deep learning enhancement framework: An end-to-end enhancement architecture, a three-level attention convolutional network, and an output layer fidelity design can suppress strong sidelobe interference.
[0214] In order to convert the optimized spectrum into a feature form more suitable for multi-layer perceptron processing and improve the prediction accuracy of the model for DOA angles and beam weight vectors, in some specific embodiments of the present application, S500 performs logarithmic-difference processing on the optimized spectrum and then inputs it into the multi-layer perceptron to obtain the DOA angle set and beam weight vector, such as Figure 4 As shown, the following steps can be taken:
[0215] S501, log-difference processing: enhances the local features of the spectrum, especially the gradient near the peak, and suppresses the flat area, making it easier to be learned by the multi-layer perceptron.
[0216] S501.1 compresses the dynamic range through logarithmic transformation and enhances low-amplitude features. The formula is as follows:
[0217] ;
[0218] in: It is a numerical stabilization term that prevents negative numbers from appearing when taking logarithms. It is used to describe different direction angles in space and is the independent variable of the spatial spectrum.
[0219] S501.2 uses differential operations to enhance local change features, suppress background noise, and simulate the sensitivity of human hearing to slope changes.
[0220] Specifically, the first-order central difference:
[0221] ;
[0222] in: is the angle sampling step, It is the target angle of the current differential spectrum value to be calculated. It is a specific direction angle scanned in spatial spectrum analysis and is used to locate the direction of the sound source.
[0223] S501.3 range normalization processing makes the input data conform to the standard normal distribution and accelerates the convergence of the neural network
[0224] ;
[0225] in: , is the mean; σ = , is the standard deviation.
[0226] S502, angle sampling and feature enhancement.
[0227] S502.1, uniform sampling: select R points at equal intervals in the range [0, 2π):
[0228] ;
[0229] S502.2 Local Window Extraction:
[0230] For each angle sampling point (in ): Take The local window at the center
[0231] ;
[0232] The window size is 11 points (5 points before and after), corresponding to Angle range (when 5.5 when
[0233] S502.3, for each window W i Calculate three statistical features:
[0234] Local mean: Indicates the average signal strength of the area near the current angle, reflects the background noise level, and distinguishes signal from noise areas
[0235] ;
[0236] Local peak: Indicates the signal dynamic range in the area near the current angle, highlights the signal mutation area, and enhances the peak detection capability.
[0237] ;
[0238] Peak position offset: indicates the maximum value point in the window relative to the center point The offset (in points) of the peak indicates the direction and distance of the true peak position.
[0239] ;
[0240] S502.4 Feature Vector Construction
[0241] For each angle point Construct a 4-dimensional feature vector:
[0242] ;
[0243] Finally, the feature matrix is formed:
[0244] ;
[0245] S503, multi-layer perceptron network topology architecture.
[0246] Shared feature layer input: feature matrix (Include R 4-dimensional features of each angle point), and gradually extract high-level features through three fully connected layers. The specific structure is:
[0247] ;
[0248] ;
[0249] ;
[0250] Where: W1, W2, W3 are the weight matrices of the first, second, and third fully connected layers respectively; b1, b2, b3 are the bias vectors of the first, second, and third fully connected layers respectively; h1, h2, h3 are the output feature vectors of the first, second, and third fully connected layers respectively. is the activation function, and the output h3 is the shared feature vector.
[0251] Parameter initialization:
[0252] ;
[0253] is the number of input neurons
[0254] DOA angle output branch (dense layer output): output vector ,in The maximum number of detectable signal sources, each Represents the estimated DOA angle, as follows:
[0255] ;
[0256] ;
[0257] Beam weight output branch (dense layer output): each element w m Indicates the m The complex weights of the antennas beamform the signal: , as follows:
[0258] ;
[0259] ;
[0260] ;
[0261] Among them, w real is the real part of the beam weight, given by W real h3+b real get;
[0262] Among them, w imag is the imaginary part of the beam weight, given by W imag h3+b imag get;
[0263] h3 is the shared feature vector in the multilayer perceptron network, W real is the weight matrix for calculating the real part of the beam weight, b real is the bias vector for calculating the real part of the beam weight; W imag is the weight matrix for calculating the imaginary part of the beam weight, b imag is the bias vector used to calculate the imaginary part of the beam weights.
[0264] Source quantity classification output branch (dense layer output), source quantity classification branch pi Indicates that there is i +1 signal source probability, the final output , the specific structure is as follows:
[0265] ;
[0266] Where: W c Used to perform linear transformation on the features of ReLU output, b c is the bias; W d、 b d are the weight matrix and bias vector for linear transformation of shared feature h3.
[0267] Probability vector ;
[0268] ;
[0269] ;
[0270] In the above embodiment of the present application, S501-S502 is a three-level feature enhancement pipeline: physical signal processing (acoustic spectrum enhancement) is deeply integrated with deep learning feature engineering, breaking through the limitation of traditional multi-layer perceptrons that directly process the original spectrum.
[0271] S503 is a multi-task output joint optimization: a single network simultaneously outputs the DOA angle, beam weight, and number of signal sources, achieving an end-to-end acoustic perception closed loop. This improves DOA resolution, reduces computational latency, and enhances the algorithm's real-time performance.
[0272] To obtain more accurate abnormal sound source positioning, in some embodiments of the present application, S600, a beamforming signal is generated according to the beam weight vector, and the spatial coordinates of the abnormal sound source in the transformer are located by spectral peak detection, as shown in Figure 4 The following steps can be used:
[0273] S601, the beamformer forms a main lobe gain in the target direction and a null in the interference direction, and the beamforming signal is as follows:
[0274] y ( t )= x( t )
[0275] Wherein: y ( t ) is the beamforming signal at time t. w beam H : the conjugate transpose of the beam weight vector. represents the array signal at time t
[0276] S602, a time-frequency spectrum matrix is generated using short-time Fourier transform:
[0277] Y ( f , τ )=
[0278] Wherein:
[0279] : the element of the time-frequency spectrum matrix, representing the time-frequency value at frequency f , time τ , which is the output result of the short-time Fourier transform
[0280] y ( t ): the original time-domain signal, which is the input signal to be analyzed
[0281] represents the value of the time-domain signal at time .
[0282] w ( m ): window function w the value at discrete point m ;
[0283] w is the window function; N is the window length; R is the frame shift; frequency resolution. fs is sampling frequency, refers to the sampling frequency of the original time-domain signal y t ) is the sampling rate at which sampling is performed.
[0284] S603, Mel spectrum conversion, obtain Mel spectrum S mel .
[0285] S604, peak detection:
[0286] exist S mel The local maximum point that satisfies the following formula is detected: and ;
[0287] Specifically, μ is the Mel spectrum S mel The overall (or within a certain statistical range) mean is used to measure the average energy level of the spectrum. 3σ is a threshold calculation term combined with the standard deviation σ.
[0288] The time difference of arrival (TDOA) is:
[0289] ;
[0290] τ p : Time frame corresponding to the spectrum peak
[0291] f 0: Spectral peak center frequency (such as 800Hz iron core vibration)
[0292] ∠: complex phase angle
[0293] Y i (f0,τ p ), Y j (f0,τ p ):respectively sensor i and sensor j at the "spectral peak center frequency f0, the time frame corresponding to the spectrum peak τ p The complex spectrum value at " contains the amplitude and phase information of the signal. The time difference is derived by calculating the phase difference through the ratio of the two.
[0294] S605, the sound source position is solved as follows:
[0295] ;
[0296] : Sound source coordinates
[0297] : sensor i coordinate
[0298] c : Speed of sound
[0299] Least squares solution:
[0300] ;
[0301] in,
[0302] ;
[0303] ;
[0304] Transformer coordinate transformation:
[0305] ;
[0306] R: 3×3 rotation matrix;
[0307] t=[ t x , t y , t z ] T : translation vector
[0308] The above-mentioned embodiment of the present application adopts a multi-domain fusion spectrum peak detection mechanism, which combines statistical threshold screening (μ + 3σ) with three-dimensional neighborhood extreme value verification (joint judgment of space, time and frequency), breaking through the limitations of traditional two-dimensional spectrum peak detection and effectively improving the anti-interference ability.
[0309] In addition, the GRU network processes the acoustic signal to generate a pseudo-covariance matrix, which is processed quickly through the time-series processing capabilities of the recurrent neural network. The entire process uses neural network parallel computing to reduce the computational delay of traditional iterative algorithms.
[0310] Peak detection is performed in the time-frequency domain, using fast STFT and Mel spectrum conversion, combined with non-maximum suppression, to quickly locate the peak, effectively improving the real-time performance of positioning.
[0311] Based on the same technical concept, in some specific implementations of the present application, such as Figure 5 As shown, a transformer abnormal sound source localization system 100 based on GRU-RNN and spectral enhancement includes:
[0312] Signal module 110: collects acoustic signals through a sensor array;
[0313] GRU-RNN module 120: Processing the acoustic signal using a recurrent neural network (RNN) implemented by a gated recurrent unit (GRU) to generate a pseudo covariance matrix;
[0314] Eigenvalue decomposition module 130: performs eigenvalue decomposition on the pseudo covariance matrix to obtain eigenvalues;
[0315] Optimization module 140: Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multi-layer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, which is then input into an adaptive resolution enhancement network to generate an optimized spectrum;
[0316] Logarithmic difference module 150: performs logarithmic-difference processing on the optimized spectrum and then inputs it into a multi-layer perceptron to obtain a DOA angle set and a beam weight vector;
[0317] Detection module 160: generates a beamforming signal according to the beam weight vector, and locates the spatial coordinates of the abnormal sound source in the transformer through spectrum peak detection.
[0318] The specific implementation technology of each module / unit in the above example of this application can refer to the steps corresponding to the transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement in the above embodiment, which will not be repeated here.
[0319] The preferred features of the above embodiments can be used alone in any embodiment, or in any combination without conflict. In addition, parts not described in detail in the embodiments can be implemented using existing technologies.
[0320] Based on the same technical concept, in other embodiments of the present application, a terminal is provided, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it can be used to execute the above-mentioned transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement, or, run the above-mentioned transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement.
[0321] Based on the same technical concept, in other embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute the above-mentioned transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement, or run the above-mentioned transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement.
[0322] Optionally, the memory is used to store programs. The memory may include volatile memory (volatile memory), such as random-access memory (RAM), such as static random-access memory (SRAM) and double data rate synchronous dynamic random access memory (DDR SDRAM). The memory may also include non-volatile memory (non-volatile memory), such as flash memory. The memory is used to store computer programs (such as applications and functional modules that implement the above-mentioned methods), computer instructions, etc. These computer programs and computer instructions may be partitioned and stored in one or more memories. Furthermore, these computer programs, computer instructions, data, etc. can be accessed by the processor.
[0323] The aforementioned computer programs, computer instructions, etc. may be partitioned and stored in one or more memories, and the aforementioned computer programs, computer instructions, data, etc. may be called by a processor.
[0324] The processor is configured to execute the computer program stored in the memory to implement the various steps of the method involved in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0325] The processor and memory can be independent structures or integrated structures. When the processor and memory are independent structures, the memory and processor can be coupled via a bus.
[0326] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0327] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0328] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0329] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0330] The above describes some specific embodiments of the present application. It should be understood that the present application is not limited to the specific embodiments described above, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the substantive content of the present application. The above preferred features may be used in any combination as long as they do not conflict with each other.
Claims
1. A transformer abnormal sound source localization method based on GRU-RNN and spectral enhancement, characterized in that: include: Acquiring acoustic signals through a sensor array; Processing the acoustic signal using a recurrent neural network (RNN) implemented by a gated recurrent unit (GRU) to generate a pseudo covariance matrix; Performing eigenvalue decomposition on the pseudo-covariance matrix to obtain eigenvalues; Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multi-layer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, which is then input into an adaptive resolution enhancement network to generate an optimized spectrum; Performing logarithmic-difference processing on the optimized spectrum and then inputting the result into a multi-layer perceptron to obtain a beam weight vector; A beamforming signal is generated according to the beam weight vector, and the spatial coordinates of an abnormal sound source in the transformer are located by spectrum peak detection.
2. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 1 is characterized in that: The method of implementing a recursive neural network (RNN) to process the acoustic signal through a gated recurrent unit (GRU) to generate a pseudo-covariance matrix includes: Processing the acoustic signal, dividing it into overlapping time frames, performing short-time Fourier transform on the signal of each time frame and converting it into the frequency domain to obtain N×F frequency point signals, where N is the number of time frames and F is the number of frequency points; For one of the frequency point signals, construct a matrix obtained by multiplying the frequency point signal and its conjugate transpose; Performing vectorization processing on the matrix to obtain an eigenvector including the real part and the imaginary part of the matrix elements; For each time frame, the feature vectors of all frequency points are concatenated into the input vector of the frame, and N frames of input vectors constitute an input sequence; The input vector of each time frame is used as the current time step, input into the gated recurrent unit GRU, and output the hidden state of the current time step; A fully connected layer is used to process the hidden state, mapping it to a real vector, which is reorganized into a conjugate symmetric matrix as the pseudo covariance matrix.
3. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 1 is characterized in that: Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multilayer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, including: Arrange the eigenvalues in ascending order and obtain corresponding eigenvector matrices; Constructing a noise subspace according to the number of signal sources and the eigenvector matrix; Inputting the eigenvectors of the noise subspace into a multilayer perceptron network, calculating the probability that each eigenvector belongs to the noise subspace, and constructing a weighted noise subspace matrix based on the probability; The orthogonality between the weighted noise subspace matrix and the signal steering vector is utilized and the MUSCI algorithm is adopted to calculate the initial spatial spectrum.
4. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 3 is characterized in that: Inputting the spatial spectrum into an adaptive resolution enhancement network for enhancement to obtain an optimized spectrum includes: Taking the initial values of the spatial spectrum at L discrete angle points as input vectors; Inputting the input vector into multiple serially connected weight learning modules to generate attention-enhanced features; wherein each weight learning module includes a one-dimensional convolution unit and an attention unit in sequence, the one-dimensional convolution unit of the first weight learning module extracts features from the input vector, the attention unit learns the importance weights of different angular regions from the features extracted by the one-dimensional convolution unit, and generates attention-enhanced features as the input vector of the one-dimensional convolution unit of the next weight learning module; For the attention enhancement feature output by the last weight learning module, a convolution layer with a convolution kernel size of 1 is used to reduce the number of channels to 1, and the enhanced spectral value is obtained through a linear activation function.
5. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 1 is characterized in that: The performing of logarithmic-difference processing on the optimized spectrum and inputting the resulting data into a multi-layer perceptron to obtain a beam weight vector includes: performing a logarithmic transformation on the optimized spectrum; Based on the logarithmic transformation results, central difference calculation and range normalization processing are performed to obtain normalized results; uniformly sampling a number of points between the angles [0-2π), extracting a local window of the sampling points, and calculating statistical features of the local window based on the normalized result; Constructing an angle feature vector based on the normalization result and the statistical features to form a final angle feature matrix; The angle feature matrix is input into a multi-layer perceptron network to obtain beam weights.
6. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 5 is characterized in that: Inputting the angle feature matrix into a multi-layer perceptron network to obtain beam weights includes: The linear layer of the multilayer perceptron network outputs a complex weight w beam= w real + j w imag , as the beam value; Among them, w real is the real part of the beam weight, given by W real h3+b real get; w imag is the imaginary part of the beam weight, given by W imag h3+b imag get; h3 is the shared feature vector in the multilayer perceptron network, W real is the weight matrix for calculating the real part of the beam weight, b real is the bias vector for calculating the real part of the beam weight; W imag is the weight matrix for calculating the imaginary part of the beam weight, b imag is the bias vector used to calculate the imaginary part of the beam weights.
7. The method for locating abnormal sound sources of transformers based on GRU-RNN and spectral enhancement according to claim 1 is characterized in that: Generating a beamforming signal according to the beam weight vector and locating the spatial coordinates of an abnormal sound source in the transformer by spectrum peak detection includes: forming a beamformed signal based on the beam weight and the acoustic signal; generating a time-frequency spectrum matrix by short-time Fourier transform for the beamforming signal; Performing Mel spectrum conversion on the time-frequency spectrum data matrix to obtain a Mel spectrum; Detecting local maximum points on the Mel spectrum; Calculating the arrival time difference based on the local maximum point; The spatial coordinates of the abnormal sound source are determined based on the arrival time difference.
8. A transformer abnormal sound source localization system based on GRU-RNN and spectral enhancement, characterized in that: include: Signal module: collects acoustic signals through sensor array; GRU-RNN module: Processes the acoustic signal using a recurrent neural network (RNN) implemented by a gated recurrent unit (GRU) to generate a pseudo-covariance matrix; Eigenvalue decomposition module: performing eigenvalue decomposition on the pseudo covariance matrix to obtain eigenvalues; Optimization module: Based on the eigenvalues, a weighted noise subspace matrix is constructed through a multi-layer perceptron network, and probability weights are introduced into the MUSCI algorithm to obtain an initial spatial spectrum, which is then input into an adaptive resolution enhancement network to generate an optimized spectrum; Logarithmic difference module: performing logarithmic-difference processing on the optimized spectrum and then inputting the result into a multi-layer perceptron to obtain a beam weight vector; Detection module: generates a beamforming signal according to the beam weight vector, and locates the spatial coordinates of the abnormal sound source in the transformer through spectrum peak detection.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it can be used to perform the method according to any one of claims 1 to 7, or run the system according to claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it can be used to perform the method according to any one of claims 1 to 7, or to run the system according to claim 8.
Citation Information
Patent Citations
Unmanned ship positioning method based on Kalman filtering
CN117268400A
Partial discharge positioning method and system based on radio waves
CN118837695A