Transformer voiceprint diagnosis method based on dual-path time-frequency fusion and related device
By constructing time-domain and frequency-domain features through a dual-path time-frequency fusion method for transformer acoustic signature diagnosis, and performing deep interaction and adaptive fusion, the problem of data scarcity and insufficient feature representation in transformer acoustic signature diagnosis models is solved, thereby improving the accuracy and robustness of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-03
AI Technical Summary
Existing transformer voiceprint diagnostic models based on Class Token suffer from insufficient diagnostic accuracy and robustness due to data scarcity and limited feature representation capabilities.
A transformer acoustic signature diagnostic method based on dual-path time-frequency fusion is adopted. By constructing dual-path Class Token features in the time and frequency domains, deep interaction and adaptive fusion are performed. Combined with multi-scale feature extraction and attention mechanism, the feature representation capability is improved.
It significantly improves the accuracy and robustness of transformer fault diagnosis, and can better integrate acoustic signature features of different resolutions to adapt to different operating conditions and equipment models.
Smart Images

Figure CN121789716A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of transformer fault diagnosis technology, and relates to a dual-path time-frequency fusion transformer acoustic signature diagnosis method and related devices. Background Technology
[0002] In recent years, deep learning-based transformer fault diagnosis technology has made significant progress, especially the VisionTransformer architecture, which has shown great potential in the field of acoustic signature signal processing. However, current Class Token-based transformer acoustic signature diagnostic models still face two major technical bottlenecks: At the data level, the long operating time and low probability of fault occurrence of transformers result in a severe scarcity of fault acoustic signature samples, making it difficult to support sufficient training of deep models; at the same time, acoustic signature data annotation is highly dependent on the experience of domain experts, with high annotation costs and strong subjectivity, which restricts the large-scale application of the model. At the model level, the existing single-path Class Token architecture has limited ability to express complex acoustic signature features and struggles to simultaneously capture complementary information from time-domain waveform features and frequency-domain spectral features, resulting in insufficient generalization ability across different operating conditions and equipment models, and consequently, low accuracy and robustness of diagnostic results. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and to provide a transformer acoustic signature diagnosis method and related device based on dual-path time-frequency fusion. This method and related device can diagnose transformer faults with high accuracy and robustness.
[0004] To achieve the above objectives, this invention discloses a transformer acoustic signature diagnostic method based on dual-path time-frequency fusion, comprising: Acquire the acoustic signature signal of the transformer; A dual-path Class Token feature is constructed based on the acoustic signature signal of the transformer; The dual-path Class Token features are subjected to deep interaction and adaptive fusion to obtain the final multi-scale fused features; The final multi-scale fusion features are input into the trained classifier, and the fault state of the transformer is determined based on the output of the classifier.
[0005] A further improvement of the dual-path time-frequency fusion transformer acoustic signature diagnostic method described in this invention is as follows: Furthermore, the process of constructing the dual-path Class Token feature based on the transformer's acoustic signature signal is as follows: The acoustic signature signal of the transformer is preprocessed to obtain a standardized acoustic signature signal frame sequence; A dual-path input feature is constructed based on the standardized voiceprint signal frame sequence, wherein the dual-path input feature includes a time-domain path input matrix and a frequency-domain path input matrix; The dual-path input features are encoded to obtain dual-path Class Token features.
[0006] Furthermore, the process of preprocessing the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence is as follows: The acoustic signature signal of the transformer is sequentially denoised, normalized, and framed. Based on the results of the framed processing, the standardized acoustic signature signal frame sequence is constructed.
[0007] Furthermore, the process of encoding the dual-path input features to obtain the dual-path Class Token features is as follows: The dual-path input features are input into a dual-path Transformer encoder for position encoding and Transformer encoding to obtain dual-path Class Token features.
[0008] Furthermore, the process of performing deep interaction and adaptive fusion on the dual-path Class Token features to obtain the final multi-scale fused features is as follows: Based on the cross-path attention mechanism, the dual-path Class Token features are deeply interacted to obtain the fused Class Token features. The fused Class Token features are subjected to multi-scale time-frequency feature adaptive fusion to obtain the final multi-scale fused features.
[0009] Furthermore, the loss function of the classifier during training is:
[0010] in, This is the loss value. The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . Represents all weight parameters of the model. This represents the number of fault categories.
[0011] This invention discloses a dual-path time-frequency fusion transformer acoustic signature diagnostic system, comprising: The acquisition module is used to acquire the acoustic signature signal of the transformer; A construction module is used to construct a dual-path Class Token feature based on the acoustic signature signal of the transformer; The fusion module is used to perform deep interaction and adaptive fusion of the dual-path Class Token features to obtain the final multi-scale fused features; The judgment module is used to input the final multi-scale fusion features into the trained classifier and judge the fault state of the transformer based on the output of the classifier.
[0012] A further improvement of the dual-path time-frequency fusion transformer acoustic signature diagnostic system of the present invention is as follows: Furthermore, the building module includes: The preprocessing unit is used to preprocess the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence. The construction unit is used to construct dual-path input features based on the standardized voiceprint signal frame sequence, wherein the dual-path input features include a time-domain path input matrix and a frequency-domain path input matrix; The encoding unit is used to encode the dual-path input features to obtain dual-path Class Token features.
[0013] Furthermore, the process of preprocessing the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence is as follows: The acoustic signature signal of the transformer is sequentially denoised, normalized, and framed. Based on the results of the framed processing, the standardized acoustic signature signal frame sequence is constructed.
[0014] Furthermore, the process of encoding the dual-path input features to obtain the dual-path Class Token features is as follows: The dual-path input features are input into a dual-path Transformer encoder for position encoding and Transformer encoding to obtain dual-path Class Token features.
[0015] Furthermore, the fusion module includes: An interaction unit is used to perform deep interaction on the dual-path Class Token features based on a cross-path attention mechanism to obtain the fused Class Token features. The fusion unit is used to perform multi-scale time-frequency feature adaptive fusion on the fused Class Token features to obtain the final multi-scale fused features.
[0016] Furthermore, the loss function of the classifier during training is:
[0017] in, This is the loss value. The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . Represents all weight parameters of the model. This represents the number of fault categories.
[0018] This invention discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the dual-path time-frequency fusion transformer acoustic signature diagnosis method.
[0019] This invention discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the dual-path time-frequency fusion transformer acoustic signature diagnosis method.
[0020] The present invention has the following beneficial effects: In practical operation, the dual-path time-frequency fusion transformer acoustic signature diagnosis method and related device of the present invention construct dual-path Class Token features based on the acoustic signature signal of the transformer, and perform deep interaction and adaptive fusion on the dual-path Class Token features to obtain the final multi-scale fused features. This solves the problem that the traditional single-path Class Token model is insufficient in expressing complex acoustic signature features. Through deep interaction and adaptive fusion, it effectively integrates acoustic signature features of different resolutions, solves the problem that single-scale features are difficult to fully characterize the transformer fault state, and significantly improves the accuracy and robustness of fault diagnosis. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system structure diagram of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0025] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0026] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.
[0027] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0028] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0030] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0031] Example 1 refer to Figure 1 The dual-path time-frequency fusion transformer acoustic signature diagnostic method of the present invention includes: 1) Acquire the acoustic signature signal of the transformer; 2) Construct a dual-path Class Token feature based on the acoustic signature signal of the transformer; The specific operation of step 2) is as follows: 21) Data preprocessing of audio signals; The acoustic signature signal of the transformer is preprocessed, specifically, the acoustic signature signal of the transformer is denoised, normalized and framed.
[0032] Specifically, the transformer acoustic signature signal is discretized and sampled at a frequency of 16000Hz, then normalized to obtain:
[0033] in, This is the normalized transformer acoustic signature data. These are the raw data values of the transformer acoustic signature signal. and The maximum and minimum values in the transformer acoustic signature signal are respectively used. The normalized transformer acoustic signature signal data is then processed by frame segmentation to obtain:
[0034] in, Let N represent the i-th frame signal, N be the frame length, L be the frame shift, and T be the total signal length. A standardized audioprint signal frame sequence is constructed based on the results of the framing process.
[0035] 22) Construct dual-path input features based on standardized speaker signal frame sequences, wherein the dual-path input features are a time-domain path input matrix and a frequency-domain path input matrix; The specific process of step 22) is as follows: 221) The time-domain path input matrix is constructed as follows:
[0036] Where M is the total number of frames; The input matrix represents the time-domain path, with each row representing a frame of time-domain signal.
[0037] 222) Obtain the frequency domain path input matrix; The specific operation of step 222) is as follows: 2221) Window each frame of signal and perform a Fast Fourier Transform:
[0038] in, For Hamming window functions, , Indicates the first Frame signal in the The amplitude spectrum at each frequency point.
[0039] 222) The results of the Fast Fourier Transform are processed using a Mel filter bank;
[0040] in, For the first The response of a Mel filter,
[0041] 2223) Construct the Mel spectrum;
[0042] in, The number of cepstral coefficients at the Mel frequency. Indicates the first The first frame of the signal Mel frequency cepstral coefficients.
[0043] 224) Based on the Mel spectrogram, construct the frequency domain path input matrix. for:
[0044] in, Indicates the first All MFCC coefficient vectors of the frame, frequency domain path input matrix Each row in the vector represents the MFCC feature vector of a frame of signal.
[0045] 23) Extract dual-path Class Token features; Construct a dual-path Transformer encoder, inputting the dual-path input features into the dual-path Transformer encoder to obtain dual-path Class Token features. Specifically, step 23) involves the following steps: Construct a two-path Transformer encoder; Class Token initialization:
[0046] in, For feature dimension, These are learnable Class Tokens for the time-domain path and the frequency-domain path, respectively, used to aggregate global feature information for their respective paths.
[0047] Location coding for:
[0048] in, Indicates the position in the sequence. Indicates a dimension index. This is the position encoding matrix, used to provide sequence order information for the Transformer.
[0049] The Transformer encoding process is as follows:
[0050]
[0051] in, The embedding matrix of the time-domain path, For the number of Transformer layers, Indicates the first The output of the time-domain path corresponds to the Class Token output at the first position. The frequency-domain path is processed similarly, using... As input, this step uses a dual-path Transformer encoder to extract the temporal dynamic features and frequency structural features of the voiceprint signal, respectively, to provide high-quality feature representations for subsequent cross-path interactions.
[0052] 3) Perform deep interaction and adaptive fusion on the dual-path Class Token features to obtain the final multi-scale fusion features; The specific operation of step 3) is as follows: 31) Based on the cross-path attention mechanism, achieve deep interaction of path Class Token features to obtain fused Class Token features; The specific process of step 31) is as follows: 311) Extract the respective Class Tokens from the dual-path Transformer encoder:
[0053]
[0054] in, and These represent the time-domain path and the frequency-domain path respectively. Output of the Class Token after layer Transformer encoding.
[0055] 312) Calculate cross-attention for:
[0056] in, For querying the matrix, The key matrix, For value matrices, For a learnable projection matrix, This refers to the dimension of attention mechanisms.
[0057] 313) Calculate the characteristics of the merged Class Token for:
[0058] in, Presentation layer normalization operation, This is a random deactivation operation. This refers to the characteristics of the merged Class Token.
[0059] 32) Based on multi-scale time-frequency feature adaptive fusion, a multi-scale feature fusion module is constructed. By fusing the fused Class Token features, time-frequency features at different levels are integrated. The specific operation of step 32) is as follows: 321) Extracting multi-scale features :
[0060] in, Indicates the first The Class Token output of the Transformer layer. It is a multi-scale feature matrix, containing feature representations of shallow, medium and high levels.
[0061] 322) Calculate attention weights for:
[0062] in, The weight matrix is a learnable matrix. For the first Attention vectors of various scales For the first Attention weights for each scale feature, satisfying .
[0063] 323) Weighted fusion is used to obtain the final multi-scale fusion features. for:
[0064] It should be noted that this invention dynamically adjusts the importance of features at different scales through an adaptive attention mechanism, enabling the model to select the most relevant feature scale according to the specific fault type, thereby enhancing the model's expressive power and generalization performance.
[0065] 4) Construct a classifier, train the classifier, and then combine the final multi-scale fused features. The input is fed into a trained classifier for classification to determine the fault state of the transformer, which is a normal state, a partial discharge fault state, an overheating fault state, a discharge fault state, or other types of fault states.
[0066] The specific operation of step 4) is as follows: Classification output results for:
[0067] in, This is the classification weight matrix. For bias vectors, Number of fault categories This represents the predicted failure probability distribution.
[0068] The loss function for this classifier during training is:
[0069] in, The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . This represents all the weight parameters of the model. The classifier is trained using the backpropagation algorithm and a gradient descent optimizer (such as AdamW), by minimizing the loss function. To update the classifier parameters, early stopping strategy and learning rate decay are used during training to improve training efficiency and model performance.
[0070] The trained classifier is used for transformer fault diagnosis, namely:
[0071] in, The final fault diagnosis result is the fault category with the highest probability.
[0072] It should be noted that this invention improves the feature representation capability of traditional single-path Class Token models by designing a time-frequency domain dual-path Class Token architecture, constructing Transformer encoders for both time-domain and frequency-domain paths, and establishing a cross-path attention interaction mechanism between the two paths. This constructs a dual-path feature complementary model, further enhancing the model's comprehensive representation capability of the time-frequency features of acoustic signature signals, solving the problem of insufficient representation capability of traditional single-path models for complex acoustic signature features, and improving the accuracy and robustness of fault diagnosis models. Furthermore, this invention constructs a multi-scale feature extraction module to extract multi-level feature representations from Transformer layers of different depths, and designs a feature-level attention weighting mechanism to adaptively integrate time-frequency features of different scales. This can dynamically adjust the importance weights of features at each scale according to the specific fault type, solving the problem that single-scale features cannot comprehensively represent the complex fault states of transformers, and significantly improving the generalization performance and adaptability of fault diagnosis models under different operating conditions.
[0073] Example 2 The dual-path time-frequency fusion transformer acoustic signature diagnostic system of the present invention includes: The acquisition module is used to acquire the acoustic signature signal of the transformer; A construction module is used to construct a dual-path Class Token feature based on the acoustic signature signal of the transformer; The fusion module is used to perform deep interaction and adaptive fusion of the dual-path Class Token features to obtain the final multi-scale fused features; The judgment module is used to input the final multi-scale fusion features into the trained classifier and judge the fault state of the transformer based on the output of the classifier.
[0074] In this embodiment, the building module includes: The preprocessing unit is used to preprocess the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence. The construction unit is used to construct dual-path input features based on the standardized voiceprint signal frame sequence, wherein the dual-path input features include a time-domain path input matrix and a frequency-domain path input matrix; The encoding unit is used to encode the dual-path input features to obtain dual-path Class Token features.
[0075] In this embodiment, the process of preprocessing the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence is as follows: The acoustic signature signal of the transformer is sequentially denoised, normalized, and framed. Based on the results of the framed processing, the standardized acoustic signature signal frame sequence is constructed.
[0076] In this embodiment, the process of encoding the dual-path input features to obtain dual-path Class Token features is as follows: The dual-path input features are input into a dual-path Transformer encoder for position encoding and Transformer encoding to obtain dual-path Class Token features.
[0077] In this embodiment, the fusion module includes: An interaction unit is used to perform deep interaction on the dual-path Class Token features based on a cross-path attention mechanism to obtain the fused Class Token features. The fusion unit is used to perform multi-scale time-frequency feature adaptive fusion on the fused Class Token features to obtain the final multi-scale fused features.
[0078] In this embodiment, the loss function of the classifier during the training process is:
[0079] in, This is the loss value. The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . Represents all weight parameters of the model. This represents the number of fault categories.
[0080] The module division in this embodiment is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in each embodiment of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0081] Example 3 A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the dual-path time-frequency fusion transformer acoustic signature diagnostic method, for example, including: acquiring the acoustic signature signal of the transformer; constructing dual-path Class Token features based on the acoustic signature signal of the transformer; performing deep interaction and adaptive fusion on the dual-path Class Token features to obtain final multi-scale fusion features; inputting the final multi-scale fusion features into a trained classifier; and determining the fault state of the transformer based on the output of the classifier. The memory may include main memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device. The processor, network interface, and memory are interconnected via an internal bus, which may be an industry standard architecture bus, a peripheral component interconnection standard bus, an extended industry standard architecture bus, etc., and the bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store the program; specifically, the program may include program code, which includes computer operation instructions. The memory may include main memory and non-volatile memory, and provides instructions and data to the processor.
[0082] Example 4 A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the dual-path time-frequency fusion transformer acoustic signature diagnostic method. For example, the method includes: acquiring the acoustic signature signal of the transformer; constructing dual-path Class Token features based on the transformer's acoustic signature signal; performing deep interaction and adaptive fusion on the dual-path Class Token features to obtain a final multi-scale fusion feature; inputting the final multi-scale fusion feature into a trained classifier; and determining the transformer's fault state based on the classifier's output. Specifically, the computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0083] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0084] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0085] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0086] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0087] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and disclosure of the invention. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.
[0088] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
[0089] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any simple modifications, alterations, or equivalent structural changes made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A transformer acoustic signature diagnostic method using dual-path time-frequency fusion, characterized in that, include: Acquire the acoustic signature signal of the transformer; A dual-path Class Token feature is constructed based on the acoustic signature signal of the transformer; The dual-path Class Token features are subjected to deep interaction and adaptive fusion to obtain the final multi-scale fused features; The final multi-scale fusion features are input into the trained classifier, and the fault state of the transformer is determined based on the output of the classifier.
2. The transformer acoustic signature diagnostic method based on dual-path time-frequency fusion according to claim 1, characterized in that, The process of constructing the dual-path Class Token feature based on the transformer's acoustic signature signal is as follows: The acoustic signature signal of the transformer is preprocessed to obtain a standardized acoustic signature signal frame sequence; A dual-path input feature is constructed based on the standardized voiceprint signal frame sequence, wherein the dual-path input feature includes a time-domain path input matrix and a frequency-domain path input matrix; The dual-path input features are encoded to obtain dual-path Class Token features.
3. The transformer acoustic signature diagnostic method based on dual-path time-frequency fusion according to claim 2, characterized in that, The process of preprocessing the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence is as follows: The acoustic signature signal of the transformer is sequentially denoised, normalized, and framed. Based on the results of the framed processing, the standardized acoustic signature signal frame sequence is constructed.
4. The transformer acoustic signature diagnostic method based on dual-path time-frequency fusion according to claim 2, characterized in that, The process of encoding the dual-path input features to obtain dual-path Class Token features is as follows: The dual-path input features are input into a dual-path Transformer encoder for position encoding and Transformer encoding to obtain dual-path Class Token features.
5. The transformer acoustic signature diagnostic method based on dual-path time-frequency fusion according to claim 1, characterized in that, The process of performing deep interaction and adaptive fusion on the dual-path Class Token features to obtain the final multi-scale fused features is as follows: Based on the cross-path attention mechanism, the dual-path Class Token features are deeply interacted to obtain the fused Class Token features. The fused Class Token features are subjected to multi-scale time-frequency feature adaptive fusion to obtain the final multi-scale fused features.
6. The transformer acoustic signature diagnostic method based on dual-path time-frequency fusion according to claim 1, characterized in that, The loss function of the classifier during training is: in, The loss value. The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . Represents all weight parameters of the model. This represents the number of fault categories.
7. A dual-path time-frequency fusion transformer acoustic signature diagnostic system, characterized in that, include: The acquisition module is used to acquire the acoustic signature signal of the transformer; A construction module is used to construct a dual-path Class Token feature based on the acoustic signature signal of the transformer; The fusion module is used to perform deep interaction and adaptive fusion of the dual-path Class Token features to obtain the final multi-scale fused features; The judgment module is used to input the final multi-scale fusion features into the trained classifier and judge the fault state of the transformer based on the output of the classifier.
8. The transformer acoustic signature diagnostic system based on dual-path time-frequency fusion according to claim 7, characterized in that, The building module includes: The preprocessing unit is used to preprocess the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence. The construction unit is used to construct dual-path input features based on the standardized voiceprint signal frame sequence, wherein the dual-path input features include a time-domain path input matrix and a frequency-domain path input matrix; The encoding unit is used to encode the dual-path input features to obtain dual-path Class Token features.
9. The transformer acoustic signature diagnostic system based on dual-path time-frequency fusion according to claim 8, characterized in that, The process of preprocessing the acoustic signature signal of the transformer to obtain a standardized acoustic signature signal frame sequence is as follows: The acoustic signature signal of the transformer is sequentially denoised, normalized, and framed. Based on the results of the framed processing, the standardized acoustic signature signal frame sequence is constructed.
10. The transformer acoustic signature diagnostic system based on dual-path time-frequency fusion according to claim 8, characterized in that, The process of encoding the dual-path input features to obtain dual-path Class Token features is as follows: The dual-path input features are input into a dual-path Transformer encoder for position encoding and Transformer encoding to obtain dual-path Class Token features.
11. The transformer acoustic signature diagnostic system based on dual-path time-frequency fusion according to claim 7, characterized in that, The fusion module includes: An interaction unit is used to perform deep interaction on the dual-path Class Token features based on a cross-path attention mechanism to obtain the fused Class Token features. The fusion unit is used to perform multi-scale time-frequency feature adaptive fusion on the fused Class Token features to obtain the final multi-scale fused features.
12. The transformer acoustic signature diagnostic system based on dual-path time-frequency fusion according to claim 7, characterized in that, The loss function of the classifier during training is: in, The loss value. The number of samples in the batch. For the first The true label of each sample The corresponding predicted probability, The regularization coefficient is . Represents all weight parameters of the model. This represents the number of fault categories.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the transformer acoustic signature diagnosis method with dual-path time-frequency fusion as described in any one of claims 1-6.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the transformer acoustic signature diagnosis method with dual-path time-frequency fusion as described in any one of claims 1-6.