Low-rank fine-tuning transformer fault diagnosis method based on adaptive attention guidance

By using an adaptive attention-guided low-rank fine-tuning method, computational resource allocation and model performance are optimized, solving the problems of high computational resource consumption and insufficient context adaptation of large language models in transformer fault diagnosis, and achieving efficient fault diagnosis.

CN120873758AActive Publication Date: 2025-10-31QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Patent Information

Application Number
CN202511366711.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-10-31
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing large language models suffer from high computational resource consumption and insufficient context adaptation in transformer fault diagnosis tasks, resulting in unstable fine-tuning effects.

Method used

We employ an adaptive attention-guided low-rank fine-tuning method, which introduces an attention scoring mechanism, a dynamic rank allocation strategy, and a hierarchical learning rate adjustment mechanism to construct a lightweight adaptation framework, thereby optimizing computational resource allocation and model performance.

Benefits of technology

While reducing computational costs, it improves the accuracy of transformer fault diagnosis and the adaptability of the model, making it suitable for industrial applications with large amounts of data and limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873758A_ABST
    Figure CN120873758A_ABST
Patent Text Reader

Abstract

The invention relates to a low-rank fine-tuning transformer fault diagnosis method based on adaptive attention guidance, and belongs to the technical field of artificial intelligence. An attention scoring mechanism, an adaptive attention scoring mechanism, a dynamic rank allocation strategy and a hierarchical learning rate adjustment mechanism are introduced, and a context-aware dynamic updating strategy is further fused, so that the low-rank fine-tuning transformer fault diagnosis method based on adaptive attention guidance is realized. According to updating of real-time performance, loss and gradient dynamic intelligent triggering key parameters in the model training process, a large-model lightweight adaptation frame suitable for a transformer fault diagnosis task is constructed, and on the premise that diagnosis accuracy is ensured, model fine adjustment and deployment cost is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning. Background Technology

[0002] With the accelerating digitalization and intelligentization of power systems, the importance of transformers, as core equipment in power grid operation, for monitoring their operational status and diagnosing faults is becoming increasingly prominent. Currently, power companies have accumulated a large amount of textual data about transformer operation, such as inspection logs, fault reports, maintenance records, and alarm records. This data contains rich tacit knowledge and expert experience. However, how to utilize this unstructured information for intelligent analysis remains a critical technical bottleneck that urgently needs to be overcome in the field of power operation and maintenance.

[0003] In recent years, Large Language Models (LLMs) have made significant progress in natural language processing tasks and have been gradually applied to scenarios such as fault question answering and operation and maintenance knowledge extraction. Although general pre-trained models have strong semantic understanding capabilities, their direct application to transformer fault diagnosis tasks still faces significant challenges: on the one hand, the terminology in the power industry is special and the context is complex, resulting in a lack of context adaptation ability in general models; on the other hand, fine-tuning large models requires a large amount of computing resources, which is not conducive to deployment in actual industrial edge devices.

[0004] Therefore, how to reduce fine-tuning costs while maintaining model performance and quickly adapting it to the data characteristics of the transformer field has become a current research hotspot and challenge. To address these issues, the industry has gradually explored lightweight fine-tuning strategies, such as Parameter Efficient Fine-tuning (PEFT) and Low-Rank Adaptation (LoRA). However, these methods often fail to fully utilize the differences in importance of input data at different levels or lack flexibility in parameter allocation, leading to unstable fine-tuning results or performance bottlenecks. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes an adaptive attention-guided low-rank fine-tuning method, particularly suitable for transformer fault diagnosis, especially in power systems, such as transformer fault detection and early warning. This method effectively improves the performance of pre-trained models in transformer fault diagnosis tasks by fine-tuning them, making it especially suitable for industrial applications with large amounts of data and limited computational resources.

[0006] This invention introduces an attention scoring mechanism, an adaptive attention scoring mechanism, a dynamic rank allocation strategy, and a hierarchical learning rate adjustment mechanism, and further integrates a context-aware dynamic update strategy. Based on the real-time performance, loss, and gradient dynamic intelligent triggering of key parameter updates during model training, it constructs a lightweight adaptation framework for large models suitable for transformer fault diagnosis tasks. This significantly reduces the cost of model fine-tuning and deployment while ensuring diagnostic accuracy.

[0007] The technical solution of this invention is as follows: The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning includes the following steps: S1, Data Acquisition and Preprocessing Stage; S101. Collect relevant operation and maintenance data of transformers in the target power system; S102. Standardize and clean the collected data and preprocess it to construct a local transformer fault diagnosis dataset; S2, Initialization phase; S201. Sample data from the constructed local transformer fault diagnosis dataset, perform forward and backward propagation on the original pre-trained language model, and collect gradients, attention activation maps and hidden state information of each layer. S202. Based on the collected information, calculate the gradient importance score, activation importance score, and value change importance score for each layer, and construct a unified attention scoring index based on the learnable weight coefficient combination. S203. Allocate the required rank and learning rate for each layer according to the obtained attention score, and initialize the LoRA parameter matrices A and B of the corresponding layers to achieve the initial configuration of the low-rank structure. S3, Fine-tuning stage; S301. Execute the standard training batch process on the fault diagnosis data, i.e. the transformer's relevant operation and maintenance data, including forward propagation, loss calculation and back propagation, to complete the local update of parameters; S302. Real-time monitoring of the training dynamics of neural network models based on attention mechanisms and low-rank representations, including performance metrics, training loss, and gradient dynamics. S303. Based on the monitored performance and training dynamics of the neural network model based on attention mechanism and low-rank representation, when the preset context-aware triggering conditions are met, dynamically update the attention score and rank allocation strategy to achieve adaptive adjustment of the low-rank structure. S304. At the end of each training round, the overall effect is evaluated based on the performance of the validation set, and the rank allocation strategy is adjusted based on the feedback from the neural network model based on the attention mechanism and low-rank representation. S4, Adaptive Adjustment Phase; S401. Monitor and verify the trend of indicator changes. If the performance improvement is small in multiple consecutive evaluations, automatically adjust key hyperparameters or implement a learning rate restart strategy. S402. If the performance still does not improve significantly after multiple adjustments, the early stop mechanism is triggered, and the model with the best verification performance is selected as the final model. S5, Deployment and Real-time Diagnostics Phase; The finely tuned neural network model based on attention mechanism and low-rank representation is deployed to an edge server to enable online parsing and diagnosis of transformer operation and maintenance logs or real-time monitoring data.

[0008] Preferably, in S101, the relevant operation and maintenance data of the transformer includes inspection records, maintenance logs, alarm information, operating parameters, and environmental parameters in unstructured or semi-structured text data.

[0009] Preferably, in S102, the collected data is standardized, cleaned, and preprocessed; this means that the collected data is sequentially subjected to noise reduction, normalization, missing value filling, classification, and labeling.

[0010] Preferably, in S201, data is sampled from the constructed local transformer fault diagnosis dataset, and forward and backward propagation is performed on the original pre-trained language model to collect gradients, attention activation maps, and hidden state information of each layer; including: Collecting gradients from each layer: During backpropagation, the original pre-trained language model, based on the error calculated by the loss function, is propagated backward from the output layer to the input layer, thereby calculating the gradients of the weight matrices of each layer and collecting gradient information. , where n represents the nth layer, gradient information This indicates the sensitivity of the nth layer parameters to the loss function under the current task; Collect attention activation maps of each layer: The original pre-trained language model synchronously collects key signals during the forward propagation process, that is, extracts attention activation information; the attention activation map, that is, the attention activation information, refers to the attention weight matrix generated by each attention head in the original pre-trained language model when processing the input sequence; Collect hidden state information of each layer; hidden state information refers to the intermediate representation output of each encoder or decoder during the forward propagation of the Transformer model, including the contextual semantic features of the input data after nonlinear transformation in the current layer.

[0011] Preferably, the original pre-trained language model is BERT.

[0012] Preferably, BERT includes 12 encoder layers. The input text is first processed by the embedding layer, and then passed through each encoder layer in sequence. The output of each encoder layer is the hidden state of that encoder layer, which is a three-dimensional tensor of shape [batch_size, seq_len, hidden_size], where batch_size represents the sample batch size; seq_len represents the length of the input sequence; and hidden_size represents the feature dimension of the current layer.

[0013] Preferably, in S202, the gradient importance score G(n) of each layer is calculated using the following formula: ; in, This represents the gradient of the weight matrix of the nth layer. This represents the Frobenius norm, which is the square root of the sum of the squares of all elements in the corresponding matrix.

[0014] Preferably, in S202, the activation importance score is calculated using the activation map of the self-attention mechanism, and the calculation steps are as follows: Standardize by the number of heads, divide the total difference by the number of attention heads in the layer, and calculate the average single-head difference; the total difference is the sum of the L1 differences between each pair of attention matrices of each attention head in the same layer; Standardize by sequence length, and divide the average single-head difference by the square of the input sequence length; The activation importance score is obtained by processing the result after the sequence length is standardized using a dynamic normalization function.

[0015] Preferably, the result after sequence length standardization is processed using a dynamic normalization function, including: In the early stages of training, i.e. t < 0.4T, where t is the current training step number and T is the total training step number, global maximum normalization is used for processing. During the middle of training, i.e., 0.4T≤t<0.6T, a mixture function is used for processing; In the later stages of training, i.e., when t≥0.6T, quantile normalization is used for processing.

[0016] Preferably, in S202, the importance score V(n) of the value change is calculated using the following formula: ; Where H(n) and H0(n) are the hidden state outputs of the nth layer in the fine-tuning and original models, respectively; H(i) and H0(i) are the hidden state outputs of the ith layer in the fine-tuning and original models, respectively. Alternatively, the importance score V(n) for value change can be calculated using the following formula: ; Where α1, α2, α3 are weighting coefficients, and V(n) represents the importance score of the value change in the nth layer; Preferably, in S202, the attention score S(n) is calculated using the following formula: Where G(n) represents the gradient importance score of the nth layer, A(n) represents the activation importance score of the nth layer, and V(n) represents the value change importance score of the nth layer. These are the learnable weight coefficients.

[0017] Preferably, in S203, resource allocation is performed based on attention score S(n), according to a dynamic rank allocation strategy and hierarchical self-allocation. The adaptive learning rate mechanism allocates the rank r(n) and learning rate lr(n) for each layer, including the following steps: Basic parameter settings and rank and learning rate allocation: Set basic parameters, including the basic rank value. Maximum rank Scaling factor Basic learning rate and learning rate adjustment factor Based on the calculated attention score S(n), a rank value r(n) and a learning rate lr(n) are assigned to each layer; Rank allocation: The rank r(n) is allocated using the following formula: ; Learning rate allocation: The hierarchical adaptive learning rate lr(n) is allocated using the following formula: ; For each layer, initialize the corresponding LoRA parameter matrices A and B according to the assigned rank value, as follows: LoRA parameter matrix initialization: The corresponding LoRA parameter matrices A and B are initialized according to the assigned rank values. Matrix A is initialized using a normal distribution, and matrix B is initialized with all zeros.

[0018] Preferably, in S203, based on the performance of the neural network model based on the attention mechanism and low-rank representation after every N training batches, the attention score and rank allocation strategy are periodically updated to achieve dynamic adaptation of the low-rank structure; including: Update the attention scores of each layer using a sliding window as the average: ; S(l) oldLet S(l) represent the attention score at the last update of layer l. current This represents the attention that was just calculated in the l-th layer during the current training cycle; Periodically update the rank allocation strategy, including: The total rank budget constraint is: ; Rank calculation logic: ; The LoRA matrix dimensions are adaptively adjusted using the following strategy: Rank expansion: When increasing from r to r+Δr, the newly added parameter matrix block is orthogonally initialized and inherits the principal components of the original matrix through SVD decomposition; Rank contraction: When increasing from r to r-Δr, retain the eigencomponents corresponding to the first Δr largest singular values ​​in the singular value decomposition.

[0019] Preferably, in S4, the adaptive adjustment phase includes: When the model performance improvement is detected to be small or stagnant during the fine-tuning process, that is, when the improvement of the evaluation index is less than the preset threshold for N consecutive times, it is determined that the performance improvement is small and the model may be trapped in a local optimum; the key hyperparameters are automatically adjusted or the learning rate restart strategy is implemented.

[0020] If the performance improvement is still minimal after multiple adjustments, the early stopping mechanism is triggered, and the model with the best verification performance is selected as the final model.

[0021] Preferably, in S5, the deployment and real-time diagnostic phase includes: ① Edge computing device deployment: The finely tuned low-rank adaptation model is deployed on the NVIDIA Jetson Xavier AGX edge computing device; ② Data Input and Real-time Monitoring: Real-time monitoring data of the transformer is collected through sensors, and the sensors transmit the real-time data to the edge server via the Modbus TCP / IP protocol; the real-time monitoring data of the transformer includes: input voltage, output current, oil temperature, load current, and alarm status; ③ Real-time fault diagnosis and early warning: The fine-tuned neural network model based on attention mechanism and low-rank representation processes the real-time monitoring data of the transformer after receiving it; specifically, it includes: receiving multi-source data input from the transformer online monitoring system, converting the multi-source data input into a standard format acceptable to the neural network model based on attention mechanism and low-rank representation according to the preset data preprocessing process, and constructing a time series input feature tensor. The neural network model based on attention mechanism and low-rank representation integrates historical data distribution with current input features, uses attention mechanism to capture the coupling relationship and dynamic change trend between key variables, and uses the enhanced feature expression ability of low-rank structure to judge the current state. At the output end, the neural network model based on attention mechanism and low-rank representation classifies or grades the operating state of transformer, and can output specific fault type and confidence score in combination with abnormal trends. Meanwhile, the neural network model based on attention mechanism and low-rank representation supports a dynamic scoring mechanism based on sliding time window, which continuously judges the health status trend at multiple time points, and realizes early identification and warning of potential faults; if the model judges the state to be abnormal multiple times in a row, the warning mechanism is triggered. The diagnostic results include: Fault prediction: If the transformer current or oil temperature exceeds the set threshold, an early warning will be automatically issued to indicate that there is a risk of transformer failure. Fault type identification: The finely tuned neural network model based on attention mechanism and low-rank representation identifies fault types, including overload, short circuit, and abnormal temperature; the specific judgment criteria are as follows: If the real-time current exceeds the maximum allowable current, it is determined to be an overload fault; If the oil temperature exceeds the set threshold, it is determined to be an abnormal temperature. ④ Real-time reasoning and decision support: When a transformer fails or has a potential failure, the edge server generates a fault diagnosis report, which includes the fault type, fault time, fault location, and fault severity.

[0022] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning.

[0023] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning.

[0024] The beneficial effects of this invention are as follows: By dynamically allocating rank values, computational resource allocation is optimized, resulting in better performance with the same number of parameters and achieving higher parameter efficiency. Based on the hierarchical learning rate mechanism, the optimization of important parameters can be accelerated, training time can be shortened, and convergence speed can be improved. By identifying key parameters of the task through the attention mechanism, the adaptability of the model to specific tasks is improved. In addition, the optimized rank allocation reduces the overall number of parameters, lowers storage requirements, and is more suitable for resource-constrained scenarios. Attached Figure Description

[0025] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below.

[0026] Figure 1 This is a flowchart illustrating an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 2 This is a flowchart illustrating the data acquisition and preprocessing stage of an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 3 This is a flowchart illustrating the initialization phase of an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 4 This is a flowchart illustrating the fine-tuning training phase of an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 5 This is a flowchart illustrating the adaptive adjustment stage in an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 6 This is a flowchart illustrating the deployment and real-time diagnosis phases of an adaptive attention-guided low-rank fine-tuning transformer fault diagnosis method proposed in this invention. Figure 7 This is a diagram of the encoder-decoder architecture in which the scoring function S(n) is applied. Figure 8 This is a diagram of the dynamic rank r(n) acting on the encoder-decoder architecture in this invention; Figure 9 This is a diagram illustrating the architecture of the encoder-decoder using the hierarchical adaptive learning rate machine in this invention. Detailed Implementation

[0027] The present invention will be further defined below with reference to the accompanying drawings and embodiments, but is not limited thereto.

[0028] Terminology Explanation: 1. Power System: This refers to the organic whole used for power generation, transmission, transformation, distribution, and consumption, including power plants, substations, transmission lines, distribution networks, and terminal electrical equipment. Its main function is to realize the production, transmission, and distribution of electrical energy, ensuring stable and reliable power supply for users in various application scenarios.

[0029] 2. Global Maximum Normalization: This is a common data preprocessing method. Its basic idea is to divide the value of a certain feature in all samples by the maximum value of that feature in the entire dataset, thereby compressing the value range of that feature to the [0, 1] interval. This method is suitable for scenarios where the feature values ​​have a large range, but a uniform scale is desired for processing.

[0030] 3. Quantile normalization: This method standardizes the data by mapping the distribution of statistical feature values ​​within a dataset (such as the median, quartiles, etc.). It maps the feature value of each sample to its quantile position in the entire dataset, effectively reducing the impact of outliers on the data distribution. It is suitable for processing data with skewed distributions or noise interference.

[0031] 4. Mixture function: This refers to a mathematical structure that combines multiple functions of different types to enhance the expressive power of a model. In algorithm or model design, mixture functions are often used to fuse multiple features or process data with different distributions to extract information more comprehensively and improve overall prediction or classification performance.

[0032] 5. PyTorch, the Python version of Torch, is an open-source neural network framework developed by Facebook, specifically designed for GPU-accelerated deep neural network (DNN) programming. Torch is a classic tensor library for manipulating multidimensional matrix data and is widely used in machine learning and other math-intensive applications. PyTorch's computation graph is dynamic and can change in real time according to computational needs.

[0033] Example 1 A transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning, such as Figure 1 As shown, it includes the following steps: S1, Data Acquisition and Preprocessing Stage; such as Figure 2 As shown: S101. Collect relevant operation and maintenance data of transformers in the target power system; S102. Standardize and clean the collected data and preprocess it to construct a local transformer fault diagnosis dataset; S2, Initialization phase; such as Figure 3 As shown: S201. Sample data from the constructed local transformer fault diagnosis dataset, perform forward and backward propagation on the original pre-trained language model, and collect gradients, attention activation maps and hidden state information of each layer. S202. Based on the collected information, calculate the gradient importance score, activation importance score, and value change importance score for each layer, and construct a unified attention scoring index based on the learnable weight coefficient combination. S203. Allocate the required rank and learning rate for each layer according to the obtained attention score, and initialize the LoRA parameter matrices A and B of the corresponding layers to achieve the initial configuration of the low-rank structure. S3, Fine-tuning stage; such as Figure 4 As shown: S301. Execute the standard training batch process on the fault diagnosis data, i.e. the transformer's relevant operation and maintenance data, including forward propagation, loss calculation and back propagation, to complete the local update of parameters; S302. Real-time monitoring of the training dynamics of neural network models based on attention mechanisms and low-rank representations, including performance metrics, training loss, and gradient dynamics. S303. Based on the monitored performance and training dynamics of the neural network model based on attention mechanism and low-rank representation, when the preset context-aware triggering conditions are met, dynamically update the attention score and rank allocation strategy to achieve adaptive adjustment of the low-rank structure. S304. At the end of each training round, the overall effect is evaluated based on the performance of the validation set, and the rank allocation strategy is adjusted based on the feedback from the neural network model based on the attention mechanism and low-rank representation. S4, Adaptive Adjustment Phase; such as Figure 5 As shown: S401. Monitor and verify the trend of indicator changes. If the performance improvement is small in multiple consecutive evaluations, automatically adjust key hyperparameters or implement a learning rate restart strategy. S402. If the performance still does not improve significantly after multiple adjustments, the early stop mechanism is triggered, and the model with the best verification performance is selected as the final model. S5, Deployment and Real-time Diagnostics Phase; such as Figure 6 As shown: S501. Deploy the finely tuned neural network model based on attention mechanism and low-rank representation to the edge server to realize online parsing and diagnosis of transformer operation and maintenance logs or real-time monitoring data.

[0034] Example 2 The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning described in Example 1 differs in that: In S101, the relevant operation and maintenance data of the transformer includes inspection records, maintenance logs, alarm information, operating parameters, and environmental parameters in unstructured or semi-structured text data. Real-time data is obtained from relevant sensors and monitoring equipment in the power system to ensure the timeliness and accuracy of the data.

[0035] In S102, the collected data undergoes standardization cleaning and preprocessing; this involves sequentially performing noise reduction, normalization, missing value imputation, classification, and labeling on the collected data. This results in a localized transformer fault diagnosis dataset that meets the task requirements.

[0036] In S201, data is sampled from the constructed local transformer fault diagnosis dataset, and forward and backward propagation is performed on the original pre-trained language model to collect gradients, attention activation maps, and hidden state information of each layer; including: Collecting gradients from each layer: During backpropagation, the original pre-trained language model, based on the error calculated by the loss function, is propagated backward from the output layer to the input layer, thereby calculating the gradients of the weight matrices of each layer and collecting gradient information. , where n represents the nth layer, gradient information This represents the sensitivity of the nth layer parameters to the loss function under the current task; the specific implementation includes: First, during the forward propagation process, the original pre-trained language model generates prediction results based on the input samples and calculates loss function values ​​(such as cross-entropy loss) to measure the error between the prediction and the true label. Subsequently, during the backpropagation phase, the chain rule is used to propagate the error signal forward sequentially, starting from the partial derivative of the loss function with respect to the output layer, and applying this to the parameters of each layer, including the weight matrix. and bias terms Take the partial derivatives separately, in the form of: ; in, Represents the loss function. Indicates the first Layer error signal, This is the output of the previous layer (or the input of the current layer). In its implementation, PyTorch automatically constructs the computation graph during model definition and backpropagation, automatically calculating partial derivatives and accumulating gradients. By calling the `.backward()` method, the system automatically calculates the gradients corresponding to the parameters of each layer along the computation graph and stores them in the `.grad` attribute of each parameter tensor. By collecting this gradient information, the sensitivity of each layer's parameters to the current task loss can be quantified, further used to analyze model structure adaptability or perform subsequent steps such as low-rank decomposition and rank allocation strategies.

[0037] For example, in a language model based on the Transformer architecture, some layers exhibit large gradients when processing text translation tasks, indicating that adjusting the parameters of these layers has a significant impact on the accuracy of the translation results, and these layers require further attention and optimization.

[0038] Collect attention activation maps for each layer: The original pre-trained language model simultaneously collects key signals during the forward propagation process, namely, extracts attention activation information; the attention activation map, or attention activation information, refers to the attention weight matrix generated by each attention head in the original pre-trained language model when processing the input sequence; it reflects the degree of attention the model pays to inputs at different positions in each layer.

[0039] Taking BERT as an example, in each Transformer encoder layer, the input sequence is first mapped into three matrices: query (Q), key (K), and value (V). The formula for calculating the attention weights is:

[0040] ; in, The generated matrix is ​​in After normalization, it becomes an attention activation map, where each term represents the attention score of the current word to words in other positions.

[0041] In practice, activation maps for each layer and each attention head can be collected by calling the model's feedforward function and registering the `attention_weights` output of intermediate layers. For example, in PyTorch, the attention map of the corresponding layer can be extracted using `model.encoder.layer[i].attention.self.get_attention_map()`. The dimensions of these attention weight tensors are typically [batch_size, num_heads, seq_len, seq_len], representing the attention relationship between each head and each position in the input sequence. Collecting these attention activation maps is crucial for subsequent analysis of the model structure and for performing low-rank structure optimization work such as attention pruning and rank adjustment.

[0042] Attention activation information records the degree of correlation between different locations and features when the model processes data. In natural language processing tasks, by analyzing attention activation maps, we can understand the model's focus on different words in the input text when generating text, thus providing a basis for evaluating the importance of each layer in information processing.

[0043] The hidden state information of each layer is collected. Hidden state information refers to the intermediate representation output of each encoder or decoder layer during the forward propagation of the Transformer model, including the contextual semantic features of the input data after nonlinear transformation in the current layer. The hidden state of each layer preserves the feature expression of the input sequence in the current representation space and is the basis for the model to construct semantic representations and make final predictions.

[0044] As the final step in collecting information for forward computation, it is necessary to systematically capture the feature representations of each layer of the model: collecting hidden state information. The hidden state is the output of each layer after feature extraction and transformation of the input information during the data processing process. Collecting the hidden state information H(n) allows us to understand how the model represents the data at different layers. During fine-tuning, by comparing the changes in the hidden state before and after fine-tuning, we can evaluate the role and contribution of each layer in adapting to the new task.

[0045] The original pre-trained language model is BERT.

[0046] For example, in a language model BERT based on the Transformer architecture, when processing English to Chinese text translation tasks, the gradient values ​​of some intermediate layers (such as layers 6 and 9) are relatively large, indicating that the parameter adjustment of these layers has a significant impact on the accuracy of the translation results, and they need to be focused on and optimized in the future.

[0047] The original pre-trained language model refers to a basic model that has already been trained on a large-scale general corpus (such as English Wikipedia or BooksCorpus). The language model used in this invention is BERT, whose network structure consists of 12 stacked Transformer encoder layers. Each layer contains substructures such as multi-head self-attention mechanism, feedforward fully connected network, residual connection and layer normalization, which can effectively model the contextual relationships between words in the input sequence.

[0048] After pre-training, the model undergoes a fine-tuning phase to optimize parameters for a specific task. During fine-tuning, the gradients of each layer's parameters are calculated using the standard backpropagation algorithm to determine the sensitivity of different layers to the task loss, thus supporting subsequent rank adjustment strategies and low-rank modeling steps.

[0049] BERT consists of 12 encoder layers. The input text is first processed by the embedding layer, and then passed through each encoder layer in turn. The output of each encoder layer is the hidden state of that encoder layer, which is a three-dimensional tensor of shape [batch_size, seq_len, hidden_size]. Here, batch_size represents the sample batch size; seq_len represents the length of the input sequence; and hidden_size represents the feature dimension of the current layer (e.g., 768 in BERT-base).

[0050] During implementation, the hidden states of all layers can be collected by calling the model's forward inference function and setting options for outputting intermediate layer results. For example, when using the HuggingFace Transformers framework in PyTorch, setting `output_hidden_states=True` will yield a list of `hidden_states` containing all layer hidden states in the model's output, with the following structure:

[0051] ; in For the output of the embedding layer, arrive These represent the hidden states of each encoder layer. This hidden state information can be used for subsequent analysis of the expressive power and task relevance of each layer, playing a crucial supporting role in structural optimization operations such as hierarchical importance assessment, low-rank dimensionality reduction, and pruning strategy formulation.

[0052] In S202, the gradient importance score G(n) of each layer is calculated using the following formula: ; in, This represents the gradient of the weight matrix of the nth layer. This represents the Frobenius norm, which is the square root of the sum of the squares of all elements in the corresponding matrix. This importance score... It reflects the sensitivity of the parameters of layer n to the loss function under the current task, specifically obtained by normalizing the overall gradient magnitude (gradient norm). The introduction of the Frobenius norm makes this evaluation metric dimensionally consistent and comparable, thus it can be used for importance ranking and low-rank allocation strategies among different layers.

[0053] In S202, the activation importance score is calculated using the activation graph of the self-attention mechanism. The calculation steps are as follows: To measure the differences in attention distribution across different layers and eliminate the influence of the number of attention heads in different layers, the total difference is calculated by standardizing by the number of attention heads in that layer and dividing the total difference by the number of attention heads in that layer. The total difference is the sum of the L1 differences between each pair of attention matrices of each attention head in the same layer. Normalize by sequence length, and divide the average single-head difference by the square of the input sequence length (i.e., the total number of elements in each attention matrix). The activation importance score is obtained by processing the result after the sequence length is standardized using a dynamic normalization function.

[0054] The results of sequence length standardization are processed using a dynamic normalization function, including: In the early stages of training, i.e., t < 0.4T, where t is the current training step and T is the total training step, global maximum normalization is used for processing; this method can preserve extreme signals and help to quickly locate key layers. During the middle of training, i.e., 0.4T≤t<0.6T, a blending function is used to achieve a smooth transition.

[0055] In the later stages of training, i.e., t≥0.6T, quantile normalization is used. The 25th and 75th quantiles of the data are used for normalization via interquartile range to suppress noise interference and prevent overfitting.

[0056] In S202, the importance score V(n) for value change is calculated using the following formula: ; Where H(n) and H0(n) are the hidden state outputs of the nth layer in the fine-tuning and original models, respectively; H(i) and H0(i) are the hidden state outputs of the ith layer in the fine-tuning and original models, respectively. Alternatively, the importance score V(n) for value change can be calculated using the following formula: ; Where α1, α2, α3 are weighting coefficients, and V(n) represents the importance score of the value change in the nth layer; In S202, the attention score S(n) is calculated using the following formula: Where G(n) represents the gradient importance score of the nth layer, A(n) represents the activation importance score of the nth layer, and V(n) represents the value change importance score of the nth layer. These are the learnable weight coefficients.

[0057] Among them, the weighting coefficient These are the learnable parameters trained and updated by the optimizer during the fine-tuning phase (i.e., step S301). This mechanism enables the system to adaptively learn how to weigh different types of importance signals based on feedback from model performance and data characteristics during training, thereby more accurately assessing the true importance of each layer. To ensure the rationality and stability of the learnable weight coefficients, nonnegativity constraints or L1 / L2 regularization can be applied to them.

[0058] Sigmoid is a commonly used activation function. The input to the sigmoid function is any real number, and the output value is between 0 and 1, exhibiting smoothness, continuity, and monotonically increasing characteristics. It compresses the model's output into a probability space and is often used in the output layer of binary classification tasks, or in modules such as mapping attention scores and gating mechanisms, giving the output probabilistic meaning.

[0059] In S203, resource allocation, based on the attention score S(n), assigns the rank value r(n) and learning rate lr(n) of each layer according to the dynamic rank allocation strategy and the hierarchical adaptive learning rate mechanism, including the following steps: Basic parameter settings and rank and learning rate allocation: Set basic parameters, including the basic rank value. (Minimum rank of all levels), maximum rank value Scaling factor Basic learning rate and learning rate adjustment factor Based on the calculated attention score S(n), a rank value r(n) and a learning rate lr(n) are assigned to each layer; After completing the basic parameter settings, the next step is the quantitative allocation of core resources. The first step is rank allocation, where the rank r(n) is allocated using the following formula: ; This formula is based on the construction of a Lagrange function and the solution using the Lagrange multiplier method under the total rank constraint (R is the preset total rank budget) and the constraints of each layer. Layers with high importance will be assigned higher rank values, thereby obtaining a larger parameter space for adjustment; layers with low importance will be assigned lower rank values ​​to achieve optimized utilization of computing resources.

[0060] After the rank allocation is completed, the adaptive allocation of the learning rate parameter is performed: the layer-adaptive learning rate lr(n) is allocated using the following formula: ; This formula introduces an importance weighting mechanism, designs the learning rate as a weakly inverse proportional function of the gradient magnitude, and balances sensitivity and stability through a linear term, so that the learning rate is positively correlated with the importance score of each layer, which can accelerate the optimization of important parameters.

[0061] After completing the differentiated allocation of rank and learning rate, we proceed to the specific configuration stage of model parameters. For each layer, we initialize the corresponding LoRA parameter matrices A and B according to the allocated rank, as follows: The core operation in the parameter initialization phase is the customized configuration of the LoRA parameter matrices, specifically as follows: LoRA parameter matrix initialization involves initializing the corresponding LoRA parameter matrices A and B according to the assigned rank values. Matrix A is initialized using a normal distribution, which allows the matrix elements to be randomly distributed around the zero mean, helping the model explore different parameter spaces in the early stages of training. Matrix B is initialized with all zeros. During training, matrix B is gradually updated based on the data and the model's learning progress, enabling the model to adapt to new tasks by adding low-rank matrices without changing most of the parameters of the original pre-trained model.

[0062] After completing the initial configuration of the model parameters, the iterative training phase based on dynamic resource allocation is initiated.

[0063] In S203, based on the performance of the neural network model based on the attention mechanism and low-rank representation after every N training batches, the attention score and rank allocation strategy are periodically updated to achieve dynamic adaptation of the low-rank structure; including: Update the attention scores of each layer using a sliding window as the average: ; S(l) old Let S(l) represent the attention score at the last update of layer l. current This represents the attention that was just calculated in the l-th layer during the current training cycle; Periodically update the rank allocation strategy, including: The total rank budget constraint is: (Remain unchanged); Rank calculation logic: ; After determining the new rank values ​​for each layer, an adaptive adjustment of the LoRA matrix dimensions is performed, with the following specific strategy: Rank expansion: When the rank increases from r to r+Δr, the newly added parameter matrix block is orthogonally initialized and inherits the principal components of the original matrix through SVD decomposition. When the rank increases from r to r+Δr (where r represents the rank value of the current layer), a new parameter matrix block needs to be introduced for the newly added Δr rank components. The new parameter block is orthogonally initialized to maintain numerical stability, and at the same time, by performing SVD (Singular Value Decomposition) on the original matrix, its principal components (i.e., the eigendirections corresponding to the largest singular values) are used as the basis for the new matrix to inherit existing knowledge.

[0064] Rank contraction: When increasing from r to r-Δr, retain the eigencomponents corresponding to the first Δr largest singular values ​​in the singular value decomposition.

[0065] By performing SVD decomposition on the original weight matrix, only the feature components corresponding to the first r−Δr largest singular values ​​are retained, that is, the most representative subspace information is retained, thereby minimizing the accuracy loss.

[0066] In step S301, after obtaining the hierarchical attention score S(n), some or all of the parameters in the LoRA module and the Transformer model are fine-tuned. During this stage, the model iteratively updates the LoRA parameters, other trainable parameters of the Transformer model, and the learnable weight coefficients from step S202, based on the loss function of the diagnostic task, using gradient descent or other optimization algorithms. This iterative process aims to minimize the losses in the diagnostic task and improve the model's performance in transformer fault diagnosis.

[0067] Specifically, the update includes the following parameters: ①LoRA Parameters: These mainly refer to the low-rank decomposition matrices A and B added to the attention layers and feedforward layers of the Transformer model. These matrices are the core of this invention for achieving efficient parameter fine-tuning. For example, for linear layers in the Transformer... Its update is frozen during LoRA fine-tuning, but by introducing two low-rank matrices, its update is modeled as follows: During fine-tuning, only the parameters of matrices A and B need to be trained, and their dimensions and initial assignments are controlled by the dynamic rank allocation strategy in step S203. The original pre-trained Transformer weights (such as BERT's self-attention weights, feedforward network weights, etc.) are usually kept frozen during fine-tuning, which greatly reduces the number of parameters that need to be updated and the computational resource consumption.

[0068] ② Additional Trainable Parameters within the Transformer Model Itself: While the core of LoRA lies in freezing the original pre-trained weights, a small number of other parameters in the Transformer model, besides the main weight matrix, are sometimes fine-tuned to further optimize performance or adapt to specific needs. These may include, but are not limited to, the bias terms of each linear layer in the Transformer (such as Q, K, V projection layers, output projection layers, and linear layers in feedforward networks); the layer normalization parameters of each layer's normalization operation in the Transformer model, i.e., their scaling factor (gamma) and offset (beta); and task-specific output layers / heads added to map the feature representation of the Transformer model to specific transformer fault diagnosis results. The weights and biases of these task-specific output layers (e.g., a linear layer followed by a Softmax activation function to output the probabilities of different fault types) are fully trainable parameters used to transform the semantic features of the Transformer output into a final fault type prediction or diagnostic score.

[0069] ③ Learnable Weight Coefficients: These are the weights used in step S202 to combine the scores for gradient importance, activation importance, and value change importance. These coefficients are also optimized as model parameters during fine-tuning, enabling the model to adaptively learn how to weigh different types of importance signals in order to more accurately assess the true importance of each layer.

[0070] Specifically, the iterative process is as follows: ① Batch Loading: Load a mini-batch of data containing multiple training samples from the constructed localized transformer fault diagnosis dataset. Each sample includes transformer operation and maintenance data (such as text logs, operating parameters, etc.) and its corresponding real fault label.

[0071] ② Forward Pass: The loaded batch of data, used as input, first passes through the embedding and positional encoding layers of the Transformer model. Subsequently, the data sequentially passes through the encoder layers (or decoder layers, depending on the specific architecture) of the Transformer model. In each layer, the frozen parameters of the original pre-trained model are used to perform core computations (such as self-attention computation and feedforward network transformation), while the LoRA module fine-tunes the feature representation by adding the product of the input and low-rank matrices A and B to the output of the original weight matrix. For example, when an input tensor X passes through a LoRA-enhanced linear layer, its output is... After processing by all Transformer layers, the semantic features of the output are passed to a task-specific output layer (such as a classification head). The task-specific output layer transforms these features into predictions or diagnostic scores for transformer fault types (e.g., probability distributions of various fault types).

[0072] ③ Loss Calculation: The model's forward propagation predictions (e.g., predicted fault type probabilities) are compared with the true fault labels in the batch of data. The current prediction error is calculated based on a predefined loss function (e.g., cross-entropy loss for multi-class tasks, which quantifies the difference between the predicted probability distribution and the true labels). The smaller the loss value, the closer the model's prediction is to reality.

[0073] ④ Backward Propagation: Based on the calculated loss value, the error signal is propagated backward from the loss function to all trainable parameters involved in the calculation through automatic differentiation and the chain rule. During this process, the system automatically calculates the loss function for all trainable parameters (including LoRA parameters A and B, task-specific output layer parameters, and learnable weight coefficients). The gradient. Note that since the original pre-trained weights are frozen, they do not participate in gradient calculation and updates, thus avoiding significant computational overhead.

[0074] ⑤ Parameter Update: The optimizer uses the gradient calculated by backpropagation, combined with a preset learning rate (including the hierarchical adaptive learning rate defined in S203), to update all trainable parameters according to its specific optimization strategy. For example, in the gradient descent algorithm, the parameters are adjusted in small steps along the negative direction of the gradient (i.e., the direction in which the loss function decreases the fastest). This step aims to reduce the value of the loss function, enabling the model to make more accurate predictions in the next iteration.

[0075] ⑥ Iterative Repetition: The process of "data batch loading" to "parameter update" described above will be repeated until all batches of data in the entire training dataset have been trained once, marking the completion of one "training epoch". This training cycle will continue for multiple training epochs. After each training epoch, the model's performance is typically evaluated on an independent validation set (as described in S304). The training process will continue until the model performance reaches a satisfactory level on the validation set, converges, or triggers an early stopping mechanism (as described in S402) to prevent overfitting.

[0076] In S302, while executing the training batch process (S301), the system monitors several key metrics of model training in real time to provide contextual information to guide subsequent dynamic adjustments. These monitored metrics include performance metrics, training loss, and gradient dynamics. Specifically, every 500 training batches, the system evaluates the model's diagnostic accuracy and F1 score on an independent validation set and records the performance data from the last five batches. Simultaneously, the system continuously monitors the value of the training loss function and its rate of change, calculating the average and standard deviation of the training loss over the last 100 training batches. Furthermore, the system monitors the average gradient norm of the LoRA modules in all encoder and decoder layers of the model to detect trends in their changes.

[0077] In S303, instead of the original fixed periodic update mechanism, this invention dynamically updates the attention score and rank allocation strategy based on the model performance, training loss, and gradient dynamics monitored in real time in S302, when preset context-aware trigger conditions are met, to achieve adaptive adjustment of the low-rank structure. Once any trigger condition is met, the system will automatically perform the following operations: recalculate the attention score (corresponding to step S202), and reallocate the rank values ​​and learning rates of each layer (corresponding to step S203), thereby achieving dynamic adaptation of the low-rank structure. The trigger conditions are as follows:

[0078] If the improvement in diagnostic accuracy is less than 0.01% in 5 consecutive evaluations on the validation set.

[0079] If the standard deviation of the training loss in the most recent 100 batches is less than 0.001.

[0080] If the validation loss starts to rise three times in a row, or if the average gradient norm of any key layer LoRA module drops sharply by more than 90% of its historical average within 50 batches.

[0081] S303 specifically includes the following update operations: Attention score update: The attention score S(n) for each layer is updated using an exponentially weighted moving average (EWMA) method, which integrates historical data and scores calculated from the latest batch. The specific formula is as follows:

[0082]

[0083] in, Set to 0.1.

[0084] Rank allocation strategy update: Total rank budget constraint: The total effective rank budget of the model is set to 256.

[0085] Rank calculation logic: Based on the updated attention scores and total rank budget, the rank values ​​of each layer are redistributed. First, the attention scores of all layers are normalized:

[0086]

[0087] Then, the rank value is assigned to each level n. The calculation is as follows:

[0088] in, The total rank budget refers to the upper limit of the sum of ranks that can be allocated to all layers modified by LoRA techniques during model fine-tuning.

[0089] Ensure that the rank of each layer is at least 4. Finally, for all layers... Make adjustments to ensure No more than .

[0090] Perform adaptive adjustment of LoRA matrix dimensions: Dynamically adjust the dimensions of the corresponding layer LoRA parameter matrices A and B based on the reallocated r_n.

[0091] Layer-by-layer learning rate adjustment: Learning rate of each LoRA module layer According to its assigned rank Adjustments will be made.

[0092] S304. Evaluation and Decision: At the end of each epoch, the overall performance is evaluated based on the diagnostic performance of an independent validation set. Based on the model's diagnostic performance feedback, determine whether any of the following conditions are met:

[0093] If the accuracy of the validation set does not improve for three consecutive epochs.

[0094] If the validation set accuracy reaches 98%.

[0095] The system makes a decision when any of the above conditions are met. If the validation set accuracy does not improve for three consecutive epochs, the current learning rate is adjusted to 0.5 times the original learning rate.

[0096] In S4, the adaptive adjustment phase includes: The system monitors and verifies the trend of the performance indicators. If, after the context-aware dynamic update in S303, the model performance improvement is still small or stagnant (i.e., the performance improvement in N consecutive evaluations is less than a preset threshold, such as the improvement in the indicator in 5 consecutive evaluations being less than the preset threshold of 0.5%), it is determined that the performance improvement is small and the model may be trapped in a local optimum. At this time, the system will automatically adjust the hyperparameters used to optimize the learnable weight coefficients (e.g., prioritize adjusting their specific learning rate or reinitialization strategy), or adjust other key hyperparameters (e.g., LoRA scaling factor α and learning rate adjustment factor β). The adjustment range can be set according to experience, or a learning rate restart strategy can be implemented to try to break out of the current predicament.

[0097] If performance improvement remains minimal after multiple adjustments, an early stopping mechanism is triggered, and the model with the best validation performance is selected as the final model; priority is given to adjusting the learnable weight coefficients used to optimize them. The hyperparameters (such as their specific learning rate or reinitialization strategy), or other key hyperparameters (such as the LoRA scaling factor α and the learning rate adjustment factor β), can be adjusted empirically. The purpose of this stage is to ensure that the model can continuously learn and optimize its performance on transformer fault diagnosis tasks through intelligent hyperparameter tuning;

[0098] In S5, the deployment and real-time diagnostics phase includes: ① Edge computing device deployment: The finely tuned low-rank adaptation model was deployed on an NVIDIA Jetson Xavier AGX edge computing device. This device has high-performance computing capabilities, using a Volta GPU architecture and 32GB of memory, which can meet the computing requirements of the transformer fault diagnosis model in real-time data processing. The edge server connects to the power system where the transformer is located via Ethernet / IP protocol to receive monitoring data in real time.

[0099] ② Data Input and Real-time Monitoring: Real-time monitoring data of the transformer is collected by sensors, and the sensors transmit the real-time data to the edge server via the Modbus TCP / IP protocol. The real-time monitoring data of the transformer includes: input voltage (unit: volt, V), output current (unit: ampere, A), oil temperature (unit: degree Celsius, ℃), load current (unit: ampere, A), and alarm status (such as overload, short circuit, abnormal temperature, etc.). ③ Real-time fault diagnosis and early warning: The fine-tuned neural network model based on attention mechanism and low-rank representation processes the real-time monitoring data of the transformer after receiving it. Specifically, this includes: receiving multi-source data input from the transformer online monitoring system, such as winding temperature, oil temperature, load current, voltage, frequency, partial discharge signal, vibration signal, and gas content (DGA); converting the multi-source data input into a standard format acceptable to the neural network model based on attention mechanism and low-rank representation according to the preset data preprocessing process (including normalization, noise reduction, outlier removal, etc.), and constructing a time series input feature tensor. The neural network model based on attention mechanism and low-rank representation integrates historical data distribution with current input features, uses attention mechanism to capture the coupling relationship and dynamic trend of key variables, and judges the current state based on the enhanced feature expression ability of low-rank structure. At the output end, the neural network model based on attention mechanism and low-rank representation classifies or grades the operating status of transformer, such as "normal", "minor abnormality", "warning", "serious fault", etc., and can output specific fault types (such as winding overheating, core grounding, electrical partial discharge, mechanical loosening, etc.) and confidence scores in combination with abnormal trends. Meanwhile, the neural network model based on attention mechanism and low-rank representation supports a dynamic scoring mechanism based on sliding time window, continuously judging the health status trend at multiple time points, and realizing early identification and warning of potential faults; if the model determines the state is abnormal multiple times in a row, the warning mechanism is triggered; an alarm signal is output and linked to the upper-level operation and maintenance system, prompting maintenance personnel to intervene or adjust the operation strategy, thereby realizing the intelligent fault management capability of "diagnosing while running".

[0100] The model performs real-time fault diagnosis based on collected data such as voltage, current, and oil temperature. The diagnosis results include: Fault prediction: If the transformer current or oil temperature exceeds the set threshold, an automatic warning will be issued, indicating that the transformer is at risk of failure. Specifically, the transformer current data usually comes from the load current sensor in the real-time monitoring system. This data reflects the transformer's working load and electrical performance. The oil temperature is obtained through an oil temperature sensor, which reflects the thermal state inside the transformer tank. This is closely related to the transformer's heat dissipation capacity, load conditions, and the efficiency of the cooling system.

[0101] Current threshold: Thresholds for current data (e.g., load current) are typically closely related to the transformer's rated current. Standard threshold ranges are usually:

[0102] Normal operation: The current should be maintained between 80% and 100% of the rated current.

[0103] Warning threshold: If the current continues to exceed 110%-120% of the rated current, there may be an overload risk.

[0104] Fault risk threshold: If the current exceeds 120% of the rated current (or the upper limit set according to the specific model), it may cause transformer overload, overheating or electrical fault, and warning measures should be taken immediately.

[0105] Oil temperature threshold: Oil temperature data (e.g., transformer oil temperature) primarily reflects the transformer's cooling effect and heat dissipation capacity. Its set threshold is generally based on the transformer's rated oil temperature and the design standards of the cooling system.

[0106] Normal oil temperature: usually between 40°C and 80°C (the specific value is related to the transformer's design specifications, operating environment and type).

[0107] Warning oil temperature: When it exceeds 85°C, it may cause the internal temperature of the transformer to be too high, affecting the insulation performance, and triggering the warning stage.

[0108] Fault risk oil temperature: If the oil temperature continues to exceed 95°C or 100°C, a high temperature fault warning needs to be triggered, indicating that there may be a risk of insulating oil degradation, local overheating, or cooling system failure.

[0109] Fault type identification: The finely tuned neural network model based on attention mechanism and low-rank representation identifies fault types, including overload, short circuit, and abnormal temperature; the specific judgment criteria are as follows: If the real-time current exceeds the maximum allowable current (e.g., 300A), it is determined to be an overload fault; If the oil temperature exceeds the set threshold (e.g., 85℃), it is determined to be an abnormal temperature. Fault Type Identification: The finely tuned neural network model based on attention mechanism and low-rank representation is used to identify transformer fault types. The specific process is as follows: The neural network model based on attention mechanism and low-rank representation receives multi-dimensional data (such as current, oil temperature, partial discharge, etc.) from the real-time monitoring system, combines historical data with features learned during training, and uses deep learning and attention mechanism to comprehensively analyze the current operating status. First, the model extracts and reduces the dimensionality of the multi-dimensional input data, and uses the low-rank adaptation mechanism to optimize the representation capability of each layer of features; then, the classifier module determines the fault type of the processed data.

[0110] Fault types include, but are not limited to: Overload fault: When the transformer load current exceeds the set threshold of the rated current, it indicates that the transformer may be at risk of overload. Short circuit fault: When the transformer current rises abnormally and remains above the rated value, accompanied by other data characteristics such as partial discharge and abnormal vibration, a short circuit fault may occur. Abnormal temperature: When the transformer oil temperature exceeds the set safety threshold (such as 85°C or 95°C), it indicates that there may be insufficient heat dissipation or a cooling system failure, resulting in abnormal temperature. Partial discharge fault: By monitoring the partial discharge signal, when the amplitude of the signal exceeds the normal range, it indicates that some insulation material inside the transformer may have broken down or been damaged. Mechanical failure: When the vibration or noise signal is abnormal, it indicates that the mechanical parts of the transformer (such as the fan, motor bearings, etc.) may be malfunctioning. Insulation fault: By monitoring gas content (such as hydrogen, ethylene gas concentration, etc.), a high concentration of gas may indicate that the transformer's insulation material has deteriorated and there is a risk of breakdown.

[0111] The specific criteria for judgment are as follows: Current threshold: Based on load current data, if the current continuously exceeds 110%-120% of the rated value, it is judged as an overload fault; if the current fluctuates drastically and matches the characteristics of a short circuit, it is judged as a short circuit fault.

[0112] Oil temperature threshold: If the oil temperature exceeds 85°C, it enters the warning stage; if it exceeds 95°C, it is considered an abnormal temperature, which may lead to a decrease in the insulation performance of the transformer.

[0113] Partial discharge: If the amplitude and frequency of the partial discharge signal exceed the set safety threshold, it is determined to be a partial discharge fault.

[0114] Vibration and noise signals: When the amplitude and frequency of the vibration signal change significantly and exceed the normal range, it is determined to be a mechanical fault.

[0115] ④ Real-time reasoning and decision support: When a transformer fails or has a potential failure, the edge server generates a fault diagnosis report, which includes the fault type, fault time, fault location, and fault severity.

[0116] Figure 7This invention illustrates the scoring function S(n) applied to the encoder-decoder architecture, which is based on a Transformer model with an encoder-decoder structure. This model consists of N stacked encoder blocks and N stacked decoder blocks. The encoder is responsible for converting the input transformer operation and maintenance data (after embedding and position encoding) into a contextual representation, while the decoder uses the contextual information output by these encoders and its own output embedding (also position-encoded) to progressively generate the target fault diagnosis result. Sub-layers within each encoder and decoder block (such as attention layers and feedforward networks) are connected through residual connections and layer normalization (Add & Normalize) operations to ensure stable training and information flow of the deep network.

[0117] From the center to each encoder and decoder block are a Multi-Head Self-Attention (MHSA) layer and a Feed-Forward Network (FFN) sublayer. The decoder also includes an additional Encoder-Decoder Attention layer to bridge information between the encoder and decoder. The structural innovation lies in the ingenious integration of a Low-Rank Adjustment (LoRA) module into the linear transformation matrices of these core layers, including the Q, K, V projections and output linear layers of the MHSA, the weight matrix of the FFN, and the encoder-decoder attention layer. LoRA works by adding a pair of small low-rank matrices (such as...) next to the original pre-trained weights. Figure 7 As indicated by the multiple "low-rank adaptation" arrows, this allows for efficient fine-tuning while freezing most of the original parameters.

[0118] Furthermore, another core structural innovation of this model lies in its adaptive attention guidance mechanism. This mechanism structurally guides the parameter configuration (e.g., the rank of the low-rank matrix) of the LoRA modules integrated in each layer through dynamically calculated attention scores S(n). This means that the attention score S(n) is not an independent layer, but rather a high-level control signal that dynamically adjusts the "scale" and "intensity" of multiple LoRA modules within the model, thereby optimizing the allocation of model resources at the structural level. This allows the model to prioritize computational resources and parameter updates for feature paths more critical to transformer fault diagnosis. Finally, the decoder output passes through a linear layer with LoRA and a Softmax function to produce the final fault diagnosis probability.

[0119] Figure 8 This is a diagram of the dynamic rank r(n) acting on the encoder-decoder architecture in this invention; Figure 8This paper details how the Low-Rank Adjustment (LoRA) mechanism of the dynamic rank r(n) plays a crucial role in the encoder-decoder architecture of the Transformer model of this invention. The figure highlights the extensive integration of the LoRA module into the model and how its core parameter (rank r) is adaptively allocated according to a dynamic policy. The model as a whole follows the encoder-decoder structure of the Transformer, using positional encoding to process input and output embeddings to capture the temporal information of the sequence.

[0120] This invention extensively integrates the low-rank fine-tuning (LoRA) module into the core linear transformation layer. For example... Figure 8 As shown, this includes the query (Q), key (K), and value (V) projection layers within the Multi-Head Attention (MHSA) layer, as well as the linear layers after the MHSA output. Furthermore, the key linear layers within the Feed Forward Network (FFN) in the encoder and decoder, and the masked MHSA and encoder-decoder attention layers in the decoder, also integrate the LoRA module. LoRA achieves efficient parameter fine-tuning by adding a pair of trainable low-rank matrices (BAs) to the original weight matrix (which remains frozen during fine-tuning).

[0121] The core innovation of this invention lies in the dynamic rank r(n) mechanism of LoRA. For example... Figure 8 The annotations next to each module clearly indicate that the LoRA parameters (especially their rank r) of these key layers are no longer fixed, but are adaptively determined based on the dynamic rank r(n) or attention score S(n). For example, "Q, K, V projection layer weights use low-rank decomposition" and are controlled by dynamic rank; various attention layers in the feedforward network and decoder are clearly marked "low-rank adaptation parameter allocation for this layer is determined based on dynamic rank r(n)" or "low-rank adaptation is applied based on dynamic rank r(n)". This mechanism ensures that the model can intelligently allocate different sizes of parameter spaces (i.e., different ranks r) to the LoRA modules according to the importance of each layer or the needs of the current training stage, thereby achieving the best balance between optimizing model resource utilization, accelerating convergence, and improving transformer fault diagnosis performance. Finally, the decoder output generates the final diagnostic probability through a linear layer that also applies low-rank adaptation and the Softmax function.

[0122] Figure 9 This is a diagram illustrating the architecture of the encoder-decoder using the hierarchical adaptive learning rate machine in this invention. Figure 9This paper further elucidates how the hierarchical adaptive learning rate mechanism proposed in this invention specifically applies to the encoder-decoder Transformer architecture, with a particular focus on the implementation of Adaptive Low-Rank Adaptation. Figure 9 exist Figure 7 and Figure 8 Building upon this, the integration of low-rank adaptation modules into the key linear layer of the Transformer is clearly demonstrated, and the adaptive characteristics of these modules are emphasized.

[0123] As shown in the figure, this invention extensively integrates the Low-Rank Adaptation (LoRA) module into the core components of the Transformer architecture: including the Q, K, and V projection layers and the final output linear layer within the Multi-Head Attention (MHSA) network, the linear layers in the Feedforward Network (FFN), and the masked MHSA and encoder-decoder attention in the decoder section. These integrations achieve efficient parameter fine-tuning by adding a pair of trainable low-rank matrices (BA) to the original weight matrix.

[0124] The core structural innovation of this invention lies in the mechanisms of "adaptive low-rank adaptation" and "adaptive fine-tuning adaptation". Figure 9 The annotations and connections explicitly state that the low-rank adaptation parameters and their learning rates of these key modules (such as multi-head attention and feedforward networks) are not fixed, but dynamically adjusted by a hierarchical adaptive learning rate mechanism. This means that the model can intelligently adjust the rank values ​​of the low-rank matrices of different layers or parameters and their learning rates based on feedback during training or a pre-set strategy, thereby achieving more refined and efficient parameter updates. This adaptive strategy ensures that the model maintains strong expressive power while minimizing fine-tuning costs and optimizing the model's convergence speed and final performance on transformer fault diagnosis tasks.

[0125] The performance data of the transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning in this embodiment are shown in Table 1. Table 1. Results Data Table;

[0126] Example 3 A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning as described in Embodiment 1 or 2.

[0127] Example 4 A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning as described in Embodiment 1 or 2.

Claims

1. A transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning, characterized in that, Includes the following steps: S1, Data Acquisition and Preprocessing Stage; S101. Collect relevant operation and maintenance data of transformers in the target power system; S102. Standardize and clean the collected data and preprocess it to construct a local transformer fault diagnosis dataset; S2, Initialization phase; S201. Sample data from the constructed local transformer fault diagnosis dataset, perform forward and backward propagation on the original pre-trained language model, and collect gradients, attention activation maps and hidden state information of each layer. S202. Based on the collected information, calculate the gradient importance score, activation importance score, and value change importance score for each layer, and then apply the learnable weight coefficients. Combine and construct a unified attention scoring metric; S203. Allocate the required rank and learning rate for each layer according to the obtained attention score, and initialize the LoRA parameter matrices A and B of the corresponding layers to achieve the initial configuration of the low-rank structure. S3, Fine-tuning stage; S301. Execute the standard training batch process on the fault diagnosis data, i.e. the transformer's relevant operation and maintenance data, including forward propagation, loss calculation and back propagation, to complete the local update of parameters; S302. Real-time monitoring of the training dynamics of neural network models based on attention mechanisms and low-rank representations, including performance metrics, training loss, and gradient dynamics. S303. Based on the monitored performance and training dynamics of the neural network model based on attention mechanism and low-rank representation, when the preset context-aware triggering conditions are met, dynamically update the attention score and rank allocation strategy to achieve adaptive adjustment of the low-rank structure. S304. At the end of each training round, the overall effect is evaluated based on the performance of the validation set, and the rank allocation strategy is adjusted based on the feedback from the neural network model based on the attention mechanism and low-rank representation. S4, Adaptive Adjustment Phase; S401. Monitor and verify the trend of indicator changes. If the performance improvement is small in multiple consecutive evaluations, automatically adjust key hyperparameters or implement a learning rate restart strategy. S402. If the performance still does not improve significantly after multiple adjustments, the early stop mechanism is triggered, and the model with the best verification performance is selected as the final model. S5, Deployment and Real-time Diagnostics Phase; The finely tuned neural network model based on attention mechanism and low-rank representation is deployed to an edge server to enable online parsing and diagnosis of transformer operation and maintenance logs or real-time monitoring data.

2. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, In S101, the relevant operation and maintenance data of the transformer includes inspection records, maintenance logs, alarm information, operating parameters, and environmental parameters in unstructured or semi-structured text data. In S102, the collected data is standardized, cleaned, and preprocessed; this means that the collected data is sequentially denoised, normalized, filled with missing values, classified, and labeled.

3. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, In S201, data is sampled from the constructed local transformer fault diagnosis dataset, and forward and backward propagation is performed on the original pre-trained language model to collect gradients, attention activation maps, and hidden state information of each layer; including: Collecting gradients from each layer: During backpropagation, the original pre-trained language model, based on the error calculated by the loss function, is propagated backward from the output layer to the input layer, thereby calculating the gradients of the weight matrices of each layer and collecting gradient information. , where n represents the nth layer, gradient information This indicates the sensitivity of the nth layer parameters to the loss function under the current task; Collect attention activation maps of each layer: The original pre-trained language model synchronously collects key signals during the forward propagation process, that is, extracts attention activation information; the attention activation map, that is, the attention activation information, refers to the attention weight matrix generated by each attention head in the original pre-trained language model when processing the input sequence; Collect hidden state information of each layer; hidden state information refers to the intermediate representation output of each encoder or decoder during the forward propagation of the Transformer model, including the contextual semantic features of the input data after nonlinear transformation in the current layer.

4. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, The original pre-trained language model was BERT; BERT consists of 12 encoder layers. The input text is first processed by the embedding layer, and then passed through each encoder layer in turn. The output of each encoder layer is the hidden state of that encoder layer, which is a three-dimensional tensor of shape [batch_size, seq_len, hidden_size]. Here, batch_size represents the sample batch size; seq_len represents the length of the input sequence; and hidden_size represents the feature dimension of the current layer.

5. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, In S202, the gradient importance score G(n) of each layer is calculated using the following formula: ; in, This represents the gradient of the weight matrix of the nth layer. This represents the Frobenius norm, which is the square root of the sum of the squares of all elements in the corresponding matrix; In S202, the activation importance score is calculated using the activation graph of the self-attention mechanism. The calculation steps are as follows: Standardize by the number of heads, divide the total difference by the number of attention heads in the layer, and calculate the average single-head difference; the total difference is the sum of the L1 differences between each pair of attention matrices of each attention head in the same layer; Standardize by sequence length, and divide the average single-head difference by the square of the input sequence length; The activation importance score is obtained by processing the result after the sequence length is standardized using a dynamic normalization function.

6. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 5, characterized in that, The results of sequence length standardization are processed using a dynamic normalization function, including: In the early stages of training, i.e. t < 0.4T, where t is the current training step number and T is the total training step number, global maximum normalization is used for processing. During the middle of training, i.e., 0.4T≤t<0.6T, a mixture function is used for processing; In the later stages of training, i.e., when t≥0.6T, quantile normalization is used for processing.

7. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, In S202, the importance score V(n) for value change is calculated using the following formula: ; Where H(n) and H0(n) are the hidden state outputs of the nth layer in the fine-tuning and original models, respectively; H(i) and H0(i) are the hidden state outputs of the ith layer in the fine-tuning and original models, respectively. Alternatively, the importance score V(n) for value change can be calculated using the following formula: ; Where α1, α2, α3 are weighting coefficients, and V(n) represents the importance score of the value change in the nth layer; In S202, the attention score S(n) is calculated using the following formula: Where G(n) represents the gradient importance score of the nth layer, A(n) represents the activation importance score of the nth layer, and V(n) represents the value change importance score of the nth layer. These are the learnable weight coefficients.

8. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 5, characterized in that, In S203, resource allocation is based on attention score S(n), according to a dynamic rank allocation strategy and hierarchical self-allocation. The adaptive learning rate mechanism allocates the rank r(n) and learning rate lr(n) for each layer, including the following steps: Basic parameter settings and rank and learning rate allocation: Set basic parameters, including the basic rank value. Maximum rank Scaling factor Basic learning rate and learning rate adjustment factor Based on the calculated attention score S(n), a rank value r(n) and a learning rate lr(n) are assigned to each layer; Rank allocation: The rank r(n) is allocated using the following formula: ; Learning rate allocation: The hierarchical adaptive learning rate lr(n) is allocated using the following formula: ; For each layer, initialize the corresponding LoRA parameter matrices A and B according to the assigned rank value, as follows: LoRA parameter matrix initialization: The corresponding LoRA parameter matrices A and B are initialized according to the assigned rank values. Matrix A is initialized using a normal distribution, and matrix B is initialized with all zeros. In S203, based on the performance of the neural network model based on the attention mechanism and low-rank representation after every N training batches, the attention score and rank allocation strategy are periodically updated to achieve dynamic adaptation of the low-rank structure; including: Update the attention scores of each layer using a sliding window as the average: ; S(l) old Let S(l) represent the attention score at the last update of layer l. current This represents the attention that was just calculated in the l-th layer during the current training cycle; Periodically update the rank allocation strategy, including: The total rank budget constraint is: ; Rank calculation logic: ; The LoRA matrix dimensions are adaptively adjusted using the following strategy: Rank expansion: When increasing from r to r+Δr, the newly added parameter matrix block is orthogonally initialized and inherits the principal components of the original matrix through SVD decomposition; Rank contraction: When increasing from r to r-Δr, retain the eigencomponents corresponding to the first Δr largest singular values ​​in the singular value decomposition.

9. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to claim 1, characterized in that, In S4, the adaptive adjustment phase; include: When the model performance improvement is detected to be small or stagnant during the fine-tuning process, that is, when the improvement of the evaluation index is less than the preset threshold for N consecutive times, it is determined that the performance improvement is small and the model may be trapped in a local optimum; the key hyperparameters are automatically adjusted or the learning rate restart strategy is implemented. If the performance improvement is still minimal after multiple adjustments, the early stopping mechanism is triggered, and the model with the best verification performance is selected as the final model.

10. The transformer fault diagnosis method based on adaptive attention-guided low-rank fine-tuning according to any one of claims 1-9, characterized in that, In S5, the deployment and real-time diagnostics phase includes: ① Edge computing device deployment: The finely tuned low-rank adaptation model is deployed on the NVIDIA Jetson Xavier AGX edge computing device; ② Data Input and Real-time Monitoring: Real-time monitoring data of the transformer is collected through sensors, and the sensors transmit the real-time data to the edge server via the Modbus TCP / IP protocol; the real-time monitoring data of the transformer includes: input voltage, output current, oil temperature, load current, and alarm status; ③ Real-time fault diagnosis and early warning: The fine-tuned neural network model based on attention mechanism and low-rank representation processes the real-time monitoring data of the transformer after receiving it; specifically, it includes: receiving multi-source data input from the transformer online monitoring system, converting the multi-source data input into a standard format acceptable to the neural network model based on attention mechanism and low-rank representation according to the preset data preprocessing process, and constructing a time series input feature tensor. The neural network model based on attention mechanism and low-rank representation integrates historical data distribution with current input features, uses attention mechanism to capture the coupling relationship and dynamic change trend between key variables, and uses the enhanced feature expression ability of low-rank structure to judge the current state. At the output end, the neural network model based on attention mechanism and low-rank representation classifies or grades the operating state of transformer, and can output specific fault type and confidence score in combination with abnormal trends. Meanwhile, the neural network model based on attention mechanism and low-rank representation supports a dynamic scoring mechanism based on sliding time window, which continuously judges the health status trend at multiple time points, and realizes early identification and warning of potential faults; if the model judges the state to be abnormal multiple times in a row, the warning mechanism is triggered. The diagnostic results include: Fault prediction: If the transformer current or oil temperature exceeds the set threshold, an early warning will be automatically issued to indicate that there is a risk of transformer failure. Fault type identification: The finely tuned neural network model based on attention mechanism and low-rank representation identifies fault types, including overload, short circuit, and abnormal temperature; the specific judgment criteria are as follows: If the real-time current exceeds the maximum allowable current, it is determined to be an overload fault; If the oil temperature exceeds the set threshold, it is determined to be an abnormal temperature. ④ Real-time reasoning and decision support: When a transformer fails or has a potential failure, the edge server generates a fault diagnosis report, which includes the fault type, fault time, fault location, and fault severity.

Citation Information

Patent Citations

  • Log anomaly detection method based on efficient fine tuning of adaptive low-rank parameters

    CN118260689A

  • Injection molding process fault diagnosis model training method and system based on large language model and fault diagnosis method

    CN120408032A

  • Fine tuning method and system based on power failure multi-modal model, and medium

    CN120493773A

  • Distribution network line facility defect detection method based on line magnetic field variable characteristics

    CN120522512A

  • Dynamic optimization system for AI model training parameters

    CN120633719A

Cited By

  • Universe automobile temperature prediction method and system based on meta transfer learning

    CN121188417A

  • A global vehicle temperature prediction method and system based on meta-transfer learning

    CN121188417B

  • Low-voltage power distribution system monitoring and protecting method and device based on AI technology

    CN121216737A

  • A method and device for monitoring and protecting low-voltage power distribution systems based on AI technology

    CN121216737B

  • Model training method and device based on parameter adjustment, equipment and medium

    CN121479308A