Transformer fault prediction method based on DeepSeek large model

The fault samples are added through hierarchical random sampling and improved SMOTE algorithm, combined with the DBSCAN algorithm to screen normal samples, use the DeepSeek model to extract deep features, and integrate multiple learning models for transformer fault prediction, solving the problems of data imbalance and insufficient timing feature extraction, and achieving efficient fault prediction and early latent fault warning.

CN120561672APending Publication Date: 2025-08-29CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510563658.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the existing transformer fault prediction models, data imbalance is serious, and the scarcity of fault samples leads to frequent misreports of traditional machine learning models, and insufficient adaptability to timing feature extraction and early warning thresholds.

Method used

The fault samples are added by hierarchical random sampling and improved SMOTE algorithm, and the normal samples are screened in combination with the DBSCAN algorithm. DeepSeek model is used to extract deep features, and feature importance evaluation is performed through the multi-head self-attention mechanism and feedforward neural network, and multiple learning models are integrated for prediction.

Benefits of technology

It effectively improves the fault recall rate, reduces the false alarm rate, can detect latent faults in the early stage, provides support for the preventive maintenance of the transformer, and improves the stability and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561672A_ABST
    Figure CN120561672A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer fault prediction method based on a DeepSeek large model. The method comprises the following steps: acquiring transformer fault related original data; the method comprises the following steps of: dividing transformer fault related original data into a training set and a verification set by adopting a layered random sampling method, layering an original training data set for the training set according to categories of fault and normal samples, and dividing the original training data set into a secondary sampling subset and a single sampling subset; an improved SMOTE algorithm is adopted for the secondary sampling subset to increase fault sample data; screening normal sample data for the single sampling subset by adopting a clustering-based screening algorithm; evaluating the data set balance degree of the secondary sampling subset and the single sampling subset; combining the subsets after secondary sampling and the subsets after single sampling to form a final training data set; identifying deep features in the data by using a DeepSeek model, and extracting feature information with high correlation on transformer fault prediction; and inputting and integrating the deep features into a plurality of learning models to predict and obtain a final prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power equipment detection, and in particular relates to a transformer fault prediction method based on a DeepSeek large model. Background Art

[0002] Transformers are critical equipment in power system operation, and their health directly impacts the stability and safety of the entire system. However, current transformer fault prediction and early warning technologies face numerous challenges. Data imbalance severely constrains model performance. Fault samples are extremely scarce in real-world scenarios, and traditional machine learning models such as support vector machines (SVMs) and random forests often experience frequent underreporting when processing such data due to uneven sample distribution. Summary of the Invention

[0003] Aiming at the problem of scarcity of fault samples in the training data set of the power equipment detection model in the existing transformer fault prediction model, the present invention provides a transformer fault prediction method based on the DeepSeek large model.

[0004] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0005] A transformer fault prediction method based on the DeepSeek large model includes the following steps:

[0006] S1. Collect raw data related to transformer faults;

[0007] S2. The original data related to transformer faults are divided into a training set and a validation set using a stratified random sampling method. The training set is stratified according to the categories of fault and normal samples into a secondary sampling subset and a single sampling subset.

[0008] S3. Use the improved SMOTE algorithm to increase the fault sample data for the secondary sampling subset;

[0009] S4. Using a clustering-based screening algorithm to screen the normal sample data for the single sampling subset;

[0010] S5. Evaluate the balance of the datasets of the secondary sampling subset and the single sampling subset;

[0011] S6, merging the subsampled subset and the single-sampled subset to form the final training data set;

[0012] S7. Use the DeepSeek model to identify deep features in the data and extract feature information that is highly relevant to transformer fault prediction;

[0013] S8. Integrate the deep feature input into multiple learning models to predict the final prediction result.

[0014] Furthermore, the improved SMOTE algorithm is a time series SMOTE (TS-SMOTE) algorithm.

[0015] Furthermore, the detailed steps of using the improved SMOTE algorithm to increase the fault sample data for the secondary sampling subset are as follows:

[0016] First calculate the nearest neighbors: for each fault sample x i , find its k nearest neighbor samples in the fault sample set; since it is time series data, dynamic time warping (DTW) distance is used to measure the similarity between samples.

[0017] Next, generate a new sample: randomly select a neighboring sample x j , and in x i and x j Perform linear interpolation between them to generate a new fault sample x new ; Considering the continuity of the time series, the interpolation process is smoothed in the time dimension. new =x i +λ(x j -x i ) where λ is a random number in the interval [0,1].

[0018] Furthermore, the detailed steps of using clustering-based screening algorithm to screen normal sample data for a single sampling subset are as follows:

[0019] The DBSCAN algorithm is used to cluster normal samples and automatically identify different clusters and noise points according to the density of the samples;

[0020] Secondly, sample screening: for each cluster, calculate its central sample;

[0021] According to the preset screening ratio, samples that are closer to the center sample are selected as retained samples, and the rest of the samples are eliminated. This can retain the main features of normal samples while reducing redundant samples.

[0022] Furthermore, the evaluation indicators for evaluating the balance degree of the datasets of the secondary sampling subset and the single sampling subset include the ratio of faulty samples to normal samples and G-mean (geometric mean);

[0023] The ratio of faulty samples to normal samples: Calculate the ratio of the number of faulty samples to normal samples after subsampling and make it close to 1:1 to achieve a balanced dataset;

[0024] G-mean: The geometric mean of recall and specificity, which comprehensively evaluates the performance of the model on faulty samples and normal samples.

[0025] The calculation formula is: Among them, recall rate indicates the proportion of fault samples correctly identified by the model, and specificity indicates the proportion of normal samples correctly identified by the model.

[0026] Furthermore, the DeepSeek model includes a multi-head self-attention mechanism and a feedforward neural network:

[0027] The multi-head self-attention mechanism captures the dependencies between different positions in the data:

[0028] For the input sequence X=[x1,x2,...,x n ], first map it to three spaces: query, key, and value, denoted as Q, K, and V respectively;

[0029] Then, by calculating the attention score, we get the weighted feature representation;

[0030] The specific calculation steps are as follows:

[0031] 1) Calculate attention score: Among them, d k is the dimension of query and key, and the softmax) function is used to normalize the attention score to the [0,1] interval.

[0032] 2) Multi-head attention: Map the input sequence to multiple different subspaces, calculate multiple attention heads in parallel, and finally concatenate the results and perform linear transformation to enhance the expressiveness of the model. MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O . Where h is the number of attention heads, W O is the weight matrix of the linear transformation.

[0033] Feedforward neural network: The feature representation of each position undergoes nonlinear transformation through a feedforward neural network to further extract feature information; the feedforward neural network consists of two fully connected layers and an activation function (such as ReLU), and its calculation formula is: FFN(x) = max(0,xW1b1)W2+b2, where W1 and W2 are weight matrices, and b1 and b2 are bias vectors.

[0034] Furthermore, a feature importance evaluation method based on game theory is used to evaluate the feature information extracted by the DeepSeek model, extracting feature information that is highly relevant to transformer fault prediction.

[0035] SHAP value calculation: For a feature x i , whose SHAP value φ i Indicates the marginal contribution of the feature to the model prediction result. In the specific calculation, all possible feature combinations are enumerated, the prediction value of the model under each combination is calculated, and its contribution is determined according to the order in which the features are added.

[0036] Feature screening: Features are sorted based on the absolute value of their SHAP values, and features with larger SHAP values ​​are selected as important for transformer fault prediction. For example, a threshold is set to retain only features with an absolute SHAP value greater than the threshold, thereby achieving feature screening and dimensionality reduction.

[0037] Furthermore, DeepSeek model training and optimization; after building the DeepSeek model and determining the feature importance evaluation method, the model needs to be trained and optimized to improve the accuracy of its fault prediction.

[0038] Loss function: The cross entropy loss function is used as the training target of the model. For the binary classification problem (fault or normal), its calculation formula is: Where N is the number of samples, y i is the true label of the sample, p i is the model's predicted probability for the sample

[0039] Optimization algorithm, using Adam optimization algorithm to update the model parameters, it combines the advantages of AdaGrad and RMSProp algorithms and can adaptively adjust the learning rate. Its update formula is:

[0040] m t =β1m t-1 +(1-β1)g t

[0041]

[0042] Among them, m t and v t are the first-order moment estimate and the second-order moment estimate, β1 and β2 are the decay rates, g is the gradient at the current moment, α is the learning rate, and ∈ is a small constant used to prevent the denominator from being zero.

[0043] Furthermore, multiple learning models are integrated: the bottom layer uses multiple different machine learning models to learn and predict data, and the upper layer model integrates and re-learns the output of the bottom layer model.

[0044] Furthermore, the detailed steps of step S8 are:

[0045] Data preparation:

[0046] The balanced data is sorted into the following format:

[0047] D={(x1,y1),(x2,y2),…,(x n ,y n )}

[0048] where x n is the eigenvector, y n Is the corresponding label. Divide the data set into training set D train and the test set D test .

[0049] Select the underlying model:

[0050] Choose different types of machine learning models as the underlying model, such as decision tree, support vector machine (SVM), random forest, K-nearest neighbor, etc. Assume that we have selected m underlying models M1, M2, ..., M m .

[0051] Underlying model training and prediction:

[0052] The cross-validation method is used to train the underlying model. The specific steps are:

[0053] 1) The training set D train Divide into k mutually disjoint subsets D1, D2, ..., D K .

[0054] 2) For each underlying model M j (j=1,2,…,m):

[0055] 3) Perform k iterations. In the i-th iteration:

[0056] 4) Use Divide by D i K-1 training model M j .

[0057] 5) Use the trained model M j To D i The samples in are predicted and the prediction result y is obtained i,j .

[0058] 6) Concatenate the prediction results obtained by k iterations to form a new feature vector yj with a length of n train (Number of training set samples).

[0059] Upper model training:

[0060] The new feature matrix y=[y1,y2,…,y m ] as input, the original label y train As output, train an upper model M top The upper model can choose logistic regression (LogisticRegression), neural network (NeuralNetwork), etc.

[0061] Test set predictions:

[0062] After completing the cross-validation of the underlying model and the training of the upper-level model, predictions are made on the test set.

[0063] First, for each underlying model M j , using the entire training set D train Retrain to allow the model to fully learn the information of the training set. Then, use the trained underlying model to train the test set D test The samples in are predicted to obtain the prediction results y corresponding to each underlying model test,j Subsequently, the prediction results of these underlying models are combined to form a new feature matrix y test =[y test,1 ,y test,2 ,…,y test,m ]. Finally, the feature matrix is ​​input into the upper model M top Make predictions to get the final prediction results.

[0064] Model Evaluation: Comprehensively evaluate the model using metrics such as precision, accuracy, recall, and F1 value. Precision reflects the percentage of samples predicted as faulty by the model that are actually faulty; accuracy measures the accuracy of the model's overall predictions; recall indicates the percentage of actual faulty samples correctly predicted as faulty by the model; and F1 value is the harmonic mean of precision and recall, comprehensively accounting for the impact of both. This multi-metric evaluation ensures that the model achieves high performance across various aspects, providing reliable support for practical applications.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] By combining the DeepSeek model with imbalanced datasets and training dataset scarcity optimization techniques, the team overcomes the limitations of existing technologies in model performance due to data imbalance, as well as the shortcomings of traditional methods in time series feature extraction and adaptability of warning thresholds. This method not only more comprehensively captures subtle changes in transformer operating conditions, providing a richer information foundation for fault prediction, but also effectively improves fault recall rates, reduces false alarm rates, and proactively identifies potential fault hazards, providing strong support for preventive maintenance of transformers. In practical applications, this method demonstrates significant technical advantages and value, providing more reliable technical support for the stable operation of power systems and equipment management.

[0067] By leveraging the powerful feature recognition and learning capabilities of the DeepSeek large model, we can effectively process the complex features in transformer operation data, mine deep information in the data, and improve the model's sensitivity to fault characteristics.

[0068] Secondary sampling technology: The data is processed through the secondary sampling method to effectively balance the data set, so that the model can fully learn the characteristics of fault samples during training, thereby improving the fault recall rate and reducing the false alarm rate.

[0069] Stacking ensemble learning method: The stacking ensemble learning method is introduced to improve the generalization ability and prediction accuracy of the overall model by building a multi-layer model combination, ensuring the stability and reliability of the system in practical applications.

[0070] Early warning of latent faults: It can support early warning of latent faults (such as the early stage of partial discharge), providing a valuable time window for transformer maintenance and management, and helping to take timely measures to avoid further deterioration of the fault. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is an overall flow chart of a transformer fault prediction method based on the DeepSeek large model in an embodiment of the present invention;

[0072] Figure 2 The figure is a detailed flow chart of a transformer fault prediction method based on the DeepSeek large model in an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and drawings. The contents mentioned in the embodiments are not intended to limit the present invention.

[0074] like Figure 1 and 2 As shown, this embodiment provides a transformer fault prediction method based on the DeepSeek large model, including the following steps:

[0075] S1. Collect raw data related to transformer faults;

[0076] S2. The original data related to transformer faults are divided into a training set and a validation set using a stratified random sampling method. The training set is stratified according to the categories of fault and normal samples into a secondary sampling subset and a single sampling subset.

[0077] S3. Use the improved SMOTE algorithm to increase the fault sample data for the secondary sampling subset;

[0078] S4. Using a clustering-based screening algorithm to screen the normal sample data for the single sampling subset;

[0079] S5. Evaluate the balance of the datasets of the secondary sampling subset and the single sampling subset;

[0080] S6. Merge the subsampled subset and the single-sampled subset to form the final training data set;

[0081] S7. Use the DeepSeek model to identify deep features in the data and extract feature information that is highly relevant to transformer fault prediction;

[0082] S8. Integrate the deep feature input into multiple learning models to predict the final prediction result.

[0083] The improved SMOTE algorithm is the time series SMOTE (TS-SMOTE) algorithm.

[0084] Detailed steps for using the improved SMOTE algorithm to increase fault sample data for the secondary sampling subset:

[0085] First calculate the nearest neighbors: for each fault sample x i , find its k nearest neighbor samples in the fault sample set; since it is time series data, dynamic time warping (DTW) distance is used to measure the similarity between samples.

[0086] Next, generate a new sample: randomly select a neighboring sample x j , and in x i and x j Perform linear interpolation between them to generate a new fault sample x new ; Considering the continuity of the time series, the interpolation process is smoothed in the time dimension. new =x i +λ(x j -x i ) where λ is a random number in the interval [0,1].

[0087] Detailed steps for screening normal sample data using a clustering-based screening algorithm for a single sampling subset:

[0088] The DBSCAN algorithm is used to cluster normal samples and automatically identify different clusters and noise points according to the density of the samples;

[0089] Secondly, sample screening: for each cluster, calculate its central sample;

[0090] According to the preset screening ratio, samples that are closer to the center sample are selected as retained samples, and the rest of the samples are eliminated. This can retain the main features of normal samples while reducing redundant samples.

[0091] The evaluation indicators for evaluating the balance degree of the datasets of the secondary sampling subset and the single sampling subset include the ratio of faulty samples to normal samples and G-mean (geometric mean);

[0092] The ratio of faulty samples to normal samples: Calculate the ratio of the number of faulty samples to normal samples after subsampling and make it close to 1:1 to achieve a balanced dataset;

[0093] G-mean: The geometric mean of recall and specificity, which comprehensively evaluates the performance of the model on faulty samples and normal samples.

[0094] The calculation formula is: Among them, recall rate indicates the proportion of fault samples correctly identified by the model, and specificity indicates the proportion of normal samples correctly identified by the model.

[0095] The DeepSeek model includes a multi-head self-attention mechanism and a feedforward neural network:

[0096] The multi-head self-attention mechanism captures the dependencies between different positions in the data:

[0097] For the input sequence X=[x1,x2,...,x n ], first map it to three spaces: query, key, and value, denoted as Q, K, and V respectively;

[0098] Then, by calculating the attention score, we get the weighted feature representation;

[0099] The specific calculation steps are as follows:

[0100] 1) Calculate attention score: Among them, d k is the dimension of query and key, and the softmax) function is used to normalize the attention score to the [0,1] interval.

[0101] 2) Multi-head attention: Map the input sequence to multiple different subspaces, calculate multiple attention heads in parallel, and finally concatenate the results and perform linear transformation to enhance the expressiveness of the model. MultiHead(Q,K,V)=Concat(head1,head2,…,head h )W O . Where h is the number of attention heads, W O is the weight matrix of the linear transformation.

[0102] Feedforward neural network: The feature representation of each position undergoes nonlinear transformation through a feedforward neural network to further extract feature information; the feedforward neural network consists of two fully connected layers and an activation function (such as ReLU), and its calculation formula is: FFN(x) = max(0,xW1b1)W2+b2, where W1 and W2 are weight matrices, and b1 and b2 are bias vectors.

[0103] The feature importance evaluation method based on game theory is used to evaluate the feature information extracted by the DeepSeek model, and feature information with high relevance to transformer fault prediction is extracted;

[0104] SHAP value calculation: For a feature x i , whose SHAP value φ i Indicates the marginal contribution of the feature to the model prediction result. In the specific calculation, all possible feature combinations are enumerated, the prediction value of the model under each combination is calculated, and its contribution is determined according to the order in which the features are added.

[0105] Feature screening: Features are sorted based on the absolute value of their SHAP values, and features with larger SHAP values ​​are selected as important for transformer fault prediction. For example, a threshold is set to retain only features with an absolute SHAP value greater than the threshold, thereby achieving feature screening and dimensionality reduction.

[0106] DeepSeek model training and optimization: After building the DeepSeek model and determining the feature importance evaluation method, the model needs to be trained and optimized to improve the accuracy of its fault prediction.

[0107] Loss function: The cross entropy loss function is used as the training target of the model. For the binary classification problem (fault or normal), its calculation formula is: Where N is the number of samples, y i is the true label of the sample, p i is the model's predicted probability for the sample

[0108] Optimization algorithm, using Adam optimization algorithm to update the model parameters, it combines the advantages of AdaGrad and RMSProp algorithms and can adaptively adjust the learning rate. Its update formula is:

[0109] m t =β1m t-1 +(1-β1)g t

[0110]

[0111] Among them, m t and v t are the first-order moment estimate and the second-order moment estimate, β1 and β2 are the decay rates, g is the gradient at the current moment, α is the learning rate, and ∈ is a small constant used to prevent the denominator from being zero.

[0112] Integrate multiple learning models: The underlying layer uses multiple different machine learning models to learn and predict data, and the upper-layer model integrates and re-learns the output of the underlying model.

[0113] Detailed steps of step S8:

[0114] Data preparation:

[0115] The balanced data is sorted into the following format:

[0116] D={(x1,y1),(x2,y2),…,(x n ,y n )}

[0117] where x n is the eigenvector, y n Is the corresponding label. Divide the data set into training set D train and the test set D test .

[0118] Select the underlying model:

[0119] Choose different types of machine learning models as the underlying model, such as Decision Tree, Support Vector Machine (SVM), Random Forest, K-Nearest Neighbors, etc. Suppose we choose m underlying models M1, M2, ..., M m .

[0120] Underlying model training and prediction:

[0121] The cross-validation method is used to train the underlying model. The specific steps are:

[0122] 1) The training set D train Divide into k mutually disjoint subsets D1, D2, ..., D K .

[0123] 2) For each underlying model M j (j=1,2,…,m):

[0124] 3) Perform k iterations. In the i-th iteration:

[0125] 4) Use Divide by D i K-1 training model M j .

[0126] 5) Use the trained model M j To D i The samples in are predicted and the prediction result y is obtained i,j .

[0127] 6) Concatenate the prediction results obtained by k iterations to form a new feature vector y j , length n train (Number of training set samples).

[0128] Upper model training:

[0129] The new feature matrix y=[y1,y2,…,y m ] as input, the original label y train As output, train an upper model M top The upper model can choose logistic regression (LogisticRegression), neural network (NeuralNetwork), etc.

[0130] Test set predictions:

[0131] After completing the cross-validation of the underlying model and the training of the upper-level model, predictions are made on the test set.

[0132] First, for each underlying model M j , using the entire training set D train Retrain to allow the model to fully learn the information of the training set. Then, use the trained underlying model to train the test set D test The samples in are predicted to obtain the prediction results y corresponding to each underlying model test,j Subsequently, the prediction results of these underlying models are combined to form a new feature matrix y test =[y test,1 ,y test,2 ,…,y test,m ]. Finally, the feature matrix is ​​input into the upper model M topMake predictions to get the final prediction results.

[0133] Model Evaluation: Comprehensively evaluate the model using metrics such as precision, accuracy, recall, and F1 value. Precision reflects the percentage of samples predicted as faulty by the model that are actually faulty; accuracy measures the accuracy of the model's overall predictions; recall indicates the percentage of actual faulty samples correctly predicted as faulty by the model; and F1 value is the harmonic mean of precision and recall, comprehensively accounting for the impact of both. This multi-metric evaluation ensures that the model achieves high performance across various aspects, providing reliable support for practical applications.

[0134] Compared with the prior art, the present invention has the following beneficial effects:

[0135] By combining the DeepSeek model with unbalanced dataset optimization techniques, the team overcomes the limitations of existing technologies in model performance due to data imbalance, as well as the shortcomings of traditional methods in time series feature extraction and adaptability of warning thresholds. This method not only more comprehensively captures subtle changes in transformer operating conditions, providing a richer information foundation for fault prediction, but also effectively improves fault recall rates, reduces false alarm rates, and proactively identifies potential fault hazards, providing strong support for preventive maintenance of transformers. In practical applications, this method demonstrates significant technical advantages and value, providing more reliable technical support for the stable operation of power systems and equipment management.

[0136] By leveraging the powerful feature recognition and learning capabilities of the DeepSeek large model, we can effectively process the complex features in transformer operation data, mine deep information in the data, and improve the model's sensitivity to fault characteristics.

[0137] Secondary sampling technology: The data is processed through the secondary sampling method to effectively balance the data set, so that the model can fully learn the characteristics of fault samples during training, thereby improving the fault recall rate and reducing the false alarm rate.

[0138] Stacking ensemble learning method: The stacking ensemble learning method is introduced to improve the generalization ability and prediction accuracy of the overall model by building a multi-layer model combination, ensuring the stability and reliability of the system in practical applications.

[0139] Early warning of latent faults: It can support early warning of latent faults (such as the early stage of partial discharge), providing a valuable time window for transformer maintenance and management, and helping to take timely measures to avoid further deterioration of the fault.

[0140] The above describes in detail the transformer fault prediction method based on the DeepSeek large model provided by this application. The description of the specific embodiments is intended only to facilitate understanding of the method and core concepts of this application. It should be noted that those skilled in the art may make various improvements and modifications to this application without departing from the principles of this application, and such improvements and modifications also fall within the scope of protection of the claims of this application.

Claims

1. A transformer fault prediction method based on DeepSeek large model, characterized in that: Including steps: S1. Collect raw data related to transformer faults; S2. The original data related to transformer faults are divided into a training set and a validation set using a stratified random sampling method. The training set is stratified according to the categories of fault and normal samples into a secondary sampling subset and a single sampling subset. S3. Use the improved SMOTE algorithm to increase the fault sample data for the secondary sampling subset; S4. Using a clustering-based screening algorithm to screen the normal sample data for the single sampling subset; S5. Evaluate the balance of the datasets of the secondary sampling subset and the single sampling subset; S6. Merge the subsampled subset and the single-sampled subset to form the final training data set; S7. Use the DeepSeek model to identify deep features in the data and extract feature information that is highly relevant to transformer fault prediction; S8. Integrate the deep feature input into multiple learning models to predict the final prediction result.

2. The transformer fault prediction method based on the DeepSeek large model according to claim 1 is characterized in that: Detailed steps for using the improved SMOTE algorithm to increase fault sample data for the secondary sampling subset: Calculate the nearest neighbor: For each fault sample x i , find its k nearest neighbor samples in the fault sample set; Generate a new sample: randomly select a neighboring sample x j , and in x i and x j Perform linear interpolation between them to generate a new fault sample x new ; Considering the continuity of the time series, the interpolation process is smoothed in the time dimension.

3. The transformer fault prediction method based on the DeepSeek large model according to claim 2 is characterized in that: Detailed steps for screening normal sample data using a clustering-based screening algorithm for a single sampling subset: The DBSCAN algorithm is used to cluster normal samples and automatically identify different clusters and noise points according to the density of the samples; Secondly, sample screening: for each cluster, calculate its central sample; According to the preset screening ratio, samples closer to the center sample are selected as retained samples, and the remaining samples are eliminated.

4. The transformer fault prediction method based on the DeepSeek large model according to claim 3 is characterized in that: The evaluation indicators for evaluating the balance degree of the datasets of the secondary sampling subset and the single sampling subset include the ratio of faulty samples to normal samples and G-mean (geometric mean); Ratio of faulty samples to normal samples: Calculate the ratio of the number of faulty samples to normal samples after subsampling and make it close to 1:1 to achieve a balanced dataset; G-mean: The geometric mean of recall and specificity, which comprehensively evaluates the performance of the model on faulty samples and normal samples.

5. The transformer fault prediction method based on the DeepSeek large model according to claim 4 is characterized in that: The DeepSeek model includes a multi-head self-attention mechanism and a feedforward neural network: The multi-head self-attention mechanism captures the dependencies between different positions in the data: For the input sequence X=[x1,x2,...,x n ], first map it to three spaces: query, key, and value, denoted as Q, K, and V respectively; Then, by calculating the attention score, we get the weighted feature representation; The specific calculation steps are as follows: Calculate the attention score: Among them, d k is the dimension of query and key, and the softmax) function is used to normalize the attention score to the [0,1] interval. Multi-head attention: Map the input sequence into multiple different subspaces, calculate multiple attention heads in parallel, and finally concatenate the results and perform a linear transformation; Feedforward neural network: The feature representation of each position undergoes nonlinear transformation through a feedforward neural network to further extract feature information; the feedforward neural network consists of two fully connected layers and an activation function (such as ReLU), and its calculation formula is: FFN(x) = max(0,xW1b1)W2+b2, where W1 and W2 are weight matrices, and b1 and b2 are bias vectors.

6. The transformer fault prediction method based on the DeepSeek large model according to claim 5 is characterized in that: The feature importance evaluation method based on game theory is used to evaluate the feature information extracted by the DeepSeek model, and feature information with high relevance to transformer fault prediction is extracted; SHAP value calculation: For a feature x i , whose SHAP value φ i Indicates the marginal contribution of the feature to the model prediction results; Feature screening: features are sorted according to the absolute value of SHAP value, and features with larger SHAP value are selected as features that are important for transformer fault prediction.

7. The transformer fault prediction method based on the DeepSeek large model according to claim 6 is characterized in that: DeepSeek model training and optimization; Loss function: The cross entropy loss function is used as the training target of the model. For the binary classification problem (fault or normal), its calculation formula is: Where N is the number of samples, y i is the true label of the sample, p i is the model's predicted probability for the sample; Optimization algorithm, uses Adam optimization algorithm to update the parameters of the model, and its update formula is: m t =β1m t-1 +(1-β1)g t v t =β2v t-1 +(1-β2)g t 2 Among them, m t and v t are the first-order moment estimate and the second-order moment estimate, β1 and β2 are the decay rates, g is the gradient at the current moment, α is the learning rate, and ∈ is a small constant used to prevent the denominator from being zero.

8. The transformer fault prediction method based on the DeepSeek large model according to claim 7 is characterized in that: Integrate multiple learning models: The bottom layer uses multiple different machine learning models to learn and predict data, and the upper layer model integrates and re-learns the output of the bottom layer model.

9. The transformer fault prediction method based on the DeepSeek large model according to claim 8 is characterized in that: Detailed steps of step S8: Data preparation: The balanced data is sorted into the following format: D={(x1,y1),(x2,y2),…,(x n ,y n )} where x n is the eigenvector, y n is the corresponding label, and the data set is divided into training set D train and the test set D test ; Underlying model training and prediction: The cross-validation method is used to train the underlying model. The specific steps are: 1) The training set D train Divide into k mutually disjoint subsets D1, D2, ..., D K . 2) For each underlying model M j (j=1,2,…,m): 3) Perform k iterations. In the i-th iteration: 4) Use Divide by D i K-1 training model M j . 5) Use the trained model M j To D i The samples in are predicted and the prediction result y is obtained i,j . 6) Concatenate the prediction results obtained by k iterations to form a new feature vector yj with a length of n train (Number of training set samples). Upper model training: The new feature matrix y=[y1,y2,…,y m ] as input, the original label y train As output, train an upper model M top ; Test set predictions: After completing the cross-validation of the underlying model and the training of the upper-level model, predictions are made on the test set.