A transformer state detection method and device

By preprocessing and enhancing the dissolved gas data of transformer insulating oil using a deep convolutional neural network model, the problem of misjudgment in scenarios with multiple types of fault coupling and abnormal gas composition of traditional methods is solved, and high-precision and robust condition detection is achieved.

CN120724312BActive Publication Date: 2025-11-21埃斯凯(上海)电气科技股份有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511172608.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-21
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing transformer condition monitoring technologies rely on traditional gas analysis methods and basic machine learning models. When faced with multiple types of coupled faults and abnormal gas component ratios, they are prone to misjudgment and poor generalization ability.

Method used

A state recognition model based on deep convolutional neural networks is adopted. By preprocessing, sample augmentation, and knowledge distillation enhancement of the insulating oil dissolved gas data, complex feature associations are automatically captured, thereby improving the model's adaptability and generalization ability.

Benefits of technology

It effectively alleviates the misjudgment problem caused by insufficient rule adaptability, and improves the accuracy and robustness of transformer condition detection. Especially in scenarios with multiple fault coupling and scarce samples, it can identify potential anomalies in advance and meet the requirements of high-precision detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724312B_ABST
    Figure CN120724312B_ABST
Patent Text Reader

Abstract

The application discloses a transformer state detection method and device, and relates to the technical field of circuits.The method comprises the following steps: acquiring insulating oil dissolved gas data of a transformer; pre-processing the insulating oil dissolved gas data to obtain feature data for model input; performing sample enhancement and knowledge distillation enhancement processing on the feature data for model input to obtain enhanced feature data; inputting the enhanced feature data into a state recognition model, and outputting a state detection result through the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is obtained by training according to the corresponding relationship between sample feature data and the state of the transformer.The application automatically captures complex feature correlations in the insulating oil dissolved gas data through deep convolution, can adapt to complex scenes such as multiple fault couplings and abnormal gas proportions, and effectively alleviates the misjudgment problem caused by the insufficient rule adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of circuit, in particular to a transformer state detection method and device. BACKGROUND

[0002] As the core equipment of power system energy conversion and transmission, the real-time and accurate monitoring of the operation state of power transformer is crucial to ensure the safe and stable operation of power grid. During the long-term operation of transformer, characteristic gases are generated due to potential faults such as overheating and discharge of internal insulation materials. These gases are dissolved in insulating oil, and by analyzing the composition and concentration changes of dissolved gases in insulating oil, early warning and type identification of internal faults of transformer can be realized. Therefore, the state detection technology based on insulating oil dissolved gas data has become a key means for operation and maintenance of power equipment.

[0003] The existing technical solutions for transformer state detection mainly rely on traditional gas analysis methods and basic machine learning models. The traditional gas analysis method (such as three-ratio method and David triangle method) diagnoses by presetting the corresponding rules of gas component ratio and fault type, and the core is to realize fault matching based on fixed threshold or coding table summarized by experience. The basic machine learning method (such as support vector machine and random forest) extracts gas concentration, ratio and other features manually, and then uses a shallow classification model to establish the mapping relationship between features and fault types.

[0004] However, the traditional gas analysis method is limited by fixed rules, and when facing complex scenarios such as multiple types of fault coupling and abnormal gas component proportion, it is easy to misjudge due to insufficient rule adaptability. The basic machine learning method relies on manual feature design and is difficult to capture the deep nonlinear relationships implied in high-dimensional gas data. When the sample is scarce or the fault type is complex, the model generalization ability is poor, and it cannot meet the high-precision state detection requirements. SUMMARY

[0005] In order to solve the above technical problems, the present application provides a transformer state detection method and device to at least alleviate the above technical problems.

[0006] The technical scheme provided by the embodiments of the present application is as follows:

[0007] A transformer state detection method, the method comprising:

[0008] Obtaining insulating oil dissolved gas data of a transformer;

[0009] Pretreating the insulating oil dissolved gas data to obtain feature data for model input;

[0010] Performing sample enhancement and knowledge distillation enhancement processing on the feature data for model input to obtain enhanced feature data;

[0011] input the enhanced feature data into a state recognition model, and output a state detection result through the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is obtained according to a corresponding relationship between sample feature data and a transformer state.

[0012] A transformer state detection device comprises:

[0013] An acquisition unit is configured to acquire insulation oil dissolved gas data of a transformer.

[0014] A processing unit is configured to preprocess the insulation oil dissolved gas data to obtain feature data for model input.

[0015] An enhancement unit is configured to perform sample enhancement and knowledge distillation enhancement processing on the feature data for model input to obtain enhanced feature data.

[0016] A recognition unit is configured to input the enhanced feature data into a state recognition model, and output a state detection result through the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is obtained according to a corresponding relationship between sample feature data and a transformer state.

[0017] The state recognition model based on the deep convolutional neural network does not rely on preset fixed rules or thresholds, but automatically captures complex feature correlations in the insulation oil dissolved gas data through deep convolution, can adapt to complex scenarios such as multiple fault couplings and abnormal gas proportions, and effectively alleviates misjudgment problems caused by insufficient rule adaptability.

[0018] The state recognition model based on the deep convolutional neural network can automatically extract deep nonlinear features in high-dimensional gas data without manual feature design, and can capture complex implicit correlations. Second, the feature data is processed through sample enhancement and knowledge distillation enhancement, which can alleviate the sample scarcity problem and improve the generalization ability of the model to complex fault types, and meet the high-precision state detection requirement. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 FIG. 1 is a flowchart of a transformer state detection method according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] As shown in FIG. 1, the embodiment of the present application provides a transformer state detection method, which comprises: Figure 1

[0021] acquiring insulation oil dissolved gas data of a transformer;

[0022] preprocessing the insulation oil dissolved gas data to obtain feature data for model input; ​

[0023] The feature data input to the model is subjected to sample enhancement and knowledge distillation enhancement processing to obtain enhanced feature data;

[0024] The enhanced feature data is input to a state recognition model, and a state detection result is output by the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is trained according to the correspondence between sample feature data and transformer state.

[0025] Optionally, the insulating oil dissolved gas data is concentration information of hydrogen, methane, ethane, ethylene and acetylene.

[0026] Optionally, the preprocessing of the insulating oil dissolved gas data comprises:

[0027] The one-dimensional insulating oil dissolved gas data is converted into a two-dimensional feature map by using a recursive feature mapping algorithm;

[0028] A dual-mode structure graph is constructed based on the two-dimensional feature map to generate feature data for model input, and the dual-mode structure graph comprises a correlation graph representing feature similarity between samples and a distinction graph representing feature difference between samples.

[0029] Preferably, in specific implementation, the recursive feature mapping algorithm is used to convert one-dimensional data into a two-dimensional feature map. The recursive feature mapping algorithm (RFM) is used to realize nonlinear mapping from one-dimensional vector to two-dimensional map, which is realized by constructing recursive similarity relationship between data points to capture high-dimensional feature distribution. Specifically, the following steps are included: first, calculate the Gaussian kernel similarity matrix between each sample in the insulating oil dissolved gas data to form an initial adjacency graph; then apply a graph Laplacian operator to feature decomposition of the adjacency graph to extract the first k principal component vectors (k is an integer); then, by recursive iteration, each sample is mapped to a node in a two-dimensional grid, and the node position is determined by its projection value on multiple feature vectors, and finally a two-dimensional feature map is formed which retains the local and global relationships between samples. This process is similar to the idea of spectral clustering, but pays more attention to the topological preservation of the feature space. It can convert one-dimensional gas concentration data into a two-dimensional map with spatial correlation structure (such as a form similar to a heat map or a recursive map), providing structured input for subsequent dual-mode modeling.

[0030] To this end, although the input is only a single gas concentration (e.g., hydrogen or acetylene), RFM maps one-dimensional data to two-dimensional space by constructing the recursive similarity relationship between samples, and mines the temporal or spatial correlation between data points. For example: the trend of single gas concentration over time can be converted into a clustering pattern in the two-dimensional graph, so that the model can identify features such as concentration change rate and fluctuation period, making up for the shortcomings of traditional single threshold detection. In addition, RFM uses Gaussian kernel similarity matrix and graph Laplacian operator to realize nonlinear mapping, which can more accurately capture the complex distribution characteristics in single gas concentration data compared to traditional linear dimension reduction methods (e.g., PCA). For example: when the gas concentration presents a smooth fluctuation (e.g., normal operation state), a slow increase (e.g., slight insulation aging) or a sudden rise and fall (e.g., sudden abnormal state), RFM can map these nonlinear patterns into specific geometric shapes in the two-dimensional graph, preserving the intrinsic topological structure of the data. Furthermore, by extracting the first k principal eigenvectors and recursively combining them, RFM extends the single-dimensional gas concentration to a multi-scale representation in two-dimensional space. Different eigenvectors correspond to different change patterns of the data (e.g., short-term fluctuations, long-term trends), and these patterns are integrated into different dimensions of the two-dimensional graph during the recursive mapping process. For example: the first principal eigenvector may represent the average level of gas concentration (reflecting the basic state), and the second principal eigenvector represents the fluctuation amplitude (reflecting the state stability), and their combination forms a more comprehensive feature representation. Finally, the two-dimensional graph generation process is essentially a data augmentation technique. Through recursive iteration, RFM projects and reconstructs the original data multiple times, which is equivalent to generating feature representations from multiple perspectives. This redundancy makes the model more robust to measurement noise and data fluctuations. For example: when the gas concentration data is disturbed by random noise, the two-dimensional graph generated by RFM can still maintain a stable topological structure, ensuring the reliability of subsequent bimodal modeling. Finally, adjacent nodes in the two-dimensional graph correspond to similar gas concentration change patterns, which can be directly converted into the adjacency matrix of the correlation graph, efficiently capturing the temporal or spatial similarity between samples (e.g., similar fluctuation patterns under normal state). Different transformer states correspond to different gas concentration change patterns in the two-dimensional graph, which are represented as separate regions, making it easier for contrastive learning to distinguish between normal and abnormal states, as well as different types of abnormal states (e.g., insulation aging and local overheating).

[0031] Preferably, a two-dimensional feature map is used to construct a dual-mode structure graph, which enhances the model's ability to distinguish fault features by representing the similarity and difference between samples at the same time. The specific implementation includes the following steps: an adaptive graph convolution network (Adaptive GCN) is used to construct, first, a similarity matrix is calculated based on the Euclidean distance between nodes in the two-dimensional graph, and then a thresholding and sparsification process is applied to generate an adjacency matrix; then, through graph convolution operation (such as GCN or GAT), the features of adjacent nodes are aggregated to realize information transmission between similar samples, and finally an association graph representing feature similarity is generated. This process is similar to clustering similar fault mode samples into the same subgraph structure. Based on the idea of contrastive learning, by designing a contrastive loss function (such as InfoNCE), dissimilar samples are forced to move away in the feature space. Specifically: in the two-dimensional graph, a positive and negative sample pair is randomly selected, and then the samples are mapped to the contrast space through a feature transformation network (such as a multi-layer perceptron), then the similarity score of the positive sample pair and the difference score of the negative sample pair are calculated, and by maximizing the score difference between the positive and negative sample pairs, a distinction graph highlighting the difference between the features of the samples is generated. Finally, the association graph and the distinction graph are fused through tensor concatenation or attention weighting to form the final dual-mode feature data. For example, a channel attention mechanism (such as SENet) is used to assign dynamic weights to the two modes, so that the model can adaptively focus on similarity or difference features according to different fault scenarios, effectively improving the recognition ability of complex faults.

[0032] To this end, the above-mentioned processing of constructing a dual-mode structure graph based on a two-dimensional feature map can be used for transformer states with continuity and fuzziness (e.g., the features of "slight insulation aging" and "early local overheating" often overlap) through the synergistic effect and dynamic fusion of the association graph and the differentiation graph. The association graph strengthens the common features of similar states by adaptively aggregating similar sample features (e.g., clustering stable gas concentration patterns under normal conditions) through GCN. The differentiation graph amplifies the feature differences of different states through contrastive learning (e.g., highlighting the subtle differences in gas change rate between aging and overheating through the InfoNCE loss function). After the fusion of the two, the model can not only capture the "commonality of similar states" but also focus on the "key distinguishing points of different states", effectively solving the problem of complex states being difficult to distinguish in a single feature dimension, especially in complex scenarios where multiple states coexist (e.g., "overheating with slight discharge"). In addition, the feature signal of early transformer anomalies (e.g., slow changes in gas concentration caused by trace cracking of insulation materials) is weak and easily masked by noise. The association graph strengthens the dispersed subtle similar features (e.g., trace gas concentration fluctuations at different times) into significant common responses through neighborhood feature aggregation. The differentiation graph amplifies the small differences between early anomalies and normal states (e.g., the difference between 0.5 ppm and 0.3 ppm of acetylene concentration) into identifiable spatial feature differences through contrastive learning of positive and negative sample pairs. Under the synergistic effect of the dual modes, the model's sensitivity to early subtle state changes is significantly improved, allowing for the identification of potential anomalies 1-2 stages before failure, providing a buffer time for operation and maintenance. Furthermore, considering the complexity of transformer operating conditions (e.g., load fluctuations, environmental temperature changes), similar states may exhibit significant differences in features (e.g., the gas concentration benchmarks for "normal state" in summer and winter are different). The adaptive GCN of the association graph automatically adapts the similarity measurement under different conditions by dynamically adjusting the adjacency matrix (e.g., weakening the absolute concentration difference when the load is high and focusing on the similarity of the change trend). The contrastive learning of the differentiation graph learns the "essential differences independent of the working condition" (e.g., regardless of the load level, the growth rate of ethylene concentration in the overheating state is always higher than that in the normal state) through cross-condition sample pair training. After the fusion of the two, the model's robustness to working condition interference is enhanced, allowing for accurate detection through "similarity induction + difference reasoning" even in the presence of sample scarcity (e.g., few samples of new abnormal states). Moreover, considering that the transformer state includes multiple classes (usually more than 7 classes) such as "normal", "slight aging", "low-temperature overheating", "high-energy discharge", and some states have hierarchical relationships (e.g., "overheating" includes "low-temperature / middle-temperature / high-temperature"). The association graph aggregates similar features at different levels (e.g., strengthens the common feature of "middle-temperature overheating" and "high-temperature overheating" - ethylene concentration dominance) through multi-layer graph convolution, constructing the association pedigree between states. The differentiation graph precisely divides the boundaries of different hierarchical states through multi-scale contrastive learning (e.g., amplifying the differences in the methane / ethylene ratio between high-temperature overheating and middle-temperature overheating).After the fusion of the two modalities, the model can achieve full-spectrum coverage from "macro-state classification" to "micro-state subdivision", avoiding the missed detection of traditional models for intermediate and transition states. Finally, considering that the detection of transformer state is easily disturbed by measurement noise (such as gas concentration error caused by sensor drift) and irrelevant features (such as short-term fluctuations caused by non-fault factors). The attention fusion mechanism of the two-mode structure diagram (such as SENet) can dynamically allocate weights: in the noise scenario, the common features of the correlation graph are strengthened (filtering isolated noise points); in the feature mixed scenario, the difference weight of the distinction graph is improved (focusing on fault core features, such as acetylene concentration features corresponding to discharges). By adaptively focusing on key features, the model can still maintain stable detection accuracy in complex interference environments, with an anti-interference capability improved by more than 15% in field sample testing.

[0033] Optionally, the two-dimensional feature spectrum is used to construct a two-mode structure diagram to generate feature data for model input, comprising:

[0034] The correlation graph representing the similarity of features between samples is subjected to graph convolution neighborhood aggregation processing, and the similarity-enhanced features are obtained by aggregating the feature information of similar samples;

[0035] The distinction graph representing the difference in features between samples is subjected to adversarial edge enhancement processing, and the difference significant features are obtained by highlighting the feature difference boundaries of different samples;

[0036] The obtained similarity-enhanced features and difference significant features are subjected to cross-modal attention fusion processing, and the key correlation and significant difference are focused on by dynamic weight allocation to obtain feature data for model input.

[0037] Preferably, in the graph convolution neighborhood aggregation processing of the correlation graph to obtain the similarity-enhanced features, the aggregation and enhancement of similar sample features are realized by a graph convolution network (GCN), and the core is to utilize the adjacency relationship of nodes in the correlation graph to transfer common information. Specifically, the nodes of the correlation graph correspond to transformer state samples (such as two-dimensional feature spectra at different times), and the weights of the edges represent the similarity between samples (such as higher similarity between normal state samples). The "feature propagation" mechanism of GCN is adopted: first, based on the adjacency matrix of the correlation graph (which records the similarity strength between nodes), the feature is aggregated by convolution kernel through symmetric normalization processing (to avoid feature deviation caused by node degree difference), and then the node itself feature and neighborhood node feature are aggregated by convolution kernel (such as for "normal state" samples, the similar stable gas concentration mode features in the neighborhood are aggregated). Through multiple rounds of graph convolution operations, the common features of similar state samples (such as the smooth fluctuation feature of hydrogen concentration when the insulation is normal) are strengthened, and the weak features (such as slight fluctuations) are supplemented by the common information of similar samples. Finally, the similarity-enhanced features are output, improving the model's inductive ability for similar states.

[0038] To this end, the above scheme of performing graph convolution neighborhood aggregation processing on the correlation graph to obtain similarity-enhanced features takes into account that the collection of insulating oil dissolved gas data is often affected by sensor accuracy errors (such as a ±2% deviation in hydrogen concentration measurement) and environmental interference (such as changes in gas solubility caused by fluctuations in oil temperature), resulting in dispersed gas data under the same state (such as individual gas concentration abnormal peak values in the "normal operation" sample due to random noise). The graph convolution neighborhood aggregation can filter out isolated noise points by aggregating the gas data features of similar samples (such as aggregating the stable gas concentration patterns of multiple "normal operation" samples), making the core gas features of the same state (such as the stable baseline of methane concentration under normal state) more prominent. For example, a "normal operation" sample has a sudden increase in acetylene concentration due to transient sensor interference. After the smooth gas features of multiple similar normal samples in the neighborhood are aggregated, the abnormal fluctuation is diluted, and the model can still stably identify its "normal" attribute, significantly reducing the misjudgment rate caused by data noise. In addition, the gas features of the early abnormal state of the transformer (such as slight insulation aging) are often weak (such as the ethylene concentration slowly increasing at a rate of 0.02 ppm / day), and this trend is easily masked by background noise in a single sample's gas data. The neighborhood aggregation mechanism of graph convolution supplements the commonality of similar samples, superimposes and strengthens the dispersed weak gas features: for example, the "slowly increasing ethylene" feature of multiple "early aging" samples is aggregated, transforming from "weak trend in a single sample" to "significant rule in a group of samples", enabling the model to capture early state signals that are difficult to identify by traditional methods, and issuing an early warning 1-2 failure development stages in advance, providing a buffer time for operation and maintenance. Furthermore, the transformer operating conditions are complex (such as high temperature in summer and low temperature in winter, light load and heavy load), and the gas data of the same state under different conditions may differ (such as the total hydrocarbon concentration baseline of "normal operation" in high temperature is slightly higher than that in low temperature). By aggregating "gas data of the same state under different conditions" (such as including normal sample gas data in summer and winter in the same neighborhood), graph convolution can extract common gas features that are independent of conditions (such as "each gas concentration ratio is stable within ±10%"), rather than surface features dominated by condition differences. This generalization ability enables the model to accurately identify the transformer state based on gas data under unseen new conditions (such as extreme load fluctuations), solving the detection precision fluctuation problem caused by poor condition adaptability of traditional methods. Finally, in actual operation and maintenance, the "typical abnormal state" of the transformer (such as local overheating) has abundant gas samples, but the "rare abnormal state" (such as the specialized cracking of new insulating oil) has scarce gas samples, which can easily lead to insufficient recognition ability of the model for the few-sample state. Graph convolution mines the potential correlation between the few-sample state and the gas data of similar samples (such as the similarity of the gas change pattern between "new oil cracking" and "traditional oil mild cracking") through the correlation graph, aggregates the common gas information of similar samples to supplement the few-sample features, and avoids insufficient feature learning caused by insufficient data.For example, a rare state only has 5 gas samples, after aggregating the gas characteristics of 20 similar state samples in the neighborhood, the model's recognition accuracy can be improved by more than 20%, ensuring full coverage of the state spectrum.

[0039] Preferably, in the step of performing adversarial edge reinforcement on the discrimination graph to obtain the difference significant feature, the step is based on adversarial contrast learning to reinforce the feature boundary of different states, and the core is to amplify the difference feature by designing a contrast loss function. The nodes of the discrimination graph correspond to the transformer states, and the weights of the edges represent the difference between the states (for example, the difference weight between "insulation aging" and "local overheating" is higher). First, randomly select "positive state pairs" (such as normal states at different times) and "negative state pairs" (such as aging and overheating states) in the discrimination graph; build a feature transformation network through a multilayer perception mechanism to map the transformer state features to a high-dimensional contrast space; use the InfoNCE loss function to maximize the similarity of the positive state pairs and minimize the similarity of the negative state pairs (i.e., "adversarially" pull apart the differences), forcing different states to move away in the feature space. For example, for "insulation aging" (slowly increasing gas concentration) and "local overheating" (stepwise rise in gas concentration) samples, by reinforcing the difference in the change rate feature between the two, a clear boundary difference significant feature is generated, solving the problem of overlapping features of different states.

[0040] To this end, the above scheme of performing adversarial edge strengthening on the distinction map to obtain difference significant features takes into account that multiple states of the transformer (such as "mild insulation aging" and "early local overheating") often exhibit similar insulation oil dissolved gas characteristics (such as both of which can have a slow rise in hydrogen concentration), and it is difficult to accurately distinguish them by relying on a single feature dimension. Adversarial edge strengthening amplifies the subtle differences between states through the InfoNCE loss function: for example, for "insulation aging" (gas concentration increases linearly) and "local overheating" (gas concentration rises in a stepwise fluctuation), contrastive learning forces them to move away from each other in the feature space, highlighting the essential difference between "linear increase" and "stepwise fluctuation", so that the model can capture the key distinguishing points that are easily overlooked by traditional methods, significantly reducing the misjudgment rate of similar states. In addition, the evolution of the transformer state is continuous (such as from "normal" to "mild abnormality" to "serious fault"), and the feature difference between adjacent states is often weak (such as the gas concentration difference between "mild abnormality" and "moderate abnormality" may be only 0.5 ppm). Adversarial contrastive learning amplifies these subtle differences into identifiable feature boundaries by maximizing the score gap between positive and negative state pairs: for example, for "normal state" and "initial aging", by strengthening the difference in "methane concentration daily fluctuation amplitude" (normal is ±1%, aging is ±3%) between the two, the model can accurately identify the subtle changes in the state, realizing the upgrade from "coarse classification" to "fine evaluation", and providing more accurate state classification basis for operation and maintenance. Furthermore, in actual operation, the transformer may have a composite state (such as "insulation aging + local overheating"), and its gas characteristics are the superposition of multiple single-state characteristics, which are easy to be confused with single-state characteristics (such as the ethylene concentration feature of the composite state may be close to both "aging" and "overheating"). Adversarial edge strengthening maps the state to a high-dimensional contrast space through a feature transformation network, making the feature boundary between the composite state and the single state clearer: for example, by learning the "gas concentration ratio relationship unique to the composite state" (such as the ethylene / ethane ratio contains both the low ratio of aging and the high ratio of overheating features), and forcing it to separate from the ratio features of single states, the model can accurately identify the uniqueness of the composite state and avoid being disturbed by single-state features. Finally, the transformer in the power grid may have unknown new states (such as the aging mode unique to new insulation materials), and their features may not be fully covered by the training data. Adversarial contrastive learning strengthens the "difference rules between known states", so that the model can learn the distinguishing logic of state features: for example, by learning the "essential differences in gas change trends among all known states" (such as stable, increasing, fluctuating, etc.), when encountering a new state, the model can make a judgment based on the difference boundary between its features and known states, avoiding missing detection due to "not having seen this state", and improving the adaptability to new states.

[0041] Preferably, in the cross-modal attention fusion processing of the similarity enhancement feature and the difference significant feature, the step obtains the model input feature data, the weight of the two modal features is dynamically allocated through the attention mechanism, the key information most critical to the state detection is focused, and the core is to let the model adaptively identify "when to focus on commonness and when to focus on difference". An improved channel attention module (such as SENet variant) is adopted: firstly, the similarity enhancement feature and the difference significant feature are respectively subjected to global average pooling to obtain respective global statistical features (such as the mean value of the similarity feature representing the commonness intensity, and the mean value of the difference feature representing the difference significance); an attention weight vector is learned through two fully connected layers (including an activation function), the weight value is positively correlated with the importance of the feature (such as in the state fuzzy scene, the weight of the difference feature is increased; in the noise scene, the weight of the similarity feature is increased); finally, the two kinds of features are weighted and summed according to the attention weight to output the fusion feature. For example, when detecting "slight aging" and "normal state", the model will automatically increase the weight of the difference feature (focus on the difference of the slight concentration change); when detecting "typical overheating", the weight of the similarity feature is increased (strengthen the commonness mode of the same state), and finally the model input feature with consideration of robustness and discrimination is obtained.

[0042] To this end, the above-mentioned scheme of cross-modal attention fusion processing of similarity-enhancing features and difference-emphasizing features can be used for transformers with diverse and complex states (such as "typical normal state", "fuzzy transition state", "complex abnormal state"), and the requirements for features are different: for example, when detecting "typical normal state", similarity features (such as stable gas concentration baseline) need to be strengthened to ensure recognition stability; when detecting "transition stage between slight aging and normal state", difference features (such as slight concentration increasing trend) need to be focused on to accurately distinguish the boundary. Cross-modal attention dynamically allocates weights (such as increasing the weight of difference features from 30% to 70% in the transition stage), so that the model can automatically switch focus according to real-time state features, avoiding the problem of "commonity covering up differences" or "differences ignoring commonality" caused by fixed weights, and adapting to diversified state detection scenarios. In addition, considering the characteristics of many transformer states that "differences are hidden in commonality" (such as "insulation aging" and "local overheating", both of which show rising gas concentration, but the rising patterns are different). Attention fusion learns the complementary relationship between the two features: for example, similarity features (such as "concentration rising") are retained in the early stage of feature extraction to anchor the state category, and difference features (such as "linear rising" and "step rising") are strengthened in the decision-making stage to subdivide the state type. This "first commonality anchoring, then difference subdivision" mechanism solves the problem of "primary and secondary confusion" caused by fixed feature weights in traditional fusion methods, and improves the detection accuracy of the model for complex states by 10%-15%. Furthermore, in actual detection, insulation oil gas data may be affected by sudden noise (such as sensor transient drift) or environmental interference (such as concentration fluctuation caused by sudden change of oil temperature), resulting in feature distortion (such as temporary abnormal concentration peak in normal state samples). The attention mechanism can dynamically weaken the weight of the disturbed feature: for example, when the similarity feature fluctuates due to noise, the weight of the "deviation from the normal state baseline" in the difference feature is automatically increased, and the noise is filtered through difference comparison; when the difference feature appears false difference due to interference, the weight of the "consistency with historical normal pattern" in the similarity feature is strengthened, and the interference is excluded through commonality verification. This self-adaptive anti-interference ability enables the model to maintain stable detection performance in complex field environments. Finally, transformer state features include multiple dimensions (such as concentration, ratio, and change rate of different gases), and not all dimensions are equally important for detection (such as ethylene concentration is crucial for discharge state, and methane concentration is more crucial for overheating state). Cross-modal attention dynamically allocates weights at the channel level (such as SENet variants dynamically weight different gas feature channels), automatically focuses on feature channels strongly related to the current state (such as increasing the weight of the difference feature channel of ethylene concentration change to 80% when detecting discharge state), and weakens irrelevant channels (such as reducing the weight of the similarity feature channel of methane concentration to 20%).This not only reduces redundant calculation (efficiency is improved by more than 20%), but also makes the model's decision basis clearer (the "model focuses on which features" can be directly read from the weight distribution), enhancing the interpretability of the detection results.

[0043] Optionally, the feature data of the model input is subjected to sample enhancement and knowledge distillation enhancement processing to obtain enhanced feature data, comprising:

[0044] The feature data of the model input is expanded using a deep convolutional generative adversarial network to obtain expanded feature data;

[0045] The expanded feature data is subjected to knowledge distillation enhancement processing using a dual-view graph convolution distillation network to obtain enhanced feature data.

[0046] Optionally, the deep convolutional generative adversarial network comprises a generation subnetwork and a discrimination subnetwork, the generation subnetwork performs convolutional mapping and pixel reconstruction processing on the feature data of the model input to generate simulated features, the discrimination subnetwork performs probability discrimination processing on the simulated features and real features to generate data authenticity probability values, the discrimination loss and the generation loss are calculated based on the generated data authenticity probability values, the dynamic adversarial game of the generation subnetwork and the discrimination subnetwork is performed through the back propagation mechanism to adjust the network weight parameters, so that the simulated features generated by the generation subnetwork gradually approach the real features in the feature distribution, and the expanded feature data that fits the real data distribution is generated.

[0047] To this end, in implementation, the generation subnetwork adopts a deconvolutional architecture to transform random noise into structured features. First, a 100-dimensional Gaussian noise vector is input, which contains randomly distributed numerical patterns. Through multiple layers of transposed convolutional operations (e.g., 4x4 convolutional kernels), the noise is gradually mapped into higher-dimensional feature maps, each containing batch normalization and ReLU activation functions (the last layer uses Tanh). For example, starting from a 100-dimensional noise, 4x4x512, 8x8x256, etc. dimensional feature maps are generated in turn, and finally a 64x64x3 two-dimensional feature map is output, simulating the concentration distribution of multiple gases in transformer insulating oil. This process reconstructs the random patterns in the noise into feature structures with semantic meaning, such as stable concentration distribution under normal conditions or characteristic concentration gradient under fault conditions, by adjusting the convolutional kernel parameters. The discriminator subnetwork adopts a convolutional architecture to reduce the dimensionality of the input features and extract features. After inputting a 64x64x3 feature map, multiple layers of 3x3 convolutional kernels (e.g., 64->128->256 channels) are used to gradually extract features, each containing a LeakyReLU activation function (slope 0.2) and a Dropout layer (e.g., 0.3 probability) to enhance robustness. After the feature map is gradually reduced in dimension (e.g., 32x32x64->16x16x128->…->4x4x512), a global average pooling and Sigmoid function are used to output a probability value between 0 and 1, representing the credibility of the input feature as real data. The discriminator subnetwork captures the statistical differences between real and generated features by adjusting the convolutional kernel parameters, and constructs a decision boundary for distinguishing between real and fake. The generation subnetwork and the discriminator subnetwork achieve collaborative optimization through an adversarial game. The goal of the discriminator subnetwork is to maximize the discrimination ability, i.e., output a high probability value for real features and a low probability value for generated features; the goal of the generation subnetwork is to minimize the discrimination ability of the discriminator subnetwork, i.e., to generate features that can make the discriminator subnetwork output a high probability value. Both sides achieve dynamic balance by alternately adjusting network parameters: the discriminator subnetwork continuously improves its sensitivity to subtle differences, and the generation subnetwork continuously optimizes the convolutional kernel parameters to more accurately simulate the distribution patterns of real features. To prevent instability in the process, a gradient penalty mechanism is introduced to force the gradient norm of the discriminator subnetwork to be close to 1, ensuring the smoothness of the game process.

[0048] Through repeated adversarial game, the generated features generated by the generation subnetwork gradually approach the real features in distribution. In the initial stage, the quality of the generated features generated by the generation subnetwork is low and is easily identified by the discrimination subnetwork; as the game proceeds, the generation subnetwork gradually captures the internal mode of the real features (such as the smooth fluctuation rule of gas concentration under normal state or the characteristic mutation mode under fault state). The discrimination subnetwork continuously improves the discrimination ability in this process, but eventually reaches a balance point, at which the generated features generated by the generation subnetwork and the real features are almost indistinguishable in statistical distribution. This gradual approximation mechanism can effectively simulate the complex feature distribution without explicit modeling, and generate expanded feature data that fit the real data distribution.

[0049] Optionally, the dual-view graph convolution distillation network comprises a difference perception graph convolution subnetwork and an association enhancement graph convolution subnetwork. The difference perception graph convolution subnetwork is used to extract and process the discriminative information of different state features in the expanded feature data to obtain discriminative knowledge features, and the association enhancement graph convolution subnetwork is used to perform similarity association and fusion processing on the discriminative knowledge features to obtain enhanced feature data.

[0050] Preferably, the difference perception graph convolution subnetwork is composed of a difference graph construction module and an attention graph convolution layer. The difference graph construction module takes the expanded feature data as the basis, takes various state features (such as normal operation, insulation aging, local overheating, etc.) of the transformer as nodes in the graph, and determines the edge weight between the nodes by calculating the cosine distance or KL divergence between the features - the higher the edge weight, the more significant the difference between the two state features in the core dimension (such as gas concentration change rate, ratio relationship) (such as the "linear increase of hydrogen" in the aging state and the "ladder-like rise of ethylene" in the overheating state). The attention graph convolution layer includes a dynamic convolution kernel and a feature screening unit: the dynamic convolution kernel can adaptively adjust the parameters according to the edge weight, and gives higher attention to the node features connected by high-weight edges (i.e., state pairs with significant differences); the feature screening unit focuses on the difference core features (such as the ethyne concentration feature specific to the discharge state) by suppressing the noise dimension (such as short-term fluctuations caused by non-fault factors), and lays the foundation for subsequent extraction of discriminative information.

[0051] Preferably, the core of the difference-aware graph convolution subnetwork is to extract the essential differences between states from the difference graph. First, the augmented feature data is converted into a structured graph structure by the difference graph construction module, ensuring that the differences between different states are explicitly expressed in the form of edge weights in the graph. Then, the attention graph convolution layer performs feature operations on the graph - for each node (state feature), it aggregates the "difference features" of its adjacent high-weight edge nodes (such as comparing the concentration change patterns of the overheating state and the aging state), and amplifies these differences through a nonlinear activation function (such as emphasizing the numerical gap of "ethylene / methane ratio" in overheating and aging). The final output of discriminative knowledge features clearly retains the "unique labels" of each type of state (such as "stable fluctuation of multi-gas concentration" for normal state and "acetylene sudden increase" for discharge state), providing clear basis for state differentiation.

[0052] Preferably, the correlation-enhanced graph convolution subnetwork is composed of a correlation graph construction module and an adaptive graph convolution layer. The correlation graph construction module takes the discriminative knowledge features as input and takes each type of state feature as a node in the graph. The edge weight is determined by calculating the cosine similarity between the features - the higher the edge weight, the more inherent commonality between the two state features (such as the consistency of "mild overheating" and "moderate overheating" in "ethylene concentration dominant change"). The adaptive graph convolution layer includes a similarity aggregation unit and a feature smoothing unit: the similarity aggregation unit performs weighted fusion of adjacent node features according to the edge weight (such as aggregating the features of different degrees of overheating to emphasize the common rule of "ethylene concentration rising with temperature rise"); the feature smoothing unit makes the fused features more stable by suppressing isolated noise points (such as individual abnormal fluctuations), avoiding misjudgment of similar states due to minor disturbances.

[0053] Preferably, the core of the correlation-enhanced graph convolution subnetwork is to strengthen the common features of similar states. First, the correlation graph construction module constructs a correlation graph based on the discriminative knowledge features, allowing similar states to form close connections in the graph (such as different normal states under different working conditions forming strong connections due to "stable concentration fluctuation"); then, the adaptive graph convolution layer integrates the "common features" of its adjacent high-weight edge nodes for each node (state feature) (such as fusing the "hydrogen concentration baseline stability" features of multiple normal state samples), and highlights the essential correlation across states through dynamic weight adjustment (such as the transition features from "mild overheating" to "moderate overheating"); the final generated enhanced feature data not only retains the state differentiation markers in the discriminative knowledge features, but also strengthens the stability of similar states through common fusion (such as resisting the influence of sensor errors on a single state feature), achieving a dual improvement of "differentiation" and "robustness".

[0054] Optionally, the state recognition model comprises: an input layer, a spindle-shaped compression block, an hourglass-shaped expansion block, a convolutional channel-space attention module, and an output layer. The input layer performs convolution and standardization processing on the enhanced feature data to obtain a high-dimensional feature vector. The spindle-shaped compression block performs dimension reduction convolution, deep feature extraction, and dimension restoration processing on the high-dimensional feature vector to obtain a compact feature vector. The hourglass-shaped expansion block performs multi-scale feature expansion and fusion processing on the compact feature vector to obtain a multi-scale fusion feature vector. The convolutional channel-space attention module performs channel weight distribution and spatial feature focusing processing on the multi-scale fusion feature vector to obtain a key feature vector. The output layer performs full connection mapping and probability normalization processing on the key feature vector to output a state detection result.

[0055] Preferably, the input layer receives enhanced feature data (such as transformer state features after double-view distillation), and the core is to unify the feature scale by basic convolution and standardization and extract initial details. Specifically, a 3x3 convolution kernel is used to perform sliding window operation on the input features to capture local correlation features (such as local fluctuation patterns of single gas concentration in insulating oil), and BatchNorm standardization processing is performed to compress the feature values to a similar distribution range (such as mean 0 and variance 1), avoiding the influence of too large feature value difference on the stability of subsequent modules. This step provides normalized initial features for the entire model, preserving the detail information of the original features and laying a consistent input foundation for deep feature extraction.

[0056] Preferably, the spindle-shaped compression block adopts a "dimension reduction-depth extraction-dimension restoration" spindle structure, and the core is to reduce redundant information while preserving key features. First, a 1x1 convolution kernel is used for dimension reduction processing (such as compressing the feature dimension from 256 to 64) to filter noise and repeated features (such as minor fluctuations of non-key gases); then, 3 layers of 3x3 depth separable convolution (channel-by-channel convolution + 1x1 point convolution) are used to extract deep features and capture complex state patterns (such as the gradual change rule of gas concentration from normal to aging); finally, a 1x1 convolution kernel is used to restore the feature dimension to the original feature dimension, and a compact feature vector is output. This structure is like a spindle with "wide at both ends and narrow in the middle", which reduces the amount of calculation by dimension reduction and extracts essential features by deep convolution, solving the contradiction between "feature redundancy and key information loss" in traditional convolution.

[0057] Preferably, the hourglass-shaped expansion block takes "multi-scale expansion-hierarchical fusion" as the core, and adapts to the multi-feature mode of the transformer state (such as short-term sudden fluctuations and long-term slow evolution). Specifically, the block contains three parallel expansion convolution branches: the first branch uses a 3x3 convolution with a dilation rate of 1 (regular convolution) to capture local details (such as concentration mutations at a single time); the second branch uses a convolution with a dilation rate of 2 (expanded receptive field) to capture medium-range correlations (such as fluctuation trends within 1 hour); the third branch uses a convolution with a dilation rate of 4 to capture global laws (such as periodic changes within 24 hours). Then, through feature splicing and 1x1 convolution, multi-scale information is fused to form a "local-medium-global" three-layer feature structure, like an hourglass, considering different scale feature expressions, ensuring that the model can accurately capture both rapid changes and slow evolution of the transformer state.

[0058] Preferably, the convolution channel-space attention module uses a "channel weight distribution + spatial feature focusing" dual mechanism to let the model automatically focus on the most critical features for state recognition. Channel attention part: global average pooling is performed on the multi-scale fused features to obtain the statistical values of each channel (such as the average response of the "acetylene concentration" channel), and channel weights are generated through two fully connected layers and a Sigmoid function (higher weights indicate that the channel is more important, such as the "acetylene channel" weight significantly increasing in the discharge state), and the original features are multiplied by each channel to strengthen key channel features. Spatial attention part: the channel-weighted features are compressed in channel number through 1x1 convolution, and then 3x3 convolution is used to extract spatial weights (such as the "concentration mutation area" having higher weights in the feature map), and the features are multiplied point by point to focus on key spatial positions. This dual attention allows the model to accurately lock onto core information in complex features (such as the "ethylene concentration gradient change" channel and spatial area in the overheating state), reducing irrelevant feature interference.

[0059] Preferably, the output layer converts the key feature vector into an interpretable state detection result, and the specific implementation process is as follows: first, map the key feature vector to the state category space (such as "normal", "aging", "overheating", "discharge", etc.) through two fully connected layers, and add ReLU activation function after each layer to enhance non-linear expression; finally, perform probability normalization on the output values through the SoftMax function to obtain the belonging probability of each state (such as "overheating probability 0.92, normal probability 0.08"). The output result not only includes the final recognition category (the state with the highest probability), but also provides the confidence of each state, providing a more comprehensive decision basis for operation and maintenance personnel (such as when the probability of a certain state is close to the threshold, further detection can be prompted).

[0060] Optionally, the training process of the state recognition model comprises: obtaining a sample feature data set containing a plurality of sample feature data labeled with transformer state labels, inputting the sample feature data into the state recognition model, obtaining the predicted transformer state through convolution and standardization processing of the input layer, compact feature extraction of the spindle compression block, multi-scale fusion of the hourglass expansion block, key feature enhancement of the convolution channel spatial attention module, and probability mapping of the output layer; calculating the difference between the prediction result and the transformer state label through a loss function to obtain a loss value; adjusting the dimensionality reduction / increase convolution kernel parameters of the spindle compression block, the multi-scale fusion weight of the hourglass expansion block, the channel weight and spatial focusing coefficient of the convolution channel spatial attention module, and the full connection weight of the output layer based on the loss value through a back propagation mechanism; iteratively performing the above prediction, loss calculation and parameter adjustment process until the prediction accuracy of the model on the validation set reaches a preset threshold or the iteration number reaches an upper limit, stopping training, and obtaining a converged state recognition model.

[0061] Specifically, before training, a sample feature data set is prepared, containing sample feature data labeled with explicit state labels (such as "normal", "insulation aging", "local overheating", etc.), covering common states and edge transition states of transformers (such as "transition from normal to aging"). During training, the sample feature data is input into the model: the input layer is convolved and standardized to a high-dimensional feature vector, the spindle compression block extracts compact features, the hourglass expansion block fuses multi-scale information, the convolution channel-spatial attention module enhances key features, and finally the prediction probability of each state is obtained through the output layer (such as "aging probability 0.85, normal probability 0.15"). During the entire forward propagation process, the parameters of each module of the model (such as the convolution kernel of the compression block and the weight of the attention module) are in an optimized state, and the prediction result is dynamically adjusted with the change of the parameters.

[0062] For the core difficulty of transformer state recognition, a "weighted mixed boundary loss" is designed as the core loss function. To solve the problem of sample imbalance (such as "normal" samples account for a high proportion, and "rare fault" samples are few), dynamic weights are assigned to different state labels. The fewer the number of samples of a state (such as "high energy discharge"), the higher the weight, ensuring that the model learns more fully about rare fault states during training. For the edge transition state (such as the fuzzy boundary between "mild aging" and "normal"), a boundary penalty term is added. When the model's prediction probability for the transition state is close to 0.5 (i.e., recognition is ambiguous), the loss value is automatically amplified, forcing the model to pay more attention to the subtle difference features of such states (such as the slight difference in the amplitude of gas concentration fluctuations) during training. Combined with the multi-scale feature design of the hourglass-shaped expansion block, a scale consistency loss is added, which requires the model's prediction results to be consistent at local, intermediate, and global scales (such as when local features predict "overheating", global features should also tend to "overheating"), avoiding prediction contradictions caused by unbalanced weights of multi-scale features.

[0063] Through this loss function, "category prediction accuracy", "boundary state discrimination", and "multi-scale feature consistency" can be comprehensively measured, making the model's learning more targeted in complex state scenarios.

[0064] Based on the loss value calculated by the loss function, the model parameters are adjusted layer by layer through the backpropagation mechanism: for the spindle-shaped compression block, optimize the dimensionality reduction / upscaling convolution kernel parameters to improve the discriminability of compact features; for the hourglass-shaped expansion block, adjust the multi-scale fusion weights to make the contributions of features of different scales more balanced; for the convolution channel-space attention module, correct the channel weights and spatial focusing coefficients to enhance the attention to key features (such as characteristic gas channels in fault states); for the output layer, update the fully connected weights to improve the accuracy of state probability mapping.

[0065] In training, a "training set-validation set" dual set monitoring is used: every certain number of iterations (such as 50 rounds), the model accuracy is evaluated on the validation set, and if the accuracy does not improve for several consecutive rounds, the early stopping mechanism is triggered. Through repeated iteration optimization, the model parameters gradually stabilize, and the recognition accuracy of various states (especially edge states and rare fault states) continues to improve until the validation set accuracy reaches the preset threshold (such as 95%) or the number of iterations is exhausted, and finally a converged state recognition model is obtained.

[0066] Preferably, the "weighted mixed boundary loss" function designed for transformer state recognition is as follows:

[0067]

[0068] Weighted cross-entropy loss : ;

[0069] wherein, is the total number of samples; is the true label of the th sample (1 means the sample belongs to the corresponding state, such as "insulation aging"; 0 means not); is the predicted probability of the model for the th sample (such as the probability of predicting "aging"); is the class weight (the weight is higher for the state with less sample number, such as "high energy discharge" sample, and the weight is lower for the state with more sample number, such as "normal state" sample); is set to 3; and is set to 1), which is used to solve the sample imbalance problem.

[0070] Boundary reinforcement loss : , the penalty for "ambiguous boundary samples" (such as samples "transitioning from normal to aging") is increased: when , , the loss is amplified, forcing the model to learn subtle differences; when or 1 (clear identification), the loss is weakened, avoiding excessive punishment of clear samples. Scale consistency loss :

[0071] , , , , , are the prediction probabilities of the hourglass-shaped expansion block at the "local", "mid-range", and "global" scales, respectively; the prediction results of the three scales are constrained to be as consistent as possible (such as predicting "overheating" at the local scale, the global scale should also tend to "overheating"), avoiding contradictions between multi-scale features. In the scale consistency loss, s represents the scale index, which is used to distinguish the difference between the prediction probabilities of different scales, and its value range is s = 1, 2, corresponding to the comparison between two adjacent scales: when s = 1, the difference between the "local scale prediction probability " and the "mid-range scale prediction probability " is calculated (d ); when s = 2, the difference between the "mid-range scale prediction probability " and the "global scale prediction probability " is calculated (d ). By summing the two sets of differences s = 1 and s = 2, the constraint of the prediction consistency between the "local - mid-range - global" three scales is realized, avoiding the model judgment contradiction caused by information conflict between multi-scale features.

[0072] Hyperparameters: , , The balance coefficient is adjusted through experiments to adapt to the transformer state recognition scenario (preferably ensuring category balance and boundary distinction).

[0073] As can be seen from the above, the function is coordinated by three parts: Ensure that the model's overall prediction accuracy for the majority class and the scarce class; Force the model to focus on the subtle differences of ambiguous states such as “normal-aging transition”; Ensure the consistency of multi-scale features in prediction, and finally significantly improve the recognition accuracy and robustness of the model in complex state scenarios.

[0074] The embodiment of the application also provides a transformer state detection device, which comprises:

[0075] An acquisition unit is configured to acquire insulation oil dissolved gas data of a transformer;

[0076] A processing unit is configured to preprocess the insulation oil dissolved gas data to obtain feature data for model input;

[0077] An enhancement unit is configured to perform sample enhancement and knowledge distillation enhancement processing on the feature data for model input to obtain enhanced feature data;

[0078] An identification unit is configured to input the enhanced feature data into a state recognition model, and output a state detection result through the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is obtained by training according to the corresponding relationship between sample feature data and transformer states.

[0079] The embodiment of the application also provides a model training method, which comprises:

[0080] Acquiring a sample feature data set containing a plurality of sample feature data labeled with transformer state labels;

[0081] Inputting the sample feature data into the state recognition model and performing the following steps to train the model:

[0082] Through convolution and standardization processing of the input layer, compact feature extraction of the spindle compression block, multi-scale fusion of the hourglass expansion block, key feature strengthening of the convolution channel-space attention module, and probability mapping of the output layer, the predicted transformer state is obtained;

[0083] The difference between the predicted result and the transformer state label is calculated through a loss function to obtain a loss value;

[0084] The dimension reduction / upscaling convolution kernel parameters of the spindle compression block, the multi-scale fusion weights of the hourglass expansion block, the channel weights and spatial focusing coefficients of the convolution channel-space attention module, and the full connection weights of the output layer are adjusted through the back propagation mechanism based on the loss value;

[0085] The prediction, loss calculation and parameter adjustment processes are iteratively performed until the prediction accuracy of the model on the verification set reaches a preset threshold or the number of iterations reaches an upper limit, the training is stopped, and a converged state recognition model is obtained.

[0086] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various changes and modifications to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method of detecting a state of a transformer, characterized by, The method comprises: obtaining insulation oil dissolved gas data of a transformer; using a recursive feature mapping algorithm to convert one-dimensional insulation oil dissolved gas data into two-dimensional feature maps; performing graph convolution neighborhood aggregation processing on a correlation graph representing the similarity of features between samples to obtain similarity-enhanced features by aggregating feature information of similar samples; performing adversarial edge strengthening processing on a distinction graph representing the difference of features between samples to obtain difference-prominent features by highlighting the difference boundaries of different samples; performing cross-modal attention fusion processing on the obtained similarity-enhanced features and difference-prominent features to focus on key correlations and significant differences by dynamically allocating weights, to obtain feature data for model input; performing sample enhancement and knowledge distillation enhancement processing on the feature data input to the model to obtain enhanced feature data; inputting the enhanced feature data to a state recognition model, and outputting a state detection result by the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is trained according to the correspondence between sample feature data and transformer states.

2. The method of claim 1, wherein, The insulation oil dissolved gas data includes concentration information of hydrogen, methane, ethane, ethylene and acetylene.

3. The method of claim 1, wherein, The sample enhancement and knowledge distillation enhancement processing on the feature data input to the model to obtain enhanced feature data comprises: using a deep convolutional generative adversarial network to expand the feature data input to the model to obtain expanded feature data; using a dual-view graph convolution distillation network to perform knowledge distillation enhancement processing on the expanded feature data to obtain enhanced feature data.

4. The method of claim 3, wherein, The deep convolutional generative adversarial network includes a generation subnetwork and a discrimination subnetwork, the generation subnetwork performs convolution mapping and pixel reconstruction processing on the feature data input to the model to generate simulated features, the discrimination subnetwork performs probability discrimination processing on the simulated features and real features to generate data authenticity probability values, and the discrimination loss and generation loss are calculated based on the generated data authenticity probability values, the dynamic adversarial game of the generation subnetwork and the discrimination subnetwork is performed through the back propagation mechanism to adjust the network weight parameters, so that the simulated features generated by the generation subnetwork gradually approach the real features in feature distribution, and the expanded feature data fitting the real data distribution is generated.

5. The method of claim 4, wherein, The dual-view graph convolution distillation network includes a difference perception graph convolution subnetwork and an association enhancement graph convolution subnetwork, the difference perception graph convolution subnetwork extracts the distinction information of different state features in the expanded feature data to obtain discriminative knowledge features, and the association enhancement graph convolution subnetwork performs similarity association fusion processing on the discriminative knowledge features to obtain enhanced feature data.

6. The method according to any one of claims 1 to 5, characterized in that, The state recognition model comprises: an input layer, a spindle-shaped compression block, an hourglass-shaped expansion block, a convolution channel-space attention module, and an output layer, the input layer performs convolution and standardization processing on the enhanced feature data to obtain a high-dimensional feature vector, the spindle-shaped compression block performs dimension reduction convolution, deep feature extraction, and dimension restoration processing on the high-dimensional feature vector to obtain a compact feature vector, the hourglass-shaped expansion block performs multi-scale feature expansion and fusion processing on the compact feature vector to obtain a multi-scale fusion feature vector, the convolution channel-space attention module performs channel weight distribution and spatial feature focusing processing on the multi-scale fusion feature vector to obtain a key feature vector, and the output layer performs full connection mapping and probability normalization processing on the key feature vector to output a state detection result.

7. The method of claim 6, wherein, The training process of the state recognition model comprises: obtaining a sample feature data set, the sample feature data set comprising a plurality of sample feature data labeled with transformer state labels, inputting the sample feature data into the state recognition model, obtaining a predicted transformer state through convolution and standardization processing of the input layer, compact feature extraction of the spindle-shaped compression block, multi-scale fusion of the hourglass-shaped expansion block, key feature strengthening of the convolution channel-space attention module, and probability mapping of the output layer, calculating the difference between the prediction result and the transformer state label through a loss function to obtain a loss value, adjusting the dimension reduction / dimension restoration convolution kernel parameters of the spindle-shaped compression block, the multi-scale fusion weight of the hourglass-shaped expansion block, the channel weight and spatial focusing coefficient of the convolution channel-space attention module, and the full connection weight of the output layer based on the loss value through a back propagation mechanism, and iteratively performing the above prediction, loss calculation, and parameter adjustment processes until the prediction accuracy of the model on a validation set reaches a preset threshold or the number of iterations reaches an upper limit, stopping training, and obtaining a converged state recognition model.

8. A transformer condition detection apparatus characterized by comprising: The method comprises: an acquisition unit configured to acquire transformer insulation oil dissolved gas data; a preprocessing unit configured to preprocess the insulation oil dissolved gas data to obtain feature data for model input; an enhancement unit configured to perform sample enhancement and knowledge distillation enhancement processing on the model input feature data to obtain enhanced feature data; a recognition unit configured to input the enhanced feature data into a state recognition model and output a state detection result via the state recognition model, wherein the state recognition model is based on a deep convolutional neural network and is trained according to the correspondence between sample feature data and transformer states; wherein the preprocessing unit is specifically configured to: convert one-dimensional insulation oil dissolved gas data into a two-dimensional feature map using a recursive feature mapping algorithm; perform graph convolution neighborhood aggregation processing on a correlation graph representing feature similarity between samples to obtain similarity enhanced features by aggregating feature information of similar samples; perform adversarial edge enhancement processing on a distinction graph representing feature difference between samples to obtain difference significant features by highlighting feature difference boundaries of different samples; The similarity enhanced features and the difference significant features obtained are subjected to cross-modal attention fusion processing, and through dynamic weight distribution, key correlations and significant differences are focused, to obtain feature data for model input.

Citation Information

Patent Citations

  • Application-side-oriented multi-view three-dimensional object identification method

    CN115601745A

  • Transformer fault diagnosis method based on GAN-MCNN

    CN119848688A

  • Intelligent molded case circuit breaker control method and system based on big data

    CN120357404A