Knowledge distillation-based electrical equipment state transition diagnosis method and related equipment
By constructing a finite element model of power equipment to generate simulation data, training the teacher network and transferring knowledge to the student network, and combining adversarial generation and domain adversarial training, the problems of data scarcity and low diagnostic accuracy in cross-device diagnosis of power equipment are solved, and high-precision state migration diagnosis is achieved.
Patent Information
- Application Number
- CN202510515121.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the power equipment state transition diagnosis model has a reduced diagnostic accuracy during cross-device diagnosis due to differences in data distribution, and the scarcity of fault data makes model training difficult, making it difficult to achieve high-accuracy cross-device diagnosis.
By constructing a finite element model of power equipment to generate simulated fault data, the teacher network is trained and knowledge is transferred to the student network through KL divergence. Combined with adversarial generative networks and domain adversarial training, the student network is adjusted to adapt to real-time working conditions, forming a working condition adaptive diagnosis model.
It achieves high-precision state transition diagnosis across devices, reduces dependence on real fault data, and improves the generalization ability and diagnostic accuracy of the student network.
Smart Images

Figure CN120633369A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a power equipment state migration diagnosis method based on knowledge distillation and related equipment. Background Art
[0002] In the advancement of intelligent power system operation and maintenance, deep learning technology based on knowledge distillation offers a new direction for power equipment state transition diagnosis. Traditional approaches often build deep neural networks and train models using large amounts of labeled data to accurately extract power equipment fault characteristics and diagnose their status. However, in real-world scenarios, acquiring power equipment fault samples is extremely difficult. On the one hand, equipment failures occur infrequently and last only a short time, making it difficult to capture sufficient fault data. On the other hand, fault data labeling requires professional technicians to analyze the equipment's operating history and fault mechanisms, which is labor-intensive and inefficient, resulting in an extremely scarce supply of labeled data for model training.
[0003] At the same time, different types of power equipment, such as generators, reactors, and relay protection devices, have significantly different distribution characteristics of their monitoring data in the time and frequency domains due to differences in operating conditions, manufacturing processes, and structural designs. Existing diagnostic models based on knowledge distillation, when attempting to transfer knowledge from a teacher model to a student model for cross-device diagnosis, fail to effectively address data distribution differences, making it difficult for the migrated model to adapt to the data characteristics of the new equipment, resulting in a significant decrease in diagnostic accuracy. This situation severely restricts the widespread application of deep learning technology based on knowledge distillation in power equipment state migration diagnosis. There is an urgent need to propose innovative methods to overcome the technical bottlenecks of data scarcity and low cross-device diagnostic accuracy.
[0004] In view of this, a power equipment state transition diagnosis method and related equipment based on knowledge distillation are needed. Summary of the Invention
[0005] To address the low cross-device diagnostic accuracy problem in existing technologies, the present invention provides a power equipment state transition diagnostic method and related equipment based on knowledge distillation, which can improve cross-device diagnostic accuracy. The specific technical solution is as follows:
[0006] In a first aspect, an embodiment of the present application provides a method for diagnosing power equipment state transition based on knowledge distillation, comprising:
[0007] Step S1: constructing a finite element model of the power equipment and generating simulated fault data based on the finite element model;
[0008] Step S2: training a teacher network based on the simulated fault data and training a student network based on the measured fault data, wherein the measured fault data is data obtained by measuring real faulty power equipment; wherein the teacher network and the student network are preset deep learning networks;
[0009] Step S3: Transfer the knowledge of the teacher network to the student network through KL divergence;
[0010] Step S4: According to the real-time working condition parameters of the equipment to be diagnosed, the parameters of the student network are adjusted to obtain a working condition adaptive diagnosis model.
[0011] Preferably, step S1 includes: inputting the physical parameters of the power equipment in a fault state into the finite element model to obtain a fault feature map output by the finite element model; performing data enhancement on the fault feature map through an adversarial generative network to obtain the simulated fault data; wherein the generator and the discriminator of the adversarial generative network are adversarially trained based on the physical field distribution error of the power equipment.
[0012] Preferably, the loss function of the adversarial generative network includes a physical constraint term and a distribution alignment term; wherein the physical constraint term is the L2 norm error between the generated data of the adversarial generative network and the solution of the Maxwell equation; and the distribution alignment term is the Wasserstein distance between the generated data and the real fault data.
[0013] Preferably, the teacher network uses a deep residual network to process the simulated fault data to output a high-dimensional feature vector and a fault probability distribution of the simulated fault data; the student network uses a convolutional network to process the measured fault data; and the student network aligns the feature distribution of the teacher network by minimizing the KL divergence.
[0014] Preferably, the calculation formula of the KL divergence includes:
[0015] L KD =T 2 *∑[P teacher (x)*log(P teacher (x) / P student (x))];
[0016] Among them, L KD represents KL divergence, T is the temperature coefficient, P teacher (x), P student (x) are the probability distributions of the teacher network and the student network output based on the input x.
[0017] Preferably, step S4 includes: constructing a domain adversarial training module, the domain adversarial training module including a feature extractor, a domain discriminator and a gradient reversal layer, the parameters of the feature extractor are the parameters of the student network, and the gradient reversal layer is used to reverse the gradient of the domain discriminator when backpropagating the feature extractor; extracting the features of the simulated fault data and the features of the real-time operating condition parameters through the feature extractor; calculating the domain difference loss between the simulated fault data and the real-time operating condition parameters based on the features of the simulated fault data and the features of the real-time operating condition parameters through the domain discriminator; wherein, the feature extractor performs domain invariant feature learning through the gradient reversal layer, and the loss function is:
[0018] L total =L classify -λ·L domain ;
[0019] λ=2·(1+e -10p ) -1 -1;
[0020] Among them, L total is the total loss, L classify is the classification loss, L domain is the domain discrimination loss, λ is the domain adaptation loss weight, and p is the training progress ratio; after the preset convergence conditions are reached, the parameters of the student network are adjusted based on the parameters of the feature extractor.
[0021] Preferably, the feature extractor further includes a multi-head self-attention module; the domain difference loss between the simulated fault data and the measured fault data is calculated based on the features of the simulated fault data and the features of the measured fault data, including: calculating the feature channel weight of the feature according to the real-time working condition parameter through the multi-head self-attention module; the calculation formula of the feature channel weight includes:
[0022]
[0023] Among them, α i is the weight of the i-th feature channel, h i is the i-th feature channel vector, W q is the working condition parameter mapping matrix, d is the number of feature channels; based on the features of the simulated fault data and the measured fault data, as well as the feature channel weights, the domain difference loss of the simulated fault data and the measured fault data is calculated.
[0024] In a second aspect, an embodiment of the present application provides a power equipment state transition diagnosis system based on knowledge distillation, which is applied to the method described in the first aspect, and the system includes:
[0025] A generation module, used to construct a finite element model of the power equipment and generate simulated fault data based on the finite element model;
[0026] A training module, configured to train a teacher network based on the simulated fault data and a student network based on measured fault data, wherein the measured fault data is data obtained by measuring real faulty power equipment; wherein the teacher network and the student network are preset deep learning networks;
[0027] The migration module is used to transfer the knowledge of the teacher network to the student network through KL divergence;
[0028] The adjustment module is used to adjust the parameters of the student network according to the real-time working condition parameters of the equipment to be diagnosed, so as to obtain a working condition adaptive diagnosis model.
[0029] In a third aspect, an embodiment of the present application provides a computing device, comprising: a memory for storing a program; and a processor for loading the program to execute the method described in the first aspect.
[0030] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the method as described in the first aspect.
[0031] Compared with the existing technology, the beneficial effects of the present invention are as follows: by constructing a finite element model of the power equipment to generate simulated fault data, the dependence on real fault data can be reduced, solving the problem of scarce fault sample data; then, the teacher network is trained based on the simulated fault data, and the student network is trained based on the measured fault data, and knowledge transfer from the teacher network to the student network is achieved through KL divergence, thereby improving the generalization ability of the student network; then, the parameters of the student network are adjusted according to the real-time operating parameters of the equipment to be diagnosed, and the resulting operating condition adaptive diagnosis model can more accurately diagnose the equipment to be diagnosed. Using the embodiments of the present application, high-precision state migration diagnosis across devices can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly describes the drawings required for the specific embodiments or the description of the prior art. Similar elements or parts are generally identified by similar reference numerals throughout the drawings. Elements or parts in the drawings are not necessarily drawn to scale.
[0033] Figure 1 A schematic diagram of a flow chart of a power equipment state transition diagnosis method based on knowledge distillation provided in an embodiment of the present application;
[0034] Figure 2A schematic diagram of the structure of a power equipment state transition diagnosis system based on knowledge distillation provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0037] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0038] It should also be understood that the terms used in the present specification are only for the purpose of describing particular embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0039] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0040] In order to solve the problem of low cross-device diagnosis accuracy in traditional methods, the present invention provides a power equipment state migration diagnosis method and related equipment based on knowledge distillation, which can improve the cross-device diagnosis accuracy.
[0041] See also Figure 1 , Figure 1 The present invention provides a flow chart of a method for diagnosing state transition of power equipment based on knowledge distillation, which is applied to computing equipment; Figure 1 As shown, the method includes:
[0042] Step S1: The computing device constructs a finite element model of the power equipment and generates simulated fault data based on the finite element model.
[0043] Among them, computing equipment can construct finite element models of power equipment through simulation software.
[0044] A finite element model is a numerical computational model based on the finite element method, which transforms actual physical systems (such as structures, fluids, and electromagnetic fields) into mathematical models. Specifically, a finite element model discretizes a continuous physical system into a finite number of grids. Mathematical equations are established for each element and assembled together to form a set of equations that describe the entire system.
[0045] Preferably, the finite element model includes an electromagnetic field control equation, a coupled heat conduction equation, and constraints, wherein the constraints include structural parameters and material property constraints; the expression of the electromagnetic field control equation includes:
[0046]
[0047] Where A is the magnetic vector potential, μ is the magnetic permeability, J is the current density vector, σ is the conductivity, and t is the time variable. is the Hamiltonian operator, represents the rate of change of magnetic vector potential with time; the expression of the coupled heat conduction equation includes:
[0048]
[0049] Where ρ is the object density, C p is the specific heat capacity at constant pressure, T is the temperature, k is the thermal conductivity, Q Joule represents the Joule heat source term, is the rate of change of temperature with time, is the gradient operator.
[0050] Preferably, when the electric power equipment includes an iron core, the structural parameters include the dimensional parameters of the iron core, and the material property constraints include the nonlinear BH curve of the magnetic permeability of the iron core; when the electric power equipment is a gas-insulated metal-enclosed switchgear GIS equipment, the structural parameters include the conductor geometric parameters, and the material property constraints include the SF6 gas discharge characteristic curve; when the electric power equipment is a permanent magnet synchronous motor, the structural parameters include the permanent magnet pole arc coefficient, and the material property constraints include the NdFeB permanent magnet demagnetization curve.
[0051] For different types of power equipment, different constraints can be established to represent typical characteristics of the power equipment, so as to directly associate the failure mechanism of the corresponding type of power equipment.
[0052] Preferably, the computing device can input the physical parameters of the power equipment in a fault state into the finite element model to obtain a fault characteristic map output by the finite element model; perform data enhancement on the fault characteristic map through an adversarial generative network to obtain the simulated fault data; wherein the generator and the discriminator of the adversarial generative network are adversarially trained based on the physical field distribution error of the power equipment.
[0053] The computing device may sequentially input physical parameters under different fault states into the finite element model to obtain fault characteristic maps corresponding to different fault types or fault modes.
[0054] The computing device can input basic simulation data and random physical disturbances into the generator to obtain simulated fault data that conforms to physical laws. The basic simulation data is the field distribution matrix of the finite element model after inputting the physical parameters under the fault state.
[0055] Among them, the computing device can input the generated data or the measured data into the discriminator to obtain the data source probability judged by the discriminator; and then calculate the Wasserstein distance between the generated data and the measured data.
[0056] Preferably, the loss function of the adversarial generative network includes a physical constraint term and a distribution alignment term; wherein the physical constraint term is the L2 norm error between the generated data of the adversarial generative network and the solution of the Maxwell equation; and the distribution alignment term is the Wasserstein distance between the generated data and the real fault data.
[0057] Maxwell's equations are fundamental equations that describe electromagnetic phenomena. By calculating the L2-norm error between the generated data and the solutions to the Maxwell equations, we ensure that the generated data is physically plausible and conforms to the fundamental laws of electromagnetism. This helps improve the quality and credibility of the generated data, enabling it to better reflect real-world physical phenomena. The physical constraint ensures that the model, during training, not only focuses on the similarity between the generated data and real data but also on satisfying the constraints of the physical equations. This additional constraint helps the model learn more essential features and laws, thereby improving its generalization ability and enabling it to better generate physically consistent results when faced with new, unseen data. The physical constraint provides the model with prior knowledge that the generated data must satisfy the physical equations. This prevents the model from overfitting to noise or irrelevant features in the training data, allowing it to learn more general and representative features, thereby reducing the risk of overfitting.
[0058] Among them, the Wasserstein distance is more sensitive to the translation and scaling of the distribution and can better capture the shape and structural differences of the distribution. Therefore, using the Wasserstein distance as a distribution alignment term can more effectively match the distribution of generated data with the distribution of real fault data. By minimizing the Wasserstein distance, the model will strive to generate samples with a distribution similar to that of real fault data, which helps to increase the diversity of generated data. Because real fault data often has certain distribution characteristics, the model will generate a variety of different samples in the process of learning this distribution, rather than just samples similar to the training data. The Wasserstein distance has good mathematical properties during the optimization process and can provide more stable gradient information, which helps the optimization algorithm converge better.
[0059] Step S2: The computing device trains a teacher network based on the simulated fault data and trains a student network based on the measured fault data.
[0060] The measured fault data is data obtained by measuring real faulty power equipment; and the teacher network and the student network are preset deep learning networks.
[0061] In the field of deep learning, the teacher network and student network represent a model architecture and training strategy. The teacher network is a large, complex model. Trained on large datasets, it achieves excellent performance in power equipment fault identification tasks and possesses extensive knowledge and experience. The student network, on the other hand, is a relatively small and simple model designed for computational efficiency or to identify faults for specific types of equipment. It has fewer parameters and is therefore easier to run on resource-constrained devices.
[0062] Preferably, the teacher network uses a deep residual network to process the simulated fault data, so as to output a high-dimensional feature vector and a fault probability distribution of the simulated fault data; and the student network uses a convolutional network to process the measured fault data.
[0063] Traditional deep neural networks can experience vanishing or exploding gradients as the number of layers increases, making them difficult to train. Deep residual networks address this problem by introducing residual blocks. The core idea of residual blocks is to enable the network to learn residual mappings through skip connections. This allows gradients to be transferred more smoothly through skip connections during backpropagation, avoiding vanishing or exploding gradients and enabling deeper training layers.
[0064] Through deep residual networks, the teacher network is able to build a deeper network structure, thereby learning more complex and abstract data features. This structure allows the teacher network to process data more accurately and provide high-quality knowledge to the student network.
[0065] The inference speed can be improved by diagnosing the power equipment to be diagnosed through a student network with a smaller number of parameters.
[0066] Step S3: The computing device transfers the knowledge of the teacher network to the student network through KL divergence.
[0067] Among them, KL divergence (Kullback-Leibler Divergence), also known as relative entropy, is a measurement method used to measure the difference between two probability distributions.
[0068] Preferably, the student network aligns the feature distribution of the teacher network by minimizing the KL divergence. By minimizing the KL divergence between the feature distribution of the teacher network and the feature distribution of the student network, the student network can learn the feature distribution of the teacher network.
[0069] Preferably, the calculation formula of the KL divergence includes:
[0070] L KD =T 2 *Σ[P teacher (x)*log(P teacher (x) / P student (x))];
[0071] Among them, L KD represents KL divergence, T is the temperature coefficient, P teacher (x), P student (x) are the probability distributions of the teacher network and the student network output based on the input x.
[0072] The temperature coefficient T decays dynamically with the training rounds, with an initial value of T0=5 and a decay coefficient η=0.95.
[0073] The loss function in the feature distribution alignment process can be expressed as:
[0074] L total =αL KD +(1-α)L CE ;
[0075] Among them, L total is the total loss, L CE is the loss of the student network, and α is the weight coefficient.
[0076] Minimizing the KL divergence (KL-divergence) achieves alignment of the teacher and student network feature distributions. Essentially, this transfers the "knowledge" (intrinsic patterns of the data) learned by the teacher network to the student network in the form of a probability distribution. This process not only addresses the performance degradation caused by the capacity limitations of small models, but also leverages the rich information in the teacher network to enhance the generalization capabilities of the student network, ultimately enabling efficient and lightweight cross-device status diagnosis.
[0077] Preferably, the computing device fixes the model parameters of the teacher network when calculating the KL divergence loss, so that it does not participate in gradient updates. The teacher network is typically a pre-trained model or one that achieves better performance than the student network during training. Freezing its parameters ensures that the "knowledge" provided by the teacher network is stable when calculating the KL divergence loss. If the teacher parameters are constantly changing, the learning objectives of the student network will also be constantly changing, which is not conducive to the student network effectively learning the knowledge and characteristics contained in the teacher network.
[0078] While training the student network, continuing to train and update the teacher network's parameters can cause the teacher network to overfit to the current training dataset, making the knowledge learned by the teacher network too specialized and losing its general description of the overall data distribution. Freezing the teacher's parameters can prevent this from happening, ensuring that the teacher network can continue to provide general and representative knowledge to the student network.
[0079] Step S4: The computing device adjusts the parameters of the student network according to the real-time working condition parameters of the device to be diagnosed, and obtains a working condition adaptive diagnosis model.
[0080] The device to be diagnosed is a target device to be diagnosed, and specifically, may refer to one or more devices of a certain type.
[0081] Among them, the student network after parameter adjustment by the computing equipment is the final working condition adaptive diagnosis model.
[0082] In step S3, the student network has learned common fault signatures from the teacher network (trained on simulation data) through knowledge distillation, but its parameters are still optimized primarily for the simulation data. At this point, the student network can leverage real-time operating data from the device to be diagnosed (such as load factor and ambient temperature). Through domain adversarial training and an attention mechanism, the student network's feature extraction weights are dynamically adjusted, enabling better generalization in real-world environments. Finally, the adjusted student network retains its original structure, but its parameters are optimized for the target device, resulting in a directly deployable operating-condition adaptive diagnosis model.
[0083] Preferably, the computing device can construct a domain adversarial training module, which includes a feature extractor, a domain discriminator and a gradient reversal layer. The parameters of the feature extractor are the parameters of the student network. The gradient reversal layer is used to reverse the gradient of the domain discriminator when backpropagating the feature extractor; the features of the simulated fault data and the features of the real-time operating condition parameters are extracted by the feature extractor; the domain difference loss between the simulated fault data and the real-time operating condition parameters is calculated by the domain discriminator based on the features of the simulated fault data and the features of the real-time operating condition parameters; wherein, the feature extractor performs domain invariant feature learning through the gradient reversal layer, and the loss function is:
[0084] L total =L classify -λ·L domain ;
[0085] λ=2·(1+e -10p ) -1 -1;
[0086] Among them, L total is the total loss, L classify is the classification loss, L domain is the domain discrimination loss, λ is the domain adaptation loss weight, and p is the training progress ratio; after the preset convergence conditions are reached, the parameters of the student network are adjusted based on the parameters of the feature extractor.
[0087] The core goal of domain adversarial training is to eliminate the distribution differences between simulated fault data (source domain) and the real-time operating parameters of the equipment to be diagnosed (target domain) through dynamic domain adaptation technology, so that the student network after knowledge distillation can adapt to the actual operating conditions of the target equipment.
[0088] Specifically, the feature extractor extracts features from the input data, such as magnetic field distortion and partial discharge pulses, by reusing the feature extraction layer of the student network, mapping the simulated data and real data into the same feature space, providing basic features for domain alignment.
[0089] The domain discriminator D is used to determine whether the feature comes from the simulation domain or the real domain. Through adversarial training, it forces the feature extractor to generate domain-invariant features, directly eliminating the distribution difference between the source domain and the target domain.
[0090] The gradient reversal layer reverses the gradient of the domain discriminator during back propagation, allowing the feature extractor to "cheat" the domain discriminator, so that the optimization direction of the feature extractor is opposite to that of the domain discriminator, ensuring that the feature space is insensitive to domain differences.
[0091] Specifically, the domain adversarial training module also includes a label classifier, which is used to classify faults based on the output of the feature extractor, and iterates by calculating the classification loss to ensure that the student network still maintains the ability to discriminate fault types while adapting to domain differences.
[0092] The domain discriminator initially accurately distinguished features from the source and target domains, indicating a significant distribution difference. The gradient reversal layer then reversely optimized the feature extractor, causing it to generate features that were indistinguishable between the two types of data. After training, the domain discriminator's accuracy dropped to 45%-55%, close to random guessing, demonstrating that the distribution difference had been eliminated. At this point, the computing device could determine that convergence conditions had been met.
[0093] Optionally, the convergence condition also includes the total loss value being less than a preset threshold, or reaching a maximum number of iterations.
[0094] Preferably, the feature extractor further includes a multi-head self-attention module; the computing device can calculate the feature channel weight of the feature according to the real-time working condition parameter through the multi-head self-attention module; the calculation formula of the feature channel weight includes:
[0095]
[0096] Among them, α i is the weight of the i-th feature channel, h i is the i-th feature channel vector, W q is the working condition parameter mapping matrix, d is the number of feature channels; based on the features of the simulated fault data and the measured fault data, as well as the feature channel weights, the domain difference loss of the simulated fault data and the measured fault data is calculated.
[0097] Through the multi-head self-attention mechanism, the weights of feature channels are dynamically adjusted according to real-time operating parameters (such as load rate and temperature), so that under complex working conditions, the fault features most relevant to the current operating status are enhanced and irrelevant or interfering features are suppressed.
[0098] In the embodiment of the present application, by constructing a finite element model of the power equipment to generate simulated fault data, the reliance on real fault data can be reduced, solving the problem of scarce fault sample data. Then, the teacher network is trained based on the simulated fault data, and the student network is trained based on the measured fault data. The knowledge transfer from the teacher network to the student network is achieved through KL divergence, which improves the generalization ability of the student network. The parameters of the student network are then adjusted according to the real-time operating parameters of the device to be diagnosed. The resulting operating condition adaptive diagnosis model can more accurately diagnose the device to be diagnosed. Using the embodiment of the present application, high-precision state migration diagnosis across devices can be achieved.
[0099] The above describes the method part provided by the embodiment of the present application. The following describes the system part provided by the embodiment of the present application.
[0100] See also Figure 2 , Figure 2 A schematic diagram of the structure of a power equipment state transition diagnosis system based on knowledge distillation provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the system 20 includes:
[0101] A generating module 201 is used to construct a finite element model of the power equipment and generate simulated fault data based on the finite element model;
[0102] A training module 202 is configured to train a teacher network based on the simulated fault data and a student network based on measured fault data, wherein the measured fault data is data obtained by measuring real faulty power equipment; wherein the teacher network and the student network are preset deep learning networks;
[0103] A migration module 203 is used to migrate the knowledge of the teacher network to the student network through KL divergence;
[0104] The adjustment module 204 is used to adjust the parameters of the student network according to the real-time working condition parameters of the device to be diagnosed, so as to obtain a working condition adaptive diagnosis model.
[0105] Preferably, the generation module 201 is specifically used to input the physical parameters of the power equipment in a fault state into the finite element model to obtain a fault feature map output by the finite element model; perform data enhancement on the fault feature map through an adversarial generative network to obtain the simulated fault data; wherein the generator and the discriminator of the adversarial generative network are adversarially trained based on the physical field distribution error of the power equipment.
[0106] Preferably, the loss function of the adversarial generative network includes a physical constraint term and a distribution alignment term; wherein the physical constraint term is the L2 norm error between the generated data of the adversarial generative network and the solution of the Maxwell equation; and the distribution alignment term is the Wasserstein distance between the generated data and the real fault data.
[0107] Preferably, the teacher network uses a deep residual network to process the simulated fault data to output a high-dimensional feature vector and a fault probability distribution of the simulated fault data; the student network uses a convolutional network to process the measured fault data; and the student network aligns the feature distribution of the teacher network by minimizing the KL divergence.
[0108] Preferably, the calculation formula of the KL divergence includes:
[0109] L KD =T 2 *Σ[Pteacher (x)*log(P teacher (x) / P student (x))];
[0110] Among them, L KD represents KL divergence, T is the temperature coefficient, P teacher (x), P student (x) are the probability distributions of the teacher network and the student network output based on the input x.
[0111] Preferably, the adjustment module 204 is specifically used to construct a domain adversarial training module, which includes a feature extractor, a domain discriminator and a gradient reversal layer. The parameters of the feature extractor are the parameters of the student network. The gradient reversal layer is used to reverse the gradient of the domain discriminator when backpropagating the feature extractor; the features of the simulated fault data and the features of the real-time operating parameters are extracted by the feature extractor; the domain difference loss between the simulated fault data and the real-time operating parameters is calculated by the domain discriminator based on the features of the simulated fault data and the features of the real-time operating parameters; wherein, the feature extractor performs domain invariant feature learning through the gradient reversal layer, and the loss function is:
[0112] L total =L classify -λ·L domain ;
[0113] λ=2·(1+e -10p ) -1 -1;
[0114] Among them, L total is the total loss, L classify is the classification loss, L domain is the domain discrimination loss, λ is the domain adaptation loss weight, and p is the training progress ratio; after the preset convergence conditions are reached, the parameters of the student network are adjusted based on the parameters of the feature extractor.
[0115] Preferably, the feature extractor further includes a multi-head self-attention module; the adjustment module 204 is specifically configured to calculate the feature channel weight of the feature according to the real-time working condition parameters through the multi-head self-attention module; the calculation formula of the feature channel weight includes:
[0116]
[0117] Among them, α i is the weight of the i-th feature channel, h i is the i-th feature channel vector, W qis the working condition parameter mapping matrix, d is the number of feature channels; based on the features of the simulated fault data and the measured fault data, as well as the feature channel weights, the domain difference loss of the simulated fault data and the measured fault data is calculated.
[0118] The power equipment state transition diagnosis system based on knowledge distillation provided in the embodiment of the present application can be understood by referring to the corresponding content of the aforementioned method embodiment part, and will not be repeated here.
[0119] like Figure 3 As shown, Figure 3 A possible logical structure diagram of a computing device provided in an embodiment of the present application. The computing device 300 includes: a processor 301, a communication interface 302, a memory 303, and a bus 304. The processor 301, the communication interface 302, and the memory 303 are interconnected via the bus 304. In the embodiment of the present application, the processor 301 is used to control and manage the actions of the computing device 300. For example, the processor 301 is used to execute Figure 1 The steps in the embodiments and / or other processes for the technology described herein. The communication interface 302 is used to support the computing device 300 to communicate. The memory 303 is used to store program codes and data of the computing device 300.
[0120] Among them, the processor 301 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a transistor logic device, a hardware component or any combination thereof. It can implement or execute the various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 304 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0121] In another embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes instructions. When the instructions are executed on a computer, the computer executes the above-mentioned Figure 1 The method described in the embodiment.
[0122] Those skilled in the art will appreciate that the units of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition of each example has been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0123] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0124] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0125] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0126] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0127] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk or optical disk, and other media that can store program codes.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.
Claims
1. A power equipment state transition diagnosis method based on knowledge distillation, characterized in that: include: Step S1: constructing a finite element model of the power equipment and generating simulated fault data based on the finite element model; Step S2: training a teacher network based on the simulated fault data and training a student network based on the measured fault data, wherein the measured fault data is data obtained by measuring real faulty power equipment; wherein the teacher network and the student network are preset deep learning networks; Step S3: Transferring the knowledge of the teacher network to the student network through KL divergence; Step S4: According to the real-time working condition parameters of the device to be diagnosed, the parameters of the student network are adjusted to obtain a working condition adaptive diagnosis model.
2. The method according to claim 1, characterized in that The step S1 comprises: Inputting physical parameters of the power equipment in a fault state into the finite element model to obtain a fault characteristic map output by the finite element model; The fault feature map is data enhanced by a generative adversarial network to obtain the simulated fault data; wherein the generator and the discriminator of the generative adversarial network are adversarially trained based on the physical field distribution error of the power equipment.
3. The method according to claim 2, characterized in that The loss function of the adversarial generative network includes a physical constraint term and a distribution alignment term; wherein the physical constraint term is the L2 norm error between the generated data of the adversarial generative network and the solution of the Maxwell equation; and the distribution alignment term is the Wasserstein distance between the generated data and the real fault data.
4. The method according to claim 1, wherein The teacher network uses a deep residual network to process the simulated fault data, and is used to output a high-dimensional feature vector and a fault probability distribution of the simulated fault data; the student network uses a convolutional network to process the measured fault data; and the student network aligns the feature distribution of the teacher network by minimizing the KL divergence.
5. The method according to claim 1, wherein The calculation formula of the KL divergence includes: L KD =T 2 *Σ[P teacher (x)*log(P teacher (x) / P student (x))]; Among them, L KD represents KL divergence, T is the temperature coefficient, P teacher (x), P student (x) are the probability distributions of the teacher network and the student network output based on the input x.
6. The method according to any one of claims 1 to 5, characterized in that The step S4 comprises: Constructing a domain adversarial training module, the domain adversarial training module including a feature extractor, a domain discriminator, and a gradient reversal layer, wherein the parameters of the feature extractor are the parameters of the student network, and the gradient reversal layer is used to reverse the gradient of the domain discriminator during back-propagation training of the feature extractor; Extracting the features of the simulated fault data and the features of the real-time operating condition parameters by the feature extractor; Calculating, by the domain discriminator, domain difference losses between the simulated fault data and the real-time operating condition parameters based on the characteristics of the simulated fault data and the characteristics of the real-time operating condition parameters; The feature extractor performs domain-invariant feature learning through the gradient reversal layer, and the corresponding loss function includes: L total =L classify -λ·L domain ; λ=2·(1+e -10p ) -1 -1; Among them, L total is the total loss, L classify is the classification loss, L domain is the domain discrimination loss, λ is the domain adaptation loss weight, and p is the training progress ratio; After the preset convergence condition is reached, the parameters of the student network are adjusted based on the parameters of the feature extractor.
7. The method according to claim 6, characterized in that The feature extractor further includes a multi-head self-attention module; the calculating the domain difference loss between the simulated fault data and the measured fault data based on the features of the simulated fault data and the features of the measured fault data includes: The multi-head self-attention module calculates the feature channel weight of the feature according to the real-time working condition parameters; the calculation formula of the feature channel weight includes: Among them, α i is the weight of the i-th feature channel, h i is the i-th feature channel vector, W q is the working condition parameter mapping matrix, d is the number of characteristic channels; Based on the features of the simulated fault data and the features of the measured fault data, and the feature channel weights, domain difference losses of the simulated fault data and the measured fault data are calculated.
8. A power equipment state transition diagnosis system based on knowledge distillation, characterized in that: The method according to any one of claims 1 to 7, wherein the system comprises: A generation module, configured to construct a finite element model of the power equipment and generate simulated fault data based on the finite element model; A training module, configured to train a teacher network based on the simulated fault data and train a student network based on measured fault data, wherein the measured fault data is data obtained by measuring real faulty power equipment; wherein the teacher network and the student network are preset deep learning networks; A migration module, configured to transfer the knowledge of the teacher network to the student network through KL divergence; The adjustment module is used to adjust the parameters of the student network according to the real-time working condition parameters of the equipment to be diagnosed, so as to obtain a working condition adaptive diagnosis model.
9. A computing device, characterized in that include: Memory, used to store programs; A processor, configured to load the program to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Knowledge distillation-based 1D-CNN online partial discharge identification method
CN121410481A
Method for predicting multi-working-condition transient physical field of fluid equipment based on residual distillation
CN121659803A
Full-life-cycle intelligent early warning and maintenance method for permanent magnet pump driven by large model
CN121961537A