Electrical equipment aging prediction model knowledge distillation method and aging prediction method
By using a dynamic masking mechanism and loss calculation to optimize the feature alignment of the student model in electrical equipment aging prediction, the problem of insufficient feature alignment in traditional knowledge distillation methods is solved, the prediction accuracy and robustness on edge devices are improved, and real-time aging state recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-13
AI Technical Summary
Existing knowledge distillation methods for predicting the aging of electrical equipment suffer from problems such as insufficient feature alignment leading to decreased student model performance, insufficient robustness, and poor generalization ability. In particular, they are difficult to achieve efficient and accurate aging state prediction on edge devices.
By inputting the feature supervision signal of the pre-trained aging prediction teacher model into the initial student model, a second student feature is generated using a dynamic masking mechanism. The encoder of the student model is updated by combining feature alignment loss, mask generation loss and task loss to ensure feature dimension alignment and robustness. The student classifier is removed after training, while the encoder and projector are retained to reduce computational complexity.
It significantly improves the robustness and classification accuracy of the student model on edge devices, achieves millisecond-level aging state recognition, and supports real-time operation and maintenance decisions for smart grids.
Smart Images

Figure CN121660158A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid technology, and in particular to a knowledge distillation method and aging prediction method for electrical equipment aging prediction models. Background Technology
[0002] In the context of smart grids, predicting the aging status of electrical equipment is crucial for ensuring the safe and stable operation of the power grid. With the high proportion of new energy sources connected to the grid and the increasing electronic nature of the power grid, critical equipment is subjected to frequent switching, harmonic impacts, and temperature cycling, significantly accelerating insulation aging. Sudden failures can trigger cascading trips and large-scale power outages, causing substantial economic losses and social impact. Traditional periodic maintenance relies on manual inspections and offline testing, making it difficult to detect latent faults in a timely manner, resulting in both over-maintenance and under-maintenance issues. Therefore, real-time aging status prediction based on online monitoring data has become a core requirement of intelligent operation and maintenance. However, field-deployed monitoring terminals can only provide low-frequency time-series data such as current, voltage, and temperature, and the computing power and storage resources of edge embedded chips are limited, making it impossible to directly run large-scale deep learning models. Knowledge distillation technology, by transferring knowledge from large teacher models to lightweight student models, can compress models while maintaining prediction accuracy, meeting the requirements of millisecond-level inference and deployment in harsh environments at the edge, thereby promoting the practical application and widespread adoption of electrical equipment aging prediction models.
[0003] Existing knowledge distillation methods typically train student models through feature alignment and soft label transfer, but these methods have significant drawbacks. First, the student model requires retraining the classifier, leading to quadratic "feature-classification" errors, particularly unstable on aging boundary samples, impacting classification reliability. Second, the intermediate feature alignment process relies on manually designed loss functions and hyperparameters, which are cumbersome to adjust and poorly adaptable to different device types or sampling rates, lacking generalization ability. Furthermore, traditional feature distillation methods require the student model to mimic teacher features pixel-by-pixel, easily resulting in overfitting and insufficient robustness, making it difficult to handle noise and missing data in real-world scenarios. These shortcomings limit the effective application of knowledge distillation in complex industrial scenarios. Summary of the Invention
[0004] This invention provides a knowledge distillation method for electrical equipment aging prediction models and a corresponding aging prediction method, which solves the problem of student model performance degradation caused by insufficient feature alignment in traditional knowledge distillation, and significantly improves the robustness and classification accuracy of student models on edge devices.
[0005] In a first aspect, embodiments of the present invention provide a knowledge distillation method for an electrical equipment aging prediction model, comprising:
[0006] The pre-acquired sample data is input into the pre-trained aging prediction teacher model to obtain teacher features; wherein, the sample data includes historical time-series signals of electrical equipment; the historical time-series signals include current time-series signals, voltage time-series signals and temperature time-series signals;
[0007] The sample data is input into an initial aging prediction student model that is only initialized. Within the initial aging prediction student model, a first student feature and a second student feature are generated sequentially. Based on the teacher feature, the first student feature, and the second student feature, loss is calculated to obtain feature alignment loss and mask generation loss. The feature dimension of the first student feature is the same as that of the teacher feature. The second student feature is generated through a preset dynamic masking mechanism.
[0008] The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction in order to obtain the task loss.
[0009] Based on the feature alignment loss, mask generation loss, and task loss, the encoder parameters of the initial aging prediction student model are updated, and after the parameter update is completed, the classifier of the initial aging prediction student model is removed to obtain the aging prediction student model.
[0010] This invention provides high-quality feature supervision signals to student models using a teacher model trained on large-scale data with strong feature extraction capabilities, ensuring a reliable starting point for knowledge transfer. By randomly initializing the student model, it avoids inheriting biases from the pre-trained model. Through a distillation process, it learns the teacher's knowledge from scratch, enhancing the model's generalization ability. Student features with lower feature dimensions than the teacher's features are aligned with the teacher's feature dimensions for easy direct comparison and loss calculation. Next, a second student feature is generated through a dynamic masking mechanism, forcing the student to recover the complete teacher features from partial features, improving feature representation ability and robustness. Finally, feature alignment loss is calculated to ensure that student features are aligned with the teacher's features. Numerically close to teacher features, reducing feature space differences, and calculating mask generation loss quantifies the student model's feature reconstruction capability; by reusing teacher classifier weights, it avoids secondary errors introduced by student classifier training, improving classification accuracy and stability; by combining three losses to update the encoder, it balances feature learning and classification tasks through multi-task learning, preventing overfitting and ensuring that the student model learns teacher features while maintaining task performance; finally, the student classifier is removed, leaving only the encoder and projector in the student model to reduce model complexity and computational overhead, facilitating edge deployment and preparing the student model to directly use the teacher classifier during the inference phase, ensuring the accuracy of the output results. Compared with existing technologies, this invention solves the problem of student model performance degradation caused by insufficient feature alignment in traditional knowledge distillation, significantly improving the robustness and classification accuracy of the student model on edge devices.
[0011] Furthermore, the initial aging prediction student model includes an encoder, a feature projector, a mask block, a generator block, and a classifier;
[0012] Within the initial aging prediction student model, the first student feature and the second student feature are generated sequentially, specifically as follows:
[0013] The sample data is input into the encoder to generate initial student features, and the initial student features are mapped to the same feature dimension as the teacher features by the feature projector to obtain the first student features.
[0014] The masking block is used to mask the first student feature to obtain the masked student feature, and the generation block is used to restore and reconstruct the masked student feature to obtain the second student feature.
[0015] This invention uses a feature projector to align student and teacher features in terms of dimensions; it introduces a dynamic masking mechanism to randomly occlude some features to simulate missing data or noise; and it generates blocks to reconstruct complete features, thereby training students' ability to infer global features from local information and improving the model's robustness.
[0016] Furthermore, the convolution kernels of the mask block and the generator block share parameters through weight binding.
[0017] The embodiments of the present invention reduce the number of model parameters and lower computation and storage requirements by sharing parameters, while enhancing feature consistency. In particular, the projector and the generator block share some parameters, which can ensure the coherence of feature mapping and avoid feature ambiguity.
[0018] Furthermore, the step of performing a masking operation on the first student feature using the mask block to obtain the masked student feature specifically involves:
[0019] For each training iteration, a cosine annealing strategy is used to adjust the dynamic mask ratio and determine the current dynamic mask ratio.
[0020] A mask matrix is randomly generated based on the current dynamic mask ratio using a preset spatiotemporal joint mechanism; wherein, the spatiotemporal joint mechanism is to independently use a Bernoulli variable to determine whether to mask each position of the first student feature; each position is determined jointly from the time step dimension and the channel dimension;
[0021] Based on the mask matrix, a masking operation is performed on the first student features to obtain the masked student features.
[0022] This invention adjusts the dynamic mask ratio using a cosine annealing strategy, allowing training to progress from high mask ratio (difficult) to low mask ratio (easy), balancing training difficulty and preventing premature overfitting. By randomly masking at the time step and channel dimensions, it simulates local data gaps in real-world scenarios, forcing students to learn the local-to-global correlation of time-series data and enhancing robustness to incomplete data.
[0023] Furthermore, the loss calculation based on the teacher features, the first student features, and the second student features yields the feature alignment loss and the mask generation loss, specifically as follows:
[0024] Calculate the Euclidean distance between the first student feature and the teacher feature, and determine the Euclidean distance as the feature alignment loss;
[0025] By using a preset attention mechanism, the mask loss calculation weights are determined, and the mask generation loss between the second student feature and the teacher feature is calculated based on the mask loss calculation weights.
[0026] This invention uses Euclidean distance to prevent feature scale differences from affecting optimization, ensuring that features are numerically closely aligned; it uses an attention mechanism for mask generation loss, with weights focused on regions that are difficult to recover, making loss calculation more targeted and improving reconstruction quality, and encouraging students to focus on key feature regions.
[0027] Furthermore, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction, thereby obtaining the task loss, specifically as follows:
[0028] The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction, obtains the classification prediction result, and calculates the loss based on the classification prediction result to obtain the task loss.
[0029] This invention improves training efficiency while ensuring the reliability of classification results by transferring teacher classifier weights during the training phase, avoiding errors introduced by student classifier training; and maintains the performance orientation of the student model on the final task by calculating the difference between the prediction and the true label based on the task loss.
[0030] Furthermore, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction and obtain the classification prediction result, specifically as follows:
[0031] The classifier weights of the aging prediction teacher model are transferred to the classifier of the initial aging prediction student model to obtain the student classifier; wherein, the student classifier is only used in the model training process.
[0032] The first student feature is input into the student classifier to generate a classification prediction result.
[0033] This invention ensures compatibility between student features and the teacher classifier by temporarily using teacher classifier parameters during the training phase. Furthermore, the student classifier is used only as an intermediate carrier during training and is removed during inference, reducing deployment complexity and ultimately achieving seamless reuse of the classifier. This ensures consistency between training and inference while simplifying the model structure.
[0034] Secondly, embodiments of the present invention provide a method for predicting the aging of electrical equipment, including:
[0035] Real-time signal data of electrical equipment in the target smart grid is continuously collected and input into a pre-trained aging prediction student model so that the aging prediction student model outputs projection features; wherein, the real-time signal data includes real-time current signal, real-time voltage signal and real-time temperature signal; the projection features are features with the same feature dimension as the output features of the pre-trained aging prediction teacher model;
[0036] The projection features are input into the classifier of the aging prediction teacher model to obtain the aging prediction result, so as to realize the real-time prediction of the aging status of electrical equipment in the target smart grid.
[0037] The aging prediction student model is obtained through the knowledge distillation method for electrical equipment aging prediction models described above.
[0038] This invention, tailored to smart grid scenarios, continuously monitors the status of electrical equipment to provide real-time input for an aging prediction student model. The student model is lightweight, retaining only the encoder and projector, resulting in fewer parameters and high computational efficiency, making it suitable for edge devices. The teacher classifier is then reused during the inference phase, avoiding additional classifier training and improving prediction accuracy and speed. This invention applies the distilled student model to actual prediction tasks, achieving millisecond-level aging status identification and providing real-time, reliable operation and maintenance decision support for smart grids.
[0039] Furthermore, the step of inputting the real-time signal data into the pre-trained aging prediction student model, so that the aging prediction student model outputs projection features, specifically involves:
[0040] The encoder of the aging prediction student model generates initial signal features, and the feature projector of the aging prediction student model maps the initial signal features to the same feature dimension as the output features of the pre-trained aging prediction teacher model to obtain projected features.
[0041] This invention uses a feature projector to map the initial signal features to the same dimension as the teacher's features, ensuring that the student features are consistent with the input dimension of the teacher classifier, avoiding inference errors, thereby ensuring the smoothness of the subsequent prediction process and improving system compatibility and stability.
[0042] Furthermore, the step of inputting the projected features into the classifier of the aging prediction teacher model to obtain the aging prediction result specifically involves:
[0043] By truncating the gradient, the projected features are input into the classifier of the aging prediction teacher model to generate aging prediction results through forward propagation.
[0044] This invention employs gradient truncation to prevent backpropagation from updating the teacher classifier parameters during inference, keeping them fixed and ensuring the consistency and reliability of predictions. This maintains the stability of the teacher classifier, avoids parameter drift during inference, and improves system robustness.
[0045] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0046] Figure 1 A schematic diagram of a knowledge distillation method for an electrical equipment aging prediction model provided in an embodiment of the present invention;
[0047] Figure 2 A flowchart illustrating a knowledge distillation method for predicting the aging of electrical equipment, as exemplified by an embodiment of the present invention;
[0048] Figure 3 This is a schematic diagram of an electrical equipment aging prediction method provided in an embodiment of the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example 1:
[0051] like Figure 1 As shown, a knowledge distillation method for an electrical equipment aging prediction model provided by an embodiment of the present invention includes the following steps:
[0052] S101, Input the pre-acquired sample data into the pre-trained aging prediction teacher model to obtain teacher features; wherein, the sample data includes historical time-series signals of electrical equipment; the historical time-series signals include current time-series signals, voltage time-series signals and temperature time-series signals;
[0053] In one specific embodiment, in a smart grid scenario, current, voltage, and temperature time-series signals are continuously collected from electrical equipment in three aging states (normal, slight aging, and severe aging) at a frequency of Δt = 1 min to construct an original dataset; samples are truncated using a sliding window method with a window length L = 24 and a step size s = 1 to obtain an input tensor with a shape of 24×3, corresponding to the label y∈{0,1,2}.
[0054] In one specific embodiment, the teacher network adopts a 1-D-VGG16 architecture, containing 13 one-dimensional convolutional layers and 3 fully connected layers, with the last layer outputting K=3 dimensions. The teacher network is semi-supervised trained on a large-scale dataset until convergence. The pre-training process uses a conventional loss function to ensure the teacher network has strong feature extraction and classification capabilities. After pre-training, the parameters of the teacher network are fixed. The encoder of the teacher model is F... t Output feature tensor Where C t H represents the number of channels. t ×W t Spatial resolution; Classifier C t It consists of a global average pooling layer and a fully connected layer, and outputs a class probability vector p. t ∈R K K is the total number of categories;
[0055] S102, the sample data is input into the initial aging prediction student model that is only initialized, so that the first student feature and the second student feature are generated sequentially within the initial aging prediction student model, and loss is calculated based on the teacher feature, the first student feature and the second student feature to obtain feature alignment loss and mask generation loss; wherein, the feature dimension of the first student feature is the same as that of the teacher feature; the second student feature is generated through a preset dynamic masking mechanism.
[0056] In this embodiment, the initial aging prediction student model includes an encoder, a feature projector, a mask block, a generator block, and a classifier. Within the initial aging prediction student model, the generation of a first student feature and a second student feature is performed sequentially. Specifically, the sample data is input into the encoder to generate initial student features, and the initial student features are mapped to the same feature dimension as the teacher features using the feature projector to obtain the first student feature. The first student feature is then masked using the mask block to obtain a masked student feature, and the masked student feature is then restored and reconstructed using the generator block to obtain the second student feature.
[0057] In one specific embodiment, the student network selects a 1-D-VGG16 with half the number of channels and randomly initializes its parameters. The encoder of the student model is F. s Output features satisfy H s =H t W s =W t , where C s C represents the number of channels for student characteristics. t H represents the number of channels for teacher characteristics. s H represents the height of the student feature map. t For the height of the teacher feature map, W s W represents the width of the student feature map. t The width of the teacher feature map.
[0058] In one specific embodiment, a feature projector is designed to match the difference in feature dimensions between the teacher model and the student model. The feature projector consists of a 1×1 convolutional layer and a batch normalization (BN) layer, as shown in the formula:
[0059]
[0060] Among them, C s and C t The number of channels corresponding to the student model and the teacher model, This is the output of the student after passing through the feature projector. During initialization, the projector parameters are randomly set.
[0061] In one specific embodiment, the feature map size is restored by zero-padding, the mask features are input into the feature projector, and mapped to the same spatial dimension as the teacher features, ensuring tensor dimension consistency:
[0062]
[0063] Where O represents a zero-filled matrix, and ⊙ represents element-wise multiplication. The student characteristics represented by the mask,
[0064] The output of the feature projector represents the mapped student features, M∈{0,1} H×W Let represent the Bernoulli mask matrix, where 0 indicates that the time step is masked and 1 indicates that it is preserved.
[0065] In this embodiment, the step of masking the first student feature using the mask block to obtain the masked student feature specifically involves: for each training iteration, a cosine annealing strategy is used to adjust the dynamic mask ratio to determine the current dynamic mask ratio; a mask matrix is randomly generated based on the current dynamic mask ratio using a preset spatiotemporal joint mechanism; wherein the spatiotemporal joint mechanism independently uses a Bernoulli variable to determine whether to mask each position of the first student feature; each position is determined jointly from the time step dimension and the channel dimension; and a mask operation is performed on the first student feature according to the mask matrix to obtain the masked student feature.
[0066] In one specific embodiment, the mask generation module (including mask blocks and generation blocks) is used to randomly generate a mask matrix to mask the student's feature map. Map pixels already contain information about neighboring pixels to some extent. Therefore, by using a subset of pixels to recover the complete feature map, this method aims to generate the teacher's features using the masked features of the student.
[0067] In one specific embodiment, the generation block uses two Conv1 d_3×1 blocks sandwiched with ReLU and introduces residual connections to reconstruct the complete temporal features of the teacher from the occluded temporal segments.
[0068]
[0069] in, To generate the reconstructed student features, The mask is a dynamic mask for the student features after masking. Conv1 d represents a 1-D convolution with a kernel size of 3×1. A ReLU activation function is introduced in the middle, and the residual is connected back to the input.
[0070] In one specific embodiment, the dynamic mask ratio τ is adjusted using a cosine annealing strategy, as shown in the formula:
[0071]
[0072] Where, τ current τ is the dynamic mask ratio for the current training epoch. final τ is the final value of the mask ratio. init =0.8 is the initial value for the mask ratio, and total_epochs is the set total number of training epochs.
[0073] In one specific embodiment, the mask generation module employs a spatiotemporal joint random mask, that is, independently sampling Bernoulli variables for each location of the 24×1 feature map:
[0074] M i ~Bernoulli(1-τ current ), i = 1, ..., 24,
[0075] Among them, M i Let be the mask variable at the i-th time step (0 masks, 1 preserves), and Bernoulli(p) denotes the Bernoulli distribution.
[0076] Furthermore, the mask pattern is refreshed randomly in each training batch. Because random pixels are used in each iteration, all pixels will be used throughout the entire training process, which means that the features will be more robust and their representational power will be improved.
[0077] In this embodiment, the convolution kernel of the mask block and the generator block share parameters through weight binding.
[0078] In one specific embodiment, the sharing of convolution kernel parameters between the projector P (feature projector) and the generated block g is achieved through weight binding, specifically:
[0079] The weight matrix of the 1×1 convolutional layer of the projector With the first 3×1 convolutional layer of the generated block Shared channels, the formula is:
[0080]
[0081] Where: indicates taking all elements of this dimension, i.e., the center position of the generated block. The convolution kernel inherits the mapping parameters of the projector, enhancing feature consistency.
[0082] In this embodiment, the step of calculating the loss based on the teacher feature, the first student feature, and the second student feature to obtain the feature alignment loss and the mask generation loss specifically involves: calculating the Euclidean distance between the first student feature and the teacher feature, and determining the Euclidean distance as the feature alignment loss; determining the mask loss calculation weight through a preset attention mechanism, and calculating the mask generation loss between the second student feature and the teacher feature based on the mask loss calculation weight.
[0083] In one specific embodiment, the feature alignment loss L align Normalized L2 (Euclidean distance) distance is used to prevent feature scale differences from affecting the optimization:
[0084]
[0085] in, For the features of students after projection, f t Characteristics of teachers.
[0086] In one specific embodiment, the mask generation loss L gen Introducing attention weights Focus on areas that are difficult to recover:
[0087]
[0088] Where c represents the feature channel, and i and j represent the time step index (i) and spatial position (j, which is 1 here due to the one-dimensional convolution). This indicates the value of the student's reconstructed feature at that location. The value of the teacher characteristic at this position, A i,j The attention coefficient is at position (i,j). The larger the error, the greater the weight, forcing the network to prioritize learning the "difficult to reconstruct" regions.
[0089] Preferably, hierarchical adaptive weights γ1 are designed for the different depth feature maps of the teacher and student encoders:
[0090]
[0091] Where μ = 0.1 is the depth coefficient, and d l This represents the depth of the l-th layer. The deeper the layer (i.e., the index d), the higher the depth. l The larger the value of γ1, the more dominant the deep semantic features are in the distillation loss; the weight of shallow detail features automatically decreases.
[0092] Furthermore, the total distillation loss can be written as a four-level weighted average:
[0093]
[0094] S103, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction in order to obtain the task loss;
[0095] In this embodiment, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction in order to obtain the task loss. Specifically, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction, obtains the classification prediction result, and calculates the loss based on the classification prediction result to obtain the task loss.
[0096] In this embodiment, the initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction and obtain classification prediction results. Specifically, the classifier weights of the aging prediction teacher model are transferred to the classifier of the initial aging prediction student model to obtain a student classifier. The student classifier is only used in the model training process. The first student feature is input into the student classifier to generate classification prediction results.
[0097] In one specific embodiment, the optimization objective of the student model is to minimize the cross-entropy loss (i.e., task loss) between the predicted result and the true label:
[0098]
[0099] Among them, y k The actual label value. These are predicted values.
[0100] Simultaneously, the encoder parameters are updated using distillation loss. The optimizer uses AdamW, and the learning rate is adjusted according to a certain number of rounds. After training for a certain number of rounds, the rate is reduced by a factor to ensure that the task loss dominates in the later stages of training, thus avoiding overfitting to teacher features.
[0101] S104, based on the feature alignment loss, mask generation loss and task loss, update the parameters of the encoder of the initial aging prediction student model, and after the parameter update is completed, remove the classifier of the initial aging prediction student model to obtain the aging prediction student model.
[0102] In one specific embodiment, the overall loss function is a weighted sum of the losses from each of the above components, with weighting coefficients α and β used to balance the contributions of each component:
[0103]
[0104] Preferably, the performance of the student model is verified every ten cycles during training. If the verification loss does not decrease for three consecutive cycles, an early stopping mechanism is triggered to save the optimal parameters.
[0105] To better illustrate the working principle and steps of this method, see [link / reference]. Figure 2 One example, Figure 2 The flowchart illustrates a knowledge distillation method for an electrical equipment aging prediction model, as exemplified by an embodiment of the present invention.
[0106] This invention provides high-quality feature supervision signals to student models using a teacher model trained on large-scale data with strong feature extraction capabilities, ensuring a reliable starting point for knowledge transfer. By randomly initializing the student model, it avoids inheriting biases from the pre-trained model. Through a distillation process, it learns the teacher's knowledge from scratch, enhancing the model's generalization ability. Student features with lower feature dimensions than the teacher's features are aligned with the teacher's feature dimensions for easy direct comparison and loss calculation. Next, a second student feature is generated through a dynamic masking mechanism, forcing the student to recover the complete teacher features from partial features, improving feature representation ability and robustness. Finally, feature alignment loss is calculated to ensure that student features are aligned with the teacher's features. Numerically close to teacher features, reducing feature space differences, and calculating mask generation loss quantifies the student model's feature reconstruction capability; by reusing teacher classifier weights, it avoids secondary errors introduced by student classifier training, improving classification accuracy and stability; by combining three losses to update the encoder, it balances feature learning and classification tasks through multi-task learning, preventing overfitting and ensuring that the student model learns teacher features while maintaining task performance; finally, the student classifier is removed, leaving only the encoder and projector in the student model to reduce model complexity and computational overhead, facilitating edge deployment and preparing the student model to directly use the teacher classifier during the inference phase, ensuring the accuracy of the output results. Compared with existing technologies, this invention solves the problem of student model performance degradation caused by insufficient feature alignment in traditional knowledge distillation, significantly improving the robustness and classification accuracy of the student model on edge devices.
[0107] Example 2:
[0108] like Figure 3 As shown, an embodiment of the present invention provides a method for predicting the aging of electrical equipment, comprising the following steps:
[0109] S201, continuously collect real-time signal data of electrical equipment in the target smart grid, and input the real-time signal data into a pre-trained aging prediction student model so that the aging prediction student model outputs projection features; wherein, the real-time signal data includes real-time current signal, real-time voltage signal and real-time temperature signal; the projection features are features with the same feature dimension as the output features of the pre-trained aging prediction teacher model;
[0110] In this embodiment, the step of inputting the real-time signal data into the pre-trained aging prediction student model so that the aging prediction student model outputs projected features specifically involves: generating initial signal features through the encoder of the aging prediction student model, and mapping the initial signal features to the same feature dimension as the output features of the pre-trained aging prediction teacher model through the feature projector of the aging prediction student model to obtain projected features.
[0111] S202, the projection features are input into the classifier of the aging prediction teacher model to obtain the aging prediction result, so as to realize the real-time prediction of the aging status of electrical equipment in the target smart grid.
[0112] The aging prediction student model is obtained through the knowledge distillation method for electrical equipment aging prediction models described above.
[0113] In this embodiment, the step of inputting the projected features into the classifier of the aging prediction teacher model to obtain the aging prediction result specifically involves: inputting the projected features into the classifier of the aging prediction teacher model through gradient truncation, so as to generate the aging prediction result through forward propagation.
[0114] In one specific embodiment, during inference, student features are passed through the projector and then truncated by gradient sg(·) before being fed into the teacher classifier:
[0115]
[0116] Where sg(·) means that the parameters of the tensor are not updated during gradient backpropagation; the three-class probability is obtained, and the argmax output aging state {normal, slightly aged, severely aged} is taken; the parameters of the entire classifier are frozen, and backpropagation only updates the student encoder and projector.
[0117] This invention, tailored to smart grid scenarios, continuously monitors the status of electrical equipment to provide real-time input for an aging prediction student model. The student model is lightweight, retaining only the encoder and projector, resulting in fewer parameters and high computational efficiency, making it suitable for edge devices. The teacher classifier is then reused during the inference phase, avoiding additional classifier training and improving prediction accuracy and speed. This invention applies the distilled student model to actual prediction tasks, achieving millisecond-level aging status identification and providing real-time, reliable operation and maintenance decision support for smart grids.
[0118] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0119] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A knowledge distillation method for predicting the aging of electrical equipment, characterized in that, include: The pre-acquired sample data is input into the pre-trained aging prediction teacher model to obtain teacher features; wherein, the sample data includes historical time-series signals of electrical equipment; the historical time-series signals include current time-series signals, voltage time-series signals and temperature time-series signals; The sample data is input into an initial aging prediction student model that is only initialized. Within the initial aging prediction student model, a first student feature and a second student feature are generated sequentially. Based on the teacher feature, the first student feature, and the second student feature, loss is calculated to obtain feature alignment loss and mask generation loss. The feature dimension of the first student feature is the same as that of the teacher feature. The second student feature is generated through a preset dynamic masking mechanism. The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction in order to obtain the task loss. Based on the feature alignment loss, mask generation loss, and task loss, the encoder parameters of the initial aging prediction student model are updated, and after the parameter update is completed, the classifier of the initial aging prediction student model is removed to obtain the aging prediction student model.
2. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 1, characterized in that, The initial aging prediction student model includes an encoder, a feature projector, a mask block, a generator block, and a classifier; Within the initial aging prediction student model, the first student feature and the second student feature are generated sequentially, specifically as follows: The sample data is input into the encoder to generate initial student features, and the initial student features are mapped to the same feature dimension as the teacher features by the feature projector to obtain the first student features. The masking block is used to mask the first student feature to obtain the masked student feature, and the generation block is used to restore and reconstruct the masked student feature to obtain the second student feature.
3. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 2, characterized in that, The convolution kernels of the mask block and the generator block share parameters through weight binding.
4. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 2, characterized in that, The process of masking the first student feature using the mask block to obtain the masked student feature specifically involves: For each training iteration, a cosine annealing strategy is used to adjust the dynamic mask ratio and determine the current dynamic mask ratio. A mask matrix is randomly generated based on the current dynamic mask ratio using a preset spatiotemporal joint mechanism; wherein, the spatiotemporal joint mechanism is to independently use a Bernoulli variable to determine whether to mask each position of the first student feature; each position is determined jointly from the time step dimension and the channel dimension; Based on the mask matrix, a masking operation is performed on the first student features to obtain the masked student features.
5. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 1, characterized in that, The loss calculation is performed based on the teacher features, the first student features, and the second student features to obtain the feature alignment loss and the mask generation loss, specifically as follows: Calculate the Euclidean distance between the first student feature and the teacher feature, and determine the Euclidean distance as the feature alignment loss; By using a preset attention mechanism, the mask loss calculation weights are determined, and the mask generation loss between the second student feature and the teacher feature is calculated based on the mask loss calculation weights.
6. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 1, characterized in that, The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction and obtain the task loss, specifically as follows: The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction, obtains the classification prediction result, and calculates the loss based on the classification prediction result to obtain the task loss.
7. The knowledge distillation method for predicting the aging of electrical equipment as described in claim 6, characterized in that, The initial aging prediction student model reuses the classifier weights of the aging prediction teacher model to perform classification prediction and obtain the classification prediction result, specifically as follows: The classifier weights of the aging prediction teacher model are transferred to the classifier of the initial aging prediction student model to obtain the student classifier; wherein, the student classifier is only used in the model training process. The first student feature is input into the student classifier to generate a classification prediction result.
8. A method for predicting the aging of electrical equipment, characterized in that, include: Real-time signal data of electrical equipment in the target smart grid is continuously collected and input into a pre-trained aging prediction student model so that the aging prediction student model outputs projection features; wherein, the real-time signal data includes real-time current signal, real-time voltage signal and real-time temperature signal; the projection features are features with the same feature dimension as the output features of the pre-trained aging prediction teacher model; The projection features are input into the classifier of the aging prediction teacher model to obtain the aging prediction result, so as to realize the real-time prediction of the aging status of electrical equipment in the target smart grid. The aging prediction student model is obtained through the knowledge distillation method for electrical equipment aging prediction models as described in claims 1 to 7.
9. The method for predicting the aging of electrical equipment as described in claim 8, characterized in that, The step of inputting the real-time signal data into a pre-trained aging prediction student model, so that the aging prediction student model outputs projection features, specifically involves: The encoder of the aging prediction student model generates initial signal features, and the feature projector of the aging prediction student model maps the initial signal features to the same feature dimension as the output features of the pre-trained aging prediction teacher model to obtain projected features.
10. The method for predicting the aging of electrical equipment as described in claim 8, characterized in that, The step of inputting the projected features into the classifier of the aging prediction teacher model to obtain the aging prediction result is as follows: By truncating the gradient, the projected features are input into the classifier of the aging prediction teacher model to generate aging prediction results through forward propagation.