Image classification and recognition system and method based on AI training model

By introducing technical means such as data monitoring module and dynamic tag correction module in the image classification recognition system, the problem of the lack of dynamic training mechanism in the existing system is solved, real-time update of the model and high-accuracy recognition are achieved.

CN120164010APending Publication Date: 2025-06-17PINGDINGSHAN PINGGAO-YASKAWA SWITCH APP CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510154510.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing image classification recognition system based on AI training models lacks a dynamic training mechanism and cannot update and learn new equipment failure characteristics in real time, resulting in a decrease in recognition accuracy and the inability to detect new defects and failures in time.

Method used

The data monitoring module is used to evaluate the distribution offset between the input data and the training data in real time through the joint distribution offset detection algorithm of maximum mean difference and KL divergence, trigger the model update request, and realize dynamic model update and optimization through the dynamic label correction module, model distillation module, model tuning module and model version control module.

Benefits of technology

It realizes the model to quickly adapt to the new environment, maintains high recognition accuracy, reduces computing resource consumption and model update cycle, and enhances the flexibility and response speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005269143800000101
    Figure BDA0005269143800000101
  • Figure BDA0005269143800000131
    Figure BDA0005269143800000131
  • Figure BDA0005269143800000141
    Figure BDA0005269143800000141
Patent Text Reader

Abstract

The invention discloses an image classification and recognition system and method based on an AI training model, relates to the technical field of image classification and recognition, and solves the problem that an existing method lacks dynamic adaptability. Comprising a data monitoring module, a dynamic label correction module, a model distillation module, a model tuning module and a model version control module. The data monitoring module evaluates input data distribution offset in real time by adopting a maximum mean value difference and KL divergence algorithm, and triggers model updating; the dynamic label correction module realizes automatic label correction through clustering active learning and semantic segmentation; the model distillation module uses adaptive incremental learning and a knowledge distillation mechanism to reduce computing resource consumption; the model adjusting and optimizing module adopts an elastic weight integration method to maintain the stability of the model; performance fluctuation is prevented through the model version control module; according to the invention, the flexibility and response speed of electric power facility inspection are greatly improved, and the defect detection precision and the long-term adaptability of the model are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image classification and recognition, and more particularly to an image classification and recognition system and method based on an AI training model. Background Art

[0002] With the rapid development of artificial intelligence (AI) and big data technologies, deep learning has become the core technology driving the transformation in the field of image classification and recognition. Especially in the industrial field, such as intelligent inspection of power systems, equipment status monitoring, and fault diagnosis, AI-based image classification and recognition technologies are widely used. This technology improves the inspection efficiency and fault discovery ability through automated and intelligent means, and has become an indispensable key link in the construction of modern smart grids.

[0003] Intelligent inspection and fault detection in power systems mainly rely on an image classification and recognition system trained by AI. This system obtains a large amount of image data through drones, inspection robots, and fixed monitoring devices. These data are preprocessed for images, features are extracted, and then combined with deep learning models (such as CNN, ResNet) and traditional image processing algorithms (such as edge detection, feature point matching) to achieve intelligent recognition and detection of problems such as power equipment defects and faults. Existing technologies build large-scale labeled data sets and train CNN models to learn image features. For example, in the literature "Research on Image Recognition Method for Power Equipment Defects Based on Deep Learning", a large number of normal and faulty images of power equipment are collected, and transfer learning is used to fine-tune the pre-trained model to improve the recognition accuracy. However, existing methods mainly train models based on fixed data sets. Although the image classification and recognition accuracy has been improved to a certain extent, due to relying on fixed data sets, there is a lack of a dynamic training mechanism. During the operation of power equipment, affected by factors such as environment and time, its appearance and fault characteristics are constantly changing. For example, for a long-term operating transmission line, the insulator may develop new deterioration characteristics due to aging, and internal components of the transformer may generate new fault modes due to temperature and humidity changes. However, existing models cannot update and learn these new features in real time. Once there are changes not covered by the training data, the recognition accuracy will drop significantly. Continuously collecting new data and retraining the model is costly and time-consuming, and it is difficult to meet the real-time requirements. This problem of lack of dynamic adaptability makes the system perform poorly when facing equipment dynamic changes, unable to detect newly emerging defects and faults in a timely and accurate manner, posing a potential threat to the safe and stable operation of the power system. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention discloses an image classification and recognition system and method based on an AI training model, aiming to solve the problem of the lack of a dynamic training mechanism to cope with equipment changes in the existing image classification and recognition system based on an AI training model.

[0005] To achieve the above technical effects, the present invention adopts the following technical solutions:

[0006] An image classification and recognition system based on an AI training model, comprising: a data monitoring module, a dynamic label correction module, a model distillation module, a model tuning module, and a model version control module;

[0007] The data monitoring module is used to evaluate the distribution shift degree between the input data and the training data in real time by using a joint distribution shift detection algorithm based on maximum mean discrepancy and KL divergence; if it is detected that the distribution deviation between the input data and the training data is greater than a preset threshold, a model update request is triggered, a distribution shift alarm is generated, and the input data is transmitted to the dynamic label correction module; if it is detected that the distribution deviation between the input data and the training data is less than or equal to the preset threshold, the model operation is maintained;

[0008] The dynamic label correction module is used to automatically correct the labels of the input data through active learning and semantic segmentation algorithms based on clustering, and output high-confidence samples and low-confidence samples; the dynamic label correction module is also used to transmit the high-confidence samples to the model distillation module, submit the low-confidence samples to an expert system for manual verification through a feedback mechanism, and output the verified samples to the model distillation module;

[0009] The model distillation module is used to perform local model fine-tuning on new data samples by using an adaptive incremental learning algorithm, and transfer the new model knowledge to the main model in combination with a knowledge distillation mechanism;

[0010] The model tuning module is used to dynamically constrain the model weight parameters of the model fine-tuned by the model distillation module by using an elastic weight consolidation method. When the model tuning module detects a conflict between new data features and the existing model, it elastically adjusts the conflicting parameters and pushes the adjusted model parameters to the model version control module in real time; when it detects that the data features are consistent with the existing model, it maintains the current parameter configuration;

[0011] The model version control module is used to monitor the model performance indicators in real time through a multi-version model management mechanism; and when it detects that the performance of the new model is better than that of the old model, it automatically deploys the new model; when the performance of the new model decreases or fluctuates beyond a preset threshold, it automatically rolls back to the previous stable version through a version rollback strategy.

[0012] As a further technical solution of the present invention, in the data monitoring module, the method for the joint distribution shift detection algorithm of the maximum mean difference and KL divergence to real-time evaluate the shift degree between the input data and the training data distribution is as follows: Based on the input image feature vector obtained by feature extraction and the training data reference distribution model with time dynamic characteristics generated by time series analysis in combination with the operation historical data of power equipment, the distance between the input data and the training data in the feature space is measured by an adaptive kernel function to obtain the maximum mean difference distance; meanwhile, the difference between the probability distributions of the input data and the training data is calculated by the KL divergence algorithm to obtain the KL divergence value; finally, the weights are obtained by the analytic hierarchy process according to the severity and occurrence probability of different types of faults of power equipment, and the maximum mean difference distance and the KL divergence algorithm are weighted and fused to obtain a comprehensive shift index.

[0013] As a further technical solution of the present invention, in the adaptive kernel function, the bandwidth is automatically adjusted according to the local density of the input feature vector through an adaptive bandwidth selection strategy; the working method of the adaptive bandwidth selection strategy is as follows:

[0014] r1. Based on the input feature vector, calculate the local data density by the kernel density estimation method;

[0015] r2. Initially determine the bandwidth by the Silverman empirical method;

[0016] r3. Fine-tune the bandwidth within a preset range by the cross-validation method;

[0017] The KL divergence algorithm corrects the noise and uncertainty in the data through an entropy correction factor to reflect the true difference in the data distribution.

[0018] As a further technical solution of the present invention, in the dynamic label correction module, the working method for automatically correcting the labels of the input data is as follows: First, calculate the attention weights for the pixels in different regions through a semantic segmentation network integrating a multi-head self-attention mechanism to generate a segmented image with semantic information; then, adopt a density clustering algorithm combined with a domain knowledge graph to integrate the information of the knowledge graph into the density and neighborhood calculations, perform a clustering operation on the image data points, and correct the labels according to the clustering results to obtain the image data with corrected labels; based on the image data with corrected labels, adopt a confidence evaluation model based on Bayesian deep learning, estimate the uncertainty of the prediction distribution through Monte Carlo dropout, and generate a confidence score for each sample.

[0019] As a further technical solution of the present invention, the adaptive incremental learning algorithm uses a gradient accumulation mechanism to divide newly added data samples into multiple small batches. In the processing of each small batch, the model parameters are updated according to the sample gradient information. At the same time, through a dynamic learning rate adjustment strategy, the learning rate is adaptively adjusted according to the model convergence situation and sample difficulty, and a locally fine-tuned model is output; when the adaptive incremental learning algorithm processes newly added data samples, the meta-learner, according to the performance feedback of different batches of samples, learns the best update step size and direction for each sample, and compares the new samples with the old samples through a contrast learning algorithm, calculates the loss using the loss function of contrast learning, and optimizes the performance of the locally fine-tuned model.

[0020] As a further technical solution of the present invention, in the model distillation module, the method for the knowledge distillation mechanism to transfer the knowledge of the new model to the main model is as follows: First, the prediction result of the new model is used as the soft label through the soft label distillation method, and the KL divergence is used to calculate the soft loss between the soft label and the prediction result of the main model, and the soft loss is used as the optimization target; then, in the processing of the intermediate layer feature representations of the new model and the main model, the feature mapping distillation mechanism is used to construct a feature mapping function to map the intermediate layer features of the new model to the corresponding intermediate layer of the main model, and the cosine similarity is used to measure the feature difference to optimize the parameters of the main model; in addition, the attention-guided distillation method is used to calculate the attention weights of different regions of the new model and the main model, and the main model is guided to learn the attention mechanism of the new model according to the weight difference to optimize the performance of the main model.

[0021] As a further technical solution of the present invention, the elastic weight integration method regards the weight parameters as random variables and uses a Bayesian network to establish the probability relationship between the weight parameters and the comprehensive feature vector; the Bayesian network calculates the posterior probability distribution of different weight parameters through Bayesian inference and outputs the optimal range of weights; in the process of weight adjustment, the conflict situation between the current data features and the existing model is considered, and an elastic adjustment range is assigned to each weight parameter according to the expectation and variance of the posterior probability distribution; when the elastic weight integration method detects a conflict between the new data features and the existing model, based on the gradient change of the model on the new data, the gradient perception conflict detection algorithm marks the parameters with abnormal gradients as conflict parameters, and adjusts the conflict parameters according to the optimal weight range obtained by Bayesian inference; during the adjustment process, the elastic weight integration method adopts an adaptive learning rate decay strategy to dynamically adjust the learning rate according to the severity of the conflict and the model performance feedback.

[0022] As a further technical solution of the present invention, the model performance metrics monitored in real time by the model version control module include accuracy, recall, precision, F1 score, average precision threshold, and inference time; when the change range of the time series data of the model performance metrics exceeds the preset threshold range within a fixed time window, it is considered that performance fluctuations occur.

[0023] As a further technical solution of the present invention, an image classification and recognition method based on an AI training model includes the following steps:

[0024] Step S1: Based on the input power equipment image data, use the joint distribution shift detection algorithm based on maximum mean discrepancy and KL divergence to compare the distribution of the input data with the training data in real time;

[0025] Step S2: Calculate the degree of difference in the distribution of the input data and the training data in the feature space by calculating the maximum mean discrepancy, and at the same time use KL divergence to measure the difference in their probability distributions;

[0026] Step S3: If the calculated distribution deviation is greater than the preset threshold, trigger a model update request, generate a distribution shift alarm, and transmit the data to the dynamic label correction module; if the deviation is less than or equal to the preset threshold, maintain the model parameters to run in the current state;

[0027] Step S4: Based on the output data of Step S3, perform clustering analysis according to the feature similarity of the data through the clustering-based active learning algorithm to obtain the potential structures and patterns in the data; at the same time, perform semantic understanding and segmentation on different regions in the image through the semantic segmentation algorithm to identify different parts of the power equipment and its components in the image;

[0028] Step S5: Automatically correct the labels of the data, and based on the label correction results, screen out high-confidence samples and low-confidence samples through confidence evaluation;

[0029] Step S6: Directly transmit the high-confidence samples to the model distillation module; submit the low-confidence samples to the expert system for verification through the feedback mechanism, and output the verified samples to the model distillation module;

[0030] Step S7: Receive the output data of Step S6 as new data samples; adopt an incremental learning strategy to perform local model fine-tuning; and combine the knowledge distillation mechanism to transfer the locally fine-tuned model for the new data to the main model in the form of soft labels;

[0031] Step S8: Adopt the elastic weight consolidation method to dynamically constrain the model weight parameters of the model output in Step S7;

[0032] Step S9: Push the adjusted model parameters to the model version control module in real time; if it is detected that the new data features are the same as the existing model, keep the current parameter configuration unchanged;

[0033] Step S10: Monitor the performance metrics of the model in real time through the multi-version model management mechanism, including accuracy, recall rate, F1 value, average precision threshold, and inference time; if the performance metrics of the new model are better than those of the old model, apply the new model to the actual image classification and recognition task; if the performance of the new model decreases or fluctuates beyond the preset threshold, automatically roll back the model to the previous stable version through the version rollback strategy.

[0034] Based on the above technical solutions, the positive and beneficial effects of the present invention are:

[0035] 1. Through the maximum mean discrepancy (MMD) and KL divergence joint distribution shift detection algorithm adopted by the data monitoring module, the real-time evaluation of the distribution deviation between the input data and the training data is realized. Once a distribution change exceeding the preset threshold is detected, the system will immediately trigger a model update request, ensuring that the model can quickly adapt to the new environment and maintain a high recognition accuracy.

[0036] 2. The dynamic label correction module combines clustering-based active learning and semantic segmentation algorithms to realize the automatic label correction of the input data, generating high-confidence and low-confidence samples. High-confidence samples are directly used for subsequent processing, while low-confidence samples are submitted to the expert system for manual verification through the feedback mechanism, greatly improving the annotation efficiency and accuracy and reducing the need for manual intervention. At the same time, the model distillation module uses an adaptive incremental learning algorithm to fine-tune the local model for new data samples, and transfers the new model knowledge to the main model through the knowledge distillation mechanism. This dynamic training mechanism reduces the consumption of computing resources, shortens the model update cycle, and enhances the flexibility and response speed of the system.

[0037] 3. The model tuning module uses the elastic weight consolidation method to dynamically constrain the model weight parameters output by the model distillation module. When the new data features conflict with the existing model, the conflicting parameters can be elastically adjusted, rather than adopting a fixed weight adjustment strategy as in the traditional method. This dynamic constraint and elastic adjustment enable the model to better adapt to the complex feature changes in the operation of power equipment, ensure that the model parameter adjustment is more targeted and flexible, improve the adaptability of the model to new features, and avoid performance degradation caused by improper parameter adjustment. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, where:

[0039] Figure 1 It is an architecture diagram of an image classification and recognition system based on an AI training model of the present invention;

[0040] Figure 2 It is a step flow chart of an image classification and recognition method based on an AI training model of the present invention. Detailed implementation manners

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] As Figure 1 shown, an image classification and recognition system based on an AI training model includes: a data monitoring module, a dynamic label correction module, a model distillation module, a model tuning module, and a model version control module; wherein, the output end of the data monitoring module is connected to the input end of the dynamic label correction module; the output end of the dynamic label correction module is connected to the input end of the model distillation module; the output end of the model distillation module is connected to the input end of the model tuning module; the output end of the model tuning module is connected to the input end of the model version control module;

[0043] As Figure 2 shown, an image classification and recognition method based on an AI training model includes the following steps:

[0044] Step S1: Based on the input power equipment image data, use the joint distribution shift detection algorithm based on maximum mean discrepancy and KL divergence to compare the distribution of the input data with the training data in real time;

[0045] Step S2: Calculate the maximum mean discrepancy to measure the degree of difference in the distribution of the input data and the training data in the feature space, and at the same time use KL divergence to measure the difference in their probability distributions;

[0046] Step S3: If the calculated distribution deviation is greater than the preset threshold, trigger a model update request, generate a distribution shift alarm, and transmit the data to the dynamic label correction module; if the deviation is less than or equal to the preset threshold, maintain the model parameters to run in the current state;

[0047] Step S4: Based on the output data of Step S3, perform clustering analysis according to the feature similarity of the data through an active learning algorithm based on clustering to obtain the potential structures and patterns in the data; at the same time, perform semantic understanding and segmentation on different regions in the image through a semantic segmentation algorithm to identify different parts of the power equipment and its components in the image;

[0048] Step S5: Automatically correct the labels of the data, and according to the label correction results, screen out high-confidence samples and low-confidence samples through confidence evaluation;

[0049] Step S6: Directly transmit the high-confidence samples to the model distillation module; submit the low-confidence samples to the expert system for verification through a feedback mechanism, and the verified samples are output to the model distillation module;

[0050] Step S7: Receive the output data of Step S6 as new data samples; adopt an incremental learning strategy to perform local model fine-tuning; and combine the knowledge distillation mechanism to transfer the locally fine-tuned model for the new data to the main model in the form of soft labels;

[0051] Step S8: Adopt the elastic weight consolidation method to dynamically constrain the model weight parameters of the model output in Step S7;

[0052] Step S9: Push the adjusted model parameters to the model version control module in real time; if it is detected that the new data features are consistent with the existing model, maintain the current parameter configuration unchanged;

[0053] Step S10: Real-time monitor the performance metrics of the model through a multi-version model management mechanism, including accuracy, recall rate, F1 value, average precision threshold, and inference time; if the performance metrics of the new model are better than those of the old model, apply the new model to the actual image classification and recognition task; if the performance of the new model decreases or fluctuates beyond the preset threshold, automatically roll back the model to the previous stable version through the version rollback strategy.

[0054] Among them, the data monitoring module is used to real-time evaluate the deviation degree of the input data from the training data distribution by using a joint distribution shift detection algorithm based on maximum mean discrepancy and KL divergence; if it is detected that the distribution deviation between the input data and the training data is greater than the preset threshold, trigger a model update request, generate a distribution shift alarm and transmit the input data to the dynamic label correction module; if it is detected that the distribution deviation between the input data and the training data is less than or equal to the preset threshold, maintain the model operation;

[0055] In the data monitoring module, the method for the joint distribution shift detection algorithm of the maximum mean difference and KL divergence to real-time evaluate the shift degree between the input data and the training data distribution is as follows: Based on the input image feature vectors obtained by feature extraction and the training data reference distribution model with time dynamic characteristics generated by time series analysis in combination with the operation history data of power equipment, the distance between the input data and the training data in the feature space is measured by an adaptive kernel function to obtain the maximum mean difference distance; at the same time, the difference between the probability distributions of the input data and the training data is calculated by the KL divergence algorithm to obtain the KL divergence value; finally, the weights are obtained by the analytic hierarchy process according to the severity and occurrence probability of different types of faults of power equipment, and the maximum mean difference distance and the KL divergence algorithm are weighted and fused to obtain a comprehensive shift index. The adaptive kernel function automatically adjusts the bandwidth according to the local density of the input feature vectors through an adaptive bandwidth selection strategy; the working method of the adaptive bandwidth selection strategy is:

[0056] r1. Calculate the local data density based on the input feature vectors by the kernel density estimation method;

[0057] r2. Initially determine the bandwidth using the Silverman empirical method;

[0058] r3. Fine-tune the bandwidth within a preset range using the cross-validation method;

[0059] The KL divergence algorithm corrects the noise and uncertainty in the data through an entropy correction factor to reflect the true difference in the data distribution.

[0060] In a specific embodiment, the kernel density estimation method uses the Parzen window method. According to the position of the input feature vectors in the feature space, a kernel function, such as a Gaussian kernel function, is set for each feature vector. By summing and normalizing the kernel functions of each feature vector, the density estimation of the entire feature space is obtained. This method can smoothly estimate the distribution of feature vectors and can reflect the local density of the data, providing density information for bandwidth selection.

[0061] The Silverman empirical method calculates the bandwidth through formula (1) according to the standard deviation and sample size of the data:

[0062] h = 0.9×min(σ, IQR / 1.34)×n -1 / 5 (1)

[0063] where σ is the sample standard deviation, IQR is the interquartile range, and n is the sample size. This formula comprehensively considers the dispersion degree of the data and the sample scale, providing a reasonable initial bandwidth estimate for subsequent fine-tuning and ensuring that the selection of the bandwidth has a certain statistical basis.

[0064] The cross-validation method divides the input data into k equal parts. For each possible bandwidth, one of them is used as the validation set and the rest as the training set in turn, and the maximum mean difference under the bandwidth is calculated. By traversing different bandwidth and cross-validation combinations, the bandwidth that minimizes the maximum mean difference on the validation set is selected to ensure that the selected bandwidth can achieve good performance on different data subsets, thereby optimizing the bandwidth selection.

[0065] In the calculation of the maximum mean difference, the adaptive kernel function makes the shape and coverage of the kernel function change dynamically according to the local density of the data according to the above strategy, ensuring that a smaller bandwidth is used for dense areas and a larger bandwidth is used for sparse areas, so as to more accurately measure the distance between the input data and the training data in the feature space. Adaptive bandwidth selection avoids the problems caused by the use of fixed bandwidth. For the image feature space of power equipment with uneven data distribution, the range of the kernel function can be adjusted according to the density of local features, so that the maximum mean difference calculation can better reflect the real data similarity and improve the sensitivity of distribution shift detection. The reliability and stability of bandwidth selection are ensured by cross-validation method, so that the optimal bandwidth can be found in different data partitioning situations, and the robustness of distribution shift detection is enhanced, especially for complex power equipment feature distribution, which can more accurately reflect the local structural differences of data.

[0066] The KL divergence algorithm quantifies the difference between two distributions by calculating the expectation of the logarithmic difference between them. For the characteristic distribution of power equipment images, it evaluates the degree of deviation of the input data relative to the training data from the perspective of probability distribution, reflecting the change of the data in the overall probability distribution. The entropy correction factor corrects the KL divergence according to the size of the entropy value. In the power equipment data, the noise may come from the interference of the image acquisition equipment, environmental factors, etc. The entropy correction factor adjusts the KL divergence result according to the entropy change caused by the noise, so that the final difference measurement is closer to the actual distribution difference of the data. In implementation, the KL divergence algorithm comprehensively evaluates the data distribution difference from a probability perspective, which can reflect the changes in the probability distribution level of the input data and the training data, and can provide certain measurements for different types of data offsets (such as the offset of the overall data or the offset of the local data). The entropy correction factor can effectively deal with the noise and uncertainty in the data, avoid the erroneous KL divergence calculation results caused by noise, make the evaluation of the distribution difference more resistant to interference, and ensure the accuracy and reliability of the distribution difference measurement.

[0067] The analytic hierarchy process first constructs a hierarchical structure model, taking the severity and occurrence probability of different types of faults in power equipment as factors at different levels. By constructing a judgment matrix, the importance of different fault types is compared. For example, the severity of the short-circuit fault of a transformer and the surface pollution fault of an insulator are quantitatively compared. Using the eigenvalue and eigenvector calculation method, the maximum eigenvalue and the corresponding eigenvector of the judgment matrix are solved, and the eigenvector is normalized to obtain the weights of each factor. The maximum mean difference distance and the KL divergence value are multiplied by their respective weights and added together to obtain a comprehensive deviation index.

[0068] When making a decision, according to the calculated comprehensive deviation index, if it is greater than the preset threshold, it indicates that the distribution deviation of the input data has exceeded the acceptable range, and there may be an abnormal situation in the power equipment, triggering a model update request, transmitting the data to the dynamic label correction module, and generating a distribution deviation alarm for subsequent processing and update operations. If the comprehensive deviation index is less than or equal to the preset threshold, it is considered that the input data is still within the normal distribution range, maintaining the normal operation of the model and ensuring the stability of the system. The analytic hierarchy process assigns reasonable weights to different measurement indicators according to the actual situation of power equipment faults, avoiding the unscientific nature of simple averaging or random weight assignment, and making the comprehensive deviation index better reflect the actual impact of data deviation on the state judgment of power equipment.

[0069] In implementation, this decision-making mechanism based on the comprehensive deviation index can reasonably judge when the model needs to be updated, ensure that the model update is necessary and targeted, avoid unnecessary update operations, and can detect abnormal data in a timely manner, ensuring the performance of the system and the accuracy of power equipment fault detection. Through these more refined technical principles and their synergistic effects, the data monitoring module can more accurately monitor the distribution deviation of power equipment image data, providing a solid foundation for the efficient operation of the entire AI training model-based image classification and recognition system and the improvement of fault detection performance, enabling the system to better adapt to the dynamic changes and different operating states of power equipment.

[0070] The dynamic label correction module is used to automatically correct the input data through clustering-based active learning and semantic segmentation algorithms, and output high-confidence samples and low-confidence samples. The dynamic label correction module is also used to transmit the high-confidence samples to the model distillation module, submit the low-confidence samples to the expert system for manual verification through a feedback mechanism, and output the verified samples to the model distillation module. In the dynamic label correction module, the working method for automatically correcting the input data is as follows: First, a semantic segmentation network integrating the multi-head self-attention mechanism calculates the attention weights for pixels in different regions to generate a segmentation image with semantic information. Then, a density clustering algorithm combined with a domain knowledge graph integrates the information of the knowledge graph into density and neighborhood calculations, performs clustering operations on the image data points, and corrects the labels according to the clustering results to obtain the image data with corrected labels. Based on the image data with corrected labels, a confidence evaluation model based on Bayesian deep learning is used to estimate the uncertainty of the prediction distribution through Monte Carlo dropout and generate the confidence score for each sample.

[0071] Among them, in the semantic segmentation network, the multi-head self-attention mechanism MHSA divides the feature representation of the input image into multiple heads, and each head independently calculates the attention weights. For each head, by calculating the similarity between the query Q vector, the key K vector, and the value V vector, the dot product attention formula is used to determine the attention degree of each position T to other positions:

[0072]

[0073] where d K is the dimension of the key vector. In this way, different heads can focus on different aspects of the image. For example, one head may focus on the contour information of the power equipment, and another head may focus on the texture details on the equipment. Through the parallel processing of multiple heads, the network can capture the importance of different regions in the image from multiple perspectives. Then, the results of each head are concatenated together and, after linear transformation, the final attention-weighted feature representation is obtained.

[0074] This attention-weighted feature representation is input into subsequent segmentation network layers, usually the decoder part, such as transposed convolutional layers or upsampling layers, to restore the low-resolution feature map to the original image size and generate a segmented image with semantic information. The transposed convolutional layer performs a convolutional operation on the feature map through a learned convolutional kernel and inserts zero elements to enlarge the size, gradually restoring to the original image resolution. The upsampling layer enlarges the feature map through methods such as interpolation while maintaining semantic information, and finally generates a segmented image with semantic information for each pixel, which can clearly divide different objects or parts of objects in the image. For example, different components of electrical equipment can be segmented. In implementation, MHSA can more effectively capture long-range dependencies in the image, overcoming the limitation that traditional convolutional neural networks (CNNs) can only process local information. For complex electrical equipment images, there may be long-range semantic connections between different components, such as a fault area and a normal area that are far apart on the device. MHSA can better associate this information, thus generating a more accurate semantic segmentation result. The setting of multiple heads allows the network to simultaneously focus on features at multiple levels, making the segmentation result more detailed and accurate, which helps with subsequent label correction because accurate segmentation provides a more reliable basis for subsequent label assignment, avoiding label errors caused by inaccurate segmentation.

[0075] A density clustering algorithm combined with a domain knowledge graph constructs a domain knowledge graph, which contains various prior knowledge about electrical equipment, such as the relationships between different equipment components, the connections between fault types and features, etc. For example, the knowledge graph will record the connection relationships between the components of a transformer (such as windings, iron cores, insulation layers) and the possible faults (such as short circuits, overheating) and their characteristic manifestations in the image. Based on a density clustering algorithm (such as DBSCAN), the information in the knowledge graph is incorporated into it. When calculating the density of data points, not only the spatial distance of the data points in the feature space is considered, but the density calculation is also adjusted according to the relationships in the knowledge graph. For example, if two data points are relatively close in the feature space and the components they represent in the knowledge graph belong to closely related components of the same equipment, their density weights will increase. When determining the neighborhood, it is also judged according to the knowledge graph information. If the components represented by two data points usually have faults simultaneously or are in similar operating states, their neighborhood range will be expanded.

[0076] For the clustering operation of image data points, based on the updated density and neighborhood information, data points with similar density and neighborhood relationships are grouped into one class. When the number of data points within the neighborhood of a data point exceeds the set minimum number of points and the distance between them meets the conditions, it will be marked as a core point, and the clustering is gradually expanded. According to the clustering results, data points of the same class are assigned the same label, and the original label is corrected according to the definition of this class in the knowledge graph to make the label more in line with the actual situation of power equipment. For example, if a cluster contains points representing a transformer winding and shows high-temperature characteristics, its label can be corrected to "overheated winding".

[0077] In implementation, the use of the domain knowledge graph can guide the clustering process, avoid blind clustering relying only on data features, and make the clustering results more meaningful. For power equipment, different fault and normal states may have similar data features, but combining the knowledge graph can distinguish them and improve the accuracy of label correction. It can handle complex image data, especially for the case of sparse or uneven data. Combining the information of the knowledge graph can make the clustering more reasonable, thereby improving the reliability of label correction, reducing incorrect labels caused by fuzzy data features, and better serving subsequent model training and optimization.

[0078] Bayesian deep learning combines deep learning and Bayesian methods. In this confidence evaluation model, the weights of the network are no longer fixed values but probability distributions. During the training process, Monte Carlo Dropout is used for uncertainty estimation. In the inference stage, Dropout is enabled multiple times for forward propagation to obtain multiple different prediction results. For example, for the class prediction of an image of a power equipment, multiple possible classes and their probability distributions will be obtained. Through the Monte Carlo method, statistical analysis is performed on these multiple prediction results, and the variance of these results is used to represent the uncertainty.

[0079] For each sample, a confidence score is generated based on the magnitude of uncertainty. The lower the uncertainty, the higher the confidence score, indicating that the model is more certain about the prediction of that sample; conversely, high uncertainty results in a low confidence score. For example, for a clear image of a power equipment in normal state, the multiple prediction results may be concentrated in the "normal" category, and its confidence score is high; while for a partially occluded or blurred image, the prediction results will be more dispersed, with high uncertainty and a low confidence score. In implementation, it is possible to quantify the uncertainty of the model's prediction for each sample, which is very important for identifying samples that are difficult to classify. For the image classification of power equipment, some images may be difficult to judge due to various factors (such as noise, occlusion, and lighting changes). This model can provide a basis for subsequent processing, screen out low-confidence samples for manual verification, and improve the overall reliability of the system. Additionally, overconfident predictions can be avoided, making the system's judgment of samples more cautious and scientific. For newly emerging abnormal situations or unseen image types, it can more accurately identify them as low-confidence samples instead of wrongly assigning high confidence, which helps the robustness and adaptability of the model in practical applications.

[0080] The model distillation module is used to perform local model fine-tuning on new data samples using an adaptive incremental learning algorithm and transfer the new model knowledge to the main model in combination with the knowledge distillation mechanism. The adaptive incremental learning algorithm uses a gradient accumulation mechanism to divide the new data samples into multiple small batches. In the processing of each small batch, the model parameters are updated according to the sample gradient information, and at the same time, the learning rate is adaptively adjusted according to the model convergence situation and sample difficulty through a dynamic learning rate adjustment strategy, and the locally fine-tuned model is output. When the adaptive incremental learning algorithm processes new data samples, the meta-learner learns the best update step size and direction for each sample according to the performance feedback of different batches of samples, and compares the new samples with the old samples through a contrastive learning algorithm, calculates the loss using the loss function of contrastive learning, and optimizes the performance of the locally fine-tuned model. In the model distillation module, the working method of the knowledge distillation mechanism for transferring the new model knowledge to the main model is as follows: First, the prediction result of the new model is used as a soft label through the soft label distillation method, and the soft loss between the soft label and the prediction result of the main model is calculated using the KL divergence, with the soft loss as the optimization target; then, in the processing of the intermediate layer feature representations of the new model and the main model, the feature mapping distillation mechanism is used to construct a feature mapping function to map the intermediate layer features of the new model to the corresponding intermediate layer of the main model, measure the feature difference using the cosine similarity, and optimize the main model parameters; in addition, the attention-guided distillation method is used to calculate the attention weights of different regions of the new model and the main model, and the main model is guided to learn the attention mechanism of the new model according to the weight difference to optimize the performance of the main model.

[0081] Among them, when faced with newly added data samples, they are divided into multiple small batches because in deep learning, processing a large amount of data at once may exceed the memory limit of the hardware. In each small batch, the gradient information of the samples is calculated. The gradient reflects the rate of change of the model parameters with respect to the error of the samples in this batch. The specific formula for calculating the gradient is:

[0082]

[0083] where L is the loss function and θ are the model parameters. By accumulating the gradients of multiple small batches, the problems of vanishing gradients or exploding gradients that may occur when using too large batches are avoided, and at the same time, more data information can be used to update the model parameters. For example, for the image classification of power equipment, newly added image samples may be divided into 32 or 64 images per batch, the gradients are calculated in turn, and the gradient values are stored. After accumulating a certain number of batches, the parameters are updated uniformly.

[0084] The dynamic learning rate adjustment strategy is adjusted based on the model convergence situation and sample difficulty. For the model convergence situation, the downward trend of the loss function after multiple iterations can be observed. If the loss function decreases slowly, it indicates that the model is approaching convergence, and at this time, the learning rate can be reduced, otherwise increased. For sample difficulty, for more difficult samples (such as images of power equipment with complex faults), it may cause larger errors in the model, and the learning rate can be appropriately increased to accelerate convergence; while for simple samples, the learning rate is reduced to avoid overfitting.

[0085] The specific implementation can adopt a learning rate decay function, such as exponential decay:

[0086] lr = lr0 × μ epoch

[0087] where lr0 is the initial learning rate, μ is the decay factor, and epoch is the number of training rounds. The value of μ is adjusted according to different situations.

[0088] The meta-learner regards the sample update as a meta-task. It receives the performance feedback of different batches of samples. For example, based on information such as the prediction accuracy and loss value of the samples, this information is used as the input of the meta-learner. Through a meta-learning network, it learns how to adjust the update step size and direction for each sample. This network can learn the sensitivity of different samples to different parameters and output the optimal update step size and direction for each sample according to the characteristics of the samples. For example, for some key power equipment fault image samples, the meta-learner may learn that they require a larger update step size and a special update direction to ensure that the model can quickly and accurately learn the characteristics of these important samples.

[0089] The contrastive learning algorithm aims to bring the representation of similar samples in new and old samples closer and to pull the representation of dissimilar samples further apart. Through the loss function of contrastive learning, such as the InfoNCE loss function:

[0090]

[0091] where q i is the query sample representation, is the positive sample representation, k j is the negative sample representation, and c is the temperature parameter. By minimizing this loss function, the model learns the similarities and differences between samples. In the classification of power equipment images, for images representing the same fault type in new and old samples, their representations in the feature space are made closer, and images of different fault types are represented farther, thereby optimizing the performance of the model after local fine-tuning and improving the model's ability to distinguish between similar and different categories of samples. In implementation, it can effectively process new data under limited hardware resources, avoid gradient problems caused by excessive data volume, and ensure the stability and effectiveness of model updates. In addition, dynamic learning rate adjustment enables the model to adaptively adjust learning according to data and convergence conditions, improves the training efficiency and performance of the model, and avoids overfitting or underfitting. Secondly, the meta-learner can adjust the update strategy for each sample individually, improves the adaptability of the model to different samples, especially for complex power equipment image data, and can better learn the characteristics of key samples. At the same time, the contrastive learning algorithm can enhance the model's ability to distinguish between samples of different categories, especially for different fault categories of power equipment, and improve the classification accuracy and robustness of the model.

[0092] Soft label distillation regards the prediction results of the new model for the newly added data samples as soft labels. Unlike traditional hard labels (true labels), soft labels contain more category probability information. The difference between the soft label and the prediction result of the main model is calculated by KL divergence. The KL divergence formula is:

[0093]

[0094] Where P is the prediction distribution of the new model and Q is the prediction distribution of the main model. Using the KL divergence as the soft loss function aims to make the prediction distribution of the main model closer to that of the new model, enabling the main model to learn the probability distribution information of the new model rather than just the information of hard labels. For power equipment images, it allows the main model to learn the probability judgments of the new model for fuzzy or uncertain samples. The feature mapping distillation mechanism constructs a feature mapping function that maps the feature representations of the intermediate layer of the new model to the corresponding positions in the intermediate layer of the main model, which involves learning complex mapping relationships. By minimizing the cosine similarity difference between the features of the intermediate layers of the new model and the main model, for images of power equipment, different intermediate layers of different models may extract features at different levels. Through feature mapping, the main model can learn better feature representations of the intermediate layer of the new model, helping the main model understand image features from a richer perspective. And calculating the attention weights of the new model and the main model in different regions, which is usually done through attention mechanisms such as self-attention mechanisms. For different regions in the image, different attention weights are obtained according to their importance to the final prediction. Then, the difference in attention weights between the two is compared, and the difference information is fed back to the main model. For example, for the key component regions of power equipment, if the new model assigns high attention weights, this method guides the main model to also focus more attention on these regions, enabling the main model to learn the attention allocation strategy of the new model, optimizing the attention to key components, and improving the feature extraction and classification capabilities for key components. In implementation, soft label distillation enables the main model to learn richer prediction information from the new model, not limited to hard labels, improving the generalization ability of the model and the prediction ability for new samples. The feature mapping distillation mechanism allows the main model to learn excellent feature representations of the intermediate layer of the new model, improving the feature extraction ability of the main model and enhancing the model's understanding of complex image features. Attention-guided distillation can guide the main model to learn the attention allocation of the new model, increasing the attention of the main model to important regions, especially for key components in power equipment, improving the recognition and classification performance of the model for different components.

[0095] The model tuning module is used to dynamically constrain the model weight parameters of the model fine-tuned by the model distillation module using the Elastic Weight Consolidation method. When the model tuning module detects a conflict between new data features and the existing model, it elastically adjusts the conflicting parameters and pushes the adjusted model parameters to the model version control module in real time; when it detects that the data features are consistent with the existing model, it maintains the current parameter configuration. The Elastic Weight Consolidation method treats the weight parameters as random variables and uses a Bayesian network to establish the probability relationship between the weight parameters and the comprehensive feature vector. The Bayesian network calculates the posterior probability distribution of different weight parameters through Bayesian inference and outputs the optimal weight range. During the weight adjustment process, the conflict situation between the current data features and the existing model is considered, and an elastic adjustment range is assigned to each weight parameter according to the expectation and variance of the posterior probability distribution. When the Elastic Weight Consolidation method detects a conflict between new data features and the existing model, it uses the gradient-aware conflict detection algorithm to mark the parameters with abnormal gradients as conflicting parameters based on the gradient change of the model on the new data, and adjusts the conflicting parameters according to the optimal weight range obtained by Bayesian inference. During the adjustment process, the Elastic Weight Consolidation method adopts an adaptive learning rate decay strategy to dynamically adjust the learning rate according to the severity of the conflict and the model performance feedback.

[0096] In the Elastic Weight Consolidation method, the weight parameters of the model are regarded as random variables rather than traditional fixed values. A Bayesian network is a probabilistic graphical model that represents the dependence relationship between variables through a directed acyclic graph. For the weight parameters of the model and the comprehensive feature vector, the Bayesian network constructs a graph structure according to their internal connection. For example, for an electric power equipment image classification system, assume that one node represents the weight of a certain convolutional layer and another node represents the feature vector extracted from the image. The Bayesian network will learn the probability relationship between them based on prior knowledge and training data.

[0097] Through a large amount of data and iterative calculations, the Bayesian network can capture the probability distribution of different weight parameters under different feature vector inputs, providing a basis for subsequent weight adjustment. Bayesian inference is an iterative calculation process based on Bayes' theorem. After receiving new feature vector data, the posterior probability distribution is continuously updated according to the existing Bayesian network structure and prior information. For example, for different types of electric power equipment image features, the probability distribution of different weight parameters will be updated. The optimal weight range is determined by calculating the expectation and variance of the posterior probability distribution. The expectation can be regarded as the most likely weight value, and the variance reflects the uncertainty of the weight. A smaller variance indicates that the estimation of this weight is more certain, while a larger variance indicates that more data is needed to accurately estimate the weight.

[0098] For a model that classifies power equipment faults, it may calculate the optimal range of different convolutional kernel weights based on the image features of different fault types. This range takes into account the diversity and uncertainty of the features and provides a flexible range for weight adjustment.

[0099] The gradient-aware conflict detection algorithm detects conflicts by observing the gradient changes of the model on new data. The gradient reflects the sensitivity of the model to errors under the current parameters. When new data is input, if the gradient change of a weight parameter is abnormal compared to the normal situation, such as a sudden increase or decrease in the gradient, it may indicate that the feature corresponding to this parameter is no longer applicable in the new data, and it is marked as a conflict parameter. For example, for new power equipment image data, if the gradient of a weight parameter used to detect insulator faults suddenly changes significantly, it may be because new aging features have appeared in the insulators in the new image, and the original parameter is no longer appropriate. For the parameters marked as conflicts, they are adjusted according to the optimal weight range obtained by Bayesian inference. Since the optimal range takes into account the uncertainty of the data and prior information, the adjustment process is more scientific and reasonable.

[0100] The adaptive learning rate decay strategy dynamically adjusts the learning rate according to the severity of the conflict and the model performance feedback. The severity of the conflict can be measured by the magnitude of the gradient change or the degree of performance degradation of the model on new data. For example, if the conflict is severe, a larger learning rate decay factor can be used to make the learning rate drop rapidly to avoid over-adjustment; if the conflict is small, a smaller decay factor is used to ensure fine-tuning of the parameters. At the same time, performance feedback, such as changes in accuracy and recall, also affects the adjustment of the learning rate. If the model performance improves after adjustment, the learning rate can be appropriately reduced for more fine-grained adjustment; if the performance drops, the learning rate may need to be increased or the parameter range readjusted.

[0101] In implementation, regarding the weight parameters as random variables and establishing probability relationships using Bayesian networks make the model tuning more flexible and adaptable, enabling better handling of the uncertainties brought by new data. For the dynamic changes in the operating environment of power equipment and the equipment status, such as new fault characteristics, the model can adjust the weights according to the probability distribution instead of remaining fixed. Additionally, the optimal weight range obtained through Bayesian inference takes into account various factors, making the weight adjustment more scientific, avoiding the problems of over-adjustment or under-adjustment, and improving the stability and performance of the model when dealing with new data. Secondly, the gradient-aware conflict detection algorithm can accurately identify the parameters that are not suitable for new data, improving the pertinence of conflict detection, avoiding blind adjustment of all parameters, and enhancing the tuning efficiency. At the same time, the adaptive learning rate decay strategy dynamically adjusts the learning rate according to the actual situation, can flexibly adjust according to different conflict situations and performance feedback, ensures the stability and convergence of the model during the adjustment process, avoids performance degradation or non-convergence of training caused by inappropriate learning rates, and improves the adaptability of the model to new data and the recognition ability for different fault types.

[0102] The model version control module is used to monitor the model performance metrics in real time through a multi-version model management mechanism; and when it detects that the performance of the new model is better than that of the old model, it automatically deploys the new model; when the performance of the new model deteriorates or fluctuates beyond the preset threshold, it automatically rolls back to the previous stable version through the version rollback strategy. The model performance metrics monitored by the model version control module in real time include accuracy, recall, precision, F1 score, average precision threshold, and inference time; the model version control module considers that performance fluctuations occur when the change range of the time series data of the model performance metrics exceeds the preset threshold range within a fixed time window.

[0103] Among them, for accuracy, it is obtained by dividing the number of correctly predicted samples of the model for the test data set by the total number of samples. For power equipment image classification, different versions of the model are used to predict the test set (including images of various power equipment in different states), and the number of correct classifications is counted. For example, assuming there are 1000 power equipment images in the test set and the model correctly classifies 850 of them, the accuracy is 850 / 1000 = 0.85.

[0104] Recall is calculated for each category. For a certain category, such as the transformer fault category in power equipment, it is the number of samples correctly predicted as transformer faults by the model divided by the actual number of transformer fault samples. Precision is the proportion of the number of samples that truly belong to a certain category (such as transformer faults) among the samples predicted by the model as a certain category.

[0105] The F1-score balances recall and precision. After calculating the F1-score for each category, the mean average precision (mAP) is then considered, which is obtained by weighted averaging the F1-scores of different categories to comprehensively evaluate the performance of the model in multi-class tasks.

[0106] The inference time is the time taken from the input image to the output prediction result when the model makes predictions on the test set, which is obtained by recording the start and end times and calculating the difference.

[0107] These performance metrics will be stored in a performance metric database. Each version of the model has corresponding performance metric records, forming a time series data for subsequent analysis and comparison. Whenever there is a new model version update or a new test set, the above performance metrics will be recalculated and the results added to the performance metric time series of the corresponding version. In actual operation, new power equipment images will be continuously input into the system, predictions will be made on them, and the corresponding performance metrics will be updated. For example, for newly acquired power equipment images, the system will use them as part of the test set to calculate the performance metrics of the model and update the storage.

[0108] For different types of power equipment images, different weights will be assigned according to their importance and occurrence frequency to more accurately reflect the performance of the model in actual applications. For example, for key transmission line equipment images, their performance metrics may be given higher weights during calculation because they are crucial for the stable operation of the power system. In implementation, the multi-version model management mechanism provides a comprehensive performance evaluation system for the system by accurately calculating and storing various performance metrics, which can clearly reflect the performance of different version models at different time points, helping system administrators comprehensively understand the performance evolution of the model. The real-time update of performance metrics enables the system to keep track of the performance changes of the model when facing continuously updated power equipment image data, providing data support for subsequent model selection and optimization, and improving the adaptability and maintainability of the system.

[0109] When a new model is generated, the performance metrics of the new model are compared with those of the old model. For accuracy, if the accuracy of the new model on the same test set is higher than that of the old model, it indicates that the new model may be better in overall classification ability. For example, if the accuracy of the old model is 80% and that of the new model is 85%, the new model is more excellent in this metric. For recall and precision, each category is compared. If the new model improves both recall and precision in key power equipment fault categories (such as capacitor faults), it means that the new model performs better in finding and accurately identifying such faults. Compare the F1-score and mAP of the new model and the old model to comprehensively evaluate the performance improvement of the new model in multi-class tasks. At the same time, consider the inference time. If the new model has shorter or the same inference time while improving the performance metrics, it indicates its greater performance advantage.

[0110] For different performance metrics, different weights are assigned according to the actual requirements of power equipment. For example, for the fault detection task, the weight of recall rate may be higher because missing a fault will bring serious consequences. While for the identification of equipment in normal state, the weight of precision rate may be higher to avoid false alarms. Through weighted calculation, a comprehensive performance score is obtained, and the comprehensive performance scores of the new model and the old model are compared.

[0111] When the comprehensive performance score of the new model is higher than that of the old model and meets a certain advantage threshold (for example, the comprehensive performance score is more than 5% higher than that of the old model), the system will trigger the operation of automatically deploying the new model. This may involve loading the parameters of the new model into the actually running system and updating the model service of the system to ensure that the system uses a model with better performance for the power equipment image classification task. In implementation, through comprehensive and detailed performance comparison and reasonable weighted evaluation, it can scientifically judge whether the new model is better, avoiding the one-sidedness of relying only on a single metric and ensuring that the deployed new model better meets the actual requirements of power equipment image classification in multiple performance dimensions.

[0112] The mechanism of automatically deploying the new model improves the timeliness of system update, enabling the system to quickly adopt a model with better performance, enhancing the overall performance of the system, and ensuring the efficiency and accuracy of equipment fault detection and identification in the power system.

[0113] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that these specific implementation manners are only illustrative. Without departing from the principles and essence of the present invention, those skilled in the art can make various omissions, substitutions, and changes to the details of the above methods and systems. For example, combining the above method steps so as to perform substantially the same function in a substantially the same way to achieve substantially the same result falls within the scope of the present invention. Therefore, the scope of the present invention is only defined by the appended claims.

Claims

1. An image classification and recognition system based on an AI training model; characterized in that: The system includes a data monitoring module, a dynamic label correction module, a model distillation module, a model tuning module and a model version control module; The data monitoring module is used to evaluate the degree of deviation between the input data and the training data distribution in real time by using a joint distribution deviation detection algorithm based on maximum mean difference and KL divergence; If the distribution deviation between the input data and the training data is detected to be greater than a preset threshold, a model update request is triggered, a distribution deviation alarm is generated, and the input data is transmitted to the dynamic label correction module; If the distribution deviation between the input data and the training data is less than or equal to the preset threshold, the model is maintained in operation; The dynamic label correction module is used to automatically correct the labels of input data through clustering-based active learning and semantic segmentation algorithms, and output high-confidence samples and low-confidence samples; the dynamic label correction module is also used to transmit the high-confidence samples to the model distillation module, submit the low-confidence samples to the expert system for manual verification through a feedback mechanism, and output the verified samples to the model distillation module; The model distillation module is used to fine-tune the local model for the newly added data samples using an adaptive incremental learning algorithm, and to transfer the new model knowledge to the main model in combination with the knowledge distillation mechanism; The model tuning module is used to dynamically constrain the model weight parameters of the model fine-tuned by the model distillation module using an elastic weight integration method. When the model tuning module detects that the new data features conflict with the existing model, it elastically adjusts the conflicting parameters and pushes the adjusted model parameters to the model version control module in real time; when it detects that the data features are consistent with the existing model, it maintains the current parameter configuration; The model version control module is used to monitor model performance indicators in real time through a multi-version model management mechanism; and automatically deploy a new model when it is detected that the performance of the new model is better than the old model; when the performance of the new model degrades or fluctuates beyond a preset threshold, it automatically rolls back to the previous stable version through a version rollback strategy.

2. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: In the data monitoring module, the method for the joint distribution offset detection algorithm of the maximum mean difference and KL divergence to evaluate the offset degree of the input data and the training data distribution in real time is: based on the input image feature vector obtained by feature extraction and the training data reference distribution model with time dynamic characteristics generated by time series analysis in combination with the historical data of power equipment operation, the distance between the input data and the training data in the feature space is measured by an adaptive kernel function to obtain the maximum mean difference distance; at the same time, the difference in the probability distribution of the input data and the training data is calculated by the KL divergence algorithm to obtain the KL divergence value; finally, the weights are obtained according to the severity and occurrence probability of different types of faults of the power equipment through the hierarchical analysis method, and the maximum mean difference distance and the KL divergence algorithm are weightedly fused to obtain a comprehensive offset index.

3. The image classification and recognition system based on the AI ​​training model according to claim 2, characterized in that: The adaptive kernel function automatically adjusts the bandwidth according to the local density of the input feature vector through an adaptive bandwidth selection strategy; the working method of the adaptive bandwidth selection strategy is: r1. Based on the input feature vector, calculate the local data density by kernel density estimation method; r2. Use the Silverman empirical method to preliminarily determine the bandwidth; r3. Use cross-validation to fine-tune the bandwidth within the preset range; The KL divergence algorithm corrects the noise and uncertainty in the data through the entropy correction factor to reflect the real difference in data distribution.

4. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: In the dynamic label correction module, the working method for automatic label correction of input data is as follows: first, the attention weights of pixels in different areas are calculated through a semantic segmentation network that integrates a multi-head self-attention mechanism to generate a segmented image with semantic information; then, a density clustering algorithm combined with a domain knowledge graph is used to integrate the information of the knowledge graph into density and neighborhood calculations, cluster the image data points, and correct the labels based on the clustering results to obtain the image data after label correction; based on the image data after label correction, a confidence assessment model based on Bayesian deep learning is used to estimate the uncertainty of the predicted distribution through Monte Carlo dropout, and generate a confidence score for each sample.

5. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: The adaptive incremental learning algorithm adopts a gradient accumulation mechanism to divide the newly added data samples into multiple small batches. In the processing of each small batch, the model parameters are updated according to the sample gradient information. At the same time, the learning rate is adaptively adjusted according to the model convergence and sample difficulty through a dynamic learning rate adjustment strategy, and the locally fine-tuned model is output. When processing the newly added data samples, the adaptive incremental learning algorithm learns the optimal update step size and direction for each sample according to the performance feedback of samples in different batches through a meta-learner, and compares the new samples with the old samples through a contrastive learning algorithm, calculates the loss using the contrastive learning loss function, and optimizes the performance of the model after local fine-tuning.

6. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: In the model distillation module, the working method of the knowledge distillation mechanism for migrating the new model knowledge to the main model is as follows: first, the prediction result of the new model is used as a soft label through the soft label distillation method, and the soft loss of the soft label and the prediction result of the main model is calculated using the KL divergence, with the soft loss as the optimization target; then, in the processing of the intermediate layer feature representation of the new model and the main model, the feature mapping distillation mechanism is used to construct a feature mapping function to map the intermediate layer features of the new model to the corresponding intermediate layer of the main model, and the feature difference is measured by cosine similarity to optimize the main model parameters; in addition, the attention-guided distillation method is used to calculate the attention weights of different regions of the new model and the main model, and the main model is guided to learn the attention mechanism of the new model according to the weight difference, thereby optimizing the performance of the main model.

7. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: The elastic weight integration method regards the weight parameters as random variables, and uses the Bayesian network to establish the probabilistic relationship between the weight parameters and the comprehensive feature vector; the Bayesian network calculates the posterior probability distribution of different weight parameters through Bayesian reasoning, and outputs the optimal range of weights; in the weight adjustment process, the conflict between the current data features and the existing model is considered, and an elastic adjustment range is allocated to each weight parameter according to the expectation and variance of the posterior probability distribution; when the elastic weight integration method detects that the new data features conflict with the existing model, the gradient-abnormal parameters are marked as conflict parameters based on the gradient changes of the model on the new data through the gradient-aware conflict detection algorithm, and the conflict parameters are adjusted according to the optimal range of weights obtained by Bayesian reasoning; during the adjustment process, the elastic weight integration method adopts an adaptive learning rate decay strategy to dynamically adjust the learning rate according to the severity of the conflict and model performance feedback.

8. The image classification and recognition system based on the AI ​​training model according to claim 1, characterized in that: The model performance indicators monitored in real time by the model version control module include accuracy, recall rate, precision, F1 score, average precision threshold and inference time; the model version control module considers that performance fluctuations occur when the range of change of the time series data of the model performance indicators within a fixed time window exceeds a preset threshold range.

9. An image classification and recognition method based on an AI training model, characterized in that: An image classification and recognition system based on an AI training model as described in any one of claims 1 to 8, comprising the following steps: Step S1: Based on the input power equipment image data, a joint distribution shift detection algorithm based on maximum mean difference and KL divergence is used to perform a real-time comparison between the distribution of the input data and the training data; Step S2, measuring the difference between the input data and the training data in the feature space distribution by calculating the maximum mean difference, and using KL divergence to measure the difference in the probability distribution of the two; Step S3: If the calculated distribution deviation is greater than a preset threshold, a model update request is triggered, a distribution deviation alarm is generated, and data is transmitted to the dynamic label correction module; if the deviation is less than or equal to the preset threshold, the model parameters are maintained to run at the current state; Step S4: Based on the output data of step S3, cluster analysis is performed according to the feature similarity of the data through a clustering-based active learning algorithm to obtain the potential structure and pattern in the data; at the same time, different regions in the image are semantically understood and segmented through a semantic segmentation algorithm to identify different parts of the power equipment and its components in the image; Step S5: automatically correct the labels of the data, and select high-confidence samples and low-confidence samples through confidence evaluation according to the label correction results; Step S6: directly transmit the high-confidence samples to the model distillation module; submit the low-confidence samples to the expert system for verification through the feedback mechanism, and output the verified samples to the model distillation module; Step S7, receiving the output data of step S6 as a new data sample; using an incremental learning strategy to fine-tune the local model; and combining the knowledge distillation mechanism to migrate the local model fine-tuned for the new data to the main model in the form of a soft label; Step S8, using an elastic weight integration method to dynamically constrain the model weight parameters of the output model of step S7; Step S9: Push the adjusted model parameters to the model version control module in real time; if it is detected that the new data features are consistent with the existing model, maintain the current parameter configuration unchanged; Step S10: Monitor the performance indicators of the model in real time through the multi-version model management mechanism, including accuracy, recall rate, F1 value, average precision threshold and inference time; if the performance indicators of the new model are better than those of the old model, apply the new model to the actual image classification and recognition task; if the performance of the new model decreases or fluctuates beyond the preset threshold, automatically roll back the model to the previous stable version through the version rollback strategy.

Citation Information

Cited By

  • Method and system for predicting on-machine wear state of diamond milling cutter based on machine learning

    CN121073934A

  • Incremental OS-ELM cable defect online identification method and system

    CN121211371A

  • Target feature intelligent identification system and method based on multi-source data

    CN121365269A

  • Power equipment identification method based on deep learning

    CN121412758A

  • A power equipment identification method based on deep learning

    CN121412758B