Cross-scene defect identification method and system driven by migration and fine tuning

By constructing a knowledge graph in the field of power defects and using CycleGAN style transfer, combined with reinforcement learning controller to optimize the transfer strategy, the problems of data scarcity and poor generalization ability in cross-scenario identification of power equipment are solved, and the model is able to be efficiently adapted to diverse inspection scenarios and the recognition accuracy is improved.

CN120976200APending Publication Date: 2025-11-18INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511286509.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing power equipment defect identification technologies suffer from poor cross-scenario generalization ability and a scarcity of labeled data in the target domain, making it impossible for models to adapt to diverse inspection scenario requirements.

Method used

By pre-training an initial model using a general power image dataset, a knowledge graph for the power defect domain is constructed. CycleGAN is used for style transfer and semantic consistency verification. A reinforcement learning controller is combined to generate a transfer fine-tuning strategy to optimize the model's adaptability to the target domain.

Benefits of technology

A large number of image samples with the same visual style as the target domain were generated, which expanded the scale of the target domain training set, improved the generalization ability and recognition accuracy of the model in cross-scene defect recognition, and solved the problems of data scarcity and cross-scene adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976200A_ABST
    Figure CN120976200A_ABST
Patent Text Reader

Abstract

The invention discloses a migration fine tuning driven cross-scene defect identification method and system, and relates to the technical field of electric power inspection, and the method comprises the steps: obtaining a basic model through employing a general electric power image data set; constructing a triple set, and forming a power defect domain knowledge graph; performing style conversion on the general power image by adopting CycleGAN, and performing semantic consistency verification on an image sample through a triple set; forming a state vector; inputting the state vector into a reinforcement learning controller, and performing fine tuning training on the basic model to obtain a fine tuning defect identification model; and evaluating the performance of the fine tuning defect identification model on the target domain image verification set, and optimizing a migration strategy selection mechanism according to a performance index. According to the method, the problem of scarcity of target domain labeling data can be solved, sufficient data support is provided for model fine adjustment, model overfitting caused by insufficient data is avoided, the model can adapt to a visual style and a defect distribution rule of a target scene, and the generalization ability of cross-scene defect recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power inspection, in particular to a cross-scene defect identification method and system driven by transfer fine-tuning. BACKGROUND

[0002] With the intelligentization and scaling development of the power system, the safe operation and efficient maintenance of power equipment have become the core demand to ensure the stability of the power grid. As a key means of equipment state monitoring, power inspection has gradually shifted from traditional manual inspection to an automated inspection mode based on image recognition technology. At present, through the deep learning model, the equipment defects in the inspection image are automatically detected, which can greatly improve the inspection efficiency, reduce the labor cost and the safety risk of operation, and has become the core research direction of intelligent operation and maintenance in the power industry.

[0003] However, the existing power equipment defect identification technology still has some problems: Poor cross-scene generalization ability: different power scenes (such as substations, transmission lines, distribution cabinets, etc.) have significant visual style differences, including environmental lighting, equipment layout, shooting angle, etc., which leads to poor generalization ability of the model in new scenes, and the model cannot adapt to the diversified inspection scene requirements; Lack of labeled data in target scenes: power equipment defect samples cannot quickly accumulate a sufficient number of defect images, and manual labeling of defect images requires professional power knowledge support, which has high labeling cost and long cycle, limiting the transfer and fine-tuning effect of the model.

[0004] Based on this, the present application provides a cross-scene defect identification method and system driven by transfer fine-tuning, which can eliminate the drawbacks of the prior art. SUMMARY

[0005] The purpose of the present application is to provide a cross-scene defect identification method and system driven by transfer fine-tuning to solve the problems of poor cross-scene generalization ability and lack of labeled data in the target domain in the background art.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions: A cross-scene defect identification method driven by transfer fine-tuning, specifically comprising the following steps: Step S1, pre-training an initial defect identification model using a general power image dataset to obtain a basic model with power field representation capability, wherein the general power image dataset includes defect types and parameter data of multiple types of power equipment; Step S2, extracting semantic information based on the general power image dataset to construct a triple set composed of entities-relation-entities to form a power defect field knowledge graph; Step S3: Use CycleGAN to perform style transfer on the general power images in the general power image dataset to generate image samples with the same style as the target domain image scene, and use the triple set to verify the semantic consistency of the image samples, and use the verified image samples as amplification samples. Step S4: Extract the style features, defect distribution features, and statistical difference features between the target domain image samples and the source domain image samples to form a state vector; Step S5: Input the state vector into the reinforcement learning controller. The reinforcement learning controller outputs a transfer fine-tuning strategy and fine-tunes the base model according to the transfer fine-tuning strategy to obtain a fine-tuned defect recognition model adapted to the target domain image. Step S6: Construct a target domain image validation set using the image set used to evaluate model performance in the target domain scene; evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set; generate a reward signal based on the performance index and feed it back to the reinforcement learning controller to optimize the transfer strategy selection mechanism.

[0007] Preferably, step S2 specifically includes: Step S2-1: Extract equipment entities, defect entities, and structural entities from the general power image dataset. The equipment entities include transformers, cables, and insulators. The defect entities include cracks, corrosion, and damage. The structural entities include terminals, shells, and supports. Step S2-2: Construct the association relationship between entities based on spatial location relationship and semantic subordination relationship. The spatial location relationship is inferred by combining image analysis and annotation information. The semantic subordination relationship includes existing, belonging to, and occurring. Step S2-3: Based on the extracted equipment entities, defect entities, and structural parts entities, as well as the constructed spatial location relationships and semantic subordination relationships, organize them into a set of entity-relationship-entity triples; Step S2-4: Standardize, merge, and deduplicate the triple set. Construct a power defect domain knowledge graph based on entity type and relation type. The power defect domain knowledge graph is stored and managed using Neo4j and RDF database.

[0008] Preferably, the specific steps for semantic consistency verification of image samples in step S3 include: Step S3-1: Identify defective entities and semantic relationships in the image samples generated by CycleGAN, and construct the triple representation corresponding to the image sample. The process of identifying the image sample is consistent with the process of extracting entities and constructing relationships in step S2. Step S3-2: Match the triple representation of the image sample with the triple set in the knowledge graph of the power defect domain. If all the triple representations of the image sample exist in the triple set, it is determined to be semantically consistent and the image sample is selected as the augmented sample. Otherwise, the image sample is discarded and CycleGAN is used to regenerate the image.

[0009] Preferably, step S4 specifically includes: Step S4-1: Use a pre-trained style encoder to obtain style features of the target domain image samples, the style features including color distribution features, texture direction features and brightness contrast features; Step S4-2: Statistically analyze the category distribution ratio of defect samples in the target domain image samples to obtain the defect distribution characteristics of the target domain image samples. The defect distribution characteristics include the quantity distribution of each type of defect in the target domain and the quantity ratio of each type of defect in the target domain scene. Step S4-3: Calculate the statistical difference measure between the target domain image sample and the source domain image sample in the feature space. The statistical difference measure includes the mean difference, covariance difference, and maximum mean difference. Step S4-4: Perform Z-score normalization on the style features, defect distribution features and statistical difference features, and concatenate the normalized features to form a fixed-dimensional state vector.

[0010] Preferably, the style encoder is used to extract style features from the target domain image, including: Color statistics unit: used to convert the target domain image to RGB color space and Lab color space respectively, and to extract color distribution features for each color channel of the color space using the color histogram method; Texture feature unit: Used to input the target domain image into the convolutional layer of a pre-trained convolutional neural network, extract feature maps and calculate the corresponding Gram matrix to obtain texture features that characterize the texture structure of the image; Brightness / Contrast Unit: Used to extract the average brightness value, standard deviation, and brightness dynamic range of the L channel in the Lab color space.

[0011] Preferably, the migration fine-tuning strategy in step S5 specifically includes: Fine-tuning level selection: controls whether to freeze the first few layers of the base model and fine-tune the subsequent layers of the base model; Feature alignment method selection: Based on the statistical difference metric between the target domain image sample and the source domain image sample, the reinforcement learning controller selects an appropriate feature alignment method, which includes maximum mean difference matching, adversarial domain discriminator training, and batch normalized statistical readjustment. Pseudo-label confidence threshold setting: In pseudo-label training of unlabeled target domain image samples, a confidence threshold is set so that only high-confidence image samples participate in training; Loss function weighting strategy: Determine the weighting coefficients for the source domain supervision loss and the target domain migration loss.

[0012] Preferably, the specific steps of optimizing the transfer strategy selection mechanism in step S6 include: testing the performance indicators of the fine-tuned defect recognition model on the target domain image validation set, the performance indicators including recognition accuracy and recall, comparing the performance indicators with preset benchmark values, calculating reward values ​​through a reward function, feeding the reward values ​​back to the reinforcement learning controller, using the Q-value update algorithm to reinforce the parameters of the reinforcement learning controller, and improving the effectiveness of the transfer strategy through multiple rounds of iterative training.

[0013] Preferably, the initial defect identification model adopts a deep learning network structure, which is one of YOLO, Faster R-CNN or Swin-Transformer.

[0014] Preferably, the abnormal image samples in the semantic consistency verification include: images of defective entities and device entities that are incorrectly combined, images of defective parts that are misplaced, and images of entities that overlap and conflict.

[0015] A transfer-fine-tuning-driven cross-scene defect identification system includes: The pre-training module is used to pre-train the initial defect identification model using a general power image dataset; The semantic extraction module is used to extract semantic information based on the general power image dataset, construct a set of triples consisting of entity-relationship-entity, and form a knowledge graph in the power defect domain. The style transfer module is used to perform style transfer on the general power image using CycleGAN and generate image samples that are consistent with the scene style of the target domain image. The feature extraction module is used to extract the style features, defect distribution features, and statistical difference features between the target domain image samples and the source domain image samples, forming a state vector; The strategy generation module is used to input the state vector into the reinforcement learning controller, the reinforcement learning controller outputs a transfer fine-tuning strategy, and fine-tunes the base model according to the transfer fine-tuning strategy to obtain a fine-tuned defect recognition model adapted to the target domain image. The strategy optimization module is used to evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set, and generate a reward signal based on the performance index and feed it back to the reinforcement learning controller to optimize the transfer strategy selection mechanism.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention achieves the transfer of general power images to the style of the target scene through CycleGAN, which can generate a large number of image samples consistent with the visual style of the target domain. It also introduces a knowledge graph of the power defect domain and performs semantic consistency verification on the generated image samples based on the triple set of entity-relationship-entity, which facilitates the filtering of semantically abnormal samples, thereby expanding the scale of the target domain training set, solving the problem of scarce labeled data in the target domain, providing sufficient data support for model fine-tuning, and avoiding model overfitting due to insufficient data. 2. This invention extracts style features such as color, texture, and brightness of the target domain through a style encoder, and constructs a state vector by combining defect distribution features and statistical differences. This enables the reinforcement learning controller to accurately perceive the degree of difference across scenes, and then outputs an adaptive transfer fine-tuning strategy, allowing the model to quickly adapt to the visual style and defect distribution patterns of the target scene, thereby improving the generalization ability of cross-scene defect recognition. Attached Figure Description

[0017] Figure 1 This is a schematic diagram illustrating the steps of the cross-scene defect identification method of the present invention.

[0018] Figure 2 This is a schematic diagram of step S2 of the present invention.

[0019] Figure 3 This is a schematic diagram of step S3 of the present invention.

[0020] Figure 4 This is a schematic diagram of step S4 of the present invention.

[0021] Figure 5 This is a schematic diagram of the cross-scene defect recognition system of the present invention.

[0022] Figure labeling: Pre-training module 10, semantic extraction module 20, style transfer module 30, feature extraction module 40, policy generation module 50, policy optimization module 60. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0024] In this embodiment, as Figures 1-4 As shown, a cross-scene defect identification method driven by transfer fine-tuning specifically includes the following steps: Step S1: Use a general power image dataset to pre-train the initial defect identification model to obtain a basic model with power field representation capabilities. The general power image dataset includes defect types and parameter data of multiple types of power equipment. Specifically, the general power image dataset covers multiple types of power equipment and their common defect types, and has good universality and representativeness. The initial defect recognition model adopts a deep learning network structure, which is one of YOLO, Faster R-CNN or Swin-Transformer. This step adopts the mainstream deep learning network structure and combines it with target detection or image classification tasks to carry out supervised training on the general power image dataset, so that the model learns the basic visual features and defect representation capabilities in power scenarios. Step S2: Extract semantic information based on the general power image dataset, construct a set of triples consisting of entity-relationship-entity, and form a knowledge graph in the power defect domain; like Figure 2 As shown, it specifically includes: Step S2-1: Extract equipment entities, defect entities, and structural entities from the general power image dataset. Equipment entities include, but are not limited to, transformers, cables, and insulators. Defect entities include, but are not limited to, cracks, corrosion, and damage. Structural entities include, but are not limited to, terminals, shells, and supports. The general power images in the general power image dataset are all labeled. Entity information, such as structural parts and their corresponding positions or functions, can be extracted from the general power images through manual labeling, semi-automatic labeling, or by using target detection technology. Step S2-2: Construct the relationships between entities based on spatial location relationships and semantic subordination relationships. Spatial location relationships are inferred by combining image analysis and annotation information, for example: "Insulators are located above power lines" and "Transformers are located in the center of substations". Semantic subordination relationships include existing, belonging to, and occurring, for example: "Cracks are defects of insulators" and "Corrosion may exist on metal casings". Use knowledge bases or natural language processing tools such as SpaCy, OpenIE, and BERT to extract semantic relationships, for example: "Transformers are part of power equipment" and "Wires have a functional relationship with current". Steps S2-3: Based on the extracted equipment entities, defect entities, and structural component entities, as well as the constructed spatial location relationships and semantic subordination relationships, organize them into a set of entity-relationship-entity triples. For example, transformer – defect exists – corrosion. These triples, described in natural language, clearly indicate the specific relationships between different entities. Step S2-4: Standardize, merge, and remove duplicates from the set of triples. Standardization avoids the same concept appearing multiple times. For example, high-voltage cable and HV cable are the same entity. A thesaurus or a knowledge graph of the power defect domain can be used for entity standardization. Multiple triples with the same semantics are merged and removed to avoid the impact of redundant data on the knowledge graph. The processed set of triples is used to construct a knowledge graph of the power defect domain according to entity type and relation type. The knowledge graph of the power defect domain is stored and managed using Neo4j and RDF database to facilitate subsequent querying and reasoning, and to ensure the accuracy and reliability of the knowledge graph of the power defect domain. Step S3: Use CycleGAN to perform style transfer on the general power images in the general power image dataset to generate image samples with the same style as the target domain image scene. Then, use the triple set to verify the semantic consistency of the image samples and use the verified image samples as amplification samples. like Figure 3 As shown, the specific steps for semantic consistency verification of image samples include: Step S3-1: Identify defective entities and semantic relationships in the image samples generated by CycleGAN, and construct the triplet representation corresponding to the image sample. The process of identifying the image sample is consistent with the process of extracting entities and constructing relationships in step S2. Step S3-2: Match the triple representation of the image sample with the triple set in the knowledge graph of the power defect domain. If all the triple representations of the image sample exist in the triple set, it is determined to be semantically consistent and the image sample is selected as the augmentation sample. Otherwise, the image sample is discarded and CycleGAN is used to generate the image again. This process is repeated up to N times (e.g., N=5). If a semantically consistent image still cannot be generated, the augmentation of the image is abandoned to avoid infinite loop. Specifically, a general power image is input into a trained CycleGAN network. The CycleGAN network, based on object detection architectures such as Faster R-CNN or DETR, identifies and locates all entities in the image. It incorporates a relationship prediction module to analyze the spatial location and visual context features between pairs of entities, determining whether semantic subordinate relationships such as "exists," "belongs to," or "occurs" exist, and thus constructs triple representations. These triple representations are then trained using the power defect domain knowledge graph constructed in step S2 or its derived labeled data to generate style-transferred images. For example, if the original image in the general power image is a rusted iron tower image against a sunny background, the generated image might be transformed into an image showing rust in a rainy, foggy, or nighttime environment, thus approximating the style characteristics of a specific target scene, such as a mountainous area, rainy season, or low-light scene. The triple representations of the generated image samples are matched and verified against the triple set. If all triple representations exist in the power defect domain knowledge graph, then... If an image is deemed semantically structurally sound and can be used as a training sample, it is considered semantically inconsistent, discarded, and regenerated. In the actual generation process, without a semantic consistency verification mechanism, the CycleGAN network will generate abnormal image samples, including: images with incorrect combinations of defective entities and equipment entities, for example: "The generated image shows 'a broken switch has a defect,' but this type of defect has never been associated with 'switch' in the knowledge graph"; images with disordered defect locations, for example: "The 'crack' in the image is marked on the edge of the 'insulator cap,' but the crack only appears on the 'porcelain skirt' in the knowledge graph, indicating that the defect location is illogical"; and images with overlapping and conflicting entities, for example: "The generated image shows 'wire' and 'surge arrester' entities superimposed at the same pixel position, but this has never co-occurred in the knowledge graph, indicating that there is a visual semantic conflict in the image generation." Specifically, this step addresses the problem of scarce defect data in the target domain by introducing semantic consistency verification based on the power defect knowledge graph. This ensures that the generated images are semantically consistent with the target scene and power defect knowledge, effectively avoiding semantic distortion problems introduced by style transfer, such as incorrect defect locations, false defect types, or unreasonable combinations between locations, thus ensuring the high quality and high credibility of the generated samples. Step S4: Extract the style features, defect distribution features, and statistical difference features between the target domain image samples and the source domain image samples to form a state vector; like Figure 4 As shown, it specifically includes: Step S4-1: Use a pre-trained style encoder to obtain style features of the target domain image samples. The style features include color distribution features, texture orientation features, and brightness contrast features. Step S4-2: Statistically analyze the category distribution ratio of defect samples in the target domain image samples to obtain the defect distribution characteristics of the target domain image samples. The defect distribution characteristics include the quantity distribution of each type of defect in the target domain and the proportion of their occurrence in the target domain scene. Step S4-3: Calculate the statistical difference measure between the target domain image sample and the source domain image sample in the feature space. The statistical difference measure includes the mean difference, covariance difference, and maximum mean difference. Step S4-4: Perform Z-score normalization on style features, defect distribution features and statistical difference features, and concatenate the normalized features to form a fixed-dimensional state vector. Specifically, style features refer to the low-level features related to the visual appearance of a scene extracted from an image by a style encoder, including but not limited to color distribution, texture direction, and brightness contrast. Color distribution features, texture direction features, and brightness contrast features can effectively capture the style features of the target domain image, helping the model understand the visual differences of the target scene. In the target domain image samples, the types and distribution of defects are key factors affecting defect recognition performance. By performing category statistics on defect samples in the target domain image samples, the quantity distribution of each defect category and the proportion of each defect category appearing in the target scene can be obtained. This ensures that the model adjusts its learning strategy according to the distribution characteristics of defects in the target domain, avoiding performance degradation caused by class imbalance. Before constructing the state vector, all extracted features are Z-score normalized to ensure that the influence of different features on the model is on the same order of magnitude. This state vector contains the style features of the target domain image, defect distribution information, and statistical differences between the source and target domains. It can comprehensively reflect the cross-scene differences between the source and target domains, helping the reinforcement learning controller to accurately perceive the transfer difficulty and improve the adaptability and recognition performance of the defect recognition model in new scenes. Specifically, the style encoder is used to extract style features from target domain images. The style encoder is a pre-trained feature extraction module with the same network structure as the first four convolutional layers of VGG-19. This encoder is pre-trained on a large natural image dataset to extract general low-level visual style features without requiring specific training for power images. Its parameters are frozen during state vector construction and are only used for forward propagation. The style encoder includes: Color statistics unit: used to convert the target domain image to RGB color space and Lab color space respectively, and to extract color distribution features for each color channel of the color space using the color histogram method; Texture feature unit: Used to input the target domain image into the convolutional layer of a pre-trained convolutional neural network, extract feature maps and calculate the corresponding Gram matrix to obtain texture features that characterize the texture structure of the image; Brightness / contrast unit: used to extract the average brightness value, standard deviation, and brightness dynamic range of the L channel in the Lab color space; Specifically, in the RGB color space, color histograms are used to statistically analyze the R, G, and B channels, dividing them into 64 color bins to obtain the color distribution characteristics of each channel. In the Lab color space, the mean and standard deviation of the L, a, and b channels are extracted to enhance the representation of brightness and perceived color. The convolutional neural network is a VGG-19 convolutional neural network, and the feature maps of the first four convolutional layers (conv1_1, conv2_1, conv3_1, conv4_1) are extracted. For each feature map... Calculate the corresponding Gram matrix: ,in Indicates the number of channels. and Indicates the height and width of the feature map. The i-th channel of the i-th layer feature map is represented at the i-th position. The style encoder comprehensively extracts the style features of the target domain image through color histogram, Gram matrix and brightness statistics, which can effectively characterize the style differences in color, texture and brightness contrast of different scenes, provide accurate basis for dynamic adjustment of transfer strategy, and improve the adaptability and stability of defect recognition model in cross-scene application. Step S5: Input the state vector into the reinforcement learning controller. The output of the reinforcement learning controller is the transfer fine-tuning strategy. The basic model is fine-tuned and trained according to the transfer fine-tuning strategy to obtain a fine-tuned defect recognition model adapted to the target domain image. Specifically, the migration fine-tuning strategy includes: Fine-tuning layer selection: controls whether to freeze the first few layers of the base model and fine-tune the subsequent layers of the base model. For example, you can choose to freeze the first 4 layers of the base model or only fine-tune the fully connected layers. Based on the task complexity and the amount of data in the target domain, you can then adjust the generalization ability and adaptability of the model. Feature alignment method selection: Based on the statistical difference measure between the target domain image samples and the source domain image samples, the reinforcement learning controller selects an appropriate feature alignment method. Feature alignment methods include maximum mean difference matching, adversarial domain discriminator training, and batch normalized statistical readjustment. Pseudo-label confidence threshold setting: In pseudo-label training of unlabeled target domain image samples, a confidence threshold is set so that only high-confidence image samples participate in training. For example, the confidence threshold is set to 0.7 or 0.8, so that only high-confidence samples participate in training, reducing the negative impact of noisy samples on the model. Loss function weighting strategy: Determine the weighting coefficients for the source domain supervision loss and the target domain transfer loss. An example of a weighted loss function is as follows: ,in, For weighted loss function, For source domain monitoring loss, For target domain migration loss, and These are the corresponding weighting coefficients; Specifically, the reinforcement learning controller adopts an architecture based on a deep Q-network (DQN), which includes a state-value network. This network takes the state vector as input and outputs the Q-value (action value) of all possible transfer fine-tuning policy actions. As is known to those skilled in the art, by modeling the selection problem of transfer fine-tuning policy parameters as a reinforcement learning task, the system can dynamically generate the optimal fine-tuning scheme according to the difference between the source domain and the target domain. This avoids the problem of policy reliance on manual experience setting in traditional transfer learning. It has the characteristics of being adaptive, adjustable, and iteratively optimized, and can effectively improve the model's generalization ability and recognition accuracy in the target domain. It is especially suitable for cross-scene defect recognition tasks with large data distribution variations. Step S6: Construct a target domain image validation set using the image set used to evaluate model performance in the target domain scene; evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set; generate a reward signal based on the performance index and feed it back to the reinforcement learning controller to optimize the transfer strategy selection mechanism. Specifically, the steps for optimizing the transfer strategy selection mechanism include: testing the performance metrics of the fine-tuned defect recognition model on the target domain image validation set, where the performance metrics include recognition accuracy and recall; comparing the performance metrics with preset benchmark values; calculating the reward value using a reward function; feeding the reward value back to the reinforcement learning controller; using the Q-value update algorithm to reinforce the parameters of the reinforcement learning controller; and improving the effectiveness of the transfer strategy through multiple rounds of iterative training. This process has two termination conditions: one is that the number of training rounds reaches a set upper limit, and the other is that in multiple consecutive rounds of training, the increase in reward value is less than a set threshold, which is considered convergence. Once either condition is met, the iteration stops. Specifically, the fine-tuned defect recognition model is evaluated on a target domain image validation set. Performance metrics are compared with preset thresholds. Based on the comparison results, a reward function is designed to quantify the degree of model performance improvement under this round of strategy. The reward function formula is as follows: ,in, and Indicates the baseline performance threshold. The reward function is defined by α and β, which are preset weighting coefficient hyperparameters used to balance the importance of accuracy and recall in the reward. Typically, α = 1.0 and β = 0.8 are set. This reward value is used as an immediate feedback input to the reinforcement learning controller. By adjusting the transfer fine-tuning policy output, the model's adaptability in the target domain is gradually improved. This avoids the problem of lack of feedback adjustment in policy settings in traditional transfer methods, and effectively improves the adaptability and generalization performance of the defect recognition model in the target domain. Among them, such as Figure 5 As shown, a transfer-fine-tuned cross-scene defect identification system includes: The pre-training module is used to pre-train the initial defect identification model using a general power image dataset; The semantic extraction module is used to extract semantic information based on a general power image dataset, construct a set of triples consisting of entity-relation-entity, and form a knowledge graph in the power defect domain. The style transfer module is used to perform style transfer on general power images using CycleGAN and generate image samples that are consistent with the scene style of the target domain image. The feature extraction module is used to extract the style features, defect distribution features, and statistical difference features between the target domain image samples and the source domain image samples, forming a state vector; The policy generation module is used to input the state vector into the reinforcement learning controller, the reinforcement learning controller outputs the transfer fine-tuning policy, and the base model is fine-tuned and trained according to the transfer fine-tuning policy to obtain a fine-tuned defect recognition model adapted to the target domain image. The strategy optimization module is used to evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set, and generate reward signals based on performance indicators and feed them back to the reinforcement learning controller to optimize the transfer strategy selection mechanism. Example 1

[0025] This embodiment uses substation outdoor inspection images as the target domain and constructs a general power image dataset using data from the State Grid public data platform or SDD (Substation Defect Dataset); Step S1: Pre-train the initial defect identification model with YOLOv5 or ResNet-50 as the backbone to learn the basic representation capabilities of power equipment, components and common defects, and obtain the basic model. Step S2: Extract entities such as "circuit breaker", "corrosion", and "fixing bolt" from the annotation information of the general power image dataset, analyze the spatial position relationship and semantic subordination relationship of the entities in the image, such as "circuit breaker - contains - bolt" and "bolt - exists - corrosion", construct a set of triples and standardize them to form a knowledge graph in the field of power defects; Step S3: Use CycleGAN to convert the original general power images in the general power image dataset into images that conform to the style of the target substation scene. For example, an image with dark lighting and an overhead power tower background. Then, use the style transfer module 30 to represent the triples of the generated image samples and compare them with the knowledge graph of power defects. If the triples of the generated image, such as "disconnect switch - exist - break", exist in the graph, they are retained as amplified samples; otherwise, they are discarded and regenerated. Step S4: Extract style features from the target domain image samples, statistically analyze the distribution of defect categories, calculate the mean difference between the target domain samples and the source domain samples in the feature space, and normalize and concatenate them into a state vector. Step S5: Input the state vector into the reinforcement learning controller, which outputs the transfer fine-tuning strategy. For example: freeze the first 3 layers of the network, use feature alignment, set the confidence threshold to 0.85, set the loss weighting parameters to 0.3 in the source domain and 0.7 in the target domain, and perform fine-tuning operation according to the above parameters to obtain the fine-tuned defect recognition model. Step S6: Test the model performance on the target domain validation set, calculate the accuracy and recall. If the recognition accuracy is higher than the baseline value (set to 0.85), give a positive reward; otherwise, give negative feedback and update the policy network. This process iterates until the policy is stable or the maximum number of rounds is reached and then terminates. In summary, CycleGAN enables the transfer of general power images to the style of the target scene, improving the visual style relevance of the generated image samples. Entity-relationship-entity triples are used to verify the semantic consistency of the generated images, effectively avoiding semantic drift or defective generation problems during style transfer, ensuring the reliability and effectiveness of data augmentation, and improving the model's recognition accuracy and generalization ability in new scenes.

[0026] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A cross-scene defect identification method driven by migration fine-tuning, characterized in that, Specifically, the following steps are included: Step S1: Use a general power image dataset to pre-train the initial defect identification model to obtain a basic model with power field representation capabilities. The general power image dataset includes defect types and parameter data of multiple types of power equipment. Step S2: Extract semantic information based on the general power image dataset, construct a set of triples consisting of entity-relationship-entity, and form a knowledge graph in the power defect domain; Step S3: Use CycleGAN to perform style transfer on the general power images in the general power image dataset to generate image samples with the same style as the target domain image scene, and use the triple set to verify the semantic consistency of the image samples, and use the verified image samples as amplification samples. Step S4: Extract the style features, defect distribution features, and statistical difference features between the target domain image samples and the source domain image samples to form a state vector; Step S5: Input the state vector into the reinforcement learning controller. The reinforcement learning controller outputs a transfer fine-tuning strategy and fine-tunes the base model according to the transfer fine-tuning strategy to obtain a fine-tuned defect recognition model adapted to the target domain image. Step S6: Construct a target domain image validation set using the image set used to evaluate model performance in the target domain scene; evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set; generate a reward signal based on the performance index and feed it back to the reinforcement learning controller to optimize the transfer strategy selection mechanism.

2. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, Step S2 specifically includes: Step S2-1: Extract equipment entities, defect entities, and structural entities from the general power image dataset. The equipment entities include transformers, cables, and insulators. The defect entities include cracks, corrosion, and damage. The structural entities include terminals, shells, and supports. Step S2-2: Construct the association relationship between entities based on spatial location relationship and semantic subordination relationship. The spatial location relationship is inferred by combining image analysis and annotation information. The semantic subordination relationship includes existing, belonging to, and occurring. Step S2-3: Based on the extracted equipment entities, defect entities, and structural parts entities, as well as the constructed spatial location relationships and semantic subordination relationships, organize them into a set of entity-relationship-entity triples; Step S2-4: Standardize, merge, and deduplicate the triple set. Construct a power defect domain knowledge graph based on entity type and relation type. The power defect domain knowledge graph is stored and managed using Neo4j and RDF database.

3. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, The specific steps for semantic consistency verification of image samples in step S3 include: Step S3-1: Identify defective entities and semantic relationships in the image samples generated by CycleGAN, and construct the triple representation corresponding to the image sample. The process of identifying the image sample is consistent with the process of extracting entities and constructing relationships in step S2. Step S3-2: Match the triple representation of the image sample with the triple set in the knowledge graph of the power defect domain. If all the triple representations of the image sample exist in the triple set, it is determined to be semantically consistent and the image sample is selected as the augmented sample. Otherwise, the image sample is discarded and CycleGAN is used to regenerate the image.

4. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, Step S4 specifically includes: Step S4-1: Use a pre-trained style encoder to obtain style features of the target domain image samples, the style features including color distribution features, texture direction features and brightness contrast features; Step S4-2: Statistically analyze the category distribution ratio of defect samples in the target domain image samples to obtain the defect distribution characteristics of the target domain image samples. The defect distribution characteristics include the quantity distribution of each type of defect in the target domain and the quantity ratio of each type of defect in the target domain scene. Step S4-3: Calculate the statistical difference measure between the target domain image sample and the source domain image sample in the feature space. The statistical difference measure includes the mean difference, covariance difference, and maximum mean difference. Step S4-4: Perform Z-score normalization on the style features, defect distribution features and statistical difference features, and concatenate the normalized features to form a fixed-dimensional state vector.

5. The cross-scene defect identification method driven by migration fine-tuning according to claim 4, characterized in that, The style encoder is used to extract style features from a target domain image, including: Color statistics unit: used to convert the target domain image to RGB color space and Lab color space respectively, and to extract color distribution features for each color channel of the color space using the color histogram method; Texture feature unit: Used to input the target domain image into the convolutional layer of a pre-trained convolutional neural network, extract feature maps and calculate the corresponding Gram matrix to obtain texture features that characterize the texture structure of the image; Brightness / Contrast Unit: Used to extract the average brightness value, standard deviation, and brightness dynamic range of the L channel in the Lab color space.

6. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, The migration fine-tuning strategy in step S5 specifically includes: Fine-tuning level selection: controls whether to freeze the first few layers of the base model and fine-tune the subsequent layers of the base model; Feature alignment method selection: Based on the statistical difference metric between the target domain image sample and the source domain image sample, the reinforcement learning controller selects an appropriate feature alignment method, which includes maximum mean difference matching, adversarial domain discriminator training, and batch normalized statistical readjustment. Pseudo-label confidence threshold setting: In pseudo-label training of unlabeled target domain image samples, a confidence threshold is set so that only high-confidence image samples participate in training; Loss function weighting strategy: Determine the weighting coefficients for the source domain supervision loss and the target domain migration loss.

7. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, The specific steps of optimizing the transfer strategy selection mechanism in step S6 include: testing the performance indicators of the fine-tuned defect recognition model on the target domain image validation set, the performance indicators including recognition accuracy and recall, comparing the performance indicators with preset benchmark values, calculating reward values ​​through reward functions, feeding the reward values ​​back to the reinforcement learning controller, using the Q-value update algorithm to reinforce the parameters of the reinforcement learning controller, and improving the effectiveness of the transfer strategy through multiple rounds of iterative training.

8. The cross-scene defect identification method driven by migration fine-tuning according to claim 1, characterized in that, The initial defect identification model adopts a deep learning network structure, which is one of YOLO, Faster R-CNN or Swin-Transformer.

9. The cross-scene defect identification method driven by migration fine-tuning according to claim 3, characterized in that, The abnormal image samples in the semantic consistency verification include: images of defective entities and equipment entities that are incorrectly combined, images of defective parts that are misplaced, and images of entities that overlap and conflict.

10. A cross-scene defect identification system driven by migration fine-tuning, characterized in that, include: The pre-training module (10) is used to pre-train the initial defect identification model using a general power image dataset; The semantic extraction module (20) is used to extract semantic information based on the general power image dataset, construct a set of triples consisting of entity-relationship-entity, and form a knowledge graph in the power defect domain; The style transfer module (30) is used to perform style transfer on the general power image using CycleGAN and generate image samples that are consistent with the scene style of the target domain image; The feature extraction module (40) is used to extract the style features, defect distribution features and statistical difference features between the target domain image samples and the source domain image samples to form a state vector; The strategy generation module (50) is used to input the state vector into the reinforcement learning controller, the reinforcement learning controller outputs a transfer fine-tuning strategy, and fine-tunes the base model according to the transfer fine-tuning strategy to obtain a fine-tuned defect recognition model adapted to the target domain image. The strategy optimization module (60) is used to evaluate the performance of the fine-tuned defect recognition model on the target domain image validation set, and generate a reward signal based on the performance index and feed it back to the reinforcement learning controller to optimize the transfer strategy selection mechanism.

Citation Information

Cited By

  • Equipment manufacturing energy consumption and state visual monitoring method based on industrial internet of things

    CN121767502A

  • Metal surface microdefect polarization detection method based on Gram determinant

    CN122016654A

  • Deep learning model migration adaptation method under small sample labeling scenario

    CN122509283A

  • Deep learning model migration adaptation method under small sample labeling scenario

    CN122509283B