Power equipment visual large model construction, training and evaluation method and system

By constructing a large visual model of power equipment and utilizing a dual-branch structure and multi-round training to optimize the network, the problem of low efficiency in traditional inspection methods is solved, achieving high precision and high efficiency in power equipment detection.

CN119692420BActive Publication Date: 2026-01-27STATE GRID INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411724941.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2026-01-27
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Traditional manual inspection methods face problems of low efficiency and insufficient accuracy in power grids, and existing professional visual models are difficult to meet the needs of massive data and complex scenarios in power equipment inspection.

Method used

A large-scale visual model of power equipment is constructed. A sample database is built based on historical power inspection images and expanded images. A dual-branch structure with adjustable fusion reparameterization and adapter is adopted. The network structure is optimized by combining multiple rounds of coarse-grained and fine-grained training. A parameter freezing loss function is designed. Multi-model hybrid chain anomaly detection and model performance quantification test are carried out.

Benefits of technology

It improves the accuracy and efficiency of power equipment inspection, enhances the model's adaptability to different power scenarios, and enables efficient identification and comprehensive evaluation of defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119692420B_ABST
    Figure CN119692420B_ABST
Patent Text Reader

Abstract

The application discloses a power equipment visual large model construction, training and evaluation method and system. The method comprises the following steps: constructing a power sample database based on historical power inspection images and expanded power inspection images; generating a power detection benchmark large model based on power sample images in the power sample database and pre-training of a pre-constructed power large model; adding a double-branch structure of adjustable fusion reparameterization and an adapter to the power detection benchmark large model to obtain a business model; and performing multi-round coarse-grained training and fine-grained training on the business model to obtain a defect detection business large model in a power scene. Through the above technical scheme, the training efficiency of the defect detection business large model is improved, and the defect detection business large model after training has better model performance in different power scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method and system for constructing, training, and evaluating large-scale visual models of power equipment. Background Technology

[0002] With the continuous expansion of the power grid and the significant increase in the proportion of new energy integration, the demands on the new power grid are constantly upgrading. This leads to the accumulation and sedimentation of massive amounts of data, putting unprecedented pressure on traditional manual inspection methods. Professional visual models built on traditional convolutional neural networks have been widely used in power grid operations. Compared to these models, large-scale models, with their large number of parameters, higher accuracy, and stronger generalization ability, have shown greater application potential in multiple professional fields such as safety supervision and equipment maintenance.

[0003] Therefore, there is an urgent need for a large-scale detection model for power equipment. Summary of the Invention

[0004] This invention provides a method and system for constructing, training, and evaluating a large visual model of power equipment, in order to improve the detection accuracy of power equipment in power scenarios.

[0005] According to one aspect of the present invention, a method for constructing, training, and evaluating a large-scale visual model of power equipment is provided. The method includes: constructing a power sample database based on historical power inspection images and expanded power inspection images; wherein the expanded power inspection images are expanded images of scarce power inspection images; generating a power detection benchmark large-scale model based on power sample images in the power sample database and pre-training the pre-constructed large-scale power model; adding an adjustable fusion reparameterization and adapter dual-branch structure to the power detection benchmark large-scale model to obtain a business model; and performing multiple rounds of coarse-grained training and fine-grained training on the business model to obtain a defect detection business large-scale model in a power scenario.

[0006] According to another aspect of the present invention, a system for constructing, training, and evaluating a large-scale visual model of power equipment is provided. The system includes: a sample database construction module for constructing a power sample database based on historical power inspection images and expanded power inspection images; wherein the expanded power inspection images are expanded images of scarce power inspection images; a benchmark large-scale model generation module for generating a power detection benchmark large-scale model based on power sample images in the power sample database and pre-training a pre-constructed power large-scale model; a business model construction module for adding an adjustable fusion reparameterization and adapter dual-branch structure to the power detection benchmark large-scale model to obtain a business model; and a business model training module for performing multiple rounds of coarse-grained training and fine-grained training on the business model to obtain a defect detection business large-scale model in a power scenario.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: at least one processor; and a memory communicatively connected to said at least one processor.

[0008] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the power equipment visual large model construction, training, and evaluation method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the method for constructing, training, and evaluating a large visual model of power equipment as described in any embodiment of the present invention.

[0010] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the method for constructing, training, and evaluating a large visual model of power equipment as described in any embodiment of the present invention.

[0011] The technical solutions of this invention propose a high-quality image screening method to achieve batch quantitative scoring of images and improve sample review efficiency. A five-dimensional information retrieval method for power vision samples is proposed, which manages a power sample database constructed from historical power inspection images and expanded power inspection images through binary retrieval tag feature indexing. This increases the number of scarce samples, resulting in a balanced distribution of sample image data in the database, laying a data foundation for model training and improving training effectiveness. A lightweight pre-training method for a large-scale power vision model is proposed, optimizing the network structure and designing a parameter freezing loss function to reduce training time. Based on the pre-trained benchmark large-scale power detection model, dual-branch structure parameter fine-tuning is performed. Through multiple rounds of coarse-grained and fine-grained training, the training efficiency of the large-scale defect detection business model is improved. Multi-model hybrid chain anomaly detection enhances the performance of the trained large-scale defect detection business model, making it more applicable to different power scenarios. Through quantitative testing of model effectiveness, a comprehensive evaluation of the model's application effect is achieved.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a method for constructing, training, and evaluating a large visual model of power equipment according to Embodiment 1 of the present invention.

[0015] Figure 2 This is a flowchart of a method for constructing, training, and evaluating a large visual model of power equipment according to Embodiment 2 of the present invention.

[0016] Figure 3 This is a flowchart of a method for constructing, training, and evaluating a large visual model of power equipment according to Embodiment 3 of the present invention.

[0017] Figure 4 This is a flowchart of a method for constructing, training, and evaluating a large visual model of power equipment according to Embodiment 4 of the present invention.

[0018] Figure 5 This is a schematic diagram of the structure of a power equipment visual large model construction, training and evaluation system provided in Embodiment 5 of the present invention.

[0019] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the method for constructing, training, and evaluating a large visual model of power equipment according to an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0021] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0022] Example 1

[0023] Figure 1 This invention provides a flowchart of a method for constructing, training, and evaluating a large visual model of power equipment, as described in Embodiment 1. This embodiment is applicable to the detection of power equipment. The method can be executed by a large visual model construction, training, and evaluation system for power equipment. This system can be implemented in hardware and / or software and can be configured in various general-purpose computing devices, such as client devices or server devices integrating a large visual model. Figure 1 As shown, the method includes...

[0024] S110. Construct a power sample database based on historical power inspection images and expanded power inspection images.

[0025] Among them, the expanded power inspection images can be expanded images of scarce power inspection images.

[0026] In this embodiment of the invention, power inspection images refer to images collected by personnel from power equipment that requires inspection and maintenance. Power inspection can refer to the regular or irregular inspection and maintenance of transmission lines, substations, and distribution equipment in a power system to ensure the safe and stable operation of the power system. Existing power inspections typically utilize drones equipped with various sensor devices, such as cameras and infrared cameras, to inspect the area to be inspected. For example, power inspection images can be acquired by controlling drones.

[0027] Specifically, a power sample database can be constructed using historically acquired power inspection images and expanded power inspection images to lay a data foundation for training a large-scale detection model for power equipment.

[0028] S120. Based on power sample images in the power sample database and pre-training a pre-built power large model, a power detection benchmark large model is generated.

[0029] The large-scale power model can be composed of an image sequence conversion layer, a feature learning layer, and a perception output layer. The image sequence conversion layer is used to convert the sample data input to the model into images. The feature learning layer is used to extract features from the converted image sample data. The perception output layer is used to integrate the features extracted by the feature learning layer and output the anomaly recognition results of the sample data. Optionally, the large-scale power model can be a visual large-scale model.

[0030] It should be noted that the large-scale power detection benchmark model can be fine-tuned on downstream tasks within specific power scenarios, and further trained using relevant task data from small-scale downstream tasks. The large-scale power equipment detection model can be used in various power scenarios. Power scenarios can encompass various facilities and environments related to power production, transmission, distribution, and use. These scenarios cover the entire power system from power plants to end users, including transmission, substation, and distribution scenarios. Transmission scenarios refer to the process of power being transmitted from power plants to substations via high-voltage transmission lines; substation scenarios mainly involve voltage transformation within the power system; and distribution scenarios refer to the process of power being distributed from substations to end users via distribution lines.

[0031] S130. Add an adjustable fusion reparameterization and adapter dual-branch structure to the power detection benchmark large model to obtain the business model.

[0032] Among them, the business model can refer to the model after fine-tuning the model parameters of the large power detection benchmark model and then retraining it based on the task data of downstream tasks in a specific power scenario.

[0033] In this embodiment of the invention, after determining the large power detection benchmark model, the model structure of the large power detection benchmark model can be improved by adding an adjustable fusion reparameterization and adapter dual-branch structure to obtain the business model.

[0034] Optionally, for the reparameterization branch, all parameters in the model structure of the power detection benchmark model are frozen. After the features are output from the perception output layer, the features are scaled using scaling and offset factors to enhance the expressive power of the original features. For the adapter branch, a network structure consisting of fully connected layers, multi-scale dilated convolutions, and activation functions is added after the multi-head attention mechanism. Then, the features output from the reparameterization branch and the adapter branch are added and fused to obtain fused features, which serve as the unified features for the next step. This retains the enhanced original features and increases the adaptability to downstream tasks, allowing it to be used as a business model for further targeted training.

[0035] Optionally, in this embodiment of the invention, the adapter structure in the dual-branch structure may include:

[0036] The first fully connected layer performs dimensionality reduction on the output features of the multi-head attention layer; a three-branch multi-scale dilated convolution kernel performs convolution on the output features of the fully connected layer; a 1×1 convolution kernel enhances the mean of the output features of the three-branch multi-scale dilated convolution kernel; and a second fully connected layer performs dimensionality increase on the output features of the 1×1 convolution kernel after they are activated.

[0037] Specifically, for the features output from the multi-head attention layer of the large-scale power detection benchmark model, a fully connected layer is used for dimensionality reduction. Then, the output features are convolved by a set of three-branch multi-scale dilated convolution kernels, each with a 3x3 convolution size and three different dilation factors of 1, 2, and 4, respectively. This increases the receptive field while capturing multi-scale information of the target. The average value of the features processed by the three-branch dilated convolution is calculated, and then fused with a 1x1 convolution kernel to enhance the features. Finally, after processing with a non-linear activation function, a fully connected layer is used to increase the dimensionality of the output features. The activation function can be ReLU, GeLU, or similar forms. To reduce feature loss caused by convolution computation, skip connections are used.

[0038] Based on the above embodiments, optionally, the adjustable fusion reparameterization structure in the dual-branch structure includes: scaling and shifting the output features of the multilayer perceptron layer sequentially based on preset scaling and shifting factors to obtain enhanced representation features.

[0039] Correspondingly, the output of the dual-branch structure is a fusion of the output features of the second fully connected layer and the enhanced expression features.

[0040] Specifically, if the features output by the sensing output layer in the large power detection benchmark model are represented by x, the scaling factor by γ, and the offset factor by β, then the formula for calculating the reparameterized branch output features can be: y = γ⊙x + β.

[0041] This involves performing a dot product using a scaling factor, followed by summation using an offset factor. Since this operation is a completely linear transformation, it can be reparameterized by absorbing the scaling and offset terms without adding any extra parameters or computational cost.

[0042] By adding a dual-branch structure of adjustable fusion reparameterization and adapter, the inefficiency and overfitting problems of full-parameter fine-tuning training are avoided, achieving efficient training of the business model and improving the fine-tuning efficiency and model performance of the business model in the power scenario.

[0043] S140. Perform multiple rounds of coarse-grained and fine-grained training on the business model to obtain a large-scale defect detection business model in the power scenario.

[0044] Specifically, the power sample images in the power sample database can be divided into a training set and a test set for the business model. First, the existing power sample images in the image database are used to perform the first coarse-grained training and fine-grained training on the business model. After training, the performance of the business model is tested, and new sample images are created based on the test results. Then, the new sample images are used to perform a second coarse-grained training and fine-grained training on the business model. The business model obtained from the second training is the final large-scale defect detection business model for the power scenario.

[0045] Optionally, the performance of the business model can be evaluated using three metrics: accuracy, FLOPs (floating-point operations), and Param (number of parameters). These metrics assess the model's recognition effectiveness after the initial coarse-grained training and subsequent fine-grained training. It's important to note that accuracy represents the proportion of the class with the highest predicted probability that matches the true label in the model's predictions for test samples; it focuses on whether the model correctly predicts the most likely class. FLOPs represent the number of floating-point operations required for a single forward propagation, measuring the model's computational cost, equivalent to its "time complexity." Param represents the total number of parameters that need to be trained, including all weights and biases, measuring the model's size, equivalent to its "space complexity."

[0046] By testing the performance of the business model, the problem of low availability of the business model in practical deployment in the power scenario was solved, the model performance of the business model was improved, and a comprehensive and fine-grained evaluation of the application effect of the business model was achieved.

[0047] The technical solution of this invention constructs a power sample database based on historical power inspection images and expanded power inspection images, thereby increasing the number of scarce samples and ensuring a balanced distribution of sample image data in the database. This lays a data foundation for model training, improving the training effect of the model. Furthermore, by conducting multiple rounds of coarse-grained and fine-grained training on the trained power detection benchmark model, the training efficiency of the defect detection business model is improved, and the performance of the trained defect detection business model is enhanced, making it more applicable to different power scenarios.

[0048] Example 2

[0049] Figure 2 This is a flowchart illustrating the construction, training, and evaluation of a large-scale visual model for power equipment, provided in Embodiment 2 of the present invention. This embodiment further refines the above embodiments, providing specific steps for performing multiple rounds of coarse-grained and fine-grained training on the business model to obtain a large-scale business model for defect detection in power scenarios. It should be noted that for parts not detailed in this embodiment, please refer to the relevant descriptions in other embodiments, which will not be repeated here. Figure 2 As shown, the method includes...

[0050] S210. The business model is trained in coarse-grained and fine-grained manner using power sample images from the power sample database.

[0051] S220. The normal sample images are identified and tested using the business model to obtain the false detection images in the normal sample images.

[0052] S230. Randomly stitch together the false detection image and the abnormal sample image into a new sample image.

[0053] S240. The business model is trained again with coarse-grained and fine-grained training through new sample images to obtain a large business model for defect detection in the power scenario.

[0054] False positive images can refer to images that do not contain defects but are incorrectly detected as having defects by the business model. Abnormal sample images refer to images that actually contain defects.

[0055] Specifically, typical model training only focuses on defect images, which can easily lead to false detections in practical applications. To reduce false detections, this invention tests the initially trained business model by inputting a large number of normal sample images and obtaining the model's output detection results. If a normal sample image is detected as defective and a defect type is output, it is identified as a false detection image. All false detection images are then processed, and abnormal sample images are randomly selected and scaled and stitched together with the false detection images to create a new image. For example, one abnormal sample image can be randomly stitched with three false detection images, or two abnormal sample images can be stitched with two false detection images, maintaining a reasonable aspect ratio for the overall new image. The labels on the new image are simultaneously calculated and adjusted to form a new sample image. This new sample image, along with existing sample images, forms a new sample image set. Based on this, the coarse and fine training phases of the model are repeated, and the optimal model is evaluated and selected as the final large-scale defect detection business model for the power industry scenario.

[0056] Based on the above embodiments, optionally, the business model can be subjected to coarse-grained training and fine-grained training sequentially using power sample images from the power sample database, including: cropping the power sample images according to at least one preset cropping method to obtain cropped sample images; performing coarse-grained training on the business model based on the cropped sample images; after coarse-grained training is completed, scaling the power sample images according to an acceptable training size for the business model to obtain scaled sample images; and performing fine-grained training on the business model based on the scaled sample images.

[0057] The cropping methods can include even cropping and center cropping.

[0058] Specifically, the coarse-grained training phase aims to improve the model's ability to focus on defect targets. First, the power sample images are cropped to ensure their size matches the acceptable training size for the business model. This invention employs a hybrid approach of equal-segment cropping and center cropping. Equal-segment cropping divides the sample image into N equal parts (typically four), while simultaneously labeling and transforming the corresponding coordinates. Center cropping involves cropping a portion of the sample image from its center, also labeling and transforming the coordinates, as equipment components in power inspection scenarios are often located at the image center. Then, the business model is coarse-grainedly trained using a target detection algorithm and the cropped sample images. High-performance target detection algorithms such as Yolov8 and RT-DETR can be employed.

[0059] The fine-grained training phase builds upon the coarse-grained training. Instead of cropping the power sample images, it scales them proportionally to the acceptable training size for the business model. Simultaneously, the labeled images undergo corresponding positional coordinate scaling transformations, creating a new set of sample images. Because this new training set uses scaled images, the targets in the images are smaller compared to the cropped images, increasing the training difficulty. However, since the business model trained in the coarse phase already possesses preliminary recognition capabilities, the difficulty of model training and optimization is significantly reduced, and it exhibits a certain degree of multi-scale target recognition capability across different inspection images. Furthermore, for defect types that are difficult to train on, such as those with varied shapes and small defect targets, the loss weight for this type is increased in the second training phase, enhancing the model's attention to and recognition ability for challenging defects.

[0060] The embodiments of this invention gradually optimize model performance using a coarse-to-fine cascade mode; at the same time, for the model's false detections and difficult defect types, a mixed positive and negative image sample iterative training strategy to remove false detections and a high-weight loss strategy for difficult types are proposed to further improve the model's fine-tuning efficiency.

[0061] Example 3

[0062] Figure 3 This is a flowchart illustrating the construction, training, and evaluation of a large-scale visual model for power equipment, provided in Embodiment 3 of the present invention. This embodiment further refines the above embodiments, providing specific steps for generating a large-scale benchmark model for power detection based on power sample images from a power sample database and pre-training a pre-constructed large-scale power model. It should be noted that for parts not detailed in this embodiment, please refer to the relevant descriptions in other embodiments, which will not be repeated here. Figure 3 As shown, the method includes...

[0063] S310. Input the power sample image into the image sequence conversion layer for image conversion processing.

[0064] The image conversion process may include: segmenting the power sample image and representing the segmented image blocks as a one-dimensional sequence vector.

[0065] Specifically, after the power sample image is input into the image sequence conversion layer, the layer performs image segmentation, dividing the power sample image into multiple image blocks of equal size. The segmented image blocks are then processed according to the diffusion equation to remove noise while preserving edge information. A nonlinear image pyramid is then constructed based on the processed image blocks; and image features within the nonlinear image pyramid are sampled and fused to convert the image blocks into one-dimensional sequence vectors. Optionally, to further reduce memory overhead, the one-dimensional sequence vectors can be binarized.

[0066] Optionally, in this embodiment of the invention, the segmented image blocks can also be masked, and the image mask region can be selected by a random algorithm to improve image quality.

[0067] S320. Input the power sample image after image conversion processing into the feature learning layer for feature extraction to determine the positional relationship between image blocks and the high-level image features of the image blocks.

[0068] The feature learning layer can be at least two visual converter modules connected in series. Optionally, the visual converter modules can include a multidimensional attention mechanism model and a feedforward neural network model.

[0069] Specifically, after inputting the converted power sample image into the feature learning layer, the multidimensional attention mechanism model can calculate the attention weight and feature subspace at each position in the input one-dimensional sequence vector to obtain the positional correlation between image patches; the feedforward neural network model can perform nonlinear feature transformation and mapping on the one-dimensional sequence vector to obtain high-level image features. Optionally, the feedforward neural network model may include fully connected layers and activation functions.

[0070] S330. Input the positional relationship between image patches and the high-level image features of the image patches into the perception output layer for feature integration to generate a predicted power image. Then, pre-train the large power model based on the predicted power image until the preset model pre-training termination condition is met.

[0071] Optionally, the model pre-training termination condition can be determined based on the image loss function between the power sample image input to the image sequence transformation layer and the predicted power image; the image loss function can be used to characterize the image difference between the power sample image and the predicted power image.

[0072] Optionally, the image loss function can be as follows:

[0073] Where N is the total number of pixels in the image patch, I i Let i be the value of the i-th pixel in the image patch of the power sample image. To predict the value of the i-th pixel in an image patch of a power image.

[0074] By iteratively training the model using the image loss function between the predicted power images and power sample images, the training effect of the model and the performance of the benchmark model for power detection are improved.

[0075] The technical solution of this invention provides a large model network structure that integrates convolutional neural networks and visual converters at the hierarchical level, and introduces deep convolutional layers for one-dimensional mapping of image slices. This enhances the large model's ability to extract features from sample data, improves the power large model's ability to recognize inspection images with complex outdoor backgrounds, and enhances the model's recognition capabilities.

[0076] Example 4

[0077] Figure 4 This is a flowchart illustrating the construction, training, and evaluation of a large-scale visual model of power equipment, provided in Embodiment 4 of the present invention. This embodiment further refines the above embodiments, providing specific steps for constructing a power sample database based on historical power inspection images and expanded power inspection images. It should be noted that for parts not detailed in this embodiment, please refer to the relevant descriptions in other embodiments, which will not be repeated here. Figure 4 As shown, the method includes...

[0078] S410. Using a pre-built pole prediction model, determine the power scene to which the historical power inspection images belong.

[0079] The first training set of the pole prediction model is obtained by data augmentation of historical power inspection images, and is used to predict the power scene to which the power inspection images belong. Data augmentation techniques such as flipping, scaling, cropping, and color transformation are used to process the historical power inspection images, and the processed image data is used as the first training set of the pole prediction model. Using data augmentation techniques to process historical power inspection images increases the sample size for training the pole prediction model, further improving its generalization ability.

[0080] Optionally, before determining the power scene to which the historical power inspection images belong, image preprocessing operations can be performed on the historical power images in advance. For example, an image edge extraction algorithm can be used to extract edge information from the historical power inspection images to obtain edge detail images; a pre-built lightweight image quality evaluation model can be used to evaluate the quality of the edge detail images and determine the quality score; historical power inspection images corresponding to edge images with quality scores higher than the quality score threshold are retained, while those with lower scores are discarded.

[0081] By evaluating the quality of historical power inspection images using a lightweight image quality assessment model, batch scoring and screening of historical power inspection images were achieved, thereby improving the sample quality of the subsequently generated power sample database.

[0082] S420. Label the defect images in the historical power inspection images under the power scenario to build an initial power sample database.

[0083] Specifically, defect sample images from historical power inspection images under different power scenarios are labeled. For example, vibration damper detachment in transmission images, silicone barrel damage in substation images, and flange rod corrosion in distribution images. It should be noted that the initial power sample database includes sample image sets from different power scenarios, such as transmission, substation, and distribution scenarios. It should also be noted that the corresponding defect conditions are different in different power scenarios. For example, vibration damper detachment in transmission scenario images, silicone barrel damage in substation scenario images, and flange rod corrosion in distribution scenario images. Performing different analyses for different power scenarios can reduce the computational load of subsequent detection model processing.

[0084] Optionally, the label naming method for power sample images in the initial power sample database can be based on the defect location information in the power sample image, the type of power equipment to which the power sample image belongs, and the type of defect present in the power sample image. For example, the label naming format can be "location information + equipment type + defect type".

[0085] Optionally, in this embodiment of the invention, search keywords can also be set for the initial sample database. Information such as the location of defects in the power sample image, the type of power equipment to which the power sample image belongs, the acquisition time of the power sample image, the type of defects in the power sample image, and the label specification version can be binarized and encoded into search keywords to quickly search and extract power sample images from different dimensions.

[0086] By designing a multi-dimensional sample retrieval strategy, lightweight and efficient retrieval and extraction of massive power sample images were achieved.

[0087] S430. Based on the power sample images in the initial power sample database, generate expanded power inspection images; and add the expanded power inspection images to the initial power sample database to generate the power sample database.

[0088] In this embodiment of the invention, expanded power inspection images can be generated from power sample images in the initial power sample database. The scarce sample images in the initial power sample database are expanded, and the expanded power inspection images are annotated in the same way as the power sample images before being added to the initial power sample database to generate a power sample database.

[0089] Optionally, based on power sample images in the initial power sample database, an expanded power inspection image is generated, including: identifying normal sample images and scarce sample images from historical power inspection images under the same power scenario, and using the normal sample images as background sample images; detecting and segmenting the scarce sample images using a dedicated detection and segmentation model to obtain scarce targets; detecting and segmenting the background sample images using a dedicated detection and segmentation model to obtain a background image; and pasting the scarce targets into the background image to obtain the expanded power inspection image.

[0090] Specifically, this can be achieved by selecting scarce class sample images with sparse sample distributions and using a dedicated segmentation model for segmentation prediction, and selecting normal sample images as the background sample set and using the same dedicated segmentation model for segmentation prediction. The former yields scarce targets, including scarce equipment and / or scarce defects, while the latter yields the background image. One scarce target is randomly selected, and a background image with a non-equipment area covering the special equipment is randomly selected. The scarce target is then pasted into a specific position on the background image to form a new expanded sample. This process is repeated multiple times to ultimately generate a large number of scarce samples and achieve class balance. Furthermore, before pasting the scarce target into the background image, it can be randomly processed, including random set transformations, random color transformations, and random noise addition, to enhance the richness of the scarce targets. For example, random set transformations include horizontal flipping, vertical flipping of special classes, and jitter scaling; random color transformations include brightness changes, contrast changes, and saturation changes; and random noise includes Gaussian noise, salt-and-pepper noise, Poisson noise, speckle noise, exponential noise, and uniform noise.

[0091] The technical solution of this invention forms a systematic data generation and augmentation for scarce samples, starting from the samples to ultimately achieve the diversity and distribution balance of large model algorithm samples, thereby improving the robustness of downstream models and increasing the target detection rate.

[0092] Optionally, the generation process of the dedicated detection and segmentation model may include: dividing the historical power inspection images into packets to obtain at least two packets; detecting the historical power inspection images in the first packet using at least two detection models to obtain a first detection result; the training frameworks of each detection model are different; inputting the first detection result and the historical power inspection images in the first packet into a general image segmentation model to obtain a first segmentation sample set; and iteratively training the general image segmentation model based on the first segmentation sample set to obtain the dedicated detection and segmentation model.

[0093] The detection targets of the detection model can include defects and devices in the image; the detection and segmentation targets of the dedicated detection and segmentation model can also include defects and devices in the image.

[0094] It should be noted that existing open-source segmentation and detection models are generally suitable for common targets, such as people, vehicles, and objects, but their detection and recognition performance is poor for industry-specific data. As an open-source model, the general image segmentation model (Segment Anything Model) cannot accurately detect and segment equipment and defects in UAV inspection images. It is necessary to first fine-tune and train the general image segmentation model based on historical power line inspection images to obtain a dedicated detection and segmentation model that can detect equipment and defects.

[0095] The dedicated detection and segmentation model is the core of subsequent sample expansion. Therefore, after obtaining the dedicated detection and segmentation model, it is necessary to compare its detection results with those of existing detection models on the same UAV inspection images, and then iteratively optimize it to obtain a dedicated detection and segmentation model with the expected detection accuracy.

[0096] During deep learning model training, various methods are employed to augment the training data. However, the number of augmented samples for scarce classes still falls short of the robustness requirements of the detection algorithm, resulting in low accuracy in detecting scarce samples. This invention, after iteration, uses a dedicated detection and segmentation model to segment scarce class samples (images), obtaining rare defects and / or rare devices within these samples. These are then pasted onto original sample images without rare defects and / or rare faults, resulting in augmented samples for the scarce class. By using a dedicated detection and segmentation model for efficient and accurate defect and device segmentation, the efficiency of sample augmentation can be improved, significantly increasing the number of augmented samples and thus meeting the robustness requirements of the algorithm.

[0097] In an optional embodiment of the present invention, historical power inspection images can be randomly divided into n packets D = {D1, D2, ..., Dn}. n}, n≥2, each sub-packet contains a number of power line inspection images. Select the power line inspection images in the first sub-packet D1, and use the existing marker detection model to detect the power line inspection images. The existing marker detection model includes two pre-trained detection models Ma and Mb with different frames.

[0098] If the same target appears at the same location in a single image and the confidence level is greater than or equal to the first threshold (e.g., 0.5), then this location is initially assumed to be the target and automatically labeled, forming an automatically labeled dataset B. If the same defect appears at the same location in a single image and the second threshold (e.g., 0.3) is less than the confidence level and less than the first threshold, then manual review and fine-tuning are performed, and the reviewed and fine-tuned samples are added to the labeled sample dataset A. If the same target appears at the same location in a single image and the confidence level is less than or equal to the second threshold, then it is placed in the inspection data package d. Datasets A and B are integrated, and the combined dataset is input into a general image segmentation model, prompting it to segment the image based on the labeled information. Then, the segmentation results are manually reviewed and fine-tuned to form the first segmentation sample set.

[0099] The system employs a hybrid detection approach combining multiple existing detection models with a confidence-based filtering strategy to automate target selection and labeling. It integrates a general segmentation model that requires only a single label to completely segment instances, eliminating the need for additional training and fine-tuning, and can be used on entirely new datasets. This semi-automated labeling system significantly reduces the workload of manual labeling. While some data still requires manual verification, this improves the overall efficiency of the data processing workflow. Later optimizations of the semi-automated labeling system on the segmentation dataset involve fine-tuning the general segmentation model using the segmentation dataset, resulting in an initial dedicated detection and segmentation model. Upgrading the semi-automated labeling system by combining existing detection models with the dedicated detection and segmentation model achieves multi-model complementarity, continuously iterating to improve the detection performance of the dedicated detection and segmentation model, and realizing the efficient development and processing of target detection and segmentation models and segmentation datasets.

[0100] Example 5

[0101] Figure 5 This is a schematic diagram of the structure of a large-scale visual model construction, training, and evaluation system for power equipment provided in Embodiment 5 of the present invention. Figure 5 As shown, the system includes...

[0102] The sample database construction module 510 is used to construct a power sample database based on historical power inspection images and expanded power inspection images; wherein, the expanded power inspection images are expanded images of scarce power inspection images.

[0103] The benchmark large model generation module 520 is used to generate a benchmark large model for power detection based on power sample images in the power sample database and pre-training a pre-built power large model.

[0104] The business model construction module 530 is used to add an adjustable fusion reparameterization and adapter dual-branch structure to the power detection benchmark large model to obtain the business model.

[0105] The business model training module 540 is used to perform multiple rounds of coarse-grained training and fine-grained training on the business model to obtain a large business model for defect detection in the power scenario.

[0106] The technical solution of this invention constructs a power sample database based on historical power inspection images and expanded power inspection images, thereby increasing the number of scarce samples and ensuring a balanced distribution of sample image data in the database. This lays a data foundation for model training, improving the training effect of the model. Furthermore, by conducting multiple rounds of coarse-grained and fine-grained training on the trained power detection benchmark model, the training efficiency of the defect detection business model is improved, and the performance of the trained defect detection business model is enhanced, making it more applicable to different power scenarios.

[0107] Optionally, the business model training module 540 includes...

[0108] The first training unit is used to perform coarse-grained training and fine-grained training on the business model using power sample images from the power sample database.

[0109] The model testing unit is used to perform recognition tests on normal sample images through the business model to obtain false detection images in the normal power inspection images.

[0110] The sample update unit randomly splices the falsely detected image and the abnormal sample image into a new sample image.

[0111] The second training unit is used to perform coarse-grained training and fine-grained training on the business model again using the new sample images to obtain a large-scale defect detection business model in the power scenario.

[0112] Optionally, the first training unit is specifically used to: crop the power sample image according to at least one preset cropping method to obtain a cropped sample image; perform coarse-grained training on the business model based on the cropped sample image; after the coarse-grained training is completed, scale the sample image according to the training size acceptable to the business model to obtain a scaled sample image; and perform fine-grained training on the business model based on the scaled sample image.

[0113] Optionally, the adapter structure in the dual-branch structure includes: a first fully connected layer that performs dimensionality reduction on the output features of the multi-head attention layer; a three-branch multi-scale dilated convolution kernel that performs convolution on the output features of the fully connected layer; a 1×1 convolution kernel that fuses and enhances the mean of the output features of the three-branch multi-scale dilated convolution kernel; and a second fully connected layer that performs dimensionality increase on the output features of the 1×1 convolution kernel after they are activated.

[0114] Optionally, the adjustable fusion reparameterization structure in the dual-branch structure includes: scaling and shifting the output features of the perceptual output layer sequentially based on preset scaling and shifting factors to obtain enhanced expression features; correspondingly, the output of the dual-branch structure is a fusion feature of the output features of the second fully connected layer and the enhanced expression features.

[0115] Optionally, the large power model includes an image sequence conversion layer, a feature learning layer, and a perception output layer.

[0116] Optionally, the benchmark large model generation module 520 includes...

[0117] The image conversion unit is used to input the power sample image into the image sequence conversion layer for image conversion processing; the image conversion processing includes: performing image segmentation on the power sample image and representing the segmented image blocks as a one-dimensional sequence vector.

[0118] The feature extraction unit is used to input the power sample image after image conversion processing into the feature learning layer for feature extraction operation, to determine the positional relationship between image blocks and the high-level image features of the image blocks; wherein, the feature learning layer consists of at least two visual converter modules connected in series;

[0119] The pre-training unit is used to input the positional relationships between image patches and the high-level image features of the image patches into the perception output layer for feature integration operations, generate a predicted power image, and pre-train the power model based on the predicted power image until the preset model pre-training termination condition is met.

[0120] Optionally, the model pre-training termination condition is determined based on the image loss function between the power sample image input to the image sequence conversion layer and the predicted power image; the image loss function is used to characterize the image difference between the power sample image and the predicted power image.

[0121] Optionally, the sample database construction module 510 includes...

[0122] The scene determination unit is used to determine the power scene to which the historical power inspection image belongs by using a pre-built tower prediction model; wherein the power scene is a transmission scene, a substation scene, and a distribution scene.

[0123] The initial sample library construction unit is used to annotate defect images in historical power inspection images under power scenarios to build an initial power sample database.

[0124] The sample library construction unit is used to generate expanded power inspection images based on the sample image data in the initial power sample database; and to add the expanded power inspection images to the initial power sample database to generate a power sample database.

[0125] Optionally, the sample library construction unit includes...

[0126] The background sample determination subunit is used to determine normal sample images and scarce sample images from historical power inspection images under the same power scenario, and to use the normal sample images as background sample images.

[0127] The scarce target determination subunit is used to detect and segment scarce class sample images using a dedicated detection and segmentation model to obtain scarce targets.

[0128] The background image determination sub-unit is used to detect and segment the background sample image using a dedicated detection and segmentation model to obtain the background image.

[0129] An expanded sample sub-unit is used to paste the scarce target onto the background image to obtain an expanded power inspection image.

[0130] Optionally, the process for generating a dedicated detection and segmentation model includes...

[0131] The historical power line inspection images are divided into at least two sub-packages.

[0132] The historical power inspection images in the first sub-package are detected by at least two detection models to obtain the first detection result; the training frameworks of each detection model are different.

[0133] The first detection result and the historical power inspection images in the first sub-package are input into a general image segmentation model to obtain the first segmentation sample set.

[0134] The general image segmentation model is iteratively trained based on the first segmentation sample set to obtain a dedicated detection and segmentation model.

[0135] The power equipment visual large model construction, training, and evaluation device provided in the embodiments of the present invention can execute the power equipment visual large model construction, training, and evaluation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0136] Example 6

[0137] Figure 6A schematic diagram of an electronic device 610 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0138] like Figure 6 As shown, the electronic device 610 includes at least one processor 611 and a memory, such as a read-only memory (ROM) 612 or a random access memory (RAM) 613, communicatively connected to the at least one processor 611. The memory stores computer programs executable by the at least one processor. The processor 611 can perform various appropriate actions and processes based on the computer program stored in the ROM 612 or loaded from storage unit 618 into the RAM 613. The RAM 613 may also store various programs and data required for the operation of the electronic device 610. The processor 611, ROM 612, and RAM 613 are interconnected via a bus 614. An input / output (I / O) interface 615 is also connected to the bus 614.

[0139] Multiple components in electronic device 610 are connected to I / O interface 615, including: input unit 616, such as keyboard, mouse, etc.; output unit 617, such as various types of displays, speakers, etc.; storage unit 618, such as disk, optical disk, etc.; and communication unit 619, such as network card, modem, wireless transceiver, etc. Communication unit 619 allows electronic device 610 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] Processor 611 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 611 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 611 performs the various methods and processes described above, such as methods for building, training, and evaluating large-scale visual models of power equipment.

[0141] In some embodiments, the method for constructing, training, and evaluating a large visual model of power equipment can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 618. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 610 via ROM 612 and / or communication unit 619. When the computer program is loaded into RAM 613 and executed by processor 611, one or more steps of the method for constructing, training, and evaluating a large visual model of power equipment described above can be performed. Alternatively, in other embodiments, processor 611 can be configured to perform the method for constructing, training, and evaluating a large visual model of power equipment by any other suitable means (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0146] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0147] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0148] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for constructing, training, and evaluating a large visual model of power equipment, characterized in that, include: A power sample database is constructed based on historical power inspection images and expanded power inspection images; wherein, the expanded power inspection images are expanded images of scarce power inspection images. Based on power sample images in the power sample database and pre-training a pre-built power large model, a power detection benchmark large model is generated. A business model is obtained by adding a dual-branch structure of adjustable fusion reparameterization and adapter to the large power detection benchmark model; wherein, the adapter structure is a network structure consisting of fully connected layers, multi-scale dilated convolutions and activation functions added after the multi-head attention mechanism; The business model is subjected to multiple rounds of coarse-grained training and fine-grained training to obtain a large-scale defect detection business model in the power scenario.

2. The method according to claim 1, characterized in that, The business model is subjected to multiple rounds of coarse-grained and fine-grained training to obtain a large-scale defect detection business model in the power sector, including: The business model is trained using coarse-grained and fine-grained methods, respectively, by using power sample images from the power sample database. The business model is used to identify and test normal sample images to obtain false detection images in the normal sample images; The falsely detected images and abnormal sample images are randomly stitched together to form a new sample image; The business model is then trained again using the new sample images, followed by coarse-grained training and then fine-grained training, to obtain a large-scale defect detection business model for the power sector.

3. The method according to claim 2, characterized in that, The business model is trained sequentially using coarse-grained and fine-grained training on power sample images from the power sample database, including: The power sample image is cropped according to at least one preset cropping method to obtain the cropped sample image; The business model is trained using the cropped sample images; After coarse-grained training is completed, the sample images are scaled according to the training size acceptable to the business model to obtain scaled sample images. The business model is trained in a fine-grained manner based on the scaled sample images.

4. The method according to claim 1, characterized in that, The adapter structure in the dual-branch structure includes: The first fully connected layer that performs dimensionality reduction on the output features of the multi-head attention layer; A three-branch multi-scale dilated convolution kernel that performs convolution operations on the output features of a fully connected layer; A 1×1 convolution kernel that fuses and enhances the mean of the output features of a three-branch multi-scale dilated convolution kernel; The second fully connected layer performs dimensionality upscaling after the output features of the 1×1 convolution kernel are activated.

5. The method according to claim 4, characterized in that, The tunable fusion reparameterization structure in the dual-branch structure includes: Based on preset scaling and offset factors, the output features of the perception output layer are scaled and offset sequentially to obtain enhanced expression features; Correspondingly, the output of the dual-branch structure is a fusion of the output features of the second fully connected layer and the enhanced expression features.

6. The method according to claim 1, characterized in that, The large-scale power model includes an image sequence conversion layer, a feature learning layer, and a perception output layer. Accordingly, based on power sample images in the power sample database and pre-training a pre-built large power model, a large power detection benchmark model is generated, including: The power sample image is input into the image sequence conversion layer for image conversion processing; the image conversion processing includes: segmenting the power sample image and representing the segmented image blocks as one-dimensional sequence vectors; The power sample image after image conversion is input into the feature learning layer for feature extraction to determine the positional relationship between image blocks and the high-level image features of the image blocks; wherein, the feature learning layer consists of at least two visual converter modules connected in series. The positional relationships between image patches and the high-level image features of the image patches are input to the perception output layer for feature integration to generate a predicted power image. The power model is then pre-trained based on the predicted power image until the preset model pre-training termination condition is met.

7. The method according to claim 6, characterized in that, The model pre-training termination condition is determined based on the image loss function between the power sample image input to the image sequence conversion layer and the predicted power image; the image loss function is used to characterize the image difference between the power sample image and the predicted power image.

8. The method according to claim 1, characterized in that, A power sample database was constructed based on historical power inspection images and expanded power inspection images, including: A pre-built pole prediction model is used to determine the power scene to which the historical power inspection images belong; wherein, the power scene is a transmission scene, a substation scene, and a distribution scene; Defect sample images from historical power inspection images in power scenarios were labeled to construct an initial power sample database. Based on the power sample images in the initial power sample database, expanded power inspection images are generated; and the expanded power inspection images are added to the initial power sample database to generate a power sample database.

9. The method according to claim 8, characterized in that, Based on the power sample images in the initial power sample database, expanded power inspection images are generated, including: Normal sample images and scarce sample images are identified from historical power inspection images under the same power scenario, and the normal sample images are used as background sample images. The scarce target is obtained by detecting and segmenting the scarce class sample image through a dedicated detection and segmentation model; The background sample image is detected and segmented using a dedicated detection and segmentation model to obtain the background image; The scarce target is pasted onto the background image to obtain an expanded power inspection image.

10. The method according to claim 9, characterized in that, The generation process of the dedicated detection and segmentation model includes: The historical power line inspection images are divided into at least two sub-packages; The historical power inspection images in the first sub-package are detected by at least two detection models to obtain the first detection result; the training frameworks of each detection model are different. The first detection result and the historical power inspection images in the first sub-package are input into a general image segmentation model to obtain the first segmentation sample set; The general image segmentation model is iteratively trained based on the first segmentation sample set to obtain a dedicated detection and segmentation model.

11. A system for constructing, training, and evaluating a large visual model of power equipment, characterized in that, include: The sample database construction module is used to construct a power sample database based on historical power inspection images and expanded power inspection images; wherein, the expanded power inspection images are expanded images of scarce power inspection images. The benchmark large model generation module is used to generate a benchmark large model for power detection based on power sample images in the power sample database and pre-training a pre-built power large model. The business model construction module is used to add a dual-branch structure of adjustable fusion reparameterization and adapter to the large power detection benchmark model to obtain the business model; wherein, the adapter structure is a network structure consisting of fully connected layers, multi-scale dilated convolutions and activation functions added after the multi-head attention mechanism; The business model training module is used to perform multiple rounds of coarse-grained training and fine-grained training on the business model to obtain a large-scale business model for defect detection in the power scenario.

12. The system according to claim 11, characterized in that, The business model training module includes: The first training unit is used to perform coarse-grained training and fine-grained training on the business model sequentially using power sample images from the power sample database. The model testing unit is used to perform recognition tests on normal sample images using the business model, and to obtain false detection images in the normal sample images; The sample update unit randomly splices the falsely detected image and the abnormal sample image into a new sample image; The second training unit is used to perform coarse-grained training and fine-grained training on the business model again using the new sample images to obtain a large-scale defect detection business model in the power scenario.

13. The system according to claim 12, characterized in that, The first training unit is specifically used for: The power sample image is cropped according to at least one preset cropping method to obtain the cropped sample image; The business model is trained using the cropped sample images; After coarse-grained training is completed, the sample images are scaled according to the training size acceptable to the business model to obtain scaled sample images. The business model is trained in a fine-grained manner based on the scaled sample images.

14. The system according to claim 11, characterized in that, The adapter structure in the dual-branch structure includes: The first fully connected layer that performs dimensionality reduction on the output features of the multi-head attention layer; A three-branch multi-scale dilated convolution kernel that performs convolution operations on the output features of a fully connected layer; A 1×1 convolution kernel that fuses and enhances the mean of the output features of a three-branch multi-scale dilated convolution kernel; The second fully connected layer performs dimensionality upscaling after the output features of the 1×1 convolution kernel are activated.

15. The system according to claim 11, characterized in that, The large-scale power model includes an image sequence conversion layer, a feature learning layer, and a perception output layer.

16. The system according to claim 15, characterized in that, The benchmark large model generation module includes: An image conversion unit is used to input an electric power sample image into an image sequence conversion layer for image conversion processing; the image conversion processing includes: segmenting the electric power sample image and representing the segmented image blocks as a one-dimensional sequence vector; The feature extraction unit is used to input the power sample image after image conversion processing into the feature learning layer for feature extraction operation, to determine the positional relationship between image blocks and the high-level image features of the image blocks; wherein, the feature learning layer consists of at least two visual converter modules connected in series; The pre-training unit is used to input the positional relationships between image patches and the high-level image features of the image patches into the perception output layer for feature integration operations, generate a predicted power image, and pre-train the power model based on the predicted power image until the preset model pre-training termination condition is met.

17. The system according to claim 16, characterized in that, The model pre-training termination condition is determined based on the image loss function between the power sample image input to the image sequence conversion layer and the predicted power image; the image loss function is used to characterize the image difference between the power sample image and the predicted power image.

18. The system according to claim 11, characterized in that, The sample database construction module includes: The scene determination unit is used to determine the power scene to which the historical power inspection image belongs by using a pre-built tower prediction model; wherein the power scene is a transmission scene, a substation scene, and a distribution scene; The initial sample library construction unit is used to annotate defect images in historical power inspection images under power scenarios and build an initial power sample database. The sample library construction unit is used to generate expanded power inspection images based on the sample image data in the initial power sample database; and to add the expanded power inspection images to the initial power sample database to generate a power sample database.

19. The system according to claim 18, characterized in that, The sample library construction unit includes: The background sample determination subunit is used to determine normal sample images and scarce sample images from historical power inspection images under the same power scenario, and to use the normal sample images as background sample images. The scarce target determination subunit is used to detect and segment scarce class sample images using a dedicated detection and segmentation model to obtain scarce targets; The background image determination sub-unit is used to detect and segment the background sample image using a dedicated detection and segmentation model to obtain the background image; An expanded sample sub-unit is used to paste the scarce target onto the background image to obtain an expanded power inspection image.

20. The system according to claim 19, characterized in that, The generation process of the dedicated detection and segmentation model includes: The historical power line inspection images are divided into at least two sub-packages; The historical power inspection images in the first sub-package are detected by at least two detection models to obtain the first detection result; the training frameworks of each detection model are different. The first detection result and the historical power inspection images in the first sub-package are input into a general image segmentation model to obtain the first segmentation sample set; The general image segmentation model is iteratively trained based on the first segmentation sample set to obtain a dedicated detection and segmentation model.

21. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the method for constructing, training, and evaluating a large visual model of power equipment as described in any one of claims 1-10.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method for constructing, training, and evaluating a large visual model of power equipment as described in any one of claims 1-10.

23. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor, implements the method for constructing, training, and evaluating a large visual model of power equipment as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Power distribution equipment infrared image fault identification method and system

    CN113688947A

  • Efficient road panorama detection method based on model fusion

    CN114882239A