Method and device for generating multi-target equipment detection model of transformer substation and computer equipment
By dividing feature maps in the multi-target equipment detection of substations and performing downsampling, nearest neighbor interpolation processing and model pruning, an efficient detection model adapted to terminal equipment memory is generated, which solves the problem of low detection efficiency caused by the large storage space of traditional models, and achieves efficient and accurate equipment detection.
Patent Information
- Application Number
- CN202510348721.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional neural network architecture occupies a large amount of storage space in the detection of multi-target equipment in substations, and cannot adapt to the memory space requirements of terminal devices, resulting in low detection efficiency.
By inputting the device image to the input layer of the object detection model, the feature map is divided into the first and second feature map sets, downsampling and nearest neighbor interpolation processing are performed, the model is trained based on the device image detection results and labeling information, and pruning is performed to generate the substation multi-objective device detection model.
Reduce computing resource consumption, improve device detection efficiency and accuracy, adapt to the memory space requirements of terminal equipment, and enhance device identification and detection capabilities.
Smart Images

Figure CN120259689A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of equipment detection, and particularly to a method, device, computer equipment, computer-readable storage medium, and computer program product for generating a substation multi-target equipment detection model. Background Art
[0002] As a core component of the power system, a substation undertakes important energy conversion and distribution functions. With the development of deep learning and artificial intelligence technologies, intelligent substation target recognition and intelligent condition diagnosis have gradually become one of the achievable means, and mobile terminal devices such as unmanned aerial vehicles, inspection robots, and augmented reality devices have gradually been popularized and applied in substations. In this context, it is becoming increasingly important to solve the problem of multi-target detection of mobile terminal devices in substations.
[0003] Traditional technologies mainly use neural network architectures with networks such as ResNet as the backbone network for substation equipment recognition and detection. However, the neural network architectures used in traditional technologies will occupy more storage space and cannot meet the memory space requirements of the terminal devices mainly used in substation equipment detection, which is not conducive to improving the detection efficiency of substation multi-target equipment detection. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium, and computer program product for generating a substation multi-target equipment detection model that can improve the detection efficiency of substation multi-target equipment detection in view of the above technical problems.
[0005] In a first aspect, this application provides a method for generating a substation multi-target equipment detection model, including:
[0006] Inputting the equipment image into the input layer of the target detection model to be trained to obtain a feature map, and dividing the feature map into a first feature map set and a second feature map set according to a preset division method;
[0007] Inputting the first feature map set into the downsampling module in the target detection model to perform downsampling processing on the first feature map set to obtain a downsampling result, and inputting the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result;
[0008] Determining the equipment image detection result according to the second feature map set and the upsampling result; the equipment image detection result represents the type and position of the equipment in the equipment image;
[0009] Train the object detection model according to the device image detection results and the device annotation information pre-annotated for the device image, and prune the object detection model to obtain a substation multi-object device detection model; the substation multi-object device detection model is used to detect and locate devices in the image.
[0010] In one embodiment, the determining the device image detection results according to the second feature map set and the upsampling result includes:
[0011] Determine a target feature map according to the second feature map set and the upsampling result; the target feature map characterizes the edge features and texture features of the device.
[0012] Divide the target feature map into at least one unit by the object detection model, and determine the detection results corresponding to each unit; the detection results include the device bounding boxes and device types corresponding to each unit.
[0013] Determine the device image detection results according to the detection results corresponding to each unit.
[0014] In one embodiment, the object detection model includes a convolutional layer, and the convolutional layer includes the downsampling module and the upsampling module. The pruning of the object detection model includes:
[0015] When there is a batch normalization layer after the convolutional layer, input the device image into the object detection model, and adjust the reduction coefficients of the input channels of the object detection model according to the output result of the object detection model and the device annotation information pre-annotated for the device image.
[0016] Determine the input channels to be deleted according to the reduction coefficients of the input channels of the object detection model.
[0017] Delete the input channels to be deleted and the corresponding reduction coefficients of the input channels to be deleted from the object detection model.
[0018] In one embodiment, after the step of obtaining the substation multi-object device detection model, the method further includes:
[0019] Adjust the substation multi-object device detection model according to the device image and the device annotation information pre-annotated for the device image to obtain a fine-tuned substation multi-object device detection model.
[0020] Input a preset validation set into the fine-tuned substation multi-object device detection model to obtain precision information, recall information, and accuracy information.
[0021] Determine the model performance evaluation result according to the precision information, the recall information, and the accuracy information, and adjust the fine-tuned substation multi-object device detection model according to the model performance evaluation result.
[0022] In one embodiment, adjusting the substation multi-object device detection model according to the device image and the device annotation information pre-annotated for the device image includes:
[0023] Determine a target model structure from the substation multi-object device detection model, and input the device image into the substation multi-object device detection model while fixing the model parameters of the target model structure;
[0024] Adjust the model parameters of the model structure other than the target model structure in the substation multi-object device detection model according to the output result of the substation multi-object device detection model and the device annotation information pre-annotated for the device image.
[0025] In one embodiment, inputting the device image into the input layer of the target detection model to be trained to obtain a feature map includes:
[0026] Preprocess the device image through the input layer to obtain a processed image;
[0027] Obtain the pixel information of the processed image, and perform convolution processing on the pixel information to obtain the feature map.
[0028] In a second aspect, the present application further provides a substation multi-object device detection model generation device, including:
[0029] A sampling module, configured to input the first feature map set into the downsampling module in the target detection model, perform downsampling processing on the first feature map set to obtain a downsampling result, and input the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result;
[0030] An identification module, configured to determine a device image detection result according to the second feature map set and the upsampling result; the device image detection result characterizes the type and position of the device in the device image;
[0031] A generation module, configured to train the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and perform pruning on the target detection model to obtain a substation multi-object device detection model; the substation multi-object device detection model is used to detect and locate devices in an image.
[0032] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the steps of the above method are implemented.
[0033] In a fourth aspect, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0034] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0035] For the above method, device, computer device, computer-readable storage medium, and computer program product for generating a substation multi-target device detection model, by inputting a device image into an input layer of a target detection model to be trained, a feature map is obtained, and according to a preset partitioning method, the feature map is partitioned into a first feature map set and a second feature map set, so as to partition the feature map into a directly transmitted part and a part that needs to be convolutionally processed, facilitating the subsequent model to learn feature representations at different levels and reducing computational resource consumption; inputting the first feature map set into a downsampling module in the target detection model to perform downsampling processing on the first feature map set to obtain a downsampling result, and inputting the downsampling result set into an upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result, thereby fusing multi-scale feature map information, enabling the model to accurately detect and identify devices in the device image by combining multi-scale feature map information and improving detection accuracy; determining a device image detection result according to the second feature map set and the upsampling result; the device image detection result represents the type and location of the devices in the device image, enabling the model to learn feature representations at different levels, reducing computational resource consumption, and improving the efficiency and accuracy of device detection; training the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and pruning the target detection model to obtain a substation multi-target device detection model; the substation multi-target device detection model is used to detect and locate devices in the image, thereby pruning the model after training the target detection model, reducing the model's computational amount, lowering the computational cost and memory occupancy, facilitating the deployment and application of the model on resource-constrained devices, making the device recognition and detection more adaptable to the memory space requirements of the terminal devices mainly used in substation device detection, and further improving the detection efficiency of substation multi-target device detection. Description of the Drawings
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0037] Figure 1 It is an application environment diagram of a method for generating a substation multi-target device detection model in an embodiment;
[0038] Figure 2 It is a flowchart of a method for generating a substation multi-target device detection model in an embodiment;
[0039] Figure 3 It is a flowchart of substation multi-target detection based on the YOLOv5 algorithm in an embodiment;
[0040] Figure 4 It is a schematic diagram of model pruning in an embodiment;
[0041] Figure 5 It is a schematic diagram of the architecture of a substation multi-target device detection model in an embodiment;
[0042] Figure 6 It is a structural block diagram of a device for generating a substation multi-target device detection model in an embodiment;
[0043] Figure 7 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0044] In order to make the objectives, technical solutions and advantages of the present application more clear, the following further details the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0045] The method for generating a substation multi-target device detection model provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The terminal 102 inputs the device image into the input layer of the target detection model to be trained, obtains the feature map, and divides the feature map into a first feature map set and a second feature map set according to a preset division method; the terminal 102 inputs the first feature map set into the downsampling module in the target detection model, performs downsampling processing on the first feature map set to obtain the downsampling result, and inputs the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain the upsampling result; the terminal 102 determines the device image detection result according to the second feature map set and the upsampling result; the device image detection result represents the type and position of the device in the device image; the terminal 102 trains the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and prunes the target detection model to obtain a substation multi-target device detection model; the substation multi-target device detection model is used to detect and locate the devices in the image. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0046] In an exemplary embodiment, as Figure 2 shown, a method for generating a substation multi-target device detection model is provided. Taking this method applied to the terminal as an example, it includes the following steps S202 to step S208. Among them:
[0047] Step S202, input the device image into the input layer of the target detection model to be trained, obtain the feature map, and divide the feature map into a first feature map set and a second feature map set according to a preset division method.
[0048] Among them, the device images can refer to images of devices (such as main transformers, lightning arresters, disconnectors, and reactors, etc.) in the substation collected at different angles, different brightness levels, and different scenes inside the substation. In practical applications, the device images can be divided into a training data set, a validation data set, and a test data set.
[0049] Among them, the object detection model can include the YOLOv5 model.
[0050] Among them, the feature map can refer to the feature representation extracted from the device image through convolution operations.
[0051] Among them, the first set of feature maps can refer to the feature maps that need to be convolved.
[0052] Among them, the second set of feature maps can refer to the feature maps that are directly passed without convolution processing.
[0053] As an example, the terminal can input the device image into the input layer of the object detection model to be trained. After the input layer performs convolution processing on the device image to obtain a feature map, the terminal can then divide the feature map into a first set of feature maps and a second set of feature maps according to a preset division method. Among them, the first set of feature maps needs to be convolved again, and the second set of feature maps is directly passed to the positioning and recognition process without convolution processing. The model can combine the results after convolution processing of the first set of feature maps and the second set of feature maps for object positioning and recognition, reducing redundant calculation steps.
[0054] Step S204: Input the first set of feature maps into the downsampling module in the object detection model, perform downsampling processing on the first set of feature maps to obtain a downsampling result, and input the downsampling result set into the upsampling module in the object detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result.
[0055] Among them, the downsampling result can refer to the information obtained by performing downsampling processing on the first set of feature maps through the downsampling module.
[0056] Among them, the upsampling result can refer to the information obtained by performing nearest neighbor interpolation processing on the downsampling result through the upsampling module.
[0057] As an example, the terminal can input the first set of feature maps into the downsampling module in the object detection model. The downsampling module (such as Feature Pyramid Network) can perform downsampling processing (such as convolution processing and pooling processing) on the first set of feature maps to obtain a downsampling result. In practical applications, the downsampling module can gradually downsample the feature maps from the shallow layer to the deep layer, gradually disassemble the information of the feature maps, reduce the resolution step by step, fuse the low-level feature maps with the high-level feature maps, and enhance the detailed information contained in the high-level feature maps. Then the terminal can input the set of downsampling results into the upsampling module in the object detection model. The upsampling module (such as Path Aggregation Network, a neural network architecture designed to improve object detection performance) can perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result. In practical applications, the upsampling module can gradually upsample the feature maps from the deep layer to the shallow layer, gradually increase the resolution of the feature values through nearest neighbor interpolation, enable the low-level feature maps to obtain high-level semantic information, and generate multi-scale features. In specific implementation, the downsampling module and the upsampling module can be used as a set of sampling modules. The object detection model can include several sets of sampling modules. The sampling modules can implement bidirectional sampling, so that the feature maps not only retain the low-level position information but also contain high-level semantic information, improving the accuracy of multi-object detection in the substation.
[0058] Step S206: Determine the device image detection result according to the second set of feature maps and the upsampling result.
[0059] Among them, the device image detection result can refer to the information indicating the type and position of the device in the device image.
[0060] As an example, the terminal can merge the second set of feature maps and the upsampling result. The merged result can be used as the target feature map. The terminal can analyze the edge features and texture features of the device in combination with the target feature map, perform target localization based on the target feature map, analyze the type and position of the device in the device image, and obtain the device image detection result.
[0061] Step S208: Train the object detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and prune the object detection model to obtain a multi-object device detection model for the substation.
[0062] Among them, the device annotation information can be pre-annotated by the user. The device annotation information can be used as label information to represent the real distribution of the devices in the device image.
[0063] Among them, the multi-object device detection model for the substation can refer to a model used to detect and locate the devices in the image.
[0064] As an example, the terminal can analyze the degree of difference between the device image detection result and the device annotation information pre-annotated for the device image, and train the object detection model according to the device image detection result and the device annotation information pre-annotated for the device image. For example, the terminal can adjust the model parameters of the object detection model according to the degree of difference between the device image detection result and the device annotation information pre-annotated for the device image until the degree of difference between the device image detection result and the device annotation information pre-annotated for the device image meets the preset difference requirement. At this time, the terminal can determine that the object detection model training is completed. The terminal can also prune the trained object detection model, streamline the model structure, and obtain a substation multi-object device detection model, thereby reducing the model calculation amount, reducing the calculation cost and memory occupancy, and facilitating the deployment and application of the model on resource-constrained devices.
[0065] In the above method for generating a substation multi-object device detection model, by inputting a device image into the input layer of a to-be-trained object detection model, a feature map is obtained, and according to a preset partitioning method, the feature map is partitioned into a first feature map set and a second feature map set, so as to partition the feature map into a directly transmitted part and a part that needs to be processed by convolution, so that the subsequent model can learn feature representations at different levels and reduce the consumption of computing resources; inputting the first feature map set into the downsampling module in the object detection model to perform downsampling processing on the first feature map set to obtain a downsampling result, and inputting the downsampling result set into the upsampling module in the object detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result, so as to fuse the feature map information of multiple scales, so that the model can combine the feature map information of multiple scales to accurately detect and identify the devices in the device image and improve the detection accuracy; determining the device image detection result according to the second feature map set and the upsampling result; the device image detection result represents the type and position of the devices in the device image, so that the model can learn feature representations at different levels, reduce the consumption of computing resources, and improve the efficiency and accuracy of device detection; training the object detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and pruning the object detection model to obtain a substation multi-object device detection model; the substation multi-object device detection model is used to detect and locate the devices in the image, so as to prune the model after training the object detection model, reduce the model calculation amount, reduce the calculation cost and memory occupancy, and facilitate the deployment and application of the model on resource-constrained devices, so that the device recognition and detection are more adapted to the memory space requirements of the terminal devices mainly used in substation device detection, and further improve the detection efficiency of substation multi-object device detection.
[0066] In an exemplary embodiment, determining a device image detection result based on a second set of feature maps and an upsampling result includes: determining a target feature map based on the second set of feature maps and the upsampling result; the target feature map characterizes the edge features and texture features of the device; dividing the target feature map into at least one unit through a target detection model, and determining the detection result corresponding to each unit; the detection result includes the device bounding box and device type corresponding to each unit; and determining the device image detection result based on the detection result corresponding to each unit.
[0067] Among them, the target feature map may refer to the information obtained by merging the second set of feature maps and the upsampling result. In practical applications, the merging methods when merging the second set of feature maps and the upsampling result may include splicing, adding, weighted fusion, etc.
[0068] Among them, the unit may refer to the information obtained by dividing the target feature map according to a preset grid size.
[0069] Among them, the device bounding box may refer to a rectangular box used to identify the position of the device in the unit.
[0070] As an example, the terminal may merge the second set of feature maps and the upsampling result according to a preset merging method to obtain a target feature map, which can characterize the edge features and texture features of the device. Then, the target detection model may divide the target feature map into at least one unit, perform target recognition on each unit, determine the device bounding box and device type of the device in each unit, and obtain the detection result corresponding to each unit. After that, the terminal may combine the detection results corresponding to each unit to determine the device image detection result corresponding to the device image.
[0071] In this embodiment, by determining a target feature map based on the second set of feature maps and the upsampling result; the target feature map characterizes the edge features and texture features of the device; dividing the target feature map into at least one unit through a target detection model, and determining the detection result corresponding to each unit; the detection result includes the device bounding box and device type corresponding to each unit; and determining the device image detection result based on the detection result corresponding to each unit, it is possible to perform target localization and recognition by combining the upsampling result obtained after convolution processing of the first set of feature maps and the second set of feature maps, reduce redundant calculation steps, enhance the feature extraction ability of the target detection model, and thus improve the detection efficiency of multi-target device detection in a substation.
[0072] In some embodiments, the object detection model includes a convolutional layer, and the convolutional layer includes a downsampling module and an upsampling module. Pruning the object detection model includes: when there is a batch normalization layer after the convolutional layer, inputting a device image into the object detection model, and adjusting the reduction coefficients of the input channels of the object detection model according to the output result of the object detection model and the device annotation information pre-annotated for the device image; determining the input channels to be deleted according to the reduction coefficients of the input channels of the object detection model; and deleting the input channels to be deleted and the corresponding reduction coefficients of the input channels to be deleted from the object detection model.
[0073] Among them, the batch normalization layer may refer to the BN (Batch Normalization) layer. In practical applications, the batch normalization layer can reduce the internal covariate shift by normalizing the data of each batch.
[0074] As an example, when there is a batch normalization layer after the convolutional layer, the terminal can set the BN layer parameters corresponding to the batch normalization layer and the reduction coefficients of the input channels of the object detection model (such as ). After the terminal inputs the device image into the object detection model, the output result of the object detection model can represent the device image detection result of the device image. At this time, the terminal can analyze and adjust the reduction coefficients of the input channels of the object detection model according to the output result of the object detection model and the device annotation information pre-annotated for the device image. Then, the terminal can analyze the input channels that need to be deleted in the input channels of the object detection model according to the reduction coefficients of the input channels of the object detection model, determine the input channels to be deleted, and delete the input channels to be deleted and the corresponding reduction coefficients of the input channels to be deleted from the object detection model, so as to achieve pruning. After pruning, the model structure of the object detection model changes, and the BN layer parameters also change.
[0075] In this embodiment, when there is a batch normalization layer after the convolutional layer, by inputting a device image into the object detection model, and adjusting the reduction coefficients of the input channels of the object detection model according to the output result of the object detection model and the device annotation information pre-annotated for the device image; determining the input channels to be deleted according to the reduction coefficients of the input channels of the object detection model; and deleting the input channels to be deleted and the corresponding reduction coefficients of the input channels to be deleted from the object detection model, the model can be pruned and compressed, the model structure can be streamlined, the model reduction of the object detection network can be realized, and the substation equipment detection speed can be improved.
[0076] In some embodiments, after the step of obtaining the substation multi-target device detection model, the above method further includes: adjusting the substation multi-target device detection model according to the device image and the device annotation information pre-annotated for the device image to obtain a fine-tuned substation multi-target device detection model; inputting a preset validation set into the fine-tuned substation multi-target device detection model to obtain precision information, recall information, and accuracy information; determining a model performance evaluation result according to the precision information, recall information, and accuracy information, and adjusting the fine-tuned substation multi-target device detection model according to the model performance evaluation result.
[0077] Among them, the fine-tuned substation multi-target device detection model may refer to the model obtained after adjusting the target device detection model.
[0078] Among them, the model performance evaluation result may refer to the information characterizing the model performance (such as device detection accuracy) of the fine-tuned substation multi-target device detection model.
[0079] As an example, the terminal can fine-tune the substation multi-target device detection model. For example, the terminal can adjust the substation multi-target device detection model according to the device image and the device annotation information pre-annotated for the device image to obtain a fine-tuned substation multi-target device detection model. Then the terminal can input a preset validation set into the fine-tuned substation multi-target device detection model, analyze the difference between the output result of the fine-tuned substation multi-target device detection model and the label of the data in the validation set, and determine the precision information, recall information, and accuracy information. After that, the terminal can determine the model performance evaluation result according to the precision information, recall information, and accuracy information, and adjust the fine-tuned substation multi-target device detection model according to the model performance evaluation result.
[0080] In practical applications, the calculation expression of the precision information can be expressed as:
[0081] .
[0082] Among them, can refer to the precision information, can refer to the number of samples (i.e., true positives) where the output result of the fine-tuned substation multi-target device detection model represents that the device type is the preset type and the device type in the device image is also the preset type, can refer to the number of samples (i.e., false positives) where the output result of the fine-tuned substation multi-target device detection model represents that the device type is the preset type and the device type in the device image is different from the preset type.
[0083] The calculation expression of the recall information can be expressed as:
[0084] 。
[0085] Among them, it can refer to recall information, and it can also refer to the number of samples (i.e., false negatives) in the output result of the multi-target device detection model for the substation after fine-tuning, where the device type characterized by the output result is different from the preset type and the device type in the device image is the preset type.
[0086] The accuracy information can be a multi-class average accuracy, and the calculation expression of the accuracy information can be expressed as:
[0087] 。
[0088] Among them, it can refer to the accuracy information, M can refer to the number of preset categories, and AP0 can refer to the category.
[0089] In this embodiment, by adjusting the multi-target device detection model for the substation according to the device image and the device annotation information pre-annotated for the device image, a fine-tuned multi-target device detection model for the substation is obtained; the preset validation set is input into the fine-tuned multi-target device detection model for the substation to obtain precision information, recall information, and accuracy information; according to the precision information, recall information, and accuracy information, the model performance evaluation result is determined, and the fine-tuned multi-target device detection model for the substation is adjusted according to the model performance evaluation result, which can use the validation set to adjust the model again after fine-tuning the model, optimize the model performance, and thus improve the accuracy of the device image detection result output by the model.
[0090] In some embodiments, adjusting the multi-target device detection model for the substation according to the device image and the device annotation information pre-annotated for the device image includes: determining a target model structure from the multi-target device detection model for the substation, and inputting the device image into the multi-target device detection model for the substation while fixing the model parameters of the target model structure; adjusting the model parameters of the model structure other than the target model structure in the multi-target device detection model for the substation according to the output result of the multi-target device detection model for the substation and the device annotation information pre-annotated for the device image.
[0091] Among them, the target model structure can refer to the model structure that needs to be frozen.
[0092] As an example, the terminal can train a specific structure in the substation multi-target device detection model by freezing a part of the structure in the model, so as to optimize the model performance. For example, the terminal can determine the target model structure to be frozen from the substation multi-target device detection model, and input the device image into the substation multi-target device detection model while fixing the model parameters of the target model structure. Then, the terminal can adjust the model parameters of the model structure in the substation multi-target device detection model except the target model structure according to the difference between the output result of the substation multi-target device detection model and the device annotation information pre-annotated for the device image, until the difference between the output result of the substation multi-target device detection model and the device annotation information pre-annotated for the device image meets the preset difference condition, thereby realizing the fine-tuning of the substation multi-target device detection model.
[0093] In this embodiment, by determining the target model structure from the substation multi-target device detection model, inputting the device image into the substation multi-target device detection model while fixing the model parameters of the target model structure, and adjusting the model parameters of the model structure in the substation multi-target device detection model except the target model structure according to the output result of the substation multi-target device detection model and the device annotation information pre-annotated for the device image, it is possible to fine-tune the model while freezing a part of the structure in the model, so as to correct the output result of the model and improve the accuracy of the device image detection result output by the model.
[0094] In some embodiments, inputting the device image into the input layer of the target detection model to be trained to obtain a feature map includes: preprocessing the device image through the input layer to obtain a processed image; acquiring the pixel information of the processed image, and performing convolution processing on the pixel information to obtain a feature map.
[0095] As an example, after the terminal inputs the device image into the input layer of the target detection model to be trained, the target detection model can perform preprocessing such as data augmentation, adaptive anchor box, random scaling, and normalization on the device image to obtain a processed image. Then, the terminal can acquire the pixel information of the processed image and perform convolution processing on the pixel information to obtain a feature map. In practical applications, the feature map obtained by the terminal after performing convolution processing on the pixel information can be a reduced feature map, and then the reduced feature map can be bilaterally sampled through the convolution layer of the target detection model for subsequent target localization and detection in combination with multi-scale feature information.
[0096] In this embodiment, the device image is preprocessed through the input layer to obtain a processed image; the pixel information of the processed image is acquired and subjected to convolution processing to obtain a feature map, which can preprocess the device image and output a feature map based on the pixel information, facilitating the improvement of the generalization ability of the model, optimizing the model performance, and thus enhancing the accuracy of multi-object device detection in a substation.
[0097] In some embodiments, as Figure 3 shown, a schematic flowchart of multi-object detection in a substation based on the YOLOv5 algorithm is provided. The terminal first obtains an image of multi-object devices in a complex substation as the device image, and then the terminal can identify multi-object images with different brightness levels, different internal scenes of the substation, and different angles, and train the YOLOv5 model in combination with the data of image annotation to obtain a pre-trained YOLOv5 multi-object recognition model. After that, the terminal can perform sparse training on the input channels of the model and prune the BN layer of the sparsely trained model to obtain a lightweight model, and then restore the training network accuracy through fine-tuning, so as to complete the size measurement of multi-object devices in the substation by using the YOLOv5 multi-object recognition model.
[0098] In practical applications, the YOLOv5 multi-object recognition model can include an input layer and a convolution layer. The convolution layer includes a CSP (Cross Stage Partial Network) module, an FPN (Feature Pyramid Network) module, and a PANet (Path Aggregation Network, a neural network architecture designed to improve the performance of object detection) module. Among them, the input layer can perform preprocessing such as data augmentation, adaptive anchor boxes, random scaling, and normalization on the image, collect and perform convolution operations on the pixel information of the image to obtain a reduced feature map. The reduced feature map can be successively input into the CSP module, the FPN module, and the PANet module for multiple up and down bidirectional samplings: the FPN module gradually down-samples the feature map from the shallow layer to the deep layer, and gradually disassembles the information of the feature map through the convolution layer and the pooling layer, reducing the resolution step by step, and fusing the low-level feature map with the high-level feature map to enhance the detailed information contained in the high-level feature map; the PANet module gradually up-samples the feature map from the deep layer to the shallow layer, and gradually increases the resolution of the feature values through nearest neighbor interpolation, enabling the low-level feature map to obtain high-level semantic information and generating multi-scale features. The bidirectional sampling enables the feature map to contain high-level semantic information while retaining the low-level position information, improving the accuracy of multi-object detection in the substation; the CSP module divides the feature map into a directly transmitted part and a part that is merged after convolution processing, extracts the edge and texture features of devices such as high-voltage lines and transformers, reduces redundant calculation steps, enhances the feature extraction ability of the proposed model, and improves the object detection efficiency.
[0099] After the above processing, a feature map containing the semantic information and location information of the target equipment in the substation is obtained. Target localization and confidence calculation are performed on the feature map to obtain a pre-trained YOLOv5 model. Specifically, the YOLOv5 algorithm divides the feature map into multiple grids, and each grid is responsible for predicting the targets within its range, including the target bounding boxes, such as the center point coordinates, width, and height, as well as the target categories. In addition, the grid also calculates the confidence of the bounding box to characterize the possibility that the bounding box contains a target.
[0100] Insert a BN layer after the convolutional layer, retain the channel scaling parameters, and use the reduction coefficient Score the importance of each input channel to distinguish important channels and unimportant channels, as follows:
[0101] The original loss function of the model can be expressed as:
[0102] .
[0103] Among them, the function consists of two parts, which are the matching degree between the predicted bounding box and the true bounding box respectively. The other part calculates the matching degree between the predicted category and the true category. The smaller the values of the two parts, the higher the matching degree. B represents the set of predicted bounding boxes, represents the set of true bounding boxes; represents the number of correctly predicted bounding boxes; represents the total set of bounding boxes; represents the set of predicted categories; represents the number of wrongly predicted categories. By minimizing this loss function, the YOLOv5 model can learn more accurate prediction results of the bounding boxes and category predictions of multi-target equipment in the substation.
[0104] Through sparse training, a model containing each weight is obtained, specifically including:
[0105] Add a K constraint term to the original loss function as follows:
[0106] .
[0107] Among them, the function consists of two parts, which are the reward function of the initial Yolov5 model respectively. n and m represent the output and input of the model, and W is the hyperparameter of the neural network; the term where the equation is located is the K constraint term for the sparsity of the normalization layer, is the penalty factor. During the process of minimizing the loss function, the weight coefficients of unimportant channels Gradually approaches 0. When most of the weight coefficients are concentrated near 0, the sparse training is successful. When pruning the model after sparse training, all unimportant channels with channel weight coefficients approaching 0 and their corresponding weights can be removed.
[0108] As Figure 4 shown, a schematic diagram of model pruning is provided. When pruning the BN layer of the sparse training model, the terminal can associate the penalty factor with each input channel in the convolutional layer, sparsify the penalty factor and then regularize it to automatically classify the importance of channels. Among them, the numerical value of the penalty factor represents the importance of the channel. Prune the channels with smaller penalty factor values to obtain a simplified YOLOv5 model.
[0109] By fine-tuning to restore the training network accuracy, the terminal can use the original dataset containing object detection labels to adjust the hyperparameters of the pruned model, and train specific layers by freezing some layers. Compare the original dataset with the training network results to correct the training network results. Then the terminal can use the validation set to test the model performance and evaluate the model effect using three metrics: precision, recall, and multi-class average accuracy. Among them, precision evaluates the accuracy of the prediction results, recall evaluates the comprehensiveness of the predicted content, and multi-class average accuracy evaluates the average value of the accuracies of each category. According to the performance evaluation results, the terminal can further adjust the hyperparameters of the model or repeat the fine-tuning process to optimize the model performance and improve the accuracy of the model output results.
[0110] As Figure 5 shown, a schematic diagram of the architecture of a substation multi-target device detection model is provided. The substation multi-target device detection model can include an object detection model constructed by the YOLOv5 algorithm. Therefore, the substation multi-target device detection model can regard object detection as a regression problem and directly predict the object position, size, and category through a convolutional neural network, which is especially suitable for real-time object detection tasks. Pruning and fine-tuning the BN layer of the YOLOv5 model, the influence of different pruning rates on the network model accuracy is compared and analyzed, and object detection is performed on important substation equipment such as main transformers, lightning arresters, disconnectors, and reactors. The detection results show that on the premise that the average precision of the pruned model only drops by 0.05%, compared with the original model, the model size can be reduced by 70%, and the detection speed is increased by 17.1%, which can meet the deployment requirements of the model on mobile terminal devices.
[0111] In this embodiment, a pre-trained model is obtained by processing multi-target images of a complex substation. Then, pruning and compression are performed on the normalization layer in the neural network to streamline the model structure, and then sparse training is carried out. Finally, fine-tuning is used to repair the network accuracy. The YOLOv5 algorithm is used to reduce the model of the multi-target detection network for the substation, which can effectively identify the size of the multi-target detection model for the substation, improve the model detection speed, and improve the work efficiency of substation inspection personnel while ensuring the detection accuracy.
[0112] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0113] Based on the same inventive concept, the embodiments of the present application also provide a substation multi-target device detection model generation device for implementing the above-mentioned substation multi-target device detection model generation method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the substation multi-target device detection model generation device provided below can refer to the limitations on the substation multi-target device detection model generation method in the above text, and will not be repeated here.
[0114] In an exemplary embodiment, as Figure 6 shown, a substation multi-target device detection model generation device is provided, including: an input module 602, a sampling module 604, an identification module 606, and a generation module 608, where:
[0115] The input module 602 is configured to input a device image into the input layer of a target detection model to be trained, obtain a feature map, and divide the feature map into a first feature map set and a second feature map set according to a preset division method.
[0116] The sampling module 604 is configured to input the first feature map set into the downsampling module in the target detection model, perform downsampling processing on the first feature map set to obtain a downsampling result, and input the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result.
[0117] The recognition module 606 is configured to determine a device image detection result according to the second feature map set and the upsampling result; the device image detection result characterizes the type and position of the device in the device image.
[0118] The generation module 608 is configured to train the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and prune the target detection model to obtain a substation multi-target device detection model; the substation multi-target device detection model is used to detect and locate devices in an image.
[0119] In one exemplary embodiment, the recognition module 606 is further specifically configured to determine a target feature map according to the second feature map set and the upsampling result; the target feature map characterizes the edge feature and texture feature of the device; divide the target feature map into at least one unit through the target detection model, and determine the detection result corresponding to each unit; the detection result includes the device bounding box and device type corresponding to each unit; determine the device image detection result according to the detection result corresponding to each unit.
[0120] In one exemplary embodiment, the target detection model includes a convolutional layer, the convolutional layer includes the downsampling module and the upsampling module, and the generation module 608 is further specifically configured to, when there is a batch normalization layer after the convolutional layer, input the device image into the target detection model, and adjust the reduction coefficient of each input channel of the target detection model according to the output result of the target detection model and the device annotation information pre-annotated for the device image; determine the input channels to be deleted according to the reduction coefficient of each input channel of the target detection model; delete the input channels to be deleted and the reduction coefficients corresponding to the input channels to be deleted from the target detection model.
[0121] In one exemplary embodiment, the device further includes a fine-tuning module, which is specifically configured to adjust the substation multi-target device detection model according to the device image and the device annotation information pre-annotated for the device image, so as to obtain a fine-tuned substation multi-target device detection model; input a preset validation set into the fine-tuned substation multi-target device detection model to obtain precision information, recall information, and accuracy information; determine a model performance evaluation result according to the precision information, the recall information, and the accuracy information, and adjust the fine-tuned substation multi-target device detection model according to the model performance evaluation result.
[0122] In one exemplary embodiment, the fine-tuning module is specifically further configured to determine a target model structure from the substation multi-target device detection model, and input the device image into the substation multi-target device detection model while fixing the model parameters of the target model structure; adjust the model parameters of the model structure other than the target model structure in the substation multi-target device detection model according to the output result of the substation multi-target device detection model and the device annotation information pre-annotated for the device image.
[0123] In one exemplary embodiment, the input module 602 is specifically further configured to preprocess the device image through the input layer to obtain a processed image; obtain pixel information of the processed image, and perform convolution processing on the pixel information to obtain the feature map.
[0124] Each module in the above-mentioned substation multi-target device detection model generation device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0125] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a method for generating a multi-object detection model for a substation. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0126] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0127] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0128] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0129] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0130] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0131] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0132] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0133] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A method for generating a multi-objective device detection model for a substation, characterized in that The method includes: Input the device image into the input layer of the target detection model to be trained, obtain a feature map, and divide the feature map into a first feature map set and a second feature map set according to a preset division method; Input the first feature map set into the downsampling module in the target detection model, perform downsampling processing on the first feature map set to obtain a downsampling result, and input the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result; Determine the device image detection result according to the second feature map set and the upsampling result; the device image detection result characterizes the type and position of the device in the device image; Train the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and prune the target detection model to obtain a substation multi-target device detection model; the substation multi-target device detection model is used to detect and locate the devices in the image.
2. The method according to claim 1, characterized in that, The determining the device image detection result according to the second feature map set and the upsampling result includes: Determine a target feature map according to the second feature map set and the upsampling result; the target feature map characterizes the edge feature and texture feature of the device; Divide the target feature map into at least one unit through the target detection model, and determine the detection result corresponding to each unit; the detection result includes the device bounding box and device type corresponding to each unit; Determine the device image detection result according to the detection results corresponding to each unit.
3. The method according to claim 1, wherein The target detection model includes a convolutional layer, the convolutional layer includes the downsampling module and the upsampling module, and the pruning of the target detection model includes: When there is a batch normalization layer after the convolutional layer, input the device image into the target detection model, and adjust the reduction coefficient of each input channel of the target detection model according to the output result of the target detection model and the device annotation information pre-annotated for the device image; Determine the input channels to be deleted according to the reduction coefficients of each input channel of the target detection model; Delete the input channels to be deleted and the corresponding reduction coefficients of the input channels to be deleted from the target detection model.
4. The method according to claim 1, characterized in that After the step of obtaining the substation multi-target device detection model, the method further includes: Adjust the substation multi-target device detection model according to the device image and the device annotation information pre-annotated for the device image to obtain a fine-tuned substation multi-target device detection model; Input a preset validation set into the fine-tuned substation multi-target device detection model to obtain precision information, recall information, and accuracy information; Determine the model performance evaluation result according to the precision information, the recall information, and the accuracy information, and adjust the fine-tuned substation multi-target device detection model according to the model performance evaluation result.
5. The method according to claim 4, wherein Adjusting the multi-object device detection model of the substation according to the device image and the device annotation information pre-annotated for the device image includes: Determining a target model structure from the multi-object device detection model of the substation, and inputting the device image into the multi-object device detection model of the substation while fixing the model parameters of the target model structure; Adjusting the model parameters of the model structure other than the target model structure in the multi-object device detection model of the substation according to the output result of the multi-object device detection model of the substation and the device annotation information pre-annotated for the device image.
6. The method according to claim 1, characterized in that, The step of inputting the device image into the input layer of the target detection model to be trained to obtain a feature map includes: Preprocessing the device image through the input layer to obtain a processed image; Obtaining the pixel information of the processed image and performing convolution processing on the pixel information to obtain the feature map.
7. A device for generating a multi-objective equipment detection model for a substation, characterized in that, The device includes: An input module, configured to input a device image into the input layer of the target detection model to be trained to obtain a feature map, and divide the feature map into a first feature map set and a second feature map set according to a preset division method; A sampling module, configured to input the first feature map set into the downsampling module in the target detection model to perform downsampling processing on the first feature map set to obtain a downsampling result, and input the downsampling result set into the upsampling module in the target detection model to perform nearest neighbor interpolation processing on the downsampling result to obtain an upsampling result; An identification module, configured to determine a device image detection result according to the second feature map set and the upsampling result; the device image detection result characterizes the type and position of the device in the device image; A generation module, configured to train the target detection model according to the device image detection result and the device annotation information pre-annotated for the device image, and prune the target detection model to obtain a multi-object device detection model of the substation; the multi-object device detection model of the substation is used to detect and locate the devices in the image.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.