Power energy equipment identification method, device and terminal equipment

By introducing ParNet and RGFPN networks into the Mask-RCNN model, combined with the joint feature extraction and fusion of feature pyramids, the missed and misidentified problems of small-scale object detection in multi-objective image recognition in the prior art are solved, and more efficient recognition accuracy and speed are achieved.

CN114399681BActive Publication Date: 2025-05-16INST OF ECONOMIC & TECH STATE GRID HEBEI ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210042791.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-14
Publication Date
2025-05-16
Estimated Expiration
2042-01-14

AI Technical Summary

Technical Problem

In the prior art, in multi-object image recognition, especially in small-object detection, there are problems of misunderstanding and misunderstanding, and the detection speed is slow, making it difficult to achieve real-time performance.

Method used

The Mask-RCNN model is used to combine ParNet and RGFPN networks to enhance the original feature information flow through joint feature extraction and fusion of feature pyramids, and optimize model parameters through loss function to improve recognition accuracy and speed.

Benefits of technology

While maintaining high recognition performance, the training speed is improved by about 30%, significantly improving the recognition accuracy of power and energy equipment, especially the recognition rate of small target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399681B_ABST
    Figure CN114399681B_ABST
Patent Text Reader

Abstract

The present application is applicable to the field of image recognition technology, and discloses a method, device and terminal device for identifying electric energy equipment. The above-mentioned electric energy equipment identification method includes: collecting electric energy equipment images, preprocessing the electric energy equipment images, and establishing an electric energy equipment image data set; the electric energy equipment image data set includes a training set and a verification set; based on the training set and the verification set, a Mask-RCNN model for electric energy equipment identification is trained; the real-life picture of the electric energy equipment to be identified is input into the Mask-RCNN model for electric energy equipment identification, and the identification result of the energy equipment on the real-life picture is obtained. The Mask-RCNN model for electric energy equipment identification established based on the electric energy equipment image data set improves the detection and identification accuracy of electric energy equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image recognition technology, and in particular, relates to a method, device and terminal device for identifying electric energy equipment. Background Art

[0002] In the problem of electric power equipment identification, multi-target image recognition has become a core technical issue. Convolutional neural networks have achieved remarkable results in image recognition and classification tasks, and the emergence of the R-CNN detection algorithm has marked the beginning of deep learning in target detection. Next, Fast-RCNN changed the shortcomings of R-CNN, such as low detection accuracy, low detection efficiency, and high resource consumption, and the detection accuracy has been greatly improved. However, the detection speed is still slow, the detection efficiency is low, and a large amount of time redundancy is caused during the detection process, which cannot achieve real-time performance. Faster-RCNN adds the RPN network on the basis of the Fast R-CNN model, which has improved the speed and detection accuracy compared to Fast-RCNN. The YOLO target detection model has also formed YOLO-v2 and YOLO-v3 on the original basis, and its performance and accuracy have been improved.

[0003] However, the above model still has the phenomenon of missed recognition and misrecognition for the problem of small target detection in multi-target image recognition. After feature extraction, the mapping ratio of low-resolution feature map to high-resolution feature map is large. This compression process causes the loss of feature information flow, resulting in the disappearance of the features of small target objects during the feature extraction process. The insufficient fusion of feature layers at each resolution leads to insufficient use of feature information, which can easily cause misrecognition. The neural network algorithm of the existing model also converges slowly, and is inefficient when detecting large amounts of data, making it difficult to achieve rapid target recognition. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a method, an apparatus and a terminal device for identifying electric energy equipment to improve the recognition accuracy of electric energy equipment.

[0005] This application is implemented through the following technical solutions:

[0006] In the first aspect, an embodiment of the present application provides a method for identifying electric energy equipment, which is characterized by comprising: collecting images of electric energy equipment, preprocessing the images of electric energy equipment, and establishing an electric energy equipment image dataset; the electric energy equipment image dataset includes a training set and a verification set; based on the training set and the verification set, training a Mask-RCNN model for identifying electric energy equipment; inputting a real-life picture of the electric energy equipment to be identified into the Mask-RCNN model for identifying electric energy equipment, and obtaining the recognition result of the energy equipment on the real-life picture.

[0007] In an embodiment of the present application, the trained Mask-RCNN model for identifying electric energy equipment can increase the training speed by about 30% while maintaining high performance, and enhance the original feature information flow by combining the features of each stage of the final feature pyramid, thereby improving the recognition accuracy of electric energy equipment.

[0008] Based on the first aspect, in some embodiments, electric energy equipment images are collected, the electric energy equipment images are preprocessed, and an electric energy equipment image dataset is established, including: collecting electric energy equipment images, annotating multi-scale electric energy equipment in the electric energy equipment images, and obtaining annotation information containing the contours and types of the multi-scale electric energy equipment; merging components belonging to the same type of electric energy equipment into a complete electric energy equipment according to the annotation information, and generating a mask; and segmenting the dataset according to the scale of the complete electric energy equipment to obtain an electric energy equipment image dataset.

[0009] Based on the first aspect, in some embodiments, based on a training set and a validation set, a Mask-RCNN model for electric energy equipment identification is trained, including: extracting features from the electric energy equipment images in the training set to obtain a global information feature image; aligning the regions of interest on the global information feature image through a RoIAlign layer to obtain an output result of a Mask-RCNN prototype model; performing classification and regression on the output result of the Mask-RCNN prototype model through a loss function to obtain the optimal parameters of the Mask-RCNN prototype model, and using the Mask-RCNN prototype model with the optimal parameters as the Mask-RCNN model for electric energy equipment identification.

[0010] Based on the first aspect, in some embodiments, feature extraction is performed on the electric energy equipment images in the training set to obtain a global information feature image, including: extracting features from the electric energy equipment images in the training set through a Par Net network and an RGFPN network to obtain a first feature image with different resolutions; synchronizing the resolution of the first feature image to obtain a second feature image with the same resolution; and performing channel fusion on the second feature image to obtain a global information feature image.

[0011] In an embodiment of the present application, the ParNet network uses a parameter renormalization method to reorganize the decoupled branch structure into a convolutional layer module with the same parameter reconstruction, compressing the network depth to 12 layers, which can improve the training speed while maintaining high performance. The present invention inputs the fused output of the traditional FPN to ParNet for a secondary cycle, enhances the original feature information flow by combining the features of each stage of the final feature pyramid, obtains a feature map of global information, and inputs it to the RolAlign layer, thereby improving the detection accuracy of the model. The feature information recalibrated by recursive FPN helps to improve the accuracy of the target detection model. The recognition rate of targets of different sizes is improved, and the recognition rate of small target objects is improved more significantly.

[0012] Based on the first aspect, in some embodiments, feature extraction is performed on the electric energy equipment images in the training set through the Par Net network and the RGFPN network to obtain a first feature image with different resolutions, including: training the electric energy equipment images through the Par Net network and the FPN network to obtain first feature information; inputting the features of the corresponding layer in the first feature into the Par Net network and the FPN network again for convolution to obtain second feature information; combining the second feature information with the first feature information to obtain a first feature image.

[0013] Based on the first aspect, in some embodiments, performing resolution synchronization on the first feature image to obtain a second feature image with the same resolution includes: dividing the first feature image into C 1 Layer to C 5 Layer, C 3 The layer is the middle layer; by downsampling, C 1 Layer and C 2 The resolution of the layer is reduced to the middle layer C 3 The resolution size of the layer; by upsampling C 4 Layer and C 5 The resolution of the layer is increased to the middle layer C 3 The resolution size of the layer is obtained to obtain a second feature image with the same resolution.

[0014] In the second aspect, an embodiment of the present application provides a device, including: a data acquisition module, used to acquire images of electric energy equipment, pre-process the images of electric energy equipment, and establish an electric energy equipment image data set; the electric energy equipment image data set includes a training set and a verification set; a training module, used to train an electric energy equipment recognition Mask-RCNN model based on the training set and the verification set; and a recognition module, used to input a real-life picture of the electric energy equipment to be identified into the electric energy equipment recognition Mask-RCNN model to obtain the recognition result of the energy equipment on the real-life picture.

[0015] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the electric energy equipment identification method as described in any one of the first aspects above are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the electric energy equipment identification method as described in any one of the first aspects above are implemented.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0019] Figure 1 It is a flowchart of the electric energy equipment identification method provided in the embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of annotating a picture of a power equipment provided in an embodiment of the present application;

[0021] Figure 3 It is a ParNet-RGFPN network model framework diagram provided in an embodiment of the present application;

[0022] Figure 4 It is a schematic diagram of the structure of the fusion module provided in the embodiment of the present application;

[0023] Figure 5 This is a comparison diagram of the effects of the Mask-RCNN model for identifying electric energy equipment provided in the embodiment of the present application and the Mask-RCNN prototype model;

[0024] Figure 6 It is a structural schematic diagram of the electric energy equipment identification device provided in an embodiment of the present application;

[0025] Figure 7 It is a structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0028] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0030] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0031] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0032] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0033] The following explains the terms used in this application.

[0034] Convolutional Neural Network: Convolutional Neural Network (CNN) is a type of feedforward neural network with deep structure and convolutional calculation. It is one of the representative algorithms of deep learning. Convolutional neural network has the ability of representation learning and can classify input information in a translation-invariant manner according to its hierarchical structure. Therefore, it is also called "translation-invariant artificial neural network".

[0035] Backbone: The backbone network of convolutional neural network.

[0036] Image recognition: The technology of using computers to process, analyze and understand images to identify targets and objects of various different patterns is a practical application of deep learning algorithms.

[0037] Object detection: Object detection, also known as object extraction, is a type of image segmentation based on the geometric and statistical features of the object. It combines the segmentation and recognition of the object into one, and its accuracy and real-time performance are important capabilities of the entire system. Automatic object extraction and recognition are particularly important in complex scenes when multiple objects need to be processed in real time.

[0038] Feature extraction: The method and process of using computers to extract characteristic information from images.

[0039] Pooling: Reduce the amount of computation.

[0040] Anchor: The mapping point of the center of the current sliding window in the original pixel space on the feature image generated by the CNN network is called the Anchor. According to the predefined Anchor, nine borders of different shapes and sizes can be generated on the original image with a point on the feature image as the center.

[0041] FPN: FPN is a feature pyramid model that combines multi-level features to solve multi-scale problems.

[0042] COCO dataset: Common Objects in Context, a dataset provided by the Microsoft team that can be used for image recognition.

[0043] Confidence: The confidence interval of a probability sample is an interval estimate of a population parameter of the sample. The confidence interval shows the degree to which the true value of the parameter has a certain probability of falling around the measurement result. The confidence interval gives the range of the degree of credibility of the measured value of the measured parameter, that is, the "certain probability" required above. This probability is called the confidence level.

[0044] Small object detection has always been a practical and common problem due to low resolution, blurred images, little information, and high noise. In the past few years, some solutions have emerged to improve the performance of small object detection.

[0045] Before deep learning methods became popular, for targets of different scales, we usually started with the original image, built an image pyramid using different resolutions, and then used a classifier to perform sliding window target detection on each layer of the pyramid. However, this method is inefficient. Although the construction of the image pyramid can be accelerated by using convolution kernel separation or direct scaling, it still requires multiple feature extractions. Therefore, in recent years, deep learning methods have been widely used for image recognition, among which convolutional neural networks are most suitable for image recognition tasks. In order to improve the performance of small target detection, some improvement methods using convolutional neural networks have also been proposed.

[0046] First, we can increase the types and number of small target samples in the training set. Deep learning algorithms often use the COCO dataset as samples for training. On the one hand, to address the problem of the small number of images containing small targets in the COCO dataset, we use a sampling strategy. Regardless of the ratio of detecting small targets, oversampling helps. On the other hand, to address the problem of the small number of small targets in the same image, we use a segmentation mask to cut out the small target image, and then use the copy and paste method, as well as rotation and scaling to achieve artificial enhancement, so that more anchors match the small targets during training.

[0047] Thirdly, since the feature maps at different stages correspond to different receptive fields, the degree of abstraction of the information they express is also different. The shallow feature map has a small receptive field and is more suitable for detecting small targets, while the deep feature map has a large receptive field and is more suitable for detecting large targets. Therefore, based on the image pyramid, it is proposed to integrate the feature maps at different stages to improve the target detection performance, namely, the feature fusion pyramid FPN. The fusion of features at different fusion layers only requires one forward calculation to complete.

[0048] In order to enhance the representation ability of features, the use of multi-level feature pyramids is of great significance for model discrimination. The basis of the FPN network model is the feature pyramid. By establishing additional top-down pathways for features of different resolutions at each level, the communication of features at each level is realized, which enhances the representation ability of feature layers at each level. However, the additional pathways established by FPN make the features of different layers of the network only receive the semantic information of adjacent layers, and cannot receive the semantic information of other layers, and the feature fusion is not sufficient. The underlying features will contain more feature information about texture, position and edges. In order to make the network recognize the target position more accurately, the existing technology uses PANet to add additional pathways between the bottom and top layers on the basis of FPN to enhance the detailed feature information of the target at the top layer. However, considering the overall feature fusion and detection accuracy, the use of PANet has not yet achieved good results.

[0049] In order to improve the recognition accuracy of electric energy equipment, this application provides a method for identifying electric energy equipment, such as Figure 1 As shown, the electric energy equipment identification method may include steps 101 to 103.

[0050] Step 101: Collect images of electric power equipment, pre-process the images of electric power equipment, and establish an electric power equipment image dataset; the electric power equipment image dataset includes a training set and a validation set.

[0051] In some embodiments, step 101 may include steps 1011 to 1013 .

[0052] Step 1011: Collect an image of electric energy equipment, and annotate the multi-scale electric energy equipment in the image of the electric energy equipment to obtain annotation information including the outline and type of the multi-scale electric energy equipment.

[0053] In some embodiments, due to the lack of open source datasets that can be used for machine learning in the field of power equipment, before applying the Mask-RCNN model, this application uses substation inspection robots to capture a variety of power energy equipment images, and uses the LabelImg labeling tool to establish a dataset containing 200 power energy equipment images. The multi-scale power equipment contained in each power energy equipment image is labeled, and the annotation of each image is as follows: Figure 2 As shown in the figure, the annotation information includes the outline and type of multi-scale equipment. After the image is annotated, an XML file will be generated, which contains all the annotation information. According to the information in the dataset, the type of the equipment is set to Telegraph poles, transformer, insulator, crossarm, and wire clip.

[0054] Step 1012: merging components belonging to the same electric energy equipment type into a complete electric energy equipment according to the labeling information, and generating a mask.

[0055] Merge the parts that belong to the same device but are marked separately, and merge the parts with the same marked name into a complete power energy device according to the marking information, and generate a mask at the same time. The mask is to use the selected image, graphic or object to cover the whole or part of the processed image to control the image processing area or processing process. Its function is to extract the energy equipment in the area of ​​interest, multiply the pre-made area of ​​interest mask with the image to be processed, and obtain the image of the area of ​​interest. The image value in the area of ​​interest remains unchanged, and the image value outside the area is 0, so as to achieve the purpose of segmentation.

[0056] Step 1013: segment the data set according to the scale of the complete electric power equipment to obtain an electric power equipment image data set.

[0057] After the merging is completed, in order to improve the effect of the subsequent Mask-RCNN model on target detection for devices of different scales, the input image is segmented according to the scale of the energy equipment. After the segmentation, the data set is randomly divided into two parts, the training set and the validation set. The training set is used to train the Mask-RCNN model for power energy equipment recognition, and the validation set is used to verify the actual effect of the neural network.

[0058] Step 102: Based on the training set and the validation set, a Mask-RCNN model for electric energy equipment recognition is trained.

[0059] In some embodiments, step 102 may include steps 1021 to 1023 .

[0060] Step 1021: extract features from the electric energy equipment images in the training set to obtain a global information feature image.

[0061] As the latest network model of R-CNN, the Mask-RCNN prototype model absorbs the advantages of the R-CNN system algorithm and makes further improvements on their basis. Mask-RCNN uses RoIAlign instead of RoI pooling. Specifically, it removes the original rounding operation, retains the calculated floating point number, and uses bilinear interpolation to complete the pixel operation. As a result, accurate pixel-level alignment is achieved in the field of instance segmentation. However, the backbone network ResNeXt-101 seriously slows down the training and reasoning speed of the model, and the application performance of the original FPN network in feature extraction of small-scale power equipment also has room for improvement.

[0062] The improved Mask-RCNN prototype model of this application adopts Par Net (Parallel Networks) as the backbone network, synchronizes the resolution of each resolution feature layer after RGFPN processing (feature pyramid networks), and then performs channel fusion on the feature map after unified resolution through the feature fusion module to obtain a feature map of global information with higher detection accuracy, which is input into the RoIAlign layer.

[0063] In some embodiments, Par Net and RGFPN are used to extract features from the electric power equipment images in the training set to obtain first feature images with different resolutions.

[0064] ParNet is used as the backbone for pre-training. The decoupled branch structure is reorganized into the same parameter-reconstructed convolutional layer module through parameter renormalization. Rep-Block and SSE structures are introduced to compress the network depth to 12 layers. After ParNet pre-training, the input image is input into the RGFPN network for feature extraction to obtain the first feature information, which contains feature maps generated at different stages. The feature map generated by the i-th stage is represented as C k , in the feature layer, C 1 The highest resolution, C 5 The resolution is the lowest. Figure 3 It is an expanded framework diagram of the feature extraction process. RGFPN sends the first feature information output by training the image of power energy equipment through the ParNet network and the FPN network to ParNet again as the feature of the corresponding layer of the backbone for convolution. After processing through FPN, the second feature information is obtained. The second feature information is combined with the first feature information to obtain the first feature image.

[0065] Adjust the size of the first feature image by downsampling C 1 Layer and C 2 The resolution of the layer is reduced to the middle layer C 3 The resolution size of the layer. By upsampling C 4 Layer and C 5 The resolution of the layer is increased to the middle layer C 3 The resolution size of the layer is obtained to obtain a second feature image with the same resolution.

[0066] In the process of resizing, in order to retain the feature information to the maximum extent, the feature maps of different resolution sizes are 1 , C 2 , C 4 , C 5}Adjust to middle layer C3 The resolution of the layer is M3*M3. The high-resolution feature layer is downsampled to match C 3 The resolution of the stage is 2000, and the low-resolution feature layer is upsampled by bilinear interpolation to achieve C 3 The size of the stage resolution. In the upsampling and downsampling stages, the low-resolution feature layer retains the feature information of large target objects by using the transposed convolution method. The high-resolution feature map generated by downsampling retains the feature information of small target objects.

[0067] Specifically, we use two downsampling to reduce C 1 , C 2 The feature layer resolution of the stage is reduced to C 3 The resolution of the stage. The feature layer with high resolution will contain more detailed information of the target object. The feature information of higher stage, such as C 5 With a more abstract feature description, C 1 , C 2 The purpose of downsampling in both stages is to extract the detailed feature information of small target objects at different stages, so that the model can have a good recognition effect on small target objects.

[0068] The resized feature maps are connected using channel connections and then enter the fusion module for feature fusion processing.

[0069] like Figure 4 As shown in the figure, the fusion module consists of two convolutional layers: a point-by-point convolutional layer (1*1*N) and a standard convolutional layer (3*3*N). The function of the fusion module is to perform feature fusion processing on the second feature image after the unified resolution and reduce the dimension of the feature layer. The feature information flow of the second feature image is expressed as M3*M3*5N, where N is the number of output channels at each stage of the feature pyramid, and M3 represents C 3 The feature map size of the layer. The point-by-point convolution layer fuses the input features and reduces the feature dimension from 5N to N. This stage will generate a feature map with global feature information. The dimension of the feature map after the point-by-point convolution layer is expressed as M3*M3*N. Using the standard convolution layer can increase the difference in feature information between adjacent pixels and reduce the feature confusion effect caused by resizing and point-by-point convolution. The standard convolution layer does not change the dimension of the feature information flow. The dimension of the new feature information flow generated by the fusion module is M3*M3*N. The feature image with the new feature information flow is the global information feature image.

[0070] The feature map after recursive FPN significantly improves the detection performance of objects of different scales. The feature information flow processed by the resize and fusion module is recalibrated and enhanced by drawing on the idea of ​​residual connection. The recalibrated feature map will contain global feature information and the detection accuracy will be more accurate. The feature map after the fusion module has balanced global feature information, and the powerful feature representation ability improves the detection accuracy of the model, and the detection effect of small targets is significantly improved.

[0071] Step 1022: Align the regions of interest on the global information feature image through the RoIAlign layer to obtain the output result of the Mask-RCNN prototype model.

[0072] RPN (region proposal networks) is used to extract the region of interest on the global information feature image. The extracted region of interest and the original feature map are then input into the RoI Align layer for feature alignment. Target objects of different sizes belong to feature layers of different resolutions. Therefore, different feature layers should be used as inputs to the RoI Align layer for regions of interest (RoI) of different scales. The region of interest of large target objects should be mapped to a low-resolution feature layer, such as C5. The corresponding region of interest of small target objects should be mapped to a high-resolution feature layer, such as C1.

[0073] Specifically, the region of interest with a width of w and a height of h on the input image is assigned to C of the feature pyramid k The relationship between layer, k and width and height is as follows:

[0074]

[0075] Among them, 224 represents the size of the input image, k 0 Indicates wh=224 2 The region of interest should be mapped to the target level, k 0 The default setting is 5, which represents the output of the C5 layer. w and h represent the length and width of the region of interest. Assuming that the RoI is 112×112 in size, then k=k 0 -1=5-1=4, which means that the region of interest should use a feature layer of size C4. From the above formula, we can see that if the scale of the region of interest becomes smaller, for example, becomes 1 / 2 of 224, then the region of interest should be mapped to a larger resolution level.

[0076] Each deep convolutional neural network is flexible and variable within a certain range, and these changes are caused by different network parameters. The model effect obtained through training is verified through the validation set. When the recognition accuracy reaches the predetermined standard, the parameters of the deep convolutional neural network are determined as follows:

[0077] Learning rate = 0.001

[0078] epochs = 50

[0079] How many parts are all samples divided into? steps per epoch = 100

[0080] RoI confidence threshold detection min condidence = 0.9

[0081] Number of images processed by each GPU images per GPU = 2

[0082] Step 1023: classify and regress the output results of the Mask-RCNN prototype model through the loss function to obtain the optimal parameters of the Mask-RCNN prototype model, and use the Mask-RCNN prototype model with the optimal parameters as the Mask-RCNN model for electric energy equipment identification.

[0083] Use RGFPN to extract the bounding box of the feature map and map it to the feature map, input RoIAlign, perform RoIAlign operation, use Hybrid Adam-SGD optimizer, classify and regress the output results according to the loss function, and obtain the optimal model parameters.

[0084] This application uses multiple loss functions in Mask-RCNN, including: rpn_class_loss (RPN network classification loss), rpn_bbox_loss (RPN network regression loss), class_loss (classification loss), bbox_loss (regression loss) and mask_loss (Mask segmentation mask regression loss). The average binary cross entropy corresponding to each point in the image is obtained by calculating the relative entropy error pixel by pixel. The improved calculation formula is as follows:

[0085]

[0086] Step 103: Input the real-scene picture of the electric energy equipment to be identified into the electric energy equipment identification Mask-RCNN model to obtain the identification result of the energy equipment on the real-scene picture.

[0087] In some embodiments, training can be performed based on Python 3.6. There are 200 images in the power energy equipment data set, of which the training set data includes 128 images and the validation set data includes 72 images. The neural network weights obtained by training the 128 images in the training set are imported into the prototype model, and then the images in the test set are tested. The final output result is as follows: Figure 5 shown.

[0088] One of the reasons why small objects cannot be recognized is that the features of small objects with a candidate box size of less than 32×32 pixels are difficult to be learned by the feature pyramid network. The root cause of this problem is that the feature extraction of the feature pyramid network is divided into five stages, and the feature map generated in each stage is 1 / 2 smaller than the feature map of the previous stage. In addition, the communication of the feature information flow of the feature pyramid is limited to adjacent feature layers, and the traditional FPN has limitations on the accuracy of multi-scale object detection. Therefore, the feature information of small objects is easily lost in the feature information flow, making the model unable to recognize such small objects. The present invention aims at the problem that the feature pyramid has low precision and the feature information flow fusion of each stage is insufficient. The non-deep network ParNet network is used as the backbone to replace the ResNeXt-101 network used in the traditional Mask-RCNN. Although the extremely deep architecture design and ingenious residual structure used in the ResNeXt-101 network can achieve higher performance, it also seriously slows down the training and reasoning speed of the model. The ParNet network used in the present invention uses a parameter renormalization method to reorganize the decoupled branch structure into the convolutional layer module reconstructed with the same parameter. By introducing Rep-Block and SSE structures, the network depth is compressed to 12 layers, and the training speed can be increased by about 30% while maintaining high performance. Secondly, the improved FPN network RGFPN, i.e., recursive global FPN, is applied to obtain a feature map of global information with higher detection accuracy. In addition, the optimizer and multiple loss functions used in the present invention significantly improve the convergence speed of the model.

[0089] See also Figure 6 The electric energy equipment identification device in the embodiment of the present application may include: a data acquisition module 610, a training module 620, and an identification module 630.

[0090] The data acquisition module 610 is used to acquire images of electric power equipment, pre-process the images of electric power equipment, and establish an electric power equipment image dataset; the electric power equipment image dataset includes a training set and a verification set.

[0091] The training module 620 is used to train a Mask-RCNN model for electric power equipment identification based on a training set and a validation set.

[0092] The recognition module 630 is used to input the real-scene picture of the electric energy equipment to be identified into the electric energy equipment recognition Mask-RCNN model to obtain the recognition result of the energy equipment on the real-scene picture.

[0093] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0094] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0095] The present application also provides a terminal device, see Figure 7 The terminal device 700 may include: at least one processor 710, a memory 720, and a computer program stored in the memory 720 and executable on the at least one processor 710, wherein the processor 710 implements the steps in any of the above-mentioned method embodiments when executing the computer program, for example Figure 1 Steps 101 to 103 in the illustrated embodiment. Alternatively, when the processor 710 executes the computer program, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 6 Functions of modules 610 to 630 are shown.

[0096] Exemplarily, the computer program may be divided into one or more modules / units, one or more modules / units are stored in the memory 720, and executed by the processor 710 to complete the present application. The one or more modules / units may be a series of computer program segments capable of completing specific functions, and the program segments are used to describe the execution process of the computer program in the terminal device 700.

[0097] Those skilled in the art will understand that Figure 7It is only an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components, such as input and output devices, network access devices, buses, etc.

[0098] The processor 710 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0099] The memory 720 may be an internal storage unit of the terminal device, or an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. The memory 720 is used to store the computer program and other programs and data required by the terminal device. The memory 720 may also be used to temporarily store data that has been output or is to be output.

[0100] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0101] The electric energy equipment identification method provided in the embodiment of the present application can be applied to terminal devices such as computers, wearable devices, vehicle-mounted devices, tablet computers, laptops, netbooks, personal digital assistants (PDA), augmented reality (AR) / virtual reality (VR) devices, mobile phones, etc. The embodiment of the present application does not impose any restrictions on the specific type of terminal devices.

[0102] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in each embodiment of the above-mentioned electric energy equipment identification method can be implemented.

[0103] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in each embodiment of the above-mentioned electric energy equipment identification method when executing the computer program product.

[0104] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0105] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0106] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0107] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0108] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for identifying electric energy equipment, characterized in that: include: Collecting electric power equipment images, preprocessing the electric power equipment images, and establishing an electric power equipment image data set; The electric power equipment image dataset includes a training set and a verification set; Based on the training set and the validation set, a Mask-RCNN model for electric energy equipment recognition is trained; Inputting the real-scene picture of the electric energy equipment to be identified into the electric energy equipment identification Mask-RCNN model to obtain the identification result of the energy equipment on the real-scene picture; The Mask-RCNN model for electric energy equipment recognition is trained based on the training set and the validation set, including: Extracting features from the electric energy equipment images in the training set to obtain a global information feature image; The RoIAlign layer is used to align the regions of interest on the global information feature image to obtain the output result of the Mask-RCNN prototype model. The output results of the Mask-RCNN prototype model are classified and regressed through the loss function to obtain the optimal parameters of the Mask-RCNN prototype model. The Mask-RCNN prototype model with the optimal parameters is used as the Mask-RCNN model for power energy equipment identification. The step of extracting features from the electric energy equipment images in the training set to obtain a global information feature image includes: Performing feature extraction on the electric power equipment images in the training set through the Par Net network and the RGFPN network to obtain first feature images with different resolutions; Performing resolution synchronization on the first feature image to obtain a second feature image with the same resolution; The second feature image is subjected to channel fusion to obtain a global information feature image.

2. The electric energy equipment identification method according to claim 1, characterized in that: The collecting of electric power equipment images, preprocessing of the electric power equipment images, and establishing of electric power equipment image data sets include: Collecting an image of electric energy equipment, annotating multi-scale electric energy equipment in the image of electric energy equipment, and obtaining annotation information including the outline and type of the multi-scale electric energy equipment; According to the annotation information, components belonging to the same type of electric energy equipment are combined into a complete electric energy equipment to generate a mask; The data set is segmented according to the scale of the complete electric power energy equipment to obtain an electric power energy equipment image data set.

3. The electric energy equipment identification method according to claim 1, characterized in that: The feature extraction of the electric power equipment images in the training set is performed by the Par Net network and the RGFPN network to obtain first feature images with different resolutions, including: The Par Net network and the FPN network are used to train the electric power equipment image to obtain the first feature information; Input the features of the corresponding layer in the first feature into the Par Net network and the FPN network again for convolution to obtain the second feature information; The second feature information is combined with the first feature information to obtain a first feature image.

4. The electric power energy equipment identification method according to claim 3, characterized in that: The step of synchronizing the resolution of the first feature image to obtain a second feature image having the same resolution includes: Divide the first feature image into layers C1 to C5 according to the resolution from high to low, with the C3 layer being the middle layer; The resolution of layers C1 and C2 is reduced to the resolution of the middle layer C3 by downsampling; The resolutions of the C4 and C5 layers are increased to the resolution of the middle layer C3 by upsampling, thereby obtaining a second feature image with the same resolution.

5. The electric energy equipment identification method according to claim 3, characterized in that: The step of performing channel fusion on the second feature image to obtain a global information feature image includes: Reducing the feature dimension of the second feature image through a point-by-point convolutional layer; The feature information difference of adjacent pixels is increased through the standard convolution layer to obtain a global information feature image.

6. An electric power equipment identification device, characterized in that: include: A data acquisition module, used for acquiring images of electric power equipment, preprocessing the images of electric power equipment, and establishing an electric power equipment image data set; The electric power equipment image dataset includes a training set and a verification set; A training module, used for training a Mask-RCNN model for electric energy equipment recognition based on the training set and the validation set; The recognition module is used to input the real-scene picture of the electric energy equipment to be recognized into the Mask-RCNN model for electric energy equipment recognition to obtain the recognition result of the energy equipment on the real-scene picture; The training module is specifically used for: Extracting features from the electric energy equipment images in the training set to obtain a global information feature image; The RoIAlign layer is used to align the regions of interest on the global information feature image to obtain the output result of the Mask-RCNN prototype model. The output results of the Mask-RCNN prototype model are classified and regressed through the loss function to obtain the optimal parameters of the Mask-RCNN prototype model. The Mask-RCNN prototype model with the optimal parameters is used as the Mask-RCNN model for power energy equipment identification. The training module is specifically used for: Performing feature extraction on the electric power equipment images in the training set through the Par Net network and the RGFPN network to obtain first feature images with different resolutions; Performing resolution synchronization on the first feature image to obtain a second feature image with the same resolution; The second feature image is subjected to channel fusion to obtain a global information feature image.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the electric energy equipment identification method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the electric energy equipment identification method according to any one of claims 1 to 5 are implemented.