Recognition Model Training, Recognition Method and Device Based on Power Data Feature Extraction

By optimizing the YOLOv5 model, replacing the convolution with reverse depth separable convolution and adding the coordinate attention submodule, the problem of insufficient identification accuracy of lightweight models in the power industry is solved, and efficient real-time analysis and accurate detection on embedded devices are achieved.

CN114693963BActive Publication Date: 2025-07-18GLOBAL ENERGY INTERCONNECTION RES INST CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111529014.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-07-18
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the prior art, lightweight models are insufficiently used in image recognition in the power industry and cannot be analyzed in real time on embedded devices, resulting in untimely defect repair and loss of accuracy.

Method used

The optimized YOLOv5 model is adopted, and the k×k convolution in the convolution module and the Bottleneck submodule is replaced with the reverse depth separable convolution, and the coordinate attention submodule is added to the C3 module to form the optimized YOLOv5 model, which improves the recognition accuracy by reducing the amount of calculation.

Benefits of technology

It significantly reduces the calculation amount, and at the same time improves the recognition accuracy of the image recognition model, realizes efficient real-time analysis on embedded devices, and improves the accuracy of defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693963B_ABST
    Figure CN114693963B_ABST
Patent Text Reader

Abstract

The present invention provides a recognition model training, recognition method and device based on power data feature extraction. Among them, the recognition model training method based on power data feature extraction includes: obtaining a training data set; inputting the training data set into an optimized YOLOv5 model, and training the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolution module and a C3 module. The C3 module includes a Bottleneck sub-module. Replace the k×k (k>1) convolution in the convolution module and the Bottleneck sub-module with a depthwise separable convolution, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form an optimized YOLOv5 model. Replacing the convolution with a depthwise separable convolution significantly reduces the computational amount of the model. Moreover, adding a coordinate attention sub-module after the Bottleneck sub-module in the C3 module enhances the position sensitivity of the feature map after spatial fusion, and significantly improves the accuracy by increasing a small amount of computational amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly relates to a recognition model training, recognition method and device based on power data feature extraction. Background Art

[0002] Currently in the power industry, there has always been a need for intelligent computing and real-time feedback on the edge side using embedded devices. However, in some scenarios, the excessive amount of computation still seriously affects the practicality of intelligent edge devices. Taking the aerial images captured by a power transmission inspection drone as an example, it is required that the model maintain high recognition accuracy while using fewer computing resources, so that the inference model can only be deployed on a server or desktop computer with strong computing power support. In order to save manpower and time, in many current power scenarios, it is necessary to use embedded devices for real-time analysis and intelligent feedback on the edge side, such as automatic drone inspections. Some detection models with high accuracy usually require a large amount of computing resources. Limited by the resources of embedded devices, the current solution is to bring the pictures back after on-site shooting, and then perform inference on a computer or server with a GPU graphics card. After detecting the defects, corresponding maintenance personnel are dispatched to repair according to the detected defects. For such a maintenance, it is necessary to arrange staff to go to the site twice, and there is a certain time interval between the picture-taking time and the repair time, which is not conducive to the timely repair of defects. Although using a lightweight model can reduce the amount of computation, using a lightweight model reduces the amount of computation at the cost of sacrificing accuracy. Therefore, when using a lightweight model to process pictures, the obtained results have poor accuracy. Summary of the Invention

[0003] Therefore, the technical problem to be solved by the present invention is to overcome the defect that the results obtained by using a lightweight model to process pictures in the prior art have poor accuracy, so as to provide a recognition model training, recognition method and device based on power data feature extraction.

[0004] The first aspect of the present invention provides a recognition model training method based on power data feature extraction, including: obtaining a training data set; inputting the training data set into an optimized YOLOv5 model, and training the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolution module and a C3 module. The C3 module includes a Bottleneck sub-module. Replace the k×k (k>1) convolution in the convolution module and the Bottleneck sub-module with a depthwise separable convolution in the reverse direction, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form an optimized YOLOv5 model.

[0005] Optionally, in the recognition model training method based on power data feature extraction provided by the present invention, the depthwise separable convolution decomposes the three-dimensional convolution kernel in the convolution module and the Bottleneck sub-module into a pointwise convolution and a two-dimensional depthwise convolution; the pointwise convolution is used to fuse the input feature map at the channel position and input the feature map after channel fusion into the two-dimensional depthwise convolution; the two-dimensional depthwise convolution is used to fuse the feature map after channel fusion at the spatial position and output the feature map after spatial fusion.

[0006] Optionally, in the recognition model training method based on power data feature extraction provided by the present invention, the attention sub-module is used to locally fuse the features in the X and Y directions at the spatial position of the feature map after spatial fusion.

[0007] Optionally, in the recognition model training method based on power data feature extraction provided by the present invention, inputting the training data set into the optimized YOLOv5 model and training the optimized YOLOv5 model to obtain an image recognition model includes: inputting the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain the current optimized YOLOv5 model; after iterating one epoch, adding 1 to the value of m, inputting the validation data set into the current optimized YOLOv5 model, and calculating the recognition accuracy of the current optimized YOLOv5 model according to the output result of the current optimized YOLOv5 model; if the recognition accuracy of the current optimized YOLOv5 model is higher than the pre-stored highest accuracy, retaining the current optimized YOLOv5 model and replacing the pre-stored highest accuracy with the recognition accuracy of the current optimized YOLOv5 model; if the current loss function of the current optimized YOLOv5 model is greater than or equal to the pre-stored loss function, adding 1 to the value of n; if the value of m is greater than or equal to the first preset value, or the value of n is greater than or equal to the second preset value, determining the current optimized YOLOv5 model as the image recognition model.

[0008] Optionally, in the recognition model training method based on power data feature extraction provided by the present invention, if the recognition accuracy of the current optimized YOLOv5 model is less than or equal to the pre-stored highest accuracy, perform the step of adding 1 to the value of n if the current loss function of the current optimized YOLOv5 model is greater than or equal to the pre-stored loss function.

[0009] Optionally, in the recognition model training method based on power data feature extraction provided by the present invention, if the value of m is less than the first preset value and the value of n is less than the second preset value, return to the step of inputting the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain the current optimized YOLOv5 model.

[0010] The second aspect of the present invention provides a recognition method based on power data feature extraction, including: obtaining an image to be recognized; inputting the image to be recognized into an image recognition model to generate a recognition result, and the image recognition model is trained by the recognition model training method based on power data feature extraction provided in the first aspect of the present invention.

[0011] The third aspect of the present invention provides a recognition model training device based on power data feature extraction, including: a training data acquisition module for obtaining a training data set; a model training module for inputting the training data set into an optimized YOLOv5 model to train the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolutional module and a C3 module. The C3 module includes a Bottleneck sub-module. Replace the k×k (k>1) convolution in the convolutional module and the Bottleneck sub-module with a depthwise separable convolution, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form an optimized YOLOv5 model.

[0012] The fourth aspect of the present invention provides a recognition device based on power data feature extraction, including: an image to be recognized acquisition module for obtaining an image to be recognized; an image recognition module for inputting the image to be recognized into an image recognition model to generate a recognition result, and the image recognition model is trained by the recognition model training method based on power data feature extraction provided in the first aspect of the present invention.

[0013] The fifth aspect of the present invention provides a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to execute the recognition model training method based on power data feature extraction provided in the first aspect of the present invention, or the recognition method based on power data feature extraction provided in the second aspect of the present invention.

[0014] The technical solution of the present invention has the following advantages:

[0015] The recognition model training, recognition method and device based on power data feature extraction provided by the present invention use the optimized YOLOv5 model with the YOLOv5 model as the basic model. The k×k (k>1) convolution in the Bottleneck sub-module of the convolution module and the C3 module in the YOLOv5 model is replaced with a depthwise separable convolution, which significantly reduces the computational complexity of the model. After adding the coordinate attention sub-module, the position information ignored in the feature extraction process is marked in the convolution channel, and the orientation and feature information are used to strengthen the localization ability of the model in the spatial dimension. By introducing the coordinate attention sub-module, while increasing the computational complexity by less than 1%, the recognition effect of the image recognition model is significantly improved, and the recognition accuracy is improved. Also, because the coordinate attention sub-module is placed after the Bottleneck sub-module in the C3 module, the two-dimensional depth convolution in the depthwise separable convolution performs spatial fusion on the features in the feature map in the neighborhood, and the coordinate attention sub-module performs local fusion on the features in the X and Y directions in the spatial position of the feature map. Introducing the coordinate attention sub-module after the depthwise separable convolution in the C3 module enhances the position sensitivity of the feature map after spatial fusion, and significantly improves the accuracy by increasing a small amount of computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a flowchart (one) of a specific example of the recognition model training method based on power data feature extraction in an embodiment of the present invention;

[0018] Figure 2 It is a schematic diagram of the YOLOv5 model structure in an embodiment of the present invention;

[0019] Figure 3 It is a schematic diagram of the structure of the C3 module in the YOLOv5 model in an embodiment of the present invention;

[0020] Figure 4 It is a schematic diagram of the processing process of the input image by the three-dimensional convolution kernel in an embodiment of the present invention;

[0021] Figure 5 It is a schematic diagram of the processing process of the input image by the depthwise separable convolution in an embodiment of the present invention;

[0022] Figure 6It is a schematic structural diagram of the convolution module Re-dsc obtained by replacing the k×k (k>1) convolution in the YOLOv5 model with a depthwise separable convolution in the embodiments of the present invention;

[0023] Figure 7 It is a schematic structural diagram of the Bottleneck sub-module after replacing the k×k (k>1) convolution in the Bottlenek sub-module of the C3 module of the YOLOv5 model with a depthwise separable convolution in the embodiments of the present invention;

[0024] Figure 8 It is a schematic structural diagram of the coordinate attention sub-module in the embodiments of the present invention;

[0025] Figure 9 It is a schematic structural diagram of the C3 module after adding the coordinate attention sub-module coordatt to the C3 module in the embodiments of the present invention;

[0026] Figure 10 It is a schematic structural diagram of the optimized YOLOv5 model in the embodiments of the present invention;

[0027] Figure 11 It is a flowchart (two) of a specific example of the method for training an identification model based on power data feature extraction in the embodiments of the present invention;

[0028] Figure 12 It is a flowchart of a specific example of the identification method based on power data feature extraction in the embodiments of the present invention;

[0029] Figure 13 It is a schematic block diagram of a specific example of the device for training an identification model based on power data feature extraction in the embodiments of the present invention;

[0030] Figure 14 It is a schematic block diagram of a specific example of the identification device based on power data feature extraction in the embodiments of the present invention;

[0031] Figure 15 It is a schematic block diagram of a specific example of the computer device in the embodiments of the present invention. Detailed implementation manners

[0032] Next, the technical solutions of the present invention will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the protection scope of the present invention.

[0033] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0034] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0035] The embodiment of the present invention provides a method for training an identification model based on power data feature extraction, as Figure 1 shown, including:

[0036] Step S11: Obtain a training data set.

[0037] In an optional embodiment, the training data set includes multiple training images, and the training images are determined according to actual needs. Exemplarily, if the trained image recognition model is used to identify defects in transmission lines, the training images can be images including 21 categories such as insulator explosion images, bird's nest images, and shock absorber damage images. And, in order to facilitate the training of the initial model, the transmission line defects in each training image are labeled. Exemplarily, an open-source labeling software Labelme can be used to label the transmission line defects, and the labeling includes 21 categories.

[0038] Step S12: Input the training data set into the optimized YOLOv5 model, and train the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolution module and a C3 module. The C3 module includes a Bottleneck sub-module. Replace the k×k (k>1) convolution in the convolution module and the Bottleneck sub-module with a depthwise separable convolution, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form the optimized YOLOv5 model.

[0039] In the embodiment of the present invention, the YOLOv5 model includes various convolutions, and feature extraction is achieved through each convolution operation. If the trained image recognition model is used to identify defects in the power system, each convolution operation in the YOLOv5 model is used to extract power data features.

[0040] The YOLOv5 model includes various convolutions. In the embodiment of the present invention, only the k×k (k>1) convolution is replaced with a depthwise separable convolution.

[0041] In an optional embodiment, the optimized YOLOv5 model uses the YOLOv5 model as shown Figure 2 as the basic model, as shown in Figure 2In the YOLOv5 model shown, Conv is the convolutional module, C3 is the C3 module, and the structure of the C3 module is as shown in Figure 3 As shown, the C3 module includes the Bottlenek sub-module and multiple Conv sub-modules. In the embodiments of the present invention, not all Convs will be replaced, only the Convs that perform k×k (k>1) convolution will be replaced, and the Conv sub-modules in the C3 module will not be replaced either. Only the k×k (k>1) convolution in the Bottlenek sub-module will be replaced. In the conventional YOLOv5 model, three-dimensional convolutional kernels are used in both the convolutional module and the Bottleneck sub-module. The process of processing the input image through the three-dimensional convolutional kernel is as shown in Figure 4 As shown.

[0042] In the YOLOv5 model as shown in Figure 2 As shown, in different convolutional modules, the convolutional operations are different. Moreover, in the Bottleneck sub-module, there are multiple different types of convolutional operations. In the embodiments of the present invention, only the k×k (k>1) convolution is replaced with the inverse depthwise separable convolution. Exemplarily, the 6×6 and 3×3 convolutions in the YOLOv5 model are replaced with the inverse depthwise separable convolution.

[0043] The inverse depthwise separable convolution decouples the k×k convolution in YOLOv5 from the spatial and channel dimensions, decomposing it into a pointwise convolution that modifies the number of channels and a depth convolution in the two-dimensional space, turning the three-dimensional input feature into independent channel dimensions and two-dimensional plane features. The replaced inverse depthwise separable convolution is as shown in Figure 5 As shown.

[0044] In an alternative embodiment, after replacing the k×k (k>1) convolution in the convolutional module Conv of the YOLOv5 model shown in Figure 2 with the inverse depthwise separable convolution, a convolutional module Re-dsc is formed. The structure of the convolutional module Re-dsc is as shown in Figure 6 As shown. After replacing the k×k (k>1) convolution in the Bottlenek sub-module of the YOLOv5 model shown in Figure 2 with the inverse depthwise separable convolution, the structure of the replaced Bottlenek sub-module is as shown in Figure 7 As shown.

[0045] After replacing the k×k (k>1) convolution with the inverse depthwise separable convolution, first, a new feature is constructed by linearly combining the input channels through the pointwise convolution. Then, the feature maps are concatenated using the Concat operation for depth convolution. During the depth convolution, a convolutional kernel is assigned to each channel, and the dimension of the convolved feature map is used as the input feature map. The number of channels of the feature map after the depth convolution is the same as that of the input feature map, and the feature map cannot be expanded.

[0046] Since depth convolution performs convolution operations independently on each channel of the input layer and does not effectively utilize the feature information of different channels at the same spatial position. Therefore, pointwise convolution is needed to combine these feature maps in the channel dimension to generate new feature maps. In the embodiments of the present invention, depthwise separable convolution, which exchanges depth convolution and pointwise convolution, is used to replace the ordinary convolution in YOLOv5. The ordinary convolution kernel with k×k×in_channel×out_channel (k = 3 or k = 6) is decomposed into two parts. The first part is a 1×1×in_channel×out_channel pointwise convolution to fuse the feature maps in the channel dimension, and the second part is a k×k×in_channel depth convolution to fuse the feature maps in the spatial dimension.

[0047] Replacing the ordinary convolution with inverse depthwise separable convolution effectively reduces the amount of multiply-accumulate operations and the number of parameters of YOLOv5. Moreover, compared with the depthwise separable convolution that decomposes the ordinary convolution into depth convolution first and then pointwise convolution, the inverse depthwise separable convolution that decomposes the ordinary convolution into depth convolution and pointwise convolution in the embodiments of the present invention further reduces the amount of computation.

[0048] In the method for training an identification model based on power data feature extraction provided in the embodiments of the present invention, in order to balance the usage amount of computing resources and the accuracy of the model, after replacing the convolution in the YOLOv5 model with inverse depthwise separable convolution, a coordinate attention sub-module as shown in Figure 8 is added. The coordinate attention sub-module is a simple and efficient module. By embedding position information into channel attention, the model can obtain information in a larger area without introducing a large overhead, enabling the model to significantly improve the accuracy of the model while increasing a small amount of computation.

[0049] The coordinate attention sub-module decomposes channel attention into two parallel one-dimensional feature encodings to efficiently integrate spatial coordinate information into the generated attention map. Specifically, the coordinate attention sub-module first uses two one-dimensional global pooling operations to aggregate the input feature map along the vertical and horizontal directions into two separate feature maps containing specific direction information respectively; then encodes these two feature maps with specific direction information into two attention maps, each of which captures the long-range dependencies of the input feature map along one spatial direction, and thus the position information is preserved in the generated attention map; finally, both attention maps are applied to the input feature map through multiplication to emphasize the representation of the attention area.

[0050] As shown in Figure 8As shown, the coordinate attention sub-module performs average pooling on the horizontal and vertical directions respectively to obtain two one-dimensional vectors, then concatenates them in the spatial dimension (performs Concat operation) and uses 1×1 convolution to compress the channels, and then encodes the spatial information of the vertical and horizontal directions through batch normalization (uses BatchNorm operation) and non-linear activation (uses Non-linear operation). Next, the obtained feature map is separated (uses Split operation), and each passes through 1×1 convolution to obtain the same number of channels as the input feature Figure 1 and performs Sigmoid activation, and finally expands and applies it to the input feature map through multiplication. The lightweight 1×1 convolution is used in the coordinate attention sub-module to operate, mark the position information ignored in the feature extraction process in the convolution channels, and strengthen the positioning ability of the model in the spatial dimension through the direction and feature information. In two independent dimensions, the coordinate attention sub-module infers the attention map along different dimensions respectively and performs adaptive optimization based on the extraction of the feature map.

[0051] The attention map obtained by using the coordinate attention sub-module is sensitive to the direction and position information of the feature map.

[0052] In an optional embodiment, the coordinate attention sub-module is placed after the Bottleneck sub-module in the C3 module. The structure of the C3 module after introducing the coordinate attention is as Figure 9 shown. In Figure 9 the embodiment shown, coordatt is the coordinate attention sub-module, and Bottlenek is the Bottleneck sub-module that replaces the three-dimensional convolution with the inverse depthwise separable convolution. That is, after performing the inverse depthwise separable convolution on the feature image, the feature map is input into the coordinate attention sub-module, and the coordinate attention sub-module incorporates the position information into the feature map that has undergone channel fusion and spatial fusion.

[0053] In the embodiment of the present invention, the coordinate attention sub-module is introduced after the inverse depthwise separable convolution in the C3 module of YOLOv5. The two-dimensional depth convolution in the inverse depthwise separable convolution performs spatial fusion on the features in the feature map in the neighborhood, and the coordinate attention sub-module performs local fusion on the features in the X and Y directions of the feature map in the spatial position. It can be seen that introducing the coordinate attention sub-module after the inverse depthwise separable convolution in the C3 module enhances the position sensitivity of the feature map after spatial fusion.

[0054] In an optional embodiment, as Figure 2The convolutional module in the YOLOv5 model shown and the k×k (k>1) convolution in the Bottleneck sub-module are replaced with depthwise separable inverse convolutions, and an optimized YOLOv5 model is formed by adding a coordinate attention sub-module after the Bottleneck sub-module in the C3 module as Figure 10 shown.

[0055] In an optional embodiment, in order to verify that placing the coordinate attention sub-module after the Bottleneck sub-module in the C3 module has better performance, the embodiments of the present invention provide the test results when the coordinate attention sub-module is placed at different positions in the YOLOv5 model:

[0056] Table 1

[0057]

[0058]

[0059] In Table 1 above, yolov5s represents the original yolov5s model, yolov5s-dsc represents the result using depthwise separable convolutions, yolov5s-dscrev represents the result using depthwise separable inverse convolutions, yolov5s-dscrev-coordattConv represents placing the coordinate attention sub-module between the SiLU layer and the add layer in the Bottleneck as Figure 7 shown, yolov5s-dscrev-corrdattSPPF represents placing the coordinate attention sub-module in the SPPF as Figure 2 shown on the basis of yolov5s-dscrev, yolov5s-dscrev-corrdattC3_position 1 represents placing the coordinate attention sub-module between the Figure 3 shown concat layer and conv layer on the basis of yolov5s-dscrev, and yolov5s-dscrev-corrdattC3_position 2 represents placing the coordinate attention sub-module at the Figure 9 shown coordatt position.

[0060] The optimized YOLOv5 model used in the recognition model training method based on power data feature extraction provided by the embodiments of the present invention. The number of parameters of the lightweight YOLOv51 is reduced by about 63.86%, the multiply-accumulate operations are reduced by about 61.25%, and the mAP is only reduced by about 0.64%. The number of parameters of the lightweight YOLOv5m is reduced by about 60.02%, the multiply-accumulate operations are reduced by about 56.06%, and the mAP is only reduced by about 0.64%. The number of parameters of the lightweight YOLOv5s is reduced by about 54.28%, the multiply-accumulate operations are reduced by about 47.11%, and the mAP is only reduced by about 0.65%.

[0061] It can be seen from the above test results that placing the coordinate attention sub-module after the Bottleneck sub-module in the C3 module has better performance.

[0062] In the recognition model training method based on power data feature extraction provided by the embodiments of the present invention, the training data set is input into the optimized YOLOv5 model, and the optimized YOLOv5 model is trained to obtain an image recognition model. The optimized YOLOv5 model is based on the YOLOv5 model. The convolution in the convolution module and the Bottleneck sub-module in the C3 module of the YOLOv5 model is replaced with a depthwise separable convolution, which significantly reduces the computational amount of the model. After adding the coordinate attention sub-module, the position information ignored in the feature extraction process is marked on the convolution channel, and the positioning ability of the model is enhanced in the spatial dimension through the direction and feature information. By introducing the coordinate attention sub-module, while increasing the computational amount by less than 1%, the recognition effect of the image recognition model is significantly improved, the recognition accuracy is improved, and because the coordinate attention sub-module is placed after the Bottleneck sub-module in the C3 module, the two-dimensional depth convolution in the depthwise separable convolution performs spatial fusion on the features in the feature map in the neighborhood, and the coordinate attention sub-module performs local fusion on the features in the X direction and Y direction in the spatial position of the feature map. Introducing the coordinate attention sub-module after the depthwise separable convolution in the C3 module enhances the position sensitivity of the feature map after spatial fusion, and significantly improves the accuracy by increasing a small amount of computational amount.

[0063] In an optional embodiment, as Figure 11 shown, in the recognition model training method based on power data feature extraction provided by the embodiments of the present invention, the process of training the optimized YOLOv5 model with the training data set specifically includes:

[0064] Step S121: Input the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain the current optimized YOLOv5 model.

[0065] Step S122: After iterating for one epoch, increment the value of m by 1. Input the validation dataset into the current optimized YOLOv5 model, and calculate the recognition accuracy of the current optimized YOLOv5 model based on the output results of the current optimized YOLOv5 model.

[0066] In an alternative embodiment, the batch size during training can be set to 64, the loss threshold size of the intersection over union can be set to 0.5, and the initial learning rate can be set to 0.01.

[0067] Step S123: Determine whether the recognition accuracy of the current optimized YOLOv5 model is higher than the pre - stored highest accuracy. If it is determined that the recognition accuracy of the current optimized YOLOv5 model is higher than the pre - stored highest accuracy, execute Step S124; if it is determined that the recognition accuracy of the current optimized YOLOv5 model is less than or equal to the pre - stored highest accuracy, directly execute Step S125. In the embodiments of the present invention, the pre - stored highest accuracy is the highest accuracy obtained during previous multiple iterations.

[0068] Step S124: Retain the current optimized YOLOv5 model, and replace the pre - stored highest accuracy with the recognition accuracy of the current optimized YOLOv5 model.

[0069] Step S125: Determine whether the current loss function of the current optimized YOLOv5 model is less than the pre - stored loss function. If the current loss function of the current optimized YOLOv5 model is greater than or equal to the pre - stored loss function, execute Step S126; if the current loss function of the current optimized YOLOv5 model is less than the pre - stored loss function, set n to 1 and replace the pre - stored loss function with the current loss function, and then execute Step S127. In the embodiments of the present invention, if the current loss function is less than the pre - stored loss function, the pre - stored loss function is updated to the current loss function; if the current loss function is greater than the pre - stored loss function, the loss function is not updated.

[0070] Step S126: Increment the value of n by 1.

[0071] Step S127: Determine whether the value of m is greater than or equal to the first preset value, and whether the value of n is greater than or equal to the second preset value.

[0072] If the value of m is greater than or equal to the first preset value, or the value of n is greater than or equal to the second preset value, determine the current optimized YOLOv5 model as the image recognition model.

[0073] If the value of m is less than the first preset value and the value of n is less than the second preset value, return to the above Step S121.

[0074] In an embodiment of the present invention, the value of m represents the number of epochs of iteration, and n represents the number of epochs during which the loss function does not decrease continuously. When the number of epochs of iteration is greater than or equal to a first preset value, or, when the number of epochs during which the loss function is continuously not lower than the same pre-stored loss function is greater than or equal to a second preset value, the training is stopped, and the currently optimized YOLOv5 model is determined as the image recognition model.

[0075] In an alternative embodiment, the first preset value and the second preset value can be set according to actual requirements. Exemplarily, the first preset value can be set to 1000, and the second preset value can be set to 100. That is, when iterating 1000 epochs, or, when the loss function does not decrease for 100 consecutive epochs, the training is stopped, and the currently optimized YOLOv5 model is determined as the image recognition model.

[0076] An embodiment of the present invention provides a recognition method based on power data feature extraction, as Figure 12 shown, including:

[0077] Step S21: Obtain the image to be recognized.

[0078] In an alternative embodiment, the image to be recognized can be an image of a transmission line.

[0079] Step S22: Input the image to be recognized into the image recognition model to generate a recognition result. The image recognition model is trained by the recognition model training method based on power data feature extraction provided in any of the above embodiments.

[0080] In an alternative embodiment, when the image of the transmission line is input into the image recognition model, the generated recognition result includes the defects existing in the transmission line in the image.

[0081] An embodiment of the present invention provides a recognition model training device based on power data feature extraction, as Figure 13 shown, including:

[0082] The training data acquisition module 11 is used to acquire the training data set. For detailed content, refer to the description in the above method embodiment, and details will not be repeated here.

[0083] The model training module 12 is configured to input a training data set into the optimized YOLOv5 model, train the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolutional module and a C3 module. The C3 module includes a Bottleneck sub-module. Replace the convolutions in the convolutional module and the Bottleneck sub-module with depthwise separable convolutions, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form the optimized YOLOv5 model. For detailed content, refer to the description in the above method embodiments and will not be elaborated here.

[0084] An embodiment of the present invention provides an identification device based on power data feature extraction, as Figure 14 shown, including:

[0085] The image to be recognized acquisition module 21 is configured to acquire the image to be recognized. For detailed content, refer to the description in the above method embodiments and will not be elaborated here.

[0086] The image recognition module 22 is configured to input the image to be recognized into the image recognition model to generate a recognition result. The image recognition model is trained by the recognition model training method based on power data feature extraction provided in any of the above embodiments. For detailed content, refer to the description in the above method embodiments and will not be elaborated here.

[0087] An embodiment of the present invention provides a computer device, as Figure 15 shown. The computer device mainly includes one or more processors 31 and a memory 32, Figure 15 Taking one processor 31 as an example in

[0088] The computer device may further include: an input device 33 and an output device 34.

[0089] The processor 31, the memory 32, the input device 33, and the output device 34 may be connected through a bus or other means, Figure 15 Taking connection through a bus as an example in

[0090] The processor 31 may be a Central Processing Unit (CPU). The processor 31 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips. The general-purpose processor may be a microprocessor or any conventional processor, etc. The memory 32 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the recognition model training device based on power data feature extraction, or the recognition device based on power data feature extraction. In addition, the memory 32 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 32 may optionally include a memory remotely provided with respect to the processor 31, and these remote memories may be connected to the recognition model training device based on power data feature extraction, or the recognition device based on power data feature extraction through a network. The input device 33 may receive a calculation request input by a user (or other digital or character information), and generate a key signal input related to the recognition model training device based on power data feature extraction, or the recognition device based on power data feature extraction. The output device 34 may include a display device such as a display screen for outputting calculation results.

[0091] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the recognition model training method based on power data feature extraction, or the recognition method based on power data feature extraction in any of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disc, a Read-Only Memory (ROM), a Random Access Memory (RAM), a Flash Memory, a Hard Disk Drive (abbreviation: HDD), or a Solid-State Drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.

[0092] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. And the obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.

Claims

1. A method for training an identification model based on power data feature extraction, characterized in that Including: Obtain a training data set; Input the training data set into the optimized YOLOv5 model, and train the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolutional module and a C3 module. The C3 module includes a Bottleneck sub-module and multiple Conv sub-modules. Replace the k×k convolution in the convolutional module and the Bottleneck sub-module with a depthwise separable convolution, where k>1, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form the optimized YOLOv5 model; The depthwise separable convolution decomposes the three-dimensional convolution kernel in the convolutional module and the Bottleneck sub-module into a pointwise convolution and a two-dimensional depth convolution; The pointwise convolution is used to fuse the input feature map at the channel position and input the feature map after channel fusion into the two-dimensional depth convolution; The two-dimensional depth convolution is used to fuse the feature map after channel fusion at the spatial position and output the feature map after spatial fusion; The attention sub-module is used to locally fuse the features in the X direction and the Y direction of the feature map after spatial fusion at the spatial position.

2. The recognition model training method based on power data feature extraction according to claim 1, wherein, Inputting the training data set into the optimized YOLOv5 model and training the optimized YOLOv5 model to obtain an image recognition model includes: Input the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain the current optimized YOLOv5 model; After iterating one epoch, add 1 to the m value, input the validation data set into the current optimized YOLOv5 model, and calculate the recognition accuracy of the current optimized YOLOv5 model according to the output result of the current optimized YOLOv5 model; If the recognition accuracy of the current optimized YOLOv5 model is higher than the pre-stored highest accuracy, retain the current optimized YOLOv5 model and replace the pre-stored highest accuracy with the recognition accuracy of the current optimized YOLOv5 model; If the current loss function of the current optimized YOLOv5 model is greater than or equal to the pre-stored loss function, add 1 to the n value; If the m value is greater than or equal to the first preset value, or the n value is greater than or equal to the second preset value, determine the current optimized YOLOv5 model as the image recognition model.

3. The method for training an identification model based on power data feature extraction according to claim 2, wherein If the recognition accuracy of the current optimized YOLOv5 model is less than or equal to the pre-stored highest accuracy, perform the step of adding 1 to the n value if the current loss function of the current optimized YOLOv5 model is greater than or equal to the pre-stored loss function.

4. The method for training an identification model based on power data feature extraction according to claim 2 or 3, wherein If the value of m is less than the first preset value and the value of n is less than the second preset value, return the step of inputting the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain the current optimized YOLOv5 model.

5. An identification method based on power data feature extraction, characterized in that Including: Obtain the image to be recognized; Input the image to be recognized into the image recognition model to generate a recognition result, and the image recognition model is trained by the recognition model training method based on power data feature extraction according to any one of claims 1-4.

6. An identification model training device based on power data feature extraction, characterized in that, Including: A training data acquisition module for acquiring a training data set; A model training module for inputting the training data set into the optimized YOLOv5 model to train the optimized YOLOv5 model to obtain an image recognition model. The optimized YOLOv5 model uses the YOLOv5 model as the basic model. The YOLOv5 model includes a convolution module and a C3 module. The C3 module includes a Bottleneck sub-module and multiple Conv sub-modules. Replace the k×k convolution in the convolution module and the Bottleneck sub-module with a depthwise separable convolution, where k>1, and add a coordinate attention sub-module after the Bottleneck sub-module in the C3 module to form the optimized YOLOv5 model; The depthwise separable convolution decomposes the three-dimensional convolution kernel in the convolution module and the Bottleneck sub-module into a pointwise convolution and a two-dimensional depth convolution; The pointwise convolution is used to fuse the input feature map at the channel position and input the feature map after channel fusion into the two-dimensional depth convolution; The two-dimensional depth convolution is used to fuse the feature map after channel fusion at the spatial position and output the feature map after spatial fusion; The attention sub-module is used to locally fuse the features in the X direction and the Y direction of the feature map after spatial fusion at the spatial position.

7. An identification device based on power data feature extraction, characterized in that Including: An image to be recognized acquisition module for acquiring an image to be recognized; An image recognition module for inputting the image to be recognized into the image recognition model to generate a recognition result, and the image recognition model is trained by the recognition model training method based on power data feature extraction according to any one of claims 1-4.

8. A computer device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to execute the recognition model training method based on power data feature extraction according to any one of claims 1-4, or the recognition method based on power data feature extraction according to claim 5.

Citation Information

Patent Citations

  • Mask detection method based on yolov4

    CN113762201A