Target detection method and device based on improved YOLOv5 model, medium and electronic equipment
By introducing the CMH attention mechanism and improved loss function in the YOLOv5 model, the accuracy and robustness of small object detection in the distribution network scenario are solved, and more efficient object detection is achieved.
Patent Information
- Application Number
- CN202510207993.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing target detection methods are difficult to accurately detect small targets in distribution network scenarios, and there are misjudgments and misjudgments, and they are poor in versatility, so they cannot effectively deal with small target detection tasks in complex contexts.
The small object detection layer of the CMH attention mechanism is added to the backbone network of the YOLOv5 model, and the region of interest pooling method and the minimum point distance intersection loss function are used for model training, and the improved loss function is used to improve the border detection accuracy.
It improves the detection accuracy and accuracy of small targets, enhances the detection performance and robustness of the model in complex distribution network scenarios, and reduces missed and missed targets.
Smart Images

Figure CN120298655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly relates to an object detection method, device, storage medium and electronic device based on an improved YOLOv5 model. Background Art
[0002] With the continuous development of modern power systems, the distribution network, as a key link in power transmission, has been increasing in scale and complexity. The distribution network scenario covers numerous power equipment, lines, and related facilities, and these elements play a crucial role in ensuring stable power supply. Therefore, in the daily operation and maintenance, fault troubleshooting, and safety monitoring of the distribution network, it has become extremely important to accurately detect various small target objects (such as small power accessories, local line fault points, tiny safety hazard signs, etc.). For example, timely detection of tiny damaged components on distribution network equipment can prevent potential faults in advance and avoid affecting power supply; accurately identifying small target fault points on the line helps quickly locate and repair problems, reducing the power outage time and scope.
[0003] However, traditional detection methods, such as template matching and edge detection, although having certain applications in some simple scenarios, still face many challenges in the small target detection task in the distribution network scenario and are difficult to meet the actual needs. The distribution network environment is complex, with many background interference factors, and the characteristics of small target objects themselves are not obvious. These methods are difficult to accurately extract the effective features of small targets, resulting in low detection accuracy and the existence of misjudgment and missed judgment situations. At the same time, the generality of the detection algorithm is poor, and the algorithm needs to be redesigned and adjusted for different types of small targets, consuming a large amount of manpower, time, and computing costs.
[0004] In recent years, general object detection models represented by the YOLO series, Faster RCNN, etc. have achieved remarkable results in many image processing fields. However, although these algorithms have the advantages of fast detection speed and high accuracy, there are still obvious limitations when dealing with the small target detection task in the distribution network scenario. On the one hand, the small targets in the distribution network are often extremely small, occupying few pixels in the image. After the feature information of the small targets is extracted and processed by conventional methods, it is easy to be weakened or even lost, resulting in difficult accurate detection. On the other hand, the distribution network scenario has its own particularity, such as the existence of a large number of similar power equipment and line backgrounds, which easily confuses the target with the background. The model structure and algorithm mechanism lack sufficient pertinence and adaptability when distinguishing small targets in such a complex background. Summary of the Invention
[0005] In view of this, the present invention provides an object detection method, device, storage medium and electronic device based on an improved YOLOv5 model, mainly aiming to solve the problem that there is a possibility of data leakage at present, which reduces the security of data verification.
[0006] To solve the above problems, the present application provides an object detection method based on an improved YOLOv5 model, including:
[0007] Adding a small object detection layer with a CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0008] Based on a pre-acquired sample data set, using the region of interest alignment pooling method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain an object detection model;
[0009] Obtaining a power distribution network image to be object-detected;
[0010] Using the object detection model to perform object detection on the power distribution network image to obtain an object detection result corresponding to the power distribution network image.
[0011] Optionally, the step of using the region of interest pooling alignment method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model based on a pre-acquired sample data set to obtain an object detection model specifically includes:
[0012] Annotating the sample data set to obtain a labeled data set;
[0013] Using the region of interest pooling alignment method to perform object detection on the labeled data set by using the improved YOLOv5 model to obtain the position of the predicted bounding box, the predicted category of the target object and the predicted confidence score corresponding to the labeled data set;
[0014] Based on the position of the real target bounding box annotated in the labeled data set, the position of the predicted bounding box, the target object category label, the predicted category and the predicted confidence score, using the minimum point distance intersection over union loss function to perform calculation and processing to obtain a loss value;
[0015] Based on the loss value, using the backpropagation algorithm to adjust the parameters of the improved YOLOv5 model to obtain the current improved YOLOv5 model;
[0016] Based on the labeled data set, iteratively updating the current improved YOLOv5 model until the model converges, and determining the updated current improved YOLOv5 model as the object detection model.
[0017] Optionally, performing object detection on the distribution network image using the object detection model to obtain an object detection result corresponding to the distribution network image specifically includes:
[0018] Performing feature extraction on the distribution network image using the backbone network of the object detection model to obtain initial feature images of different scales;
[0019] Performing fusion on the initial feature images of different scales using the neck network of the object detection model to obtain fused images of different scales;
[0020] Performing object detection on the fused image using the detection head network of the object detection model to obtain an object detection result corresponding to the distribution network image.
[0021] Optionally, performing feature extraction on the distribution network image using the backbone network of the object detection model to obtain initial feature images of different scales specifically includes:
[0022] Performing slicing operation on the distribution network image to obtain a first feature image;
[0023] Performing feature extraction on the first feature image using multiple convolutional neural networks to obtain multiple second feature images of different scales;
[0024] Performing feature extraction on the second feature image output by the convolutional neural network adjacent to the small object detection layer using the channel attention module of the small object detection layer to obtain a channel feature image;
[0025] Performing feature extraction on the second feature image output by the convolutional neural network adjacent to the small object detection layer using the spatial attention module of the small object detection layer to obtain a spatial feature image;
[0026] Performing splicing based on the channel feature image and the spatial feature image to obtain a third feature image;
[0027] The initial feature images include the first feature images of different scales, each of the second feature images, and the third feature image.
[0028] Optionally, performing feature extraction on the second feature image using the channel attention module of the small object detection layer to obtain a channel feature image specifically includes:
[0029] Performing global average pooling on the second feature image to obtain a first feature vector;
[0030] Performing global max pooling on the second feature image to obtain a second feature vector;
[0031] Perform one-dimensional convolution calculation based on the first feature vector, the second feature vector, and a preset activation function to obtain a channel attention weight vector;
[0032] Perform a multiplication operation based on the channel attention weight vector and the second feature image to obtain the channel feature image.
[0033] Optionally, the feature extraction is performed on the second feature image by using the spatial attention module of the small object detection layer to obtain a spatial feature image, which specifically includes:
[0034] Separate the channels of the second feature image to obtain each channel-separated feature image;
[0035] Perform global average pooling on the same channel-separated feature image to obtain a first separated feature vector;
[0036] Perform global max pooling on the same channel-separated feature image to obtain a second separated feature vector;
[0037] Perform shared convolution calculation based on the first separated feature vector and the second separated feature vector to obtain the spatial attention weight corresponding to the same channel-separated feature image;
[0038] Perform calculation based on the channel-separated feature image and the spatial attention weight respectively to obtain an initial spatial feature image corresponding to each channel-separated feature image;
[0039] Perform splicing on each initial spatial feature image to obtain the spatial feature image.
[0040] Optionally, the detection head network of the object detection model is used to perform object detection on the fusion image to obtain an object detection result corresponding to the power distribution network image, which specifically includes:
[0041] Identify the fusion image to obtain a region of interest;
[0042] Use the bilinear interpolation algorithm to perform interpolation calculation on the region of interest to obtain an initial target region;
[0043] Perform average pooling on each sub-region in the initial target region to obtain an initial fusion feature vector corresponding to each sub-region;
[0044] Perform feature fusion based on each initial fusion feature vector to obtain a fusion feature vector;
[0045] Perform detection based on the fusion feature vector by using the detection head network of the object detection model to obtain an object detection result corresponding to the power distribution network image;
[0046] Among them, the object detection result includes the position coordinates of the object prediction box, the object category probability, and the prediction confidence.
[0047] To solve the above problems, the present application provides an object detection device based on an improved YOLOv5 model, including:
[0048] An adding module, configured to add a small object detection layer with a CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0049] A training module, configured to perform model training on the improved YOLOv5 model based on a pre-acquired sample data set by using the region of interest alignment pooling method and the minimum point distance intersection over union loss function to obtain an object detection model;
[0050] An obtaining module, configured to obtain a power distribution network image to be object-detected;
[0051] A detection module, configured to perform object detection on the power distribution network image by using the object detection model to obtain an object detection result corresponding to the power distribution network image.
[0052] To solve the above problems, the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above object detection method based on the improved YOLOv5 model are implemented.
[0053] To solve the above problems, the present application provides an electronic device including at least a memory and a processor, where a computer program is stored on the memory, and when the processor executes the computer program on the memory, the steps of the above object detection method based on the improved YOLOv5 model are implemented.
[0054] The beneficial effects in the present application: In the present application, a small object detection head with a CMH attention mechanism is introduced through the present invention, which is used to analyze and process the features of tiny objects, assign different weights to different channels and regions, enhance the expression of important features and the suppression of irrelevant features by the model, thereby effectively improving the detection accuracy and precision of tiny objects; the loss function is changed to the MPDIoU regression loss function, and a more appropriate aspect ratio is used to improve the detection accuracy of the bounding box, avoid problems such as missed detection and false detection of objects caused by traditional processing methods, and enhance the detection performance and robustness in complex power distribution network scenarios.
[0055] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. Description of the Drawings
[0056] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Also, throughout the drawings, the same reference numerals are used to denote the same components. In the drawings:
[0057] Figure 1 A schematic flowchart of a target detection method based on an improved YOLOv5 model provided by an embodiment of the present application is shown;
[0058] Figure 2 A model structure diagram of the improved YOLOv5 model provided by an embodiment of the present application is shown;
[0059] Figure 3 A schematic flowchart of a target detection method based on an improved YOLOv5 model provided by another embodiment of the present application is shown;
[0060] Figure 4 A structural block diagram of a target detection method based on an improved YOLOv5 model provided by another embodiment of the present application is shown. Detailed Embodiments
[0061] Reference is made herein to the various aspects and features of the present application with reference to the drawings.
[0062] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be considered as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.
[0063] The drawings included in the specification and forming a part of the specification show the embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, are used to explain the principles of the present application.
[0064] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the drawings.
[0065] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.
[0066] When combined with the drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.
[0067] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments claimed are merely examples of the present application and can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but are merely a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.
[0068] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", each of which may refer to one or more of the same or different embodiments according to the present application.
[0069] An embodiment of the present application provides an object detection method based on an improved YOLOv5 model, as Figure 1 shown, including:
[0070] Step S101: Add a small object detection layer with a CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0071] In a specific implementation process, the basic YOLOv5 model includes a backbone network, a neck network, and a detection head network. As Figure 2 shown, based on the basic YOLOv5 model, the present application adds a small object detection layer introducing a CMH attention mechanism after the last feature extraction layer to obtain an improved YOLOv5 model. A small object detection layer applying a CMH attention mechanism is added at the end position of the basic YOLOv5 backbone network to process and analyze the features of tiny objects. By adjusting the weights of channel and spatial information, important feature information is focused, and unimportant information is suppressed, so that the model can effectively extract and express key features, improving the detection accuracy and precision of tiny objects.
[0072] Step S102: Based on a pre-acquired sample data set, use the region of interest pooling method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain an object detection model;
[0073] In a specific implementation process, the sample data set is labeled to obtain a labeled data set;
[0074] Perform object detection on the label dataset based on the improved YOLOv5 model to obtain object prediction bounding boxes, object prediction categories, and confidence levels corresponding to the label dataset; calculate and process the real object bounding boxes annotated in the label dataset and the object prediction bounding boxes using the minimum point distance intersection over union loss function to obtain a loss value; adjust the parameters of the improved YOLOv5 model based on the loss value using the backpropagation algorithm to obtain the current improved YOLOv5 model; iteratively update the current improved YOLOv5 model based on the label dataset until the model converges, and determine the updated current improved YOLOv5 model as the object detection model.
[0075] Step S103: Obtain a power distribution network image to be object-detected.
[0076] In a specific implementation process, the power distribution network image can be an image of the power distribution area collected by a drone, or a power distribution network image collected by a fixed camera, an inspection robot, a remote sensing satellite, an infrared camera, etc. The power distribution network image includes images of one or more target objects such as power equipment, lines, and power facilities. Performing object detection on the power distribution network image can detect and identify small targets such as small power accessories, local line fault points, and tiny safety hazard signs during the operation of the power system, which helps to quickly locate and repair problems and reduce the power outage time and scope.
[0077] Step S104: Perform object detection on the power distribution network image using the object detection model to obtain an object detection result corresponding to the power distribution network image.
[0078] In a specific implementation process, use the backbone network of the object detection model to extract features from the power distribution network image to obtain initial feature images of different scales; use the neck network of the object detection model to fuse the initial feature images of different scales to obtain fused images of different scales; perform a pooling operation on the fused images using the region of interest alignment pooling method to obtain fused feature vectors corresponding to the fused images; use the detection head network of the object detection model to perform object detection on the fused feature vectors to obtain an object detection result corresponding to the power distribution network image. The object detection result includes the position coordinates of each object prediction bounding box, the object category probability corresponding to each object prediction bounding box, and the prediction confidence level for accurately predicting each target object.
[0079] This application introduces a small target detection head with CMH attention mechanism through the present invention, which is used to analyze and process the features of small targets, assign different weights to different channels and regions, enhance the expression of important features and the suppression of irrelevant features by the model, so as to effectively improve the detection accuracy and precision of small targets; change the loss function to the MPDIoU regression loss function, use a more appropriate aspect ratio, improve the detection accuracy of the bounding box, avoid problems such as target missed detection and misdetection caused by traditional processing methods, and enhance the detection performance and robustness in complex distribution network scenarios.
[0080] Another embodiment of this application provides another object detection method based on the improved YOLOv5 model, as Figure 3 shown, including:
[0081] Step S201: Add a small target detection layer with CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0082] In the specific implementation process, the basic YOLOv5 model includes a backbone network, a neck network, and a detection head network. Based on the basic YOLOv5 model, a small target detection layer CMH layer with CMH attention mechanism is added after the last feature extraction layer to obtain an improved YOLOv5 model.
[0083] Step S202: Based on the pre-obtained sample data set, use the region of interest pooling alignment method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain an object detection model;
[0084] In the specific implementation process, label the sample data set to obtain a label data set; each sample image in the sample data set is a distribution network image containing one or more target objects, and the target objects include power equipment, power facilities, and power lines, etc. The sample data set can be labeled using annotation tools such as LabelImg or CVAT or by manual annotation. The annotation information includes the target object bounding box and the category label corresponding to the target object; based on the improved YOLOv5 model, use the region of interest pooling alignment method to perform object detection on the label data set to obtain the target prediction box, target prediction category, and confidence corresponding to the label data set; input the label data set into the improved YOLOv5 model, and the model outputs the target prediction box, target prediction category, and confidence corresponding to each annotated image in the label data set; calculate and process the true target box labeled based on the label data set and the target prediction box using the minimum point distance intersection over union loss function to obtain a loss value; the mathematical expression of the minimum point distance intersection over union loss function MPDIoU loss function can be shown by the following formula (1):
[0085]
[0086] Among them, IoU is the intersection over union of the predicted bounding box and the target bounding box, and α and β are preset weight coefficients used to adjust the relative importance of the two factors of the center distance and the aspect ratio in the loss function. d is the center distance between the predicted bounding box and the target bounding box, and r p and r are the aspect ratios of the width and height of the predicted bounding box and the target bounding box respectively, and w t and h t are the width and height of the target bounding box. Based on the loss value, the parameters of the improved YOLOv5 model are adjusted using the backpropagation algorithm to obtain the current improved YOLOv5 model; the current improved YOLOv5 model is iteratively updated based on the label dataset until the model converges, and the updated current improved YOLOv5 model is determined as the target detection model.
[0087] Step S203: Obtain the power distribution network image to be target-detected;
[0088] In the specific implementation process, the power distribution network image can be an image of the power distribution area collected by a drone, or a power distribution network image collected by a fixed camera, an inspection robot, a remote sensing satellite, an infrared camera, etc. The power distribution network image includes images of one or more target objects such as power equipment, lines, and power facilities. Target detection of the power distribution network image can detect and identify small targets such as small power accessories, local line fault points, and tiny safety hazard signs during the operation of the power system, which helps to quickly locate and repair problems and reduce the power outage time and scope.
[0089] Step S204: Perform a slicing operation on the power distribution network image to obtain a first feature image;
[0090] In the specific implementation process, the power distribution network image is sliced at the focus layer of the main feature extraction network of the target detection model to obtain a first feature image. Specifically, the spatial size of the input image is halved, and the number of channels is doubled at the same time to obtain the first feature image, which prepares for subsequent feature extraction.
[0091] Step S205: Based on the first feature image, use multiple convolutional neural networks for feature extraction to obtain multiple second feature images of different scales;
[0092] In the specific implementation process, the feature map with more channel information and reduced size undergoes feature extraction through a series of ordinary convolutional layers and a CBL (standard convolutional module) composed of Conv+BN+Leaky_relu, enabling the network to learn a more advanced feature representation and obtaining multiple second feature images of different scales.
[0093] Step S206: Based on the second feature image output by the convolutional neural network adjacent to the small target detection layer, use the channel attention module of the small target detection layer to extract features to obtain a channel feature image;
[0094] In the specific implementation process, perform global average pooling on the second feature image in the CMH layer, which is the small target detection layer of the main feature extraction network, to obtain a first feature vector; the mathematical formula for calculating the first feature vector can be shown by the following formula (2):
[0095]
[0096] Where, is the feature value after global average pooling for the c-th channel, and F c,i,j is the feature value of the c-th channel, the i-th row, and the j-th column in the original feature map F; perform global max pooling on the second feature image to obtain a second feature vector; the mathematical expression of the second feature vector can be shown by the following formula (3):
[0097]
[0098] Where, F c ' ,i,j is the feature value after global max pooling for the c-th channel. Perform one-dimensional convolution calculation based on the first feature vector, the second feature vector, and a preset activation function to obtain a channel attention weight vector; the two feature vectors obtained through global average pooling and global max pooling are aggregated in a tensor of the same size and are respectively input into a fast one-dimensional convolution. Feature extraction is performed through convolution operations, and a sigmoid activation function is applied for non-linear transformation: the mathematical formula for calculating the channel attention weight vector can be shown by the following formula (4):
[0099]
[0100] Where, σ is the activation function, is a one-dimensional convolution with a convolution kernel size of 1.
[0101] Perform multiplication operation based on the channel attention weight vector and the second feature image to obtain the channel feature image. Multiply the original input feature map F element-wise with the calculated channel attention weight vector M c to obtain the feature map F' after being processed by the channel attention module. The feature map F' after being processed by the channel attention module is used as the input of the spatial attention module, and its dimension is still C×H×W: the mathematical formula for the channel feature image is shown by the following formula (5):
[0102]
[0103] Step S207: Based on the second feature image output by the convolutional neural network adjacent to the small target detection layer, use the spatial attention module of the small target detection layer to extract features to obtain a spatial feature image;
[0104] In the specific implementation process, the second feature image is separated by channels according to the importance degree to obtain each channel-separated feature image; the channel-separated feature image includes an important feature image F1 and a secondary feature image F2; perform global average pooling on the same channel-separated feature image to obtain a first separated feature vector F avg,t ; perform global max pooling on the same channel-separated feature image to obtain a second separated feature vector F max,t ; The calculation mathematical expressions of the first separated feature vector and the second separated feature vector can be shown by the following formula (6):
[0105]
[0106] where F avg,t , F max,t are the feature values after global average pooling and max pooling of the spatial positions, and c is the number of channels; perform shared convolution calculation processing based on the first separated feature vector and the second separated feature vector to obtain the spatial attention weight corresponding to the same channel-separated feature image; the spatial attention weight includes a first spatial attention weight M s,1 calculated based on the first separated feature vector and a second spatial attention weight M s,2 calculated based on the second separated feature vector; the calculation mathematical expression of the spatial attention weight can be shown by the following formula (7):
[0107]
[0108] where b is the bias weight;
[0109] Calculate based on the channel-separated feature image and the spatial attention weight respectively to obtain an initial spatial feature image corresponding to each channel-separated feature image; specifically, calculate based on the important feature image and the first spatial attention weight to obtain an initial spatial feature image F1' corresponding to the important feature image; the mathematical expression is shown by the following formula (8):
[0110]
[0111] Calculation and processing are performed based on the secondary feature image and the second spatial attention weight to obtain an initial spatial feature image F2' corresponding to the secondary feature image; the mathematical expression is shown in the following formula (9):
[0112]
[0113] The initial spatial feature images are stitched to obtain the spatial feature image. The mathematical expression of the spatial feature image can be shown in the following formula (10):
[0114]
[0115] Step S208: The channel feature image and the spatial feature image are stitched to obtain a third feature image, so as to obtain the initial feature images of different scales;
[0116] In a specific implementation process, after parallel processing by the channel attention module and the spatial attention module, the obtained feature map F out is the output feature map after applying the CMH attention mechanism. By screening and strengthening key feature information, it helps the model to more effectively support the small target detection task: the mathematical expression of the third feature image can be shown in the following formula (11):
[0117]
[0118] The initial feature images include the first feature images of different scales, each of the second feature images, and the third feature image. For example: when the input distribution network image of 640×640×3 is first sliced by the Focus module, the pixel size becomes half of the original image, and the number of channels becomes 4 times the original, and the size of the obtained first feature image is 160 by 160; the feature map with more channel information and reduced size undergoes feature extraction through a series of ordinary convolutional layers and a CBL (standard convolutional module) composed of Conv+BN+Leaky_relu, enabling the network to learn more advanced feature representations to obtain the second feature images of different scales; the sizes of each of the second feature images are 80×80, 40×40, etc.; based on the channel feature image and the spatial feature image, the size of the stitched third feature image is 20×20; these initial feature images of different sizes are used to detect targets of different sizes.
[0119] Step S209: The neck network of the object detection model is used to fuse the initial feature images of different scales to obtain fused images of different scales;
[0120] In the specific implementation process, the neck network performs upsampling on the initial feature image to obtain a feature vector and conducts feature fusion, so as to better provide rich and effective feature information for the subsequent detection head module, thereby improving the accuracy of object detection. Specifically, the backbone network outputs an initial feature image with a smaller scale. The neck network first performs an upsampling operation on the initial feature image to enlarge the size of the feature map so that it can be stitched with other initial feature images of different scales output from the backbone network. On the stitched feature image, the CBL module further extracts and refines features. To better adapt to the object detection requirements of different scales, through downsampling and stitching operations with other feature maps processed at different scales, further feature fusion is performed to obtain global information, so that the model can more comprehensively utilize features of multiple scales for object detection.
[0121] Step S210: Use the detection head network of the object detection model to perform object detection on the fused image to obtain an object detection result corresponding to the network configuration image.
[0122] In the specific implementation process, based on the fused image, a pooling operation is performed using the region of interest alignment pooling method to obtain a fused feature vector corresponding to the fused image; the fused image is recognized to obtain a region of interest; the region of interest is the region covered by the predicted bounding box; the bilinear interpolation algorithm is used to perform interpolation calculation on the region of interest to obtain an initial target region. Specifically, each initial target region is evenly divided into k×k sub-regions; within each sub-region, the bilinear interpolation method is used to calculate the feature values at the 4 vertices of the sub-region. Based on the coordinates and feature values of the 4 feature points around the center, the feature value f(x,y) of the regional target is calculated, thereby avoiding the error caused by direct rounding. For each sub-region, let the feature values at the 4 vertices be f(0,0), f(0,1), f(1,0), and f(1,1) respectively. Then the mathematical expression of the bilinear interpolation algorithm can be shown by the following formula (12):
[0123] f(x,y) = (1 - x)(1 - y)f(0,0) + (x - 0)(1 - y)f(0,1) + (1 - x)(y - 0)f(1,0) + (x - 0)(y - 0)f(1,1) (12)
[0124] Complete bilinear interpolation to calculate the eigenvalues of the vertices of each sub-region, thereby reducing quantization errors. Perform average pooling on the initial target region to obtain an initial fused feature vector; specifically, after calculating the eigenvalues of the sub-region vertices, perform average pooling on these eigenvalues to obtain an initial fused feature vector corresponding to each sub-region. Perform feature fusion based on each of the initial fused feature vectors to obtain the fused feature vector. Concatenate the pooling results of all sub-regions into a fused feature vector of a fixed size. This fused feature vector contains rich information about the target region in the original feature map. The object detection model predicts the position of the target through the anchor box mechanism. Each grid will predict multiple anchor boxes of different scales and ratios. These anchor boxes are analogized to candidate regions to a certain extent. The model will screen and adjust these anchor boxes according to the predicted confidence and bounding box offset. Calculate the eigenvalues of the sub-region vertices through bilinear interpolation and perform average pooling to reduce quantization errors. The ROI Align pooling operation concatenates the pooling results of all sub-regions into a fused feature vector of a fixed size, meeting the requirements of the detection head for the input feature format and providing it as a feature output to the detection head. Perform detection using the detection head network of the object detection model based on the fused feature vector to obtain an object detection result corresponding to the distribution network image; wherein, the object detection result includes the position coordinates of the target prediction box of the target object, the target category probability, and the prediction confidence.
[0125] In this application, a small object detection head with a CMH attention mechanism is introduced through the present invention to analyze and process the features of small objects, assign different weights to different channels and regions, enhance the model's expression of important features and suppression of irrelevant features, thereby effectively improving the detection accuracy and precision of small objects; change the loss function to the MPDIoU regression loss function, use a more appropriate aspect ratio, improve the detection accuracy of the bounding box, avoid problems such as missed detection and misdetection of targets caused by traditional processing methods, and enhance the detection performance and robustness in complex distribution network scenarios.
[0126] Another embodiment of this application provides an object detection device based on an improved YOLOv5 model, as Figure 4 shown, including:
[0127] An adding module 1, configured to add a small object detection layer with a CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0128] A training module 2, configured to perform model training on the improved YOLOv5 model based on a pre-acquired sample data set by using the region of interest alignment pooling method and the minimum point distance intersection over union loss function to obtain an object detection model;
[0129] An obtaining module 3, configured to obtain a distribution network image to be object-detected;
[0130] The detection module 4 is configured to perform object detection on the distribution network image by using the target detection model, so as to obtain a target detection result corresponding to the distribution network image.
[0131] In a specific implementation process, the training module 2 is specifically configured to: label the sample data set to obtain a labeled data set; use the improved YOLOv5 model to perform object detection on the labeled data set by using the region of interest pooling alignment method, so as to obtain the position of the predicted bounding box, the predicted category of the target object, and the predicted confidence score corresponding to the labeled data set; perform calculation processing by using the minimum point distance intersection over union loss function based on the position of the true target bounding box labeled in the labeled data set, the position of the predicted bounding box, the target object category label, the predicted category, and the predicted confidence score, so as to obtain a loss value; perform parameter adjustment on the improved YOLOv5 model by using the backpropagation algorithm based on the loss value to obtain the current improved YOLOv5 model; perform iterative update on the current improved YOLOv5 model based on the labeled data set until the model converges, and determine the updated current improved YOLOv5 model as the target detection model.
[0132] In a specific implementation process, the detection module 4 is specifically configured to: extract features from the distribution network image by using the backbone network of the target detection model to obtain initial feature images of different scales, where the initial feature images are detected by the small object detection layer with the CMH attention mechanism; fuse the initial feature images of different scales by using the neck network of the target detection model to obtain fused images of different scales; perform object detection on the fused images by using the detection head network of the target detection model to obtain a target detection result corresponding to the distribution network image.
[0133] In a specific implementation process, the detection module 4 is further configured to: perform slicing operation on the distribution network image to obtain a first feature image; extract features from the first feature image by using a plurality of convolutional neural networks to obtain second feature images of different scales; extract features from the second feature image output by the convolutional neural network adjacent to the small object detection layer by using the channel attention module of the small object detection layer to obtain a channel feature image; extract features from the second feature image output by the convolutional neural network adjacent to the small object detection layer by using the spatial attention module of the small object detection layer to obtain a spatial feature image; splice the channel feature image and the spatial feature image to obtain a third feature image; the initial feature images include the first feature images of different scales, each of the second feature images, and the third feature image.
[0134] In the specific implementation process, the detection module 4 is further configured to: perform global average pooling on the second feature image to obtain a first feature vector; perform global max pooling on the second feature image to obtain a second feature vector; perform one-dimensional convolution calculation based on the first feature vector, the second feature vector, and a preset activation function to obtain a channel attention weight vector; perform multiplication operation based on the channel attention weight vector and the second feature image to obtain the channel feature image.
[0135] In the specific implementation process, the detection module 4 is further configured to: perform channel separation on the second feature image to obtain each channel-separated feature image; perform global average pooling on the same channel-separated feature image to obtain a first separated feature vector; perform global max pooling on the same channel-separated feature image to obtain a second separated feature vector; perform shared convolution calculation based on the first separated feature vector and the second separated feature vector to obtain the spatial attention weight corresponding to the same channel-separated feature image; perform calculation based on the channel-separated feature image and the spatial attention weight respectively to obtain an initial spatial feature image corresponding to each channel-separated feature image; perform splicing processing on each initial spatial feature image to obtain the spatial feature image.
[0136] In the specific implementation process, the detection module 4 is further configured to: identify the region of interest from the fused image; perform interpolation calculation on the region of interest using the bilinear interpolation algorithm to obtain an initial target region; perform average pooling on each sub-region in the initial target region to obtain an initial fusion feature vector corresponding to each sub-region; perform feature fusion based on each initial fusion feature vector to obtain a fusion feature vector; perform detection using the detection head network of the target detection model based on the fusion feature vector to obtain a target detection result corresponding to the power distribution network image; wherein, the target detection result includes the position coordinates of the target prediction box, the target category probability, and the prediction confidence.
[0137] This application introduces a small target detection head with CMH attention mechanism through the present invention, which is used to analyze and process the features of small targets, assign different weights to different channels and regions, enhance the model's expression of important features and suppression of irrelevant features, thereby effectively improving the detection accuracy and precision of small targets; change the loss function to the MPDIoU regression loss function, use a more appropriate aspect ratio, improve the detection accuracy of the bounding box, avoid problems such as target missed detection and misdetection caused by traditional processing methods, and enhance the detection performance and robustness in complex power distribution network scenarios.
[0138] Another embodiment of the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the following method steps are implemented:
[0139] Step 1: Add a small target detection layer with CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0140] Step 2: Based on the pre-acquired sample data set, use the region of interest alignment pooling method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain a target detection model;
[0141] Step 3: Obtain a power distribution network image to be target-detected;
[0142] Step 4: Use the target detection model to perform target detection on the power distribution network image to obtain a target detection result corresponding to the power distribution network image.
[0143] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0145] For the specific implementation process of the above method steps, reference can be made to the embodiments of any of the above object detection methods based on the improved YOLOv5 model, and this embodiment will not be repeated here.
[0146] In this application, a small target detection head with a CMH attention mechanism is introduced through the present invention to analyze and process the features of tiny targets, assign different weights to different channels and regions, enhance the model's expression of important features and suppression of irrelevant features, thereby effectively improving the detection accuracy and precision of tiny targets; the loss function is changed to the MPDIoU regression loss function, using a more appropriate aspect ratio to improve the detection accuracy of the bounding box, avoiding problems such as missed detection and misdetection of targets caused by traditional processing methods, and enhancing the detection performance and robustness in complex distribution network scenarios.
[0147] Another embodiment of this application provides an electronic device, which can be a server. The electronic device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external client through a network connection. When the electronic device program is executed by the processor, it realizes the functions or steps of the server side of an object detection method based on the improved YOLOv5 model.
[0148] In one embodiment, an electronic device is provided, which can be a client. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes non-volatile storage media and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external server through a network connection. When the electronic device program is executed by the processor, it realizes the functions or steps of the client side of an object detection method based on the improved YOLOv5 model.
[0149] Another embodiment of this application provides an electronic device, at least including a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, the following method steps are realized:
[0150] Step 1: Add a small target detection layer with CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model;
[0151] Step 2: Based on the pre-obtained sample data set, use the region of interest alignment pooling method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain a target detection model;
[0152] Step 3: Obtain a network configuration image to be target-detected;
[0153] Step 4: Use the target detection model to perform target detection on the network configuration image to obtain a target detection result corresponding to the network configuration image.
[0154] For the specific implementation process of the above method steps, refer to the embodiments of any of the above target detection methods based on the improved YOLOv5 model. This embodiment will not be repeated here.
[0155] In this application, by introducing a small target detection head with CMH attention mechanism in the present invention, it is used to analyze and process the features of small targets, assign different weights to different channels and regions, enhance the model's expression of important features and suppression of irrelevant features, thereby effectively improving the detection accuracy and precision of small targets; change the loss function to the MPDIoU regression loss function, use a more appropriate aspect ratio, improve the detection accuracy of the bounding box, avoid problems such as missed detection and misdetection of targets caused by traditional processing methods, and enhance the detection performance and robustness in complex network configuration scenarios.
[0156] The above embodiments are only exemplary embodiments of this application and are not used to limit this application. The protection scope of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of this application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of this application.
Claims
1. A target detection method based on an improved YOLOv5 model, characterized in that, Including: Adding a small target detection layer with the CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model; Based on the pre-acquired sample data set, using the region of interest alignment pooling method and the minimum point distance intersection over union loss function to train the improved YOLOv5 model to obtain a target detection model; Obtaining a network configuration image to be target-detected; Using the target detection model to perform target detection on the network configuration image to obtain a target detection result corresponding to the network configuration image.
2. The method according to claim 1, characterized in that, The step of training the improved YOLOv5 model based on the pre-acquired sample data set using the region of interest pooling alignment method and the minimum point distance intersection over union loss function to obtain a target detection model specifically includes: Annotating the sample data set to obtain a labeled data set; Using the region of interest pooling alignment method to perform target detection on the labeled data set using the improved YOLOv5 model to obtain the predicted box positions, the predicted categories of the target objects, and the predicted confidence scores corresponding to the labeled data set; Based on the true target box positions annotated in the labeled data set, the predicted box positions, the target object category labels, the predicted categories, and the predicted confidence scores, using the minimum point distance intersection over union loss function to perform calculation and processing to obtain a loss value; Based on the loss value, using the backpropagation algorithm to adjust the parameters of the improved YOLOv5 model to obtain the current improved YOLOv5 model; Based on the labeled data set, iteratively updating the current improved YOLOv5 model until the model converges, and determining the updated current improved YOLOv5 model as the target detection model.
3. The method according to claim 1, wherein The step of using the target detection model to perform target detection on the network configuration image to obtain a target detection result corresponding to the network configuration image specifically includes: Using the backbone network of the target detection model to perform feature extraction on the network configuration image to obtain initial feature images of different scales; Using the neck network of the target detection model to fuse the initial feature images of different scales to obtain fused images of different scales; Using the detection head network of the target detection model to perform target detection on the fused image to obtain a target detection result corresponding to the network configuration image.
4. The method according to claim 3, wherein The step of using the backbone network of the target detection model to perform feature extraction on the network configuration image to obtain initial feature images of different scales specifically includes: Performing a slicing operation on the network configuration image to obtain a first feature image; Based on the first feature image, using multiple convolutional neural networks to perform feature extraction to obtain multiple second feature images of different scales; Based on the second feature image output by the convolutional neural network adjacent to the small target detection layer, using the channel attention module of the small target detection layer to perform feature extraction to obtain a channel feature image; Based on the second feature image output by the convolutional neural network adjacent to the small target detection layer, using the spatial attention module of the small target detection layer to perform feature extraction to obtain a spatial feature image; Stitch the channel feature image and the spatial feature image to obtain a third feature image; The initial feature image includes the first feature images of different scales, each of the second feature images, and the third feature image.
5. The method according to claim 4, wherein The feature extraction based on the second feature image using the channel attention module of the small object detection layer to obtain a channel feature image specifically includes: Perform global average pooling on the second feature image to obtain a first feature vector; Perform global max pooling on the second feature image to obtain a second feature vector; Perform one-dimensional convolution calculation based on the first feature vector, the second feature vector, and a preset activation function to obtain a channel attention weight vector; Perform multiplication operation based on the channel attention weight vector and the second feature image to obtain the channel feature image.
6. The method according to claim 4, characterized in that, The feature extraction based on the second feature image using the spatial attention module of the small object detection layer to obtain a spatial feature image specifically includes: Separate the channels of the second feature image to obtain each channel-separated feature image; Perform global average pooling on the same channel-separated feature image to obtain a first separated feature vector; Perform global max pooling on the same channel-separated feature image to obtain a second separated feature vector; Perform shared convolution calculation based on the first separated feature vector and the second separated feature vector to obtain the spatial attention weight corresponding to the same channel-separated feature image; Perform calculation based on the channel-separated feature image and the spatial attention weight respectively to obtain an initial spatial feature image corresponding to each channel-separated feature image; Perform stitching on each of the initial spatial feature images to obtain the spatial feature image.
7. The method according to claim 3, wherein The target detection of the fusion image using the detection head network of the target detection model to obtain a target detection result corresponding to the power distribution network image specifically includes: Identify the fusion image to obtain a region of interest; Use the bilinear interpolation algorithm to perform interpolation calculation on the region of interest to obtain an initial target region; Perform average pooling on each sub-region in the initial target region to obtain an initial fusion feature vector corresponding to each sub-region; Perform feature fusion based on each of the initial fusion feature vectors to obtain a fusion feature vector; Perform detection based on the fusion feature vector using the detection head network of the target detection model to obtain a target detection result corresponding to the power distribution network image; Among them, the target detection result includes the position coordinates of the target prediction box, the target category probability, and the prediction confidence.
8. An object detection device based on an improved YOLOv5 model, characterized in that, It includes: An adding module for adding a small object detection layer with a CMH attention mechanism to the backbone network of the basic YOLOv5 model to obtain an improved YOLOv5 model; A training module for training the improved YOLOv5 model based on a pre-acquired sample data set using the region of interest alignment pooling method and the minimum point distance intersection over union loss function to obtain a target detection model; An obtaining module for obtaining a power distribution network image to be target detected; The detection module is used to perform object detection on the distribution network image by using the target detection model, and obtain the object detection result corresponding to the distribution network image.
9. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the object detection method based on the improved YOLOv5 model described in any one of claims 1-7 above are implemented.
10. An electronic device, characterized in that, It includes at least a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, the steps of the object detection method based on the improved YOLOv5 model described in any one of claims 1-7 above are implemented.
Citation Information
Cited By
Target fine granularity identification method and device
CN121392258A