A lightweight model based on YOLOv7
The lightweight YOLOv7 model addresses the precision-speed tradeoff by integrating a high-efficiency mobile neural backbone and inverse convolutions, improving detection speed and accuracy for real-time applications.
Patent Information
- Application Number
- CN202310497613.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-05-06
AI Technical Summary
The existing YOLO network has low detection accuracy and slow speed on embedded devices, making it difficult to achieve real-time object detection in complex environments.
The efficient mobile neural backbone network is used to replace the backbone network of YOLOv7, and the inverse feature convolution neural network operator is introduced to replace traditional convolution, and the model is reparameterized and lightweighted.
Without reducing detection accuracy, the detection speed and efficiency of the model are improved, and are suitable for real-time object detection of embedded devices.
Smart Images

Figure CN116580184B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object detection, and specifically to a lightweight model based on YOLOv7. Background Technique
[0002] Object detection is to locate and identify the objects of interest in an image. It is an important research direction in computer vision, and also the premise and foundation of many computer vision tasks, having important application value in fields such as autonomous driving and video surveillance. With the development of computer vision, a large number of studies have been carried out on object detection technologies based on computer vision, and more and more image processing and recognition technologies have emerged. Especially in recent years, with the popularization of the application of artificial intelligence technologies represented by deep learning, it has provided important new ideas for object detection.
[0003] Object detection technology based on deep learning no longer requires manual extraction of object features. Only by building a suitable network model and training through a dataset can it automatically find suitable object features. However, object detection technology based on deep learning also faces some problems. As the network deepens, the model becomes more and more complex, the required computational amount increases continuously, and it is difficult for the algorithm model to achieve a balance between detection accuracy and detection speed. At present, the YOLO network has advantages such as fast detection speed and strong real-time performance, and is widely used in the field of real-time object detection. However, the existing YOLO algorithms still cannot meet the application scenarios mainly based on embedded devices in terms of accuracy and speed, and are prone to problems such as missed detection and false detection in complex environments. In view of the above problems, the present invention improves the YOLOv7 network to make the model lightweight while ensuring detection accuracy. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] In view of the deficiencies of the prior art, the present invention provides a lightweight model based on YOLOv7, which solves the problems of low detection accuracy and slow detection speed existing in the traditional YOLO network.
[0006] (2) Technical Solutions
[0007] The present invention specifically adopts the following technical solutions to achieve the above object:
[0008] A lightweight model based on YOLOv7, the method includes the following steps:
[0009] Step 1, dataset preparation, dividing the target dataset into two parts, a training set and a validation set, and all images contain the position information of the target boxes and key points manually marked.
[0010] Step 2: Construct the YOLOv7 network structure. Replace the backbone network of YOLOv7 with an efficient mobile neural backbone network, and at the same time introduce an inverse convolutional neural network operator to replace the traditional convolution, obtaining the improved YOLOv7 network;
[0011] Step 3: Input the training set divided in Step 1 into the improved YOLOv7 network for training to obtain a lightweight model.
[0012] Step 4: Use the validation set images divided in Step 1 to input into the lightweight model obtained in Step 3 to obtain the finally predicted object detection bounding boxes and coordinates.
[0013] Furthermore, in Step 2, the YOLOv7 network is used as the basic framework for object detection. YOLOv7 mainly consists of an input end, a backbone network, and a prediction network. The backbone network is a convolutional neural network that forms image features, and the prediction network predicts the features of the image and generates bounding boxes and predicted classes. Each stage contains different extracted features.
[0014] Furthermore, in Step 2, an efficient neural backbone network for mobile devices is introduced to replace the feature extraction network in the YOLOv7 backbone network.
[0015] Furthermore, in Step 2, an inverse feature convolutional neural network operator is introduced to replace the traditional convolution in the backbone network and the prediction network.
[0016] Furthermore, in Step 2, specifically, the original image is subjected to feature extraction and feature fusion through the feature extraction network, and shallow feature maps, middle feature maps, and deep feature maps are respectively output. After passing through the inference convolutional layer, three types of tasks for image detection are predicted, and finally the prediction results are output.
[0017] (III) Beneficial effects
[0018] Compared with the prior art, the present invention provides a lightweight model based on YOLOv7, having the following
[0019] Beneficial effects:
[0020] In the present invention, aiming at the high real-time requirement of the network for video object detection, an efficient mobile neural backbone network is introduced to replace the backbone network of the YOLOv7 network, and the network is lightweighted through model re-parameterization, improving the detection speed of the network.
[0021] The present invention introduces an inverse feature convolutional neural network operator to replace the traditional convolution, overcomes the limitations of the traditional convolution, is lighter and more efficient than the traditional convolution, and can achieve double improvements in accuracy and efficiency on the model.
[0022] Based on the YOLOv7 network, the present invention introduces an efficient mobile neural backbone network and an inverse feature convolutional neural network operator, which can improve the detection efficiency of the model without reducing the accuracy of object detection, optimize the model, and make it have a broader application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flowchart of the present invention;
[0024] Figure 2 It is a basic module diagram of the efficient mobile neural backbone network introduced by the present invention;
[0025] Figure 3 It is a schematic diagram of the inverse feature convolutional neural network operator introduced by the present invention;
[0026] Figure 4 It is the network structure diagram of the YOLOv7 network used by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] Embodiment
[0029] As Figures 1-4 shown, a lightweight model based on YOLOv7 proposed in an embodiment of the present invention includes the following steps.
[0030] Step 1, data set preparation. The data set for object detection is divided into two parts: a training set and a validation set, and all images contain manually annotated object boxes and the position information of each key point; n object detection boxes are annotated in each image, and each detection box corresponds to a coordinate position, which is the coordinate of the center position of the detection box.
[0031] Step 2, construct the YOLOv7 network structure. YOLOv7 mainly consists of an input end, a backbone network, and a prediction network. The backbone network is a convolutional neural network that forms image features, and the prediction network predicts the features of the image and generates bounding boxes and predicted categories. Each stage contains different extracted features, and the network structure is as Figure 4 shown;
[0032] First, preprocess the input image. Perform slicing operations on the image, obtaining a value every other pixel in an image to get four images. Combine parts of the four images to form an input picture of a certain size. Then, perform a convolution operation on the newly obtained spliced image to get a two - fold down - sampled feature map without information loss, and input it into the backbone network. Generally, the input image is 640*640*3; the anchor boxes are set according to the detection layers. Each layer of anchor boxes is applied to different feature maps. In object detection tasks, generally, it is hoped to detect small objects on large feature maps because large feature maps contain more information about small objects. Therefore, the values of anchor boxes on large feature maps are usually set to small values, while the values on small feature maps are set to large values for detecting large objects; there are three detection layers in the network, so the anchor boxes are set in three rows, corresponding to the shallow, middle, and deep layers respectively. The image input into the backbone network passes through four 3*3 convolutional layers, and then through an efficient feature extraction network to increase the number of channels and extract features. Through three pooling layers and the feature extraction network, downsampling and feature extraction are performed, respectively outputting three feature maps of different sizes C3(80*80*512), C4(40*40*1024), C5(20*20*1024); C5 passes through a pooling layer and a feature processing network to obtain the feature map P5(20*20*512). Different receptive fields are obtained through max - pooling to adapt to images of different resolutions. Different pooling layers correspond to different receptive fields, thus distinguishing small objects and large objects; the feature processing network is divided into two branches. One branch performs conventional processing on the features, and the other branch processes the features of the pooling layer. Finally, the two parts are fused to output the result. C5 is fused with C4 and C3 in the order from top to bottom. Through upsampling and the feature extraction network, P3(80*80*256) and P4(40*40*512) are obtained, and then fused with P4 and P5 in the order from bottom to top. Finally, three feature maps of different sizes (20*20*255, 40*40*255, 80*80*255) are output; after passing through the inference convolutional layer, predictions are made for the three types of tasks in image detection (classification, foreground - background classification, bounding box), and finally the prediction results are output.
[0033] In the backbone network part of the present invention, an efficient neural backbone network for mobile devices is introduced to replace the feature extraction network in the YOLOv7 backbone network. This efficient mobile neural backbone network uses re - parameterization to achieve model lightweighting. Model re - parameterization uses a multi - branch complex network during training to enable the model to obtain better feature representation. During testing, the multi - branches are merged into one branch for testing, reducing the computational amount and the number of parameters, thereby improving the speed; the basic modules of this network are as Figure 2As shown, the basic module is built on top of the MobileNet-V1 block with 3*3 depthwise convolutions and 1*1 pointwise convolutions. It uses a normalization layer and a branch with a copy structure to introduce reparameterizable residual connections. There are two different structures during training time and testing time. On the left is the training-time mobile network module with reparameterizable branches, and on the right is the inference module of the reparameterized branches, using ReLU or SE-ReLU as the activation function; the introduction of the efficient mobile neural backbone network improves the speed of the model and achieves state-of-the-art performance in the efficient architecture.
[0034] Introduce the inverse feature convolutional neural network operator to replace the traditional convolutions in the backbone network and the prediction network; a set of inverse feature convolutional neural network operator kernels can be expressed as For pixel Its inverse feature convolutional neural network operator kernel is g = 1, 2, …, G is the grouping of the inverse feature convolutional neural network operator kernels, calculating the number of groups where each group shares the same inverse feature convolutional neural network operator kernel, sharing the kernel within the group, and using the inverse feature convolutional neural network operator kernel to perform multiplication and addition operations on the input to obtain the output feature map of the inverse feature convolutional neural network operator as:
[0035]
[0036] k is the channel number, and the size of the inverse feature convolutional neural network operator kernel depends on the size of the input feature map and is dynamically generated by the kernel generation function φ:
[0037] H i,j = φ(X ψi,j )
[0038] where ψ i,j is the set of input pixels corresponding to H i,j . Define the kernel generation function φ:
[0039] H i,j = φ(X i,j ) = W 1σ (W0X i,j )σ
[0040] and
[0041] are linear transformations, the intermediate channel dimension is controlled by the compression factor r, and σ represents the non-linear activation function for the two linear transformations after batch normalization.
[0042] The schematic diagram of the inverse feature convolutional neural network operator is as shown in Figure 3; For the feature vector at a coordinate point of the input feature map, first use the φ function (usually a certain linear transformation, a combination of 1x1 convolutions to generate a vector of a specific size) to generate a weight vector of a specific size, and then use a transformation H (the most general form is rearrangement) to expand the weight into a kernel, and then perform multiplication and addition with the feature vectors in the neighborhood of this coordinate point on the input feature map to obtain the final output feature map.
[0043] The inverse feature convolutional neural network operator can aggregate context semantic information in a broader space, thus overcoming the difficulty of modeling long-range interactions, and can adaptively allocate weights at different positions, so as to prioritize the visually richest elements of information in the spatial domain, overcome the shortcomings of traditional convolutions, reduce the computational amount and the number of parameters of the network, and make the model more lightweight while maintaining the accuracy.
[0044] Step 3: Input the training set divided in Step 1 into the improved YOLOv7 network for training to obtain a lightweight model.
[0045] Step 4: Input the validation set images divided in Step 1 into the lightweight model obtained in Step 3 to obtain the finally predicted object detection boxes and coordinates, etc.
[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A lightweight model based on YOLOv7, characterized in that: It includes the following steps: Step 1, dataset preparation. The target dataset is divided into a training set and a validation set. All images contain the position information of the target boxes and key points manually annotated. Step 2, construct the YOLOv7 network structure. Introduce an efficient mobile neural backbone network to replace the backbone network of YOLOv7. At the same time, introduce the inverse convolutional neural network operator to replace the traditional convolution to obtain the improved YOLOv7 network. Introduce the inverse feature convolutional neural network operator to replace the traditional convolution in the backbone network and the prediction network; a set of inverse feature convolutional neural network operator kernels can be expressed as For pixel Its inverse feature convolutional neural network operator kernel is g = 1, 2,..., G is the grouping of the inverse feature convolutional neural network operator kernels, calculate the number of groups where each group shares the same inverse feature convolutional neural network operator kernel, the kernels within the group are shared, and use the inverse feature convolutional neural network operator kernel to perform multiplication and addition operations on the input, and the output feature map of the inverse feature convolutional neural network operator is obtained as: k is the channel number. The size of the kernel of the inverse feature convolutional neural network operator depends on the size of the input feature map and is dynamically generated by the kernel generation function φ. H i,j = φ(X ψi,j ) where ψ i,j is the set of input pixels corresponding to H i,j ; define the kernel generation function φ: H i,j = φ(X i,j ) = W 1σ (W0X i,j )σ and is a linear transformation. The intermediate channel dimension is controlled by the compression factor r. σ represents the non-linear activation function for 2 linear transformations after batch normalization. For the feature vector at a coordinate point on the input feature map, first use the φ function, generally a certain linear transformation, a combination of 1x1 convolutions to generate a vector of a specific size; generate a weight vector of a specific size, and then use a transformation H, the most general form is rearrangement; expand the weight into a kernel, and then perform multiply-add with the feature vectors in the neighborhood of this coordinate point on the input feature map to obtain the final output feature map. Step 3, input the training set divided in Step 1 into the improved YOLOv7 network for training to obtain the target detection algorithm model. Step 4, use the validation set images divided in Step 1 to input into the target detection algorithm model obtained in Step 3 to obtain the finally predicted target detection boxes and coordinates.
2. The lightweight model based on YOLOv7 according to claim 1, characterized in that: In Step 2, the YOLOv7 network is used as the basic framework for target detection. YOLOv7 mainly consists of an input end, a backbone network, and a prediction network. The backbone network is a convolutional neural network that forms image features. The prediction network predicts the features of the image and generates bounding boxes and prediction categories. Each stage contains different extracted features.
3. The lightweight model based on YOLOv7 according to claim 1, wherein: In Step 2, an efficient neural backbone network for mobile devices is introduced to replace the feature extraction network in the YOLOv7 backbone network.
4. A lightweight model based on YOLOv7 according to claim 1, wherein: In Step 2, the inverse feature convolutional neural network operator is introduced to replace the traditional convolution in the backbone network and the prediction network.
5. A lightweight model based on YOLOv7 according to claim 1, characterized in that: Specifically in Step 2, the original image is subjected to feature extraction and feature fusion through the feature extraction network, and shallow feature maps, middle feature maps, and deep feature maps are respectively output. After passing through the inference convolutional layer, three types of tasks for image detection are predicted, and the final prediction results are output.
Citation Information
Patent Citations
Multistage target detection method and model based on CNN multistage feature fusion
CN108509978A
Signal generation method and device based on inverse scaling convolutional layer
CN113222113A