A target detection method based on a lightweight neural network

The lightweight neural network MobileOne, with enhanced feature extraction modules, addresses the computational demands of YOLOv7, improving target detection speed and accuracy on mobile devices.

CN116452900BActive Publication Date: 2025-07-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310448848.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-15
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

The existing YOLOv7 object detection method is difficult to meet the computing power requirements on mobile devices or embedded devices, resulting in insufficient accuracy and real-time performance of detection tasks.

Method used

The improved lightweight neural network D-MobileOne is adopted, combining deep separable convolution and multiple feature enhancement modules to replace YOLOv7's backbone network, and improve model detection speed by reducing the amount of parameters and calculations.

Benefits of technology

Significantly improve the target detection speed on low-computing equipment, enhance the model's perception of targets of different sizes, and improve the detection accuracy and robustness in extreme scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452900B_ABST
    Figure CN116452900B_ABST
Patent Text Reader

Abstract

The present invention discloses an object detection method based on a lightweight neural network, comprising the following steps; Step 1: Obtain an object detection data set to obtain an input image; Step 2: Build an object detection network based on a lightweight neural network, and the structure sequence of the object detection network is, in sequence, an input layer, a feature extraction layer, a feature enhancement layer, and an output layer; Step 3: Train the object detection network to obtain a trained network; Step 4: Input the image sample to be detected into the network trained in Step 3 for object detection and output the detection result. The present invention improves the detection speed of the model by significantly reducing the number of parameters and the amount of computation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection in computer vision, and particularly relates to an object detection method based on a lightweight neural network. Background Art

[0002] Object detection is one of the important tasks in the field of computer vision, and its research status can be summarized from two aspects: traditional methods and deep learning methods. Traditional methods mainly rely on hand-designed features and strategies of sliding windows or candidate regions. Among them, the Viola Jones detector, HOG detector, and DPM model are relatively typical representatives. These methods have good detection effects for simple scenes, but lack robustness and generalization ability for complex and variable scenes. Deep learning methods mainly rely on deep convolutional neural networks and end-to-end training frameworks. Among them, the R-CNN series, YOLO series, and SSD series are relatively typical representatives. These methods can automatically learn image features and achieve fast and accurate object detection.

[0003] Currently, deep learning methods have become the mainstream methods in the field of object detection, and YOLOv7 has achieved good results in object detection tasks. However, when deployed on mobile devices or embedded devices, the computing power requirements of YOLOv7 are difficult to meet, and thus the accuracy and real-time performance of object detection tasks cannot be guaranteed. Therefore, it is necessary to perform lightweight improvement on YOLOv7. Summary of the Invention

[0004] In order to overcome the above problems existing in the prior art, the purpose of the present invention is to provide an object detection method based on a lightweight neural network, which uses depthwise separable convolution to improve the MobileOne network and replaces the original backbone network in YOLOv7, and improves the detection speed of the model by significantly reducing the number of parameters and the amount of computation.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0006] An object detection method based on a lightweight neural network, comprising the following steps;

[0007] Step 1: Obtain an object detection data set, and the data in the data set is an input image;

[0008] Step 2: Build an object detection network based on a lightweight neural network, and the structure order of the object detection network is, in sequence, an input layer, a feature extraction layer, a feature enhancement layer, and an output layer;

[0009] 2.1) The input layer is responsible for preprocessing the input image of the object detection data set, and preprocessing and aligning the picture into a 640*640 RGB image;

[0010] 2.2) Feature extraction layer, using the improved lightweight neural network D-MobileOne as the feature extraction network, which is used to extract features from the RGB image as the input image. The feature extraction layer is responsible for performing local feature extraction and fusion on various types of information extracted, and outputting feature maps at three different levels;

[0011] 2.3) Feature enhancement layer, which enhances the feature maps at three different levels output by the feature extraction layer, including the CBAAM module, SPPCSPC module, HDCBS module, UPSample module, ELAN-Q module, and MPConv module;

[0012] 2.4) Output network part, which passes the feature maps of the three enhanced sizes in the feature enhancement layer through the REP module and convolutional layer, and finally outputs the result;

[0013] Step 3: Train the object detection network to obtain a trained network;

[0014] Step 4: Input the image samples to be detected into the network trained in Step 3 for object detection, and output the detection results.

[0015] In the above Step 1, the object detection dataset is divided according to the sign category, size, and weather conditions; classified by sign category, traffic sign types are divided into warning categories, prohibited categories, and mandatory categories; classified by weather conditions, they are divided into sunny days, nights, cloudy days, rainy days, snowy days, and foggy days.

[0016] In the above 2.2), the improved lightweight neural network D-MobileOne is composed of the D-CB-M module and the CB-N module, which extracts features from the input image and outputs feature maps at three different levels;

[0017] Regarding the combination of the D-CB-M module and the CB-N module as a unit of D-MobileOne, then there are five such units in the D-MobileOne network, and the connection method is cascaded; divided according to each D-MobileOne unit, the feature extraction layer is divided into five parts; the feature extraction layer outputs feature maps of three different sizes to the feature enhancement layer through the CBAAM module;

[0018] Among them, the third unit outputs a feature map with a size of 80×80×512, the fourth unit outputs a feature map with a size of 40×40×1024, and the fifth unit outputs a feature map with a size of 20×20×1024.

[0019] Specifically:

[0020] 2.2.1) The D-CB-M module contains three branches. The first branch contains k depthwise separable convolution modules (k is a hyperparameter) composed of 3×3 depthwise convolution and 1×1 pointwise convolution. The second branch is a CB module containing 1×1 convolution. The third branch is a BN layer. The reparameterized branch of D-CB-M contains 3×3 convolution, 1×1 convolution, BN layer, and activation function.

[0021] 2.2.2) The CB-N module contains two branches. The first branch of the CB-N module is k CB modules composed of 1×1 convolution and BN layer. The second branch is a BN layer. The reparameterized branch of CB-N contains 1×1 convolution, BN layer, and activation function.

[0022] The 2.3) is specifically as follows;

[0023] 2.3.1) The CBAAM module consists of two parts: channel attention module and spatial attention module. The channel attention module obtains channel attention weights by performing adaptive average pooling and adaptive max pooling on the input feature map, then adding them and passing through a fully connected layer, and then multiplying element-wise with the input feature map. The spatial attention module obtains spatial attention weights by performing adaptive average pooling and adaptive max pooling on the feature map with channel attention applied, then concatenating and passing through a convolutional layer, and then multiplying element-wise with the feature map with channel attention applied to finally obtain the output feature map.

[0024] 2.3.2) The SPPSCPC module contains a Spatial Pyramid Pooling (SPP) module and a CSP module. Among them, the SPP obtains different receptive fields by the method of max pooling and enables the algorithm to adapt to images of different resolutions. Among the four branches of the SPP module, it contains three max pooling branches and one skip connection branch, and the output feature map has the same size as the input feature map. The CSP module divides the features into two parts, one part is processed in the conventional way, and the other part is processed with the SSP structure.

[0025] 2.3.3) The HDCBS module contains a hybrid dilated convolution layer, a BN layer, and an activation function layer. The hybrid dilated convolution HDC layer is used to replace the convolutional layer in the original feature enhancement layer.

[0026] 2.3.4) The UPSample module uses the nearest neighbor interpolation method for upsampling. This module consists of an upsampling layer and a convolutional layer. The upsampling layer increases the size of the feature map, and the convolutional layer extracts high-level feature information from the upsampled feature map.

[0027] 2.3.5) The structure of the ELAN-Q module is similar to that of ELAN. The first branch changes the number of channels through a 1×1 convolution, and the second branch continues to process through four 3×3 convolutions after the number of channels is changed; the ELAN-Q selects five outputs to be added in the Concat layer;

[0028] 2.3.6) The MPConv module contains two branches to achieve the downsampling function; the first branch first passes through a max pooling layer (MaxPool) to achieve downsampling, and then passes through a 1×1 convolution to change the number of channels, so as to extract the low-level features of the input data for learning; the second branch first changes the number of channels through a 1×1 convolution, and then passes through a 3×3 convolution kernel with a stride of 2 for downsampling to extract the high-level features of the input data; the downsampling information obtained by the two branches is fused in the Concat layer to obtain the super downsampling result.

[0029] The 2.4) is specifically as follows:

[0030] 2.4.1) The REP module has two modes: training mode and deploy mode;

[0031] In the training mode, this module contains three branches: a 3×3 convolution branch for feature extraction, a 1×1 convolution branch for smoothing features, and an Identity branch without convolution operation;

[0032] The outputs of the three branches are added through the Concat module to obtain the final result. In the inference mode, this module only contains a 3×3 convolution layer with a stride of 1, which is formed by reparameterizing the 1×1 convolution kernel and the Identity branch of the training module, converting it into a 3×3 convolution layer, and adding its weights to form a branch that only contains a 3×3 convolution layer and a BN layer.

[0033] The specific steps of step 3 are as follows:

[0034] 3.1) Set up the environment required for the object detection network;

[0035] 3.2) Download the YOLOv7.pt pre-trained model;

[0036] 3.3) Load the pre-trained model in step 3.2), and input the data obtained in step 1 after processing into an object detection network based on a lightweight neural network built in step 2, and perform training.

[0037] An electronic device includes a processor, a memory, and a communication bus. Among them, the processor and the memory complete mutual communication through the communication bus;

[0038] The memory is used to store computer programs;

[0039] A processor, when executing a program stored in a memory, implements the above-mentioned object detection method based on a lightweight neural network.

[0040] A computer-readable storage medium stores a computer program therein, and when the computer program is executed by a processor, the above-mentioned object detection method based on a lightweight neural network is implemented.

[0041] Advantages of the present invention:

[0042] First: The present invention adds a hybrid dilated convolution module to the YOLOv7 network to obtain more context semantic information, thereby enhancing the model's perception ability for targets of different sizes and improving the model's small target detection ability.

[0043] Second: The present invention uses an improved convolutional block attention module (CBAAM) to enhance the feature extraction ability from multiple channel dimensions, enhance the model's detection ability in extreme scenarios such as night, rain, and fog, improve the object detection accuracy, and increase the model's robustness.

[0044] Third: The present invention improves the lightweight network model MobileOne and designs a lightweight object detection network based on D-MobileOne. During training, the branch structure before re-parameterization is used, and during inference, the re-parameterized structure is used. The lightweight object detection network based on D-MobileOne significantly improves the object detection speed on low-computing-power mobile devices. Description of the Drawings

[0045] Figure 1 Is a flowchart for the implementation of the present invention.

[0046] Figure 2 Is a flowchart for the model construction of the present invention.

[0047] Figure 3 Is a flowchart for training the network of the present invention. Detailed Embodiments

[0048] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0049] First, the noun terms involved in one or more embodiments of the present invention are explained.

[0050] Object detection: A computer vision technology used to find and locate objects of interest in images or videos. Object detection has many application scenarios, such as face detection, pedestrian detection, text detection, etc. Object detection algorithms can be divided into traditional methods and methods based on machine learning and deep learning.

[0051] Feature Map: In computer vision, a feature map refers to a set of data representing the content of an image obtained after a certain transformation of the image. Feature maps can reflect structural information such as points, edges, lines, and curves in the image, as well as semantic information such as color, texture, and shape of the image. Feature maps can be used for tasks such as object detection, classification, and segmentation, and can also be used to visualize and analyze the learning effect of the model.

[0052] As Figures 1 - 3 shown: It includes the following steps;

[0053] S1: Obtain an object detection dataset; taking the Chinese Traffic Sign Dataset CCTSDB 2021 as an example, the CCTSDB 2021 dataset contains 17,856 images, including 16,356 training images, 1,500 positive test images, and 500 negative test images. The test dataset is divided according to sign categories, sizes, and weather conditions; classified by category, traffic sign types are divided into warning, prohibition, and mandatory categories; divided by weather conditions, they are divided into sunny, night, cloudy, rainy, snowy, and foggy days.

[0054] S2: Build an object detection network based on a lightweight neural network, referring to Figure 2 , the flowchart of the model construction of the present invention. This network consists of four parts. According to the connection order of each part, they are:

[0055] S201: Input part, resize the image size to 640×640 and send the resized image into the feature extraction network;

[0056] S202: Feature extraction network part. In this method, the improved lightweight neural network D-MobileOne is used as the feature extraction network, and its function is to extract features from the input image, mainly composed of the D-CB-M module and the CB-N module;

[0057] S202a: The D-CB-M module contains three branches. The first branch contains k depthwise separable convolution modules composed of 3×3 depthwise convolution and 1×1 pointwise convolution (k is a hyperparameter), the second branch is a CB module containing 1×1 convolution, and the third branch is a BN layer. The reparameterized branch of D-CB-M contains 3×3 convolution, 1×1 convolution, BN layer, and activation function.

[0058] S202b: The first branch of the CB-N module is k CB modules composed of 1×1 convolution and BN layer, and the second branch is a BN layer. The reparameterized branch of CB-N contains 1×1 convolution, BN layer, and activation function.

[0059] S203: Feature Enhancement Network, which is used to enhance the three-sized feature maps output by the Feature Extraction Network, including CBAAM module, SPPCSPC module, HDCBS module, UPSample module, ELAN-Q module, and MPConv module;

[0060] S203a: The CBAAM module is an improved convolutional attention module of this method, which consists of an adaptive channel attention module and an adaptive spatial attention module. Among them, the adaptive channel attention module first performs adaptive average pooling or adaptive max pooling on the input feature map to obtain the average or maximum value of each channel in the spatial dimension; then, a small fully-connected neural network (FCN) is used to learn a set of weights, which can be adjusted according to the importance of each channel, and they are used to generate the attention map of each channel; then, the attention map and the input feature map are multiplied element-wise to highlight the channels with large amounts of information and suppress the channels with small amounts of information, improving the accuracy and robustness of the model. The input feature map of the adaptive spatial attention module first undergoes adaptive max pooling and adaptive average pooling to obtain two feature maps of 1×1×C. Then, the two feature maps are respectively fed into two different small convolutional neural networks for processing. The adaptive max pooling feature map passes through a 1×1 convolutional layer and a Sigmoid activation function to obtain an attention branch, and the adaptive average pooling feature map passes through a 1×1 convolutional layer, a Sigmoid activation function, and a 3×3 convolutional layer to obtain another attention branch. The two attention branches are multiplied, and the final attention map is generated through a 1×1 convolutional layer.

[0061] S203b: The SPPSCPC module contains a Spatial Pyramid Pooling (SPP) module and a CSP module. Among them, SPP obtains different receptive fields through max pooling and enables the algorithm to adapt to images of different resolutions. Among the four branches of the SPP module, there are three max pooling branches and one skip connection branch, and the output feature map has the same size as the input feature map. The CSP module divides the features into two parts, one part is processed in the conventional way, and the other part is processed in the SSP structure. By combining the two parts, while halving the computational cost, the speed and accuracy are improved.

[0062] S203c: The HDCBS module is an improved hybrid dilated convolution module of this method. The HDCBS contains a hybrid dilated convolution layer, a normalization layer, and an activation function layer. The hybrid dilated convolution HDC layer is used to replace the convolutional layer in the original feature enhancement network to enhance the accuracy of the model in small target detection by expanding the receptive field.

[0063] S203d: The UPSample module performs upsampling using the nearest neighbor interpolation method. This module consists of an upsampling layer and a convolutional layer. The upsampling layer increases the size of the feature map, and the convolutional layer extracts high-level feature information from the upsampled feature map.

[0064] S203e: The structure of the ELAN-Q module is similar to that of ELAN. The first branch changes the number of channels through a 1×1 convolution, and the second branch continues to be processed through four 3×3 convolutions after the number of channels is changed. Different from ELAN, ELAN-Q selects the sum of five outputs at the Concat layer, so it has a stronger perception ability for objects of different scales than ELAN.

[0065] S203f: The MPConv module contains two branches to achieve the downsampling function. The first branch first passes through a max pooling layer (MaxPool) to achieve downsampling, and then passes through a 1×1 convolution to change the number of channels, so as to extract the low-level features of the input data. The second branch first changes the number of channels through a 1×1 convolution, and then passes through a 3×3 convolution kernel with a stride of 2 for downsampling to extract the high-level features of the input data. The downsampling information obtained by the two branches is fused at the Concat layer to obtain a super downsampling result, enhancing the network's feature extraction ability and robustness.

[0066] S204: In the output network part, the feature maps of three sizes after feature enhancement pass through the REP module and the convolutional layer, and finally the output result is obtained;

[0067] S204a: The REP module has two modes: training mode and deploy mode. In the training mode, this module contains three branches: a 3×3 convolution branch for feature extraction, a 1×1 convolution branch for smoothing features, and an Identity branch that does not perform convolution operations. The outputs of the three branches are added through the Concat module to obtain the final result. In the inference mode, this module only contains a 3×3 convolutional layer with a stride of 1. This layer is formed by reparameterizing the 1×1 convolution kernel and the Identity branch of the training module, converting it into a 3×3 convolutional layer, and adding its weights to form a branch that only contains a 3×3 convolutional layer and a normalization layer.

[0068] S3: Train the object detection network;

[0069] S301: Set the environment required for the object detection network;

[0070] S302: Download the YOLOv7.pt pre-trained model;

[0071] S303: Load the pre-trained model in step S302, process the data obtained in step S1, input it into an object detection network based on a lightweight neural network built in step S2, and perform training;

[0072] S4: Input the image sample to be detected into the trained network for object detection and output the detection result.

[0073] Furthermore, if the combination of the D-CB-M module in step S202a and the CB-N module in step S202b is regarded as a unit of D-MobileOne, then there are five such units in the D-MobileOne network, and the connection method is cascading. Divide the feature extraction network into five parts according to each D-MobileOne unit. The feature extraction network outputs three feature maps of different sizes to the feature enhancement network through the CBAAM module. Among them, the third unit outputs a feature map with a size of 80×80×512, the fourth unit outputs a feature map with a size of 40×40×1024, and the fifth unit outputs a feature map with a size of 20×20×1024.

[0074] Furthermore, the combination method of the CBAAM module used in step S203a and the HDCBS module in step S203c is as follows: The feature extraction network extracts features from the input image and outputs three feature maps of different levels. The feature enhancement network containing the HDCBS module continues to output three feature maps of different sizes according to the input. These feature maps pass through the REP module and the convolutional layer and are used to detect three different tasks in the image: classification, foreground / background classification, and bounding box. The final result will be returned as the output, representing information such as the category, foreground / background, bounding box position and size of each object in the image.

[0075] The combination of the D-CB-M module in step S202a and the CB-N module in step S202b forms a unit of D-MobileOne. There are five such units in the D-MobileOne network, and the connection method is cascading. Divide the feature extraction network into five parts according to each D-MobileOne unit. The feature extraction network outputs three feature maps of different sizes to the feature enhancement network through the CBAAM module. Among them, the third unit outputs a feature map with a size of 80×80×512, the fourth unit outputs a feature map with a size of 40×40×1024, and the fifth unit outputs a feature map with a size of 20×20×1024.

[0076] The structure of the CBAAM module used in step S203a is:

[0077] CBAAM consists of two parts: a channel attention module and a spatial attention module. Channel attention is obtained by performing adaptive average pooling and adaptive max pooling on the input feature map, then adding them together and passing through a fully connected layer to obtain the channel attention weights, which are then multiplied element-wise with the input feature map. Spatial attention is obtained by performing adaptive average pooling and adaptive max pooling on the feature map after applying channel attention, then concatenating them and passing through a convolutional layer to obtain the spatial attention weights, which are then multiplied element-wise with the feature map after applying channel attention. Finally, the output feature map is obtained.

[0078] The HDCBS module used in step S203c has the following structure:

[0079] The hybrid dilated convolutional layer is followed by a normalization layer and finally passes through the SiLU activation function. The expression of the SiLU activation function is:

[0080]

[0081] Perform hybrid dilated convolution operations on the input feature map, and use dilated convolution kernels of different sizes and dilation rates to perform batch normalization operations on the feature map after hybrid dilated convolution, which can accelerate convergence and improve stability. Performing activation function operations on the feature map after batch normalization can increase non-linearity and expressive power.

Claims

1. A target detection method based on a lightweight neural network, characterized in that, Including the following steps; Step 1: Obtain a target detection dataset, and the data in the dataset are input images; Step 2: Build a target detection network based on a lightweight neural network. The structure sequence of the target detection network is, in turn, an input layer, a feature extraction layer, a feature enhancement layer, and an output layer; Step 3: Train the target detection network to obtain a trained network; Step 4: Input the image samples to be detected into the network trained in Step 3 for target detection and output the detection results; The specific content of Step 2 is as follows: 2.1) The input layer is responsible for preprocessing the input images of the target detection dataset and preprocessing and aligning the pictures into RGB images; 2.2) The feature extraction layer uses the improved lightweight neural network D-MobileOne as the feature extraction network. Its function is to extract features from the RGB images serving as input images. The feature extraction layer is responsible for performing local feature extraction and fusion on various types of information extracted and outputting three feature maps of different levels; 2.3) The feature enhancement layer enhances the three feature maps of different levels output by the feature extraction layer, including the CBAAM module, the SPPCSPC module, the HDCBS module, the UPSample module, the ELAN-Q module, and the MPConv module; 2.4) The output network part passes the three feature maps of enhanced features in the feature enhancement layer through the REP module and the convolutional layer and finally outputs the results; In the above 2.2), the improved lightweight neural network D-MobileOne is composed of the D-CB-M module and the CB-N module, extracts features from the input RGB images, and outputs three feature maps of different levels; Regarding the combination of the D-CB-M module and the CB-N module as a unit of D-MobileOne, then there are five such units in the D-MobileOne network, and the connection method is cascaded; divided according to each D-MobileOne unit, the feature extraction layer is divided into five parts; the feature extraction layer outputs three feature maps of different sizes to the feature enhancement layer through the CBAAM module.

2. The object detection method based on a lightweight neural network according to claim 1, characterized in that, In the above Step 1, the target detection dataset is divided according to the sign category, size, and weather conditions; classified according to the sign category, the traffic sign types are divided into warning categories, prohibition categories, and mandatory categories; divided according to the weather conditions, it is divided into sunny, night, cloudy, rainy, snowy, and foggy days.

3. A target detection method based on a lightweight neural network according to claim 1, characterized in that, Specifically: 2.2.1) The D-CB-M module contains three branches. The first branch contains k depthwise separable convolution modules composed of 3×3 depthwise convolution and 1×1 pointwise convolution (k is a hyperparameter), the second branch is a CB module containing a 1×1 convolution, and the third branch is a BN layer; the reparameterization branch of D-CB-M contains a 3×3 convolution, a 1×1 convolution, a BN layer, and an activation function; 2.2.2) The CB-N module contains two branches. The first branch of the CB-N module is k CB modules composed of 1×1 convolution and BN layer, and the second branch is a BN layer; the reparameterization branch of CB-N contains a 1×1 convolution, a BN layer, and an activation function.

4. A target detection method based on a lightweight neural network according to claim 1, characterized in that, The 2.3) is specifically as follows: 2.3.1) The CBAAM module consists of two parts: a channel attention module and a spatial attention module; the channel attention module performs adaptive average pooling and adaptive max pooling on the input feature map, then adds them and obtains the channel attention weights through a fully connected layer, and then multiplies them element-wise with the input feature map; the spatial attention module performs adaptive average pooling and adaptive max pooling on the feature map with channel attention applied, then concatenates them and obtains the spatial attention weights through a convolutional layer, and then multiplies them element-wise with the feature map with channel attention applied to finally obtain the output feature map; 2.3.2) The SPPSCPC module includes a Spatial Pyramid Pooling (SPP) module and a CSP module; among them, the SPP obtains different receptive fields through max pooling and enables the algorithm to adapt to images of different resolutions. Among the four branches of the SPP module, there are three max pooling branches and one skip connection branch, and the output feature map has the same size as the input feature map. The CSP module divides the features into two parts, one part is processed in the conventional way, and the other part is processed with the SSP structure; 2.3.3) The HDCBS module includes a hybrid dilated convolutional layer, a BN layer, and an activation function layer; the hybrid dilated convolutional HDC layer is used to replace the convolutional layer in the original feature enhancement layer; 2.3.4) The UPSample module uses the nearest neighbor interpolation method for upsampling; this module consists of an upsampling layer and a convolutional layer. The upsampling layer increases the size of the feature map, and the convolutional layer extracts high-level feature information from the upsampled feature map; 2.3.5) The structure of the ELAN-Q module is similar to that of ELAN. The first branch changes the number of channels through a 1×1 convolution, and the second branch continues to be processed through four 3×3 convolutions after the number of channels changes; the ELAN-Q selects the addition of five outputs at the Concat layer; 2.3.6) The MPConv module includes two branches to achieve the downsampling function; the first branch first passes through a max pooling layer (MaxPool) to achieve downsampling, and then passes through a 1×1 convolution to change the number of channels, so as to extract the low-level features of the input data; The second branch first changes the number of channels through a 1×1 convolution, and then passes through a 3×3 convolutional kernel with a stride of 2 for downsampling to extract the high-level features of the input data; The downsampling information obtained from the two branches is fused at the Concat layer to obtain the super downsampling result.

5. A target detection method based on a lightweight neural network according to claim 1, characterized in that, The 2.4) is specifically as follows: 2.4.1) The REP module has two modes: the training mode and the deploy mode; In the training mode, this module includes three branches: a 3×3 convolution branch for feature extraction, a 1×1 convolution branch for smoothing features, and an Identity branch that does not perform convolution operations; The outputs of the three branches are added together through a Concat module to obtain the final result. In inference mode, this module only contains a 3×3 convolutional layer with a stride of 1, which is formed by reparameterizing the 1×1 convolutional kernel of the training module and the Identity branch, converting it into a 3×3 convolutional layer, and adding its weights to form a branch that only contains a 3×3 convolutional layer and a BN layer.

6. The object detection method based on a lightweight neural network according to claim 1, characterized in that The specific steps of step 3 are as follows: 3.1) Set up the environment required for the object detection network; 3.2) Download the YOLOv7.pt pre-trained model; 3.3) Load the pre-trained model in step 3.2), and input the data obtained in step 1 after processing into the object detection network based on a lightweight neural network built in step 2, and perform training.

7. An electronic device, characterized in that, It includes a processor, a memory, and a communication bus. Among them, the processor and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the object detection method based on a lightweight neural network described in any one of claims 1-6 when executing the program stored on the memory.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the object detection method based on a lightweight neural network described in any one of claims 1-6.