Lightweight target detection method, and lightweight method and device of target detection model
By introducing lightweight feature extraction module PGELAN and cross-layer sorting pruning technology into the object detection model, the real-time problem of deep learning object detection algorithm on edge devices is solved, and efficient object detection is achieved.
Patent Information
- Application Number
- CN202510034356.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Deep learning-based object detection algorithms have high computing and storage requirements, and are difficult to meet real-time requirements, especially on edge devices with weak computing capabilities.
The lightweight feature extraction module PGELAN and cross-layer sorting pruning technology are used to reduce the amount of model parameters and calculations while maintaining detection accuracy.
It significantly improves the detection efficiency of the object detection model and meets the need for real-time detection on edge devices.
Smart Images

Figure CN119942137A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and specifically relates to a lightweight target detection method, a lightweight method and a device for a target detection model. Background Art
[0002] Object detection is a very important core direction of computer vision. Its main tasks are object positioning and object classification. As one of the basic problems of computer vision, object detection forms the basis of many other visual tasks, such as instance segmentation, image annotation and object tracking. In recent years, the rapid development of deep learning technology has greatly promoted the progress of object detection technology, making it achieve significant breakthroughs and has been widely used in fields such as autonomous driving, video surveillance, medical image analysis and industrial inspection.
[0003] Before the emergence of deep learning, traditional image processing relied on manual feature extraction, and it was necessary to design corresponding feature extractors for different detection tasks based on prior knowledge. In complex scenes, the matching degree between prior features and real targets was low, and the model robustness was poor. With the continuous deepening of deep learning research, convolutional neural network model algorithms have gradually been applied to tasks such as image recognition, classification, and detection, and have become the main research algorithm. Convolutional neural network models have stronger generalization and resolution capabilities than manually extracted features, and have become the mainstream technology in the field of target detection.
[0004] However, as the complexity of computer vision tasks increases, convolutional neural networks continue to develop, and the structure of convolutional neural networks becomes more and more complex through recursive and reuse designs, the number of network layers continues to increase, and the number of parameters is also increasing, so the requirements for computing and storage hardware facilities continue to increase. The target detection algorithm based on deep learning has good results, but its computational workload is large and it is difficult to meet real-time requirements, especially on edge devices with weak computing power. It is difficult to achieve real-time application. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a lightweight target detection method, a lightweight method and device for a target detection model. The present invention combines a lightweight feature extraction module and cross-layer sorting pruning, which can significantly reduce the number of model parameters and calculation amount while maintaining detection accuracy, thereby improving the detection frame rate and meeting the needs of real-time detection of target detection models on edge devices.
[0006] In a first aspect of the present invention, the present invention provides a lightweight target detection method, the method comprising:
[0007] Get the target image;
[0008] The target image is input into the backbone network of the target detection model to obtain the target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected with a second convolutional layer;
[0009] Inputting the target image into the neck network of the target detection model to obtain target fusion information of the target image;
[0010] The target image is input into the head network of the target detection model to obtain the target object information of the target image.
[0011] In a second aspect of the present invention, the present invention further provides a lightweight method for a target detection model, the method comprising:
[0012] Obtain a target detection model, the target detection model includes a backbone network, a neck network and a head network, the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected to multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected to a second convolutional layer; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters;
[0013] According to the values of all filters in the current module and the average filter weight value of the convolutional layer, the first importance of each filter in the current module is calculated;
[0014] Based on the global pruning rate of the current module, remove the first less important filter through soft pruning;
[0015] Determine the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer;
[0016] According to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer, the number of weight parameters reduced after soft pruning is calculated;
[0017] The pruning rate of each convolutional layer in the current module is calculated based on whether the filter of each convolutional layer in the current module is removed and the number of filters of each convolutional layer;
[0018] According to the difference in feature maps of the current module before and after soft pruning, the second importance of each filter that is soft pruned in the current module is calculated;
[0019] Based on the pruning rate of each convolutional layer in the current module, the second least important filter is removed by hard pruning.
[0020] In a third aspect of the present invention, the present invention further provides a lightweight target detection device, the device comprising:
[0021] A first acquisition module, used for acquiring a target image;
[0022] A first processing unit is used to input the target image into the backbone network of the target detection model to obtain target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolution layers and multiple lightweight convolution modules PGConv; the first convolution layer is connected to multiple lightweight convolution modules PGConv that decrease layer by layer, and the multiple lightweight convolution modules PGConv are connected to a second convolution layer;
[0023] A second processing unit is used to input the target image into the neck network of the target detection model to obtain target fusion information of the target image;
[0024] The third processing unit is used to input the target image into the head network of the target detection model to obtain the target object information of the target image.
[0025] In a fourth aspect of the present invention, the present invention further provides a lightweight device for a target detection model, the device comprising:
[0026] The second acquisition module is used to acquire a target detection model, wherein the target detection model includes a backbone network, a neck network and a head network, wherein the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected with a second convolutional layer; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters;
[0027] A first calculation unit, used for calculating the first importance of each filter in the current module according to the values of all filters in the current module and the average filter weight value of the convolution layer in which the filter is located;
[0028] A first pruning unit, configured to remove a first filter with less importance through soft pruning based on a global pruning rate of a current module;
[0029] A second calculation unit is used to determine the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer;
[0030] A third calculation unit is used to calculate the number of weight parameters reduced after the soft pruning process according to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer;
[0031] a fourth calculation unit, configured to calculate a pruning rate of each convolutional layer in the current module according to whether the filter of each convolutional layer in the current module is removed and the number of filters of each convolutional layer;
[0032] a fifth calculation unit, configured to calculate the second importance of each filter that is soft pruned in the current module according to a difference in feature maps of the current module before and after soft pruning;
[0033] The second pruning unit is used to remove the second less important filter through hard pruning based on the pruning rate of each convolutional layer in the current module.
[0034] Beneficial effects of the present invention:
[0035] The lightweight target detection method and device of the present invention replace the feature extraction module in the standard target detection model with the lightweight feature extraction module PGELAN. PGELAN uses lightweight convolution PGConv, which can output convolution features from different receptive fields while reducing the amount of parameters to obtain richer feature information, thereby maintaining a strong feature extraction capability. The lightweight method and device of the target detection model of the present invention prunes the optimized target detection model through a global filter weight cross-layer sorting algorithm and a minimum norm change filter selection algorithm, and significantly prunes redundant features with only a slight decrease in accuracy, thereby reducing the amount of model parameters and computational overhead, and significantly improving the detection efficiency of the target detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flow chart of a lightweight target detection method according to an embodiment of the present invention;
[0037] Figure 2 Schematic diagram of the structure of a target detection model according to an embodiment of the present invention;
[0038] Figure 3 It is a schematic diagram of the structure of the lightweight feature extraction module PGELAN according to an embodiment of the present invention;
[0039] Figure 4 Schematic diagram of the lightweight convolution PGConv structure of an embodiment of the present invention;
[0040] Figure 5 A flowchart of a lightweight method for a target detection model according to an embodiment of the present invention;
[0041] Figure 6 This is a schematic diagram of the structure of a lightweight target detection device according to an embodiment of the present invention;
[0042] Figure 7 Schematic diagram of the lightweight device structure of the target detection model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] The lightweight target detection method and the lightweight target detection model provided in the embodiment of the present application can be implemented based on artificial intelligence (AI) technology. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline, which involves a wide range of fields, including both hardware-level technology and software-level technology. AI basic technology generally includes technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0045] To facilitate understanding of the embodiments of the present application, the specific implementation methods of the lightweight target detection method, the lightweight method of the target detection model, and the device of the present invention are described in detail below, taking the execution of an image target detection task as an example.
[0046] Figure 1is a flow chart of a lightweight target detection method according to an embodiment of the present invention, such as Figure 1 As shown, the method includes:
[0047] 101. Acquire a target image;
[0048] The target image is an image of the target object information that needs to be detected. The target image can cover many different fields and channels. For example, in the field of security monitoring, the target image usually comes from widely distributed surveillance cameras. Each target image may contain important information about the target object, such as the appearance characteristics, behavioral movements, activity trajectories, etc. of the target object. For example, in the field of medical image processing, the target image can be image data obtained by advanced medical imaging equipment such as X-rays, CT, and MRI. These target images present the internal structure and physiological conditions of the human body in an intuitive way. Each pixel may contain important information about the target object, such as internal lesions of the patient's body, organ status, etc.
[0049] 102. Input the target image into the backbone network of the target detection model to obtain target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected with a second convolutional layer;
[0050] In the embodiment of the present invention, Figure 2 As shown, the target detection model can be a standard target detection algorithm network structure of the YOLO series, which includes a backbone network, a neck network and a head network; taking YOLOv7-tiny as an example, the backbone network includes 4 feature extraction modules and 3 downsampling modules, which are used to extract features of different scales and generate multi-scale feature maps. Subsequently, these feature information are fused in the neck network and the head network to obtain the final target detection result. The neck network and the head network include 4 feature extraction modules, 2 upsampling modules and 1 detection module, which can be used to extract deeper feature information, thereby improving the detection effect.
[0051] Among them, the backbone network is responsible for deep feature extraction of the input image, and can capture feature information at various levels in the image, from basic texture and color distribution to more abstract object contours and structures, providing a basis for subsequent target detection; the neck network can further process and fuse the features extracted by the backbone network, comprehensively consider feature information at different levels, and cleverly combine and adjust the features of different stages of the backbone network in the future, to generate a richer, more comprehensive and more semantically expressive feature representation, providing better feature input for the head network to accurately detect targets; the head network can ultimately complete the target detection task based on the feature information provided by the backbone network and the neck network, and is responsible for predicting key information such as the target category, position, size, etc. It will also use some advanced technologies, such as the non-maximum suppression (NMS) algorithm, to remove duplicate detection results to ensure that each target is detected only once.
[0052] Among them, considering the problem of redundant calculation when the feature extraction module of the traditional backbone network processes complex image information, the present invention uses a lightweight feature extraction module PGELAN to replace the feature extraction module in the backbone network, which can capture the key feature information in the image in a shorter time. Compared with the original module, more image data can be processed in the same time, thereby improving the feature extraction efficiency of the entire model and providing more timely feature support for subsequent target detection, classification and other tasks.
[0053] Figure 3 FIG. 4 is a schematic diagram of the structure of the lightweight feature extraction module PGELAN according to an embodiment of the present invention. Figure 3 As shown, the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected to multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected to the second convolutional layer. The input size of this module is W×H×Cin, where the feature map size is W×H and the number of channels is Cin. The module contains 4 convolution operations. First, the number of channels is halved through convolution Conv1, and the output channel number is C_, and the convolution kernel size is 1, which is used to integrate input information, reduce dimensions and extract shallow features. The feature map passes through two lightweight convolutions PGConv in turn. The convolution kernel size of PGConv2 is 3, and the receptive field size is 3; the convolution kernel size of PGConv3 is 3, and the receptive field size is 5, and different receptive fields are used to obtain multi-scale feature information. Afterwards, the outputs of Conv1, PGConv2 and PGConv3 are fused through the Concat operation to generate a fused feature map as the input of Conv4. The convolution kernel size of Conv4 is 1, and the number of output channels is Cout. By linearly combining the features of different channels, the relationship between channels is learned to achieve cross-channel information integration.
[0054] in, Figure 4 PGConv is a schematic diagram of the lightweight convolution structure of an embodiment of the present invention. Figure 4 As shown in the figure, PGConv consists of 3 convolutions, and the input feature map size is W×H×Cin. First, the number of channels is halved through convolution Conv1, and the convolution kernel size is 3, which is used to extract the main channel features and integrate the information; then it goes through two partial convolutions PConv, PConv2 convolution kernel size is 3, receptive field is 5; PConv3 convolution kernel size is 3, receptive field size is 7; the output channels of PConv2 and PConv3 are added and fused with the output of Conv1 through the Concat operation to form the final output feature map. While reducing the amount of parameters, PGConv obtains richer hierarchical feature representation by expanding the receptive field of some channels.
[0055] Based on the above lightweight feature extraction module PGELAN and the lightweight convolution PGConv, the step 102 may include:
[0056] Extracting a first feature map of the target image using a first convolutional layer;
[0057] Extracting multiple second feature maps corresponding to the first feature map of the target image using multiple lightweight convolution modules PGConv that decrease layer by layer;
[0058] In this embodiment, according to a preset ratio, the first feature map of the target image or the second feature map of the target image is split into a first channel feature map and a second channel feature map according to the number of input channels; the channel features of the first channel feature map are extracted using the current lightweight convolution module PGConv; the channel features of the first channel feature map are concat-joined with the second channel feature map to obtain one of the second feature maps of the target image.
[0059] Using concat connection to concatenate the first feature map of the target image and the plurality of second feature maps to obtain a third feature map of the target image;
[0060] A fourth feature map corresponding to the third feature map of the target image is extracted using a second convolutional layer; the fourth feature map is used to indicate target feature information of the target image.
[0061] It can be understood that the sizes of the first channel feature map and the second channel feature map of the target image may be consistent or inconsistent. For example, when the preset ratio is 0.25, the first feature map to be split is a 64-channel feature map, and 16 filters are used to convolve the feature maps of the first 16 channels to obtain the first channel feature map, and the feature maps of the following 48 channels are directly used as the second channel feature map; the first channel feature map and the second channel feature map are concat-connected to obtain one of the second feature maps. If the second feature map to be split is the second feature map, the splitting process is similar, and this embodiment will not be repeated.
[0062] In a preferred embodiment of the present invention, the activation function used in the lightweight feature extraction module PGELAN is the HardSwish function. The performance of HardSwish is similar to that of SiLU, but because it adopts piecewise linear calculation, it avoids the sigmoid operation in SiLU, so the calculation is simpler, faster, and has lower memory usage, which is more suitable for application in edge devices and embedded systems.
[0063] 103. Input the target image into the neck network of the target detection model to obtain target fusion information of the target image;
[0064] In an embodiment of the present invention, the neck network of the target detection model may be a common structure including PANFPN, SPPCSPC, Elan-w, upsample, etc.; this embodiment does not specifically limit this. This embodiment uses the neck network to further fuse the target feature information extracted from the head network of the target detection model to obtain the target fusion information of the target image.
[0065] 104. Input the target image into a head network of a target detection model to obtain target object information of the target image.
[0066] In an embodiment of the present invention, the head network of the target detection model may be a common one including multiple branches, using a 1×1 convolution operation to obtain a corresponding number of categories, which is not specifically limited in the present invention. The embodiment of the present invention utilizes the head network to further classify the target fusion information extracted from the neck network of the target detection model to obtain the target object information of the target image.
[0067] The embodiment of the present invention adopts a lightweight feature extraction module PGELAN in the backbone network of the target detection model, which can output convolution features from different receptive fields while reducing the number of parameters to obtain richer feature information, thereby maintaining a strong feature extraction capability.
[0068] Figure 5is a flow chart of a lightweight method for a target detection model according to an embodiment of the present invention, such as Figure 5 As shown, the method includes:
[0069] 201. Obtain a target detection model, wherein the target detection model includes a backbone network, a neck network and a head network, wherein the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; a first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and a second convolutional layer is connected with multiple lightweight convolutional modules PGConv; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters;
[0070] In an embodiment of the present invention, the target detection model is the same model as the target detection model used in the lightweight target detection method. Since the target detection model has more redundant parameters, this embodiment needs to prune the filters in the target detection model so that the pruned target detection model has the characteristics of lightweight.
[0071] 202. Calculate the first importance of each filter in the current module according to the values of all filters in the current module and the average filter weight value of the convolutional layer in which the filter is located;
[0072] In this embodiment of the present invention, for each filter of each convolutional layer i in the target detection model, The first importance of each filter is determined using the following formula:
[0073]
[0074] in, represents the first importance of filter j of convolutional layer i, Represents the initial value of filter j of convolution layer i, #AvgW i Represents the average filter weight of convolution layer i, so that when calculating the importance, not only the magnitude of the weight is considered, but also the weight ratio of the layer where the weight is located, making the importance calculation more accurate.
[0075] 203. Based on the global pruning rate of the current module, remove the first filter with less importance through soft pruning;
[0076] In this embodiment, the first importance of each filter can be globally sorted, and filters with less first importance can be removed through soft pruning; for these filters, soft pruning with weights reset to 0 is performed. Repeat the above operation until all convolutional layers in the current module have completed soft pruning. Subsequently, hard pruning is performed to remove the least important filter. 204. Determine the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolutional layer;
[0077] In an embodiment of the present invention, when a filter is removed by soft pruning, the size of the corresponding convolution kernel and the number of input channels and the number of output channels of the convolution layer where the filter is located can be determined. This information can be used to calculate the number of weight parameters.
[0078] 205. Calculate the number of weight parameters reduced after soft pruning according to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer.
[0079] In an embodiment of the present invention, the calculation method of the number of weight parameters reduced after the soft pruning process includes:
[0080] The number of first weight parameters is calculated according to the product of the square of the convolution kernel of the removed filter and the number of input channels of the convolution layer;
[0081] The number of second weight parameters is calculated according to the product of the square of the convolution kernel of the removed filter and the number of output channels of the convolution layer;
[0082] The number of weight parameters reduced after the soft pruning process is calculated according to the sum of the first weight parameter number and the second weight parameter number.
[0083] Exemplarily, the amount of parameters that can be reduced after the current filter is trimmed is expressed as:
[0084] Params=Cin i ×K i ×K i +Cout i+1 ×K i+1 ×K i+1
[0085] Among them, Cin i ×K i ×K i Represents the number of first weight parameters, Cout i+1 ×K i+1 ×K i+1 Represents the number of second weight parameters, Cin i Indicates the number of input channels of the convolutional layer i where the removed filter is located, Cout i+1K represents the number of output channels of the convolutional layer i+1 where the removed filter is located. i represents the convolution kernel size of the convolution layer i where the removed filter is located, K i+1 Represents the convolution kernel size of the convolution layer i+1 where the removed filter is located.
[0086] Since the embodiment of the present invention considers the current module as a whole, the filters of different convolutional layers in the same module can be removed across layers, which can significantly cut redundant features, thereby reducing the number of model parameters and computational overhead, and significantly improving the detection efficiency of the target detection model.
[0087] It can be understood that the global pruning rate refers to the ratio of the number of parameters removed from the entire network to the total number of parameters of the original network when the neural network is pruned. Assuming that the total number of parameters of the original target detection model is N, and the number of parameters removed after the soft pruning operation is M, then the global pruning rate P is the ratio of M to N. The embodiment of the present invention can remove filters with smaller weights by repeatedly executing steps 202-205, and continue to remove until the number of parameters reduced by the soft pruned filters meets the parameter number requirement corresponding to the global pruning rate, and then stop.
[0088] The global pruning rate can be preset according to the actual situation. For example, the total number of parameters of the original target detection model is 10M, and the set global pruning rate is 30%. Through the process of steps 202-205, filters with smaller weights are removed, and the number of parameters of the filters that have been soft-pruned is reduced to 3M, and then soft pruning can be stopped. 206. According to whether the filter of each convolution layer in the current module is removed and the number of filters of each convolution layer, the pruning rate of each convolution layer in the current module is calculated;
[0089] In an embodiment of the present invention, the calculation method of the pruning rate of each convolutional layer in the current module includes:
[0090] Calculate a first indication quantity according to whether a filter of each convolutional layer in the current module is removed;
[0091] Calculate a second indication quantity according to the number of filters of each convolutional layer in the current module;
[0092] According to the ratio of the first indicated quantity to the second indicated quantity, the pruning rate of each convolutional layer in the current module is calculated.
[0093] For example, the pruning rate p of each convolutional layer in the current module is i The calculation formula is expressed as:
[0094]
[0095] Among them, <(·) is the indicator function used to judge the filter Whether to be removed in the global sort, n i Indicates the number of filters in convolution layer i, by judging n i The first indication number is obtained by determining whether a filter is removed, and the number of filters of convolution layer i is taken as the second indication number. The pruning rate of each convolution layer in the current module can be measured by the ratio of the first indication number to the second indication number.
[0096] 207. Calculate the second importance of each soft-pruned filter in the current module according to the difference in feature maps of the current module before and after soft pruning;
[0097] In this embodiment of the present invention, the calculation method of the second importance of each soft-pruned filter includes:
[0098] Pass the input image through the current module before soft pruning to obtain the first feature map;
[0099] Pass the input image through the current module after soft pruning to obtain a second feature map;
[0100] According to the similarity distance between the first feature map and the second feature map, the second importance of each filter that is soft pruned is obtained.
[0101] For example, for each module, the feature map Feao generated by the current module can be calculated. For the convolution layer i and filter j from the bottom of the group upward, the filter parameters are set to 0 for soft pruning, and the feature map generated by the current module is recalculated. The change amplitude of the feature map is taken as the second importance of the current filter. The calculation formula of the second importance is expressed as:
[0102]
[0103] in, represents the second importance of filter j in the soft-pruned convolutional layer i, Distance functions such as Manhattan distance, Euclidean distance, Chebyshev distance, and Minkowski distance can be used.
[0104] For example, assuming that the current module is the 20th to 25th convolutional layers, the soft pruned filter is located in the first filter of the 20th convolutional layer of the current module, and the input image passes through the 1st to 19th convolutional layers to obtain the output feature map of the 19th convolutional layer; before the first filter of the 20th convolutional layer is soft pruned, the output feature map of the 19th convolutional layer passes through the 20th to 25th convolutional layers, and the result is Feao; after the first filter of the 20th convolutional layer is soft pruned, the output map of the 19th convolutional layer passes through the 20th to 25th convolutional layers. Since the processing of the first filter of the 20th convolutional layer is missing at this time, a result different from Feao may be obtained. Correspondingly, Feao and The distance can reflect the second importance of the first filter of the 20th convolutional layer that is soft pruned.
[0105] 208. Based on the pruning rate of each convolutional layer in the current module, the second least important filter is removed by hard pruning.
[0106] This embodiment can globally sort the second importance of filters and remove filters with lower second importance through hard pruning.
[0107] The present invention can reduce the computational cost involved in pruning operations by combining soft pruning and hard pruning; it can significantly prune redundant features, thereby reducing the number of model parameters and computational overhead, and significantly improving the detection efficiency of the target detection model.
[0108] Figure 6 FIG. 1 is a schematic diagram of the structure of a lightweight target detection device according to an embodiment of the present invention; Figure 6 As shown, the device comprises:
[0109] A first acquisition module 301 is used to acquire a target image;
[0110] The first processing unit 302 is used to input the target image into the backbone network of the target detection model to obtain the target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolution layers and multiple lightweight convolution modules PGConv; the first convolution layer is connected to multiple lightweight convolution modules PGConv that decrease layer by layer, and the multiple lightweight convolution modules PGConv are connected to a second convolution layer;
[0111] The second processing unit 303 is used to input the target image into the neck network of the target detection model to obtain target fusion information of the target image;
[0112] The third processing unit 304 is used to input the target image into the head network of the target detection model to obtain the target object information of the target image.
[0113] Figure 7 FIG. 1 is a schematic diagram of a lightweight device structure of a target detection model according to an embodiment of the present invention. Figure 7 As shown, the device comprises:
[0114] The second acquisition module 401 is used to acquire a target detection model, wherein the target detection model includes a backbone network, a neck network and a head network, wherein the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected to multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected to a second convolutional layer; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters;
[0115] A first calculation unit 402 is used to calculate the first importance of each filter in the current module according to the values of all filters in the current module and the average filter weight value of the convolution layer in which the filter is located;
[0116] A first pruning unit 403, configured to remove the first less important filter by soft pruning based on the global pruning rate of the current module;
[0117] A second calculation unit 404 is used to determine the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer;
[0118] The third calculation unit 405 is used to calculate the number of weight parameters reduced after the soft pruning process according to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer;
[0119] A fourth calculation unit 406 is used to calculate the pruning rate of each convolution layer in the current module according to whether the filter of each convolution layer in the current module is removed and the number of filters of each convolution layer;
[0120] A fifth calculation unit 407, configured to calculate the second importance of each filter that is soft pruned in the current module according to a difference in feature graphs of the current module before and after soft pruning;
[0121] The second pruning unit 408 is used to remove the second less important filter through hard pruning based on the pruning rate of each convolutional layer in the current module.
[0122] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which can include: ROM, RAM, disk or CD, etc.
[0123] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A lightweight target detection method, characterized in that: The method comprises: Get the target image; The target image is input into the backbone network of the target detection model to obtain the target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected with a second convolutional layer; Inputting the target image into the neck network of the target detection model to obtain target fusion information of the target image; The target image is input into the head network of the target detection model to obtain the target object information of the target image.
2. A lightweight target detection method according to claim 1, characterized in that: The step of inputting the target image into a backbone network of a target detection model to obtain target feature information of the target image comprises: Extracting a first feature map of the target image using a first convolutional layer; Extracting multiple second feature maps corresponding to the first feature map of the target image using multiple lightweight convolution modules PGConv that decrease layer by layer; Using concat connection to concatenate the first feature map of the target image and the plurality of second feature maps to obtain a third feature map of the target image; A fourth feature map corresponding to the third feature map of the target image is extracted using a second convolutional layer; the fourth feature map is used to indicate target feature information of the target image.
3. A lightweight target detection method according to claim 2, characterized in that: The method of extracting multiple second feature maps corresponding to the first feature map of the target image using multiple lightweight convolution modules PGConv that decrease layer by layer includes: According to a preset ratio, splitting the first feature map of the target image or the second feature map of the target image into a first channel feature map and a second channel feature map according to the number of input channels; Use the current lightweight convolution module PGConv to extract the channel features of the first channel feature map; Concat the channel features of the first channel feature map with the second channel feature map to obtain one of the second feature maps of the target image.
4. A lightweight target detection method according to any one of claims 1 to 3, characterized in that: The lightweight convolution module PGConv adopts the HardSwish activation function.
5. A lightweight method for a target detection model, characterized in that: The method comprises: Obtain a target detection model, the target detection model includes a backbone network, a neck network and a head network, the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected to multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected to a second convolutional layer; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters; According to the values of all filters in the current module and the average filter weight value of the convolutional layer, the first importance of each filter in the current module is calculated; Based on the global pruning rate of the current module, remove the first less important filter through soft pruning; Determine the size of the convolution kernel where the removed filter is located and the number of input channels and output channels of the convolution layer where it is located; According to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer, the number of weight parameters reduced after soft pruning is calculated; The pruning rate of each convolutional layer in the current module is calculated based on whether the filter of each convolutional layer in the current module is removed and the number of filters of each convolutional layer; According to the difference in feature maps of the current module before and after soft pruning, the second importance of each filter that is soft pruned in the current module is calculated; Based on the pruning rate of each convolutional layer in the current module, the second least important filter is removed by hard pruning.
6. The lightweight method for target detection model according to claim 5, characterized in that: The calculation method of the number of weight parameters reduced after the soft pruning process includes: The number of first weight parameters is calculated according to the product of the square of the convolution kernel of the removed filter and the number of input channels of the convolution layer; The number of second weight parameters is calculated according to the product of the square of the convolution kernel of the removed filter and the number of output channels of the convolution layer; The number of weight parameters reduced after the soft pruning process is calculated according to the sum of the first weight parameter number and the second weight parameter number.
7. The lightweight method for target detection model according to claim 5, characterized in that: The calculation method of the pruning rate of each convolutional layer in the current module includes: Calculate a first indication quantity according to whether a filter of each convolutional layer in the current module is removed; Calculate a second indication quantity according to the number of filters of each convolutional layer in the current module; According to the ratio of the first indicated quantity to the second indicated quantity, the pruning rate of each convolutional layer in the current module is calculated.
8. The lightweight method for target detection model according to claim 5, characterized in that: The second importance of each soft pruned filter is calculated by: Pass the input image through the current module before soft pruning to obtain the first feature map; Pass the input image through the current module after soft pruning to obtain a second feature map; According to the similarity distance between the first feature map and the second feature map, the second importance of each filter that is soft pruned is obtained.
9. A lightweight target detection device, characterized in that: The device comprises: A first acquisition module, used for acquiring a target image; A first processing unit is used to input the target image into the backbone network of the target detection model to obtain target feature information of the target image; the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolution layers and multiple lightweight convolution modules PGConv; the first convolution layer is connected to multiple lightweight convolution modules PGConv that decrease layer by layer, and the multiple lightweight convolution modules PGConv are connected to a second convolution layer; A second processing unit is used to input the target image into the neck network of the target detection model to obtain target fusion information of the target image; The third processing unit is used to input the target image into the head network of the target detection model to obtain the target object information of the target image.
10. A lightweight device for a target detection model, characterized in that: The device comprises: The second acquisition module is used to acquire a target detection model, wherein the target detection model includes a backbone network, a neck network and a head network, wherein the backbone network adopts a lightweight feature extraction module PGELAN; the lightweight feature extraction module PGELAN includes two convolutional layers and multiple lightweight convolutional modules PGConv; the first convolutional layer is connected with multiple lightweight convolutional modules PGConv that decrease layer by layer, and the multiple lightweight convolutional modules PGConv are connected with a second convolutional layer; wherein the backbone network, the neck network and the head network respectively include multiple modules, each module includes multiple convolutional layers, each convolutional layer includes multiple filters, each filter includes multiple convolutional kernels, and each convolutional kernel includes multiple weight parameters; A first calculation unit, used for calculating the first importance of each filter in the current module according to the values of all filters in the current module and the average filter weight value of the convolution layer in which the filter is located; A first pruning unit, configured to remove a first filter with less importance through soft pruning based on a global pruning rate of a current module; A second calculation unit is used to determine the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer; A third calculation unit is used to calculate the number of weight parameters reduced after the soft pruning process according to the size of the convolution kernel of the removed filter and the number of input channels and output channels of the convolution layer; a fourth calculation unit, configured to calculate a pruning rate of each convolutional layer in the current module according to whether the filter of each convolutional layer in the current module is removed and the number of filters of each convolutional layer; a fifth calculation unit, configured to calculate the second importance of each filter that is soft pruned in the current module according to a difference in feature maps of the current module before and after soft pruning; The second pruning unit is used to remove the second less important filter through hard pruning based on the pruning rate of each convolutional layer in the current module.
Citation Information
Patent Citations
FPGA-oriented deep convolutional neural network accelerator and design method
CN113487012A
Light-weight pest target detection method based on depth separable convolution
CN118279718A
Cited By
Lightweight target detection method and system based on convolutional neural network
CN120258049A
A lightweight target detection method and system based on convolutional neural network
CN120258049B