Method for designing lightweight network structure
By performing deep convolution, network pruning and structural optimization on the yolov7-tiny network, the problem of lack of guidelines for lightweight neural network design is solved, and the rapid operation and low memory usage of high-precision lightweight networks are achieved.
Patent Information
- Application Number
- CN202410191539.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-08-22
AI Technical Summary
In the prior art, the design of lightweight neural networks lacks wide and common guidelines, resulting in different requirements of different chip platforms. The existing methods such as distillation, quantization and NAS are easy to cure, which cannot effectively improve operation speed and reduce memory usage.
By changing the normal convolution to deep convolution, network pruning, modifying network structure blocks, reducing the number of network outputs and modifying the network structure, it specifically includes changing the ordinary convolution with a convolution kernel of 3x3 to 3x3 depth convolution and 1x1 ordinary convolution, the number of pruning channels is 0.25 times, modifying the concat operator and output number, adjusting anchors using the K-means clustering algorithm, and optimizing the yolov7-tiny network structure.
With an uncomplex network structure, a high-precision lightweight network is realized, which improves operating speed and reduces memory usage, meeting the time requirements within 200ms.
Smart Images

Figure CN120524992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent video processing, and in particular relates to a method for designing a lightweight network structure. Background Art
[0002] With the rapid development of artificial intelligence (AI), neural network models are widely used. Lightweight networks, which offer fewer parameters, less computational effort, and faster inference times than their heavier counterparts, are gaining increasing attention. Lightweight networks are designed to reduce both size and speed while maintaining accuracy.
[0003] There are currently no widely accepted guidelines for manually designing lightweight networks. Different designs are generated to meet the needs of different chip platforms (different chip architectures), and deployment and model performance testing are then performed on each hardware platform.
[0004] Designed based on human experience, common approaches include: reducing the number of output channels per layer; replacing a single large convolution with multiple small convolutions; and replacing regular convolutions with grouped convolutions or depthwise separable convolutions. MobileNet V1 and V2 fall into this category.
[0005] These are obtained using reinforcement learning using the Neural Architecture Search (NAS) method. MobileNet V3 falls into this category. NAS surpasses previously designed networks in image classification and language modeling tasks.
[0006] Common technologies for lightweight neural networks in existing technologies include distillation, pruning, quantization, weight sharing, low-level partitioning, lightweight attention modules, dynamic network architecture / training methods, lightweight networks including architecture design, NAS (neural architecture exploration), hardware support, etc.
[0007] However, attempts at distillation have shown little effect on lightweight networks; quantization has little effect when the model is running on the CPU alone; and NAS tends to solidify easily, and if a search network is found that fails to meet the requirements, it cannot be used well.
[0008] Additionally, commonly used technical terms include:
[0009] 1. Yolov7: It recognizes and locates objects based on a deep neural network and runs very fast. Its main contributions are:
[0010] Model reparameterization;
[0011] Label allocation strategy;
[0012] ELAN efficient network architecture;
[0013] Training with an assisted head.
[0014] 2. The k-means clustering algorithm (K-means clustering algorithm) is an iterative clustering analysis algorithm. Summary of the Invention
[0015] In order to solve the above problems, the purpose of this application is to improve the overall operating speed, reduce memory usage, and avoid unnecessary hardware settings by designing a lightweight network.
[0016] Specifically, the present invention provides a method for designing a lightweight network structure, the method comprising the following steps:
[0017] S1. Change ordinary convolution to depth convolution:
[0018] The ordinary convolution with a convolution kernel of 3x3 is changed to a depthwise convolution with a convolution kernel of 3x3 plus a normal convolution with a convolution kernel of 1x1. The depthwise convolution with a convolution kernel of 3x3 is used to improve the receptive field, and the 1x1 normal convolution is used to change the number of channels.
[0019] S2. Network pruning, that is, pruning the number of channels in the network:
[0020] Change width_multiple to 0.25, that is, the number of channels is 0.25 times the original one. The width_multiple is a parameter given in the yolo framework. This parameter can control the scaling of the number of channels. The value of 0.25 is based on the requirements of the lightweight network design, which requires the time to be within 200ms.
[0021] S3. Modify the network structure block:
[0022] Change the concat operator of the four layers of the backbone of the network structure to the concat operator of the two layers, that is, change [-1, -2, -3, -4] to [-1, -4], and only retain the results of the adjacent previous layer and the four layers above the current row. The convolution operator time and parameter amount after concat are directly halved. The number of input channels before modification is 4*C in , after modification it becomes 2*C in ;
[0023] S4. Reduce the number of network outputs:
[0024] The network has three outputs, for strides of 8, 16, and 32. Since the lightweight network cannot obtain much information, the three outputs of the network are changed to two outputs, that is, outputs for strides of 16 and 32. Since the number of outputs is reduced to 2, the number of anchors must also be changed from 3x3 to 3x2. The changed anchors can be obtained by the K-means clustering algorithm based on the training data set. S5. Modify the network structure:
[0025] Since the network output is reduced to 2, the original network structure is stride downsampled to 32, that is, the size of the current layer feature map is 1 / 32 of the input image, and then two upsamples are upsampled to stride 8 for output, and then stride to 16 and 32, and output at stride 16 and 32 respectively;
[0026] After the modification, the network structure is strided to 32, then upsampled to a stride of 16 and output, and then strided to 32.
[0027] The network includes the yolov7 framework.
[0028] In step S1, the modified network is yolov7-tiny. It is assumed that the ordinary convolution with a convolution kernel of 3x3, 32 input channels, and 64 output channels can be changed to a depthwise convolution with a convolution kernel of 3x3, 32 input channels, and 32 output channels plus an ordinary convolution with a convolution kernel of 1x1, 32 input channels, and 64 output channels;
[0029] The original ordinary convolution with a convolution kernel of 3x3 has a parameter value of W1=Cin*Cout*3*3. After changing it to a depthwise convolution with a convolution kernel of 3x3 plus an ordinary convolution with a convolution kernel of 1x1, the parameter value is W1=Cin*3*3+Cin*Cout*1*1. After simplifying W2 / W1, it becomes 1 / 9+1 / Cout>1 / 9.
[0030] In step S2, a channel pruning operation is performed on the yolov7-tiny network.
[0031] In step S3, the network of the yolov7 framework is written in a cfg file. In order, one operator occupies one line, the previous line of the current line is the adjacent previous layer, and the 4th layer from the current line is 4 layers above the current line, which is also the fourth layer from the bottom.
[0032] In step S4, the K-means clustering algorithm comprises the following steps:
[0033] Cluster the data into K groups, and randomly select K objects as the initial class centers;
[0034] Then calculate the distance between each object and each seed cluster center, and assign each object to the cluster center closest to it;
[0035] The cluster centers and the objects assigned to them represent a cluster;
[0036] Each time a sample is assigned, the cluster center is recalculated based on the existing objects in the cluster; this process will be repeated until a termination condition is met;
[0037] The termination condition is that no or a minimum number of objects are reassigned to different clusters, or no or a minimum number of cluster centers change again, and the sum of squared errors is locally minimized.
[0038] In step S5, the input image size is 640x640, and the feature map size is 20x20.
[0039] Therefore, the advantage of this application is that, without designing a complex network structure, a high-precision lightweight network can be obtained by simply manually modifying a classic open-source high-precision network, thereby improving the overall running speed and reducing memory usage. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention.
[0041] Figure 1 This is a schematic diagram of the original yolov7 network structure.
[0042] Figure 2 It is a schematic diagram of the network structure after modification of this application.
[0043] Figure 3 It is a flow chart of the present application method.
[0044] Figure 4 It is a schematic diagram of the original network structure in step S5 of this method.
[0045] Figure 5 It is a schematic diagram of the network structure after modification in step S5 of this method. DETAILED DESCRIPTION
[0046] In order to more clearly understand the technical content and advantages of the present invention, the present invention is now further described in detail with reference to the accompanying drawings.
[0047] like Figure 3As shown, the present invention provides a method for designing a lightweight network structure. This application takes the yolov7 framework as an example, and the main implementation steps are as follows:
[0048] S1. Change ordinary convolution to depth convolution;
[0049] S2. Network pruning;
[0050] S3. Modify the network structure block;
[0051] S4. Reduce the number of network outputs;
[0052] S5. Modify the network structure.
[0053] Furthermore, the method specifically includes the following steps:
[0054] Step S1. Change ordinary convolution to depth convolution
[0055] The network of the model here is a modified network, namely yolov7-tiny. First, the ordinary convolution with a convolution kernel of 3x3 is changed to a depthwise convolution with a convolution kernel of 3x3 plus a normal convolution with a convolution kernel of 1x1. The depthwise convolution with a convolution kernel of 3x3 is used to improve the receptive field, and the 1x1 normal convolution is used to change the number of channels.
[0056] For example, a normal convolution with a 3x3 kernel, 32 input channels, and 64 output channels can be changed to a depthwise convolution with a 3x3 kernel, 32 input channels, and 32 output channels plus a normal convolution with a 1x1 kernel, 32 input channels, and 64 output channels.
[0057] After modification, the parameter amount is more than 1 / 9 of the original, and the minimum is close to 1 / 9;
[0058] The original ordinary convolution with a convolution kernel of 3x3 has a parameter of: W1 = Cin * Cout * 3 * 3. After changing to a depthwise convolution with a convolution kernel of 3x3 plus a normal convolution with a convolution kernel of 1x1, its parameter is: W1 = Cin * 3 * 3 + Cin * Cout * 1 * 1. After simplifying W2 / W1, 1 / 9 + 1 / Cout> 1 / 9;
[0059] Step S2. Network pruning
[0060] Perform channel pruning on the yolov7-tiny network:
[0061] Due to the accuracy and time requirements, width_multiple is changed to 0.25. The 0.25 here is determined according to the lightweight network requirements of the design. The time requirement here is within 200ms, which can be met by 0.25, that is, the number of channels is 0.25 times the original number.
[0062] Step S3. Modify the network structure block
[0063] Put the structure block of yolov7-tiny into Figure 1 As shown, the concat operator of the backbone's four layers is changed to a concat operator of two layers, that is, [-1, -2, -3, -4] is changed to [-1, -4], and only the results of the adjacent previous layer and the 4 layers above the current row are retained. The network of the yolov7 framework is written in a cfg file. In order, one operator occupies one line, and the previous line of the current line is the adjacent previous layer. The 4th layer from the current line upward is the 4th layer above the current line, which is also the 4th layer from the bottom. The convolution operator time and parameter amount after concat are directly halved. The number of input channels before modification is 4*C in , after modification it becomes 2*C in ; The modified network structure is as follows Figure 2 shown.
[0064] Step S4. Reduce the number of network outputs
[0065] The yolov7-tiny network has three outputs, which are output when the stride is 8, 16, and 32. Since the lightweight network cannot obtain too much information, the three outputs of the network are changed to two outputs, that is, output when the stride is 16 and 32. Since the number of outputs is reduced to 2, the number of anchors is also changed from 3x3 to 3x2. The changed anchors can be obtained by the K-means clustering algorithm based on the training data set.
[0066] The steps of the K-means clustering algorithm are:
[0067] Cluster the data into K groups, and randomly select K objects as the initial class centers;
[0068] Then calculate the distance between each object and each seed cluster center, and assign each object to the cluster center closest to it;
[0069] The cluster centers and the objects assigned to them represent a cluster;
[0070] Each time a sample is assigned, the cluster center is recalculated based on the existing objects in the cluster; this process will be repeated until a termination condition is met;
[0071] The termination condition is that no or a minimum number of objects are reassigned to different clusters, or no or a minimum number of cluster centers change again, and the sum of squared errors is locally minimized.
[0072] Step S5. Modify the network structure, such as Figure 4 、 Figure 5 As shown:
[0073] Since the network output is reduced to 2, the original network structure is stride downsampled to 32, that is, the size of the feature map of this layer is 1 / 32 of the input image. For example, the input image size is 640x640, the size of the feature map is 20x20, and two upsamples are upsampled to a stride of 8 and then stride to 16 and 32, and output at stride 16 and 32 respectively; Figure 4 As shown, the circled out1 can be modified to change 3 outputs into 2 outputs and remove it.
[0074] like Figure 5 As shown in the figure, the modified network structure is strided to 32, then upsampled to a stride of 16 and output, and then output when strided to 32. This saves an upsampling block and a downsampling block, directly saving 12 layers, and about 1 / 9 of the time and parameters.
[0075] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for designing a lightweight network structure, characterized in that: The method comprises the following steps: S1. Change ordinary convolution to depth convolution: The ordinary convolution with a convolution kernel of 3x3 is changed to a depthwise convolution with a convolution kernel of 3x3 plus a normal convolution with a convolution kernel of 1x1. The depthwise convolution with a convolution kernel of 3x3 is used to improve the receptive field, and the 1x1 normal convolution is used to change the number of channels. S2. Network pruning, that is, pruning the number of channels in the network: Change width_multiple to 0.25, that is, the number of channels is 0.25 times the original one. The width_multiple is a parameter given in the yolo framework. This parameter can control the scaling of the number of channels. The value of 0.25 is based on the requirements of the lightweight network design, which requires the time to be within 200ms. S3. Modify the network structure block: Change the concat operator of the four layers of the backbone of the network structure to the concat operator of the two layers, that is, change [-1, -2, -3, -4] to [-1, -4], and only retain the results of the adjacent previous layer and the four layers above the current row. The convolution operator time and parameter amount after concat are directly halved. The number of input channels before modification is 4*C in , after modification it becomes 2*C in ; S4. Reduce the number of network outputs: The network has three outputs, which are output when the stride is 8, 16, and 32. Since the lightweight network cannot obtain much information, the three outputs of the network are changed to two outputs, that is, output when the stride is 16 and 32. Since the number of outputs is reduced to 2, the number of anchors must also be changed from 3x3 to 3x2. The changed anchors can be obtained by the K-means clustering algorithm based on the training data set; S5. Modify the network structure: Since the network output is reduced to 2, the original network structure is stride downsampled to 32, that is, the size of the current layer feature map is 1 / 32 of the input image, and then two upsamples are upsampled to stride 8 for output, and then stride to 16 and 32, and output at stride 16 and 32 respectively; After the modification, the network structure is strided to 32, then upsampled to a stride of 16 and output, and then strided to 32.
2. A method for designing a lightweight network structure according to claim 1, characterized in that: The network includes the yolov7 framework.
3. The method for designing a lightweight network structure according to claim 2, wherein: In step S1, the modified network is yolov7-tiny. It is assumed that the ordinary convolution with a convolution kernel of 3x3, 32 input channels, and 64 output channels can be changed to a depthwise convolution with a convolution kernel of 3x3, 32 input channels, and 32 output channels plus an ordinary convolution with a convolution kernel of 1x1, 32 input channels, and 64 output channels; The original ordinary convolution with a convolution kernel of 3x3 has a parameter value of W1=Cin*Cout*3*3. After changing it to a depthwise convolution with a convolution kernel of 3x3 plus an ordinary convolution with a convolution kernel of 1x1, the parameter value is W1=Cin*3*3+Cin*Cout*1*1. After simplifying W2 / W1, it becomes 1 / 9+1 / Cout>1 / 9.
4. The method for designing a lightweight network structure according to claim 3, wherein: In step S2, a channel pruning operation is performed on the yolov7-tiny network.
5. The method for designing a lightweight network structure according to claim 4, characterized in that: In step S3, the network of the yolov7 framework is written in a cfg file. In order, one operator occupies one line, the previous line of the current line is the adjacent previous layer, and the 4th layer from the current line is 4 layers above the current line, which is also the fourth layer from the bottom.
6. The method for designing a lightweight network structure according to claim 5, characterized in that: In step S4, the K-means clustering algorithm comprises the following steps: Cluster the data into K groups, and randomly select K objects as the initial class centers; Then calculate the distance between each object and each seed cluster center, and assign each object to the cluster center closest to it; The cluster centers and the objects assigned to them represent a cluster; Each time a sample is assigned, the cluster center is recalculated based on the existing objects in the cluster; this process will be repeated until a termination condition is met; The termination condition is that no or a minimum number of objects are reassigned to different clusters, or no or a minimum number of cluster centers change again, and the sum of squared errors is locally minimized.
7. The method for designing a lightweight network structure according to claim 6, characterized in that: In step S5, the input image size is 640x640, and the feature map size is 20x20.