Structured channel pruning target recognition lightweight method and device and medium
Through the structured channel pruning method, the batch normalized layer scaling factor is used to measure the importance of the convolution channel, the redundant connections of the bomb-load platform model are cut off, and lightweight object detection is realized, which solves the model adaptability problem under the limitation of hardware resources and improves detection efficiency and accuracy.
Patent Information
- Application Number
- CN202510464245.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
The hardware resources of the bomb-loaded platform are limited, and the traditional target detection algorithm model is huge and time-consuming to calculate, making it difficult to directly apply to the bomb-loaded platform.
The structured channel pruning method is adopted, and the scaling factor of the batch normalization layer is used as the standard to measure the importance of convolutional channels. The model is lightweighted through pruning redundant weight connection, combined with sparse training and cropping threshold calculation.
Significantly reduce the amount of model parameters and calculations, maintain detection accuracy, adapt to the hardware resource limitations of the bomb-mounted platform, and improve the deployment efficiency and reliability of the model in resource-constrained environments.
Smart Images

Figure CN120298674A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and specifically relates to a lightweight method for structured channel pruning target recognition based on improved Slimming. Background Art
[0002] Missile-borne platforms are usually carried on aircraft such as missiles. Their hardware resources such as computing power, storage space, power supply, etc. are strictly limited. These limiting conditions make it difficult to directly apply traditional target detection algorithm models to missile-borne platforms. Therefore, it is necessary to lightweight the model while ensuring the algorithm performance to adapt to the hardware resource limitations of missile-borne platforms.
[0003] Traditional target detection algorithm models usually contain a large number of parameters and complex calculation processes, resulting in a large model volume and long calculation time. Lightweight processing can reduce model parameters and calculation amount through methods such as pruning, quantization, knowledge distillation, etc., thereby reducing the model's demand for hardware resources.
[0004] Therefore, in order to ensure the lightweight requirements of the target detection algorithm model under the limited hardware resources of the missile-borne platform, the present application specifically proposes a lightweight method for structured channel pruning target recognition based on improved Slimming to solve the above technical problems. Summary of the Invention
[0005] The main purpose of the present invention is to provide a lightweight method for structured channel pruning target recognition to solve the technical problems proposed in the background art.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A lightweight method for structured channel pruning target recognition, used for lightweight reconstruction of the target detection algorithm model, and the following steps are executed by a computer device:
[0008] S1. During the training process of the network model, the scaling factor of the batch normalization layer is used as an important criterion for measuring the importance of convolution channels, and a linear transformation is performed on each output feature image pixel of the convolution layer to accelerate the convergence of the model network;
[0009] S2. Calculate the pruning threshold according to the maximum value of the batch normalization layer scaling factor, and delete the redundant weight connections with a model performance influence lower than the pruning threshold through channel structure pruning compression operation to reduce the model parameter quantity and calculation amount;
[0010] S3. Achieve model lightweight through network retraining fine-tuning.
[0011] Preferably, the calculation expression for performing a linear transformation on the output feature image pixels in step S1 is:
[0012]
[0013] Among them, BN(x) is the output of the channel batch normalization layer when the input is x; μ and σ are the mean and variance of the input data of each batch, and γ and β are the scaling factor and translation factor of the batch-normalized data after the model training is completed, which are used to adjust the data feature distribution.
[0014] Preferably, when γ and β approach 0, the output of the channel convolution is close to 0, indicating that the contribution of the input of this channel to the output is small.
[0015] Preferably, in the S1 step, during the training process of the model network, γ is also made to approach 0 through sparse training, and the specific operation process includes:
[0016] Adding a regularization constraint term for the scaling factor γ to the loss function, and jointly training the network weights and the scaling factor γ to make the scaling factor γ sparse. The loss function L s has the following expression:
[0017]
[0018] where L is the normal training loss of the model, s is a hyperparameter that balances the network loss and the sparsity loss, and is used to obtain a suitable sparse network through control training. R(γ) is the sparse regularization term for the scaling factor, taking the L1 regularization, and Γ is an array composed of the scaling factors γ arranged from small to large.
[0019] Preferably, the sparse regularization term R(γ) adopts the L1 regularization constraint, and there is: R(γ) = ‖γ‖1.
[0020] Preferably, the specific operation process of the channel structure pruning and compression operation in the S2 step includes:
[0021] S21. Set the number of pruning layers to N cut_l layers, the number of channels of the i-th pruning layer is N if each, and each pruning layer corresponds to a batch normalization layer containing N if scaling factor γ parameters;
[0022] S22. Calculate the pruning threshold according to the scaling factor γ, and perform the channel structure pruning operation based on the pruning threshold;
[0023] S23. Constrain the number of channels N cut after pruning to a multiple of 8, and the calculation expression for the pruning channel constraint is:
[0024]
[0025] where N cut_8 is the final number of channels, Denotes rounding up.
[0026] Preferably, during the pruning operation of the channel structure in step S22, the influence of the value of the translation factor β is superimposed on the connected activation function and convolutional layer.
[0027] Preferably, the specific calculation method of the pruning threshold in step S22 includes:
[0028] Set the preset global pruning ratio as κ, and set the pruning threshold for each channel as
[0029] Statistically calculate the maximum value of the scaling factor γ of each batch normalization layer and take the minimum value as the upper limit thresh of the pruning ratio threshold;
[0030] The pruning threshold Is constrained within the range less than thresh to maintain the integrity of the network structure and avoid all channels of the batch normalization layer being pruned. The specific expression of the upper limit of the pruning threshold is:
[0031]
[0032] where N Γ Is the number of parameters of the scaling factor γ of all batch normalization layers, and Γ is an array composed of the scaling factors γ arranged from small to large.
[0033] Preferably, during the execution of the channel structure pruning and compression operation in step S2, the sparsified model is also pruned by adjusting different pruning ratios and fine-tuned to restore the accuracy, and the optimal pruning ratio is selected by analyzing indicators such as model size, number of parameters, model GFLOPS, forward inference time, and average detection accuracy.
[0034] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the above method.
[0035] On yet another aspect, the present invention also discloses a computer device including a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the above method.
[0036] From the above technical solutions, the present invention provides a lightweight method for object recognition with structured channel pruning. Compared with the prior art, the present invention has the following advantages:
[0037] 1. The method of the present invention does not require operations such as encoding, facilitating software and hardware implementation and deployment. Using the scaling factor of the batch normalization layer as an important measure standard for convolutional channels, redundant weight connections with little impact on performance are deleted, reducing the number of model parameters and computational complexity, significantly reducing the difficulty of hardware deployment in resource-constrained environments. At the same time, it can keep the detection accuracy almost unaffected, improve the comprehensive performance of the model, and achieve the effect of model lightweighting.
[0038] 2. By sparsely training the network and adding a regularization constraint term for the scaling factor in the loss function, the present invention can achieve the overall sparsity of the network, thereby realizing effective model compression without affecting the model detection accuracy.
[0039] 3. The pruning threshold of the present invention is calculated based on the maximum value of the scaling factor of the batch normalization layer during the calculation process, which can dynamically determine the pruning criteria for each layer, thereby accurately identifying and removing redundant channels, ensuring the retention of key feature channels while reducing the model complexity, and further optimizing the structure and performance of the model.
[0040] 4. By constraining the number of channels after pruning to a multiple of 8 during the channel structure pruning operation based on the pruning threshold, counting the maximum value of the scaling factor of each batch normalization layer and taking the minimum value as the upper limit of the pruning ratio threshold, and finally constraining the pruning threshold within a range less than the threshold upper limit, the present invention can maintain the integrity of the network structure, avoid all channels of the batch normalization layer being pruned, thus preventing the loss of key information caused by excessive pruning, ensuring the stability of the model detection accuracy, and making the lightweighted model more reliable and efficient in practical applications.
[0041] 5. By analyzing the model size, number of parameters, GFLOPS, forward inference time, and average detection accuracy metrics under different pruning ratios to select the optimal pruning ratio, the present invention can balance the relationship between model lightweighting and maintaining sufficient detection accuracy, maximize the mAP effect while meeting the real-time requirements, and ensure that the model can efficiently and accurately perform object detection tasks under resource-limited conditions.
[0042] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The specification drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0044] Figure 1 Schematic diagram of the pruning process of the present invention;
[0045] Figure 2 γ histogram distribution diagram of the present invention;
[0046] Figure 3 Schematic diagram of channel pruning of the present invention;
[0047] Figure 4 Schematic diagram for comparing the curves of different indicators of the present invention changing with the pruning ratio;
[0048] Figure 5 Schematic diagram of the change in the number of channels in each layer before and after pruning of the present invention. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] In the embodiment, refer in detail to Figures 1 to 5 .
[0051] In order to meet the requirements for lightweighting the target detection algorithm model under the condition of limited hardware resources on the missile-borne platform, as Figure 1 shown. The embodiment of the present invention proposes a lightweight method for target recognition of structured channel pruning based on improved Slimming, including the following steps: using the scaling factor γ of the batch normalization layer as the standard for measuring the importance of convolutional channels, deleting redundant weight connections with little impact on performance, reducing the model parameter quantity and computational amount; through network retraining and fine-tuning to achieve model lightweighting.
[0052] The batch normalization (BN) layer is a common standard in current network designs and is usually located after the convolutional layer. The BN layer performs a linear transformation on each output feature image pixel of the convolutional layer, which can accelerate the convergence of the network and improve the network performance. The expression is:
[0053]
[0054] Where: BN(x) is the output when the input of the channel batch normalization layer is x; μ and σ are the mean and variance of the input data for each batch, and γ and β are the scaling factor and translation factor of the batch-normalized data after the model training is completed, which are generally constants, indicating scaling and translation of the batch-normalized data for adjusting the data feature distribution.
[0055] When γ and β approach 0, the output of this channel convolution will also approach 0, indicating that the contribution of the input of this channel to the output is very small. Therefore, it can be considered that removing redundant channels with a contribution of 0 has little impact on the model performance.
[0056] Since after network training, γ usually follows a normal distribution and there are few parameters close to 0, the model cannot directly perform pruning operations. Therefore, in order to make γ tend to 0, a sparse training method is adopted. An L1 regularization constraint term for γ is added to the loss function, and the network weights and the scaling factor γ are jointly trained to make γ sparse. The loss function can be expressed as follows:
[0057]
[0058] R(γ) = ‖γ‖1
[0059] where L is the normal training loss of the model, s is a hyperparameter that balances the network loss and the sparsity loss, and is used to obtain a suitable sparse network through training. R(γ) is the sparse regularization term for the scaling factor, taking L1 regularization. Γ is an array composed of the scaling factors γ arranged from small to large.
[0060] By performing sparse training on the network and adding a regularization constraint term for the scaling factor to the loss function, the overall sparsity of the network can be achieved, thereby enabling effective model compression without affecting the detection accuracy of the model.
[0061] Figure 2 In (a) and (b), they are the histogram distribution of the scaling factor γ of the BN layer corresponding to the layer to be pruned before and after sparse training. In the figure, the colors from dark to light respectively correspond to the network layers from shallow to deep. The y-axis is the serial number of the pruning layer, the x-axis is the range of γ values, and the total number of γ values is the number of channels of this layer. When drawing the histogram, the γ values of each layer are divided into 30 blocks, and the number of γ values falling into each block is counted as the z-axis to obtain the two-dimensional histogram distribution of the γ values of each layer.
[0062] Figure 2 As can be seen from (a) in, since γ is initialized to 1 in the training initialization stage, after normal training, the distribution mean of γ of each layer is near 1.
[0063] From Figure 2 As can be seen from (b) in, the distribution of γ in the shallow network is more uniform and does not all approach 0, indicating that the weights of the shallow network are more important for target feature extraction and the channel redundancy is low. Therefore, in subsequent pruning, it corresponds to a lower pruning number. As the network depth increases, the values and distribution of the scaling factor γ are closer to 0, indicating that after sparse training, the channel redundancy of the deep network increases significantly, and a higher pruning rate can be set.
[0064] Figure 2In (c), the histograms corresponding to the γ values of all layers are counted, indicating that after sparse training, the network has achieved overall sparsification and can perform subsequent channel pruning operations.
[0065] Set the pruning layer to N cut_l layers, and the number of channels of the i-th pruning layer is N if Each pruning layer corresponds to a BN batch normalization layer containing N if scaling factor γ parameters. There are a total of N Γ γ parameters for all BN layers. After arranging them from small to large, they form an array denoted as Γ. Set the global pruning ratio κ for pruning, then the pruning threshold for each corresponding channel is To avoid all channels in a layer being pruned, count the maximum value of γ for each BN layer and take the minimum value as the upper limit thresh of the pruning ratio threshold. By constraining the pruning threshold within the range less than thresh to maintain the integrity of the network structure. At the same time, considering the requirements for hardware deployment acceleration after the reconstruction of the pruned network, the number of channels N cut after pruning is constrained to be a multiple of 8. The upper limit of the pruning threshold and the channel constraint are as shown in the formula:
[0066]
[0067] where N cut_8 is the final number of channels, represents rounding up.
[0068] The pruning threshold is calculated based on the maximum value of the scaling factor of the batch normalization layer during the calculation process, which can dynamically determine the pruning criteria for each layer, thereby accurately identifying and removing redundant channels. While reducing the model complexity, it ensures that the key feature channels are retained, further optimizing the structure and performance of the model. It not only improves the efficiency and effect of model lightweighting but also makes the lightweighted model more adaptable to the deployment requirements in resource-constrained environments. At the same time, it maintains high detection accuracy and reliability, enabling different network structures to flexibly adjust the pruning strategy according to their own characteristics, enhancing the universality and practicality of the algorithm.
[0069] At the same time, by constraining the number of channels after pruning to be a multiple of 8 during the process of performing channel structure pruning operations based on the pruning threshold, counting the maximum value of the scaling factor of each batch normalization layer and taking the minimum value as the upper limit of the pruning ratio threshold, and finally constraining the pruning threshold within the range less than the threshold upper limit, it can maintain the integrity of the network structure, avoid all channels in the batch normalization layer being pruned, thereby preventing the loss of key information caused by excessive pruning, ensuring the stability of the model detection accuracy, and making the lightweighted model more reliable and efficient in practical applications.
[0070] At this time, it is noted that in the batch normalization (BN) layer, when γ approaches 0, the corresponding β parameter is not necessarily 0 and still has a certain impact on the network. Therefore, when pruning channels, the impact of the β value is superimposed on the connected activation function and convolutional layer.
[0071] To further lightweight the model, while pruning and compressing the trained airborne target detection model to meet the hardware deployment requirements and ensuring the model detection accuracy, different pruning ratios κ are used to prune the sparse model of this application and fine-tuning training is performed to restore the accuracy. By analyzing the model size, the number of parameters, the model GFLOPS, the forward inference time, and the average detection accuracy index, the optimal pruning ratio is selected, which can balance the relationship between model lightweight and maintaining sufficient detection accuracy, maximize the mAP effect while meeting the real-time requirements, and ensure that the model can efficiently and accurately perform the target detection task under limited resources. At this time, the channel pruning is as Figure 3 shown.
[0072] Figure 3 In the channel pruning in [reference], different from the single sparse model based on the scaling factor γ, a structured channel pruning model based on multi-level pruning is adopted according to the different influence degrees of the weight parameters in each layer of the network. First, the values of γ in each layer of the network are counted. Then, all the γ values are sorted according to the absolute value size, and a pruning threshold is set. Finally, according to the pruning threshold σ x -σ z range for structured pruning to obtain the channels to be pruned.
[0073] As the pruning ratio increases, the number of pruned channels increases, resulting in a gradual decrease in the number of model parameters, the amount of computation, and the model size, making the model more lightweight. When the pruning ratio is 0, it corresponds to the initial model of sparse training. At this time, the model size is 14.1MB, the amount of computation is 16.1GFLOPS, and the model running time is 10.8ms. When the pruning ratio reaches 90%, the model size is reduced to 600KB, the amount of computation is only 3.5GLOPS, and the corresponding model running time is 8.2ms.
[0074] Figure 4 In (1), (2), and (3) of [reference], the number of model parameters, the model volume size, and the model computation amount decrease significantly with the pruning ratio, improving the feasibility and efficiency of the airborne target detection algorithm for embedded operation. Figure 4 In (4) of [reference], the number of model channels decreases significantly with the model volume. Figure 4 In (5) of [reference], the running time decreases accordingly.
[0075] Figure 4In (6), when the pruning ratio is lower than 50%, the accuracy of the fine-tuned model still drops significantly. The missile-borne target detection task needs to balance the real-time performance of the model while ensuring the detection accuracy, that is, under the conditions of meeting the real-time performance of the algorithm and having a smaller number of model parameters, size, and computational complexity, the pruning ratio corresponding to the maximum mAP should be selected as much as possible. Therefore, the final selected pruning ratio is 50%.
[0076] When the global pruning ratio κ is 0.5, the pruning layers of the channel numbers of each layer before and after pruning are the convolutional layers in the feature extraction network from layer 0 to layer 51, and layers 55 - 95 correspond to the multi-scale prediction feature fusion module.
[0077] The pruning ratio of the shallow features in the feature extraction network is low, indicating that the weights of the shallow network are relatively important and the scaling factor γ is relatively larger. Therefore, the number of pruned channels is small. As the network depth increases, the pruning ratio increases after layer 37.
[0078] In the prediction feature fusion module, the deep network features usually correspond to larger channel numbers and smaller resolutions. Referring to Figure 5 , it can be seen that a large proportion of channel pruning is also performed on the convolutional layers with large channel numbers, indicating that there is channel information redundancy in the network. After channel pruning, the channels that contribute less to feature extraction in the convolutional layer are removed, reducing the number of network parameters and computational complexity, and at the same time not overly affecting the model's feature extraction ability. Pruning may cause a loss of accuracy, and the pruned network is retrained to fine-tune the network weights to achieve the effect of restoring the network accuracy.
[0079] In summary, this method does not require operations such as encoding, which is convenient for software and hardware implementation and deployment. Using the scaling factor of the batch normalization layer as the standard for measuring the importance of convolutional channels, redundant weight connections with little impact on performance are deleted, reducing the number of model parameters and computational complexity, significantly reducing the difficulty of hardware deployment in resource-constrained environments, and at the same time being able to keep the detection accuracy almost unaffected, improving the comprehensive performance of the model, and achieving the effect of model lightweight.
[0080] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.
[0081] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, the memory storing a computer program, which when executed by the processor causes the processor to execute the steps of the above method.
[0082] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which, when running on a computer, enables the computer to execute any one of the above-mentioned lightweight object recognition methods based on improved Slimming structured channel pruning in the above embodiments.
[0083] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. For the explanations, examples, and beneficial effects of related content, reference can be made to the corresponding parts in the above method.
[0084] The embodiments of the present application also provide an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.
[0085] The memory is used to store a computer program.
[0086] The processor is used to implement the above-mentioned lightweight object recognition method based on improved Slimming structured channel pruning when executing the program stored in the memory.
[0087] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0088] The communication interface is used for communication between the above electronic device and other devices.
[0089] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0090] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0091] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, user equipment, mobile station, mobile terminal, etc. The terminal device can be a mobile phone, smart TV, wearable device, tablet computer, computer with wireless transceiver function, virtual reality terminal device, augmented reality terminal device, wireless terminal in industrial control, wireless terminal in driverless, wireless terminal in remote surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, and so on. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.
[0092] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive), etc.
[0093] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
[0094] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.
[0095] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, scenario B, or the scenario where both A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality of" means more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
Claims
1. A lightweight method for structured channel pruning object recognition, characterized in that For the lightweight reconstruction of the target detection algorithm model, the following steps are performed by a computer device: During the training process of the network model, the scaling factor of the batch normalization layer is used as an important criterion for measuring the importance of convolutional channels. Each output feature image pixel of the convolutional layer is linearly transformed to accelerate the convergence of the model network; Calculate the pruning threshold according to the maximum value of the scaling factor of the batch normalization layer. Through the channel structure pruning and compression operation, redundant weight connections with a model performance influence lower than the pruning threshold are deleted to reduce the number of model parameters and the computational amount; Through network retraining and fine-tuning to achieve model lightweighting.
2. The lightweight method for structured channel pruning target recognition according to claim 1, characterized in that The calculation expression for linearly transforming the output feature image pixels in step S1 is: Among them, BN(x) is the output when the input of the channel batch normalization layer is x; μ and σ are the mean and variance of each batch of input data, and γ and β are the scaling factor and translation factor of the batch normalization data after the model training is completed, which are used to adjust the data feature distribution; When γ and β approach 0, the convolutional output of this channel approaches 0, indicating that the input of this channel contributes little to the output.
3. The lightweight method for identifying structured channel pruning targets according to claim 2, characterized in that In step S1, during the training process of the model network, γ is also made to approach 0 through sparse training. The specific operation process includes: Add a regularization constraint term for the scaling factor γ to the loss function, and jointly train the network weights and the scaling factor γ to sparsify the scaling factor γ. The loss function L s is expressed as: Among them, L is the normal training loss of the model, s is a hyperparameter that balances the network loss and the sparsity loss, which is used to obtain a suitable sparse network through control training, R(γ) is the sparse regularization term for the scaling factor, taking L1 regularization, and Γ is an array composed of the scaling factors γ arranged from small to large.
4. The lightweight method for identifying structured channel pruning targets according to claim 3, wherein, The sparse regularization term R(γ) adopts L1 regularization constraint, and there is: R(γ) = ‖γ‖1.
5. The lightweight method for identifying structured channel pruning targets according to claim 3, wherein The specific operation process of the channel structure pruning and compression operation in step S2 includes: S21. Set the cropping layer to be the N cut_l th layer, and the number of channels of the i-th cropping layer is N if ; each cropping layer corresponds to a batch normalization layer containing N if scaling factor γ parameters; S22. Calculate the pruning threshold according to the scaling factor γ, and perform the channel structure pruning operation based on the pruning threshold; S23. Constrain the number of channels N after pruning cut to be a multiple of 8. The calculation expression for the pruning channel constraint is: Where N cut_8 is the final number of channels, represents rounding up.
6. The lightweight method for identifying structured channel pruning targets according to claim 5, wherein, During the execution of the channel structure pruning operation in step S22, the influence of the value of the translation factor β is superimposed on the connected activation function and convolutional layer.
7. The lightweight method for identifying structured channel pruning targets according to claim 5, characterized in that The specific calculation method of the pruning threshold in step S22 includes: The preset global cropping ratio is κ, and the cropping thresholds corresponding to each channel are set to Statistically calculate the maximum value of the scaling factor γ of each batch normalization layer and take the minimum value as the upper limit thresh of the pruning ratio threshold; Clip the clipping threshold to a range less than thresh to maintain the integrity of the network structure and prevent all channels of the batch normalization layer from being clipped. The specific expression for the upper limit of the clipping threshold is as follows: Among them, N Γ is the number of scaling factor γ parameters of all batch normalization layers, and Γ is an array composed of the scaling factors γ arranged from small to large.
8. A computer-readable storage medium, characterized in that, There is a computer program, which when executed by a processor causes the processor to execute the steps of the method according to any one of claims 1 to 7.
9. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program, which when executed by the processor causes the processor to execute the steps of the method according to any one of claims 1 to 7.
Citation Information
Cited By
Endoscope image detection method, device, equipment, medium and product
CN121481930A