Monolithic calculation imaging edge reconstruction method based on end-to-end sensitivity analysis
By introducing end-to-end sensitivity analysis into monolithic computing imaging technology, customized model compression strategies solve the problems of time delay and high power consumption in edge environments of monolithic computing imaging technology, and significantly improve edge inference speed and adaptability.
Patent Information
- Application Number
- CN202510415259.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
Monolithic computing imaging technology faces problems of time delay and high power consumption in practical applications, especially in high-speed and low-power application scenarios, the adaptability and complexity control of existing image reconstruction algorithms in edge chip environments is insufficient.
A single-chip computing imaging edge reconstruction method based on end-to-end sensitivity analysis is proposed. By performing performance evaluation, pruning sensitivity and quantization sensitivity analysis on pre-trained deep learning models, customizing model compression strategies, differentiating different sensitivity modules, and balancing reconstruction quality and computing efficiency.
It significantly improves the speed of edge reasoning, improves the adaptability of the model in the edge environment, realizes differentiated processing of different sensitivity modules, and effectively balances the reconstruction quality and computing efficiency.
Smart Images

Figure CN119941541A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computational imaging technology, and in particular to a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis. Background Art
[0002] As an emerging field that deeply integrates optics and computing technologies, computational imaging has broken many limitations of traditional optical system design and greatly expanded the dimensional boundaries of information perception by virtue of the collaborative design of optical acquisition and computing algorithms. In recent academic exploration and technical practice, computational imaging has not only significantly improved the performance indicators of basic optical imaging, such as resolution and contrast, but has also successfully expanded the dimension of information perception to more complex and sophisticated levels such as phase, spectrum, polarization, light field, and depth of field. This technological breakthrough makes it possible to obtain imaging quality comparable to that of complex optical systems using simple and compact imaging devices.
[0003] However, monolithic computational imaging technology faces severe challenges in its practical application. Due to the introduction of special optical designs specifically for information encoding, it is highly dependent on image reconstruction algorithms, requiring highly complex and precise image reconstruction algorithms to achieve high-quality image restoration. This requirement directly leads to additional time delay problems, while the amount of computation and power consumption requirements increase exponentially. In high-speed, low-power optical application scenarios such as drone remote sensing and biomedical imaging, which have extremely demanding requirements on speed and power consumption, these problems have become key bottlenecks restricting the further development of the technology.
[0004] At present, the research hotspots of image reconstruction algorithms are mainly focused on improving reconstruction performance, such as pursuing higher image clarity, more accurate color reproduction, etc. However, this single research orientation ignores the adaptability of key parameters and overall complexity control of the algorithm in the edge chip environment. Due to the lack of customized design for edge chip characteristics, many reconstruction algorithms that perform well on high-performance GPUs are difficult to achieve their expected results in actual edge application scenarios and have poor adaptability. This is mainly because edge chips have natural limitations in power consumption and computing power, which makes it difficult for computational imaging in actual applications to meet actual needs in terms of latency and frame rate, which seriously hinders the industrialization process and widespread application of the technology.
[0005] As an important means to alleviate the pressure of edge computing, model compression technology has shown certain application potential in related fields such as computer vision and remote sensing. Among them, lightweight compression technologies such as pruning, quantization, distillation, and neural architecture search (NAS) have attracted much attention. As the earliest and relatively mature methods in this field, pruning and quantization have made preliminary attempts in the edge acceleration of non-computational imaging restoration tasks and have achieved some results. However, the optimization goals in these studies often focus on the reduction of the number of model parameters and floating-point operations per second (FLOPs), and fail to fully consider the actual inference time of edge artificial intelligence chips (such as neural processing units NPUs and field-programmable gate arrays FPGAs). This is because deploying algorithms to edge chips is a systematic project involving many complex factors, including operator configuration, chip architecture, hardware design, and memory access bottlenecks. Reducing the number of model parameters and the amount of computation does not mean that the inference speed at the edge can be improved. Although there are pruning optimization solutions for memory consumption, compression optimization directly targeting edge inference speed and latency still faces huge challenges. Therefore, it is of great significance to develop an efficient model compression method to improve the practicality of computational imaging technology. Summary of the invention
[0006] The purpose of the present invention is to propose a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis. By introducing end-to-end sensitivity analysis in computational imaging edge reconstruction, the model compression strategy can be customized according to the characteristics of the edge chip, and differentiated processing of different sensitivity modules can be achieved, effectively balancing the reconstruction quality and computational efficiency, and significantly improving the edge reasoning speed.
[0007] To achieve the above object, the present invention proposes a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis, and the specific steps are as follows: Step S1: Evaluate the performance of the pre-trained deep learning model at the edge, and replace or remove operators with poor hardware support or excessive time consumption; Step S2: analyzing the pruning sensitivity and quantization sensitivity of the deep learning model on the edge chip; Step S3: According to the pruning sensitivity analysis result, a pruning operation with a high pruning rate is performed on the low-sensitivity module, and a pruning operation with a low pruning rate is performed on the high-sensitivity module; Step S4, retraining the deep learning model after sensitivity pruning; Step S5: According to the quantization sensitivity analysis result, low quantization precision processing is performed on the low-sensitivity module, and high quantization precision operation is performed on the high-sensitivity module; Step S6: determine whether the edge reasoning performance of the deep learning model can meet the actual application requirements. If not, execute step S7; if yes, execute step S8; Step S7, reduce the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps S2 to S6 until the edge reasoning performance of the deep learning model meets the actual application requirements; Step S8, determining whether the edge inference speed of the deep learning model can meet the actual application requirements, if not, executing step S9, if yes, executing step S10; Step S9, increase the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps S2 to S8 until the edge inference speed of the deep learning model meets the actual application requirements; Step S10: Obtain a compressed deep learning model.
[0008] Preferably, the specific operations in step S1 are as follows: Step S11: deploy the pre-trained deep learning model to the edge using the tool chain corresponding to the edge hardware platform deployed on the generation; Step S12: Use corresponding performance analysis tools to analyze the operating efficiency and speed of each operator in the deep learning model network structure; Step S13: Find out the operators with poor operating efficiency and speed, and directly delete them from the network structure or replace them with other operators.
[0009] Preferably, in step S2, the specific analysis steps of pruning sensitivity and quantization sensitivity are as follows: Step S21: Based on the trained deep learning model, sensitivity analysis is performed with layers or blocks as the minimum unit; assuming that Deep learning models composed of blocks , the edge performance of the deep learning model is recorded as , pruning or quantizing deep learning models The performance of the block is ,definition ,pass , calculate pruning sensitivity and quantitative sensitivity ;in, The deep learning model piece, For deep learning models Degraded performance after block pruning or quantization, ; Step S22: According to the pruning sensitivity , pruning deep learning models at a fixed rate The block gets the pruned deep learning model , fine-tune and deploy on target AI chips to evaluate edge performance, repeat Get the pruning sensitivity of all blocks again; Step S23: According to the quantization sensitivity , for the deep learning model The block is low-bit quantized to obtain the quantized deep learning model ,The quantized sensitivity of all blocks is also obtained after fine-tuning, deployment and evaluation.
[0010] Preferably, in step S3, the specific steps of the pruning operation are as follows: Step S31, define The input and output channels of the block are and , No. Block The filter is , pruning rate ;in, is the number of output channels remaining after pruning; Step S32: Pruning sensitivity obtained in step S21 , determine the The target pruning rate of the block is as follows: ; in, For deep learning models The target pruning rate of the block, For the The pruning sensitivity of the block, for The target pruning rate function is is the conversion factor from sensitivity to pruning rate, is a positive integer; Step S33: The number is the importance criterion, which filters are pruned in ascending order. For a specific filter, The norm is defined as ;in, K is the convolution kernel size, is the weight, is a floating point weight, o is the input channel index, is the coordinate of the horizontal axis of the convolution kernel in the spatial dimension, is the coordinate of the vertical axis of the convolution kernel in the spatial dimension.
[0011] Preferably, in step S4, model retraining specifically includes: using the same training data set as the pre-trained deep learning model, adopting the same hyperparameters as the pre-trained deep learning model, including optimizer, learning rate and number of iterations, so that the loss of the deep learning model on the training data set is gradually reduced, and some performance lost due to pruning operation is restored.
[0012] Preferably, in step S5, the specific operations include: Step S51: Based on the sensitivity analysis result obtained in step S22, different quantization accuracies are used for modules with different sensitivities to balance reconstruction quality and computational efficiency; Step S52: Before the deep learning model is deployed, a mixed precision post-training static quantization (PTQ) method is used to convert the floating point weights Quantized to fixed-point weights , the quantitative formula is: ; in, For rounding operation, is the quantization scale factor, To quantize the zero point, is the quantization bit width, For cutting operation; Step S53: Use the MinMax strategy to perform quantization calibration, determine the activation value and weight of the deep learning model, and according to the sensitivity analysis results, use a lower bit width quantization for insensitive blocks and a higher bit width quantization for sensitive blocks.
[0013] Therefore, the present invention proposes a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis, which has the following beneficial effects: (1) The present invention proposes a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis. This method introduces end-to-end sensitivity analysis into computational imaging edge reconstruction for the first time. It can customize the model compression strategy according to the characteristics of the edge chip and improve the adaptability of the model in the edge environment.
[0014] (2) The present invention proposes a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis, which adopts an innovative pruning and quantization calculation method to achieve differentiated processing of different sensitivity modules based on the sensitivity analysis results, effectively balance the reconstruction quality and computational efficiency, and significantly improve the edge inference speed.
[0015] (3) The strategy proposed in the present invention for a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis can be applied to a variety of computational imaging tasks and has wide applicability, providing strong support for the promotion of computational imaging technology in practical scenarios.
[0016] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flow chart of a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis provided by the present invention; Figure 2A specific flow chart of step S3 in a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis provided by the present invention; Figure 3 A specific flow chart of step S5 in a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis provided by the present invention; Figure 4 A schematic diagram of a U-Net network structure provided in an embodiment of the present invention; Figure 5 Schematic diagram of comparison between the image reconstruction result of the compressed model provided in the embodiment of the present invention and the image reconstruction results of the input image, the label image and the original model; wherein, Figure 5 (a) in the figure is the label image. Figure 5 (b) in the figure is the input image. Figure 5 (c) in the figure is the image reconstruction result of the original model. Figure 5 (d) in the figure is the image reconstruction result of the compressed model. DETAILED DESCRIPTION
[0018] In order to make the technical solutions, advantages and purposes of the present invention clearer, the technical solutions of the embodiments of the present invention are clearly and completely described below. The described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of this application.
[0019] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.
[0020] like Figure 1 As shown, the present invention provides a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis, which can be applied to image reconstruction deep learning model compression such as infrared single-lens computational imaging system, including the following steps: Step S1: Evaluate the performance of the pre-trained deep learning model at the edge, and replace or remove operators with poor hardware support or excessive time consumption. The specific operations are as follows: Step S11: deploy the pre-trained deep learning model to the edge using the tool chain corresponding to the edge hardware platform deployed on the generation; Step S12: Use corresponding performance analysis tools to analyze the operating efficiency and speed of each operator in the deep learning model network structure; Step S13: Find out the operators with poor operating efficiency and speed, and directly delete them from the network structure or replace them with other operators.
[0021] Step S2: Analyze the pruning sensitivity and quantization sensitivity of the deep learning model on the edge chip. The specific analysis steps are as follows: Step S21: Based on the trained deep learning model, sensitivity analysis is performed with layers or blocks as the minimum unit; assuming that Deep learning models composed of blocks , the edge performance of the deep learning model is recorded as , pruning or quantizing deep learning models The performance of the block is ,definition ,pass , calculate pruning sensitivity and quantitative sensitivity ;in, The deep learning model piece, For deep learning models Degraded performance after block pruning or quantization, ; Step S22: According to the pruning sensitivity , pruning deep learning models at a fixed rate The block gets the pruned deep learning model , fine-tune and deploy on target AI chips to evaluate edge performance, repeat Get the pruning sensitivity of all blocks again; Step S23: According to the quantization sensitivity , for the deep learning model The block is low-bit quantized to obtain the quantized deep learning model ,The quantized sensitivity of all blocks is also obtained after fine-tuning, deployment and evaluation.
[0022] Step S3: According to the pruning sensitivity analysis result, a pruning operation with a high pruning rate is performed on the low-sensitivity module, and a pruning operation with a low pruning rate is performed on the high-sensitivity module. The specific steps of the pruning operation are as follows: Step S31, define The input and output channels of the block are and , No. Block The filter is , pruning rate ;in, is the number of output channels remaining after pruning; Step S32: Pruning sensitivity obtained in step S21 , determine the The target pruning rate of the block is as follows: ; in, For deep learning models The target pruning rate of the block, For the The pruning sensitivity of the block, for The target pruning rate function is is the conversion factor from sensitivity to pruning rate, is a positive integer; Step S33: The number is the importance criterion, which filters are pruned in ascending order. For a specific filter, The norm is defined as ;in, K is the convolution kernel size, is the weight, is a floating point weight, o is the input channel index, is the coordinate of the horizontal axis of the convolution kernel in the spatial dimension, is the coordinate of the vertical axis of the convolution kernel in the spatial dimension.
[0023] Step S4, retraining the deep learning model after sensitivity pruning, specifically including: using the same training data set as the pre-trained deep learning model, adopting the same hyperparameters as the pre-trained deep learning model, including optimizer, learning rate and number of iterations, so that the loss of the deep learning model on the training data set is gradually reduced, and the partial performance lost due to the pruning operation is restored.
[0024] Step S5: According to the quantization sensitivity analysis result, low quantization precision processing is performed on the low-sensitivity module, and high quantization precision operation is performed on the high-sensitivity module. The specific operations include: Step S51: Based on the sensitivity analysis result obtained in step S22, different quantization accuracies are used for modules with different sensitivities to balance reconstruction quality and computational efficiency; Step S52: Using mixed precision post-training static quantization before deep learning model deployment Method, floating point weights Quantized to fixed-point weights , the quantitative formula is: ; in, For rounding operation, is the quantization scale factor, To quantize the zero point, is the quantization bit width, For cutting operation; Step S53: Use the MinMax strategy to perform quantization calibration, determine the activation value and weight of the deep learning model, and according to the sensitivity analysis results, use a lower bit width quantization for insensitive blocks and a higher bit width quantization for sensitive blocks.
[0025] Step S6: deploy the optimized model to the target AI chip, and comprehensively evaluate the edge acceleration effect to determine whether the edge reasoning performance of the deep learning model can meet the actual application requirements. If not, execute step S7; if yes, execute step S8. Step S7, reduce the pruning rate and the ratio of low-bit quantization in hybrid quantization, and repeat steps S2 to S6 until the edge reasoning performance of the deep learning model meets the actual application requirements; this adjustment process is based on the evaluation results of the edge performance of the deep learning model, and by gradually optimizing the pruning rate and quantization ratio, the edge reasoning performance of the deep learning model can meet the actual application requirements; Step S8, deploy the optimized deep learning model to the target AI chip, and comprehensively evaluate the edge acceleration effect to determine whether the edge inference speed of the deep learning model can meet the actual application requirements. If not, execute step S9; if yes, execute step S10; Step S9: Increase the pruning rate and the ratio of low-bit quantization in hybrid quantization, and repeat steps S2 to S8 until the edge inference speed of the deep learning model meets the actual application requirements; this adjustment process is based on the evaluation results of the edge performance of the deep learning model, and by gradually optimizing the pruning rate and quantization ratio, the edge inference speed of the deep learning model can meet the actual application requirements; Step S10: Obtaining a compressed deep learning model can significantly improve the inference time while maintaining similar reconstruction quality. After the above series of processing and optimization steps, a deep learning model that can run efficiently on edge devices and maintain good reconstruction effects is finally obtained. This model is suitable for single-lens computational imaging systems and effectively improves the overall performance of the system.
[0026] Embodiment 1 The deep learning model used in this embodiment is an infrared computational imaging image reconstruction model, and the network structure is a UNet type, including a downsampling module, an intermediate module, and an upsampling module. The specific network structure is as follows: Figure 4As shown. The edge hardware platform selected in this embodiment is the Rockchip RK3588 platform. The data set used for model training is designed based on the characteristics of a single-lens camera. The PSF (Point Spread Function) of a calibrated single-lens camera is used to simulate the degradation process of a clear image. Through convolution operations, a corresponding blurred image is generated from a clear image, and the clear image is used as GT (Ground Truth), and the blurred image is used as the input of the model. The data set covers multiple categories, including 9 scenes such as buildings, vehicles, cities, and humans, with a total of 10,068 images. The data set is divided into training and test sets in a ratio of 9:1 to ensure the effectiveness of model training and evaluation. The specific steps are as follows: 1. Evaluate the performance of pre-trained deep learning models at the edge and remove or replace operators that are poorly supported by hardware, including the following steps: 1.1. First, deploy the pre-trained original U-Net network model directly to the Rockchip RK3588 platform. The deployment process is completed using the rknn-toolkit2 tool chain provided by Rockchip.
[0027] 1.2. Use the rknn.eval_perf() interface in the rknn-toolkit2 toolchain to analyze the inference time of each operator. In this embodiment, under the quantization accuracy of INT8, the specific time consumption of each operator is shown in Table 1.
[0028] Table 1 The time consumption of each operator in the pre-trained U-Net network model ;
[0029] 1.3. The MaxPooling operation in the downsampling module in this embodiment is considered to be too time-consuming, so this embodiment removes all maxpooling layers in the U-Net network model and sets the step size of the previous Conv layer to 2 to achieve the purpose of downsampling. In addition, this embodiment also removes the LeakyReLU layer after the ConvT layer, and the performance does not drop significantly. After these optimizations, the inference time of the U-Net network model is reduced by about 10 milliseconds.
[0030] 2. Use the evaluation indicators PSNR (Peak signal-to-noise ratio) and SSIM (Structure Similarity Index Measure) of the reconstructed image output by the U-Net network model as the indicators for evaluating the performance of the U-Net network model, and analyze the pruning and quantization sensitivity of the U-Net network model. Use three different pruning ratios (25%, 50%, and 75%) to prune each block in the U-Net network model, and then count the performance of the pruned U-Net network model and compare it with the unpruned U-Net network model to obtain the pruning sensitivity of each block. Use INT8 precision to quantize one block in the U-Net network model, and use FP16 precision to quantize other blocks. Statistic the performance of the quantized U-Net network model and compare it with the U-Net network model that uses FP16 precision to obtain the quantization sensitivity of each block.
[0031] 3. If Figure 2 As shown, according to the pruning sensitivity analysis result, a pruning operation with a high pruning rate is performed on the low-sensitivity module, and a pruning operation with a low pruning rate is performed on the high-sensitivity module, which specifically includes the following steps: 3.1. First, according to the obtained pruning sensitivity, different pruning rates are set for different blocks of the U-Net network model, which are 0.5, 0.5, 0.5, 0.75, 0.75, 0.75, 0.5, 0.5, and 0.5 respectively; 3.2. Use the Torch-Pruning library to automatically calculate the weights of all filters that need to be pruned in the U-Net network model Norm.
[0032] 3.3. According to the pruning rate set for each block, use the Torch-Pruning library to remove the corresponding proportion of each block. The filter with a smaller norm is the pruned U-Net network model.
[0033] 4. Retrain the pruned U-Net network model using the same dataset as the pre-training one. The hyperparameter settings of the re-training are kept consistent with those of the pre-training. Specifically, this embodiment uses the Adam optimizer, where , The initial learning rate is set to , the batch size is 2, and the number of iterations is 100. The learning rate remains constant throughout the training process, and the loss function combines loss and a VGG-based perceptual loss. In addition, random detector noise was added to the dataset during training. Gaussian noise with mean 0 and variance 0.006 was added to the natural scene images of the test set.
[0034] 5. If Figure 3 As shown, according to the results of the quantization sensitivity analysis, low quantization precision processing is performed on low-sensitivity modules, and high quantization precision operations are performed on high-sensitivity modules. The specific operations include: 5.1. According to the obtained quantization sensitivity, different quantization bit numbers are set for different blocks of the U-Net network model, namely INT8, INT8, INT8, INT8, INT8, INT8, INT8, INT8, and FP16; 5.2. Use the hybrid_quantization_step1() and hybrid_quantization_step2() interfaces provided by the rknn-toolkit2 tool chain, use the test set as the quantization calibration set, and perform hybrid quantization on the U-Net network model to obtain the final compressed U-Net network model.
[0035] 6. Deploy the optimized U-Net network model to the target AI chip and conduct a comprehensive evaluation of the edge acceleration effect to determine whether the edge performance of the model can meet the actual application requirements. If not, proceed to step 7; if yes, proceed to step 8. If the PSNR of the reconstructed image is greater than 35 and the SSIM is greater than 0.9, it indicates that the quality of the restored image meets the actual application requirements.
[0036] 7. Reduce the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps 2-6 until the edge reasoning performance of the U-Net network model can meet the actual application requirements. This adjustment process is based on the evaluation results of the model's edge performance, and by gradually optimizing the pruning rate and quantization ratio, the model's edge reasoning performance can meet the actual application requirements.
[0037] 8. Deploy the optimized U-Net network model to the target AI chip, and comprehensively evaluate the edge acceleration effect to determine whether the edge inference speed of the model can meet the actual application requirements. If not, execute step 9; if yes, execute step 10. If the inference frame rate of the model is greater than 25FPS, it means that the image reconstruction speed meets the actual application requirements.
[0038] 9. Reduce the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps 2-6 until the edge reasoning performance of the U-Net network model can meet the actual application requirements. This adjustment process is based on the evaluation results of the model's edge performance, and by gradually optimizing the pruning rate and quantization ratio, the model's edge reasoning speed can meet the actual application requirements.
[0039] 10. The compressed U-Net network model can significantly improve the inference time while maintaining similar reconstruction quality. The image reconstruction performance and inference frame rate results of the compressed U-Net network model are shown in Table 2.
[0040] Table 2 Image reconstruction performance and inference frame rate results of the compressed U-Net network model ;
[0041] like Figure 5 As shown, the image reconstruction results of the compressed U-Net network model and the image reconstruction results of the input image, label image and original model show that the present invention can greatly improve the inference speed on the edge platform while maintaining the model performance.
[0042] Therefore, the present invention provides a single-chip computational imaging edge reconstruction method based on end-to-end sensitivity analysis, which can customize the model compression strategy according to the characteristics of the edge chip, realize differentiated processing of different sensitivity modules, effectively balance the reconstruction quality and computational efficiency, and significantly improve the edge reasoning speed.
[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis, characterized in that: The specific steps are as follows: Step S1: Evaluate the performance of the pre-trained deep learning model at the edge, and replace or remove operators with poor hardware support or excessive time consumption; Step S2: analyzing the pruning sensitivity and quantization sensitivity of the deep learning model on the edge chip; Step S3: According to the pruning sensitivity analysis result, a pruning operation with a high pruning rate is performed on the low-sensitivity module, and a pruning operation with a low pruning rate is performed on the high-sensitivity module; Step S4, retraining the deep learning model after sensitivity pruning; Step S5: According to the quantization sensitivity analysis result, low quantization precision processing is performed on the low-sensitivity module, and high quantization precision operation is performed on the high-sensitivity module; Step S6: determine whether the edge reasoning performance of the deep learning model can meet the actual application requirements. If not, execute step S7; if yes, execute step S8; Step S7, reduce the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps S2 to S6 until the edge reasoning performance of the deep learning model meets the actual application requirements; Step S8, determining whether the edge inference speed of the deep learning model can meet the actual application requirements, if not, executing step S9, if yes, executing step S10; Step S9, increase the pruning rate and the proportion of low-bit quantization in hybrid quantization, and repeat steps S2 to S8 until the edge inference speed of the deep learning model meets the actual application requirements; Step S10: Obtain a compressed deep learning model.
2. The method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis according to claim 1, characterized in that: The specific operations in step S1 are as follows: Step S11: deploy the pre-trained deep learning model to the edge using the tool chain corresponding to the edge hardware platform deployed on the generation; Step S12: Use corresponding performance analysis tools to analyze the operating efficiency and speed of each operator in the deep learning model network structure; Step S13: Find out the operators with poor operating efficiency and speed, and directly delete them from the network structure or replace them with other operators.
3. The method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis according to claim 1, characterized in that: In step S2, the specific analysis steps of pruning sensitivity and quantization sensitivity are as follows: Step S21: Based on the trained deep learning model, sensitivity analysis is performed with layers or blocks as the minimum unit; assuming that Deep learning models composed of blocks , the edge performance of the deep learning model is recorded as , pruning or quantizing deep learning models The performance of the block is ,definition ,pass , calculate pruning sensitivity and quantitative sensitivity ;in, The deep learning model piece, For deep learning models Degraded performance after block pruning or quantization, ; Step S22: According to the pruning sensitivity , pruning deep learning models at a fixed rate The block gets the pruned deep learning model , fine-tune and deploy on target AI chips to evaluate edge performance, repeat Get the pruning sensitivity of all blocks again; Step S23: According to the quantization sensitivity , for the deep learning model The block is low-bit quantized to obtain the quantized deep learning model ,The quantized sensitivity of all blocks is also obtained after fine-tuning, deployment and evaluation.
4. The method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis according to claim 3, characterized in that: In step S3, the specific steps of the pruning operation are as follows: Step S31, define The input and output channels of the block are and , No. Block The filter is , pruning rate ;in, is the number of output channels remaining after pruning; Step S32: Pruning sensitivity obtained in step S21 , determine the The target pruning rate of the block is as follows: ; in, For deep learning models The target pruning rate of the block, For the The pruning sensitivity of the block, for The target pruning rate function is is the conversion factor from sensitivity to pruning rate, is a positive integer; Step S33: The number is the importance criterion, which filters are pruned in ascending order. For a specific filter, The norm is defined as ;in, K is the convolution kernel size, is the weight, is a floating point weight, o is the input channel index, is the coordinate of the horizontal axis of the convolution kernel in the spatial dimension, is the coordinate of the vertical axis of the convolution kernel in the spatial dimension.
5. The method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis according to claim 1, characterized in that: In step S4, model retraining specifically includes: using the same training data set as the pre-trained deep learning model, adopting the same hyperparameters as the pre-trained deep learning model, including optimizer, learning rate and number of iterations, so that the loss of the deep learning model on the training data set is gradually reduced, and some performance lost due to pruning operation is restored.
6. The method for edge reconstruction of single-chip computational imaging based on end-to-end sensitivity analysis according to claim 3, characterized in that: In step S5, the specific operations include: Step S51: Based on the sensitivity analysis result obtained in step S22, different quantization accuracies are used for modules with different sensitivities to balance reconstruction quality and computational efficiency; Step S52: Using mixed precision post-training static quantization before deep learning model deployment Method, floating point weights Quantized to fixed-point weights , the quantitative formula is: ; in, For rounding operation, is the quantization scale factor, To quantize the zero point, is the quantization bit width, For cutting operation; Step S53: Use the MinMax strategy to perform quantization calibration, determine the activation value and weight of the deep learning model, and according to the sensitivity analysis results, use a lower bit width quantization for insensitive blocks and a higher bit width quantization for sensitive blocks.
Citation Information
Patent Citations
Global rank perception neural network model compression method based on filter feature map
CN114037844A
Deep learning network model optimization method based on parameter quantization
CN116524173A
Neural network iteration pruning method and device based on feature perception and electronic equipment
CN118536573A