U-Net model optimization method based on dynamic interpretable pruning

Through dynamic interpretable pruning technology, the U-Net model is optimized, which solves the problems of high redundancy, inefficiency and insufficient real-time performance of existing models in medical image segmentation tasks, and achieves reduction of parameter quantity and improvement of inference speed, meeting clinical real-time requirements and maintaining high accuracy.

CN120163189APending Publication Date: 2025-06-17HEFEI CHART MEDICAL INSTR CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510247445.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing U-Net model has problems such as high redundancy, inefficient encoder-decoder structure, poor hardware adaptability and insufficient real-time performance in medical image segmentation tasks.

Method used

U-Net model optimization method based on dynamic interpretable pruning is adopted. Through hierarchical differentiated pruning strategies, multi-dimensional evaluation model and dynamic threshold adjustment, combined with the dynamic correlation between hardware characteristics and pruning granularity, the structural improvement and parameter reduction of the model are achieved.

Benefits of technology

It effectively reduces the amount of model parameters, improves the inference speed of edge devices, meets the requirements of clinical real-time, and maintains high accuracy, which has practical application value in medical imaging segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163189A_ABST
    Figure CN120163189A_ABST
Patent Text Reader

Abstract

The invention provides a U-Net model optimization method based on dynamic interpretable pruning, and the method comprises the steps: employing a hierarchical differential pruning strategy, employing structured channel pruning at an encoder, and employing unstructured intra-kernel pruning at a decoder; introducing a multi-dimensional evaluation model, and constructing a ternary evaluation model comprising a weight absolute value, gradient saliency and a feature map entropy value; dynamically associating the hardware characteristics with the pruning granularity, and presetting the corresponding pruning granularity according to the computing power characteristics of the target deployment platform; an existing medical image data set is adopted as an input feature map, and the improved U-Net model is trained through three stages of processes of pre-training, dynamic pruning training and precision recovery in sequence. According to the method, the parameter quantity is effectively reduced, and the reasoning speed of edge equipment is increased, so that the clinical real-time requirement is met; meanwhile, high precision is kept, and the method has practical application value in medical image segmentation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of model pruning optimization, and particularly relates to a method for optimizing a U-Net model based on dynamic interpretable pruning. Background Art

[0002] The U-Net model is a deep learning model widely used in tasks such as medical image segmentation. It has a unique symmetric encoder-decoder structure and skip connections structure, and is widely applicable to scenarios such as clinical real-time auxiliary diagnosis.

[0003] In the prior art, the U-Net model has the following defects in specific use: high redundancy, with a large number of inefficient convolutional kernels in the encoder-decoder structure; poor hardware adaptability, and the fixed model structure cannot adapt to different computing power platforms (such as GPUs and edge devices); insufficient real-time performance, and large-scale parameters result in the time-consuming of single-frame medical image segmentation exceeding the clinical real-time requirement (>500 ms). Summary of the Invention

[0004] The purpose of the present invention is to provide a coal pyrolysis high-temperature coal gas oil and dust separation device and method to solve the above-mentioned deficiencies of the prior art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A method for optimizing a U-Net model based on dynamic interpretable pruning, the driving method specifically includes the following steps:

[0007] S1. Improve the structure based on the U-Net model to obtain an improved U-Net model;

[0008] S11. Adopt a hierarchical differential pruning strategy, use structured channel pruning in the encoder and unstructured in-kernel pruning in the decoder;

[0009] S12. Introduce a multi-dimensional evaluation model, and construct a ternary evaluation model including the absolute value of weights, gradient significance, and feature map entropy value;

[0010] S2. Dynamically associate the hardware characteristics with the pruning granularity, and preset the corresponding pruning granularity according to the computing power characteristics of the target deployment platform;

[0011] S3. Use the existing medical image dataset as the input feature map, and train the improved U-Net model through three-stage processes of pre-training, dynamic pruning training, and accuracy recovery.

[0012] Furthermore, the ternary evaluation model is represented by the following formulas (1) and (2):

[0013]

[0014] H(F k ) = -∑ x,y p(x,y) log p(x,y) (2);

[0015] In formulas (1) and (2): α, β, and γ represent learnable adjustment coefficients used to optimize the model performance; w ijk represents the weight; Score(w ijk ) represents the importance score of the weight w ijk ; |w ijk | represents the absolute value of the weight w ijk , and usually its value is proportional to the contribution to the model output; represents the absolute value of the gradient, that is, the gradient of the loss function with respect to the weight w ijk , and the value of the weight w ijk is proportional to the degree of influence on the model training; F k represents the input feature map; H(F k ) represents the entropy of the input feature map, which is used to measure the complexity or information content of the feature map; p(x,y) represents the probability of the result (x,y).

[0016] Furthermore, the dynamic pruning training specifically includes the following process: comprehensively considering the absolute value of the weight, the gradient significance, and the entropy value of the feature map to obtain the importance score of each weight; adjusting the pruning threshold according to the training progress and the importance score of each weight, and gradually increasing and approaching the target sparsity; zeroing the weights with scores lower than the threshold while retaining the high-scoring weights.

[0017] Furthermore, the adjustment of the pruning threshold according to the training progress and the importance score of each weight is dynamically determined through the following formula (3) and the threshold is adjusted through the following formula (4):

[0018]

[0019] In formulas (3) and (4): Prune(w ijk ) represents the dynamic decision of pruning, which is used to evaluate the importance of the weight. When it is equal to 1, it means the weight is unimportant and needs to be removed; when it is equal to 0, it means the weight is important and needs to be retained; τ(t) represents the pruning threshold depending on the parameter t, and the parameter t can be the training step or epoch; τ mim and τ max respectively represent the minimum and maximum values of τ(t); T max represents the maximum value of the parameter t; specifically, when t = 0, τ(0) = τ min ; when t = T max , τ(T max) = τ max 。

[0020] As can be seen from the above technical solutions, compared with the prior art, the present invention has the following technical advantages:

[0021] 1. The present invention proposes a ternary evaluation model of weight absolute value, gradient significance and feature map entropy value, which synthesizes weight importance, training dynamics and feature layer information entropy, incorporates the uncertainty of the model output into the pruning decision, and enhances the interpretability of the pruning process; the dynamic threshold adjustment function τ(t) can adaptively change with the training progress, avoiding irreversible performance loss caused by excessive pruning in the early stage;

[0022] 2. The present invention adopts a hierarchical differential pruning strategy, using structured channel pruning (directly deleting the entire channel) in the encoder and unstructured in-kernel pruning (retaining the channel but sparsifying the convolutional kernel) in the decoder, which fits the characteristics of the encoding and decoding architecture of U-Net, enabling the encoder to focus on feature extraction, and structured pruning can maximize the computational efficiency, while the decoder needs to retain detailed information, and unstructured pruning can effectively avoid destroying the spatial continuity;

[0023] 3. The present invention dynamically associates the hardware characteristics with the pruning granularity through predefined device configurations, realizing seamless adaptation of the same model to different computing power platforms, solving the pain point of repeated optimization required for cross-platform deployment of traditional pruning models, and having industrial practicability;

[0024] 4. The present invention adopts a progressive pruning training process, proposing a three-stage process of pre-training, dynamic pruning, and accuracy recovery, allowing the model to continuously learn the weight distribution during the pruning process, and adjusting the pruning threshold to gradually approach the target sparsity, avoiding a decrease in accuracy caused by excessive pruning at one time;

[0025] 5. The optimization method of the U-Net model based on dynamic interpretable pruning of the present invention effectively reduces the number of parameters, improves the inference speed of edge devices to meet the clinical real-time requirement (<50 ms); at the same time, it maintains a high accuracy and has practical application value in medical image segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the step flow of the optimization method of the U-Net model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The following is a detailed description of a preferred embodiment of the present invention with reference to the accompanying drawings.

[0028] As Figure 1 shown, the optimization method of the U-Net model based on dynamic interpretable pruning specifically includes the following steps:

[0029] S1. Improve the structure based on the U-Net model to obtain the improved U-Net model;

[0030] S11. Adopt a hierarchical differential pruning strategy. Use structured channel pruning (directly deleting the entire channel) in the encoder and unstructured in-kernel pruning (retaining the channel but sparsifying the convolutional kernel) in the decoder;

[0031] Through this pruning strategy, the encoder can focus on feature extraction, and structured pruning can maximize the computing efficiency. While the decoder needs to retain detailed information, and unstructured pruning can effectively avoid destroying the spatial continuity;

[0032] S12. Introduce a multi-dimensional evaluation model and construct a ternary evaluation model including the absolute value of weights, gradient significance, and feature map entropy value;

[0033] Specifically, the ternary evaluation model is represented by the following formulas (1) and (2):

[0034]

[0035] H(F k )=-∑ x,y p(x,y)logp(x,y) (2);

[0036] In formulas (1) and (2): α, β, γ represent learnable adjustment coefficients for optimizing the model performance; w ijk represents the weight; Score(w ijk ) represents the importance score of the weight w ijk ; |w ijk | represents the absolute value of the weight w ijk , and usually its value is proportional to the contribution to the model output; represents the absolute value of the gradient, that is, the gradient of the loss function with respect to the weight w ijk , and the value of the weight w ijk is proportional to the influence on the model training; F k represents the input feature map; H(F k ) represents the entropy of the input feature map, which is used to measure the complexity or information content of the feature map; p(x,y) represents the probability of the result (x,y).

[0037] S2. Dynamically associate the hardware characteristics with the pruning granularity and preset the corresponding pruning granularity according to the computing power characteristics of the target deployment platform;

[0038] Dynamically associate hardware characteristics (such as memory bandwidth and parallel computing units) with pruning granularity (block size, sparsity) to achieve seamless adaptation of the same model to different computing power platforms, solving the pain point of repeated optimization required for cross-platform deployment of traditional pruning models.

[0039] S3. Use the existing medical image dataset as the input feature map, and train the improved U-Net model through three stages: pre-training, dynamic pruning training, and accuracy recovery.

[0040] The pre-training in this preferred embodiment is to use large-scale image classification datasets such as ImageNet to pre-train the encoder part of the UNet model; ImageNet contains millions of images and thousands of categories, and the UNet model can learn rich image features on this dataset.

[0041] The accuracy recovery in this preferred embodiment adopts the fine-tuning strategy. Since the weight distribution and network structure of the UNet model have changed after dynamic pruning, fine-tuning is required to re-adapt to the target task and compensate for the accuracy loss caused by pruning:

[0042] Learning rate adjustment: Adopt a lower learning rate. After pruning, the UNet model is already near a relatively good parameter space. Therefore, a lower learning rate than the initial training is usually used during fine-tuning. An overly high learning rate may cause the UNet model to jump out of the optimal region and it is difficult to recover the accuracy.

[0043] Learning rate decay strategy, such as cosine annealing, step decay, exponential decay and other learning rate decay strategies, can use a slightly higher learning rate at the beginning of fine-tuning to quickly adapt, and gradually decrease the learning rate in the later stage to finely adjust the weights, helping the UNet model converge to a better point.

[0044] Adopt a longer fine-tuning period. Dynamic pruning may change the loss landscape of the UNet model (the visualization representation of the loss function in the parameter space). Therefore, a longer training period is required for the UNet model to fully explore the new parameter space and find a better solution. At the same time, monitor the performance metrics on the validation set, such as Dice coefficient, IoU, etc. When the metrics tend to be stable or no longer improve, the fine-tuning can be stopped.

[0045] In the fine-tuning stage, continuing to use appropriate data augmentation strategies can improve the generalization ability of the UNet model, help the UNet model better adapt to the pruned structure, and improve the accuracy; stronger data augmentation strategies such as CutMix and MixUp can be tried to improve the robustness of the UNet model.

[0046] The dynamic pruning training described in this preferred embodiment specifically includes the following process: combining the absolute value of the weight, gradient significance, and feature map entropy value to obtain the importance score of each weight; adjusting the pruning threshold according to the training progress and the importance score of each weight, gradually increasing and approaching the target sparsity; zeroing the weights with scores lower than the threshold while retaining the high-scoring weights. Specifically, the adjustment of the pruning threshold according to the training progress and the importance score of each weight is dynamically determined by the following formula (3) and the threshold is adjusted by the following formula (4):

[0047]

[0048] In formulas (3) and (4): Prune(w ijk ) represents the dynamic decision of pruning, which is used to evaluate the importance of the weight. When it is equal to 1, it means the weight is unimportant and needs to be removed. When it is equal to 0, it means the weight is important and needs to be retained; τ(t) represents the pruning threshold depending on the parameter t, and the parameter t can be the training step or epoch; τ min and τ max respectively represent the minimum and maximum values of τ(t); T max represents the maximum value of the parameter t; specifically, when t = 0, τ(0) = τ min ; when t = T max , τ(T max ) = τ max .

[0049] In specific use, the pruning threshold τ(t) changes dynamically with the training epoch. An epoch refers to the number of times the entire training dataset is completely traversed by the neural network; in each epoch, the model will use different samples in the dataset again and again for training to update the weights of the model.

[0050] This step incorporates the uncertainty of the model output into the pruning decision, enhancing the interpretability of the pruning process; the dynamic threshold adjustment function τ(t) can adaptively change with the training progress, avoiding irreversible performance losses caused by excessive pruning in the early stage.

[0051] Effect evaluation: In this preferred embodiment, a comparative experiment is conducted on the traditional U-Net model and the improved model respectively from the number of parameters and the cross-layer pruning effect. The experimental data is shown in Table 1-2 below.

[0052] Index Traditional U-Net Improved model Reduction amplitude Number of parameters (M) 31.2 9.8 68.6% Inference latency (ms EINVAL) 42.3 15.7 62.9% CPU memory occupancy (MB) 680 210 69.1%

[0053] Table 1 Comparison Table of Parameter Quantities

[0054]

[0055]

[0056] Table 2 Comparison Table of Cross-Layer Pruning Effects

[0057] From the above experimental data, it can be seen that the optimization method of the U-Net model based on dynamic interpretable pruning of the present invention effectively reduces the parameter quantity and improves the inference speed of edge devices to meet the clinical real-time requirement (<50 ms); at the same time, it maintains a high accuracy and has practical application value in medical image segmentation tasks.

[0058] The above-described embodiments are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A U-Net model optimization method based on dynamic interpretable pruning, characterized in that: The driving method specifically comprises the following steps: S1. Improve the structure based on the U-Net model to obtain an improved U-Net model; S11, adopt hierarchical differentiated pruning strategy, use structured channel pruning in encoder and unstructured kernel pruning in decoder; S12, introduce a multi-dimensional evaluation model and construct a ternary evaluation model including the absolute value of weight, gradient significance and feature map entropy value; S2. Dynamically associate hardware characteristics with pruning granularity, and preset corresponding pruning granularity according to the computing power characteristics of the target deployment platform; S3. Use the existing medical imaging dataset as the input feature map, and train the improved U-Net model through three stages of pre-training, dynamic pruning training, and accuracy recovery.

2. A U-Net model optimization method based on dynamic interpretable pruning according to claim 1, characterized in that: The ternary evaluation model is expressed by the following formulas (1) and (2): H(F k )=-∑ x,y p(x,y)logp(x,y) (2); In formulas (1) and (2), α, β, and γ represent learnable adjustment coefficients used to optimize model performance; w ijk Represents weight; Score(w ijk ) represents the weight w ijk Importance score of ijk | represents the weight w ijk The absolute value of , which is usually proportional to its contribution to the model output; Represents the absolute value of the gradient, that is, the loss function For weight w ijk The gradient of weight w ijk The value of F is proportional to the degree of influence on model training; k represents the input feature map; H(F k ) represents the entropy of the input feature map, which is used to measure the complexity or information content of the feature map; p(x,y) represents the probability of the result (x,y).

3. A U-Net model optimization method based on dynamic interpretable pruning according to claim 1, characterized in that: The dynamic pruning training specifically includes the following processes: The importance score of each weight is obtained by combining the absolute value of the weight, the gradient significance and the entropy value of the feature map. The pruning threshold is adjusted according to the training progress and the importance score of each weight to gradually increase and approach the target sparsity. Zero out weights with scores below a threshold, while retaining high-scoring weights.

4. A U-Net model optimization method based on dynamic interpretable pruning according to claim 3, characterized in that: The pruning threshold is adjusted according to the training progress and the importance score of each weight by dynamically deciding through the following formula (3), and the threshold is adjusted through the following formula (4): In formulas (3) and (4): Prune(w ijk ) represents the dynamic decision of pruning, which is used to evaluate the importance of weights. When it is equal to 1, it means that the weight is not important and needs to be removed. When it is equal to 0, it means that the weight is important and needs to be retained. τ9t) represents the pruning threshold that depends on parameter t. Parameter t can be a training step or period. τ min and τ mae Respectively represent the minimum and maximum values ​​of τ(t); T max represents the maximum value of parameter t; specifically, when t = 0, τ(0) = τ min ; When t = T max When τ(T max )=τ max .

Citation Information

Cited By

  • Flight control end intelligent algorithm deployment method based on adaptive pruning

    CN120725086A

  • Model pruning method for heterogeneous cloud edge-end cooperative system

    CN120806024A

  • A model pruning method for heterogeneous cloud-edge-device collaborative systems

    CN120806024B

  • Lithium battery collector plate bonding wire segmentation model generation method, system and device and medium

    CN121190480A

  • Lithium battery current collector disc welding wire segmentation model generation method, system, device and medium

    CN121190480B