Disease grading system based on image uncertainty sensing distillation
Through a multi-expert collaborative distillation framework, combining shallow and compact feature alignment and uncertainty-sensing decoupled distillation, the data imbalance and feature coupling problems in medical image grading are solved, improving the accuracy and robustness of disease grading, and supporting more accurate clinical diagnosis.
Patent Information
- Application Number
- CN202510417709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-15
AI Technical Summary
The existing medical image grading system has problems with data distribution imbalance and domain offset in the grading tasks of diabetic retinopathy and prostate cancer, resulting in insufficient feature decoupling ability and low knowledge transfer efficiency. The traditional distillation method fails to effectively deal with the coupling between lesion areas and background noise and the uncertainty of fuzzy areas in medical images, affecting the grading accuracy.
The multi-expert collaborative distillation framework is adopted, combining shallow feature alignment, compact feature alignment and uncertainty-aware decoupling distillation modules, and dynamically adjust the knowledge transfer weights through multi-scale low-pass filtering and high-dimensional spherical mapping to improve the robustness and accuracy of the disease image grading model.
It significantly improves the generalization ability and grading performance of the disease image grading model, reduces missed detection rates and errors, and provides more reliable clinical decision support.
Smart Images

Figure CN120495719A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image processing and artificial intelligence technology, and in particular relates to a method and system for automatic classification of disease images based on image uncertainty perception distillation. Background Art
[0002] Automatic disease image grading is a core breakthrough in medical AI-enabled clinical decision-making. By quantitatively analyzing morphological differences in pathology / imaging data, it provides an objective basis for disease staging and treatment planning. Compared to subjective assessments that rely on physician experience, deep learning-based automatic grading systems can overcome the limits of human visual perception, capturing micron-level structural heterogeneity and significantly improving the detection rate of early lesions. However, the inherent inter-class similarities and intra-class differences in medical images pose a significant challenge to the model's feature decoupling capabilities.
[0003] In current medical image analysis scenarios, diabetic retinopathy (DR) and prostate cancer grading are two core scenarios with high clinical urgency. DR, a major complication of diabetes, has a grading accuracy that directly impacts the timing of laser photocoagulation treatment (intervention at the macular edema stage can reduce the risk of blindness by 70%). Prostate cancer is the leading malignant tumor in men, and histological grading (such as the ISUP grade) determines the suitability for radical resection (grading errors lead to overtreatment in up to 18%). Existing systems such as DeepDR utilize convolutional neural networks to localize DR lesions, but they suffer from a miss detection rate exceeding 30% for small sample sizes of early-stage lesions (e.g., the non-proliferative stage, which only accounts for 12% of the dataset). The prostate cancer grading model Morpho-Grader, while utilizing multi-scale feature fusion, suffers a 22.6% drop in cross-center generalization performance due to differences in tissue staining (H&E staining color difference coefficient exceeding 15%). The core root of these issues lies in the combined effects of data imbalance and domain shift, necessitating the development of innovative frameworks that simultaneously address feature decoupling, robust knowledge transfer, and consistent predictions.
[0004] To address the inefficiency of knowledge transfer, existing knowledge distillation methods typically rely on a single expert model for knowledge transfer. However, medical image data exhibits significant domain differences, such as inconsistent imaging equipment and labeling standards across different medical institutions. This makes it difficult for a single expert model to capture the feature distribution of multi-source heterogeneous data. In particular, in diabetic retinopathy and prostate cancer grading tasks, data imbalance issues, such as a scarcity of early-stage lesion samples and large intra-class variance due to staining differences, prevent a single model from fully capturing cross-domain feature correlations. Traditional methods directly use the output probabilities of the expert model as soft labels to supervise the student model, but fail to consider the heterogeneity and complementary knowledge between expert models, resulting in inefficient knowledge transfer.
[0005] However, multi-expert collaborative distillation still needs to address the underlying contradictions of feature space coupling. Traditional feature alignment methods use global matching strategies (such as L2 distance constraints), but lesion regions and background noise in medical images are strongly coupled in feature space. For example, vascular structures and camera artifacts in DR fundus images are highly intertwined in shallow feature layers, while prostate cancer glandular morphology and staining noise are difficult to distinguish in deep semantic space.
[0006] Furthermore, blurred regions and data imbalance in medical images lead to uncertainty in expert predictions. Traditional distillation methods treat expert predictions as deterministic supervisory signals. However, blurred regions and severe data imbalance in medical images can cause expert models to produce highly uncertain predictions. Directly forcing the student model to fit such noisy signals can exacerbate model bias. Summary of the Invention
[0007] In view of the above, the purpose of the present invention is to provide a disease grading system based on image uncertainty perception distillation, so as to effectively solve the problems existing in the existing distillation methods in disease image grading, significantly improve the prediction performance, and provide more powerful decision support for disease treatment.
[0008] To achieve the above-mentioned object of the invention, an embodiment provides a disease grading system based on image uncertainty perception distillation, comprising:
[0009] A data acquisition unit, which is used to acquire disease image data and perform preprocessing;
[0010] a disease grading unit, configured to predict the disease grading of the pre-processed disease image using a disease image grading model;
[0011] The disease image grading model is constructed in the following way: based on the student model, multiple expert models, shallow feature alignment module, compact feature alignment module, and uncertainty-aware decoupling distillation module are introduced to construct a multi-expert knowledge distillation framework. The student model and multiple expert models are used to extract shallow and task-independent structural features and deep and task-related semantic features in disease images and perform disease grading task predictions. The shallow feature alignment module is used to align the structural features extracted by the expert model and the student model. The compact feature alignment module is used to align the semantic features extracted by the expert model and the student model. The uncertainty-aware decoupling distillation module is used to dynamically adjust the knowledge transfer weight according to the uncertainty of the expert model to realize dynamic knowledge distillation from the expert model to the student model, and then combine the disease grading task to perform uncertainty-aware distillation learning. The learned student model is used as the disease image grading model.
[0012] Preferably, in the shallow feature alignment module, the alignment operation on the structural features extracted by the expert model and the student model includes:
[0013] A multi-scale low-pass filter is used to process the structural features of the expert model to retain the structural features of the disease image. At the same time, a learnable low-pass filter is used to process the structural features of the student model to enable the student model to learn and generalize the structural features of the expert model.
[0014] Preferably, according to the structural characteristics of the expert model, average pooling is used as a low-pass filter, and multi-scale filters are constructed by adjusting different kernel sizes and strides to adapt to different cutoff frequencies. For the mth group of filters, the projection characteristics of the expert model in the frequency domain are It is expressed as follows through a multi-scale low-pass filter (msLF):
[0015]
[0016] in, Indicates that the kernel size is k m ×k m The average pooling function of , Φ(·) represents the bilinear interpolation operation;
[0017] For the shallow structural features F of the student model S The designed learnable low-pass filter consists of multi-scale low-pass filtering, convolution downsampling module and depth-wise separable convolution. The projection features of the student model in the frequency domain Expressed as:
[0018]
[0019] Among them, DownSample s×s Represents the convolution downsampling module, Concat represents the feature concatenation operation, and Conv represents the convolution operation.
[0020] Preferably, for the shallow feature alignment module, the projection features based on the expert model and the projection features of the student model Total feature alignment loss By maximum mean difference loss and reconstruction losses Composition, expressed as It is used to guide the student model to align with the expert model in the feature space, while ensuring that the features of the expert model remain unchanged during the distillation process;
[0021] The maximum mean difference loss Measuring the structural characteristics of the student model Projected features of each expert model Distribution differences in feature space;
[0022] Reconstruction losses Measures the change in the expert model before and after feature alignment.
[0023] Preferably, in the compact feature alignment module, the alignment operation on the semantic features extracted by the expert model and the student model includes:
[0024] The semantic features of the expert model and the student model are projected into a compact high-dimensional spherical space Z. In the high-dimensional spherical space Z, the student model can learn task-related semantic information from different expert models through spatial domain feature alignment. The corresponding total feature alignment loss is Including maximum mean difference loss and reconstruction losses Calculate and express it as:
[0025] Preferably, in the uncertainty-aware decoupling distillation module, the knowledge transfer weight is dynamically adjusted according to the uncertainty of the expert model, including:
[0026] Uncertainty is calculated by measuring the deviation between the graded probability distribution output by the expert model and the ideal probability distribution. Based on the uncertainty, the student model can dynamically adjust the weight of knowledge transfer, that is, for fuzzy areas where the expert model predicts high uncertainty, the supervision weight of the area will be increased, while for task-irrelevant areas where the expert model predicts low uncertainty, precise alignment will be maintained.
[0027] Preferably, the uncertainty is expressed by the uncertainty coefficient Indicates that its construction process is:
[0028] For the expert model T t And the logical probability prediction output of the student model S and C, H, W represent the length, height, and width of the image. In the multi-scale w∈W={1,2,4,…,w max}, and perform spatial partitioning P(w,w) on each scale w, for each partition unit Z(m,n), where n∈N w ={1,4,16,…,w 2}, the calculation formula of the cumulative logical probability prediction value is as follows:
[0029]
[0030] in, and L S (j,k) represents the expert model T t The logical probability prediction value of the (j, k) position in the partition unit Z(m,n) corresponding to the student model S, ψT t(w,n) and ψS(w,n) represent the expert model T t The cumulative logical probability of the partition unit Z(m,n) corresponding to the student model S;
[0031] Design uncertainty coefficient based on prediction confidence of expert model The calculation formula is: Among them, σ represents the softmax function, max represents the maximum value function, and the coefficient uncertainty The uncertainty of a forecast is quantified by measuring the deviation from an ideal probability distribution.
[0032] Preferably, the uncertainty-based student model is able to dynamically adjust the weight of knowledge transfer by defining the uncertainty-aware decoupling distillation loss To achieve this, it is expressed as:
[0033]
[0034] Among them, the loss component is defined as: The item is used to enhance the blur area supervision, and Items are used to keep task-independent areas The exact logit alignment of Represents the square of the L2 norm.
[0035] Preferably, the total training loss function used in the uncertainty-aware distillation learning of the entire multi-expert knowledge distillation framework includes the classification loss of the image classification task Total feature alignment loss of the shallow feature alignment module Total feature alignment loss of the compact feature alignment module and uncertainty-aware decoupling distillation loss Expressed as:
[0036]
[0037] Among them, α and β are the balance coefficients between various losses.
[0038] To achieve the above-mentioned object of the invention, an embodiment further provides a computing device including a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, the device is used to implement a disease grading method based on image uncertainty perception distillation. The method uses the above-mentioned disease grading system based on image uncertainty perception distillation and includes the following steps:
[0039] Acquire disease image data through a data acquisition unit and perform preprocessing;
[0040] The disease grading unit adopts the disease image grading model to predict the disease grading of the preprocessed disease image.
[0041] To achieve the above-mentioned object of the invention, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method is used to implement a disease grading method based on image uncertainty perception distillation. The method adopts the above-mentioned disease grading system based on image uncertainty perception distillation, and includes the following steps:
[0042] Acquire disease image data through a data acquisition unit and perform preprocessing;
[0043] The disease grading unit adopts the disease image grading model to predict the disease grading of the preprocessed disease image.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) In feature decoupling learning, the present invention proposes two mechanisms, shallow feature alignment (SFA) and compact feature alignment (CFA), which effectively decouple the structural information and semantic information of the image through multi-scale low-pass filtering and spherical space mapping, thereby improving the quality and generalization ability of disease image representation.
[0046] (2) In uncertainty-aware decoupling distillation (UDD), the present invention automatically detects the uncertainty of the expert model caused by factors such as class imbalance, dynamically adjusts the knowledge transfer weight, reduces bias propagation, and ensures that the knowledge transfer process is more robust and reliable, significantly outperforming existing multi-expert knowledge distillation methods.
[0047] (3) The disease image classification model based on (1) and (2) learned through uncertainty-aware distillation can limit the improvement of image classification performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0049] Figure 1 is a schematic structural diagram of a disease grading system based on image uncertainty perception distillation provided by an embodiment;
[0050] Figure 2 Schematic diagram of the structure and flow of the disease image classification model provided in the embodiment;
[0051] Figure 3 This is a flowchart of a disease grading method based on image uncertainty perception distillation provided in an embodiment. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0053] The inventive concept of the present invention is to provide a disease grading system based on image uncertainty perception distillation. When constructing a disease image grading model, in order to solve the technical problem of low knowledge transfer efficiency due to failure to consider the heterogeneity and complementary knowledge between expert models, the present invention breaks through the limitation of a single expert and proposes a multi-expert collaborative distillation framework. By integrating the complementary features of multiple heterogeneous expert models, a dynamic knowledge fusion mechanism is constructed, and the differences between experts are used to enhance the robustness of the student model to data distribution deviation.
[0054] To address the technical problem of difficulty in regional segmentation due to the strong coupling between lesion areas and background noise in the feature space of medical images, this paper innovatively designs a hierarchical decoupling alignment mechanism, dividing the feature alignment process into two parts: shallow feature alignment (SFA) and compact feature alignment (CFA). Shallow feature alignment refers to separating high-frequency details (such as microaneurysm edges) from low-frequency structures (such as the overall retinal vascular topology) in the frequency domain, and retaining task-related morphological features through bandpass filtering; compact feature alignment refers to constructing an expert-student spherical mapping in the latent space, using contrastive learning to constrain the feature distribution consistency of semantically similar samples and eliminate inter-domain distribution offsets. This method breaks through the coarse-grained limitations of traditional alignment strategies and achieves refined knowledge transfer across model feature spaces.
[0055] Aiming to address the technical issue that blurred regions and severe data imbalance in medical images can cause expert models to produce highly uncertain predictions, and that directly forcing the student model to fit such noisy signals can exacerbate model bias, this paper proposes dynamic uncertainty decoupling distillation (UDD). This method derives the confidence distribution of expert predictions based on evidence-based learning theory and identifies highly ambiguous regions. Strict logical value alignment is applied to high-confidence regions to strengthen classification boundary learning, while feature reconstruction constraints are applied to low-confidence regions to mitigate interference from erroneous signals. This method dynamically adjusts knowledge transfer weights based on uncertainty, ensuring the robustness and reliability of distillation.
[0056] Based on the above invention concept, Figure 1As shown, the disease grading system 10 based on image uncertainty perception distillation provided by the embodiment includes a data acquisition unit 11 and a disease grading unit 12.
[0057] In the embodiment, the data acquisition unit 11 is used to acquire disease image data and perform preprocessing. Specifically, the disease image data includes fundus images and pathological slices, etc. The preprocessing includes denoising and normalization.
[0058] The disease grading unit 12 is used to predict the disease grade of the pre-processed disease image using the disease image grading model. The disease image grading model is constructed in the following way: Figure 2 As shown in the figure, based on the student model, multiple expert models, shallow feature alignment module, compact feature alignment module, and uncertainty-aware decoupling distillation module are introduced to construct a multi-expert knowledge distillation (Multi-expert Knowledge Distillation) framework, and uncertainty-aware distillation learning is performed on the multi-expert knowledge distillation framework. Multiple expert models extract feature information of disease images from different angles. The student model simulates the feature representation of the corresponding layer of the expert model through the shallow feature alignment module (SFA) and the compact feature alignment module (CFA), and dynamically adjusts the knowledge transfer weight according to the uncertainty output by the expert model through the uncertainty-aware decoupling distillation module (UDD). Ultimately, the student model can combine the advantages of multiple expert models to output high-precision disease image grading results, providing strong support for clinical diagnosis. The learned student model serves as a disease image grading model. This framework improves the generalization ability and robustness of the student model by integrating the diverse knowledge of multiple expert models, specifically including:
[0059] (1) Collect data and preprocess and deploy expert models.
[0060] We collected datasets of disease image classification samples, such as the SICAPv2 dataset for histological prostate grading and the APTOS dataset for diabetic retinopathy grading. The final annotations are the image classification labels, which represent the disease classification category. For the prostate cancer dataset, 0 represents no cancer, 1 represents GG3, 2 represents GG4, and 3 represents GG5. For the diabetic retinopathy dataset, 0 represents no retinopathy, 1 represents mild non-proliferative diabetic retinopathy, 2 represents moderate non-proliferative diabetic retinopathy, 3 represents severe non-proliferative diabetic retinopathy, and 4 represents proliferative diabetic retinopathy.
[0061] The present invention was experimentally verified in tasks under two real-world scenarios: source imbalanced distillation and target imbalanced distillation. In source imbalanced distillation, the expert model is trained on a source dataset with imbalanced categories, while the distillation process is performed on a target dataset with balanced categories. In target imbalanced distillation, the expert model is trained on a source dataset with balanced categories, while the distillation process is performed on a target dataset with imbalanced categories. In the specific implementation, balanced sub-datasets were generated by randomly sampling the original imbalanced dataset, including: SICAPv2-balanced (2500, 2222, 2500, 948): the number of samples for each hierarchical category. APTOS-balanced (600, 370, 300, 193, 295): the number of samples for each hierarchical category. The training set, validation set, and test set of all datasets are divided into a ratio of 8:1:1. In order to expand the dataset, data augmentation techniques were adopted, including Z-score normalization, random cropping, and flipping. It is worth noting that color dithering was not used because pathological images are sensitive to color changes and random color injection may destroy pathological features.
[0062] During model training, two ResNet50 models pre-trained on ImageNet were used as expert models T1 and T2. These expert models T1 and T2 were pre-trained on the corresponding imbalanced datasets. A ResNet18 model pre-trained on ImageNet was used as the student model S. Both the student model and multiple expert models were used to extract shallow, task-independent structural features and deep, task-relevant semantic features from disease images. Disease classification prediction was then performed. This invention aims to distill the knowledge from multiple expert models into a single, target student model S to solve the imbalanced disease classification task.
[0063] (2) Constructing a shallow feature alignment module
[0064] This paper introduces feature decoupling learning into feature alignment to improve performance in disease image classification tasks. Traditional knowledge distillation methods directly employ joint representation learning, which fails to effectively distinguish between task-relevant and irrelevant features, leading to performance degradation. To address this, this paper combines shallow feature alignment (SFA) and compact feature alignment (CFA) to address different levels of features.
[0065] Specifically, a shallow feature alignment (SFA) module is constructed to align the structural features extracted by the expert model and the student model in the shallow feature space, while retaining the main structural information of the image and removing noise and high-frequency details, thereby ensuring the consistency of the structural feature learning between the expert model and the student model. Specifically, the shallow feature alignment (SFA) module uses a multi-scale low-pass filter to process the features of the expert model, retaining the structural features of the disease image (i.e., features that are not related to the task), and at the same time processes the structural features of the student model through a learnable low-pass filter to achieve the student model's learning and generalization of the expert model's structural features. The specific implementation method is as follows:
[0066] Shallow structural features of expert models The traditional average pooling is used as a low-pass filter, and multi-scale filters are constructed by adjusting different kernel sizes and strides to adapt to different cutoff frequencies. For the mth group of filters, the projection features of the expert model in the frequency domain It is expressed as follows through a multi-scale low-pass filter (msLF):
[0067]
[0068] in, Indicates that the kernel size is k m ×k m The average pooling function of , Φ(·) represents the bilinear interpolation operation.
[0069] For the shallow structural features F of the student model S , a learnable low-pass filter is designed, which consists of multi-scale low-pass filtering, convolution downsampling module and depth-wise separable convolution (DSConv). The projection features of the student model in the frequency domain Expressed as:
[0070]
[0071] Among them, DownSample s×s Represents the convolution downsampling module, Concat represents the feature splicing operation, Conv 3×3 Represents a 3×3 convolution operation.
[0072] Constructing feature alignment loss function for shallow feature alignment (SFA) module To measure the distribution difference between the student model and the expert model in the feature space, the feature alignment is achieved through the maximum mean difference (MMD) and reconstruction loss (MSE). Among them, the feature alignment loss function It is reused in both the shallow feature alignment (SFA) module and the compact feature alignment (CFA) module.
[0073] First, in order to measure the distribution difference between the structural features of the student model and the structural features of each expert model in the feature space, the maximum mean difference (MMD) is used as a measurement tool, and the MMD loss The calculation formula is as follows: in, is an explicit mapping function, Represents the t-th expert model T in the i-th sample in a single batch t The structural features of the projected features after alignment, represents the projected features of the structural features of the student model S in the j-th sample in a single batch after alignment, B is the batch size, and N is the number of expert models.
[0074] At the same time, in order to ensure that the expert model remains unchanged before and after feature alignment (for example, due to privacy constraints), the present invention uses reconstruction loss (MSE) to measure the changes of the expert model before and after feature alignment. The calculation formula is: Among them, F Tt Represents the expert model T before alignment t The original structural characteristics of Represents the aligned expert model T t The projected features are aggregated using the maximum mean difference loss and reconstruction losses Total feature alignment loss It can be expressed as The loss function It is used to guide the student model to align with the expert model in the feature space, while ensuring that the features of the expert model remain unchanged during the distillation process.
[0075] (3) Constructing a compact feature alignment module
[0076] A Compact Feature Alignment (CFA) module is constructed to align the semantic features extracted by the expert model and the student model, thereby ensuring consistency in semantic feature learning between the expert and student models. The CFA module maps the features of the expert and student models to a common high-dimensional spherical space. Utilizing the feature alignment mechanism in the spatial domain, the student model learns task-related semantic features from each expert. This approach ensures that the student model can focus on the important semantic information in disease images, thereby effectively improving performance in disease image grading tasks. The specific implementation method is as follows:
[0077] The Compact Feature Alignment (CFA) module projects the feature set of the fourth layer (the penultimate layer before the fully connected layer) of the model into a compact high-dimensional spherical space Z. In this space Z, the student model is able to learn task-related semantic information from different expert models through spatial domain feature alignment. Since the feature dimensions of the feature extraction outputs of different models may be different, the present invention adds a 1×1 convolution kernel at the end of the output features before performing compact feature alignment to adjust the different output features to the same dimension. Finally, in the spherical space Z, the total feature alignment loss of the Compact Feature Alignment module (CFA) is The maximum mean difference (MMD) loss and reconstruction loss (MSE) are also used for calculation, which is expressed as: The specific calculation method is the same as that in shallow feature alignment.
[0078] (4) Constructing uncertainty-aware decoupling distillation module
[0079] By constructing an uncertainty-aware decoupled distillation (UDD) module, global and local knowledge is dynamically transferred from each expert model to the student network to address the problem of deviation in the expert model's output prediction. Specifically, the knowledge transfer weight can be dynamically adjusted according to the output uncertainty of the expert model, realizing dynamic knowledge distillation from the expert model to the student model. This module can automatically detect the uncertainty in the expert model caused by factors such as class imbalance, and specifically calculate the uncertainty by measuring the deviation between the expert model output and the ideal distribution. Based on this uncertainty, the student model can dynamically adjust the weight of knowledge transfer, that is, for fuzzy areas where the expert model prediction uncertainty is high, the system will increase the supervision weight of the area, while for task-irrelevant areas where the expert model prediction uncertainty is low, the system will maintain precise alignment. This setting ensures that the weight adjustment in the knowledge transfer process is more flexible and accurate, avoids the amplification of deviations caused by factors such as class imbalance in the knowledge transfer process, makes the distillation process more robust and reliable, and thus improves the performance and stability of the student model. The specific implementation method is as follows:
[0080] For the expert model T t and the logits (logistic) output of the student model S and C, H, W represent the length, height, and width of the image. The present invention uses multi-scale w∈W={1,2,4,…,w max}, and for each partition unit Z(m,n) under scale w (where n∈N w ={1,4,16,…,w 2}), the calculation formula for cumulative logits is as follows:
[0081]
[0082] in, and L S (j,k) represents the expert model T t The logits output value of the (j,k) position in the partition unit Z(m,n) corresponding to the student model S, ψT t (w,n) and ψS(w,n) represent the expert model T t The cumulative logits of the partition unit Z(m,n) corresponding to the student model S.
[0083] The present invention also designs the uncertainty coefficient in combination with the prediction confidence of the expert model The calculation formula is: Among them, σ represents the softmax function, max represents the maximum value function, and the coefficient The uncertainty of the prediction is quantified by measuring the deviation from the one-hot distribution.
[0084] Based on the decoupling knowledge distillation paradigm, based on the uncertainty coefficient and logits cumulative value ψT t (w,n) and ψS(w,n) define the uncertainty-aware disentangled distillation loss for: Among them, the loss component is defined as: The item is used to enhance the blur area supervision, and Items are used to keep task-independent areas The exact logit alignment of Represents the square of the L2 norm.
[0085] (5) Perform uncertainty-aware distillation learning training on the multi-expert knowledge distillation framework constructed in steps (1)-(4).
[0086] When conducting uncertainty-aware distillation learning training, the total training loss function used includes the classification loss of the image classification task Feature alignment loss and uncertainty-aware decoupling distillation loss Expressed as: Among them, α and β are the balance coefficients between the various losses, and the classification loss It means that the features extracted by the student model are predicted through the fully connected layer to obtain the predicted probability of each category, and the cross entropy loss is calculated between the predicted probability and the real label. Used to optimize the classification performance of the student model.
[0087] During training, the model was trained using a five-fold cross-validation approach, with each fold undergoing 50 epochs. Each time, a batch of data was input, the loss was calculated and backpropagated, and the model parameters were updated until training was complete. The model with the best validation results was saved during the iterations. Prediction performance was measured on the validation set using overall accuracy (OA), mean accuracy (mAcc), weighted F1-score (F1), and mean absolute error (MAE).
[0088] In practical applications, the disease grading system based on image uncertainty perception distillation can accurately predict the current stage of disease progression by analyzing a patient's medical images (such as fundus images) or pathological slides. Based on these predictions, doctors can more accurately assess the condition and optimize subsequent treatment plans, thereby improving treatment efficacy and patient prognosis.
[0089] Based on the same inventive concept, an embodiment further provides a computing device comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the device is used to implement a disease grading method based on image uncertainty perception distillation, wherein the method adopts the above-mentioned disease grading system based on image uncertainty perception distillation, such as Figure 3 As shown, the following steps are included:
[0090] S1, acquiring disease image data through a data acquisition unit and performing preprocessing;
[0091] S2, using the disease image classification model through the disease classification unit to predict the disease classification of the preprocessed disease image.
[0092] The computing device provided in the embodiment, at the hardware level, includes not only a processor and memory, but also hardware required for other services such as an internal bus, a network interface, and memory. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the disease grading method based on image uncertainty perception distillation described in S1-S2 above. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0093] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a disease grading method based on image uncertainty perception distillation is implemented. The method adopts the above-mentioned disease grading system based on image uncertainty perception distillation, and includes the following steps:
[0094] S1, acquiring disease image data through a data acquisition unit and performing preprocessing;
[0095] S2, using the disease image classification model through the disease classification unit to predict the disease classification of the preprocessed disease image.
[0096] In the embodiment, computer-readable media includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.
[0097] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A disease grading system based on image uncertainty perception distillation, characterized by: include: A data acquisition unit, which is used to acquire disease image data and perform preprocessing; a disease grading unit, configured to predict the disease grading of the pre-processed disease image using a disease image grading model; The disease image grading model is constructed in the following way: based on the student model, multiple expert models, shallow feature alignment module, compact feature alignment module, and uncertainty-aware decoupling distillation module are introduced to construct a multi-expert knowledge distillation framework. The student model and multiple expert models are used to extract shallow and task-independent structural features and deep and task-related semantic features in disease images and perform disease grading task predictions. The shallow feature alignment module is used to align the structural features extracted by the expert model and the student model. The compact feature alignment module is used to align the semantic features extracted by the expert model and the student model. The uncertainty-aware decoupling distillation module is used to dynamically adjust the knowledge transfer weight according to the uncertainty of the expert model to realize dynamic knowledge distillation from the expert model to the student model, and then combine the disease grading task to perform uncertainty-aware distillation learning. The learned student model is used as the disease image grading model.
2. The disease grading system based on image uncertainty perception distillation according to claim 1 is characterized in that: In the shallow feature alignment module, the structural features extracted by the expert model and the student model are aligned, including: A multi-scale low-pass filter is used to process the structural features of the expert model to retain the structural features of the disease image. At the same time, a learnable low-pass filter is used to process the structural features of the student model to enable the student model to learn and generalize the structural features of the expert model.
3. The disease grading system based on image uncertainty perception distillation according to claim 2 is characterized in that: According to the structural characteristics of the expert model, average pooling is used as a low-pass filter, and multi-scale filters are constructed by adjusting different kernel sizes and strides to adapt to different cutoff frequencies. For the mth group of filters, the projection characteristics of the expert model in the frequency domain are It is expressed as follows through a multi-scale low-pass filter (msLF): in, Indicates that the kernel size is k m ×k m The average pooling function of , Φ(·) represents the bilinear interpolation operation; For the shallow structural features F of the student model S The designed learnable low-pass filter consists of multi-scale low-pass filtering, convolution downsampling module and depth-wise separable convolution. The projection features of the student model in the frequency domain Expressed as: Among them, DownSample s×s Represents the convolution downsampling module, Concat represents the feature concatenation operation, and Conv represents the convolution operation.
4. The disease grading system based on image uncertainty perception distillation according to claim 3 is characterized in that: For shallow feature alignment module, projection features based on expert model and the projection features of the student model Total feature alignment loss By maximum mean difference loss and reconstruction losses Composition, expressed as It is used to guide the student model to align with the expert model in the feature space, while ensuring that the features of the expert model remain unchanged during the distillation process; The maximum mean difference loss Measuring the structural characteristics of the student model Projected features of each expert model Distribution differences in feature space; Reconstruction losses Measures the change in the expert model before and after feature alignment.
5. The disease grading system based on image uncertainty perception distillation according to claim 1 is characterized in that: In the compact feature alignment module, the semantic features extracted by the expert model and the student model are aligned, including: The semantic features of the expert model and the student model are projected into a compact high-dimensional spherical space Z. In the high-dimensional spherical space Z, the student model can learn task-related semantic information from different expert models through spatial domain feature alignment. The corresponding total feature alignment loss is Including maximum mean difference loss and reconstruction losses Calculate and express it as:
6. The disease grading system based on image uncertainty perception distillation according to claim 1 is characterized in that: In the uncertainty-aware decoupling distillation module, the knowledge transfer weight is dynamically adjusted according to the uncertainty of the expert model, including: Uncertainty is calculated by measuring the deviation between the graded probability distribution output by the expert model and the ideal probability distribution. Based on the uncertainty, the student model can dynamically adjust the weight of knowledge transfer, that is, for fuzzy areas where the expert model predicts high uncertainty, the supervision weight of the area will be increased, while for task-irrelevant areas where the expert model predicts low uncertainty, precise alignment will be maintained.
7. The disease grading system based on image uncertainty perception distillation according to claim 6 is characterized in that: The uncertainty is expressed by the uncertainty coefficient Indicates that its construction process is: For the expert model T t And the logical probability prediction output of the student model S and C, H, W represent the length, height, and width of the image. In the multi-scale w∈W={1,2,4,…,w max }, and perform spatial partitioning P(w,w) on each scale w, for each partition unit Z(m,n), where n∈N w ={1,4,16,…,w 2 }, the calculation formula of the cumulative logical probability prediction value is as follows: in, and L S (j,k) represents the expert model T t The logical probability prediction value of the (j, k) position in the partition unit Z(m,n) corresponding to the student model S, ψT t (w,n) and ψS(w,n) represent the expert model T t The cumulative logical probability of the partition unit Z(m,n) corresponding to the student model S; Design uncertainty coefficient based on prediction confidence of expert model The calculation formula is: Among them, σ represents the softmax function, max represents the maximum value function, and the coefficient uncertainty The uncertainty of a forecast is quantified by measuring the deviation from an ideal probability distribution.
8. The disease grading system based on image uncertainty perception distillation according to claim 7 is characterized in that: The uncertainty-based student model can dynamically adjust the weight of knowledge transfer by defining uncertainty-aware decoupling distillation loss. To achieve this, it is expressed as: Among them, the loss component is defined as: The item is used to enhance the blur area supervision, and Items are used to keep task-independent areas The exact logit alignment of Represents the square of the L2 norm.
9. The disease grading system based on image uncertainty perception distillation according to claim 8, characterized in that: The total training loss function used in the uncertainty-aware distillation learning of the entire multi-expert knowledge distillation framework includes the classification loss of the image classification task Total feature alignment loss of the shallow feature alignment module Total feature alignment loss of the compact feature alignment module and uncertainty-aware decoupling distillation loss Expressed as: Among them, α and β are the balance coefficients between various losses.
10. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, the method is used to implement a disease grading method based on image uncertainty perception distillation, wherein the method adopts the disease grading system based on image uncertainty perception distillation according to any one of claims 1 to 9, and includes the following steps: Acquire disease image data through a data acquisition unit and perform preprocessing; The disease grading unit adopts the disease image grading model to predict the disease grading of the preprocessed disease image.