Medical image classification method, system, electronic device and storage medium

Through the methods of trust evaluation and dynamic weight adjustment, the problems of high computational overhead and uncertainty processing in medical image classification in resource-constrained environments are solved, and the classification accuracy and robustness are improved.

CN120318610BActive Publication Date: 2025-09-19CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510805946.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-19
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing medical image classification methods have high computational overhead in resource-constrained environments and have difficulty handling the uncertainty levels of different medical images, which affects classification accuracy.

Method used

The confidence and uncertainty are obtained through the target teacher model, and the trust weight is obtained by trust evaluation. The loss data weight is adjusted, and the initial student model is trained through backpropagation to dynamically adapt to the uncertainty of different medical images.

Benefits of technology

The accuracy of medical image classification is improved, enabling the student model to achieve performance close to that of the teacher model with less computing resources and adapt to the uncertainty levels of different medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318610B_ABST
    Figure CN120318610B_ABST
Patent Text Reader

Abstract

The present application provides a medical image classification method, system, electronic device and storage medium, which belongs to the field of image processing. The method includes obtaining a sample medical image; performing a first image prediction on the sample medical image through a target teacher model to obtain teacher image prediction data; obtaining confidence from the teacher image prediction data, and obtaining uncertainty based on the confidence; performing a trust evaluation on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight; performing a second image prediction on the sample medical image through a pre-built initial student model to obtain student image prediction data; performing loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; weighting the loss data according to the target trust weight, thereby back-propagating and training the initial student model to obtain a target student model for medical image classification. The present application can improve the accuracy of medical image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a medical image classification method, system, electronic device and storage medium. Background Art

[0002] High-quality medical images are crucial for biomedical research and clinical diagnosis, enabling more efficient and accurate analysis and decision-making. With the rapid growth of medical image data from multiple imaging modalities, such as X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, medical image classification has become increasingly important in healthcare systems. Furthermore, medical image classification plays a key role in the diagnosis of a wide range of diseases across multiple imaging modalities. For example, skin diseases can be detected through dermatoscopy, cervical cancer lesions can be identified through histopathological images, respiratory diseases can be assessed through chest X-rays, neurological diseases can be diagnosed through brain MRI scans, and gastrointestinal diseases can be diagnosed through endoscopy and colonoscopy. These imaging techniques provide important clues for cell structure analysis and tumor detection, but they often present diagnostic challenges due to factors such as cell overlap and small object size. Therefore, while medical image classification has improved diagnostic accuracy and therapeutic outcomes, it still faces numerous challenges that affect system reliability. Therefore, continuous advancement of related technologies is necessary to overcome these limitations. Effective classification and timely disease detection are crucial for improving patient survival and developing appropriate treatment plans.

[0003] Automating the medical image classification process allows medical staff to spend more time on patient care rather than the tedious task of image interpretation. This shift not only improves diagnostic accuracy but also reduces the cognitive burden on physicians, enabling faster decision-making and improving overall medical outcomes. Furthermore, automated classification of medical imaging modalities is crucial for physicians to quickly access relevant image information, contributing to improved medical education and the diagnostic process. This automated classification process leverages advanced computer vision and machine learning (ML) techniques to accurately detect and diagnose a wide range of diseases. Compared to traditional methods that rely on manually designed features, computer-aided diagnosis (CAD) systems based on deep learning (DL) can automatically extract key features from large datasets, overcoming the limitations of previous approaches. Deep learning methods effectively capture more complex patterns in medical images through hierarchical representations. Furthermore, transfer learning (TL) can further optimize this process by leveraging pretrained convolutional neural networks (CNNs), reducing the reliance on large datasets and improving training efficiency.

[0004] To improve the accuracy and robustness of medical image classification, researchers have proposed a variety of innovative methods, including the use of convolutional neural networks (CNNs), ensemble learning frameworks, feature fusion strategies, and Transformer-based architectures. However, these advanced models often have high computational overhead, limiting their application in resource-constrained environments. Knowledge distillation provides an effective solution to this problem. By compressing the knowledge of a complex "large model" (teacher model) into a simple "small model" (student model), the student model can achieve similar performance with fewer computational resources. However, current knowledge distillation methods rely on static weights, which are insufficient to handle the varying levels of uncertainty in medical images. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a medical image classification method, system, electronic device and storage medium, aiming to improve the accuracy of medical image classification.

[0006] To achieve the above-mentioned objectives, the first aspect of an embodiment of the present application proposes a medical image classification method, which includes: obtaining a sample medical image; performing a first image prediction on the sample medical image through a target teacher model to obtain teacher image prediction data; obtaining confidence from the teacher image prediction data, and obtaining uncertainty based on the confidence; performing a trust evaluation on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight; performing a second image prediction on the sample medical image through a pre-constructed initial student model to obtain student image prediction data; performing loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; weighting the loss data according to the target trust weight, thereby back-propagating and training the initial student model to obtain a target student model for medical image classification.

[0007] To achieve the above-mentioned purpose, the second aspect of an embodiment of the present application proposes a medical image classification system, which includes: a data acquisition module for acquiring sample medical images; a teacher model prediction module for performing a first image prediction on the sample medical image through a target teacher model to obtain teacher image prediction data; an indicator acquisition module for obtaining confidence from the teacher image prediction data and obtaining uncertainty based on the confidence; a weight acquisition module for performing a trust evaluation on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight; a student model prediction module for performing a second image prediction on the sample medical image through a pre-built initial student model to obtain student image prediction data; a loss calculation module for performing loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; a reverse training module for weighting the loss data according to the target trust weight, thereby backpropagating and training the initial student model to obtain a target student model for medical image classification.

[0008] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the method described in the first or second aspect above when executing the computer program.

[0009] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first or second aspect above.

[0010] The medical image classification method, system, electronic device, and storage medium proposed in this application obtain sample medical images; perform a first image prediction on the sample medical images using a target teacher model to obtain teacher image prediction data; obtain confidence from the teacher image prediction data and, based on the confidence, obtain uncertainty, which helps quantify the target teacher model's confidence level in the prediction results and provides a basis for trust assessment. Furthermore, a trust assessment is performed on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight, which helps identify and utilize highly reliable prediction results, thereby improving the training effect of the student model. Furthermore, a second image prediction is performed on the sample medical images using a pre-built initial student model to obtain student image prediction data; loss is calculated based on the student image prediction data and the teacher image prediction data to obtain loss data; and the loss data is weighted according to the target trust weight, thereby backpropagating the initial student model to obtain a target student model for medical image classification. This step can dynamically adjust the student model's reliance on the target teacher model based on the uncertainty of the medical image using the target trust weight, allowing the student model to adapt to the different uncertainty levels of different medical images, thereby improving the accuracy of medical image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a flowchart of the medical image classification method provided in an embodiment of the present application;

[0012] Figure 2 is another flow chart of the medical image classification method provided in an embodiment of the present application;

[0013] Figure 3 is another flow chart of the medical image classification method provided in an embodiment of the present application;

[0014] Figure 4 yes Figure 3 Flowchart of step S305;

[0015] Figure 5 yes Figure 1 Flowchart of step S104;

[0016] Figure 6 yes Figure 5 Flowchart of step S503;

[0017] Figure 7 yes Figure 1 Flowchart of step S105;

[0018] Figure 8 is a flowchart of a medical image classification system provided by an embodiment of the present application;

[0019] Figure 9This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should be noted that although the functional modules are divided in the system schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the system or the order in the flow chart. The terms "first", "second" and the like in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0021] First, let’s analyze some of the terms used in this application:

[0022] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0023] High-quality medical images are crucial for biomedical research and clinical diagnosis, enabling more efficient and accurate analysis and decision-making. With the rapid growth of medical image data from multiple imaging modalities, such as X-rays, computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, medical image classification has become increasingly important in healthcare systems. Furthermore, medical image classification plays a key role in the diagnosis of a wide range of diseases across multiple imaging modalities. For example, skin diseases can be detected through dermatoscopy, cervical cancer lesions can be identified through histopathological images, respiratory diseases can be assessed through chest X-rays, neurological diseases can be diagnosed through brain MRI scans, and gastrointestinal diseases can be diagnosed through endoscopy and colonoscopy. These imaging techniques provide important clues for cell structure analysis and tumor detection, but they often present diagnostic challenges due to factors such as cell overlap and small object size. Therefore, while medical image classification has improved diagnostic accuracy and therapeutic outcomes, it still faces numerous challenges that affect system reliability. Therefore, continuous advancement of related technologies is necessary to overcome these limitations. Effective classification and timely disease detection are crucial for improving patient survival and developing appropriate treatment plans.

[0024] Automating the medical image classification process allows medical staff to spend more time on patient care rather than the tedious task of image interpretation. This shift not only improves diagnostic accuracy but also reduces the cognitive burden on physicians, enabling faster decision-making and improving overall medical outcomes. Furthermore, automated classification of medical imaging modalities is crucial for physicians to quickly access relevant image information, contributing to improved medical education and the diagnostic process. This automated classification process leverages advanced computer vision and machine learning (ML) techniques to accurately detect and diagnose a wide range of diseases. Compared to traditional methods that rely on manually designed features, computer-aided diagnosis (CAD) systems based on deep learning (DL) can automatically extract key features from large datasets, overcoming the limitations of previous approaches. Deep learning methods effectively capture more complex patterns in medical images through hierarchical representations. Furthermore, transfer learning (TL) can further optimize this process by leveraging pretrained convolutional neural networks (CNNs), reducing the reliance on large datasets and improving training efficiency.

[0025] To improve the accuracy and robustness of medical image classification, researchers have proposed a variety of innovative methods, including the use of convolutional neural networks (CNNs), ensemble learning frameworks, feature fusion strategies, and Transformer-based architectures. However, these advanced models often have high computational overhead, limiting their application in resource-constrained environments. Knowledge distillation provides an effective solution to this problem. By compressing the knowledge of a complex "large model" (teacher model) into a simple "small model" (student model), the student model can achieve similar performance with fewer computational resources. However, current knowledge distillation methods rely on static weights, which are insufficient to handle the varying levels of uncertainty in medical images.

[0026] Based on this, the embodiments of the present application provide a medical image classification method, system, electronic device and storage medium, aiming to improve the accuracy of medical image classification.

[0027] The medical image classification method, system, electronic device, and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the medical image classification method in the embodiments of the present application is described.

[0028] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology and machine learning / deep learning. The medical image classification method provided in the embodiments of the present application relates to the field of image processing. The medical image classification method provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, etc.; the server can be configured as a standalone physical server, or as a server cluster or distributed system consisting of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements medical image classification methods, etc., but is not limited to the above forms. The present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example, personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media including memory storage devices.

[0029] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the identity or characteristics of the object, such as object information, object behavior data, object historical data, and object location information, the permission or consent of the object will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the sensitive personal information of the object, the separate permission or consent of the object will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the separate permission or consent of the object, the necessary object-related data for the normal operation of the embodiment of the present application will be obtained.

[0030] See also Figure 1 , Figure 1 This is an optional flowchart of the medical image classification method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S107: step S101, obtaining a sample medical image; step S102, performing a first image prediction on the sample medical image through a target teacher model to obtain teacher image prediction data; step S103, obtaining confidence from the teacher image prediction data, and obtaining uncertainty based on the confidence; step S104, performing a trust evaluation on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight; step S105, performing a second image prediction on the sample medical image through a pre-constructed initial student model to obtain student image prediction data; step S106, performing loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; step S107, weighting the loss data according to the target trust weight, thereby back-propagating and training the initial student model to obtain a target student model for medical image classification.

[0031] In the embodiment of the present application, steps S101 to S107 are as follows: obtaining a sample medical image; performing a first image prediction on the sample medical image using a target teacher model to obtain teacher image prediction data; obtaining confidence from the teacher image prediction data, and obtaining uncertainty based on the confidence, which helps to quantify the target teacher model's confidence level in the prediction result and provide a basis for trust evaluation. Further, trust evaluation is performed on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight, which helps to identify and utilize highly reliable prediction results, thereby improving the training effect of the student model. Further, a second image prediction is performed on the sample medical image using a pre-built initial student model to obtain student image prediction data; loss calculation is performed based on the student image prediction data and the teacher image prediction data to obtain loss data; the loss data is weighted according to the target trust weight, thereby backpropagating the training of the initial student model to obtain a target student model for medical image classification. This step can dynamically adjust the student model's dependence on the target teacher model based on the uncertainty of the medical image using the target trust weight, so that the student model can adapt to the different uncertainty levels of different medical images, thereby improving the classification accuracy of medical images.

[0032] In step S101 of some embodiments, the sample medical images are used as samples for model training. They can be multi-disease medical images, meaning one image corresponds to one disease, and multiple images can correspond to multiple diseases. Medical images can be acquired using various imaging modalities (such as X-ray imaging, computed tomography, magnetic resonance imaging, ultrasound imaging, and endoscopic imaging), reflecting different tissues, organs, structures, or pathological conditions. These sample medical images have corresponding ground truth labels. Ground truth labels are typically provided by experienced clinical experts and indicate the type of disease associated with the image. Ground truth labels can also include lesion location, lesion outline, or other diagnostic information. For example, the ground truth label for a lung CT image acquired using CT imaging might be pneumonia.

[0033] See also Figure 2In some embodiments, before step S102, that is, before performing a first image prediction on the sample medical image through the target teacher model, the medical image classification method further includes updating the sample medical image, specifically including: step S201, adjusting the resolution of the sample medical image to obtain a standard sample medical image; step S202, performing correction processing on the standard sample medical image to obtain a first image; step S203, performing histogram equalization on the standard sample medical image to obtain a second image; step S204, using a predetermined function to decompose the first image and the second image respectively to obtain a first multi-scale frequency component corresponding to the first image and a second multi-scale frequency component corresponding to the second image; step S205, fusing the first multi-scale frequency component and the second multi-scale frequency component to obtain a fused frequency component; step S206, reconstructing the fused frequency component through an inverse wavelet transform to obtain a reconstructed sample medical image; step S207, linearly mapping the pixel values ​​of the reconstructed sample medical image to obtain an updated sample medical image.

[0034] In step S201 of some embodiments, the standard sample medical image is a resized sample medical image with a uniform spatial dimension or pixel value range. Resolution adjustment refers to scaling the image to a preset standard, resulting in a uniform width and height. Specifically, the resolution of the sample medical image can be adjusted to 224×224 using methods such as bilinear interpolation or nearest neighbor interpolation. This step prepares for subsequent image processing and model input, eliminating the impact of image size on processing results and model performance.

[0035] In step S202 of some embodiments, the first image is obtained by performing correction processing on the standard sample medical image. The correction processing can adjust the image brightness. Specifically, the correction processing can be performed using gamma correction technology, which is a nonlinear image enhancement technology that adjusts the image brightness by applying a power law transformation to the pixel values. The enhancement process follows the formula:

[0036] , formula (1);

[0037] in, and are the input and output pixel values, respectively. is the scaling constant, (Gamma) determines the type of brightness adjustment. When <1, the pixel value of the middle gray area is increased, and the image becomes brighter as a whole, which is suitable for underexposed images; when When the value is greater than 1, the middle gray is suppressed, and the overall image becomes darker, which is suitable for overexposed images. Compared with the linear enhancement method, gamma correction can adjust the image brightness while preserving the details in the shadows and highlights.

[0038] In step S203 of some embodiments, the second image is an image obtained by performing histogram equalization on a standard sample medical image. Histogram equalization can adjust the contrast of an image. Specifically, the contrast-limited adaptive histogram equalization (CLAHE) technology can be used to perform histogram equalization on an image. This is a contrast enhancement method based on local areas. It divides the image into multiple small blocks, performs histogram equalization on each small block separately, and limits the enhancement range of the histogram to avoid noise amplification. This method is particularly suitable for processing medical scan images, satellite images, and low-light images with uneven lighting problems. This method can enhance the details of areas with uneven lighting. Its core mechanism includes histogram cropping and redistribution: for each block , its histogram (have gray levels) at the threshold is cut, where is the total number of pixels in each image block, is the number of gray levels, is the crop factor. The extra pixels is evenly redistributed to all gray levels. Then the cumulative distribution function (CDF) is calculated and mapped. Finally, each block is interpolated by bilinear interpolation. Perform smooth fusion to avoid edge artifacts.

[0039] In step S204 of some embodiments, a predetermined function performs discrete wavelet transform (DWT) processing on the image to obtain coefficients (approximate and detailed components) obtained by decomposing the image at multiple scales and directions, i.e., obtaining multi-scale frequency components. Specifically, the wavedec2 function can be used to decompose the first and second images into low-frequency (approximate LL) and high-frequency (including horizontal LH, vertical HL, and diagonal HH) coefficients.

[0040] In step S205 of some embodiments, after obtaining the first multiscale frequency component and the second multiscale frequency component, the first multiscale frequency component and the second multiscale frequency component are fused at the same decomposition level and direction using an averaging operation to obtain fused multiscale wavelet coefficients, i.e., fused frequency components. If the first multiscale frequency component is [LL1, LH1, HL1, HH1] and the second multiscale frequency component is [LL2, LH2, HL2, HH2], then the fused frequency component obtained by the fusion process at the same decomposition level and direction is [(LL1+LL2) / 2, (LH1+LH2) / 2, (HL1+HL2) / 2, (HH1+HH2) / 2]. In other embodiments, the two multiscale frequency components can also be fused using a maximum value or weighted average operation.

[0041] In step S206 of some embodiments, the inverse wavelet transform (IDWT), as its name implies, is the inverse of the DWT. It is the process of recombining the frequency components (wavelet coefficients) after wavelet transformation to restore the original image or signal. Therefore, the inverse wavelet transform can be used to reconstruct the image based on the fused frequency components [(LL1+LL2) / 2, (LH1+LH2) / 2, (HL1+HL2) / 2, (HH1+HH2) / 2] to obtain a reconstructed sample medical image.

[0042] In step S207 of some embodiments, linear mapping is used to adjust the image pixel values ​​to the target range according to a linear function. Specifically, the pixel values ​​of the reconstructed sample medical image can be scaled to the intensity range of [0, 255] using Unit8 normalization technology to maintain consistency.

[0043] The above steps S201 to S207 unify the image resolution, perform brightness correction and contrast enhancement, and combine wavelet transform to achieve fusion and reconstruction of multi-scale frequency features, so as to fuse the complementary information of the two input images in the spatial and frequency domains, and ultimately generate updated sample medical images with higher quality, richer details and stronger diagnostic value, thereby providing a better data foundation for subsequent model training and medical analysis, and making the model learning consistent, avoiding hindering the generalization ability of the model.

[0044] Before step S102 in some embodiments, the sample medical image may be normalized, that is, the input image may be normalized to the range of [0, 1] to ensure the consistency of pixel value distribution, which is conducive to improving the convergence speed of the model during training.

[0045] In some embodiments, data augmentation can also be performed on the sample medical image before step S102. Specifically, data augmentation includes 90° rotation, horizontal flipping, and vertical flipping, which enhance the robustness of the model. 90° rotation enables the model to recognize features in different orientations, while flipping the image introduces greater data diversity, thereby improving the model's generalization and reducing overfitting. By exposing the model to a wider range of variation, these augmentation techniques ultimately improve its performance on unseen data (i.e., data not encountered during training).

[0046] Before step S102 in some embodiments, it is necessary to obtain a target teacher model. Figure 3 In some embodiments, the step of obtaining the target teacher model may include but is not limited to steps S301 to S305: step S301, obtaining multiple candidate teacher models; step S302, obtaining a set of medical images of multiple diseases; step S303, performing index calculations on each candidate teacher model based on the set of medical images of multiple diseases to obtain performance indicators corresponding to each candidate teacher model; wherein the performance indicators at least include verification accuracy and model calculation cost; step S304, initializing the pheromone matrix and the heuristic matrix; wherein the pheromone values ​​corresponding to each candidate teacher model in the pheromone matrix are the same, the heuristic matrix is ​​composed of multiple element values, and each element value is generated based on the performance indicators of each candidate teacher model; step S305, according to the set ant colony parameters, multiple rounds of iterative optimization are performed based on the initialized pheromone matrix and the heuristic matrix. After reaching the preset number of iterations, the highest target element value is screened out from the highest element value corresponding to each round of iteration, and the model corresponding to the target element value is used as the target teacher model.

[0047] In step S301 of some embodiments, the candidate teacher model is a model participating in the candidate evaluation. The candidate teacher model can be at least two of the following models: Densenet169, Densenet121, Densenet201, InceptionV3, MobileNetV1, MobileNetV2, ResNet50V2, ResNet101V2, ResNet152V2, VGG16, VGG19, or Xception.

[0048] In step S302 of some embodiments, the multi-disease medical image set refers to a collection of medical images containing multiple diseases, such as pneumonia, breast cancer, brain tumors, etc. Each image is labeled, and the label indicates the disease type of the image. The images in the multi-disease image set all have ImageNet features. ImageNet features refer to vectors obtained by extracting features from multi-disease medical images using a model pre-trained on ImageNet, which are used to supplement the original features of multi-disease medical images. This method of applying ImageNet features to multi-disease medical images can effectively alleviate the problem of medical data scarcity, improve semantic expression capabilities, accelerate model training and enhance generalization performance.

[0049] In step S303 of some embodiments, the validation accuracy is the classification accuracy of the model on an unseen validation set. The model computational cost is a metric that measures the model's computational resource consumption, such as the number of floating-point operations or inference time. Specifically, the validation accuracy and model computational cost of each candidate teacher model are calculated on the multi-disease medical image set.

[0050] In step S304 of some embodiments, the pheromone matrix reflects the initial value of the probability that each model is selected as a "path", indicating the exploration tendency. And the pheromone matrix includes pheromone values ​​corresponding to each candidate teacher model. The heuristic matrix is ​​an evaluation matrix constructed based on the model performance index, which guides the ant colony to be more inclined to models with better performance. And the heuristic matrix includes element values ​​corresponding to each candidate teacher model, and each element value is generated based on the performance index of each candidate teacher model. The higher the element value, the better the performance of the model. When initializing the pheromone matrix and the heuristic matrix, the pheromone values ​​corresponding to each candidate teacher model in the pheromone matrix are all set to 1. At the same time, the heuristic matrix is ​​constructed based on the verification accuracy and model calculation cost of each candidate teacher model. The value of each element in the heuristic matrix can be the value obtained by dividing the verification accuracy by the model calculation cost, or it can be a value obtained in other ways, which is not limited here.

[0051] In step S305 of some embodiments, the set ant colony parameters include the target number of ants, the number of iterations, the influence factor of pheromone, the influence factor of heuristic information, the pheromone evaporation rate, and the balance factor between exploration and utilization.

[0052] See also Figure 4In step S305 of some embodiments, a medical image classification method includes but is not limited to steps S401 to S403: in each round of iteration: step S401, a target number of ants calculate the selection probability of each candidate teacher model based on the pheromone matrix and heuristic matrix corresponding to this round of iteration, and use a roulette wheel selection method to select the target model corresponding to each ant; wherein the probability of the roulette wheel selection method is proportional to the selection probability; step S402, record the highest element value among the element values ​​corresponding to each target model in this round of iteration; step S403, update the pheromone matrix corresponding to this round of iteration based on the selection results of each ant on the model and the element value of each target model.

[0053] In each iteration, steps S401 to S403 are executed. In some embodiments, steps S401 to S403 begin with the first ant. The first ant calculates its selection probability for each candidate teacher model based on the pheromone matrix and heuristic matrix corresponding to that iteration. Then, using a roulette wheel selection method, it randomly selects a model from the candidate teacher models based on the calculated selection probabilities corresponding to each candidate teacher model, thereby obtaining the target model selected by the first ant. The roulette wheel selection method is a probability-based random selection method that constructs a "roulette wheel" based on the calculated selection probabilities, with the selected model being determined by the random landing point. Subsequent ants repeat a similar selection process as the first ant, thereby obtaining target models selected by a target number of ants. Then, the element value corresponding to each selected target model is obtained from the heuristic matrix, and the highest element value among the element values ​​corresponding to the target models selected in that iteration is recorded. Finally, the pheromone matrix corresponding to that iteration is updated based on the model selection results of each ant and the element value of each target model. The algorithm flow corresponding to the target teacher model acquisition step can be found in Table 1.

[0054] Table 1

[0055]

[0056]

[0057] In formula (2), Indicates in In the pheromone matrix of the first iteration The pheromone value corresponding to the model Power, Indicates the The heuristic matrix of the iteration The pheromone value corresponding to the model Power, and respectively control the relative importance of pheromone and heuristic information, is the total number of models.

[0058] For ease of understanding, assume =3, =2, is any value, =1, =2, the pheromone matrix in this round is [2,1,4], and the heuristic matrix is ​​[3,5,2]. The values ​​in the pheromone matrix and the heuristic matrix correspond to the first model, the second model, and the third model respectively. ,but , , At this time, the ants The probability of selecting the second model is 42.4%.

[0059] In formula (3), Indicates the In the pheromone matrix of the first iteration The pheromone value of each model; Indicates the In the pheromone matrix of the first iteration The pheromone value of each model; Indicates in The model in the iteration Whether it is ants Select, its value is 0 or 1; Indicates in Ants in the iteration The performance score of the selected model, that is, the value of the element corresponding to the selected model in the heuristic matrix.

[0060] For ease of understanding, assume v=0.1, =2, =3, the pheromone matrix in this round is [2,1,4], and the heuristic matrix is ​​[3,5,2]. The values ​​in the pheromone matrix and the heuristic matrix correspond to the first model, the second model, and the third model respectively. The first ant chooses model 2, the second ant chooses model 1, and the third ant chooses model 2. , similarly we can get , , then the pheromone matrix of the third iteration is [6.8, 12.9, 3.6].

[0061] In other embodiments, the performance index of each model can be calculated on a data set, and the calculated performance index can be used to generate a heuristic matrix, which can then be used in the ant After selecting the model, obtain the performance score of the selected model on another dataset and use the performance score as .

[0062] In steps S401 to S403, multiple ants use a roulette wheel selection method to probabilistically select candidate teacher models in each iteration based on the current pheromone matrix and heuristic matrix, simulating a multi-path exploration process. The system then records the element value of the best-performing model selected in each round to determine the quality of the current search state. Finally, according to the pheromone reinforcement principle, the pheromone matrix is ​​dynamically updated to strengthen high-quality model paths and suppress low-quality model paths, guiding the ant colony to prefer higher-performing models in subsequent iterations, gradually converging to the global optimal solution.

[0063] Steps S301 to S305 above effectively select the optimal teacher model using the Ant Colony Optimization (ANCO) algorithm, which intelligently explores and balances performance and computational cost among different models. Through a pheromone-guided exploration mechanism, ANCO identifies the model that strikes the best balance between maximizing performance and minimizing computational cost. Compared to the traditional approach of sequentially evaluating models and selecting the optimal one, ANCO dynamically adjusts the search process over multiple iterations, combining actual model performance with previous selection experience for continuous optimization, resulting in a more efficient and optimal model selection process. This approach not only conserves computing resources but also improves the accuracy of model selection, especially when faced with a large number of candidate models.

[0064] In step S102 of some embodiments, the first image prediction refers to the prediction performed on the sample medical image by the target teacher model, resulting in teacher image prediction data. The teacher image prediction data is a soft label, representing the probability distribution of the sample medical image across multiple categories. If the multiple categories are, in order, Disease A, Disease B, and Disease C, and the teacher image prediction data is [0.7, 0.2, 0.1], this means that the sample medical image has a 70% probability of belonging to Disease A, a 20% probability of belonging to Disease B, and a 10% probability of belonging to Disease C.

[0065] In one embodiment, the target teacher model can be MobileNetV1, whose structure may include a two-dimensional convolution layer (Conv2D), multiple depthwise separable convolution (DepthSepConv3×3) modules, a global average pooling layer (GAP), a fully connected layer (FC), and a classification layer (Softmax). The DepthSepConv3×3 module includes a depthwise convolution layer (DepthwiseConvolution) and a 1×1 pointwise convolution layer (PointwiseConvolution). Specifically, Conv2D is used for low-level feature extraction, the DepthSepConv3×3 module is used to process mid- / high-level feature extraction, the depthwise convolution layer is used for channel-by-channel spatial feature extraction, the pointwise convolution layer is used for cross-channel semantic fusion, GAP is used to compress the spatial dimension of the feature vector output by the previous convolution layer, FC is used to map the output of GAP to a classification vector, and Softmax is used to convert the classification vector into a probability distribution. In a more specific embodiment, there are five DepthSepConv3×3 modules, namely DepthSepConv3×3×2, DepthSepConv3×3×2, DepthSepConv3×3×2, DepthSepConv3×3×6, and DepthSepConv3×3×1. Taking DepthSepConv3×3×2 as an example, 3×3 refers to the spatial size of the convolution kernel is 3 rows × 3 columns, and ×2 refers to the channel expansion factor of 2, that is, the number of output channels is twice the number of input channels. When a sample image is input into the model, it is processed sequentially through a two-dimensional convolutional layer (Conv2D), a DepthSepConv3×3×2 layer, a DepthSepConv3×3×2 layer, a DepthSepConv3×3×2 layer, a DepthSepConv3×3×6 layer, a DepthSepConv3×3×1 layer, a global average pooling layer (GAP), a fully connected layer (FC), and a classification layer (Softmax). The output is a soft label, or probability distribution. This target teacher model replaces standard convolution with depthwise separable convolution, significantly reducing computational complexity while maintaining functionality.

[0066] In step S103 of some embodiments, the confidence level is the target teacher model's degree of certainty in its prediction results, and the uncertainty level is the uncertainty of the sample medical image. Specifically, the maximum probability value in the teacher image's prediction data can be used as the confidence level of the sample medical image. Methods such as entropy, Monte Carlo (dropout), and multiple sampling can then be used to evaluate the distribution of the prediction results and calculate their uncertainty. In some embodiments, the maximum probability value in the teacher image's prediction data can be used as the confidence level of the sample medical image, and (1-confidence level) can be used as the uncertainty level.

[0067] In step S104 of some embodiments, the target trust weight represents the degree of trust that the student model has in the teacher model during training, that is, the importance weight that the teacher model output should have in training the student model.

[0068] See also Figure 5 In step S104 of some embodiments, the medical image classification method may include but is not limited to steps S501 to S503: step S501, performing a level membership evaluation on the confidence according to a preset number of confidence membership functions to obtain a confidence level membership for each confidence level; wherein, the confidence levels of any two confidence membership functions are different, and the preset number is greater than or equal to 2; step S502, performing a level membership evaluation on the uncertainty according to a preset number of uncertainty membership functions to obtain an uncertainty level membership for each uncertainty level; wherein, the uncertainty levels of any two uncertainty membership functions are different; step S503, performing membership fusion according to the confidence level membership of each confidence level and the uncertainty level membership of each uncertainty level to obtain a target trust weight.

[0069] In some embodiments, in step S501 and step S502, the preset number is greater than or equal to 2, and can be 3 or 4. The preset number can be set according to needs and is not limited here. The confidence membership function refers to a function that fuzzily divides the confidence. The confidence levels of any two confidence membership functions are different. The confidence level membership refers to the degree of membership of the confidence in a certain confidence level, and the numerical range is usually [0,1]. The uncertainty membership function refers to a function that fuzzily divides the uncertainty. The uncertainty levels of any two uncertainty membership functions are different. The uncertainty level membership refers to the degree of membership of the uncertainty in a certain uncertainty level, and the numerical range is usually [0,1]. Specifically, the confidence can be divided into three levels, namely low, medium, and high, where low corresponds to a confidence between 0 and 0.5, medium corresponds to a confidence between 0.2 and 0.8, and high corresponds to a confidence between 0.5 and 1. The specific confidence membership function can be shown as formulas (4)-(6). At the same time, uncertainty can also be divided into three levels, namely low, medium and high, where low corresponds to uncertainty between 0 and 0.4, medium corresponds to uncertainty between 0.3 and 0.9, and high corresponds to uncertainty between 0.7 and 1. The specific uncertainty membership function can be shown as formulas (7)-(9).

[0070] , formula (4);

[0071] , formula (5);

[0072] , formula (6);

[0073] in, Refers to confidence level; refers to the low confidence level membership function; refers to the medium confidence level membership function; refers to the high confidence level membership function. Meanwhile, the low confidence level membership can be recorded as ; The confidence level membership can be recorded as ; The high confidence level membership can be recorded as .

[0074] , formula (7);

[0075] , formula (8);

[0076] , formula (9);

[0077] in, refers to uncertainty; refers to the low uncertainty level membership function; refers to the membership function of the medium uncertainty level; Refers to the high uncertainty level membership function. Meanwhile, the low uncertainty level membership can be recorded as ; The medium uncertainty level membership can be recorded as ; The high uncertainty level membership can be recorded as .

[0078] In step S503 of some embodiments, membership fusion refers to merging the two membership spaces of confidence and uncertainty.

[0079] See also Figure 6 In step S503 of some embodiments, the medical image classification method may include but is not limited to steps S601 to S603: step S601, orderly combining the confidence levels and the uncertainty levels to obtain level combinations, and determining the combination weight of each level combination; step S602, performing minimum selection on the confidence level membership and the uncertainty level membership in each level combination to obtain the fuzzy membership of each level combination; step S603, performing weight calculation based on the combination weight of each level combination and the fuzzy membership of the level combination to obtain the target trust weight.

[0080] In step S601 of some embodiments, the level combination refers to the pairing between the confidence level and the uncertainty level. The combination weight refers to the credibility evaluation weight assigned to each level combination, which is used to measure the degree of influence of the combination on the final trust weight. The order of the combination is determined according to the fuzzy rules actually set. Specifically, the fuzzy rules may include: Rule 1 (R1), if the confidence level is low and the uncertainty level is high, the output weight level is low; Rule 2 (R2), if the confidence level is medium and the uncertainty level is medium, the output weight is medium; Rule 3 (R3), if the confidence level is high and the uncertainty level is low, the output weight is high. Rule 1 means that when there is high uncertainty, the student model should reduce its dependence on the teacher model's prediction; Rule 2 represents a balanced dependence on the teacher model's prediction; Rule 3 means that when the teacher model is confident and the data uncertainty is low, the student model should strongly rely on the teacher model. According to this fuzzy rule, low confidence level and high uncertainty level can be grouped as group A, medium confidence level and medium uncertainty level can be grouped as group B, and high confidence level and low uncertainty level can be grouped as group C. The combination weight of each level combination is usually set based on increasing the influence of the teacher model on the student model when the teacher prediction is credible, and reducing the influence of the teacher model on the student model when the teacher prediction is uncredible. Under this constraint, group A The combined weight Can be set to 0.2, group B The combined weight Can be set to 0.5, group C The combined weight Can be set to 0.8.

[0081] In other embodiments, the fuzzy rules may include: if the confidence level is low and the uncertainty level is medium, the output weight level is generally low to medium (lower weight); if the confidence level is low and the uncertainty level is high, the output weight level is generally low (lowest weight); if the confidence level is medium and the uncertainty level is medium, the output weight level is generally medium (medium weight); if the confidence level is medium and the uncertainty level is high, the output weight level is generally medium to low (medium to lower weight); and so on.

[0082] It should be noted that the fuzzy rules are set according to the actual situation, but generally encourage reducing the reliance on the loss between the student model output and the teacher model output when the uncertainty is high, and reducing the reliance on the loss when the confidence is medium.

[0083] In step S602 of some embodiments, the fuzzy membership refers to the result obtained by selecting the minimum value based on the confidence level membership and the uncertainty level membership in each level combination. Specifically, based on R1, the minimum value in group A is obtained. ; Based on R2, get the minimum value in group B ; Based on R3, get the minimum value in group C .

[0084] In step S603 of some embodiments, after obtaining the combined weights of each level combination and the fuzzy membership of the level combination, a weight calculation is performed based on a preset fuzzy weight function to obtain a target trust weight. The preset fuzzy weight function is shown in formula (10), where: is the target trust weight.

[0085] , formula (10).

[0086] The above steps S601 to S603 calculate the combination of confidence and uncertainty levels through fuzzy rules to achieve refined calculation of trust weights, thereby ensuring that the model has adaptive learning and adjustment capabilities when facing different sample qualities.

[0087] Steps S501 to S503 above implement a comprehensive trust assessment of confidence and uncertainty information through a fuzzy logic mechanism. The system first categorizes the input confidence and uncertainty into fuzzy levels based on multiple preset confidence and uncertainty membership functions, obtaining the corresponding membership degrees for each level. Next, by integrating the membership degrees for each confidence and uncertainty level, a target trust weight is generated to quantify the student model's reliance on the teacher's output during the learning process for the current sample medical image.

[0088] In step S105 of some embodiments, a second image prediction is performed on the sample medical image using a pre-built initial student model to obtain student image prediction data.

[0089] See also Figure 7The initial student model includes an edge texture feature processing block, a shape feature processing block, an organ boundary feature processing block, a lesion feature processing block, a feature flattening block, and a classification layer. In step S105 of some embodiments, the medical image classification method may include but is not limited to steps S701 to S706: Step S701, performing basic feature processing on the sample medical image through the edge texture feature processing block to obtain edge texture features; Step S702, performing lightweight feature processing on the edge texture features through the shape feature processing block to obtain shape features; Step S703, performing complex feature processing on the shape features through the organ boundary feature processing block to obtain organ boundary features; Step S704, performing deep feature processing on the organ boundary features through the lesion feature processing block to obtain lesion features; Step S705, performing feature processing on the lesion features through the feature flattening block to obtain flattened features; Step S706, performing classification processing on the flattened features through the classification layer to obtain student image prediction data.

[0090] In step S701 of some embodiments, the edge texture feature processing block is used to obtain edge texture features. Edge texture features refer to areas in an image where brightness or grayscale changes suddenly, and the repetitive pattern and arrangement of pixel grayscale values ​​in the image. In order of execution, the edge texture feature processing block includes, in sequence, a first two-dimensional convolutional layer (Conv2D layer) with 32 filters, a first leaky ReLU layer (Leaky ReLU layer), and a first 2×2 maximum pooling layer (Max-Pooling2×2 layer). Among them, when the size of the sample medical image is 256×256×3, the first Conv2D layer is used to extract low-level features (such as edges and textures) from the image 256×256×3; the first Leaky ReLU layer is used to perform nonlinear processing on the low-level features. The negative slope α of the Leaky ReLU layer can be set to 0.1. The Leaky ReLU layer helps alleviate the problem of gradient disappearance by setting a small negative slope; the first Max-Pooling2×2 layer is used to halve the spatial dimension of the features output by the activation layer to obtain edge texture features of 128×128×32, thereby reducing computational complexity and preventing overfitting.

[0091] In step S702 of some embodiments, the shape feature processing block is used to obtain shape features. Shape features can be basic geometric structures, contours, etc. The shape feature processing block sequentially includes a second Conv2D layer with 32 filters, a second Leaky ReLU layer, a second Max-Pooling2×2 layer, and a first regularization layer (Dropout layer). The second Conv2D layer is used to extract mid-level features (such as basic shapes) from the 128×128×32 edge texture features; the second Leaky ReLU layer is used to perform nonlinear processing on the mid-level features, with α set to 0.1; the second Max-Pooling2×2 layer is used to halve the spatial dimensions of the features output by the activation layer to obtain 64×64×32 shape features; the first Dropout layer is used to prevent overfitting, and the dropout rate of this first Dropout layer can be set to 0.25, which means that 25% of the neurons are randomly disabled during training.

[0092] In step S703 of some embodiments, the organ boundary feature processing block is used to obtain organ boundary features. Organ boundary features reflect organ boundaries in the image. The organ boundary feature processing block includes, in sequence, a third Conv2D layer with 64 filters, a third Leaky ReLU layer, a fourth Conv2D layer with 64 filters, a fourth Leaky ReLU layer, a third Max-Pooling2×2 layer, and a second Dropout layer. Among them, the third Conv2D layer is used to extract initial high-level anatomical features (such as organ boundaries) from the shape features 64×64×32; the third Leaky ReLU layer is used to perform nonlinear processing on the initial high-level anatomical features, and α can be set to 0.1; the fourth Conv2D layer is used to extract intermediate high-level anatomical features from the initial high-level anatomical features; the fourth Leaky ReLU layer is used to perform nonlinear processing on the intermediate high-level anatomical features, and α can be set to 0.1; the third Max-Pooling2×2 layer is used to halve the spatial dimension of the features output by the activation layer to obtain organ boundary features 32×32×64; the second Dropout layer is used to prevent overfitting, and the dropout rate of the second Dropout layer can be set to 0.25.

[0093] In step S704 of some embodiments, the lesion feature processing block is used to obtain lesion features. Lesion features are high-level visual features that contain information about possible abnormalities or diseases. The lesion feature processing block includes, in sequence, a fifth Conv2D layer with 126 filters, a fifth Leaky ReLU layer, a sixth Conv2D layer with 128 filters, a sixth Leaky ReLU layer, a fourth Max-Pooling2×2 layer, and a third Dropout layer. Among them, the fifth Conv2D layer is used to separate the complex initial specific disease features (such as lesions) from the organ boundary features 32×32×64; the fifth Leaky ReLU layer is used to perform nonlinear processing on the initial specific disease features, and α can be set to 0.1; the sixth Conv2D layer is used to extract intermediate specific disease features from the initial specific disease features; the sixth Leaky ReLU layer is used to perform nonlinear processing on the intermediate specific disease features, and α can be set to 0.1; the fourth Max-Pooling2×2 layer is used to halve the spatial dimension of the features output by the activation layer to obtain lesion features 16×16×128; the third Dropout layer is used to prevent overfitting, and the dropout rate of the third Dropout layer can be set to 0.25.

[0094] In step S705 of some embodiments, a feature flattening block is used to obtain flattened features, thereby effectively transitioning from spatial data to high-level feature representation. Flattening converts lesion features into one-dimensional vectors suitable for input to a fully connected neural network. The feature flattening block sequentially includes a flattening layer (Flatten layer), a first dense layer (Dense128 layer), a seventh LeakyReLU layer, a first batch normalization layer (Batch Normalization layer), a fourth Dropout layer, a second dense layer (Dense64 layer), an eighth LeakyReLU layer, a second Batch Normalization layer, and a fifth Dropout layer. Among them, the Flatten layer is used to flatten the lesion feature 16×16×128 into a one-dimensional vector 1×32768; the Dense128 layer has 128 neurons, which is used to compress the one-dimensional vector 1×32768 into 128 dimensions; the seventh Leaky ReLU layer is used to perform nonlinear processing on the 128-dimensional output of the Dense128 layer, and α can be set to 0.1; the first Batch Normalization layer is used to standardize the features output by the activation layer, helping to accelerate training and improve stability by reducing internal covariate shift. The momentum M of the moving average can be set to 0.99, and the constant Ep to prevent zero division errors can be set to 0.01; the fourth Dropout layer is used to prevent overfitting, and the dropout rate of the fourth Dropout layer can be set to 0.25; the Dense64 layer has 64 neurons, which is used to compress the 128-dimensional vector into 64 dimensions; the eighth Leaky ReLU layer is used to perform nonlinear processing on the 64-dimensional output of the Dense64 layer, and α can be set to 0.1; the second Batch The Normalization layer is used to normalize the features output by the activation layer to help accelerate training and improve stability by reducing internal covariate shift. The momentum M of the moving average can be set to 0.99, and the constant Ep to prevent division by zero errors can be set to 0.01. The fifth Dropout layer is used to prevent overfitting, and the dropout rate of the fifth Dropout layer can be set to 0.5.

[0095] In step S706 of some embodiments, the student image prediction data is the class probability distribution obtained after the initial student model performs multi-disease prediction on the sample medical image, representing the initial student model's prediction of the likelihood of each disease. This is achieved by flattening features through the classification layer. This student image prediction data is similar to the teacher prediction data.

[0096] In steps S701 to S706, a lightweight and optimized initial student model is constructed for classification prediction, resulting in more efficient student image prediction data. This model is called lightweight and optimized because it strikes a good balance between network depth and complexity: by gradually increasing the number of filters in the convolutional layers while employing max pooling to reduce spatial dimensionality and computational cost. To prevent overfitting and stabilize the training process, this architecture applies techniques such as dropout and batch normalization to ensure efficient learning. The use of the Leaky ReLU activation function enhances the network's nonlinear representation capabilities without incurring a significant computational burden. The two smaller fully connected layers (with 128 and 64 neurons, respectively) further reduce model complexity, improving computational efficiency while maintaining good representation of the complex characteristics of various medical diseases. In contrast, more complex models may face the following challenges: excessive consumption of computing resources, prolonged training time, increased risk of overfitting, and reduced generalization, ultimately leading to reduced performance and efficiency.

[0097] In step S106 of some embodiments, loss data is calculated based on the student image prediction data and the teacher image prediction data. This loss data is used to measure the difference between the initial student model prediction data and the target teacher model prediction data, and is typically calculated using a knowledge distillation loss function.

[0098] In step S107 of some embodiments, weight adjustment refers to using the target trust weight to adjust the degree of influence of the loss data on the total loss data. Taking the total loss data as the target, the back propagation algorithm is used to calculate the gradient of each layer of the initial student model, and the optimizer is used to adjust the parameters of each layer in the initial student model. Through iterative training, the parameters of the convolutional layer, the fully connected layer, etc. in the model structure are continuously optimized so that it can learn to better express the representation of medical image features such as edges, shapes, organ boundaries, and lesions. After the training is completed, the obtained student model is the target student model, which has the ability to classify diseases in real medical image scenes and has the ability to adapt to the output quality of the target teacher model under weight control, thereby improving the stability and generalization of the model. The calculation formula of the total loss data is shown in formula (11).

[0099] , formula (11).

[0100] in, Indicates total loss data; It is a loss function that measures the difference between the initial student model prediction data and the target teacher model prediction data, also known as soft loss, usually knowledge distillation loss; It is a loss function that measures the difference between the initial student model prediction data and the true label (Ground Truth) of the sample medical image, also known as hard loss, usually cross entropy loss; is a hyperparameter used to balance and importance.

[0101] It is important to note that the target trust weight is used to balance the contributions of the soft loss (teacher-student prediction alignment) and the hard loss (true label error) to the overall training objective. During backpropagation, this target trust weight adjusts the gradients of the edge texture feature processing block, shape feature processing block, organ boundary feature processing block, and lesion feature processing block in the student model, optimizing the filters of these processing blocks to mimic the teacher's high-confidence features. This also enables the student to use labeled data to correct for uncertainty in medical images. For example, if the teacher's confidence is low (uncertainty is high), the target trust weight is reduced, forcing the student's 2D convolutional layers to rely more on the true labels to learn robust and noise-resistant features. Simultaneously, the batch normalization and regularization layers in the feature flattening block stabilize the training process and ensure good generalization of the adaptively weighted features. This fuzzy-guided interaction enables the student to learn discriminative hierarchical representations from edges (edge ​​texture features) to complex pathologies (lesion features) while maintaining computational efficiency, making it well-suited for deployment in resource-constrained medical imaging environments.

[0102] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned medical image classification method or pathological image processing method. The electronic device can be any intelligent terminal, such as a tablet computer or an in-vehicle computer.

[0103] See also Figure 8, an embodiment of the present application also provides a medical image classification system that can implement the above-mentioned medical image classification method, the system including: a data acquisition module 801, used to acquire sample medical images; a teacher model prediction module 802, used to perform a first image prediction on the sample medical image through a target teacher model to obtain teacher image prediction data; an indicator acquisition module 803, used to obtain confidence from the teacher image prediction data, and obtain uncertainty based on the confidence; a weight acquisition module 804, used to perform a trust evaluation on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight; a student model prediction module 805, used to perform a second image prediction on the sample medical image through a pre-built initial student model to obtain student image prediction data; a loss calculation module 806, used to perform loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; a back-propagation training module 807, used to adjust the weight of the loss data based on the target trust weight, thereby back-propagating the initial student model to obtain a target student model for medical image classification.

[0104] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned medical image classification method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0105] See also Figure 9 , Figure 9The hardware structure of an electronic device in another embodiment is illustrated. The electronic device includes: a processor 901, which can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; a memory 902, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the medical image classification method or pathological image processing method of the embodiments of this application; the input / output interface 903 is used to implement information input and output; the communication interface 904 is used to implement communication interaction between this device and other devices, and communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); the bus 905 transmits information between the various components of the device (such as the processor 901, memory 902, input / output interface 903 and communication interface 904); wherein the processor 901, memory 902, input / output interface 903 and communication interface 904 are connected to each other within the device through the bus 905.

[0106] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned medical image classification method or pathological image processing method. The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0107] The medical image classification method, electronic device, and storage medium provided in the embodiments of the present application obtain sample medical images; perform a first image prediction on the sample medical images using a target teacher model to obtain teacher image prediction data; obtain confidence from the teacher image prediction data, and obtain uncertainty based on the confidence, which helps to quantify the target teacher model's confidence level in the prediction results and provide a basis for trust assessment. Furthermore, a trust assessment is performed on the teacher image prediction data based on the confidence and uncertainty to obtain a target trust weight, which helps to identify and utilize highly reliable prediction results, thereby improving the training effect of the student model. Furthermore, a second image prediction is performed on the sample medical image through the pre-built initial student model to obtain student image prediction data; loss calculation is performed based on the student image prediction data and the teacher image prediction data to obtain loss data; the loss data is weighted according to the target trust weight, and the initial student model is trained by backpropagation to obtain a target student model for medical image classification. This step can dynamically adjust the degree of dependence of the student model on the target teacher model based on the uncertainty of the medical image and using the target trust weight, so that the student model can adapt to different uncertainty levels of different medical images, so that the student model focuses on high confidence areas and ignores low confidence areas, thereby improving the classification accuracy of medical images.

[0108] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. It will be appreciated by those skilled in the art that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems. It will be appreciated by those skilled in the art that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application and may include more or fewer steps than shown, or a combination of certain steps, or different steps. The system embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solutions of the present embodiment. It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the above-disclosed methods may be implemented as software, firmware, hardware, and appropriate combinations thereof. The terms "first," "second," "third," "fourth," etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. It should be understood that in this application, "at least one (item)" refers to one or more, and "a plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three types of relationships can exist. For example, "A and / or B" can represent: only A exists, only B exists, and A and B exist simultaneously, wherein A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are a kind of "or" relationship. "At least one of the following (individual)" or its similar expressions refer to any combination in these items, including any combination of singular (individual) or plural (individual). For example, at least one of a, b, or c can represent: a, b, c, "a and b," "a and c," "b and c," or "a and b and c," where a, b, and c can be single or multiple. It should be understood that the disclosed systems and methods can be implemented in other ways in the several embodiments provided herein.For example, the system embodiments described above are merely illustrative. For example, the division of units described above represents only a logical functional division. In actual implementation, alternative divisions may be employed. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented. The coupling or direct coupling or communication connection shown or discussed may be through interfaces, or indirect coupling or communication connection between systems or units, and may be electrical, mechanical, or other. The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in a single location or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of the present embodiments as needed. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. These integrated units may be implemented in either hardware or software functional units. If the integrated units are implemented as software functional units and sold or used as standalone products, they may be stored on a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, but the scope of the rights of the embodiments of the present application is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application should be within the scope of the rights of the embodiments of the present application.

Claims

1. A medical image classification method, characterized in that: The method comprises: Obtaining sample medical images; Performing a first image prediction on the sample medical image using a target teacher model to obtain teacher image prediction data; Obtaining a confidence level from the teacher image prediction data, and obtaining an uncertainty level based on the confidence level; Performing a level membership evaluation on the confidence level according to a preset number of confidence level membership functions to obtain a confidence level membership for each confidence level; wherein the confidence levels of any two of the confidence level membership functions are different, and the preset number is greater than or equal to 2; Performing a grade membership evaluation on the uncertainty according to the preset number of uncertainty membership functions to obtain an uncertainty grade membership degree for each uncertainty grade; wherein the uncertainty grades of any two of the uncertainty membership functions are different; Performing membership fusion according to the confidence level membership of each of the confidence levels and the uncertainty level membership of each of the uncertainty levels to obtain a target trust weight; Performing a second image prediction on the sample medical image using a pre-built initial student model to obtain student image prediction data; Perform loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; The loss data is weighted according to the target trust weight, thereby backpropagating and training the initial student model to obtain a target student model for medical image classification.

2. The method according to claim 1, characterized in that The performing membership fusion according to the confidence level membership of each of the confidence levels and the uncertainty level membership of each of the uncertainty levels to obtain the target trust weight includes: Orderly combining the confidence levels and the uncertainty levels to obtain level combinations, and determining a combination weight for each of the level combinations; performing minimum selection on the confidence level membership and the uncertainty level membership in each of the level combinations to obtain a fuzzy membership of each of the level combinations; A weight calculation is performed according to the combination weight of each of the level combinations and the fuzzy membership of the level combinations to obtain the target trust weight.

3. The method according to any one of claims 1 to 2, characterized in that The initial student model includes an edge texture feature processing block, a shape feature processing block, an organ boundary feature processing block, a lesion feature processing block, a feature flattening block, and a classification layer. The second image prediction is performed on the sample medical image using the pre-built initial student model to obtain student image prediction data, including: Performing basic feature processing on the sample medical image by the edge texture feature processing block to obtain edge texture features; Performing light feature processing on the edge texture feature by the shape feature processing block to obtain a shape feature; Performing complex feature processing on the shape feature by the organ boundary feature processing block to obtain an organ boundary feature; Performing deep feature processing on the organ boundary feature by the lesion feature processing block to obtain lesion features; Performing feature processing on the lesion feature by the feature flattening block to obtain a flattened feature; The flattened features are classified by the classification layer to obtain the student image prediction data.

4. The method according to claim 1, wherein Before performing the first image prediction on the sample medical image by the target teacher model, the sample medical image is also updated, specifically including: Adjusting the resolution of the sample medical image to obtain a standard sample medical image; Performing correction processing on the standard sample medical image to obtain a first image; Performing histogram equalization on the standard sample medical image to obtain a second image; Decomposing the first image and the second image using a predetermined function to obtain a first multi-scale frequency component corresponding to the first image and a second multi-scale frequency component corresponding to the second image; fusing the first multi-scale frequency component and the second multi-scale frequency component to obtain a fused frequency component; Performing image reconstruction on the fused frequency components by inverse wavelet transform to obtain a reconstructed sample medical image; Linear mapping is performed on the pixel values ​​of the reconstructed sample medical image to obtain an updated sample medical image.

5. The method according to claim 1, wherein The target teacher model is obtained in the following way: Obtain multiple candidate teacher models; Obtain medical image collections for multiple diseases; Calculating an index for each candidate teacher model based on the multi-disease medical image set to obtain a performance index corresponding to each candidate teacher model; wherein the performance index includes at least a verification accuracy rate and a model calculation cost; Initializing a pheromone matrix and a heuristic matrix; wherein the pheromone values ​​corresponding to each candidate teacher model in the pheromone matrix are the same, and the heuristic matrix is ​​composed of multiple element values, each of which is generated based on the performance indicator of each candidate teacher model; According to the set ant colony parameters, multiple rounds of iterative optimization are performed based on the initialized pheromone matrix and the heuristic matrix. After reaching the preset number of iterations, the highest target element value is screened out from the highest element values ​​corresponding to each round of iteration, and the model corresponding to the target element value is used as the target teacher model.

6. The method according to claim 5, characterized in that The ant colony parameters include a target number of ants. The multiple rounds of iterative optimization based on the set ant colony parameters and the initialized pheromone matrix and the heuristic matrix include: In each iteration: The target number of ants calculates the selection probability of each candidate teacher model based on the pheromone matrix and the heuristic matrix corresponding to this iteration, and selects the target model corresponding to each ant using a roulette wheel selection method; wherein the probability of the roulette wheel selection method is proportional to the selection probability; Record the highest element value among the element values ​​corresponding to each target model in this round of iteration; The pheromone matrix corresponding to this round of iteration is updated according to the selection results of each ant pair model and the element values ​​of each target model.

7. A medical image classification system, characterized in that: The system comprises: A data acquisition module, used to acquire sample medical images; a teacher model prediction module, configured to perform a first image prediction on the sample medical image using a target teacher model to obtain teacher image prediction data; An indicator acquisition module, configured to obtain a confidence level from the teacher image prediction data and obtain uncertainty based on the confidence level; A weight acquisition module, configured to perform a grade membership evaluation on the confidence level according to a preset number of confidence membership functions to obtain a confidence grade membership for each confidence grade; wherein the confidence grades of any two of the confidence membership functions are different, and the preset number is greater than or equal to 2; perform a grade membership evaluation on the uncertainty level according to the preset number of uncertainty membership functions to obtain an uncertainty grade membership for each uncertainty grade; wherein the uncertainty grades of any two of the uncertainty membership functions are different; and perform membership fusion based on the confidence grade membership of each of the confidence grades and the uncertainty grade membership of each of the uncertainty grades to obtain a target trust weight; a student model prediction module, configured to perform a second image prediction on the sample medical image using a pre-built initial student model to obtain student image prediction data; a loss calculation module, configured to perform loss calculation based on the student image prediction data and the teacher image prediction data to obtain loss data; A reverse training module is used to adjust the weight of the loss data according to the target trust weight, thereby backpropagating the training of the initial student model to obtain a target student model for medical image classification.

8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image processing method and device

    CN113160230A

  • Lightweight image classification neural network architecture system based on knowledge distillation

    CN119514594A