Diabetic nephropathy retinopathy recognition method and device based on double distillation
By adopting the dual distillation method in the recognition of retinopathy of diabetic nephropathy, combining pre-trained models and category adaptive attention model, the problem of small samples and category imbalance is solved, significantly improving the recognition performance, and assisting doctors in accurate diagnosis.
Patent Information
- Application Number
- CN202510290456.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has poor performance in identifying retinal diabetic nephropathy in diabetic patients, especially in small sample conditions, where there are problems of strong subjectivity and low efficiency.
Using a double distillation-based recognition method, by acquiring optical coherence tomography images and performing normalization, a pre-trained 3D-ResNet network and a dual distillation framework, including two ResNet18 network models and a category adaptive attention model, is used to carry out knowledge transfer and loss function optimization, and improve the discrimination ability and feature learning ability of the recognition model.
Effectively dealing with the problem of small samples and category imbalances has significantly improved the recognition performance of diabetic nephropathy and assisted doctors in more accurate diagnosis and treatment.
Smart Images

Figure CN120219915A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for identifying diabetic nephropathy retinopathy based on double distillation, belonging to the cross technical field of artificial intelligence and medical image processing. Background Art
[0002] Diabetes is a chronic metabolic disease. Long-term hyperglycemia can lead to various complications. Among them, diabetic nephropathy and diabetic retinopathy are two common and serious complications. Early detection and intervention are crucial for delaying the progression of diabetic nephropathy. The retina is the only part of the human body where blood vessels and nerve tissues can be directly observed, and there is a close association with overall health. Research shows that through retinal examination, not only can eye diseases be diagnosed, but also overall health problems can be reflected, especially diseases related to blood vessels, metabolism, and the nervous system.
[0003] Traditional diagnostic methods mainly rely on the experience of clinicians and manual examinations, suffering from strong subjectivity and low efficiency. With the wide application of artificial intelligence in medical image analysis, due to the lack of data privacy protection and high-quality labeled data, existing deep learning-based methods perform poorly in identifying diabetic retinopathy in diabetic patients.
[0004] Therefore, it is necessary to design a double-distillation method for identifying diabetic retinopathy with nephropathy by means of the idea of knowledge transfer, which takes into account influencing factors such as lesion size and spatial position relationship, and constructs a double knowledge transfer model to improve the discriminative ability and feature learning ability of the classifier, so as to better handle the identification of lesions under small sample conditions. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a method and device for identifying diabetic nephropathy retinopathy based on double distillation, which can effectively handle the problems of small samples and class imbalance, thereby improving the identification performance of diabetic nephropathy retinopathy.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] In a first aspect, a method for identifying diabetic nephropathy retinopathy based on double distillation provided by an embodiment of the present invention includes the following steps:
[0008] Step 1, obtaining an optical coherence tomography image of a patient;
[0009] Step 2, performing normalization processing on the optical coherence tomography image;
[0010] Step 3, pre-training a 3D-ResNet network using a publicly available large medical image dataset to obtain a pre-trained model;
[0011] Step 4: Input the preprocessed optical coherence tomography (OCT) images into a dual-distillation framework to obtain a dual-distillation model. The dual-distillation framework includes two ResNet18 network models, one focusing on the categories with less data volume and the other paying attention to the category-adaptive attention model.
[0012] Step 5: Combine the loss functions of the pre-trained model and the dual-distillation model, and distill the knowledge of the pre-trained model into the dual-network model to obtain the final retinopathy recognition model.
[0013] Step 6: Use the retinopathy recognition model to identify diabetic nephropathy retinopathy.
[0014] As a possible implementation of this embodiment, Step 1 includes the following steps:
[0015] Use an optical coherence tomography imaging device to scan and image a patient to obtain an optical coherence tomography image I, where I ∈ R H×W×D , R represents the entire image area, H and W are the length and width of the image respectively, and D is the depth of the image.
[0016] As a possible implementation of this embodiment, Step 2 includes the following steps:
[0017] Perform normalization on the optical coherence tomography image I ∈ R H×W×D :
[0018]
[0019] where I(h, w, d) represents the pixel value of the optical coherence tomography image at the coordinate (h, w, d), h ∈ [0, H), w ∈ [0, W), d ∈ [0, D), I max and I min represent the maximum and minimum values among all pixel values in the optical coherence tomography image respectively.
[0020] As a possible implementation of this embodiment, Step 4 includes the following steps:
[0021] By designing a class balance loss function, adjust the weights of the identified lesion samples of the ResNet18 network model focusing on the categories with less data volume;
[0022] By constructing a category-adaptive attention model, dynamically balance the sample weights of the category-adaptive attention model.
[0023] As a possible implementation of this embodiment, the class balance loss function is:
[0024] L fl= -(1 - p t ) γ log(p t )
[0025] where p t is the probability predicted by the model, (1 - p t ) γ is the adjustment factor, and γ is the balance parameter.
[0026] As a possible implementation of this embodiment, the class - adaptive attention model obtains the sensitivity score for each class through 3D 1×1 convolution and global average pooling:
[0027]
[0028] where f' i,j is the input image feature, GP is the global average pooling operation, k represents the number of feature map channels, i ∈ {1, 2, …, L}, and L is the number of classes;
[0029] The class - adaptive attention model performs class - wise channel average pooling on the feature map after 3D 1×1 convolution to obtain the feature map F”:
[0030]
[0031] The class - adaptive attention model combines the attention region and the original image to enhance the input feature image and discriminates the retinal lesion region CAM_OUT:
[0032]
[0033] where F is the input feature map.
[0034] As a possible implementation of this embodiment, step 5 includes the following steps:
[0035] The class - balance loss function is expressed as:
[0036] L fl = -(1 - p t ) γ log(p t )
[0037] where p t is the probability predicted by the model, (1 - p t ) γ is the adjustment factor, and γ is the balance parameter;
[0038] The cross - entropy loss function is expressed as:
[0039] L CE = -log(pt)
[0040] Among them, p t is the probability predicted by the model;
[0041] The knowledge transfer loss function is expressed as:
[0042]
[0043] Among them, P i is the probability predicted by the double-distillation model, Q i is the probability of the pre-trained network model, and Y i is the true label distribution;
[0044] Use the Adam optimization algorithm to optimize and train the loss function, distill the knowledge of the pre-trained model into the double network model, and obtain the final diabetic retinopathy recognition model.
[0045] In a second aspect, a diabetic retinopathy recognition device based on double distillation provided by an embodiment of the present invention includes:
[0046] An image acquisition module for acquiring an optical coherence tomography image of a patient;
[0047] An image processing module for normalizing the optical coherence tomography image;
[0048] A pre-trained model acquisition module for pre-training the 3D-ResNet network using a publicly available large medical image dataset to obtain a pre-trained model;
[0049] A double-distillation model acquisition module for inputting the pre-processed optical coherence tomography image into a double-distillation framework to obtain a double-distillation model, where the double-distillation framework includes two ResNet18 network models, one focusing on categories with less data volume and the other focusing on the category adaptive attention model;
[0050] A loss function joint module for jointly using the loss functions of the pre-trained model and the double-distillation model, distilling the knowledge of the pre-trained model into the double network model, and obtaining the final diabetic retinopathy recognition model;
[0051] A lesion recognition module for recognizing diabetic retinopathy using the diabetic retinopathy recognition model.
[0052] In a third aspect, an electronic device provided by an embodiment of the present invention includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of any of the above-mentioned diabetic nephropathy retinopathy recognition methods based on double distillation.
[0053] In a fourth aspect, a storage medium provided by an embodiment of the present invention stores a computer program. When the computer program is run by a processor, it performs the steps of any of the above-mentioned diabetic nephropathy retinopathy recognition methods based on double distillation.
[0054] The beneficial effects of the technical solutions of the embodiments of the present invention are as follows:
[0055] A diabetic nephropathy retinopathy recognition method based on double distillation according to the technical solution of the embodiment of the present invention includes the following steps: Step 1, obtain the optical coherence tomography image of a patient; Step 2, perform normalization processing on the optical coherence tomography image; Step 3, pre-train the 3D-ResNet network using a publicly available large medical image dataset to obtain a pre-trained model; Step 4, input the pre-processed optical coherence tomography image into a double distillation framework to obtain a double distillation model, where the double distillation framework includes two ResNet18 network models, one focusing on the category with less data volume, and the other focusing on the category adaptive attention model; Step 5, combine the loss functions of the pre-trained model and the double distillation model, and distill the knowledge of the pre-trained model into the double network model to obtain a final retinopathy recognition model; Step 6, use the retinopathy recognition model to recognize diabetic nephropathy retinopathy. The present invention effectively addresses the problems of small samples and class imbalance. By training a pre-trained model on a large medical image dataset and transferring the knowledge of the pre-trained model to a small model, and at the same time using a lesion-aware attention mechanism to improve the attention degree of the recognition model to different diseases, the recognition performance of diabetic nephropathy retinopathy is improved.
[0056] The present invention significantly improves the detection performance of retinopathy under small sample conditions, thereby assisting doctors in making more accurate diagnoses and treatments. By constructing a category-aware learning model, the present invention improves the attention degree of the model to different lesion regions, and on this basis, by constructing a teacher learning model and a student learning model, a double distillation detection framework is designed, showing excellent performance and theoretical advantages in the recognition of diabetic nephropathy retinopathy.
[0057] The present invention also proposes a deep learning model based on double distillation, which is dedicated to retinopathy recognition. Thanks to the class attention mechanism and the knowledge transfer model, the present invention significantly improves the effect of diabetic nephropathy retinopathy recognition in clinical application scenarios.
[0058] A diabetic nephropathy retinopathy recognition device with double distillation according to the technical solution of the embodiment of the present invention has the same beneficial effects as a diabetic nephropathy retinopathy recognition method with double distillation according to the technical solution of the embodiment of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flowchart of a diabetic nephropathy retinopathy recognition method based on double distillation shown according to an exemplary embodiment;
[0060] Figure 2 is a schematic structural diagram of a diabetic nephropathy retinopathy recognition device based on double distillation shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To more clearly illustrate the technical features of the solution of the present invention, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its drawings.
[0062] As Figure 1 shown, a diabetic nephropathy retinopathy recognition method based on double distillation provided by an embodiment of the present invention includes the following steps:
[0063] Step 1, obtaining an optical coherence tomography image of a patient;
[0064] Step 2, performing normalization processing on the optical coherence tomography image;
[0065] Step 3, pre-training a 3D-ResNet network using a publicly available large medical image dataset to obtain a pre-trained model;
[0066] Step 4, inputting the pre-processed optical coherence tomography image into a double distillation framework to obtain a double distillation model, where the double distillation framework includes two ResNet18 network models, one focusing on the category with less data volume and the other focusing on the category adaptive attention model;
[0067] Step 5, combining the loss functions of the pre-trained model and the double distillation model, and distilling the knowledge of the pre-trained model into the double network model to obtain a final retinopathy recognition model;
[0068] Step 6, using the retinopathy recognition model to recognize diabetic nephropathy retinopathy.
[0069] As a possible implementation of this embodiment, step 1 includes the following steps:
[0070] Use an optical coherence tomography imaging device to scan and image the patient to obtain an optical coherence tomography image I, where I ∈ R H×W×D , R represents the entire image area, H and W are the length and width of the image respectively, and D is the depth of the image.
[0071] As a possible implementation of this embodiment, step 2 includes the following steps:
[0072] Normalize the optical coherence tomography image I ∈ R H×W×D :
[0073]
[0074] where I(h, w, d) represents the pixel value of the optical coherence tomography image at the coordinate (h, w, d), h ∈ [0, H), w ∈ [0, W), d ∈ [0, D), I max and I min represent the maximum and minimum values among all pixel values in the optical coherence tomography image respectively.
[0075] As a possible implementation of this embodiment, step 4 includes the following steps:
[0076] By designing a class balance loss function, adjust the recognition lesion sample weights of the ResNet18 network model focused on the categories with less data volume;
[0077] By constructing a class adaptive attention model, dynamically balance the sample weights of the class adaptive attention model.
[0078] As a possible implementation of this embodiment, the class balance loss function is:
[0079] L fl = -(1 - p t ) γ log(p t )
[0080] where p t is the probability predicted by the model, (1 - p t ) γ is the adjustment factor, and γ is the balance parameter.
[0081] As a possible implementation of this embodiment, the class adaptive attention model obtains the sensitivity score of each category through three-dimensional 1×1 convolution and global average pooling:
[0082]
[0083] Among them, f' i,j is the input image feature, GP is the global average pooling operation, k represents the number of feature map channels, i ∈ {1, 2, …, L}, and L is the number of categories;
[0084] The class - adaptive attention model performs class - wise channel average pooling on the feature map after 3D 1×1 convolution to obtain the feature map F":
[0085]
[0086] The class - adaptive attention model combines the attention region and the original image to enhance the input feature image and discriminates the retinal lesion region CAM_OUT:
[0087]
[0088] Among them, F is the input feature map.
[0089] As a possible implementation manner of this embodiment, step 5 includes the following steps:
[0090] The class - balanced loss function is expressed as:
[0091] L fl = -(1 - p t ) γ log(p t )
[0092] Among them, p t is the probability predicted by the model, (1 - p t ) γ is the adjustment factor, and γ is the balance parameter;
[0093] The cross - entropy loss function is expressed as:
[0094] L CE = -log(pt)
[0095] Among them, p t is the probability predicted by the model;
[0096] The knowledge transfer loss function is expressed as:
[0097]
[0098] Among them, P i is the probability predicted by the dual - distillation model, Q i is the probability of the pre - trained network model, and Y i is the true label distribution;
[0099] The loss function is optimized and trained using the Adam optimization algorithm, and the knowledge of the pre-trained model is distilled into the dual network model to obtain the final diabetic retinopathy recognition model.
[0100] As Figure 2 shown, an apparatus for recognizing diabetic retinopathy based on dual distillation provided by an embodiment of the present invention includes:
[0101] An image acquisition module, configured to acquire optical coherence tomography images of a patient;
[0102] An image processing module, configured to perform normalization processing on the optical coherence tomography images;
[0103] A pre-trained model acquisition module, configured to pre-train a 3D-ResNet network using a publicly available large medical image dataset to obtain a pre-trained model;
[0104] A dual distillation model acquisition module, configured to input the pre-processed optical coherence tomography images into a dual distillation framework to obtain a dual distillation model, where the dual distillation framework includes two ResNet18 network models, one focusing on categories with less data volume and the other focusing on a category adaptive attention model;
[0105] A loss function joint module, configured to jointly optimize the loss functions of the pre-trained model and the dual distillation model, and distill the knowledge of the pre-trained model into the dual network model to obtain the final diabetic retinopathy recognition model;
[0106] A lesion recognition module, configured to recognize diabetic retinopathy using the diabetic retinopathy recognition model.
[0107] As a possible implementation manner of this embodiment, the image acquisition module uses an optical coherence tomography imaging device to perform computed tomography imaging on a patient to obtain optical coherence tomography images. The image processing module performs normalization processing on the optical coherence tomography retinal images.
[0108] The specific process of recognizing diabetic retinopathy in the present invention includes the following steps.
[0109] Step 1, obtain relevant information such as optical tomography images.
[0110] Use an optical coherence tomography imaging device to scan and image a patient to obtain an optical coherence tomography image, which is defined as I ∈ R H×W×D , where R represents the entire image area, H and W are the length and width of the image respectively, and D is the depth of the image.
[0111] Step 2, perform data preprocessing operations on the images in Step 1.
[0112] Normalize the optical coherence tomography image \(I\in\mathbb{R}\) H×W×D as follows:
[0113]
[0114] where \(I(h, w, d)\) represents the pixel value of the optical coherence tomography image at the coordinate \((h, w, d)\), \(h\in[0, H)\), \(w\in[0, W)\), \(d\in[0, D)\). \(I\) max and \(I\) min represent the maximum and minimum values among all pixel values in the optical coherence tomography image, respectively.
[0115] Step 3: Obtain a robust pre-trained model using a large amount of medical images.
[0116] Pre-train 3D-ResNet using a publicly available large medical image dataset. The pre-trained model can better capture the structural and semantic features in medical images and retain the pre-trained model parameters.
[0117] Step 4: Input the obtained image data into the dual distillation framework.
[0118] The dual distillation framework consists of two ResNet18 networks, which respectively focus on the classes with less data and the class-adaptive attention model.
[0119] The first model adjusts the weights of difficult-to-recognize lesion samples by designing a class balance loss function and focuses on samples with fewer classes. The specific loss function is as follows:
[0120] \(L\) fl \(=-(1 - p\) t ) γ \(\log(p\) t )
[0121] where \(p\) t is the probability predicted by the model, \((1 - p\) t ) γ is the adjustment factor, and \(\gamma\) is the balance parameter.
[0122] The second model dynamically balances the sample weights by constructing a class-adaptive attention model, improving the sensitivity of the model to samples.
[0123] The second model first obtains the sensitivity scores for each class through three-dimensional 1×1 convolution and global average pooling, as follows:
[0124]
[0125] where \(f'\) i,jLet \(F\) be the input feature map, \(GP\) be the global average pooling operation, \(k\) represent the number of channels of the feature map, \(i\in\{1,2,\ldots,L\}\), and \(L\) be the number of categories;
[0126] Secondly, perform category - channel average pooling on the feature map after 3D \(1\times1\) convolution to obtain a feature map, as follows:
[0127]
[0128] Finally, combine the attention region and the original image to enhance the input feature image, so as to better distinguish the retinal lesion area, as follows:
[0129]
[0130] where \(F\) is the input feature map.
[0131] Step 5: Optimize the model jointly through the distillation loss function and the cross - entropy loss function.
[0132] The class - balance loss function is expressed as:
[0133] \(L\) fl \(=-(1 - p\) t ) γ \(\log(p\) t )
[0134] where \(p\) t is the probability predicted by the model, \((1 - p\) t ) γ is the adjustment factor, and \(\gamma\) is the balance parameter;
[0135] The cross - entropy loss function is expressed as:
[0136] \(L\) CE \(=-\log(p_t)\)
[0137] where \(p\) t is the probability predicted by the model;
[0138] The knowledge transfer loss function is expressed as:
[0139]
[0140] where \(P\) i is the probability predicted by the double - distillation model, \(Q\) i is the probability of the pre - trained network model, and \(Y\) i is the true label distribution;
[0141] Use the Adam optimization algorithm to optimize and train the loss function to obtain a trained retinal lesion recognition model.
[0142] Step 6: Use the trained retinal lesion recognition model to assist in diagnosing the disease.
[0143] The trained double distillation network model is used to identify new data to obtain the recognition result.
[0144] Taking the optical coherence tomography retinal image as input, the retinal lesion recognition method of the present invention is used for disease detection. During training, the registered low-dose CT images and conventional-dose CT images are used. First, the low-dose CT image data is acquired and preprocessed; then, the preprocessed image is input into the prior knowledge extraction module to generate the corresponding prior mask image. Next, the prior mask is discretized using a discrete coding network, and a negative sample set is constructed based on this. At the same time, a joint loss function is designed to enhance the ability of the denoising network in retaining key boundary information. Finally, the prior mask of the test image is introduced into the denoising network through the knowledge fusion module. After multiple iterative processes, the denoised image is finally output.
[0145] The present invention proposes a controllable diffusion model based on SAM prior knowledge, which is specially used for low-dose CT denoising; its basic process includes: firstly, obtaining and preprocessing low-dose CT image data, then inputting it into the prior knowledge extraction module to generate the corresponding prior mask image; then, discretizing the prior mask through the discrete coding network, and on this basis, constructing a negative sample set and designing a joint loss function to enhance the denoising network's ability to retain key boundary information; finally, the prior mask of the test image is introduced into the denoising network through the knowledge fusion module, and after multiple iterations, the denoised image is finally output. Thanks to the prior knowledge provided by SAM, the present invention significantly improves the effect of low-dose CT denoising in clinical application scenarios, and the generated image boundaries are clearer and the noise is easier to suppress. By implementing the technical solution of the present invention, the denoising performance of low-dose CT images can be significantly improved, and clearer tissue boundaries can be retained, thereby assisting doctors in more accurate diagnosis and treatment.
[0146] An embodiment of the present invention provides an electronic device, including a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform any step of the above-mentioned diabetic nephropathy retinopathy identification method based on double distillation.
[0147] Specifically, the above-mentioned memory and processor can be general-purpose memory and processor, which are not specifically limited here. When the processor runs the computer program stored in the memory, the above-mentioned diabetic nephropathy retinopathy identification method based on double distillation can be executed.
[0148] Those skilled in the art can understand that the structure of the computer device does not constitute a limitation on the computer device, and it may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements.
[0149] In some embodiments, the computer device may further include a touch screen, which can be used to display a graphical user interface (e.g., the startup interface of an application) and receive user operations on the graphical user interface (e.g., the startup operation for an application). Specifically, the touch screen may include a display panel and a touch panel. Among them, the display panel can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. The touch panel can collect contact or non-contact operations of the user on or near it and generate preset operation instructions. For example, the user uses any suitable object such as a finger, a stylus, or an accessory to operate on or near the touch panel. In addition, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation and posture of the user, and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into information that can be processed by the processor, then sends it to the processor, and can receive and execute the commands sent by the processor. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel, and any technology developed in the future can also be used to implement the touch panel. Further, the touch panel can cover the display panel, and the user can operate on or near the touch panel covering the display panel according to the graphical user interface displayed on the display panel. After the touch panel detects the operation on or near it, it transmits it to the processor to determine the user input, and then the processor provides a corresponding visual output on the display panel in response to the user input. In addition, the touch panel and the display panel can be implemented as two independent components or integrated.
[0150] Corresponding to the above application startup method, an embodiment of the present invention further provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of any of the above-mentioned dual-distillation-based diabetic nephropathy retinopathy recognition methods.
[0151] The startup device of the application program provided by the embodiments of the present application can be specific hardware on the device, or software or firmware installed on the device, etc. For the device provided by the embodiments of the present application, its implementation principle and the resulting technical effects are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing described systems, devices, and units can all refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0153] In the embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of devices or modules can be in an electrical, mechanical, or other form.
[0154] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, the functional modules in the embodiments provided by the present application can be integrated into one processing module, or each module exists physically alone, or two or more modules are integrated into one module.
[0156] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for identifying diabetic nephropathy retinopathy based on double distillation, characterized in that: The steps include: Step 1, obtaining an optical coherence tomography image of a patient; Step 2, normalizing the optical coherence tomography image; Step 3, pre-training the 3D-ResNet network using a large public medical image dataset to obtain a pre-trained model; Step 4, inputting the preprocessed optical coherence tomography image into a double distillation framework to obtain a double distillation model, wherein the double distillation framework includes two ResNet18 network models, one focusing on categories with less data volume, and the other focusing on a category adaptive attention model; Step 5, combining the loss functions of the pre-trained model and the double distillation model, distilling the knowledge of the pre-trained model into the double network model, and obtaining the final retinal lesion recognition model; Step 6: Use the retinopathy recognition model to identify diabetic nephropathy retinopathy.
2. The method for identifying diabetic nephropathy retinopathy based on double distillation according to claim 1, characterized in that: The step 1 comprises the following steps: Use an optical coherence tomography imaging device to scan and image the patient to obtain an optical coherence tomography image I, where I∈R H×W×D , R represents the entire image area, H and W are the length and width of the image respectively, and D is the depth of the image.
3. The method for identifying diabetic nephropathy retinopathy based on double distillation according to claim 2, characterized in that: The step 2 comprises the following steps: For the optical coherence tomography image I∈R H×W×D Perform normalization: Where I(h,w,d) represents the pixel value of the optical coherence tomography image at the coordinate (h,w,d), h∈[0,H), w∈[0,W), d∈[0,D), I max and I min They represent the maximum and minimum values of all pixel values in the optical coherence tomography image, respectively.
4. The method for identifying diabetic nephropathy retinopathy based on double distillation according to claim 1, characterized in that: The step 4 comprises the following steps: By designing a category-balanced loss function, the weights of the identified lesion samples of the ResNet18 network model focusing on categories with less data are adjusted; By constructing a category-adaptive attention model, the sample weights of the category-adaptive attention model are dynamically balanced.
5. The method for identifying diabetic nephropathy retinopathy based on double distillation according to claim 4, characterized in that: The class balance loss function is: L fl =-(1-p t ) γ log(p t ) Among them, p t is the probability predicted by the model, (1-p t ) γ is the adjustment factor and γ is the balance parameter.
6. The method for identifying diabetic nephropathy retinopathy based on double distillation according to claim 4, characterized in that: The category-adaptive attention model obtains the sensitivity score of each category through three-dimensional 1×1 convolution and global average pooling: Among them, f' i,j is the input image feature, GP is the global average pooling operation, k represents the number of feature map channels, i∈{1,2,…,L}, and L is the number of categories; The category-adaptive attention model performs average pooling on the feature map after the three-dimensional 1×1 convolution by category channel to obtain the feature map F": The category-adaptive attention model combines the attention region with the original image, enhances the input feature image, and identifies the retinal lesion area CAM_OUT: Among them, F is the input feature map.
7. The method for identifying diabetic nephropathy retinopathy based on double distillation according to any one of claims 1 to 4, characterized in that: The step 5 comprises the following steps: The class balance loss function is expressed as: L fl =-(1-p t ) γ log(p t ) Among them, p t is the probability predicted by the model, (1-p t ) γ is the adjustment factor, γ is the balance parameter; The cross entropy loss function is expressed as: L CE =-log(pt) Among them, p t is the probability predicted by the model; The knowledge transfer loss function is expressed as: Among them, P i is the probability predicted by the double distillation model, Q i is the probability of the pre-trained network model, Y i is the true label distribution; The Adam optimization algorithm is used to optimize the loss function and distill the knowledge of the pre-trained model into the dual network model to obtain the final retinal lesion recognition model.
8. A diabetic nephropathy retinopathy identification device based on double distillation, characterized in that: include: An image acquisition module, used for acquiring an optical coherence tomography image of a patient; An image processing module, used for normalizing the optical coherence tomography image; A pre-trained model acquisition module is used to pre-train the 3D-ResNet network using a large public medical image dataset to obtain a pre-trained model; A double distillation model acquisition module, used for inputting the preprocessed optical coherence tomography image into a double distillation framework to obtain a double distillation model, wherein the double distillation framework includes two ResNet18 network models, one focusing on a category with less data volume, and the other focusing on a category adaptive attention model; The loss function joint module is used to combine the loss functions of the pre-trained model and the double distillation model, distill the knowledge of the pre-trained model into the double network model, and obtain the final retinal lesion recognition model; The lesion recognition module is used to identify diabetic nephropathy retinopathy using a retinopathy recognition model.
9. An electronic device, characterized in that: It includes a processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the diabetic nephropathy retinopathy identification method based on double distillation as described in any one of claims 1-7.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for identifying diabetic nephropathy retinopathy based on double distillation as described in any one of claims 1 to 7.