Intelligent diagnosis system for diabetic eye diseases based on multi-model cluster and causal enhanced LLM cooperation

The intelligent diagnostic system for diabetic retinopathy, which combines multi-model clustering with causal enhancement LLM, solves the problems of lack of causal logic and insufficient diagnostic reliability in existing technologies, and achieves accurate lesion detection and generation of structured diagnostic reports that meet clinical standards.

CN122117343APending Publication Date: 2026-05-29UESTC (SHENZHEN) ADVANCED RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UESTC (SHENZHEN) ADVANCED RES INST
Filing Date
2026-04-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing diagnostic systems for diabetic retinopathy rely on end-to-end deep learning models, lack medical causal logic, cannot explain the causal relationship between lesion characteristics and diagnostic grading, and the visual model output lacks explicit correlation with clinical diagnostic criteria, resulting in insufficient diagnostic reliability.

Method used

By employing a multi-model cluster and a causal-enhanced large language model (LLM) in synergy, the system identifies pathological features such as microaneurysms, bleeding points, exudates, and neovascularization through a modular detection cluster layer. The causal-enhanced LLM is then used for comprehensive diagnosis and interpretation, and combined with a standardized medical description knowledge base, a structured report that meets clinical diagnostic criteria is generated.

Benefits of technology

It achieves accurate lesion detection, possesses causal reasoning capabilities, ensures seamless alignment of diagnostic results with clinical standards, and boasts advantages such as high detection accuracy, interpretability, and standardized output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117343A_ABST
    Figure CN122117343A_ABST
Patent Text Reader

Abstract

The application discloses a kind of intelligent diagnosis system of diabetic eye disease based on multi-model cluster and cause-effect enhanced LLM cooperation, and it is related to artificial intelligence medical auxiliary diagnosis technical field.The system includes modular detection cluster layer, microaneurysm, bleeding point, exudate, neovascular detection is carried out to input fundus image simultaneously, and the visual feature corresponding to diabetic eye disease is output;Medical text modal alignment layer, the diabetic eye disease visual feature output by modular detection cluster layer is converted into standardized medical description;Large language model diagnosis layer, according to the standardized medical description output by medical text modal alignment layer, combined with the construction of standardized medical description knowledge base system, the input fundus image is carried out diabetic eye disease diagnosis reasoning, and structured diagnostic report is generated.The system has the advantages of high overall detection accuracy, knowledge evolution, diagnosis interpretability, output standardization and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence-assisted medical diagnosis technology, and in particular to an intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy. Background Technology

[0002] Ophthalmic medical imaging (such as fundus photography, OCT, OCTA, fluorescence / indocyanine green angiography, etc.) is crucial for screening and diagnosing diseases such as diabetic retinopathy (DR), age-related macular degeneration (AMD), and glaucoma. In recent years, deep learning (convolutional networks, Transformers) and multimodal vision-language models (VLMs) have made progress in tasks such as lesion detection, grading, and report generation.

[0003] Currently, the diagnosis of diabetic retinopathy mainly relies on end-to-end deep learning models, which lack medical causal logic and cannot explain the causal relationship between lesion characteristics and diagnostic grading. Moreover, the visual model output lacks explicit correlation with the Clinical Diagnostic Rating Scale (ICDR), resulting in fragmented multimodal information and insufficient diagnostic reliability.

[0004] Therefore, there is an urgent need for an intelligent diagnostic system that can accurately detect diabetic retinopathy, has causal reasoning capabilities, and meets clinical diagnostic criteria. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy, aiming to solve the aforementioned problems and provide an intelligent diagnostic system capable of accurate lesion detection, possessing causal reasoning ability, and meeting clinical diagnostic criteria. The numerous technical effects of the preferred solutions among the various technical solutions provided by this invention are detailed below.

[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention provides an intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy, comprising: The modular detection cluster layer simultaneously performs microaneurysm detection, hemorrhage detection, exudate detection, and neovascularization detection on the input fundus image, and outputs the visual features corresponding to diabetic retinopathy. The medical text modality alignment layer transforms the visual features of diabetic retinopathy output by the modular detection cluster layer into standardized medical descriptions. The large language model diagnostic layer, based on the standardized medical descriptions output by the medical text modality alignment layer and combined with the constructed standardized medical description knowledge base system, performs diagnostic reasoning for diabetic retinopathy on the input fundus images and generates a structured diagnostic report.

[0007] In some embodiments, the modular detection cluster layer includes a microaneurysm detection sub-layer, which includes a backbone network, a neck network, a multi-scale detection head, an auxiliary segmentation attention head, and a context enhancement module connected in sequence. The backbone network is a YOLOv8 backbone network, which employs a cross-stage local attention module. The neck network consists of a feature pyramid and dynamic sparse convolution. The loss function of the microaneurysm detection sub-layer includes a spatial constraint term, which regularizes and penalizes the overlap relationship of multiple detection boxes within the same region.

[0008] In some embodiments, the modular detection cluster layer includes a bleeding point detection sub-layer, which adopts the U-Net++ architecture. A multi-scale feature fusion module is embedded in the skip connection between the encoder and decoder of U-Net++, and a boundary-aware attention mechanism is added to the output layer of the U-Net++ architecture. The multi-scale feature fusion module includes dilated convolutional groups and deformable convolutional layers that operate in parallel.

[0009] In some embodiments, the modular detection cluster layer includes an exudate detection sub-layer, which is a Mask R-CNN network with a dual-branch detection head added after the ROI Align layer of the Mask R-CNN network; one branch detection head uses a circularity-constrained convolution kernel to detect hard exudates with clear boundaries, and the other branch detection head uses a Gaussian diffusion convolution kernel to detect soft exudates with blurred boundaries.

[0010] In some embodiments, the modular detection cluster layer includes a neovascularization detection sub-layer, which employs a hierarchical feature extraction strategy. The bottom feature extraction stage uses a convolutional neural network to obtain local texture features. The mid-to-high-level feature processing stage employs a multi-head self-attention mechanism to establish long-distance dependencies between vascular branches. The output layer uses a multi-scale detection head to detect vascular structures with different diameter ranges.

[0011] In some embodiments, the medical text modality alignment layer converts the visual features of diabetic retinopathy into standardized text cue words, including: The system receives the visual feature data of the lesions output from the modular detection cluster layer; automatically selects the corresponding medical description template according to the type of diabetic retinopathy; generates structured prompt words according to the selected medical description template, and outputs structured prompt words that meet the input requirements of the large language model.

[0012] In some embodiments, the output consists of structured prompts that meet the input requirements of a large language model, including role definition instructions, detailed descriptions of lesion features, diagnostic task requirements, and output format specifications.

[0013] In some embodiments, the medical description template includes a microaneurysm description template, a bleeding point description template, an exudate description template, and a neovascularization description template.

[0014] In some embodiments, generating a structured diagnostic report includes: Based on the ICDR grading standard of the standardized medical description knowledge base system, a diagnostic conclusion that conforms to the international clinical grading standard for diabetic retinopathy is output. The lesion feature quantification report based on lesion visual feature data provides a lesion feature quantification data table based on image analysis; A causal explanation mechanism for hierarchical diagnosis is provided based on causal explanation and diagnostic reasoning; Generate personalized clinical management recommendations.

[0015] In some embodiments, before the fundus images are input into the modular detection cluster layer, the fundus images are further subjected to image quality assessment and standardization preprocessing. The image quality assessment includes a comprehensive analysis of the fundus image sharpness, exposure, illumination uniformity, visual field localization, and noise level to obtain a fusion score. Fundus images with a fusion score below a threshold will be automatically rejected or prompted for re-acquisition. Fundus images with a fusion score above the threshold will be subjected to standardization preprocessing.

[0016] Implementing one of the above-described technical solutions of the present invention has the following advantages or beneficial effects: This invention identifies pathological features such as microaneurysms, hemorrhage, exudates, and neovascularization through a modular detection cluster. It utilizes a large-scale language model (LLM) with causal enhancement for comprehensive diagnosis and interpretation generation. The causal reasoning engine provides interpretive paths and structured medical descriptions consistent with clinical thinking. An expert knowledge base supports continuous learning and multi-center experience fusion, exhibiting self-optimization capabilities to ensure seamless alignment of diagnostic results with clinical standards. The system of this invention possesses advantages such as high overall detection accuracy, evolvable knowledge, interpretable diagnosis, and standardized output. Attached Figure Description

[0017] The accompanying drawings used are briefly described below. It is obvious that the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. In the drawings: Figure 1This is a block diagram of an intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy, according to an embodiment of the present invention. Figure 2 This is a layered framework diagram of a microaneurysm detection subgroup according to an embodiment of the present invention; Figure 3 This is a hierarchical framework diagram of a new blood vessel detection subgroup according to an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments, illustrating various exemplary embodiments that may be used to implement the present invention. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of the present invention disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of the present invention.

[0019] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0020] To illustrate the technical solution described in this invention, specific embodiments are described below, showing only the parts related to the embodiments of this invention.

[0021] Example 1: like Figure 1 As shown, this invention provides an intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy, comprising: The modular detection cluster layer simultaneously performs microaneurysm detection, hemorrhage detection, exudate detection, and neovascularization detection on the input fundus images, and outputs visual features corresponding to diabetic retinopathy. Diabetic retinopathy includes microaneurysms, hemorrhages, exudates, and neovascularization, and the visual features include the number of lesions, their location coordinates, area measurements, and confidence levels.

[0022] A medical text modal alignment layer, connected to a modular detection cluster layer, transforms the visual features of diabetic retinopathy into standardized medical descriptions.

[0023] The large language model diagnostic layer is connected to the medical text modality alignment layer. Based on standardized medical descriptions and combined with the constructed standardized medical description knowledge base system, it performs diagnostic reasoning for diabetic retinopathy on the input fundus images and generates a structured diagnostic report.

[0024] This system identifies lesion features such as microaneurysms, hemorrhage, exudates, and neovascularization through modular detection clusters. It utilizes a large-scale language model (LLM) with enhanced causality for comprehensive diagnosis and interpretation. The causal reasoning engine provides clinically relevant explanatory paths and structured medical descriptions. An expert knowledge base supports continuous learning and multi-center experience fusion, exhibiting self-optimization capabilities to ensure seamless alignment of diagnostic results with clinical standards. This system boasts advantages such as high overall detection accuracy, evolvable knowledge, interpretable diagnosis, and standardized output.

[0025] It should be noted that fundus images are acquired through ophthalmic medical imaging equipment. Before being input into the modular detection cluster layer, the fundus images undergo image quality assessment and standardized preprocessing to ensure the stability of the input data and the reliability of the diagnosis.

[0026] In a specific embodiment, image quality assessment mainly includes a comprehensive analysis of fundus image sharpness, exposure, illumination uniformity, visual field localization, and noise level. Specifically, a sharpness detection method based on Laplacian variance and gradient energy is used to identify blurred or defocused images; overexposure or underexposure is determined using grayscale histogram statistics and average brightness shift; uneven illumination and reflective areas are detected through Gaussian illumination field fitting; simultaneously, the geometric localization of the optic disc and macula is combined to determine whether the imaging deviates from the standard visual field range; and a lightweight deep learning network is used to fuse and score multiple indicators, outputting a three-level quality rating of "excellent," "acceptable," and "poor." Fundus images below the threshold will be automatically discarded or prompted for re-acquisition.

[0027] Images that have passed quality screening enter a standardized preprocessing workflow. This system first unifies the scale and resolution of the images, scaling the input images to a fixed size (e.g., 512×512 or 1024×1024), and removes the black background through circular cropping or U-Net segmentation, retaining only the effective retinal area. Subsequently, brightness and contrast correction are performed, using Adaptive Histogram Equalization (CLAHE) or Retinex illumination compensation algorithms to correct uneven illumination. A color normalization algorithm is then used to align the image's color distribution to a standard reference template, reducing color deviations caused by different devices or acquisition centers. To reduce imaging noise, the system further employs non-local mean filtering for denoising, smoothing background textures while maintaining clear boundaries between blood vessels and lesions. Finally, the system calculates the rotation angle based on the optic disc and macula positions to achieve geometric alignment, ensuring standardization of the visual field direction and consistency with the spatial distribution of lesions.

[0028] like Figure 2 As shown, in one or more embodiments, the modular detection cluster layer includes a microaneurysm detection sub-layer, which employs a modified YOLO (You Only Look Once) v8 architecture. This sub-layer includes a backbone network, a neck network, a multi-scale detection head, an auxiliary segmentation (boundary) attention head, a context enhancement module, and a post-processing output module, all connected in sequence.

[0029] The backbone network is a YOLOv8 backbone network, which employs a Cross-Stage Local Attention Module (CSLA) to enhance the feature extraction capability for small targets. The neck network consists of a feature pyramid and dynamic sparse convolution, specifically using the neck of YOLOv8 and replacing standard convolution with dynamic sparse convolution to significantly improve the detection sensitivity for microaneurysms smaller than 50 μm. The loss function for the microaneurysm detection subgroup layer adds a spatial constraint term to the original classification and regression loss function of YOLOv8, forcing this subgroup layer to focus on the spatial localization accuracy of small lesions.

[0030] The loss function of YOLOv8 mentioned above adds a spatial constraint term to the original classification and regression loss, that is, the loss function of YOLOv8 is improved to take into account the characteristics of microaneurysms such as small size, complex morphology and strong background interference.

[0031] In some specific embodiments, the overall loss function of the improved YOLOv8 microaneurysm detection subgroup layer is defined as follows: ; in, : Classification loss, used to determine the target category (lesion type); : This is the bounding box regression loss, which measures the coordinate deviation between the predicted box and the ground truth box; : Objectness Loss, used to distinguish lesion areas from the background; : This is the spatial constraint loss term, used to constrain the spatial consistency between predicted boxes. Its definition is as follows: ; in, This represents the intersection-over-union ratio (IoU) between the i-th predicted bounding box and its neighboring predicted bounding boxes. This is the spatial constraint coefficient, used to adjust the weight of this term in the total loss.

[0032] This spatial constraint term regularizes and penalizes the overlap of multiple detection boxes within the same region, prompting the network to automatically learn the optimal detection box distribution during training, thus avoiding redundant or discrete prediction results in densely populated microaneurysm regions.

[0033] In one or more embodiments, the modular detection cluster layer includes a bleed point detection sub-cluster, which adopts the U-Net++ architecture. Specifically, a multi-scale feature fusion module is embedded in the skip connections between the encoder and decoder of U-Net++. This multi-scale feature fusion module includes parallel-operating dilated convolutional groups (dilation rates = 1, 3, 5) and deformable convolutional layers, thereby effectively capturing multi-scale features of different shaped bleed regions, such as point-like and flame-like patterns. A boundary-aware attention mechanism is added to the output layer of the U-Net++ architecture to enhance the segmentation accuracy of bleed boundaries by calculating pixel-level boundary weight maps.

[0034] In some specific embodiments, the correspondence between the dilation rate and the dilated convolutional groups is shown in Table 1 below. In the multi-scale feature fusion module, the proposed dilated convolution group includes three parallel branches (meaning that the module consists of three dilated convolution branches): ; Table 1. Correspondence between dilation rate and dilated convolution group The output of a multi-scale dilated convolution group can be expressed as: ; in, This represents a dilated convolution with an expansion rate of k; The input feature map can be a standardized preprocessed fundus image. This indicates a splicing operation at the channel dimension.

[0035] After concatenation, a 1x1 convolution is used for channel compression and fusion: ; This allows the hemorrhage detection subgroup layer to simultaneously "see" the features of different receptive fields on the same layer, thus preserving the local highlight features of the hemorrhage points while integrating the morphological and background information of the surrounding area.

[0036] To make it easier to understand, the three dilated convolution branches achieve parallel modeling of multi-scale receptive fields through different dilation rates, where the dilation rate is linearly positively correlated with the feature receptive field. The larger the dilation rate, the wider the spatial range covered by the convolution kernel, and the larger the scale of the detected hemorrhage region.

[0037] The bleeding detection subgroup layer enables U-Net++ to capture complex bleeding patterns such as point-like, patch-like, and flame-like bleeds at different resolution layers by fusing these multi-scale responses at skip connections.

[0038] In one or more embodiments, the modular detection cluster layer includes an exudate detection sub-layer, which is a Mask R-CNN network. Specifically, a dual-branch detection head is added after the ROI Align layer of the Mask R-CNN network; one branch detection head uses a circularity-constrained convolutional kernel to specifically detect hard exudates with well-defined boundaries, while the other branch detection head uses a Gaussian diffusion convolutional kernel to detect soft exudates with blurred boundaries.

[0039] In a specific embodiment, the above-mentioned exudate detection subgroup layer is designed with a dual-branch detection head after the ROI Align layer of Mask R-CNN, which is designed for hard exudates (with clear boundaries) and soft exudates (with blurred boundaries).

[0040] In one or more specific embodiments, as shown in Table 2, the modular detection cluster layer includes an exudate detection sub-cluster layer, which employs an improved Mask R-CNN network architecture. The overall structure of this sub-cluster layer includes a feature extraction layer, a region candidate generation layer (RPN), an ROI Align feature alignment layer, and a dual-branch detection head connected in sequence.

[0041] The dual-branch detection head added after the ROI Align layer is used for targeted detection and segmentation based on the characteristic differences of different types of exudates. Specifically, First detection branch (hard exudation detection branch): By employing a circularity-constrained convolution kernel, the weight distribution is constrained by a circular template mask during convolution operations, enabling the kernel to have a stronger response to hard exudation regions with closed boundaries and prominent brightness.

[0042] In the output stage, a boundary detection loss function is incorporated to enhance contour accuracy.

[0043] Second detection branch (soft exudation detection branch): By employing a Gaussian-Diffusion Convolution kernel, a Gaussian diffusion function is introduced into the kernel weight space. This allows the exudate detection subgroup layer to retain fuzzy boundary information in the feature response, making it more suitable for detecting soft exudate regions with gradual edge transitions and smooth grayscale changes.

[0044] Blur-Consistency Loss is introduced during the training phase to enhance the robustness of the exudate detection subgroup layer to uncertain boundary regions.

[0045] The output of the dual-branch detection head is jointly evaluated by the branch fusion module to generate a comprehensive output that includes hard and soft exudation segmentation masks. This structure can simultaneously maintain the precise boundary of hard exudation detection and the edge sensitivity of soft exudation detection, enabling comprehensive identification of multiple types of exudation features.

[0046] Table 2. Overview of Module Numbers, Names, and Functions of the Effluent Detection Subgroup Layer In one or more embodiments, the modular detection cluster layer includes a neovascularization detection sub-layer, which adopts a multi-scale feature pyramid Transformer architecture and a hierarchical feature extraction strategy. Specifically, the bottom-level feature extraction stage uses a convolutional neural network to acquire local texture features; the mid-to-high-level feature processing stage introduces a multi-head self-attention mechanism to establish long-distance dependencies between vascular branches; the output layer designs multi-scale detection heads to detect vascular structures with different diameter ranges, such as: main trunk vessels >100μm, branch vessels 50-100μm, and neovascularization <50μm.

[0047] Furthermore, the low-level feature extraction stage uses convolutional neural networks (CNNs), such as ResNet or EfficientNet, to extract local texture features of blood vessels through layer-by-layer convolution and pooling operations. Let the input image be... The features extracted by the CNN are represented as C2, C3, C4, and C5, respectively, where each feature map has a dimension of 1. .

[0048] The convolution operation can be represented as: ; Represents convolution. It is a non-linear activation function. These represent the kernel weights and biases.

[0049] To accommodate both large and small blood vessels, a top-down Feature Pyramid (FPN) is constructed. First, each CNN layer's features are convolved using a 1×1 model to map them to a unified channel dimension. Then, they are sequentially upsampled and fused with features from the next lower layer. ; Where L(⋅) is a lateral 1×1 conv, and U(⋅) represents an upsampling operation that generates multi-scale features {P2,P3,P4,P5} corresponding to different blood vessel diameter ranges.

[0050] To establish long-range dependencies between vascular branches, tokenization (through 1×1 convolution or patch segmentation) is performed on the feature map (image features extracted by CNN) at each scale to form a token sequence. And input it into the Transformer Encoder. The self-attention mechanism captures long-range relevance: ; in, These are query, key, and value vectors, respectively. Dimensions for each head , These are the learnable projection matrices used to generate query, key, and value vectors, respectively. The projection matrix ( ; ; Learnable parameters used to linearly map input features to the Q, K, V space.

[0051] Multi-Head Self-Attention (MHSA) concatenates and linearly maps the outputs of multiple attention heads: ; in, It is a learnable "feature fusion matrix", where h is the number of attention heads.

[0052] Furthermore, Its functions include: Information integration: Compress the heterogeneous features (such as orientation, scale, and topological relationships) captured by h attention heads back into a unified feature space; Adaptive weighting: learning the importance of different heads in the final diagnosis; Dimension alignment: Ensures that the output can be added to the residual connection, maintaining network depth.

[0053] In this process, multi-head self-attention concatenates the outputs of h parallel attention heads and then applies the result to a learnable projection matrix. A linear transformation is performed to achieve adaptive fusion of multi-view features (such as blood vessel direction, branch topology, and tube diameter scale), and the output dimension remains consistent with the input features, which facilitates residual connection and hierarchical stacking.

[0054] In cross-scale fusion, cross-attention can be used to establish information exchange between low-scale microvascular tokens and high-scale vascular tokens: ; in, = Low-scale microvascular features are used as a query to find locations where details need to be enhanced. For low-scale microvascular tokens; = High-scale vascular trunk features serve as keys, providing spatial location information; For high-scale vascular tokens; = High-scale vascular trunk features are used as values ​​to provide contextual information to be transmitted.

[0055] It should be noted that the above mechanism ensures that the microvascular structure information can be globally correlated with the main direction and branch information of the blood vessels, thereby improving the topological consistency of the network structure.

[0056] like Figure 3 As shown, the neovascularization detection subgroup layer includes a multi-scale pyramid sampling module, a multi-scale feature extraction module, a multi-scale tokenization module, a cross-scale self-attention convolution, a segmentation head, a detection head, and an output module, which are connected in sequence.

[0057] It should be noted that the modular detection cluster layer breaks down the key diagnoses of diabetic retinopathy, improves or fine-tunes existing visual models for local retinal lesions, integrates multiple detection models for microaneurysms, hemorrhages, exudates, and neovascularization, and fuses the lesion features of multiple models. This means that the method does not require each model to achieve extremely high accuracy, but rather improves the overall diagnostic performance through the complementarity and synergy between models.

[0058] In one or more embodiments, the medical text modality alignment layer converts visual features of diabetic retinopathy into standardized text cue words, including: It receives lesion visual feature data from the modular detection cluster layer; automatically selects the corresponding medical description template according to the type of diabetic retinopathy; generates structured prompt words based on the selected medical description template, and outputs structured prompt words that meet the input requirements of the large language model.

[0059] More specifically, it receives lesion feature data output from the modular detection cluster, including: Microaneurysm characteristic data: number of lesions, spatial distribution coordinates, diameter measurement, confidence score; Bleeding point feature data: bleeding area, morphological classification, distribution density, and location information; Exudate characteristics: exudate type, cumulative area, distance from the center of the macula, and boundary clarity; Neovascularization characteristic data: vessel classification, extent score, activity assessment, and topological characteristics.

[0060] Furthermore, the medical description templates include templates for microaneurysms, bleeding points, exudates, and neovascularization. Among them, Microaneurysm description template: "Microaneurysms: {number}, mainly distributed in {anatomical location}, maximum diameter {diameter} μm, confidence level {confidence}%"; Bleeding point description template: "Bleeding points: Total area {area} mm², {morphological characteristics} distribution, mainly located in {retinal region}"; Exudate description template: "{Exudate type} Exudate: Cumulative area {area} mm², distance from macular center {distance} mm, boundary {clarity}"; New blood vessel description template: "New blood vessel: {vascular type}, range score {score} points, activity {activity level}".

[0061] Furthermore, the output includes structured prompts that conform to the input requirements of the large language model, including role definition instructions, detailed descriptions of lesion characteristics, diagnostic task requirements, and output format specifications. Among these, Role definition instructions, such as, "As an ophthalmology diagnostic expert, please analyze the following lesion characteristics:"; A detailed description of the lesion characteristics, such as the various lesion parameters arranged in the ICDR standard format; Diagnostic task requirements, such as, "Please make a diagnosis according to the ICDR classification criteria and provide a detailed reasoning process"; Output format specifications, such as "The diagnostic report should include the grading results, main evidence, differential diagnosis, and follow-up recommendations".

[0062] To achieve clinically logical and interpretable diagnoses, this embodiment constructs a causal reinforcement-based big language diabetes grading diagnostic model. By building a multi-level knowledge base system, the big language model is guided to possess multi-source knowledge fusion and advanced reasoning capabilities. The sources for constructing the standardized medical description knowledge base system include: ICDR Grading Standards Library: Structured storage of international clinical grading standards for diabetic retinopathy, including disease thresholds, grading rules, and clinical interpretations; Medical knowledge base and clinical guideline base: Integrating multi-dimensional knowledge on the pathophysiological mechanisms, typical clinical manifestations, imaging features and treatment plans of diabetic retinopathy; Expert Experience Rule Base: Collects and refines the diagnostic experience and typical cases of ophthalmologists from multiple centers to form calculable diagnostic rules and abnormality handling strategies.

[0063] It is understood that the large language model in this embodiment can be used in existing technologies, such as GPT-3, GPT-4, BERT, etc., which will not be elaborated here.

[0064] In one or more embodiments, generating a structured diagnostic report includes: The diagnostic conclusions should conform to the International Clinical Diabetic Retinopathy (ICDR) grading criteria, including: No diabetic retinopathy (No DR); Mild non-proliferative diabetic retinopathy (Mild NPDR); Moderate non-proliferative diabetic retinopathy (Moderate NPDR); Severe non-proliferative diabetic retinopathy (Severe NPDR); Proliferative diabetic retinopathy (PDR).

[0065] The lesion feature quantification report based on lesion visual feature data provides a lesion feature quantification data table based on image analysis, including: Microaneurysms: Quantity statistics, spatial distribution heatmap, diameter distribution range, and detection confidence level; Bleeding points: calculation of total bleeding area, morphological classification results, regional distribution density, and location coordinate mapping; Exudates: type identification (hard / soft), cumulative area measurement, distance from the center of the macula calculation, and boundary sharpness score; Neovascularization: genotyping identification (NVD / NVE), range scoring (0-10 points), activity assessment index, and topological feature description.

[0066] A causal explanation mechanism for graded diagnosis is provided based on causal explanation and diagnostic reasoning, including: Analysis of the conformity between key lesion features and ICDR grading criteria; medical logic explanation of the combination effect of lesions; explanation of differential diagnosis criteria and exclusion criteria; assessment of diagnostic confidence and uncertainty analysis.

[0067] Generate personalized clinical management recommendations, including: Treatment recommendations: Based on the grading results, appropriate treatment options will be recommended (observation, laser photocoagulation, anti-VEGF therapy, surgery). Follow-up plan: Develop a differentiated follow-up schedule (3-12 months) based on the severity of the lesions; Examination items: Recommended next necessary auxiliary examinations (OCT, FFA, OCTA); Risk warning: Indicates the risk of disease progression and clinical symptoms that require attention.

[0068] The above description is merely a preferred embodiment of the present invention. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.

Claims

1. A smart diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy, characterized in that, include: The modular detection cluster layer simultaneously performs microaneurysm detection, hemorrhage detection, exudate detection, and neovascularization detection on the input fundus image, and outputs the visual features corresponding to diabetic retinopathy. The modular detection cluster layer includes a microaneurysm detection subgroup layer and a multi-scale feature fusion module. The loss function of the microaneurysm detection subgroup layer includes a spatial constraint term, which regularizes and penalizes the overlap relationship of multiple detection boxes in the same region. The multi-scale feature fusion module includes parallel dilated convolution groups and deformable convolution layers. The medical text modality alignment layer transforms the visual features of diabetic retinopathy output by the modular detection cluster layer into standardized medical descriptions. The large language model diagnostic layer, based on the standardized medical descriptions output by the medical text modality alignment layer and combined with the constructed standardized medical description knowledge base system, performs diagnostic reasoning for diabetic retinopathy on the input fundus images and generates a structured diagnostic report.

2. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, The microaneurysm detection subgroup layer includes a backbone network, a neck network, a multi-scale detection head, an auxiliary segmentation attention head, and a context enhancement module connected in sequence. The backbone network is a YOLOv8 backbone network, which employs a cross-stage local attention module. The neck region consists of a feature pyramid and dynamic sparse convolution.

3. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, The modular detection cluster layer includes a bleeding point detection sub-layer, which adopts the U-Net++ architecture. The multi-scale feature fusion module is embedded in the skip connection between the encoder and decoder of U-Net++, and a boundary-aware attention mechanism is added to the output layer of the U-Net++ architecture.

4. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, The modular detection cluster layer includes an exudate detection sub-layer, which is a Mask R-CNN network with a dual-branch detection head added after the ROI Align layer of the Mask R-CNN network. One branch detection head uses a circularity-constrained convolution kernel to detect hard exudates with clear boundaries, while the other branch detection head uses a Gaussian diffusion convolution kernel to detect soft exudates with blurred boundaries.

5. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, The modular detection cluster layer includes a neovascularization detection sub-layer, which employs a hierarchical feature extraction strategy. The low-level feature extraction stage uses a convolutional neural network to obtain local texture features; middle The high-level feature processing stage employs a multi-head self-attention mechanism to establish long-distance dependencies between vascular branches; the output layer uses a multi-scale detection head to detect vascular structures of different diameter ranges.

6. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, The medical text modality alignment layer converts the visual features of diabetic retinopathy into standardized text cue words, including: The system receives the visual feature data of the lesions output from the modular detection cluster layer; automatically selects the corresponding medical description template according to the type of diabetic retinopathy; generates structured prompt words according to the selected medical description template, and outputs structured prompt words that meet the input requirements of the large language model.

7. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 6, characterized in that, The output should conform to the input requirements of the large language model, including role definition instructions, detailed description of lesion characteristics, diagnostic task requirements, and output format specifications.

8. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 6, characterized in that, The medical description templates include microaneurysm description templates, bleeding point description templates, exudate description templates, and neovascularization description templates.

9. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, Generate structured diagnostic reports, including: Based on the ICDR grading standard of the standardized medical description knowledge base system, a diagnostic conclusion that conforms to the international clinical grading standard for diabetic retinopathy is output. The lesion feature quantification report based on lesion visual feature data provides a lesion feature quantification data table based on image analysis; A causal explanation mechanism for hierarchical diagnosis is provided based on causal explanation and diagnostic reasoning; Generate personalized clinical management recommendations.

10. The intelligent diagnostic system for diabetic retinopathy based on multi-model clustering and causal enhancement LLM synergy as described in claim 1, characterized in that, Before the fundus images are input into the modular detection cluster layer, the fundus images also undergo image quality assessment and standardization preprocessing. The image quality assessment includes a comprehensive analysis of the fundus image's sharpness, exposure, illumination uniformity, visual field localization, and noise level to obtain a fusion score. Fundus images with a fusion score below a threshold will be automatically discarded or prompted for re-acquisition. Fundus images with a fusion score above the threshold undergo standardization preprocessing.