Glaucoma lesion feature screening system and method based on artificial intelligence, electronic equipment and storage medium

By employing a multimodal data acquisition and fusion method, and utilizing convolutional neural networks and Transformer encoders to extract glaucoma lesion features, this approach addresses the issues of information isolation and reliance on physician experience in existing diagnostic models, enabling efficient and objective early screening and grading of glaucoma.

CN122024977APending Publication Date: 2026-05-12SHANGHAI MIRROR MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MIRROR MEDICAL TECHNOLOGY CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Current glaucoma diagnostic models rely on single-modal information, lack effective comprehensive assessment tools, are highly dependent on physician experience, have low diagnostic efficiency, insufficient sensitivity in early screening, and lack efficient and objective quantitative grading methods.

Method used

The system employs a multimodal data acquisition and preprocessing module, combined with the multimodal feature extraction and fusion module described in the manual. It utilizes convolutional neural networks and Transformer encoders to extract and fuse high-dimensional features from fundus color images, optical coherence tomography images, and visual field examination images. Through a multi-task learning framework, it achieves the screening and classification of glaucoma lesion features.

Benefits of technology

It significantly improves the sensitivity and grading accuracy of early glaucoma lesion feature screening, realizes automated and objective multimodal information fusion, outputs structured diagnostic reports and visualized heat maps, and improves the objectivity and efficiency of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024977A_ABST
    Figure CN122024977A_ABST
Patent Text Reader

Abstract

The invention discloses a glaucoma lesion feature screening system and method based on artificial intelligence, electronic equipment and a storage medium. The lesion feature screening system comprises a multi-modal data acquisition and preprocessing module, a multi-modal feature extraction and fusion module and a screening and grading decision module. The multi-modal data acquisition and preprocessing module is used for acquiring clinical data and performing standardization processing on the acquired clinical data to generate input data with a unified specification; the multi-modal feature extraction and fusion module is used for extracting high-dimensional features from the standardized multi-modal images by using pre-trained convolutional neural networks with different structures, and generating information enhanced joint feature representation through a fusion network based on an attention mechanism; and the screening and grading decision module is used for screening and grading glaucoma lesion features through a classifier based on the joint feature representation. According to the method, the sensitivity of glaucoma early lesion feature screening and the accuracy of grading can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing and artificial intelligence technology, and relates to a lesion feature screening system, and more particularly to an artificial intelligence-based glaucoma lesion feature screening system, method, electronic device and storage medium. Background Technology

[0002] Glaucoma is the leading cause of irreversible blindness worldwide, characterized by progressive damage to the optic nerve and visual field defects. Because early-stage glaucoma patients often experience no symptoms, by the time visual function is significantly impaired, the disease has usually progressed to the middle or late stages, missing the optimal treatment window. Therefore, early and accurate screening and severe severity grading are crucial for slowing disease progression and preserving patients' visual function.

[0003] Currently, the clinical diagnosis of glaucoma relies on a series of multimodal examinations, each with its own advantages and inherent limitations:

[0004] (1) Fundus color photography: used to visually assess optic disc morphology, such as cup-to-disc ratio (C / D ratio), disc rim notch, hemorrhage, and retinal nerve fiber layer (RNFL) defects. However, this method has limited ability to identify early subtle changes, and the diagnostic results are highly dependent on the physician's clinical experience, exhibiting a strong degree of subjectivity.

[0005] (2) Optical coherence tomography (OCT): It can measure the thickness of the retinal nerve fiber layer (RNFL) and macular ganglion cell complex (GCC) with high resolution and quantitative analysis, providing objective quantitative evidence for structural damage in glaucoma. Nevertheless, the interpretation of its data still requires professional experience, and as a single modality of information, there is a risk of false positives and false negatives.

[0006] (3) Intraocular pressure measurement: As a core risk factor for glaucoma, it is an important part of routine screening. However, due to its lack of diagnostic specificity, people with normal intraocular pressure may still develop glaucoma (normal-tension glaucoma), so it cannot be used as an independent exclusion criterion.

[0007] (4) Visual field examination: It is currently regarded as the "gold standard" for judging glaucoma functional impairment. However, this examination is highly subjective, and functional deficits usually lag behind structural damage, which is not conducive to true early diagnosis.

[0008] (5) OCT angiography (OCTA): It can non-invasively observe the blood flow density (VD) of the optic disc and macular area, providing diagnostic basis from the perspective of microcirculation. However, the correlation analysis between OCTA parameters and structural parameters is relatively complex.

[0009] The existing technology has the following drawbacks:

[0010] (1) Reliance on a single modality: Diagnosis based solely on a single imaging modality (such as only looking at fundus color photos or only analyzing OCT) lacks effective comprehensive assessment tools, resulting in incomplete information and easily leading to missed or misdiagnosis. For example, early glaucoma may appear normal on fundus color photos, but thinning of the RNFL may be observed on OCT, easily leading to missed diagnosis of early or atypical cases.

[0011] (2) Highly dependent on expert experience: The accuracy of diagnosis is greatly affected by the doctor's clinical experience and subjective judgment, and there may be differences in diagnosis between different doctors.

[0012] (3) Low diagnostic efficiency: Manual quantitative measurement and comparison of massive image data is time-consuming and laborious, and is difficult to popularize in primary hospitals and large-scale screening.

[0013] (4) Lack of precise grading: Some existing computer-aided diagnostic systems focus on “screening” (i.e. binary classification: yes or no), and lack the ability to automate and refine the grading of glaucoma severity (preclinical, early, intermediate, late and terminal stages), which is crucial for the selection of treatment options.

[0014] In summary, the existing glaucoma diagnosis model faces the following core challenges: isolated information from various examination modalities, reliance on doctors' experience for subjective integration in the diagnostic process, insufficient sensitivity of early screening, and a lack of efficient and objective quantitative grading methods.

[0015] In view of this, there is an urgent need to design a new auxiliary diagnostic method for glaucoma in order to overcome at least some of the aforementioned shortcomings of existing methods. Summary of the Invention

[0016] This invention provides an artificial intelligence-based glaucoma lesion feature screening system, method, electronic device, and storage medium, which can significantly improve the sensitivity and accuracy of early glaucoma lesion feature screening and classification. It has the advantages of being objective, efficient, and highly interpretable, and is suitable for clinical auxiliary diagnosis and large-scale population screening.

[0017] To solve the above-mentioned technical problems, according to one aspect of the present invention, the following technical solution is adopted:

[0018] An artificial intelligence-based glaucoma lesion feature screening system, the lesion feature screening system comprising:

[0019] The multimodal data acquisition and preprocessing module is used to acquire clinical data and standardize the acquired clinical data to generate input data of uniform specifications; the clinical data includes fundus color images, optical coherence tomography images, and visual field examination images.

[0020] The multimodal feature extraction and fusion module is connected to the multimodal data acquisition and preprocessing module. It is used to extract high-dimensional features from the standardized multimodal images using pre-trained convolutional neural networks with different structures, and to generate information-enhanced joint feature representations through an attention-based fusion network.

[0021] The screening and grading decision module is connected to the multimodal feature extraction and fusion module, and is used to screen and grade glaucoma lesion features based on the joint feature representation through a classifier.

[0022] In one embodiment of the present invention, the multimodal feature extraction and fusion module includes:

[0023] The first convolutional neural network submodule is used to extract morphological features from fundus color photographs;

[0024] The second convolutional neural network submodule is used to extract structural quantization features of OCT images;

[0025] The third convolutional neural network submodule is used to extract visual field inspection functional features;

[0026] The feature fusion unit is used to perform dimensional alignment and cross-modal attention fusion on the features output by the above sub-modules;

[0027] The feature fusion unit is specifically used to increase the dimension of the feature vector output by the third convolutional neural network submodule through a fully connected layer so that it is consistent with the dimension of the feature vectors of other modalities. The multiple feature vectors after dimension alignment are concatenated, and the concatenated feature sequence is input into the Transformer encoder. Its self-attention mechanism is used to calculate the correlation weight between modalities, and the joint feature representation is output.

[0028] As one embodiment of the present invention, the screening and grading decision module further includes a fine classification task of grading the severity of glaucoma lesion characteristics;

[0029] The screening and grading decision-making module adopts a multi-task learning framework, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L total =α·L screen +β·L grade , where α and β are adjustable hyperparameters;

[0030] The optical coherence tomography images are thickness maps of the retinal nerve fiber layer and the macular ganglion cell complex.

[0031] As one embodiment of the present invention, severity grading is determined by a deep learning model based on the joint feature representation F. fusedAutomatic assessment; assesses multimodal information on structural damage, functional damage, and morphological changes, and outputs a comprehensive severity assessment, categorized into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information;

[0032] The lesion feature screening system further includes a data output module, which outputs a structured diagnostic report, including at least screening results, grading results, and a heat map of abnormal areas for visualizing the diagnostic basis.

[0033] The data output module uses gradient-weighted class activation mapping technology to generate diagnostic heatmaps that identify abnormal areas on the input fundus color images and OCT images.

[0034] According to another aspect of the present invention, the following technical solution is adopted: a glaucoma lesion feature screening method based on artificial intelligence, the lesion feature screening method comprising:

[0035] Multimodal data acquisition and preprocessing steps: Acquire and standardize the acquired clinical data such as fundus color images, optical coherence tomography images and visual field examination images to generate input data of uniform specifications;

[0036] Multimodal feature extraction and fusion steps: High-dimensional features are extracted from the standardized multimodal image using pre-trained convolutional neural networks with different structures, and an information-enhanced joint feature representation is generated through an attention-based fusion network;

[0037] Screening and grading decision-making steps: Based on the joint feature representation, the glaucoma lesion features are screened and graded using a classifier.

[0038] As one embodiment of the present invention, the multimodal feature extraction and fusion step includes:

[0039] The first convolutional neural network submodule extracts morphological features from fundus color photographs;

[0040] The second convolutional neural network submodule extracts structural quantization features from OCT images;

[0041] The third convolutional neural network submodule extracts visual field inspection functional features;

[0042] The feature fusion step involves dimensional alignment and cross-modal attention fusion of the features output by the above sub-modules. In the feature fusion step, the feature vector output by the third convolutional neural network sub-module is dimension-enhanced through a fully connected layer to make it consistent with the dimension of the feature vectors of other modalities. The dimension-aligned feature vectors are then concatenated, and the concatenated feature sequence is input into the Transformer encoder. Its self-attention mechanism is used to calculate the inter-modal correlation weights, and the joint feature representation is output.

[0043] As one embodiment of the present invention, the screening and grading decision-making step further includes a fine classification task of grading the severity of glaucoma lesion characteristics;

[0044] In the screening and grading decision-making steps, a multi-task learning framework is adopted, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L. total =α·L screen +β·L grade , where α and β are adjustable hyperparameters;

[0045] The optical coherence tomography images are preferably retinal nerve fiber layer thickness maps and macular ganglion cell complex thickness maps.

[0046] As one embodiment of the present invention, the severity classification is determined by the deep learning model based on the joint feature representation F. fused Automatic assessment; assesses multimodal information on structural damage, functional damage, and morphological changes, and outputs a comprehensive severity assessment, categorized into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information;

[0047] The lesion feature screening method further includes a data output step: outputting a structured diagnostic report, which includes at least screening results, grading results, and a heat map of abnormal areas for visualizing the diagnostic basis;

[0048] In the data output step, gradient-weighted class activation mapping technology is used to generate a diagnostic heatmap that identifies abnormal areas on the input fundus color photograph and OCT image.

[0049] According to another aspect of the present invention, the following technical solution is adopted: an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above method.

[0050] According to another aspect of the present invention, the following technical solution is adopted: a storage medium storing computer program instructions thereon, which, when executed by a processor, implement the steps of the above-described method.

[0051] The beneficial effects of this invention are as follows: The artificial intelligence-based glaucoma lesion feature screening system, method, electronic device and storage medium proposed in this invention can significantly improve the sensitivity and accuracy of early glaucoma lesion feature screening and classification. It has the advantages of objectivity, efficiency and strong interpretability, and is suitable for clinical auxiliary diagnosis and large-scale population screening. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the composition of an artificial intelligence-based glaucoma lesion feature screening system in one embodiment of the present invention.

[0053] Figure 2 This is a flowchart of an artificial intelligence-based glaucoma lesion feature screening method in one embodiment of the present invention.

[0054] Figure 3 This is a schematic diagram of a glaucoma lesion feature screening method based on artificial intelligence in one embodiment of the present invention.

[0055] Figure 4 This is a schematic diagram of the composition of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0056] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0057] To further understand the present invention, preferred embodiments of the present invention are described below in conjunction with examples. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, and not for limiting the scope of the claims of the present invention.

[0058] The description in this section pertains to only a few typical embodiments, and the present invention is not limited to the scope of the embodiments described. Substitution of identical or similar prior art methods with some technical features in the embodiments is also within the scope of the description and protection of this invention.

[0059] The steps described in the various embodiments in the specification are for illustrative purposes only, and the implementation of this application is not limited by the order of the steps.

[0060] The term "connection" in the specification includes both direct and indirect connections, such as connections made through active devices, passive devices, or electrical conduction media; it may also include connections made by other active or passive devices that are known to those skilled in the art and can achieve the same or similar functional purpose, such as connections made through circuits or components such as switches or follower circuits.

[0061] This invention discloses an artificial intelligence-based glaucoma lesion feature screening system. Figure 1 This is a schematic diagram of the composition of an artificial intelligence-based glaucoma lesion feature screening system according to an embodiment of the present invention; please refer to [link / reference]. Figure 1 The lesion feature screening system includes: a multimodal data acquisition and preprocessing module 1, a multimodal feature extraction and fusion module 2, and a screening and grading decision module 3.

[0062] The multimodal data acquisition and preprocessing module 1 is used to acquire clinical data and standardize the acquired clinical data to generate input data of uniform specifications; the clinical data includes fundus color images, optical coherence tomography images and visual field examination images.

[0063] The multimodal feature extraction and fusion module 2 is connected to the multimodal data acquisition and preprocessing module, and is used to extract high-dimensional features from the standardized multimodal images using pre-trained convolutional neural networks with different structures, and generate information-enhanced joint feature representations through a fusion network based on an attention mechanism.

[0064] The screening and grading decision module 3 is connected to the multimodal feature extraction and fusion module, and is used to screen and grade glaucoma lesion features based on the joint feature representation through a classifier.

[0065] In one embodiment of the present invention, the multimodal feature extraction and fusion module 2 includes: a first convolutional neural network submodule, a second convolutional neural network submodule, a third convolutional neural network submodule, and a feature fusion unit.

[0066] The first convolutional neural network submodule is used to extract morphological features from fundus color photographs; the second convolutional neural network submodule is used to extract structural quantization features from OCT images; and the third convolutional neural network submodule is used to extract functional features for visual field examination.

[0067] The feature fusion unit is used to perform dimension alignment and cross-modal attention fusion on the features output by the above sub-modules. Specifically, the feature fusion unit is used to increase the dimension of the feature vector output by the third convolutional neural network sub-module through a fully connected layer so that it is consistent with the dimension of the feature vectors of the other modalities, concatenate the multiple dimension-aligned feature vectors, input the concatenated feature sequence into the Transformer encoder, use its self-attention mechanism to calculate the inter-modal correlation weights, and output the joint feature representation.

[0068] In one embodiment of the present invention, the screening and grading decision module 3 further includes a fine classification task of grading the severity of glaucoma lesion characteristics.

[0069] The screening and grading decision module 3 adopts a multi-task learning framework, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L total =α·L screen +β·L grade α and β are adjustable hyperparameters. The optical coherence tomography images are retinal nerve fiber layer thickness maps and macular ganglion cell complex thickness maps.

[0070] In one embodiment, the severity rating is determined by a deep learning model based on the joint feature representation F. fused Automatic assessment; assesses multimodal information including structural damage (such as the thickness of the retinal nerve fiber layer and ganglion cell complex), functional damage (such as visual field defects), and morphological changes (such as optic disc morphology), and outputs a comprehensive severity assessment, divided into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information;

[0071] The lesion feature screening system further includes a data output module 4, which outputs a structured diagnostic report. The report includes at least screening results, grading results, and a heatmap of abnormal areas for visualizing the diagnostic basis. The data output module 4 uses gradient-weighted class activation mapping technology to generate a diagnostic heatmap identifying abnormal areas on the input fundus color images and OCT images.

[0072] This invention further discloses an artificial intelligence-based method for screening glaucoma lesion characteristics. Figure 2 This is a flowchart of an artificial intelligence-based glaucoma lesion feature screening method according to an embodiment of the present invention. Figure 3 This is a schematic diagram of an artificial intelligence-based glaucoma lesion feature screening method according to an embodiment of the present invention; please refer to [link / reference]. Figure 2 , Figure 3 The lesion feature screening method includes:

[0073]

Step S1

[0074]

Step S2

[0075]

Step S3

[0076] In one embodiment of the present invention, step S2, the multimodal feature extraction and fusion step includes:

[0077] The first convolutional neural network submodule extracts morphological features from fundus color photographs;

[0078] The second convolutional neural network submodule extracts structural quantization features from OCT images;

[0079] The third convolutional neural network submodule extracts visual field inspection functional features;

[0080] The feature fusion step involves dimensional alignment and cross-modal attention fusion of the features output by the above sub-modules. In the feature fusion step, the feature vector output by the third convolutional neural network sub-module is dimension-enhanced through a fully connected layer to make it consistent with the dimension of the feature vectors of other modalities. The dimension-aligned feature vectors are then concatenated, and the concatenated feature sequence is input into the Transformer encoder. Its self-attention mechanism is used to calculate the inter-modal correlation weights, and the joint feature representation is output.

[0081] In step S3, the screening and grading decision-making step further includes a fine classification task of grading the severity of glaucoma lesion characteristics; in the screening and grading decision-making step, a multi-task learning framework is adopted, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L total =α·L screen +β·L grade α and β are adjustable hyperparameters. The optical coherence tomography images are preferably retinal nerve fiber layer thickness maps and macular ganglion cell complex thickness maps.

[0082] Severity grading is determined by the deep learning model based on the joint feature representation F. fused Automatic assessment; assesses multimodal information including structural damage (such as the thickness of the retinal nerve fiber layer and ganglion cell complex), functional damage (such as visual field defects), and morphological changes (such as optic disc morphology), and outputs a comprehensive severity assessment, divided into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information;

[0083] In one embodiment of the present invention, the lesion feature screening method further includes step S4, a data output step: outputting a structured diagnostic report, the report including at least screening results, grading results, and a heatmap of abnormal areas for visualizing the diagnostic basis. In the data output step, gradient-weighted class activation mapping technology is used to generate a diagnostic heatmap identifying abnormal areas on the input fundus color images and OCT images.

[0084] In one application scenario of the present invention, the glaucoma lesion feature screening method based on artificial intelligence of the present invention includes the following steps:

[0085] Step S1. Multimodal data acquisition and preprocessing steps.

[0086] Step S11, Data Source and Labeling Steps;

[0087] In one embodiment, an ethics-approved retrospective clinical dataset was collected, containing complete multimodal imaging data from over 12,500 patients, specifically including fundus photography, optical coherence tomography (OCT) images, and visual field examination data. All data were anonymized.

[0088] A labeling team composed of three or more glaucoma experts with senior professional titles used visual field examination and comprehensive clinical diagnosis as the gold standard to perform double-blind labeling on each sample, generating classification labels (normal, early stage of glaucoma, early stage of glaucoma, intermediate stage of glaucoma, late stage of glaucoma, and end stage of glaucoma). The complete dataset was randomly divided into training, validation, and test sets in a ratio of 7:1.5:1.5 to ensure consistent data distribution.

[0089] Step S12: Multimodal data preprocessing step;

[0090] This step aims to standardize heterogeneous data from different sources into a unified format and space, laying the foundation for subsequent feature extraction.

[0091] Steps for processing fundus photos:

[0092] Optic disc localization and segmentation: The U-Net deep learning network is used to accurately segment the optic disc and optic cup regions, and quantitative morphological parameters such as the vertical cup-to-disc ratio (vCDR) are calculated.

[0093] Region of Interest (ROI) clipping: Using the geometric center of the view disk as a reference, a region of interest (ROI) of a fixed size is clipped out.

[0094] Image normalization: The cropped image is uniformly scaled to 512×512 pixels and Z-score normalization is performed. The normalized image I norm .

[0095]

[0096] Where I represents the original image, and μ and σ represent the mean and standard deviation of the image dataset, respectively.

[0097] OCT image processing steps:

[0098] Thickness map generation steps: Extract the thickness data of the retinal nerve fiber layer (RNFL) and ganglion cell complex (GCC) from the OCT circular scan, and generate a standard-sized (512×512) two-dimensional pseudo-color thickness map using a bicubic interpolation algorithm.

[0099] Numerical normalization step: Normalize the pixel values ​​in the thickness map to the [0,1] interval to eliminate the influence of dimensions.

[0100] Visual field inspection procedure:

[0101] The numerical matrix of visual field examination (e.g., 24-2 or 30-2 pattern) is converted into a normalized grayscale image through linear mapping, where the brightness of each pixel represents the sensitivity value at that site. This transforms the functional examination results into an image format that can be processed by the model. min S is the minimum brightness value. max This is the maximum brightness value.

[0102]

[0103] Where P ij S is the normalized pixel value. ij The sensitivity value is the value for the corresponding site, and the two are directly proportional.

[0104] Step S13, Data Augmentation;

[0105] To avoid model overfitting and improve its generalization ability, a series of random transformations are applied to the training set images during the training phase, including random rotation within ±15°, horizontal / vertical flipping, and brightness and contrast adjustment.

[0106] Step S2. Multimodal feature extraction and fusion.

[0107] This system constructs a deep neural network model based on multimodal feature fusion and multi-task learning. The model takes the preprocessed multimodal data as input, and its core architecture includes three modules in sequence: feature extraction, feature fusion, and cascaded decision.

[0108] Step S21, Multimodal Feature Extraction Step:

[0109] This module consists of three parallel deep subnetworks, designed to extract the most discriminative features from data of different modalities.

[0110] The Color Image Feature Subnetwork (NetF) uses a ResNet-50 model pre-trained on the large natural image dataset ImageNet as the backbone for feature extraction. Its last fully connected layer is removed, and a global average pooling layer is used. The input is a pre-processed image of the optic disc region. Through deep convolutional operations, it outputs a 1024-dimensional feature vector F. fundus This vector is used to encode the macroscopic morphological information of the visual disk.

[0111] F fundus ∈R 1024

[0112] OCT Feature Subnetwork (NetO): Employs a pre-trained DenseNet-121 model as the backbone network. Its densely connected structure facilitates gradient flow and feature reuse, making it particularly suitable for processing complex texture information in OCT thickness maps. The input is an OCT thickness pseudo-color image, and the output is a 1024-dimensional feature vector F. oct It is used to encode the microstructural changes of each layer of the retina.

[0113] F oct ∈R 1024

[0114] The Vision Feature Subnetwork (NetV) employs a lightweight 4-layer convolutional neural network (CNN) with the structure: Conv-BN-ReLU-Pooling×3+FC. The input is the grayscale image of the vision area, and the output is a 512-dimensional feature vector F. vf , used to encode functional defect patterns in the field of vision.

[0115] F vf ∈R 512

[0116] Step S22, Multimodal Feature Fusion Step:

[0117] This module employs a strategy of "feature alignment + feature concatenation + attention enhancement" to achieve deep complementarity and fusion of cross-modal information. It takes the outputs of the aforementioned three feature sub-networks as input to generate an information-enhanced joint feature representation.

[0118] Feature alignment: In order to enable features from different modalities to be efficiently computed in the same vector space, features with mismatched dimensions are first aligned.

[0119] The 512-dimensional feature vector F output by the vision feature subnetwork is processed through a fully connected layer. vf Upgrading to 1024 dimensions yields the aligned feature F'.vf This ensures that all modal features reside in the same vector space.

[0120] F' vf =ReLU(W vf ·F vf +b vf )

[0121] Among them W vf ∈R 512×1024 , is the weight matrix, responsible for linear transformation; b vf ∈R 1024 , which is the bias term, is responsible for providing translational degrees of freedom for the output features; the ReLU activation function is introduced to increase the nonlinear expressive power of the model.

[0122] Feature concatenation: The three aligned 1024-dimensional feature vectors are concatenated to obtain a 3072-dimensional pre-fusion feature F containing all modal information. concat .

[0123] F concat =[F fundus ;F oct ;F' vf ]

[0124] Cross-modal attention fusion: Leveraging the powerful global dependency modeling capabilities of the Transformer encoder, the correlation between features of different modalities is dynamically calculated to achieve information complementarity and enhancement.

[0125] Serialization and positional encoding: Serialization of 3072-dimensional F concat Remodeling into sequence X∈R 3×1024 (The original input sequence consists of feature vectors from three modalities), with the addition of a learnable positional encoding P∈R 3×1024 (Injecting the sequence order information of the modalities into the model). The sequence is added to the position information to form the initial input Z0 of the Transformer encoder.

[0126] Z0 = X + P

[0127] The initial state Z0 is then input into a 2-layer Transformer encoder. The computation of each layer (layer l) consists of the following two core sublayers:

[0128] Multi-head self-attention quantum layer:

[0129] The calculation process of the l-th layer Transformer encoder is as follows:

[0130] Z ' l =LayerNorm(Z) l-1+MultiHeadAttention(Z l-1 Multi-Head Attention allows each modality's features to focus on features from all other modalities. It dynamically and adaptively learns the complex dependencies between fundus photography, OCT, and visual field examination modalities, thereby computing a dynamic, context-sensitive set of relevance weights to enhance important features and suppress redundant information. LayerNorm, short for LayerNormalization, is a standardization technique used to stabilize and accelerate the training process of deep neural networks.

[0131] Feedforward sublayer:

[0132] Z l =LayerNorm(Z) ' l +FFN(Z ' l ))

[0133] FFN stands for Feedforward Network, which typically consists of two fully connected layers and a ReLU activation function.

[0134] After two layers of encoding, the output vector of the first token output from the last Transformer layer (corresponding to the [CLS] token, where CLS stands for Classification Token, a special "information collector" inserted at the beginning of the sequence to aggregate the global context information of the entire sequence at the end of the process) is taken as the final, information-enhanced fusion feature F. fused ∈R 1024 .

[0135] Step S3. Screening and tiered decision-making steps.

[0136] This module receives the joint feature representation F from the fusion module. fused It achieves cascaded decision-making from screening to classification through two parallel specific branches.

[0137] Screening branch: Consists of a fully connected layer and a sigmoid activation function, outputting the probability P that the patient has glaucoma. glaucoma .

[0138] P glaucoma =Sigmoid(W s ·F fused +b s ),W s ∈R 1×1024 ,b s ∈R

[0139] Among them, W sIt is a weight matrix, which can be viewed as an "importance scorer" responsible for determining the importance of each dimension in the fused feature vector to the final screening decision. s It is the bias term, a scalar that provides a reference offset for linear transformations.

[0140] Hierarchical branch: It consists of a fully connected layer and a Softmax activation function, and the output is a probability distribution G = [g0, g1, g2, g3, g4, g5] belonging to six severity levels (normal, early, early, middle, late and terminal).

[0141] G = Softmax(W g ·F fused +b g ),W g ∈R 6×1024 ,b g ∈R 6

[0142] Among them, W g This is a weight matrix, its function is to simultaneously calculate "evidence scores" for each of the six severity levels (normal, early, early, middle, late, and terminal). Each row of the matrix corresponds to a classifier for a severity level. g It is a bias vector, which is a vector containing 6 values, each of which provides a prior bias for the corresponding level.

[0143] Multi-task loss function: The overall loss function for model training is a weighted sum of the screening task loss and the hierarchical task loss.

[0144] L total =α·L screen +β·L grade

[0145] in,

[0146] For binary cross-entropy loss in binary classification, y i This is the true label for sample i (0 indicates normal, 1 indicates suspected glaucoma), p i α is the predicted probability of sample i, N is the number of samples in the batch, and α and β are adjustable hyperparameters used to balance task importance, which are adjusted according to task importance, for example, set to α = 0.4, β = 0.6.

[0147]

[0148] For multi-class classification, the cross-entropy loss is used, where c is the class index (0 to 5, corresponding to normal, pre-, early, mid, late, and terminal stages). I(y) i =c) is an indicator function; its value is 1 when the true rank of sample i equals c, and 0 otherwise. gi,c This represents the predicted probability that sample i belongs to class c.

[0149] Step S4. System deployment and interpretability report generation steps.

[0150] To facilitate clinical integration and application, the trained model is exported to open formats such as ONNX and encapsulated as a high-performance RESTful API service. Hospital Information Systems (HIS / PACS) can call this API to upload anonymized multimodal image data of patients, such as DICOM or standard image formats. After system analysis, the screening results (positive / negative) and grading results (and corresponding confidence levels) in JSON format can be returned within seconds. The interface displays structured reports and visualizations, including {"screening_result":"positive","probability":0.92,"grading_stage":"early","confidence":0.85}. Simultaneously, a corresponding physician workstation web interface is developed for data submission and intuitive display of the structured reports and visualized heatmaps.

[0151] For cases that test positive, Grad-CAM++ technology is used to trace back to the last convolutional layer of the fundus color image and OCT thickness map to generate a heat map that identifies abnormal areas (such as disc rim narrowing and RNFL wedge defects), and then overlays it onto the original image to provide doctors with intuitive diagnostic information.

[0152] In one application scenario of the present invention, the glaucoma lesion feature screening system based on artificial intelligence of the present invention includes the following modules:

[0153] The multimodal data acquisition and preprocessing module can be implemented in accordance with the implementation process of step S1 above.

[0154] The multimodal feature extraction and fusion module can be implemented by referring to the implementation process of step S2 above.

[0155] The screening and grading decision-making module can be implemented in accordance with the implementation process of step S3 above.

[0156] The system deployment and interpretability report generation module can be implemented in accordance with the implementation process of step S4 above.

[0157] This invention also discloses an electronic device, Figure 4 This is a schematic diagram of the composition of an electronic device according to an embodiment of the present invention; please refer to [link / reference]. Figure 4At the hardware level, the electronic device includes a memory, a processor, and at least one communication interface; the processor may be a microprocessor, and the memory may include main memory, such as random access memory (RAM) or non-volatile memory. Of course, the electronic device may also include other hardware as needed.

[0158] The processor, communication interface, and memory can be interconnected via an internal bus. The memory stores programs (including operating system programs and application programs); the programs may include program code, which may include computer operation instructions. The memory may include main memory and non-volatile memory, and provides instructions and data to the processor.

[0159] In one embodiment, the processor can read the corresponding program from non-volatile memory into memory and then run it; the processor can execute the program stored in memory and specifically perform the following operations (e.g. Figure 2 As shown):

[0160]

Step S1

[0161]

Step S2

[0162]

Step S3

[0163] This invention further discloses a storage medium storing computer program instructions, which, when executed by a processor, implement the following steps of the method of this invention (e.g. Figure 2 As shown):

[0164]

Step S1

[0165]

Step S2

[0166]

Step S3

[0167] In summary, the artificial intelligence-based glaucoma lesion feature screening system, method, electronic device, and storage medium proposed in this invention can significantly improve the sensitivity and accuracy of early glaucoma lesion feature screening and grading. It has the advantages of objectivity, efficiency, and strong interpretability, and is suitable for clinical auxiliary diagnosis and large-scale population screening. This invention can automatically, accurately, and efficiently fuse multimodal image information to achieve intelligent auxiliary diagnosis for early glaucoma screening and objective grading.

[0168] It should be noted that this application can be implemented in software and / or a combination of software and hardware; for example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium; for example, RAM memory, magnetic or optical drives, floppy disks, and similar devices. In addition, some steps or functions of this application can be implemented in hardware; for example, as circuitry that cooperates with a processor to perform the various steps or functions.

[0169] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0170] The description and application of the present invention herein are illustrative and not intended to limit the scope of the invention to the embodiments described above. Effects or advantages involved in the embodiments may not be apparent due to various factors, and the description of effects or advantages is not intended to limit the embodiments. Variations and modifications of the embodiments disclosed herein are possible, and various substitutions and equivalents of the components in the embodiments are well known to those skilled in the art. It should be apparent to those skilled in the art that the invention can be implemented in other forms, structures, arrangements, proportions, and with other components, materials, and parts without departing from the spirit or essential characteristics of the invention. Other variations and modifications can be made to the embodiments disclosed herein without departing from the scope and spirit of the invention.

Claims

1. A glaucoma lesion feature screening system based on artificial intelligence, characterized in that, The lesion feature screening system includes: The multimodal data acquisition and preprocessing module is used to acquire clinical data and standardize the acquired clinical data to generate input data of uniform specifications; the clinical data includes fundus color images, optical coherence tomography images, and visual field examination images. The multimodal feature extraction and fusion module is connected to the multimodal data acquisition and preprocessing module. It is used to extract high-dimensional features from the standardized multimodal images using pre-trained convolutional neural networks with different structures, and to generate information-enhanced joint feature representations through an attention-based fusion network. The screening and grading decision module is connected to the multimodal feature extraction and fusion module, and is used to screen and grade glaucoma lesion features based on the joint feature representation through a classifier.

2. The glaucoma lesion feature screening system based on artificial intelligence according to claim 1, characterized in that: The multimodal feature extraction and fusion module includes: The first convolutional neural network submodule is used to extract morphological features from fundus color photographs; The second convolutional neural network submodule is used to extract structural quantization features of OCT images; The third convolutional neural network submodule is used to extract visual field inspection functional features; The feature fusion unit is used to perform dimensional alignment and cross-modal attention fusion on the features output by the above sub-modules; The feature fusion unit is specifically used to increase the dimension of the feature vector output by the third convolutional neural network submodule through a fully connected layer so that it is consistent with the dimension of the feature vectors of other modalities. The multiple feature vectors after dimension alignment are concatenated, and the concatenated feature sequence is input into the Transformer encoder. Its self-attention mechanism is used to calculate the correlation weight between modalities, and the joint feature representation is output.

3. The artificial intelligence-based glaucoma lesion feature screening system according to claim 1, characterized in that: The screening and grading decision-making module further includes a sub-classification task of grading the severity of glaucoma lesion characteristics; The screening and grading decision-making module adopts a multi-task learning framework, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L total =α·L screen +β·L grade , where α and β are adjustable hyperparameters; The optical coherence tomography images are thickness maps of the retinal nerve fiber layer and the macular ganglion cell complex.

4. The glaucoma lesion feature screening system based on artificial intelligence according to claim 3, characterized in that: Severity grading is performed by a deep learning model based on the joint feature representation F. fused Automatic assessment; assesses multimodal information on structural damage, functional damage, and morphological changes, and outputs a comprehensive severity assessment, categorized into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information; The lesion feature screening system further includes a data output module, which outputs a structured diagnostic report, including at least screening results, grading results, and a heat map of abnormal areas for visualizing the diagnostic basis. The data output module uses gradient-weighted class activation mapping technology to generate diagnostic heatmaps that identify abnormal areas on the input fundus color images and OCT images.

5. A method for screening glaucoma lesion characteristics based on artificial intelligence, characterized in that, The lesion feature screening method includes: Multimodal data acquisition and preprocessing steps: Acquire and standardize the acquired clinical data such as fundus color images, optical coherence tomography images and visual field examination images to generate input data of uniform specifications; Multimodal feature extraction and fusion steps: High-dimensional features are extracted from the standardized multimodal image using pre-trained convolutional neural networks with different structures, and an information-enhanced joint feature representation is generated through an attention-based fusion network; Screening and grading decision-making steps: Based on the joint feature representation, the glaucoma lesion features are screened and graded using a classifier.

6. The artificial intelligence-based glaucoma lesion feature screening method according to claim 5, characterized in that: The multimodal feature extraction and fusion steps include: The first convolutional neural network submodule extracts morphological features from fundus color photographs; The second convolutional neural network submodule extracts structural quantization features from OCT images; The third convolutional neural network submodule extracts visual field inspection functional features; The feature fusion step involves dimensional alignment and cross-modal attention fusion of the features output by the above sub-modules. In the feature fusion step, the feature vector output by the third convolutional neural network sub-module is dimension-enhanced through a fully connected layer to make it consistent with the dimension of the feature vectors of other modalities. The dimension-aligned feature vectors are then concatenated, and the concatenated feature sequence is input into the Transformer encoder. Its self-attention mechanism is used to calculate the inter-modal correlation weights, and the joint feature representation is output.

7. The artificial intelligence-based glaucoma lesion feature screening method according to claim 5, characterized in that: The screening and grading decision-making steps further include a detailed classification task of grading the severity of glaucoma lesion characteristics; In the screening and grading decision-making steps, a multi-task learning framework is adopted, and its overall loss function is the weighted sum of the screening task loss function and the grading task loss function, i.e., L. total =α·L screen +β·L grade , where α and β are adjustable hyperparameters; The optical coherence tomography images are preferably retinal nerve fiber layer thickness maps and macular ganglion cell complex thickness maps.

8. The artificial intelligence-based glaucoma lesion feature screening method according to claim 7, characterized in that: Severity grading is determined by the deep learning model based on the joint feature representation F. fused Automatic assessment; assesses multimodal information on structural damage, functional damage, and morphological changes, and outputs a comprehensive severity assessment, categorized into early, intermediate, late, and terminal stages; the automated grading results are highly consistent with the comprehensive diagnostic conclusions made by clinical experts based on multimodal information; The lesion feature screening method further includes a data output step: outputting a structured diagnostic report, which includes at least screening results, grading results, and a heat map of abnormal areas for visualizing the diagnostic basis; In the data output step, gradient-weighted class activation mapping technology is used to generate a diagnostic heatmap that identifies abnormal areas on the input fundus color photograph and OCT image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 5 to 8.

10. A storage medium storing computer program instructions thereon, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 5 to 8.