A cataract vision intelligent prediction system and method based on slit lamp images

Through deep learning technology, analyzing the anterior segment slit lamp images of cataract patients, and constructing multimodal and multitasking models, solving the problem of inaccurate preoperative and postoperative vision assessment in the prior art, and achieving efficient and accurate vision prediction.

CN119174583BActive Publication Date: 2025-05-30SHANTOU UNIV·CHINESE UNIV OF HONG KONG JOINT SHANTOU INT OPHTHALMOLOGY CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411668062.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-30
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately evaluate preoperative and postoperative vision in cataract patients, especially in the case of other fundus diseases, resulting in inaccurate assessment of surgical risk and prediction of vision recovery.

Method used

Using an intelligent prediction system based on deep learning, a multimodal and multitasking model is constructed to predict preoperative and postoperative Best-Corrected Visual Acuity (BCVA) in cataract patients by pre-processing and feature extraction.

Benefits of technology

It achieves rapid and accurate assessment of preoperative and postoperative vision in cataract patients, reduces doctors' subjectivity, improves the accuracy of surgical risk assessment and visual recovery prediction, and is suitable for environments with limited medical resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119174583B_ABST
    Figure CN119174583B_ABST
Patent Text Reader

Abstract

The present invention discloses a cataract vision intelligent prediction system and method based on slit lamp images, which relates to the field of medical image processing. The system includes: an information acquisition module, a data processing module, a model construction module, and a vision prediction module. The information acquisition module is used to obtain the preoperative BCVA and the BCVA one month after surgery of cataract patients. At the same time, the information acquisition module is also used to obtain the anterior segment slit lamp images of cataract patients. The data processing module is used to preprocess the anterior segment slit lamp images to obtain processed data. The model construction module is used to construct a vision prediction model based on the preoperative BCVA, the BCVA one month after surgery, and the anterior segment slit lamp images. The vision prediction module is used to predict the vision of cataract patients before and after surgery by using the vision prediction model. The present invention can quickly and accurately predict the BCVA of cataract patients under different degrees of lens opacity and their postoperative BCVA, improving work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a cataract vision intelligent prediction system and method based on slit lamp images. Background Art

[0002] Currently, surgery is the only effective way to treat cataracts. In the case where visual function is not impaired, most patients have good postoperative vision recovery, while some patients have poor results due to combined other fundus diseases. However, in the case of severe lens opacity, the refractive medium is turbid, and the fundus before surgery is often unclear, and doctors cannot fully understand the fundus conditions before surgery.

[0003] Preoperative vision assessment of cataract patients is of great significance, which provides key information for doctors. Traditional assessment methods usually rely on doctors' experience. Ophthalmologists use a slit lamp microscope to examine, and use the Lens Opacity Classification System III (LOCS III) to judge the degree of lens opacity, evaluate the consistency between the best corrected visual acuity (BCVA) corresponding to this degree of opacity and the actual BCVA, judge whether cataract is the main cause of the patient's vision decline, identify potential fundus diseases, assist in evaluating the patient's surgical risk, and formulate personalized treatment plans for the patient.

[0004] Meanwhile, due to the clinical experience of doctors, traditional vision assessment methods are highly subjective. In addition, in patients with combined other fundus diseases, their postoperative vision is usually limited by the fundus conditions, which makes it difficult for doctors to accurately predict the postoperative vision recovery amplitude of such patients.

[0005] The advent of the era of artificial intelligence (AI) provides the possibility to solve the above problems and contradictions. As a popular field of artificial intelligence, deep learning simulates the physiological functions of the human brain through multi-layer and multi-functional artificial neural networks, and understands complex patterns from a wide range of data sets. Compared with traditional algorithms, it can better process large-scale and high-dimensional data, and is widely used in the fields of computer vision, image processing, natural language processing, etc. The diagnosis and treatment of ophthalmic diseases highly rely on imaging technology and visual data, which makes deep learning have great prospects in the field of ophthalmology. Through the learning of large-scale medical image data, accurate segmentation of the lesion area and extraction of image features, deep learning realizes the automated analysis of medical images and provides accurate and efficient medical services.

[0006] Currently, for the diagnosis and grading of cataracts, the narrow-band photography method, diffuse illumination photography method, and retroillumination method of the anterior segment slit lamp examination image are mainly used, and the evaluation is carried out with reference to the LOCS III standard image. At the same time, a rough prediction of visual acuity is made according to the different shapes of the lens in the image. However, there is no artificial intelligence system that uses the anterior segment slit lamp examination image to evaluate the current BCVA and predict the BCVA after surgery. Therefore, in view of the above problems, it is of great significance to provide a method and system for quickly, efficiently, simply, and accurately predicting the visual acuity of cataract patients before and after surgery. Summary of the Invention

[0007] The present invention aims to overcome the above-mentioned deficiencies of the prior art and provides an intelligent cataract visual acuity prediction system and method based on slit lamp images. According to the anterior segment slit lamp examination image, the preoperative BCVA and the postoperative BCVA are quickly and accurately evaluated, and important information is efficiently provided for ophthalmologists in the case of limited medical resources, so as to formulate a reasonable treatment plan for patients.

[0008] To achieve the above object, the present invention provides an intelligent cataract visual acuity prediction system based on slit lamp images, including: an information acquisition module, a data processing module, a model construction module, and a visual acuity prediction module;

[0009] The information acquisition module is used to obtain the preoperative BCVA of cataract patients and the BCVA one month after surgery; at the same time, the information acquisition module is also used to obtain the anterior segment slit lamp images of cataract patients, and the anterior segment slit lamp images include: slit lamp narrow-band photography, slit lamp diffuse illumination, and slit lamp retroillumination;

[0010] The data processing module is used to preprocess the anterior segment slit lamp image to obtain the processed data. The process includes: first, screening the anterior segment slit lamp image to obtain a screened image; then, constructing an ROI segmentation model and inputting the screened image into the ROI segmentation model to extract the ROI image; adjusting the resolution of the ROI image to 224*224, and at the same time normalizing the pixel values of the ROI image; converting the visual acuity data of the preoperative BCVA and the BCVA one month after surgery into the logMAR recording method and performing normalization processing;

[0011] The model construction module is used to construct a visual acuity prediction model based on the preoperative BCVA, the BCVA one month after surgery, and the anterior segment slit lamp images. The visual acuity prediction model includes: a multi-modal model and a multi-task model. Among them, the multi-modal model is based on EfficientNetV2-M as the basic model, processes the images of each modality in the anterior segment slit lamp images, and extracts the corresponding high-dimensional feature vectors. In the multi-task model, it is used to output the preoperative predicted visual acuity and the postoperative predicted visual acuity simultaneously. The loss functions of the two tasks are both mean square error MSELoss, and the total loss is the weighted sum of the losses of the two tasks.

[0012] The visual acuity prediction module is used to predict the visual acuity of cataract patients before and after surgery by using the visual acuity prediction model.

[0013] Preferably, the narrow slit light illumination method or the posterior retroillumination method is used to obtain the anterior segment slit lamp images. The shooting requirements of the anterior segment slit lamp images include: the patient is fully dilated before taking the photo, the pupil diameter is dilated to more than 5 mm, the position of the pupil area in the image is centered and unobstructed, and the focus point is adjusted to make the examination area clearly visible.

[0014] Preferably, the data processing module uses ResNet34 as the front-end network for feature extraction and U-Net as the back-end network for feature fusion to construct the ROI segmentation model based on Mask R-CNN.

[0015] Input the screened images into the ROI segmentation model. First, pass through the ResNet34 front-end, use Relu as the activation function, perform residual learning between two layers during the feature extraction process. ResNet34 performs 5 times of downsampling in total, consisting of one 7×7 convolution operation and 4 convolution blocks. After the last convolution in the encoding area, the obtained feature map will be used as the input of the back-end U-Net. In the decoder path, retain the 4 upsampling processes of the original U-Net. Each decoding layer will splice the feature map of the current layer with the feature map of the corresponding encoding layer to achieve feature fusion at the channel level. Finally, perform 1 more upsampling to restore the original image size and obtain the feature map.

[0016] Input the generated feature map into the RPN region proposal network of Mask R-CNN to generate candidate regions; perform alignment and normalization on the candidate regions through the ROI Align layer to ensure that the size pixels within each candidate region are consistent. Finally, perform classification and regression through the fully connected layer and generate mask predictions through the fully convolutional network. In the prediction stage, when the confidence level output by the MaskR-CNN model is greater than the threshold, it indicates that the ROI has been detected.

[0017] Preferably, the EfficientNetV2-M convolutional neural network includes: a Fused-MBConv module and an MBConv module;

[0018] First, the input ROI image is pre-processed for features through a standard convolutional layer with a stride of 2, a size of 3x3, and 24 channels, quickly reducing the image resolution, and then passing through 4 groups of Fused-MBConv modules and 4 groups of MBConv modules for feature learning and representation; the Fused-MBConv module is composed of a 3x3 convolution, an SE module, and a 1x1 convolution, and the MBConv module is stacked by a 1x1 convolution, a 3x3 depth convolution, an SE module, and a 1x1 convolution;

[0019] After that, the ROI image passes through a 1x1 convolution to increase the number of channels, and through a global average pooling layer, the feature maps of each channel are averaged to obtain a 1x1 feature vector, and the feature vector is input into the fully connected layer;

[0020] Finally, at the backend, different modality features are fused by splicing the fully connected layer and the ReLU activation function; taking the fused feature vector as the input, it is respectively passed to two multi-layer MLPs, one MLP is used to predict the preoperative vision, and the other MLP is used to predict the postoperative vision.

[0021] Preferably, in the multi-task model, the preoperative predicted vision and the postoperative predicted vision are output simultaneously, and the loss functions of both tasks are the mean square error MSELoss, and the total loss is the weighted sum of the losses of the two tasks, that is:

[0022] Loss = 0.5*Loss1 + 0.5*Loss2,

[0023] In the formula, Loss1 is the loss function of task 1; Loss2 is the loss function of task 2; the coefficient 0.5 is the weighting coefficient; Loss is the total loss function, that is, the weighted sum of the losses of the two tasks, and the total loss function is used to measure the overall error in the model training process. The model adjusts the parameters to minimize the loss function and improve the model performance.

[0024] Preferably, the vision prediction model uses the Adam optimizer, and the batch size is 16; under the Ubuntu20 operating system, programming is based on the Pytorch framework, and the training set and the validation set are input to complete the training and internal validation of the model.

[0025] The present invention also provides a method for intelligent cataract vision prediction based on slit lamp images, and the method is applied to the above system, and the steps include:

[0026] S1. Obtain the preoperative BCVA and the BCVA one month after surgery of cataract patients; at the same time, obtain the anterior segment slit lamp images of cataract patients, and the anterior segment slit lamp images include: narrow-band slit lamp illumination, diffuse slit lamp illumination, and retroillumination slit lamp illumination.

[0027] S2. Preprocess the anterior segment slit lamp images to obtain the processed data. The process includes: first, screen the anterior segment slit lamp images to obtain screened images; then, construct an ROI segmentation model and input the screened images into the ROI segmentation model to extract ROI images; adjust the resolution of the ROI images to 224*224, and at the same time normalize the pixel values of the ROI images; convert the visual acuity data of the preoperative BCVA and the BCVA one month after surgery into the logMAR recording method and perform normalization processing at the same time.

[0028] S3. Construct a visual acuity prediction model based on the preoperative BCVA, the BCVA one month after surgery, and the anterior segment slit lamp images. The visual acuity prediction model includes: a multi-modal model and a multi-task model; among them, the multi-modal model is based on EfficientNetV2-M as the basic model, processes the images of each modality in the anterior segment slit lamp images, and extracts the corresponding high-dimensional feature vectors; in the multi-task model, it is used to output the preoperative predicted visual acuity and the postoperative predicted visual acuity at the same time, and the loss functions of the two tasks are both mean square error MSELoss, and the total loss is the weighted sum of the losses of the two tasks.

[0029] S4. Use the visual acuity prediction model to predict the visual acuity of cataract patients before and after surgery.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] (1) The present invention can, according to the preoperative anterior segment slit lamp images, train the model by using deep learning algorithms for a large amount of data, quickly and accurately predict the corresponding BCVA and postoperative BCVA under different lens opacity degrees, and assist clinicians to quickly and accurately evaluate the main causes and surgical effects.

[0032] (2) The present invention can intelligently predict the postoperative visual acuity of cataract patients complicated with fundus diseases, better understand the different degrees of postoperative visual acuity recovery under different fundus conditions, assist doctors in clinical decision-making, give personalized diagnosis and treatment plans, help patients establish reasonable surgical expectations, avoid unnecessary doctor-patient disputes, and has important scientific research value.

[0033] (3) By using computer vision technology, the present invention can quickly process a large amount of image data, significantly improve work efficiency, reduce the workload of ophthalmologists, lower medical costs, and shorten the patient's visit time. Due to its strong universality and ease of understanding and mastery, the present invention can be widely applied to medical and health institutions at all levels, helping to narrow the difference in the diagnosis and treatment levels between different medical institutions and improve the overall medical environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] To more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 It is a schematic diagram of the system structure of an embodiment of the present invention;

[0036] Figure 2 It is a schematic diagram of the ROI segmentation model framework of an embodiment of the present invention;

[0037] Figure 3 It is a schematic diagram of the multi-model network framework of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0040] Embodiment 1

[0041] As Figure 1 shown, it is a schematic diagram of the system structure of this embodiment, including: an information acquisition module, a data processing module, a model construction module, and a vision prediction module;

[0042] The following will, in combination with this embodiment, detail how the present invention solves the technical problems in actual work.

[0043] First, use the information acquisition module to obtain the BCVA of cataract patients before surgery and the BCVA one month after surgery; at the same time, the information acquisition module is also used to obtain the anterior segment slit lamp images of cataract patients. In this embodiment, the anterior segment slit lamp images include: slit lamp narrow-band illumination, slit lamp diffuse illumination, and slit lamp retroillumination.

[0044] Specifically, use a Zeiss slit lamp microscope to examine the anterior segment of the subject's eye to obtain anterior segment slit lamp images, which include slit lamp narrow-band illumination, slit lamp diffuse illumination, and slit lamp retroillumination. Specifically, adjust the chin rest of the slit lamp and the position of the patient's head, adjust the picture position so that the pupil area is located in the center of the examination interface, and adjust the focus point so that the examination site is clearly visible, and obtain one image each of slit lamp narrow-band illumination, slit lamp diffuse illumination, and slit lamp retroillumination. Among them, the shooting requirements for the anterior segment slit lamp images are as follows: the patient needs to be fully dilated before taking pictures, the pupil diameter is dilated to more than 5 mm, the position of the pupil area in the image is centered and unobstructed, and the focus point is adjusted so that the examination site is clearly visible. When using the narrow slit light illumination method, set the slit width to 0.2 - 0.3 mm, and the mirror arm angle is about 30°; when using posterior retroillumination, the slit is in the shape of a short rectangle or a half moon, with a height of 2 mm, a width of a narrow slit, about 1 mm, and the mirror arm angle is about 90°. This embodiment includes 1603 samples and a total of 4809 images. After that, collect the preoperative BCVA of the subject and the BCVA one month after surgery as the visual acuity data before and after surgery for subsequent processing.

[0045] After that, the data processing module preprocesses the eye image data to obtain the processed data.

[0046] First, screen the anterior segment slit lamp images, and it is required to select images with clear focus of the eye structure, the position of the pupil area in the image is centered and unobstructed, the pupil diameter is dilated to more than 5 mm, and the lens structure is clearly visible.

[0047] The degree of cataract lesions and the impact on vision in the anterior segment slit lamp images mainly depend on the lens in the pupil area, and this part accounts for a small proportion in the slit lamp narrow-band images. Therefore, it is necessary to first segment the meaningful areas (i.e., regions of interest, ROI) in the selected images (hereinafter referred to as screening images). The specific steps include:

[0048] First, use labelme software to label the ROI of 100 slit lamp narrow-band illuminations, slit lamp diffuse illuminations, and slit lamp retroilluminations respectively; treat the extracted ROI as an instance segmentation problem, use ResNet34 as the front-end network for feature extraction, and use U-Net as the back-end network for feature fusion to construct an ROI (Region of Interest) segmentation model based on Mask R-CNN.

[0049] The screened images are input into the ROI segmentation model. First, they pass through the ResNet34 front-end, with Relu as the activation function, and residual learning between two layers is performed during the feature extraction process. ResNet34 performs downsampling 5 times in total, consisting of one 7×7 convolution operation and 4 convolution blocks. After the last convolution in the encoding area, the obtained feature map will be used as the input of the backend U-Net. In the decoder path, the original 4 upsampling processes of U-Net are retained. Each decoding layer will concatenate the feature map of the current layer with the feature map of the corresponding encoding layer to achieve feature fusion at the channel level. Finally, one more upsampling is performed to restore the original image size and obtain the feature map.

[0050] The generated feature map is input into the RPN (Region Proposal Network) of Mask R-CNN to generate candidate regions. The candidate regions are aligned and normalized through the ROI Align layer to ensure that the pixel sizes within each candidate region are consistent. Finally, classification and regression are performed through the fully connected layer, and mask predictions are generated through the fully convolutional network. In the prediction stage, if the confidence level output by the Mask R-CNN model is greater than a certain threshold (set to 0.75 in this system), it indicates that the ROI has been detected. Using the developed ROI segmentation model, the ROI images of all datasets are segmented and extracted as the input of the prediction model. The framework diagram of the ROI segmentation model is as Figure 2 shown.

[0051] The resolution of the segmented ROI images is adjusted to 224*224, and at the same time, the pixel values of the images are normalized. Then, the visual acuity of the prediction data (pre-operative and post-operative visual acuity data) is converted to the logMAR recording method and also normalized. The datasets obtained from the above two steps are randomly segmented into a training set, a validation set, and a test set at the subject level, with a segmentation ratio of 75:10:15.

[0052] Use the model construction module to construct a visual acuity prediction model based on deep learning.

[0053] This visual acuity prediction model includes a multi-modal model and a multi-task model. In this embodiment, there are three modal data, namely slit lamp narrow-band illumination, slit lamp diffuse illumination, and slit lamp retroillumination, as inputs, so a multi-modal model needs to be constructed. At the same time, the output prediction data includes pre-operative visual acuity and post-operative visual acuity, so a multi-task model needs to be constructed. The specific process is as follows:

[0054] Select EfficientNetV2-M as the base model to process the images of each modality and extract the corresponding high-dimensional feature vectors. The EfficientNetV2-M convolutional neural network includes: Fused-MBConv modules and MBConv modules.

[0055] First, the input ROI image undergoes feature preprocessing through a standard convolutional layer with a stride of 2, a size of 3x3, and 24 channels, quickly reducing the image resolution, and then undergoes feature learning and representation through 4 groups of Fused-MBConv modules and 4 groups of MBConv modules as shown in the figure; the Fused-MBConv module is composed of a 3x3 convolution, an SE module, and a 1x1 convolution, and the MBConv module is stacked by a 1x1 convolution, a 3x3 depth convolution, an SE module, and a 1x1 convolution.

[0056] After that, the ROI image passes through a 1x1 convolution to increase the number of channels and goes through a global average pooling layer to average the feature maps of each channel, obtaining a 1x1 feature vector. The feature vector is input into the fully connected layer.

[0057] Finally, at the backend, different modality features are fused by concatenating the fully connected layer and the ReLU activation function; the fused feature vector is used as the input and is respectively passed to two multi-layer perceptrons (MLPs), one MLP for predicting preoperative vision and the other MLP for predicting postoperative vision. The multi-model network framework diagram is as Figure 3 shown.

[0058] In the multi-task model, the preoperative BCVA and the BCVA one month after surgery are input into the multi-task model for training.

[0059] The outputs of the multi-task model are the predicted preoperative vision and the predicted postoperative vision. The loss functions for both tasks are the mean squared error MSELoss, and the total loss is the weighted sum of the losses of the two tasks, that is:

[0060] loss = 0.5*loss1 + 0.5*loss2,

[0061] where Loss1 is the loss function of task 1; Loss2 is the loss function of task 2; the coefficient 0.5 is the weighting coefficient; Loss is the total loss function, that is, the weighted sum of the losses of the two tasks. The total loss function is used to measure the overall error during the model training process. The model adjusts the parameters to minimize the loss function and improve the model performance.

[0062] In this embodiment, the vision prediction model uses the Adam optimizer with a batch size of 16; under the Ubuntu20 operating system, programming is based on the Pytorch framework. The training set and the validation set are input, and the model is run on two GPUs with 24G of video memory for 500 epochs to complete the training and internal validation of the model.

[0063] The performance of the model is evaluated using the test set, and the evaluation metric is the mean absolute error MAE, the coefficient of determination R2 , the mean squared error MSE.

[0064] The final visual acuity prediction module uses a visual acuity prediction model to predict the visual acuity of cataract patients before and after surgery.

[0065] Embodiment 2

[0066] This embodiment also provides an intelligent cataract visual acuity prediction method based on slit lamp images. The steps include:

[0067] S1. Obtain the BCVA of cataract patients before surgery and the BCVA one month after surgery; at the same time, obtain the anterior segment slit lamp images of cataract patients. The anterior segment slit lamp images include: slit lamp narrow-band illumination, slit lamp diffuse illumination, and slit lamp retroillumination;

[0068] S2. Preprocess the anterior segment slit lamp images to obtain the processed data. The process includes: first, screen the anterior segment slit lamp images to obtain screened images; then, construct an ROI segmentation model and input the screened images into the ROI segmentation model to extract ROI images; adjust the resolution of the ROI images to 224*224, and at the same time normalize the pixel values of the ROI images; convert the visual acuity data of BCVA before surgery and BCVA one month after surgery into the logMAR recording method and perform normalization processing at the same time;

[0069] S3. Construct a visual acuity prediction model based on the BCVA before surgery, the BCVA one month after surgery, and the anterior segment slit lamp images. The visual acuity prediction model includes: a multi-modal model and a multi-task model; among them, the multi-modal model uses EfficientNetV2-M as the basic model to process the images of each modality in the anterior segment slit lamp images and extract the corresponding high-dimensional feature vectors; in the multi-task model, it is used to simultaneously output the predicted visual acuity before surgery and the predicted visual acuity after surgery. The loss functions of the two tasks are both mean squared error MSELoss, and the total loss is the weighted sum of the losses of the two tasks;

[0070] S4. Use the visual acuity prediction model to predict the visual acuity of cataract patients before and after surgery.

[0071] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A cataract vision intelligent prediction system based on slit lamp images, characterized in that: include: Information collection module, data processing module, model building module and vision prediction module; The information acquisition module is used to obtain the BCVA of cataract patients before surgery and the BCVA one month after surgery; at the same time, the information acquisition module is also used to obtain the slit lamp image of the anterior segment of the eye of the cataract patient, and the slit lamp image of the anterior segment of the eye includes: slit lamp narrow broadband illumination, slit lamp diffuse illumination and slit lamp back illumination; The data processing module is used to pre-process the anterior segment slit lamp image to obtain processed data, and the process includes: firstly, screening the anterior segment slit lamp image to obtain a screening image; then, constructing a ROI segmentation model and inputting the screening image into the ROI segmentation model and extracting the ROI image; adjusting the resolution of the ROI image to 224×224, and normalizing the pixel value of the ROI image; converting the visual acuity data of the BCVA before surgery and the BCVA one month after surgery into the logMAR recording method, and performing normalization processing at the same time; The model building module is used to build a vision prediction model based on the BCVA before surgery, the BCVA one month after surgery and the anterior segment slit lamp image, and the vision prediction model includes: a multimodal model and a multi-task model; wherein the multimodal model uses EfficientNetV2-M as a basic model, processes the images of each modality in the anterior segment slit lamp image, and extracts the corresponding high-dimensional feature vector; in the multi-task model, the loss functions of the two tasks for simultaneously outputting the preoperative predicted vision and the postoperative predicted vision are both mean square error MSELoss, and the total loss is the weighted sum of the two task losses; the EfficientNetV2-M convolutional neural network includes: a Fused-MBConv module and an MBConv module; First, the input ROI image is preprocessed through a standard convolution layer with a stride of 2, a size of 3x3, and a channel number of 24 to quickly reduce the resolution of the image. Then, it is passed through 4 groups of Fused-MBConv modules and 4 groups of MBConv modules for feature learning and representation. The Fused-MBConv module is composed of a combination of 3x3 convolution, SE module, and 1x1 convolution, and the MBConv module is composed of a stack of 1x1 convolution, 3x3 depth convolution, SE module, and 1x1 convolution. After that, the ROI image passes through a 1x1 convolution to increase the number of channels, and passes through a global average pooling layer to average the feature maps of each channel to obtain a 1x1 feature vector, which is input into the fully connected layer; Finally, the features of different modalities are fused at the back end by splicing fully connected layers and ReLU activation functions. The fused feature vector is passed as input to two multi-layer MLPs, one for predicting preoperative visual acuity and the other for predicting postoperative visual acuity. The vision prediction module is used to predict the vision of cataract patients before and after surgery using the vision prediction model.

2. The cataract vision intelligent prediction system based on slit lamp images according to claim 1, characterized in that: The anterior segment of the eye slit lamp image is obtained using a narrow slit light illumination method or a rear reflective illumination method. The requirements for shooting the anterior segment of the eye slit lamp image include: the patient's pupil is fully dilated before shooting, the pupil diameter is dilated to more than 5 mm, the pupil area of ​​the image is centered and unobstructed, and the focus is adjusted so that the inspection area is clearly visible.

3. The cataract vision intelligent prediction system based on slit lamp images according to claim 1, characterized in that: The data processing module uses ResNet34 as the front-end network for feature extraction and U-Net as the back-end network for feature fusion to construct the ROI segmentation model based on Mask R-CNN; The screened image is input into the ROI segmentation model. First, it is passed through the ResNet34 front end, with Relu as the activation function, and residual learning between the two layers is performed during the feature extraction process. ResNet34 performs a total of 5 downsamplings, consisting of a 7×7 convolution operation and 4 convolution blocks. After the last convolution in the coding area, the obtained feature map will be used as the input of the back-end U-Net. In the decoder path, the 4 upsampling processes of the original U-Net are retained. Each decoding layer will splice the feature map of the current layer with the feature map of the corresponding coding layer to achieve feature fusion at the channel level. Finally, upsampling is performed again to restore the original image size to obtain the feature map; The generated feature map is input into the RPN region proposal network of Mask R-CNN to generate candidate regions; The candidate regions are aligned and normalized through the ROIAlign layer to ensure that the size pixels in each candidate region are consistent. Finally, classification and regression are performed through the fully connected layer, and mask prediction is generated through the fully convolutional network. In the prediction stage, when the confidence output by the Mask R-CNN model is greater than the threshold, it means that the ROI is detected.

4. The cataract vision intelligent prediction system based on slit lamp images according to claim 1, characterized in that: In the multi-task model, the preoperative predicted visual acuity and the postoperative predicted visual acuity are output simultaneously. The loss functions of the two tasks are both mean square error MSELoss, and the total loss is the weighted sum of the two task losses, that is: Loss = 0.5*Loss1 + 0.5*Loss2, In the formula, Loss1 is the loss function of task 1; Loss2 is the loss function of task 2; the coefficient 0.5 is the weighting coefficient; Loss is the total loss function, that is, the weighted sum of the losses of the two tasks. The total loss function is used to measure the overall error in the model training process. The model minimizes the loss function and improves model performance by adjusting parameters.

5. The cataract vision intelligent prediction system based on slit lamp images according to claim 1, characterized in that: The vision prediction model uses the Adam optimizer with a batch size of 16. Under the Ubuntu 20 operating system, it is programmed based on the Pytorch framework, and the training set and validation set are input to complete the training and internal validation of the model.

6. A method for intelligent prediction of cataract vision based on slit lamp images, the method being applied to the system according to any one of claims 1 to 5, characterized in that the steps include: S1. Obtain the BCVA of the cataract patient before surgery and the BCVA one month after surgery; and simultaneously obtain the anterior segment slit lamp image of the cataract patient, wherein the anterior segment slit lamp image includes: slit lamp narrow broadband illumination, slit lamp diffuse illumination, and slit lamp back illumination; S2. Preprocessing the anterior segment slit lamp image to obtain processed data, the process comprising: firstly screening the anterior segment slit lamp image to obtain a screening image; then, constructing a ROI segmentation model and inputting the screening image into the ROI segmentation model and extracting the ROI image; adjusting the resolution of the ROI image to 224×224, and normalizing the pixel values ​​of the ROI image; converting the visual acuity data of the BCVA before surgery and the BCVA one month after surgery into the logMAR recording method, and performing normalization processing at the same time; S3. A vision prediction model is constructed based on the BCVA before surgery, the BCVA one month after surgery, and the anterior segment slit lamp image. The vision prediction model includes: a multimodal model and a multi-task model; wherein the multimodal model uses EfficientNetV2-M as a basic model, processes the images of each modality in the anterior segment slit lamp image, and extracts the corresponding high-dimensional feature vector; in the multi-task model, the loss functions of the two tasks for simultaneously outputting the preoperative predicted vision and the postoperative predicted vision are both mean square error MSELoss, and the total loss is the weighted sum of the two task losses; the EfficientNetV2-M convolutional neural network includes: a Fused-MBConv module and an MBConv module; First, the input ROI image is preprocessed through a standard convolution layer with a stride of 2, a size of 3x3, and a channel number of 24 to quickly reduce the resolution of the image. Then, it is passed through 4 groups of Fused-MBConv modules and 4 groups of MBConv modules for feature learning and representation. The Fused-MBConv module is composed of a combination of 3x3 convolution, SE module, and 1x1 convolution, and the MBConv module is composed of a stack of 1x1 convolution, 3x3 depth convolution, SE module, and 1x1 convolution. After that, the ROI image passes through a 1x1 convolution to increase the number of channels, and passes through a global average pooling layer to average the feature maps of each channel to obtain a 1x1 feature vector, which is input into the fully connected layer; Finally, the features of different modalities are fused at the back end by splicing fully connected layers and ReLU activation functions. The fused feature vector is passed as input to two multi-layer MLPs, one for predicting preoperative visual acuity and the other for predicting postoperative visual acuity. S4. Use the vision prediction model to predict the vision of cataract patients before and after surgery.

Citation Information

Patent Citations

  • High myopia cataract surgery prognosis intelligent pre-judging system

    CN109998477A

  • Eye image recognition method based on multi-task learning and related equipment

    CN116563932A