Deep learning-based intraocular lens size selection method for posterior chamber of eye with lens

By building a dual-trunk multimodal feature fusion network model based on deep learning and combining it with Pentacam and UBM images, the problem of inaccuracy in ICL lens size selection was solved, efficient and explainable intelligent decision-making was achieved, surgical risks were reduced, and the precision of refractive correction surgery was promoted.

CN120600330APending Publication Date: 2025-09-05CHONGQING PURUI EYE HOSPITAL CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510571158.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing ICL lens size selection methods rely on manual measurement and experience, which are inaccurate, incomplete and subjective, resulting in a high risk of surgical complications and a lack of intelligent decision-making assistance.

Method used

A deep learning-based method was used, combined with Pentacam and UBM images, to construct a dual-backbone multimodal feature fusion network model. Through local feature enhancement and multimodal feature cross-fusion, key structural information was automatically extracted to achieve accurate selection of ICL crystal size.

Benefits of technology

It improves the accuracy and explainability of ICL lens size selection, provides an efficient and reliable intelligent decision-making solution, reduces surgical risks, and promotes the precision and intelligence of refractive correction surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600330A_ABST
    Figure CN120600330A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based intraocular lens size selection method for a posterior chamber of a lens, and relates to the technical field of medical image processing and intelligent aid decision making. The method at least comprises the following steps: S1, tracking, collecting and sorting anterior segment image historical data, marking a corresponding ICL crystal size category for an image data set through an operation success medical record of a patient, and performing division and data preprocessing on the marked image data set to ensure high quality and accuracy of the data; according to the method, Pentacam and UBM images are fused for the first time, the limitation of an existing ICL prediction technology is broken through, multi-modal image fusion, global-local information joint modeling and visual interpretation are achieved, higher prediction accuracy, better interpretability and wider clinical applicability are achieved, and a set of efficient and reliable intelligent auxiliary decision-making scheme is provided for ICL wafer selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing and intelligent auxiliary decision-making, and is mainly used for selecting the size of phakic posterior chamber intraocular lenses in ophthalmic refractive surgery. Background Art

[0002] Phakic posterior chamber intraocular lens (ICL) implantation is a widely used refractive correction procedure that primarily corrects refractive errors by implanting a miniature ICL lens into the myopic eye. Before surgery, specialists need to conduct preoperative planning and select the size of the ICL lens based on the patient's biological information about the operated eye. This process may be related to many factors, including the patient's age, corneal endothelium, iris morphology, ciliary body shape, axial length, and eye diseases. If the selected lens size is inappropriate, it may cause a series of surgical complications, which in turn may harm the patient's eye health.

[0003] In current clinical practice, although numerous tools exist to assist in selecting ICL lens size, existing analytical methods are often complex and rely heavily on the accuracy of manual measurement and the reliability of physician experience, lacking intelligent decision-making support. More often than not, physicians rely solely on their own experience to select the appropriate ICL lens. Typically, physicians rely on a combination of imaging information from the anterior segment assessment system (Pentacam), ultrasound biomicroscopy (UBM), and anterior segment optical coherence tomography (AS-OCT) in their decision-making process. Pentacam images can provide information on anterior segment structures such as corneal curvature and anterior chamber depth, but they cannot accurately visualize the ciliary body. UBM images can clearly visualize structures such as the ciliary sulcus and iris, but their overall stability is limited, potentially leading to image distortion and significant anteroposterior measurement errors. While AS-OCT can clearly visualize the implanted ICL, it also cannot accurately measure the ciliary body. Therefore, AS-OCT is generally used primarily for postoperative ICL examinations.

[0004] Therefore, the current method, which relies on manual measurement of key parameters combined with human judgment, suffers from drawbacks such as incompleteness, inaccuracy, and subjectivity. Therefore, in order to effectively improve the success rate of refractive surgery and better protect the rights of patients, it is imperative to seek and develop a new solution that can overcome the shortcomings of existing methods and promote the development of refractive surgery in a more accurate and efficient direction. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning to solve the problems raised in the background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning, comprising at least the following steps:

[0007] S1: Track, collect, and organize anterior segment imaging data. Use the patient's successful surgical records to label the image dataset with the corresponding ICL lens size category. Then, perform data segmentation and data preprocessing on the labeled image dataset to ensure high data quality and accuracy.

[0008] S2: Building a local feature augmentation module (Local Feature Augmentation Module), which enhances the local details of the ciliary body in the UBM image to improve the ability to capture details;

[0009] S3: Build a dual-stream multimodal feature fusion network model (Dual-Stream Multimodal Feature Fusion Network), which is used to fuse the two modal images and extract features at different levels of the two images;

[0010] S4: Design a multimodal feature fusion module, which aims to effectively combine global features and local detail features to ensure their complementarity and synergy.

[0011] S5: A selection network model is constructed based on the local feature enhancement module, the dual-backbone multimodal feature fusion network model, and the multimodal feature cross-fusion module. The processed multimodal images are input into the preset network model for training, and the optimal model weights are saved after training is completed.

[0012] S6: Introduce an interpretability module at the final stage of selecting the network model to map the fused deep features back to the original pixel space to help doctors understand the basis of the prediction results and the decision-making process.

[0013] Furthermore, the image history data in S1 includes at least Pentacam horizontal section images, UBM global images and UBM local image data sets, of which 70% are used as training sets, and the remaining 20% ​​and 10% are used as validation sets and test sets respectively; the UBM global images are radial scanning section images taken from 3 to 9 o'clock, and all medical image data are derived from Purui Eye Hospital.

[0014] Furthermore, the feature enhancement module in S2 at least includes the following steps:

[0015] First, the UBM local image is rotated and scaled to the corresponding position of the UBM global image using the SIFT image matching algorithm.

[0016] Then, the UBM global and local images are decomposed through the Laplacian pyramid:

[0017] P UBM全局 ={L1,L2,L3,G3}

[0018] P UBM局部 ={L′1,L′2,L′3,G3′}

[0019] Among them: L1, L2, L3 are the low-frequency layers of the global image; G3 is the highest-frequency layer of the global image; L′1, L′2, L′3 are the low-frequency layers of the local image; G3′ is the highest-frequency layer of the local image;

[0020] Then, only the decomposed high-frequency layers are fused to enhance the details. The formula is as follows:

[0021] L fina =αL UBM局部 +(1-α)L UBM整体

[0022] Among them, α is a weight coefficient;

[0023] Finally, the reconstructed overall image is calculated to obtain the enhanced UBM image, which is formulated as follows:

[0024]

[0025] Among them, G n It is the high-frequency layer of the global image.

[0026] The local feature enhancement module can enhance the richness of ciliary sulcus detail features while preserving the overall structure of the UBM image.

[0027] First, the UBM local image is rotated and scaled to the corresponding position of the UBM global image using the SIFT image matching algorithm;

[0028] Then, the UBM image is decomposed into global and local images through the Laplacian pyramid;

[0029] Then, only the decomposed high-frequency layers are fused to enhance the details;

[0030] Finally, the reconstructed overall image is calculated to obtain the enhanced UBM image

[0031] Furthermore, the S3 at least includes the following steps:

[0032] First, the input Pentacam image, UBM global image and UBM local image are used. The Pentacam image is enhanced by rotation, flipping, color transformation, etc., and the UBM global and local images are fused and enhanced through the local feature enhancement module.

[0033] Subsequently, the two processed images are sent to the global and local backbone networks respectively for global and local feature extraction. The feature extraction is divided into four stages, and the network depth of each stage is progressively increased layer by layer.

[0034] Then, the different deep features extracted from each stage of the two backbone networks are subjected to a global-dominated multimodal cross-feature fusion.

[0035] Finally, the dual-backbone multimodal feature fusion network model is converged through the supervision of the ICL crystal size classification head.

[0036] The global backbone network uses Mamba to capture long-range dependencies and stabilize the overall morphology of intraocular structures. The local backbone network utilizes ResNet to refine edge details and enhance the model's ability to capture ciliary sulcus details. Multimodal cross-feature fusion, dominated by global features, ensures that the Pentacam global features effectively control all fused features, while local features are only used as auxiliary information for matching. This reduces computational effort, avoids the additional computational overhead of cross-attention, and improves model stability. Pentacam images are generally of higher quality, while UBM images may exhibit deformation and noise. Using the UBM as a key / value pair can reduce the impact on the overall feature network.

[0037] Furthermore, the multimodal feature cross fusion module includes at least the following steps in the process of combining features:

[0038] First, the self-attention mechanism is used to calculate the Pentacam global feature F P With UBM local feature F U The correlation between them is as follows:

[0039]

[0040] Where: Q = W Q F P , is the query matrix; K = W K F U , is the bond matrix; V=W V F U , is the value matrix;

[0041] Then, the calculated attention scores are used to weight local features and finally fused to generate F fused .

[0042] Furthermore, the S5 at least includes the following steps:

[0043] Input images of different modalities into the established selection network model, train the network model with the labeled ICL crystal size target, adjust the hyperparameters, and complete the training;

[0044] After training is completed, save the optimal weight model, load the weight model, and input the test set for accuracy testing;

[0045] Furthermore, the S6 at least includes the following steps:

[0046] Through deconvolution, the original pixel space is reconstructed for the deep fusion semantic features of the multimodal feature cross-fusion module, and the fused image information is viewed and the regions of interest (RoI) are highlighted to provide interpretable explanations of the model, helping doctors understand the prediction basis of the model and enhance clinical applicability.

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] This invention integrates Pentacam and UBM images for the first time, breaking through the limitations of existing ICL prediction technology, realizing multimodal image fusion, global-local information joint modeling and visual interpretation, with higher prediction accuracy, better interpretability and wider clinical applicability, providing a set of efficient and reliable intelligent decision-making assistance solutions for ICL chip selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0051] Figure 2 Schematic diagram of a local feature enhancement module of the present invention;

[0052] Figure 3 Schematic diagram of the dual-backbone multimodal feature fusion network module of the present invention. DETAILED DESCRIPTION

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0054] The present invention provides a deep learning-based method for selecting the size of the posterior chamber intraocular lens in phakic eyes. By combining Pentacam and UBM images, the method automatically extracts key structural information, optimizes surgical planning, and improves prediction accuracy.

[0055] The present invention innovatively combines Pentacam and UBM images of two different modalities to construct a dual-stream multimodal feature fusion network model (Dual-Stream Multimodal Feature Fusion Network) for simultaneously extracting the global ocular surface structure and local ciliary body structure. By deeply fusing the feature information of the two different modalities, a multimodal cross-fusion deep learning model is constructed, thereby improving the accuracy of ICL lens size classification.

[0056] The details are as follows:

[0057] Reference Figures 1 to 3 , a deep learning-based method for selecting the size of intraocular lenses in the posterior chamber of phakic eyes, wherein, Figure 1 The overall structure diagram of the intelligent decision-making method for ICL surgical lens size provided by a specific embodiment of the present invention is shown. The specific steps are as follows:

[0058] S1: Track, collect, and organize historical anterior segment imaging data. Use the patient's successful surgical records to label the image dataset with the corresponding ICL lens size category. Then, perform data segmentation and data preprocessing on the labeled image dataset to ensure high data quality and accuracy.

[0059] The historical imaging data used in this embodiment was collected by Purui Eye Hospital. The historical imaging data includes Pentacam horizontal section images, UBM global images, and UBM local image datasets, of which 70% are used as training sets, and the remaining 20% ​​and 10% are used as validation sets and test sets, respectively; the UBM global images are radial scanning section images taken from 3 to 9 o'clock.

[0060] Then, the imaging data are labeled with ICL size categories with reference to the patient's medical records, and the labeled imaging dataset is divided and preprocessed.

[0061] S2: Building a local feature augmentation module (Local Feature Augmentation Module), which enhances the local details of the ciliary body in the UBM image to improve the ability to capture details. The specific steps are as follows:

[0062] First, the UBM local image is rotated and scaled to the corresponding position of the UBM global image using the SIFT image matching algorithm.

[0063] Then, the UBM global and local images are decomposed through the Laplacian pyramid:

[0064] P UBM全局 ={L1,L2,L3,G3}

[0065] P UBM局部 ={L′1,L′2,L′3,G3′}

[0066] Among them: L1, L2, L3 are the low-frequency layers of the global image; G3 is the highest-frequency layer of the global image; L′1, L′2, L′3 are the low-frequency layers of the local image; G3′ is the highest-frequency layer of the local image;

[0067] Then, only the decomposed high-frequency layers are fused to enhance the details. The formula is as follows:

[0068] L fina =αL UBM局部 +(1-α)L UBM整体

[0069] Among them, α is a weight coefficient;

[0070] Finally, the reconstructed overall image is calculated to obtain the enhanced UBM image, which is formulated as follows:

[0071]

[0072] Among them, G n It is the high-frequency layer of the global image.

[0073] The local feature enhancement module can enhance the richness of ciliary sulcus detail features while preserving the overall structure of the UBM image.

[0074] S3: Design a dual-backbone multimodal feature fusion network model for model training. The specific steps are as follows:

[0075] First, the input Pentacam image, UBM global image and UBM local image are used. The Pentacam image is enhanced by rotation, flipping, color transformation, etc., and the UBM global and local images are fused and enhanced through the local feature enhancement module.

[0076] Subsequently, the two processed images are sent to the global and local backbone networks respectively for global and local feature extraction. The feature extraction is divided into four stages, and the network depth of each stage is progressively increased layer by layer.

[0077] Then, the different deep features extracted from each stage of the two backbone networks are subjected to a global-dominated multimodal cross-feature fusion.

[0078] Finally, the dual-backbone multimodal feature fusion network model is converged through the supervision of the ICL crystal size classification head.

[0079] The global backbone network uses Mamba to capture long-range dependencies and stabilize the overall morphology of intraocular structures. The local backbone network utilizes ResNet to refine edge details and enhance the model's ability to capture ciliary sulcus details. Multimodal cross-feature fusion, dominated by global features, ensures that the Pentacam global features effectively control all fused features, while local features are only used as auxiliary information for matching. This reduces computational effort, avoids the additional computational overhead of cross-attention, and improves model stability. Pentacam images are generally of higher quality, while UBM images may exhibit deformation and noise. Using the UBM as a key / value pair can reduce the impact on the overall feature network.

[0080] S4: Design a multimodal feature cross fusion module for fusing global and local features. The multimodal feature cross fusion module includes at least the following steps in the process of combining features:

[0081] First, the self-attention mechanism is used to calculate the Pentacam global feature F P With UBM local feature F U The correlation between them is as follows:

[0082]

[0083] Where: Q = W Q F P , is the query matrix; K = W K F U , is the bond matrix; V=W V F U , is the value matrix;

[0084] Then, the calculated attention scores are used to weight local features and finally fused to generate F fused .

[0085] S5: A selection network model is constructed based on the local feature enhancement module, the dual-backbone multimodal feature fusion network model, and the multimodal feature cross-fusion module. The processed multimodal image is input into the preset network model for training, and the optimal model weights are saved after the training is completed.

[0086] Input images of different modalities into the established selection network model, train the network model using the labeled ICL crystal size target, adjust the hyperparameters, and complete the training;

[0087] After training is completed, save the optimal weight model, load the weight model, and input the test set for accuracy testing;

[0088] S6: Introducing an interpretability module at the final stage of selecting the network model to map the fused deep features back to the original pixel space to help doctors understand the basis of the prediction results and the decision-making process

[0089] The original pixel space is reconstructed for the deep fusion semantic features of the multimodal feature cross fusion module through deconvolution;

[0090] View the fused image information and highlight the regions of interest (RoI) for interpretable explanations, helping doctors understand the model's prediction basis and enhancing clinical applicability.

[0091] In summary:

[0092] The method of the present invention combines the global corneal structure information provided by Pentacam images and the local ciliary body structure information provided by UBM images, and adopts a dual-trunk multimodal feature fusion network model to perform intelligent prediction of ICL crystal size. In data processing, a local feature enhancement module is used to enhance the local details of the ciliary body in the UBM image, and a multimodal feature cross-fusion module is designed to optimize the combination of global and local information. Through the multimodal fusion of training data, the network model can efficiently extract key structural information in the image and achieve accurate ICL crystal size prediction. In addition, the present invention also visualizes the prediction results through an interpretability module to help doctors understand the decision-making basis of the model, thereby improving the transparency and reliability of clinical applications. This method has high prediction accuracy, good interpretability, and can effectively promote the intelligence and precision of refractive correction surgery, providing a set of efficient and reliable intelligent auxiliary decision-making solutions for ICL lens selection.

[0093] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning, characterized by: At least the following steps are included: S1: Track, collect, and organize historical anterior segment imaging data. Use the patient's successful surgical records to label the image dataset with the corresponding ICL lens size category. Then, perform data segmentation and data preprocessing on the labeled image dataset to ensure high data quality and accuracy. S2: Building a local feature enhancement module, which enhances the local details of the ciliary body in the UBM image to improve the ability to capture details; S3: Construct a dual-backbone multimodal feature fusion network model. The dual-backbone multimodal feature fusion network model is used to perform fusion training on two modal images to extract features at different levels of the two images. S4: Design a multimodal feature cross-fusion module, which aims to effectively combine global features and local detail features to ensure their complementarity and synergy, thereby obtaining multimodal images. S5: A selection network model is constructed based on the local feature enhancement module, the dual-backbone multimodal feature fusion network model, and the multimodal feature cross-fusion module. The processed multimodal images are input into the preset network model for training, and the optimal model weights are saved after training is completed. S6: Introduce an interpretability module at the final stage of selecting the network model to map the fused deep features back to the original pixel space to help doctors understand the basis of the prediction results and the decision-making process.

2. The method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning according to claim 1, characterized in that: The image history data in S1 includes at least Pentacam horizontal section images, UBM global images and UBM local image data sets, of which 70% are used as training sets, and the remaining 20% ​​and 10% are used as validation sets and test sets respectively; the UBM global images are radial scanning section images taken from 3 to 9 o'clock.

3. The method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning according to claim 1, characterized in that: The application of the local feature enhancement module comprises at least the following steps: First, the UBM local image is rotated and scaled to the corresponding position of the UBM global image using the SIFT image matching algorithm; Subsequently, the UBM image is decomposed into global and local images using the Laplacian pyramid, as shown in the following formula: <h2 style=";text-align:left;direction:ltr">P<h2 style=";text-align:left;direction:ltr"> UBM全局 <h2 style=";text-align:left;direction:ltr"> (L1,L2,L3,G3) P UBM局部 ={L′1,L′2,L′3,G′3} Among them, L1, L2, L3 are the low-frequency layers of the global image; G3 is the highest-frequency layer of the global image; L′1, L′2, L′3 are the low-frequency layers of the local image; G′3 is the highest-frequency layer of the local image; Then, only the decomposed high-frequency layers are fused to enhance the details; L fina =αP UBM局部 +(1-α)P UBM整体 Among them, α is a weight coefficient; Finally, the reconstructed overall image is calculated to obtain the enhanced UBM image, as shown in the following formula: Among them, G n It is the high-frequency layer of the global image; The local feature enhancement module can enhance the richness of ciliary sulcus detail features while preserving the overall structure of the UBM image.

4. The method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning according to claim 2, characterized in that: The S3 at least includes the following steps: First, the dual-backbone multimodal feature fusion network model is fed with Pentacam images, UBM global images, and UBM local images. The Pentacam images are enhanced, which includes at least rotation, flipping, and color transformation. The UBM global images and UBM local images are fused and enhanced using a local feature enhancement module. Subsequently, the enhanced image in the previous step is fed into the global and local backbone networks for global and local feature extraction. Feature extraction is divided into four stages, with the network depth increasing layer by layer in each stage. Then, the different deep features extracted by the two backbone networks at each stage are fused into a global-dominated multimodal cross-feature. Finally, the dual-backbone multimodal feature fusion network model is converged through the supervision of the ICL crystal size classification head.

5. The method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning according to claim 1, characterized in that: The multimodal feature cross fusion module includes at least the following steps in the process of combining features: First, the self-attention mechanism is used to calculate the Pentacam global feature F P With UBM local feature F U The correlation between them is as follows: Where: Q = W Q F P , is the query matrix; K = W K F U , is the bond matrix; V=W V F U , is the value matrix; Then, the calculated attention scores are used to weight local features and finally fused to generate F fused .

6. The method for selecting the size of a phakic posterior chamber intraocular lens based on deep learning according to claim 5, characterized in that: The S5 at least includes the following steps: Input images of different modalities into the established selection network model, train the network model with the labeled ICL crystal size target, adjust the hyperparameters, and complete the training; After training is completed, save the optimal weight model, load the weight model, and input the test set for accuracy testing.

7. The method for selecting phakic posterior chamber intraocular lens size based on deep learning according to claim 6, characterized in that: The S6 at least includes the following steps: The original pixel space is reconstructed for the deep fusion semantic features of the multimodal feature cross fusion module through deconvolution; Viewing fused image information and highlighting regions of interest for interpretable explanations helps doctors understand the model's predictions and enhances clinical applicability.