Method for generating anterior segment tomography of corneal topography after ICL surgery based on latent space diffusion
Patent Information
- Application Number
- CN202511622253.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-11-07
AI Technical Summary
然而传统方法在全局优化过程中,中心区域的细节特征易被整体图像质量目标所稀释,导致该区域预测结果无法满足临床精细评估需求
1、该方法有效填补了现有技术在利用术前角膜地形图或前节断层图进行术后预测方面的空白,不再局限于输出单一参数,而是能够生成完整的可视化术后预测图像。这为医生提供了直观的视觉判据,极大地方便了手术效果的预评估和方案制定。
Smart Images

Figure CN121439119B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of corneal topography technology, specifically to a method for generating anterior segment tomographic maps of corneal topography after ICL surgery based on latent space diffusion. Background Technology
[0002] In the field of refractive surgery, preoperative assessment and postoperative outcome prediction for ICL (implantable phakic intraocular lens) are crucial for surgical safety and visual correction. Traditional postoperative prediction methods mainly rely on statistical parametric models or finite element biomechanical simulations, which have several inherent limitations. These methods typically only output discrete postoperative parameter indicators, such as corneal thickness or local curvature values at specific points, and cannot provide complete postoperative corneal topography or anterior segment tomography visualizations. Without intuitive morphological references, surgeons can only infer the overall postoperative corneal morphological changes through abstract data and experience when evaluating surgical plans, which can easily lead to judgment bias.
[0003] Current technologies do not fully utilize routine preoperative imaging. While preoperative images such as corneal topography and anterior segment tomography are widely acquired, the rich spatial information and biomechanical features they contain have not been systematically used in predictive models. Most methods extract only a small number of feature parameters as input, failing to fully explore the corneal structural patterns contained in high-resolution images. This insufficient utilization of information limits the accuracy and generalization ability of predictive models, especially when dealing with patients with complex corneal morphology or special refractive errors, where predictive reliability decreases significantly.
[0004] At the technical implementation level, while deep learning-based generative models have some applications in the field of medical imaging, their direct application to high-resolution corneal image generation faces computational efficiency bottlenecks. Pixel-level diffusion models, when processing corneal topography maps with millions of pixels, can take tens of seconds or even minutes for a single prediction, far exceeding the time requirements for real-time assessment in clinical outpatient settings. This high latency characteristic makes it difficult for existing generative technologies to be integrated into actual diagnostic and treatment processes, limiting them to offline analysis scenarios.
[0005] Furthermore, existing prediction methods generally lack targeted optimization for key anatomical structures. The central corneal optical zone, as a core area affecting postoperative visual quality, has a decisive impact on surgical outcome assessment due to the accuracy of its morphological prediction. However, in the global optimization process of traditional methods, the detailed features of the central region are easily diluted by the overall image quality target, resulting in prediction results for this region that fail to meet the needs of precise clinical assessment.
[0006] To address this, we propose a method for generating anterior segment tomographic maps of corneal topography after ICL surgery based on latent space diffusion. Summary of the Invention
[0007] One of the technical problems this application aims to solve is that the existing technology lacks research on anterior segment tomography of corneal topography, but anterior segment tomography of corneal topography is the most commonly used preoperative examination for refractive surgery.
[0008] Most studies focus on predicting individual postoperative data, such as the arch height and anterior chamber angle. Direct, intuitive criteria are not readily available, hindering medical researchers' comprehensive assessment of risk. Furthermore, there is a lack of information on postoperative visual observations.
[0009] Pixel spatial diffusion models directly process images, resulting in long training cycles and high generation time per image during inference, which cannot meet the real-time requirements of clinical preoperative planning.
[0010] To address the aforementioned technical problems, this application provides a method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion, comprising the following steps: S1: Establish a dedicated latent space encoding-decoding framework to map preoperative corneal topography / anterior segment tomography to a low-dimensional latent space; S2: In the latent space, combined with the preoperative image conditions, the structured change pattern of "preoperative → postoperative" is learned through diffusion denoising training to obtain the diffusion model; S3: Based on the trained diffusion model, input the preoperative image of a new patient and generate a visualized postoperative corneal topography / anterior segment tomography prediction result; S4: Outputs a complete predicted image, providing intuitive visual criteria for postoperative corneal morphology.
[0011] In some embodiments, the latent space coding-decoding framework includes a consistency reconstruction mechanism to ensure that the original image can be reconstructed with high precision by the decoder after the preoperative image is encoded and compressed into the latent space, thus ensuring the reliability of feature extraction.
[0012] In some embodiments, the diffusion model training employs a stochastic differential equation (SDE)-based modeling method to simulate the image noise addition and denoising process in the latent space, efficiently learning the postoperative morphological change patterns. The forward SDE noise addition process is as follows: ,in This represents a preoperative image. The corresponding postoperative image (mean state). and These are the time-varying regression coefficient and the diffusion intensity, respectively. This is the Wiener process.
[0013] In some embodiments, a central region enhancement module (CEM) is introduced to prioritize optimizing the detailed features of the central corneal region during the generation process, thereby improving the clarity and accuracy of key regions in the predicted image.
[0014] In some embodiments, the generation time of a single postoperative predicted image is ≤8.5 seconds. The computation time is reduced to less than 1 / 10 of the original pixel-level diffusion model by latent space dimensionality reduction, which meets the real-time clinical needs.
[0015] In some embodiments, preoperative image conditions include, but are not limited to, corneal curvature, thickness, and anterior chamber depth data, providing more comprehensive image information than single-value information.
[0016] In some embodiments, the generated postoperative prediction images can be directly used for clinical efficacy evaluation, with a peak signal-to-noise ratio (PSNR) ≥28dB and structural similarity (SSIM) ≥0.92, which is superior to traditional parameter prediction methods.
[0017] In some embodiments, diffusion denoising training employs a conditional noise scheduling strategy, dynamically adjusting the noise addition intensity based on preoperative image features to improve the model's adaptability to complex cases.
[0018] In some embodiments, the decoder output layer employs a residual connection structure to preserve latent space feature details and reduce information loss during image reconstruction.
[0019] In some embodiments, the postoperative predictive image includes a three-dimensional visualization of corneal curvature distribution, thickness changes, and anterior chamber structure, covering all dimensions of refractive surgery assessment needs.
[0020] This invention has at least the following beneficial effects: 1. This method effectively fills the gap in existing technologies for postoperative prediction using preoperative corneal topography or anterior segment tomography. It is no longer limited to outputting a single parameter but can generate complete, visualized postoperative prediction images. This provides doctors with intuitive visual criteria, greatly facilitating the pre-assessment of surgical outcomes and the formulation of treatment plans.
[0021] 2. By constructing a dedicated latent space encoder-decoder framework, high-dimensional preoperative images are mapped to a low-dimensional latent space for processing, significantly reducing computational complexity. This drastically reduces the generation time of a single postoperative prediction image from approximately 89 seconds required by the traditional pixel-level diffusion model to about 8.1 seconds. This improved generation efficiency meets the urgent real-time requirements of clinical outpatient clinics, making this technology feasible for practical application.
[0022] 3. In terms of technical implementation, this method incorporates latent space feature extraction and consistency reconstruction mechanisms to ensure the reliability and integrity of preoperative image information during compression and reconstruction. The introduced central region enhancement module pays special attention to and optimizes the generation of detailed features in the central corneal region, which has the most critical impact on vision, thereby improving the clarity and accuracy of the predicted image in important areas.
[0023] 4. This method demonstrates superior image quality. The generated postoperative predicted images outperform traditional parametric prediction methods in objective evaluation metrics such as peak signal-to-noise ratio and structural similarity, meaning the prediction results are clearer and closer to the actual postoperative condition. This high-quality image output provides a more reliable basis for clinical decision-making.
[0024] 5. This method demonstrates strong generalization ability and clinical applicability. It can effectively learn and simulate complex structured changes from preoperative to postoperative stages, and the generated predictions cover key dimensions such as corneal curvature distribution, thickness changes, and anterior chamber structure, meeting the comprehensive needs of refractive surgery assessment. The selection of techniques such as diffusion process modeling based on stochastic differential equations also enhances the model's adaptability and robustness to different cases. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the method steps of the present invention; Figure 2 This is a schematic diagram of the decoder stage steps of the present invention; Figure 3 This is a schematic diagram of the noise reduction network stage steps of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] Example 1, see Figure 1 This invention provides a technical solution: a method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion, comprising the following steps: S1: Establish a dedicated latent space encoding-decoding framework to map preoperative corneal topography / anterior segment tomography to a low-dimensional latent space; S2: In the latent space, combined with preoperative image conditions, the structured change pattern of "preoperative → postoperative" is learned through diffusion denoising training; S3: Based on the trained diffusion model, input the preoperative image of a new patient and generate a visualized postoperative corneal topography / anterior segment tomography prediction result; S4: Outputs a complete predicted image, providing intuitive visual criteria for postoperative corneal morphology.
[0028] Specifically, the design principle of this technical solution is based on three key points. First, it achieves data dimensionality reduction through a dedicated latent space encoding-decoding framework, mapping high-resolution preoperative corneal topography or anterior segment tomography images into a low-dimensional latent space. This transformation preserves the core features of the image while significantly reducing data processing complexity. Second, it utilizes the learning characteristics of a diffusion model in the low-dimensional latent space, using the preoperative image as a conditional constraint, and simulates the structured evolution from the preoperative state to the postoperative state through denoising training. Finally, the decoder reconstructs the prediction results from the latent space into a visualized image, completing the transformation from abstract features to clinically usable images.
[0029] The main objectives of this technology are twofold. Firstly, it addresses the insufficient utilization of preoperative images in existing technologies by fully leveraging the inherent information in preoperative image data such as corneal topography. Secondly, it overcomes the limitation of traditional prediction methods, which can only output single postoperative parameters (such as thickness or curvature at a specific point), by achieving panoramic visualization and prediction of the entire postoperative corneal morphology. Simultaneously, it must overcome the excessively high computational cost of pixel-level diffusion models to meet the time requirements for real-time outpatient assessment.
[0030] The advantages of this design are twofold: Firstly, in terms of prediction quality, the generated complete predictive images provide doctors with intuitive visual criteria, covering key dimensions such as corneal curvature distribution, thickness variations, and anterior chamber structure, far superior to abstract representations based on single parameters. The prediction results also exhibit higher reliability in objective evaluation metrics. Secondly, in terms of computational efficiency, latent space dimensionality reduction reduces the single-image generation time from approximately 89 seconds using traditional methods to the 8-second level, enabling real-time outpatient applications. The system specifically includes a detailed optimization mechanism for the central corneal region, which has the most significant impact on visual correction; this regional enhancement further improves the accuracy of clinical decision-making. Through structured learning of the pre-operative-post-operative mapping, the system meets the stringent clinical requirements for operational timeliness while maintaining predictive accuracy. Finally, the output visualized predictive images directly serve surgical planning and outcome evaluation, forming a closed-loop support system from data to decision.
[0031] Example 2, see Figure 2 The latent space coding-decoding framework includes a consistent reconstruction mechanism to ensure that the original image can be reconstructed with high accuracy by the decoder after the preoperative image is encoded and compressed into the latent space, thus ensuring the reliability of feature extraction.
[0032] Specifically, in the latent space model for generating postoperative images from preoperative images, encoder-decoder consistency is the fundamental guarantee that the model can complete cross-domain generation. The encoder is the module responsible for ensuring encoder-decoder consistency.
[0033] Encoder-decoder consistency refers to the ability of a decoder to accurately reconstruct the original input image from the latent space information after the input image (such as a preoperative image) is compressed into the latent space by the encoder to obtain latent vectors and intermediate features.
[0034] During the encoding stage, information compression and preservation are performed. The encoder's role is to convert the high-dimensional image into low-dimensional latent vectors and a series of intermediate features (feature maps output by each layer of the encoder). The specific process is as follows: The input image is first reduced to the basic number of channels through an initial convolution. The input image is then resized using reflection padding to ensure alignment during downsampling. The preprocessed image enters the initial convolution, converting the number of channels from the input channels to the basic number of channels, resulting in an initial feature map. Multiple encoders progressively compress the spatial dimension and increase the channel dimension. Each encoder consists of four stages, each containing two sets of residual blocks (ResBlock), an attention module (only in encoder 4), and a downsampling layer, progressively compressing the spatial dimension and increasing the channel dimension. Residual block processing: The first two sets of residual blocks in each stage perform deep feature extraction on the current feature map, and retain detailed information through convolution, non-linear activation and skip connections.
[0035] With attention enhancement, only the deepest encoder will have a LinearAttention module (normalized by PreNorm and with residual connections) added to focus on key regions. Shallow layers, due to their larger feature space size, will not use an attention module.
[0036] Downsampling compression, except for the last stage, halves the feature map spatial size (e.g., H×W→H / 2×W / 2) through the downsampling module (convolution with stride of 2) in each stage, while increasing the number of channels to the next level; the last stage directly adjusts the number of channels through convolution without changing the spatial size.
[0037] Feature caching involves storing the output feature map of each residual block in a list, which serves as skip connection information for subsequent decoding stages.
[0038] Finally, the deepest features are compressed into latent vectors (in the latent space) through latent projection, while all intermediate features (stored in each encoder) are preserved. The high-dimensional features output from the deepest layer of the encoder are compressed into low-dimensional latent vectors through latent space convolution (i.e., latent projection), completing the mapping from image space to latent space.
[0039] The key to this step is to preserve the core information of the image (such as organ structure and location in preoperative images) while compressing dimensions; otherwise, crucial details will be lost during subsequent decoding.
[0040] During the decoding stage, information is restored and reconstructed. The decoder's role is to reconstruct the original image from latent vectors and intermediate features. The specific process is as follows: The latent vectors are first restored to the deep feature dimension through latent backprojection; the latent vectors are then restored to the number of channels of the deepest feature of the encoder through inverse latent space convolution (latent backprojection), providing initial high-level features for decoding.
[0041] Multiple decoders are used, and their intermediate features from the corresponding layers of the encoder are concatenated to fuse high- and low-level information. The decoder also contains four stages, symmetrical to the encoder. Each stage fuses the intermediate features of the encoder through skip connections, and gradually increases the spatial dimension and decreases the channel dimension. Feature concatenation: The input features of each decoding stage are first concatenated with the feature maps cached in the corresponding stage of the encoder in the channel dimension, fusing high-level semantic (decoder) and low-level detail (encoder) information.
[0042] The residual blocks and attention, the concatenated features are further processed by two sets of residual blocks, and the deepest decoder (decoder 4) also contains a LinearAttention module to enhance the restoration of key features.
[0043] Upsampling restoration: Except for the first stage, each stage doubles the feature map space size (e.g., H×W→2H×2W) through the upsampling module (interpolation + convolution), and reduces the number of channels to the previous level; the first stage (decoder 1) directly adjusts the number of channels to the basic number of channels through convolution.
[0044] Finally, the features are mapped back to the image channels via a final convolution and cropped to the original size. The feature map output by the decoder is added to the initial convolution result of the encoder (residual connection), and then the final convolution maps the number of channels back to the output channels (RGB3 channels). Finally, it is cropped to the original size of the input image to obtain the reconstructed image. This step achieves accurate image reconstruction by fusing intermediate features from the encoder (preserving details) and latent vectors (preserving global information).
[0045] The purpose of this design is to ensure the reliability of the entire prediction system, with the core objective of the "latent space encoding-decoding framework including a consistency reconstruction mechanism." Specifically, its design logic addresses the feature representation distortion problem commonly encountered by deep learning models when processing high-dimensional medical images. When high-resolution images such as preoperative corneal topography or anterior segment tomography are compressed and mapped to a low-dimensional latent space, crucial structural information may be lost or distorted due to the dimensionality reduction operation. The consistency reconstruction mechanism, by forcing the encoded latent space features to be able to accurately reconstruct the original preoperative image through the decoder, essentially sets up a self-verification step for the feature extraction process. This closed-loop design compels the encoder to selectively retain anatomical features crucial for image reconstruction during compression, rather than irrelevant noise or redundant information.
[0046] The direct benefits of this mechanism are multifaceted. First, it fundamentally ensures the quality of data input into the subsequent diffusion model. If latent space features cannot accurately reconstruct the original image, it indicates information loss or bias in the feature extraction process, making postoperative predictions based on this unreliable. Consistent reconstruction provides an objective verification standard for the fidelity of feature extraction. Second, this mechanism enhances the model's generalization ability. By reconstructing constraints, the encoder learns generalized feature representations with clear anatomical significance, rather than overfitted features specific to the training dataset. This allows the model to maintain stable performance when faced with preoperative images acquired by different devices or from different patient groups. Finally, this design enhances trust in clinical applications. Doctors can intuitively understand the mechanism's function—the system can "understand" the original image and "remember" key details, rather than performing unexplainable black-box operations, providing technical assurance for the reliability of the prediction results.
[0047] This mechanism also indirectly optimizes the training efficiency of subsequent diffusion models. Since the latent space features input to the diffusion model have already been reconstructed and validated, they contain higher information density and consistency. The model does not need to painstakingly identify valid signals amidst noise, allowing it to focus more on learning the "pre-operative → post-operative" mapping pattern. This phased reliability assurance provides a solid underlying foundation for the final generated visual prediction image in terms of detail reproduction and structural rationality, especially enabling more accurate prediction of the corneal central region morphology, which is crucial for refractive surgery outcomes.
[0048] Example 3, see Figure 3 The diffusion model training adopts a modeling method based on stochastic differential equations (SDE) to simulate the process of adding and removing image noise in the latent space, and efficiently learn the postoperative morphological change pattern.
[0049] Specifically, the image is encoded into the latent space by the encoder trained in the previous step, and then noise is added to the latent variables in the latent space. After the diffusion denoising process is completed in the latent space, it is decoded back into the pixel space to restore the post-generated image, thereby reducing computational complexity and improving generation stability.
[0050] For denoising, an image restoration framework based on mean-reverting stochastic differential equations (SDEs) is employed. This aims to convert preoperative images into postoperative images through a denoising-denoising mechanism driven by stochastic differential equations and the noise prediction capabilities of the denoising network. The workflow of the image generation model is as follows: Figure 3 As shown.
[0051] The positive SDE noise addition process is as follows: ; in This represents a preoperative image. The corresponding postoperative image (mean state). and These are the time-varying regression coefficient and the diffusion intensity, respectively. This is the Wiener process.
[0052] This SDE asymptotically degenerates the latent space feature map of the preoperative image into a Gaussian distribution using a closed-loop solution. Noisy images Its transition probability follows .
[0053] The reverse generation stage is achieved by solving the inverse SDE. A noise prediction U-Net network is introduced. Estimate score function The following reconstruction equation is constructed: The denoising network plays a central role here, taking conditional information (preoperative image) and time step as input. The SDE generates noise based on the current time step and then learns the denoising function in the latent space. The network employs a typical encoder-decoder structure, with the degraded image at time t as input. (Here is the noisy preoperative image at time t in the figure) and the target image. (Here, the spliced tensor of the postoperative noisy image at time t in the figure).
[0054] The encoder extracts multi-scale features through multi-level convolution and downsampling, while the decoder fuses shallow details and deep semantic information through skip connections. Specifically, a center region enhancement module (CEM) is added to decoder 3 to dynamically locate and strengthen central region features, improving the generation quality of surgically relevant areas. The overall process balances global structure and local details. The final output is optimized for residual noise. The accurate estimate is obtained. This predicted value is used for state updates via residual connections: The training process employs a maximum likelihood trajectory optimization strategy. To avoid the instability of traditional score-matching targets, this strategy maximizes the posterior probability. The optimal reverse state transition is derived as follows: The training objective function is defined as the predicted state. Compared to the ideal state The sum of L1 distances significantly improves the stability and fidelity of the solution. By iteratively performing implicit sampling and noise removal, this method gradually eliminates degradation noise in the original image and injected Gaussian noise, ultimately outputting a latent space feature map of the generated postoperative image that is highly similar to the real postoperative state. Then, the latent space decoder trained in the previous step is used to reconstruct the generated postoperative image from this latent space feature map.
[0055] A diffusion model was trained using a stochastic differential equation (SDE)-based modeling approach, designed primarily to address core challenges in medical image generation. Simulating image noise addition and denoising in the latent space essentially aims to establish a mathematical expression that better reflects the morphological changes of biological tissues. Postoperative corneal morphological changes are not simple pixel displacements, but rather continuous, gradual processes influenced by biomechanical properties. The continuous-time modeling characteristics of SDE naturally describe this gradual deformation, avoiding the abrupt distortion problems caused by discretized modeling. Its differential equation form essentially treats image state changes as continuous trajectories in the time dimension, which highly aligns with the gradual evolution characteristics of real corneal tissue after surgery.
[0056] The substantial benefits of this method are primarily reflected in its learning efficiency. By uniformly describing noise addition (forward process) and denoising prediction (backward process) through the SDE framework, the model can establish a more stable optimization objective during latent space training. Systematically learning the probabilistic path from the noise distribution to the target postoperative image avoids the pattern collapse risk common in traditional generative models. This continuous learning approach is particularly suitable for capturing subtle patterns in corneal structural changes, such as the curvature transition gradient from the peripheral to the central area or the continuous characteristics of thickness changes.
[0057] Secondly, SDE modeling significantly enhances the robustness of the computational process. Drift and diffusion terms in the equations can be finely controlled through parameter adjustment to achieve precise noise modulation, enabling the model to adaptively handle preoperative corneal morphology of varying complexity. For important cases such as highly irregular corneal topography or complex anterior chamber structures, this mechanism can dynamically adjust the learning difficulty to ensure the anatomical rationality of the predicted results. This characteristic directly improves the system's applicability in diverse clinical scenarios.
[0058] In terms of prediction quality, the SDE framework ensures the determinism of the generation process. Although the training process is based on probabilistic modeling, solving the inverse SDE yields a deterministic output, which is crucial for medical images requiring precise structural information. The final generated postoperative predicted images are more consistent with the biological characteristics of the real cornea in terms of visual metrics such as surface smoothness and boundary continuity, especially avoiding local structural breaks or unreasonable mutations that may occur with traditional methods.
[0059] Performing SDE operations in low-dimensional space significantly reduces the computational overhead of traditional pixel-level diffusion while retaining the ability to capture complex morphological changes. This highly efficient continuous spatial modeling lays the mathematical foundation for providing clinical visualization and prediction tools that conform to both biomechanical principles and real-time requirements.
[0060] Example 4 introduces a Central Region Enhancement Module (CEM) to prioritize optimizing the detailed features of the central corneal region during the generation process, improving the clarity and accuracy of key areas in the predicted image. The generated postoperative predicted image can be directly used for clinical efficacy evaluation, with a peak signal-to-noise ratio (PSNR) ≥28dB and structural similarity (SSIM) ≥0.92, outperforming traditional parametric prediction methods. The generation time for a single postoperative predicted image is ≤8.5 seconds, and the computation time is reduced to less than 1 / 10 of the original pixel-level diffusion model through latent space dimensionality reduction, meeting real-time clinical needs.
[0061] Specifically, the design principle of the central region enhancement module (CEM) in this technical solution stems from the core position of the central corneal region in vision correction. The optical characteristics of this region directly affect postoperative visual quality, and the accuracy of its morphological prediction is crucial. When generating predictive images, the system assigns higher weight to the feature representation of the central region through this module, implementing targeted optimizations in feature extraction, latent space mapping, and image reconstruction. This targeted enhancement is not simply about enlarging local pixels, but rather, based on the biomechanical characteristics of the cornea, it strengthens the ability to capture the morphological changes of the central region during model learning, ensuring the prediction accuracy of key details such as curvature gradients and thickness changes in this region.
[0062] The primary purpose of this design is to address the issue of traditional generative models potentially neglecting critical regions during global optimization. Because the central corneal region occupies a small area but has extremely high clinical value, it is easily marginalized during the overall image generation process. The CEM module establishes a region priority mechanism, tilting model resources towards key anatomical structures and avoiding sacrificing prediction fidelity in the core region in pursuit of overall image quality. Simultaneously, this design also addresses the practical needs of clinicians evaluating surgical outcomes, ensuring that the details of the central optical zone, which are of utmost concern to them, are fully represented.
[0063] The core benefit of this module lies in the improved clinical usability of the prediction results. Doctors can clearly identify subtle features of curvature changes in the central region, such as the smoothness of the pupil area's morphological transition and the offset of the apex position—indicators directly affecting visual prognosis. This targeted optimization ensures that the predicted images not only meet visual integrity requirements but also possess clinical decision support value. Simultaneously, because the optimization focuses on a specific region, it does not significantly increase overall computational complexity, maintaining its generation efficiency advantage.
[0064] The optimization design regarding generation time is based on the principle of reconstructing the computational path by utilizing the dimensionality reduction properties of the latent space. Traditional pixel-level diffusion models directly perform noise addition and denoising operations at the million-pixel level, resulting in a geometrically increasing computational load. In contrast, this scheme compresses the high-dimensional image into the latent space, potentially reducing the data dimensionality by hundreds of times, significantly reducing the number of parameters and iterative computations required by the diffusion model. This dimensionality transformation essentially changes the mathematical scale of the problem, reconstructing the computational framework while maintaining predictive functionality.
[0065] Outpatient surgical consultations typically require assessments to be completed within minutes, and traditional methods, with a single image generation time of nearly 90 seconds, cannot meet the process requirements. By compressing the time to the 8-second level through latent space dimensionality reduction, doctors can obtain predictive results in real time during patient consultations, completely changing the limitation of this technology being confined to offline research scenarios. This breakthrough in timeliness directly determines whether the technology can move from the laboratory to the operating room.
[0066] The core value of time optimization lies in achieving a balance between accuracy and efficiency. The 8-second generation speed is not achieved at the expense of prediction quality, but rather through a fundamental speedup achieved by altering the dimensional characteristics of the computational space. This efficiency allows physicians to incorporate visual predictions into their routine assessment processes without disrupting their patient flow. Simultaneously, rapid iteration capabilities provide the technical possibility for real-time intraoperative adjustments, laying the foundation for future development of dynamic surgical planning systems. Ultimately, this design transforms high-quality postoperative prediction from a theoretical concept into a practical tool that can be seamlessly integrated into clinical workflows.
[0067] The quality standards of PSNR ≥ 28dB and SSIM ≥ 0.92 are designed based on the diagnostic-grade clarity requirements of medical images. Corneal topography needs to display curvature changes at the 0.1mm level, while anterior segment tomography requires clear resolution of corneal interlaminar structures. These two thresholds have been clinically validated to ensure that the generated predicted images meet the level of visual clarity and structural fidelity required for direct evaluation by physicians. PSNR ensures pixel-level grayscale accuracy, while SSIM constrains the degree of preservation of tissue texture and boundary continuity.
[0068] The design of this quality indicator is clearly aimed at clinical application. Traditional parametric prediction methods only provide abstract numerical values, requiring doctors to rely on experience to mentally visualize morphological changes, which can easily lead to misjudgment. However, the predicted images that meet the above indicators can directly replace actual postoperative examination images for preoperative assessment of key indicators such as corneal apex deviation and optical zone regularity. This visualization capability fundamentally changes the surgical planning model, enabling doctors to visually identify potential problems (such as central island or off-center ablation risks) preoperatively.
[0069] Example 5: Preoperative image conditions include, but are not limited to, corneal curvature, thickness and anterior chamber depth data, and the image information is more comprehensive than single numerical information.
[0070] Specifically, this technical solution uses multimodal preoperative images, including corneal curvature maps, thickness maps, and anterior chamber depth maps, as input conditions. Its design principle lies in the spatial dependence of corneal morphological changes. Real postoperative corneal deformation is not a numerical change at isolated points, but rather a continuous surface reconstruction process governed by biomechanical laws. Single-point numerical parameters cannot express the gradient characteristics of curvature changes, the topological correlation of thickness distribution, and the spatial configuration of the anterior chamber structure—all key factors affecting the optical performance after ICL lens implantation. By preserving complete spatial distribution information, the image data allows the model to learn, for example, the curvature evolution pattern of the superior cornea being steeper than that of the temporal side, or the linkage changes in thickness between the nasal peripheral and central regions.
[0071] The primary objective of this design is to overcome the information loss inherent in traditional parametric modeling. Numerical tables or scatter plots lose the smooth continuity of corneal surface curvature transitions and fail to represent the propagation trend of thickness changes along the XY-axis. Using original topographic and tomographic maps as input is equivalent to directly injecting a complete biomechanical information field into the prediction system. From the preprocessing stage, the model addresses the hotspot distribution of corneal curvature, the contour features of the thickness map, and the three-dimensional spatial relationships of the anterior chamber structure, establishing a realistic physical basis for subsequent simulations of tissue response after lens implantation.
[0072] The core benefit of this choice lies first in ensuring the anatomical rationality of the prediction results. Once the model has learned the curvature gradient field expressed by a complete corneal topography map, its postoperative predictions naturally maintain a smooth transition between the corneal apex and peripheral curvature changes, avoiding non-physiological curvature abrupt changes that may occur with traditional methods. Similarly, predictions based on thickness distribution maps accurately reflect the diffusion pattern of postoperative stromal edema in different regions of the cornea, rather than simply outputting the mean changes of a few discrete points.
[0073] Secondly, this design significantly enhances the model's clinical interpretability. When evaluating the predicted results, physicians can directly analyze the curvature distribution patterns or thickness variation areas in the predicted image using their familiar image interpretation experience from preoperative examinations. Without the need for additional learning of the conversion rules between abstract parameters and clinical representations, this seamless diagnostic pathway significantly lowers the cognitive threshold for technology application. For example, the crucial central optical zone diameter delineation in corneal surgery can be directly completed on the predicted topographic map using the measurement procedures from preoperative assessment.
[0074] At the algorithmic level, using raw image input avoids the subjective risks of feature engineering. Traditional methods rely on manual selection of input parameters, which may miss key features affecting postoperative responses. The complete information preserved in the image data allows the model to autonomously discover, for example, the spatial correlation between anterior chamber depth and corneal endothelial cell density distribution, or the potential association between specific curvature patterns and postoperative higher-order aberrations. This data-driven feature mining capability is particularly important for improving the predictive robustness of complex cases.
[0075] The design also leaves room for future technological upgrades. With the clinical application of new examination methods such as high-resolution corneal biomechanical imaging, the system can seamlessly integrate new modal images into the existing framework without reconstructing the feature extraction process. This future-oriented compatibility allows the technology to continuously absorb advancements in medical imaging equipment, forming a self-improving closed-loop system.
[0076] Example 6: The diffusion denoising training employs a conditional noise scheduling strategy, dynamically adjusting the noise addition intensity based on preoperative image features to improve the model's adaptability to complex cases. The decoder output layer uses a residual connection structure to preserve latent space feature details and reduce information loss during image reconstruction. Postoperative predicted images include three-dimensional visualizations of corneal curvature distribution, thickness changes, and anterior chamber structure, covering all dimensions of refractive surgery assessment needs.
[0077] Specifically, the design principle of the conditional noise scheduling strategy stems from the highly individualized differences in corneal morphology. Different patients exhibit significant variations in preoperative corneal topography complexity; for example, high astigmatism or irregular surfaces require special handling. Traditional fixed-noise strategies struggle to balance the learning needs of various cases. This strategy dynamically determines the intensity and extent of noise addition during diffusion by analyzing features such as local curvature gradients and thickness distribution in preoperative images. Stronger perturbations are applied to structurally complex regions, forcing the model to learn more robust patterns of change in these areas; while the perturbation intensity is reduced in uniformly flat regions to prevent excessive noise from damaging the original structural features. This dynamic adjustment exposes the model to a training environment that more closely resembles real clinical challenges.
[0078] The primary purpose of this design is to address the issue of uneven model adaptability to cases of varying difficulty. When faced with rare and complex corneal morphologies, standard diffusion models may exhibit structural distortion or loss of detail. The dynamic noise strategy essentially customizes the training difficulty for each case, ensuring that highly heterogeneous regions are adequately learned. This targeted approach allows the model to maintain predictive stability even when encountering unconventional cases, avoiding systematic prediction errors caused by biases in the distribution of training data.
[0079] The benefits of this strategy are reflected in the improved inclusivity of clinical practice. For patients with special refractive errors such as keratoconus, the model will not suffer catastrophic prediction failure and can maintain acceptable prediction accuracy. Meanwhile, dynamic adjustment avoids excessive increase of the overall model capacity to adapt to complex cases, and maintains prediction efficiency for conventional cases. This adaptability provides a universal guarantee for clinical promotion, enabling the technology to cover a broader patient population.
[0080] The principle of adopting a residual connection structure at the output layer of the decoder addresses the problem of high-frequency detail loss during image reconstruction. When latent space features are decoded and restored into pixel images, transmission through multiple layers of the network may cause the gradual attenuation of corneal micro-structure information layer by layer. The residual structure establishes a direct feature channel from the bottom layer to the top layer, enabling key information such as fine textures and edge gradients extracted in the early stage to directly participate in the final image reconstruction. This pathway bypasses the intermediate processing links that are prone to information loss, which is equivalent to setting up a dedicated transmission path for important micron-level features.
[0081] The core purpose of this design is to maintain the biological structure consistency between the predicted image and the original image. Micro features such as interlayer tissue demarcation and subepithelial haze in corneal tomography directly affect the judgment of postoperative effect, and these details are easily blurred during conventional decoding. Residual connection ensures that such biomarker information is completely transmitted to the output layer, avoiding information attenuation caused by model structure design.
[0082] The corneal biological texture features retained by the residual structure allow doctors to clearly identify key signs such as changes in edema degree of the postoperative corneal stroma or micro-folds in Bowman's layer. These micro-structural information are crucial for evaluating tissue reaction after ICL implantation, and traditional reconstruction methods are difficult to maintain this level of fineness. This technology ensures that the predicted image meets the requirements of clinical diagnosis for tissue micro-structure identification.
[0083] The design principle that requires the postoperative predicted image to include three-dimensional visualization results of corneal curvature, thickness and anterior chamber structure is based on the multi-dimensional evaluation system for the efficacy of refractive surgery. Postoperative visual quality depends not only on the correction accuracy of corneal curvature, but also is closely related to anterior chamber space changes caused by lens position and the biomechanical response of corneal thickness. These three types of data have physiological interactions: for example, changes in the vault of an ICL will change the anterior chamber depth, which in turn triggers morphological adjustment of corneal endothelial cells. Comprehensive visualization is essentially a multi-system linkage simulation of the postoperative state.
[0084] The main purpose of this design is to bridge the gap between single-parameter predictions and actual clinical needs. When evaluating surgical plans, doctors must simultaneously observe corneal curvature maps to assess refractive correction effects, refer to thickness maps to evaluate surgical safety, and combine anterior chamber structure data to confirm the rationality of lens placement. Generating prediction results covering all dimensions allows doctors to make systematic and comprehensive judgments from a unified perspective, avoiding spatial registration errors caused by separate predictions from multiple systems.
[0085] Doctors can directly observe quantitative relationships such as the "distance between the peripheral corneal thickening area and the lens" or the "indirect impact of changes in anterior chamber depth on the corneal apex position" in 3D images. This multi-parameter spatial correlation analysis solves the pain point of traditional methods that require manual comparison of multiple independent prediction data, significantly improving the quality and efficiency of preoperative assessment for complex cases.
[0086] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0087] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.
Claims
1. A method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion, characterized in that, Includes the following steps: S1: Establish a dedicated latent space encoding-decoding framework to map preoperative corneal topography / anterior segment tomography to a low-dimensional latent space. The latent space encoding-decoding framework includes an encoder and a decoder. The encoder consists of four stages, each containing two sets of residual blocks and a downsampling layer. The attention module is only set in the fourth stage of the encoder. The decoder consists of four stages, symmetrical to the encoder. Each stage fuses the intermediate features of the encoder through skip connections and gradually increases the spatial dimension and decreases the channel dimension. The latent space encoding-decoding framework includes a consistency reconstruction mechanism to ensure that the original image can be reconstructed by the decoder after the preoperative image is encoded and compressed to the latent space, thus ensuring the reliability of feature extraction. The decoder output layer adopts a residual connection structure to preserve latent space feature details and reduce information loss during image reconstruction. S2: In the latent space, combined with the preoperative image conditions, the structured change pattern of "preoperative → postoperative" is learned through diffusion denoising training to obtain the diffusion model; S3: Based on the trained diffusion model, input the preoperative images of new patients who did not participate in the diffusion model training, and generate the prediction results of the anterior segment tomography of the visualized postoperative corneal topography. A central region enhancement module is introduced to prioritize optimizing the detailed features of the central corneal region during the generation process. The U-Net network output is optimized to reduce residual noise. The predicted state, the predicted state The expression is: ; To predict the state, Degraded image at time t, For the target image, Predict noise for the network; S4: Outputs a complete predicted image, providing intuitive visual criteria for postoperative corneal morphology; The diffusion model training employs a stochastic differential equation (SDE)-based modeling method to simulate the image noise addition and denoising process in the latent space, learning the postoperative morphological change patterns. The positive SDE noise addition process is as follows: ,in, Degraded image at time t, For the target image, and These are the time-varying regression coefficient and the diffusion intensity, respectively. For Wiener process; The reverse generation stage is achieved by solving the inverse SDE, and a noise prediction U-Net network is introduced. Estimating the score function The following reconstruction equation is constructed: ; This is a Wiener process in reverse time.
2. The method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion according to claim 1, characterized in that: Preoperative imaging conditions include corneal curvature, thickness, and anterior chamber depth data.
3. The method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion according to claim 2, characterized in that: The diffusion denoising training employs a conditional noise scheduling strategy, dynamically adjusting the noise addition intensity based on preoperative image features to improve the model's adaptability to cases.
4. The method for generating anterior segment tomographic images of corneal topography after ICL surgery based on latent space diffusion according to claim 1, characterized in that: Postoperative predictive images include three-dimensional visualizations of corneal curvature distribution, thickness changes, and anterior chamber structure, covering all dimensions of refractive surgery assessment needs.
Citation Information
Patent Citations
Double-view industrial CT fault online reconstruction method based on hidden space condition diffusion model
CN117635745A
Remote sensing image super-resolution reconstruction method and system based on diffusion model
CN118735785A
SAR (Synthetic Aperture Radar) image ship detection system, method and terminal based on center feature enhancement
CN120259881A
Image reconstruction method and device based on potential diffusion model, and medium
CN120510039A