Methods, systems, and computer readable media for structure-preserving image translation for depth estimation in endoscopy

A neural network translates synthetic endoscopy images to real images while preserving depth, addressing the lack of realistic reflectance in synthetic data, enhancing depth and camera pose estimation for improved 3D reconstruction in endoscopy.

WO2025250665A1PCT designated stage Publication Date: 2025-12-04THE UNIV OF NORTH CAROLINA AT CHAPEL HILL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/031232
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-05-28
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing synthetic endoscopy images lack realistic reflectance properties of mucus and living tissue, leading to inaccurate depth and camera pose estimation in endoscopy procedures due to insufficient real depth-annotated video frames for training.

Method used

A neural network is trained to translate synthetic endoscopy images to a real image domain while preserving depth information using a cycle generative adversarial network (CycleGAN) with a mutual information loss constraint, incorporating oblique and en face views, and simulating reflectance properties from real endoscopy images.

Benefits of technology

The method generates realistic-looking synthetic images that retain depth information, improving single-frame depth estimation and camera pose estimation accuracy, bridging the gap between synthetic and clinical data for enhanced 3D reconstruction in endoscopy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025031232_04122025_PF_FP_ABST
    Figure US2025031232_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A method for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images includes providing synthetic endoscopy images and corresponding depth maps as source inputs to a neural network. The method further includes providing real endoscopy images as target outputs of the neural network. The method further includes training the neural network to translate the synthetic endoscopy images to a real image domain using preservation of depth information as a constraint on the training and to translate the real endoscopy images to a synthetic image domain. The method further includes, after the training, providing a synthetic endoscopy as input to the trained neural network and generating, as output of the trained neural network, a synthetic image augmented with information from the real endoscopy images while retaining its depth map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket No.421 / 546 PCT METHODS, SYSTEMS, AND COMPUTER READABLE MEDIA FOR STRUCTURE-PRESERVING IMAGE TRANSLATION FOR DEPTH ESTIMATION IN ENDOSCOPY PRIORITY CLAIM This application claims the priority benefit of U.S. Provisional Patent Application Serial No. 63 / 652,649 filed May 28, 2024 and U.S. Provisional Patent Application Serial No.63 / 665,250 filed June 27, 2024, the disclosure of each of which is incorporated herein by reference in its entirety. TECHNICAL FIELD The subject matter described herein relates to training a neural network to generate enhanced synthetic endoscopy images to include reflectance properties learned from real endoscopy images while preserving depth information from depth maps associated with the original synthetic endoscopy images. BACKGROUND Synthetic endoscopy images are computer generated images of surfaces of endoscoped organs that simulate real images captured by an endoscopy camera. Real or clinical endoscopy images are images captured by an endoscopy camera during a real endoscopy procedure. Synthetic endoscopy images can be annotated with depth maps indicating depth values for generated pixels in the synthetic images of the surfaces viewed through the endoscopic camera. As used herein, “depth” associated with a pixel is the distance from the camera to the location on the surface being imaged corresponding to the pixel. Due to the size of endoscopes and the nature of the procedure, the inclusion of depth sensors in real endoscopes is uncommon. As a result, real endoscopy images usually cannot be annotated with depth maps from sensors. Depth estimators estimate depths of endoscopically viewed surfaces captured in video frames from real endoscopic procedures. It is desirable to Attorney Docket No.421 / 546 PCT train depth estimators to be as accurate as possible so that the location and depths in the real images can be accurately reported in real time during an endoscopy procedure. Camera pose estimators can also rely on accurate depth information to estimate camera pose. One way to train a depth or camera pose estimator for an endoscopy is to use real image images and corresponding depth maps. However, there may be an insufficient volume of real depth-annotated endoscopic video frames to train the dept or camera pose estimator. Because of the unavailability of a sufficient number of real video frames for training a depth or camera pose estimator, synthetic images can be used. However, one problem with synthetic endoscopic images is that they do not appear realistic when compared to real images of endoscopically viewed surfaces because the synthetic images do not accurately model the reflectance properties of mucus and living tissue. As a result, a depth or camera pose estimator trained on synthetic images may fail to generate accurate depth or camera pose estimates where real images show specularity, for example, due to wetness of tissue in the real images. Accordingly, in light of these and other difficulties, there exists a need for improved methods, systems, and computer readable media for generating synthetic endoscopy images augmented with effects from real endoscopy images while preserving depth information associated with the synthetic images. SUMMARY A method for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images includes providing synthetic endoscopy images and corresponding depth maps source inputs to a neural network. The method further includes providing real endoscopy images as target outputs of the neural network. The method further includes training the neural network to translate the synthetic endoscopy images to a real image domain using preservation of depth information as a constraint on the training and to translate the real endoscopy Attorney Docket No.421 / 546 PCT images to a synthetic image domain. The method further includes, after the training, providing a synthetic endoscopy image as input to the trained neural network and generating, as output of the trained neural network, a synthetic image augmented with information from the real endoscopy images while remaining consistent with its depth map. According to another aspect of the subject matter described herein, providing the synthetic endoscopy images and the real endoscopy images as inputs to a neural network includes providing the synthetic and real endoscopy images as inputs to a cycle generative adversarial network (CycleGAN). According to another aspect of the subject matter described herein, training the neural network to translate the synthetic endoscopy images to the real image domain using preservation of depth information as a constraint includes when evaluating differences between untranslated depth information and translated images. According to another aspect of the subject matter described herein, providing the synthetic endoscopy images as source inputs to the neural network includes providing synthetic images that include oblique views and en face views of endoscopic surfaces. According to another aspect of the subject matter described herein, providing the real endoscopy images as target outputs of the neural network includes providing real images that include oblique views and en face views of endoscopic surfaces. According to another aspect of the subject matter described herein, the synthetic image augmented with the information from the real endoscopy images includes a synthetic image of an oblique or an en face view of an endoscopic surface. According to another aspect of the subject matter described herein, the synthetic image generated as output is modified, via a translation process of generating the synthetic image, to simulate reflectance properties of mucus and tissue from the real endoscopy images. Attorney Docket No.421 / 546 PCT According to another aspect of the subject matter described herein, the reflectance properties include specularity, subsurface scattering, translucence, and non-Lambertian / non-uniform reflectance. According to another aspect of the subject matter described herein, the method for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images includes using the augmented synthetic image from the trained neural network and the depth map to train a single frame depth estimator to estimate depth from a single frame of endoscopic video. According to another aspect of the subject matter described herein, structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images includes using the single frame depth estimator in a processing pipeline to generate reconstructed images of endoscopic surfaces. According to another aspect of the subject matter described herein, structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images includes using pairs of augmented synthetic images from the trained neural network and their corresponding synthetic depth maps to train a camera pose estimator to estimate change in camera pose from a pair of frames of endoscopic video. The subject matter described herein can be implemented in software in combination with hardware and / or firmware. For example, the subject matter described herein can be implemented in software executed by a processor. In one exemplary implementation, the subject matter described herein can be implemented using a non-transitory computer readable medium having stored thereon computer executable instructions that when executed by the processor of a computer control the computer to perform steps. Exemplary computer readable media suitable for implementing the subject matter described herein include non-transitory computer-readable media, such as disk memory devices, chip memory devices, programmable logic devices, and application specific integrated circuits. In addition, a computer readable medium that implements the subject matter described herein may Attorney Docket No.421 / 546 PCT be located on a single device or computing platform or may be distributed across multiple devices or computing platforms. BRIEF DESCRIPTION OF THE DRAWINGS Exemplary implementations of the subject matter described herein will now be explained with reference to the accompanying drawings, of which: Figure 1 illustrates the overall framework for image translation, depth estimation, and global depth fusion; Figures 2A and 2B illustrate sample frames from various datasets. Figure 2A illustrates textures from SimCol3D

[0010] (left), C3VD (center), and proposed oblique dataset (right). The images in Figure 2B illustrate Viewpoint categories in colonoscopy: axial (left, Colon 10K [7], oblique (center) and en face (right); Figure 3 illustrates the losses used in an image translation frameworkwith image domains ^ and ^, generators ^: ^ → ^ and ^: ^ → ^, anddiscriminators ^^ and ^^. Let ^ ∈ ^, ^ ∈ ^ denote data samples and let ^^^denote the depth map corresponding to sample ^. Downstream depthestimation uses output of generator ^^^^;Figure 4 illustrates image translation examples comparing the original SimCol3D input frame, our translated image, closest image by SSIM in oblique dataset, and translated image with CycleGAN

[0019] ; Figure 5A and 5B illustrate depth estimation on an oblique dataset. The boxes highlight differences. Our image translation framework improves monocular depth estimation in general, with best performance using our proposed dataset as the translation target; Figures 6A and 6B illustrate depth estimation on en face dataset. Boxes highlight differences. Notable improvements from our framework on frames with few geometric features; Figure 7 is a block diagram illustrating an exemplary system for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images; Attorney Docket No.421 / 546 PCT Figure 8 is a flow chart illustrating an exemplary process for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images; Figure 9 is a block diagram illustrating a framework for an experiment comparing reconstructed colonoscopic video frames generated from depth and camera pose estimates from different camera pose and depth estimators; Figure 10A-10C respectively illustrate a real colonoscopic video frame (Figure 10A), a reconstruction of the colonoscopic video frame using depth and camera pose estimates from the Colonoscopy Depth Estimation (ColDE) depth and pose estimator (Figure 10B), and a reconstruction of the colonoscopic video frame using depth and camera pose estimates generated from DroidSlam trained based on enhanced synthetic colonoscopic video frames (Figure 10C); and Figure 11A-11C respectively illustrate another example of a real colonoscopic video frame (Figure 11A), a reconstruction of the colonoscopic video frame using depth and camera pose estimates from the ColDE depth and pose estimator (Figure 11B), and a reconstruction of the colonoscopic video frame using depth and camera pose estimates generated from DroidSlam trained based on enhanced synthetic colonoscopic video frames (Figure 11C). DETAILED DESCRIPTION Real-time 3D reconstruction from endoscopy video can enable the detection and measurement of missed surface areas (blind spots). In commonly-assigned, co-pending PCT Patent Application No. PCT / US24 / 18372, filed March 3, 2024 and entitled, “Methods, Systems, and Computer Readable Media for Colonoscopic Blind Spot Detection,” the disclosure of which is incorporated herein by reference in its entirety, we have described the pipeline to convert the raw video frames from the endoscope into “chunks” (short sequences of frames) suitable for reconstruction, the subsequent use of deep neural networks (DNNs) to estimate depths and surface normals for each frame, and the fusion of the frame-wise depth maps Attorney Docket No.421 / 546 PCT into a global surface that can be used for blind spot detection and further downstream tasks. The example below illustrates a framework of structure- preserving image translation enabling improved single-frame depth estimation for the colonoscopic environment. Single-frame depth estimation refers to depth estimation from a single frame of endoscopic video using a depth estimator trained from enhanced synthetic endoscopic video frames and their corresponding depth information. This framework allows us to replace the previous method for estimation of frame-wise depths such that we are now able to successfully utilize frames from challenging perspectives that view the colon surface from oblique and en face directions. The image translation and depth estimation components are evaluated and described in detail below. Figure 1 illustrates the overall framework for image translation, depth estimation, and global depth fusion. This same framework can be used for realistic synthetic image generation, depth estimation, and global depth fusion for other types of endoscopic video, especially types where the anatomic structures viewed through the endoscopic camera are tubular and exhibit reflectance properties due to surface wetness of tissue. Examples of endoscopic video images that can be processed using a depth estimator trained with the enhanced synthetic images described herein include endoscopic video images of the esophagus, vaginal canal, or any other tubular structure that can be imaged with an endoscopic camera. This framework addresses the issue of the lack of realistic-looking data with accurate depth maps that can be used to train neural networks for depth estimation. The colon surface is challenging to model due to the complex light reflectance properties of biological tissue and mucus. This framework implicitly describes these reflectance properties and bridges the gap between unrealistic computer-generated models and clinical frames. We note that while data augmentation using synthetic data is a standard approach to improve performance in deep learning, the challenges of the endoscopic environment (including complex tissue reflectance, low geometric texture, co-located light source and camera, and proximity of the camera to the surface) makes this Attorney Docket No.421 / 546 PCT framework a novel augmentation method in comparison to direct use of synthetic images. In general, image translation describes the task of modifying an input image from a source distribution to look like images from a target distribution. We can define a distribution of images via a discrete set of images. Here, our source distribution is a set of images from computer-generated colonoscopies (with corresponding depth maps for each image) and our target distribution of a set of image frames from clinical colonoscopies. We train a neural network to perform image translation in this way and thus are able to modify an entire set of synthetic images to appear realistic like clinical images. The image translation model builds upon an existing neural network in the literature modified to explicitly preserve structure (depth information) via the introduction of an optimization target to enforce information similarity (in our implementation, we use mutual information loss) between the translated result (RGB image) and the depth map corresponding to the source image. Similarity metrics other than mutual information loss can be used without departing from the scope of the subject matter described herein. Structure-preservation ensures that depth information is not lost or distorted during the translation process, allowing us to use the depth information corresponding to the original synthetic image for downstream tasks. Previous approaches using image translation to improve depth estimation in the colonoscopy domain either directly estimate depths from RGB images or require a pre-existing depth estimation or feature extraction method. Our approach performs the generic task of translation between RGB image distributions and does not require depth estimation or feature extraction on the target distribution. As a result of this design, the framework can be easily adapted to applications beyond colonoscopy video and the output of our translation model can be used for estimation of parameters associated with a pixel other than depth, such as camera pose. The output of our image translation model is the set of translated images (now appearing realistic to clinical images but retaining the depth information of the original synthetic image). In the Figure 4, we demonstrate Attorney Docket No.421 / 546 PCT the differences between the synthetic images (first row), our translated result (second row) and similar clinical images (third row). For single-frame depth estimation, we use the output of our image translation model by pairing each translated image with the corresponding depth map from the synthetic dataset. Subsequently, we train a convolutional neural network (CNN) to take these translated RGB images as input and produce the corresponding depth maps as output. The architecture chosen for this CNN is widely used in the literature for depth estimation and we make no modifications to it and train it simply using a common objective function of mean squared error. The novelty here lies in the framework design of using translated image / depth pairs for single-frame endoscopic depth estimation. This design approach allows us to easily substitute arbitrary architectures for the CNN we have chosen should the need arise or technology continue to develop, to simplify the process of training such models by enabling straightforward optimization targets, and to improve performance in challenging perspectives that are under-served by alternative (un- or self- supervised) optimization targets. We find that this approach provides superior frame-wise depth predictions to previous methods, especially in frames with challenging perspectives. When we use those depth predictions in our global fusion pipeline (described in a previous patent), we find improvements in the accuracy and reliability of the resulting reconstructions. It is believed that these translated image and depth pairs can be used to improve prediction of the change in camera pose, which would differ in the use of this image translation framework from previous methods directly utilizing synthetic images or self-supervised approaches. The combination of these improved pose and depth prediction results would further improve surface reconstructions over previous methods and bring us closer to our overall goal of providing reliable and accurate 3D reconstructions of colonoscopies in real-time. Monocular depth estimation in endoscopy video aims to overcome the unusual lighting properties of the endoscopic environment that necessitate specialized methods. One of the major challenges in this area is the domain Attorney Docket No.421 / 546 PCT gap between annotated but unrealistic synthetic data and unannotated but realistic clinical data. Previous attempts to bridge this domain gap directly target the depth estimation task itself. We propose a general pipeline of structure-preserving synthetic-to-real (sim2real) image translation (producing a modified version of the input image) to retain depth geometry through the translation process. In this way, we can generate large quantities of realistic- looking synthetic images for supervised or semi-supervised training of arbitrary networks for depth estimation with improved generalization to the clinical domain. We also propose a dataset of hand-picked sequences from clinical colonoscopies to improve the image translation process. We demonstrate the simultaneous realism of the translated images and preservation of depth maps via the performance of downstream depth estimators on various data. The following Sections 1-6 describe a methodology for structure- preserving image translation to generate augmented synthetic colonoscopy images while preserving depth information. Section 7 describes the overall methodology and a computer implementation of the methodology. 1. Introduction Colorectal cancer (CRC) is one of the leading causes of cancer mortality in the United States; the American Cancer Society estimates that there will be over 150,000 new cases and 50,000 deaths in 2024. Increased screening is one of the factors contributing to reductions in mortality

[0013] . Optical colonoscopy is the gold standard method for CRC screening but its effectiveness is highly dependent upon the skill of the physician performing the examination [9]. In particular, around 20% of potentially pre-cancerous polyps are missed during colonoscopies

[0012]

[0014] . 3D reconstruction from optical colonoscopy video can improve efficacy via guidance and visualization to the physician, automatic measurements, and autonomous navigation. One of the major challenges in this area is the lack of realistic data suitable for training neural networks to perform depth and pose estimation. While synthetic

[0010] and phantom [2] datasets exist, they do not Attorney Docket No.421 / 546 PCT accurately represent the reflectance properties of in vivo tissue. Previous approaches towards closing the domain gap [8]

[0011]

[0015] do not target challenging viewpoints making up the majority of colonoscopy videos. In this work, we propose an image translation method that generates realistic-looking video frames from synthetic colonoscopies while preserving the depth information and without requiring complex modeling of mucus and in vivo tissue. In this way, we are able to bridge the gap between unrealistic synthetic data with dense ground truth depth annotation and realistic but un-annotated clinical data to improve depth estimation on unseen clinical data. In addition, we introduce two new datasets of manually selected frames from clinical colonoscopies representing viewpoints that are particularly challenging for depth estimation and downstream reconstruction. This data both improves the realism of our image translation results and provides a dataset against which to test the quality of depth estimation results. 2. Related Work Prior datasets targeting reconstruction from colonoscopy come from clinical procedures (EndoMapper [1], Colon10K [7]), fully synthetic procedures (SimCol3D

[0010] , Zhang et al.

[0018] ), or robotic colonoscopy of a silicone phantom model of the colon (C3VD [2]). Clinical data by nature does not have per-frame depth or pose annotations; while synthetic and phantom data have such annotations, the geometry and light reflectance properties of human tissue is challenging to replicate synthetically and therefore the textures present in the synthetic and phantom data are notably different from those observed in clinical practice (Figure 2A). While the use of image translation to bridge the synthetic to clinical domain gap has been addressed previously (Section 2.1), we propose a general modular framework particularly targeting depth estimation (Section 2.2) on challenging viewpoints. This is the first work that performs structure-preserving image translation from the synthetic to clinical colonoscopy domain without requiring a pre-trained depth estimator or feature extractor in the target clinical domain. Attorney Docket No.421 / 546 PCT 2.1^Domain^gap^ Using image translation for colonoscopic depth estimation, Rau et al.

[0011] propose image-to-depth translation to directly estimate depths from images. In contrast, Mahmood and Durr [8] combine synthetic depth estimation with real-to-synthetic image translation at inference. For other tasks, Chen et al. [3] propose a structure-preserving image- to-image generative adversarial network (GAN) to improve segmentation using mutual information in the latent encoding. Similarly, Yoon et al.

[0017] propose using GAN-based dataset augmentation to boost performance. For general-purpose image translation, many previous works build upon CycleGAN

[0019] due to the structure preservation implicit in the cyclical architecture. Cheng et al. [4] present a structure-preserving alternative that decomposes style (extracted via a pretrained autoencoder) from structure (extracted via a pretrained monocular depth estimator). ^ 2.2^Depth^estimation^ In order to demonstrate the effectiveness of our image translation approach, we use performance on monocular depth estimation as the metric for comparison. Wang et al.

[0015] propose a self-supervised extension of Monodepth2 [6] for the colonoscopy domain with an iterative refinement step. For general depth estimation, modern Transformer-based methods [5]

[0016] demonstrate high-quality depth estimation results on non-medical data but rely on large training datasets. 3. Data Generally, we can categorize the viewpoint of a single frame as axial, oblique, or en face (Figure 2B). Oblique and en face viewpoints can be challenging for depth estimation due to the lack of strong geometric features. However, they make up about 70% of non-obfuscated frames within a colonoscopy video so reliable depth estimation from these views and their subsequent incorporation into reconstruction provides significant additional information about surface geometry over reconstruction from axial views alone. Attorney Docket No.421 / 546 PCT In this work, we introduce two distinct datasets: the first of oblique views and the second of en face views. Both consist of sequences of consecutive frames manually selected from a library of video recordings of full colonoscopy procedures. The datasets have been curated on the basis of the viewpoint of each frame such that a sequence extends as long as each consecutive frame is of the same viewpoint category modulo gaps of up to 30 consecutive frames with excessive obfuscation (e.g., water drops on the lens). All frames are pre-processed in the same manner. Using computed camera intrinsics and the MATLAB undistortFisheyeImage function, we warp fisheye lens projection into a pinhole projection. We then crop the image toremove the unused image area and resize to 270 ൈ 216 pixels. The originalvideos were recorded using CF and PCF series Olympus colonoscopes witha raw image size of 1350 ൈ 1080. Office of Human Research Ethics hasdetermined that this work does not constitute human subject research and does not require Internal Review Board approval. ^ 3.1^Oblique^dataset^ The first dataset, which we call the oblique dataset, consists of sequences manually selected to exclude obfuscated frames, fully axial views, and fully en face views. There are 93 sequences totaling 16,756 frames. Each sequence has between 2 and 586 frames, averaging 180 frames per sequence. We randomly divide this dataset into 90% train and 10% test partitions with divisions being made at the sequence (rather than frame) level. ^ 3.2^En^face^dataset^ The second dataset, which we call the en face dataset, consists of sequences manually selected to exclude obfuscated frames and contain only fully en face views. There are 14 sequences totaling 816 frames. Each sequence has been 15 and 136 frames, averaging 58 frames per sequence. In this work, we only use this dataset for evaluation due to its small size. Attorney Docket No.421 / 546 PCT 4. Methods We demonstrate the realism of the image translation result and effectiveness of our proposed structure-preserving loss term via downstream depth estimation. Figure 3 illustrates our framework. Our image translation result is additionally improved with the use of our proposed data over pre- existing datasets. 4.1^Image^translation^ We use the standard CycleGAN losses with generator ^: ^ → ^ anddiscriminator ^^ (and similarly generator ^: ^ → ^ and discriminator ^^):^GAN^^, ^^, ^, ^^ ൌ ^^∼^data^^^^log^^^^^^ ^ ^^∼^data^^^ ^log ^1 െ ^^൫^^^^൯^^ (1) In order to explicitly constrain the translation to preserve depth information so that the depths and translated image pairs can be used to train supervised depth estimation, we add mutual information loss. Mutual information circumvents the recursive problem of depth estimation or feature extraction on challenging clinical data. where ^^^is the ground truth depth map corresponding to the input sample ^, ^^^⋅^ is image intensity (average of all color channels), and ^ is the number of total combinations. We discretize the data into 256 bins for both depths and intensity. This loss is only applied for A → B translation. Our full objective is: We train CycleGAN to perform image translation with SimCol3D as domain A and our proposed oblique dataset as domain B. Table 2 describes the ablation experiments for this portion. The subject matter described herein is not limited to using synthetic images from the SimCol3D dataset. The synthetic images Attorney Docket No.421 / 546 PCT used to train the model may originate from any simulation platform or dataset, proprietary or public, 3D rendered or otherwise. 4.2 Depth estimation Here we are interested in model-agnostic depth estimation performance as a metric for the structure preservation through the image translation process rather than depth estimation itself and therefore note any architecture could be used. A significant mismatch between the translated image and the original depth map (lack of structure preservation during translation) will result in poor depth estimation generalization for any model. We use the Monodepth2 [6] architecture trained fully supervised from scratch. We pair the RGB result from image translation with the depth map from the original synthetic data for labels. In order to avoid data overlap, we measure performance of all models on C3VD [2]. We convert the fisheye projection to a pinhole projection using the OpenCV undistort function. ^ 4.3^Implementation^details^ For image translation, we train the modified CycleGAN using four NVIDIA Titan Xp GPUs for 30 epochs. We use the Adam optimizer and initiallearning rate of 2e-4. We use weights ^GAN ൌ 10.0, ^cyc ൌ 0.5, and ^MI ൌ 1.0.For depth estimation with Monodepth2, we use a Resnet34 backbone and train the model using a NVIDIA Quadro RTX 5000 for 20 epochs with mean squared error loss, Adam optimizer, and initial learning rate of 1e-4. We usedata augmentations of random cropping to 256 ൈ 256 and random horizontaland vertical flipping. At inference, we rescale images to 256 ൈ 256. All code isimplemented using PyTorch. 5. Results 5.1^Image^translation^ Overall, we find that the translation result (Figure 4) has both improved texture realism and retains the overall geometry of the input image. Most notably, the translation adds the specular points missing from SimCol3D Attorney Docket No.421 / 546 PCT without explicit representation. The specularity is distributed in a manner consistent with our expectation that surfaces closer to the colonoscope tip and having surface normal directions parallel with the viewing direction will more likely exhibit specular effects than those either farther from the tip or with surface normal direction different from that of the viewing direction. In Table1, we compare translation metrics against translation with ^MI ൌ 0 (vanillaCycleGAN) and find the metrics support our perception of improved translationresults when ^MI ^ 0.5.2^Depth^estimation^ While we measure depth estimation performance on C3VD [2] for comparison against baseline models, the textures and geometries represented in that dataset remain different from those observed in clinical practice. Due to the lack of annotated clinical data, we select this dataset for its better realism compared to other options. Additionally, we provide qualitative assessment of depth predictions on the test partition of our proposed oblique dataset and full en face dataset. Table 1: Image translations metrics against oblique dataset. Using ^MIhelps the model produce images more similar to the distribution of test images. C3VD^ In Table 3, we provide metrics computed after median rescaling to adjust depth scale across models. For zero-shot models (relying on generalization), we find that our framework produces the best performance in most metrics. We also find that the performance is very similar to that achieved by NormDepth

[0015] , which uses a similar Monodepth2-based architecture but is trained upon a much larger clinical dataset. We conclude that the Attorney Docket No.421 / 546 PCT performance on this dataset is satisfactory given the architecture and simplicity of the evaluation dataset and look for a larger performance gap on more challenging clinical frames. Oblique^ In Figure 5A, we show a few examples of depth estimation using NormDepth and our framework evaluated on images from the proposed oblique test partition (additional examples in Figure 5B). We have not used masking to prevent depth distortions at specular points. Overall, we see that NormDepth is biased towards predicting a depth depression near the center of the frame and poor predictions near occlusion boundaries. Meanwhile, the baseline model produces significant and repeated errors in the depth map at specular points. Our proposed model produces the best representation of rounded haustral ridges and better distinction between structures. Compared to OursC10Kand OursC3VD, our model produces depths with stronger discontinuities at occlusion boundaries and overall captures a more nuanced and accurate surface geometry. Table 2: Ablations on translation target dataset and use of MI loss Table 3: Depth evaluation on C3VD (mm). The best categorical performance is highlighted in bold. Multi-shot models train on C3VD while zero-shot rely on generalization. On easy data, our framework, baselines, and ablations perform similarly. Attorney Docket No.421 / 546 PCT ^ En face In Figures 6A and 6B, we show a few examples of depth estimation using NormDepth and our framework on images from the proposed en face dataset. In these examples, the bias of NormDepth towards predicting a center depth depression is particularly evident, as are the failures of the baseline model in specular areas. Due to the nature of this dataset, there is greater representation of surfaces with strong visual texture from vasculature. Thus, we can see that our proposed method has overall improved representation of the overall surface geometry compared to ablations but can also produce distortions to the depth map at regions with strong vascular texture. 6. Conclusions We have demonstrated that structure-preserving sim2real image translation improves monocular depth estimation in challenging colonoscopic frames. To aid this task, we introduce two datasets of hand-picked sequences from clinical data focusing on viewpoints that are under-represented in existing datasets. The image translation results improve texture realism (especially for specular points) while retaining sufficient depth geometry for successful subsequent training of depth estimator networks. We provide evaluation of depth estimation on C3VD and qualitative evaluations on our proposed datasets, finding significant performance improvements on challenging frames using this framework. Attorney Docket No.421 / 546 PCT ^ 6.1^Limitations^and^Future^Work^ Depth distortions in areas with strongly visible vasculature and few geometric features could be ameliorated by incorporating additional data into the translation target. Future work could focus on applying this approach to pose estimation or larger depth estimation models such as Transformer-based models given a larger synthetic dataset. 7. Overall Methodology and Computer Implementation Figure 7 is a block diagram illustrating an exemplary system for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images. Referring to Figure 7, the system includes a computing platform 700 including at least one processor 702 and memory 704. Computing platform 700 further includes a trained structure-preserving image translation neural network 706 that is trained to receive, as input, synthetic endoscopy images and depth maps and to generate, as output, augmented synthetic endoscopy images and depth maps, where the augmented images are augmented or changed based on real endoscopy images to simulate reflectance properties of surfaces shown in the real endoscopy images. The augmented synthetic endoscopy images output from trained structure-preserving image translation neural network 706 can be used as training inputs to a depth estimator 708 or a camera pose estimator 710. Depth estimator 708 may be a neural network that is trained to receive, as input, real endoscopy images and to generate, as output, depth maps for the endoscopy images. In our implementation, we use a Monodepth2 architecture incorporating skip connections and train the network by minimizing the mean squared error between the model output and the correct depth map. Camera pose estimator 710 is trained to estimate the pose of an endoscopic camera from pairs of input video frames and corresponding depth maps. We recommend an encoder-decoder architecture for this. The augmented synthetic endoscopy images and depth maps can be used as training inputs to either of these neural networks and it is believed that they Attorney Docket No.421 / 546 PCT will increase the accuracy of the depth or camera pose estimation over training based on purely synthetic images. Figure 8 is a flow chart illustrating an exemplary process for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images. Referring to Figure 8, in step 800, the process includes providing synthetic endoscopy images and corresponding depth maps as source inputs to a neural network. For example, synthetic endoscopy images and depth maps can be provided as training source inputs to a neural network, such as a CycleGAN. In step 802, the process further includes providing real endoscopy images as target outputs of the neural network. For example, real images from endoscopic video frames can be provided to the CycleGAN. The real endoscopy images can be images of any internal structure that can be obtained using an endoscope, including a colonoscope, a bronchoscope, and a capsule endoscope. The images may be of the colon, esophagus, airway, vaginal canal, urinary tract, etc. The synthetic images may be computer- generated images of any of the real image types. In step 804, the process further includes training the neural network to translate the synthetic endoscopy images to a real image domain using preservation of depth information as a constraint on the training and to translate the real endoscopy images to a synthetic image domain. For example, the CycleGAN may be trained in the manner described herein with depth constraints on the synthetic to real image translations. In step 806, the process further includes, after the training, providing a synthetic endoscopy image as input to the trained neural network and generating, as output of the trained neural network, a synthetic image augmented with information from the real endoscopy images while retaining its depth map. For example, after the CycleGAN is trained, synthetic endoscopy images and depth maps can be provided as inputs, and the network will generate, as outputs, versions of the synthetic images modified to simulate reflectance properties of the tissue being imaged. Attorney Docket No.421 / 546 PCT One benefit or feature of the subject matter described herein is that the quantity and types of frames that can be reconstructed from clinical endoscopic video are increased. For example, the enhanced synthetic image frames can be used to train a depth estimator, and the depth estimator can then be used as part of the processing pipeline to reconstruct endoscopically viewed surfaces for blind spot detection or other applications. In one set of experiments, a CNN trained with enhanced synthetic images as described herein is used to generate depth maps and camera pose estimates that are used to reconstruct images of colonic surfaces. In the experiments, the CNN trained with the enhanced synthetic images is DroidSlam. For comparison, a second CNN, referred to as Colonoscopy Depth Estimation (ColDE), described in Zhang, et al., ColDE: A Depth Estimation Framework for Colonoscopy Reconstruction, arXiv:2111.10371 [eess.IV] (November 19, 2021), the disclosure of which is incorporated herein by reference in its entirety, is used to generate depth and camera pose estimates, which are also used to reconstruct images of colonic surfaces. Figure 9 illustrates the framework for the experiments. Referring to Figure 9, a computing platform 900 includes at least one processor 902 and memory 904. A depth and camera pose estimator implemented using DroidSlam trained with the enhanced synthetic images described herein is illustrated by enhanced DroidSlam depth / pose estimator 906. For comparison, depth and camera pose are also obtained using a CoIDE depth / camera pose estimator 908. The depth and camera pose estimates obtained using depth / pose estimators 906 and 908 are input to an image reconstruction pipeline 910, which generates 3D models of endoscopic surfaces from the camera pose and depth estimates. Image reconstruction pipeline 910 may be implemented using the framework described in the above-referenced commonly assigned co-pending patent application. It is understood that enhanced DroidSlam depth / pose estimator 906, CoIDE depth / pose estimator 908, and image reconstruction pipeline 910 may be implemented using computer executable instructions stored in memory 904 and executed by processor 902. Attorney Docket No.421 / 546 PCT Figure 10A-10C respectively illustrate a real colonoscopic video frame (Figure 10A), a reconstruction of the colonoscopic video frame using the depth and camera pose estimates from the ColDE depth and pose estimator (Figure 10B), and a reconstruction of the colonoscopic video frame using depth and pose estimates generated using DroidSlam trained based on enhanced synthetic colonoscopic video frames (Figure 10C). Figure 11A-11C respectively illustrate another example of a real colonoscopic video frame (Figure 11A), a reconstruction of the colonoscopic video frame using the depth and camera pose estimates from the ColDE depth and pose estimator (Figure 11B), and a reconstruction of the colonoscopic video frame using depth and pose estimates generated using DroidSlam trained based on enhanced synthetic colonoscopic video frames (Figure 11C). It is believed that the use of the enhanced synthetic endoscopic video images to train depth and camera pose estimators will increase the amount and types of endoscopic video frames that can be used for image reconstruction. By enhancing the synthetic video frames used to train the depth and camera pose estimators, clinical or real endoscopic video frames showing oblique and en face views can be used for reconstruction, and, as a result, curved portions of the anatomic structures shown in the endoscopic video can be reconstructed. The disclosure of each of the following references is hereby incorporated herein by reference in its entirety. References 1. Azagra, P., Sostres, C., Ferrandez, A., Riazuelo, L., Tomasini, C., Barbed, O.L., Morlana, J., Recasens, D., Batlle, V.M., Gómez-Rodríguez, J.J., Elvira, R., López, J., Oriol, C., Civera, J., Tardós, J.D., Murillo, A.C., Lanas, A., Montiel, J. M.M.: Endomapper dataset of complete calibrated endoscopy procedures. Scientific Data 10(1) (October 2023). https: / / doi.org / 10.1038 / s41597-023-02564-7, http: / / dx.doi.org / 10.1038 / s41597-023-02564-7 Attorney Docket No.421 / 546 PCT 2. Bobrow, T.L., Golhar, M., Vijayan, R., Akshintala, V.S., Garcia, J.R., Durr, N.J.: Colonoscopy 3D video dataset with paired depth from 2D-3D registration. Medical Image Analysis p.102956 (2023) 3. Chen, J., Zhang, Z., Xie, X., Li, Y., Xu, T., Ma, K., Zheng, Y.: Beyond mutual information: Generative adversarial network for domain adaptation using information bottleneck constraint. IEEE Transactions on Medical Imaging 41(3), 595–607 (2022). https: / / doi.org / 10.1109 / TMI.2021.3117996 4. Cheng, M.M., Liu, X.C., Wang, J., Lu, S.P., Lai, Y.K., Rosin, P.L.: Structure-Preserving Neural Style Transfer. IEEE Transactions on Image Processing 29,909–920(2020).https: / / doi.org / 10.1109 / TIP.2019.2936746, https: / / ieeexplore.ieee.org / document / 8816670 / 5. Eftekhar, A., Sax, A., Bachmann, R., Malik, J., Zamir, A.: Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans (2021) 6. Godard, C., Mac Aodha, O., Firman, M., Brostow, G.J.: Digging into self-supervised monocular depth prediction (October 2019) 7. Ma, R., McGill, S.K., Wang, R., Rosenman, J., Frahm, J.M., Zhang, Y., Pizer, S.: Colon10k: A benchmark for place recognition in colonoscopy. In: 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI). pp. 1279–1283 (2021). https: / / doi.org / 10.1109 / ISBI48211.2021.9433780 8. Mahmood, F., Durr, N.J.: Deep learning and conditional randomfields- based depth estimation and topographical reconstruction from conventional endoscopy. Medical Image Analysis 48,230–243 (2018). https: / / doi.org / https: / / doi.org / 10.1016 / j.media.2018.06.005, https: / / www.sciencedirect.com / science / article / pii / S1361841518303761 9. Nierengarten, M.B.: Colonoscopy remains the gold standard for screening despite recent tarnish. Cancer 129 (2023). https: / / doi.org / 10.1002 / cncr.34622 10. Rau, A., Bano, S., Jin, Y., Stoyanov, D.: Simcol3D - 3D Reconstruction during Colonoscopy Challenge Dataset (September 2023). https: / / doi.org / 10.5522 / 04 / 24077763.v1 11. Rau, A., Edwards, P.E., Ahmad, O.F., Riordan, P., Janatka, M., Lova, Attorney Docket No.421 / 546 PCT L. B., Danail, S.: Implicit domain adaptation with conditional generative adversarial networks for depth prediction in endoscopy. International Journal of Computer Assisted Radiology and Surgery 14,1167–1176(April 2019). https: / / doi.org / 10.1007 / s11548-019-01962-w 12. van Rijn, J.C., Reitsma, J.B., Stoker, J., Bossuyt, P.M., van Deventer, S.J., Dekker, E.: Polyp miss rate determined by tandem colonoscopy: a systematic review. The American Journal of Gastroenterology (2006). https: / / doi.org / 10.1111 / j.1572-0241.2006.00390.x 13. Siegel, R.L., Giaquinto, A.N., Jemal, A.: Cancer statistics. CA: A Cancer Journal for Clinicians 74 (2024). https: / / doi.org / 10.3322 / caac.21820 14. Vemulapalli, K.C., Lahr, R.E., Rex, D.K.: Most large colorectal polyps missed by gastroenterology fellows at colonoscopy are sessile serrated lesions. Endoscopy International Open (2022). https: / / doi.org / 10.1055 / a- 1784-0959 It will be understood that various details of the subject matter described herein may be changed without departing from the scope of the subject matter described herein. Furthermore, the foregoing description is for the purpose of illustration only, and not for the purpose of limitation, as the subject matter described herein is defined by the claims as set forth hereinafter.

Claims

Attorney Docket No.421 / 546 PCT CLAIMS What is claimed is:

1. A method for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images, the method comprising: providing synthetic endoscopy images and corresponding depth maps as source inputs to a neural network; providing real endoscopy images as target outputs of the neural network; training the neural network to translate the synthetic endoscopy images to a real image domain using preservation of depth information as a constraint on the training and to translate the real endoscopy images to a synthetic image domain; and after the training, providing a synthetic endoscopy image as input to the trained neural network and generating, as output of the trained neural network, a synthetic image augmented with information from the real endoscopy images while retaining its depth map.

2. The method of claim 1 wherein providing the synthetic endoscopy images and the real endoscopy images as the source inputs to the neural network includes providing the synthetic and real endoscopy images as the source inputs to a cycle generative adversarial network (CycleGAN).

3. The method of claim 2 wherein training the neural network to translate the synthetic endoscopy images to the real image domain using preservation of depth information as a constraint when evaluating differences between untranslated depth information and translated images.

4. The method of claim 1 wherein providing the synthetic endoscopy images as source inputs to the neural network includes providing synthetic images that include oblique views and en face views of endoscopic surfaces.Attorney Docket No.421 / 546 PCT 5. The method of claim 4 wherein providing the real endoscopy images as target outputs of the neural network includes providing real images that include oblique views and en face views of endoscopic surfaces.

6. The method of claim 5 wherein the synthetic image augmented with the information from the real endoscopy images includes a synthetic image of an oblique or an en face view of an endoscopic surface.

7. The method of claim 1 wherein the synthetic image generated as output is modified, via a translation process of generating the synthetic image, to simulate reflectance properties of mucus and tissue from the real endoscopy images.

8. The method of claim 7 wherein the reflectance properties include specularity, subsurface scattering, translucence, and non- Lambertian / non-uniform reflectance.

9. The method of claim 1 comprising using the augmented synthetic image from the trained neural network and the depth map to train a single frame depth estimator to estimate depth from a single frame of endoscopic video.

10. The method of claim 9 comprising using the single frame depth estimator in a processing pipeline to generate reconstructed images of endoscopic surfaces.

11. The method of claim 1 comprising using pairs of augmented synthetic images from the trained neural network and their corresponding synthetic depth maps to train a camera pose estimator to estimate change in camera pose from a pair of frames of endoscopic video.

12. A system for structure preserving image translation for augmenting synthetic endoscopy images to include information from real endoscopy images, the system comprising: a computing platform including at least one processor and a memory; and a neural network implemented by the at least one processor for receiving, as input, synthetic endoscopy images and depth information and for generating, as output, synthetic endoscopy images augmentedAttorney Docket No.421 / 546 PCT with information from real endoscopy images while preserving the depth information.

13. The system of claim 12 wherein the neural network comprises a cycle generative adversarial network (CycleGAN).

14. The system of claim 13 wherein the CycleGan is trained to translate synthetic endoscopy images to real image domain using preservation of depth information as a constraint when evaluating differences between untranslated depth information and translated images.

15. The system of claim 12 wherein the synthetic endoscopy images include oblique views and en face views of endoscopic surfaces.

16. The system of claim 15 wherein the synthetic image augmented with the information from the real endoscopy images includes a synthetic image of an oblique or an en face view of an endoscopic surface.

17. The system of claim 12 wherein the synthetic image generated as output is modified, via a translation process of generating the synthetic image, to simulate reflectance properties of mucus and tissue from the real endoscopy images.

18. The system of claim 17 wherein the reflectance properties include specularity, subsurface scattering, translucence, and non- Lambertian / non-uniform reflectance.

19. The system of claim 12 comprising a single frame depth estimator for using the augmented synthetic image from the neural network and the depth map as a training input.

20. The system of claim 19 comprising a processing pipeline that uses the single frame depth estimator to generate reconstructed images of endoscopic surfaces.

21. The system of claim 12 comprising a camera pose estimator for using pairs of augmented synthetic images generated by the neural network and their corresponding depths as training inputs.

22. A non-transitory computer readable medium having stored thereon executable instructions that when executed by a processor of a computer control the computer to perform step comprising:Attorney Docket No.421 / 546 PCT providing synthetic endoscopy images and corresponding depth maps as source inputs to a neural network; providing real endoscopy images as target outputs of the neural network; training the neural network to translate the synthetic endoscopy images to a real image domain using preservation of depth information as a constraint on the training and to translate the real endoscopy images to a synthetic image domain; and after the training, providing a synthetic endoscopy image as input to the trained neural network and generating, as output of the trained neural network, a synthetic image augmented with information from the real endoscopy images while retaining its depth map.

Citation Information

Patent Citations

  • Stereo display system and method for endoscope using shape-from-shading algorithm

    US20170035268A1

  • Methods, systems, and computer readable media for deriving a three-dimensional (3D) textured surface from endoscopic video

    US20200219272A1

  • Methods and apparatuses for generating anatomical models using diagnostic images

    US20220230303A1

  • Robust surgical scene depth estimation using endoscopy

    US20240156325A1