Endoscopic Viewpoint Estimation Using Style-Transfer Depth Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing endoscopic navigation methods face challenges in accurately estimating the viewpoint difference between real and virtual images of a tubular structure, particularly due to noise and fine textures not captured in virtual endoscopic images, which can hinder precise navigation within narrow and complex structures like the bronchus.
Innovation Solution
An information processing apparatus and method that uses a transformation model, trained on real and virtual images, to transform images into different styles and generate depth images, allowing for accurate estimation of viewpoint differences by comparing similarity between input and transformed depth images, thereby improving navigation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If virtual endoscopic images are generated based on three-dimensional images for navigation, then the overall navigation capability is improved, but measurement precision of viewpoint difference deteriorates due to noise and fine textures not captured in virtual images
Solution Approach 1:
The patent introduces a transformation model as an intermediary that translates real endoscopic images into virtual endoscopic image style. This mediator enables comparison between real and virtual images by converting the real image domain to match the virtual image domain, thereby resolving the incompatibility caused by noise and texture differences while maintaining accurate viewpoint difference estimation
Solution Approach 2:
The transformation model creates a stylized copy of the real endoscopic image that mimics the virtual endoscopic image appearance. By generating this copied version with matched noise characteristics and texture patterns, the system enables accurate comparison and viewpoint difference calculation without being hindered by domain differences
2Measurement precision
If real endoscopic images are used for navigation, then measurement precision of viewpoint difference is improved, but object-generated harmful factors worsen due to noise and fine textures not captured in virtual images
Solution Approach 1:
The patent converts the harmful noise and texture differences into a benefit by using them as training data characteristics for the transformation model. The model learns to replicate these noise patterns and texture variations, transforming what was previously harmful interference into a matching feature that enables accurate image comparison and viewpoint estimation
3Measurement precision
If transformation model is trained to transform real images into virtual image style, then viewpoint difference estimation accuracy is improved, but device complexity increases
Solution Approach 1:
The patent replaces complex mechanical or manual image alignment methods with a data-driven transformation model based on deep learning. This substitution automates the style transfer process, reducing manual intervention while achieving accurate viewpoint difference estimation through learned transformations rather than manual calibration
Data Source
AI summary
An information processing apparatus including at least one processor, wherein the processor is configured to: acquire an input image including a real image, which represents an interior wall of a tubular structure, or a virtual image, which represents, in a pseudo manner, the interior wall as viewed from a virtual viewpoint; use a transformation model to transform the input image into a transformed image having an image style that is not included in the input image; acquire an input depth image, which represents a distance for each pixel from a viewpoint of the input image to the interior wall, and a transformed depth image, which represents a distance for each pixel from a viewpoint of the transformed image to the interior wall; and train the transformation model by using a loss function including a degree of similarity between the input depth image and the transformed depth image.


