Method and apparatus for generating an anatomical model using diagnostic images
By training a computational model to generate depth images and confidence maps, the problem of lack of depth perception in navigation and positioning of endoscopic systems is solved. This enables efficient generation of three-dimensional anatomical models based on monocular endoscopic images, supporting the accurate and efficient performance of minimally invasive surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOSTON SCIENTIFIC SCIMED INC
- Filing Date
- 2022-01-14
- Publication Date
- 2026-07-31
AI Technical Summary
Existing endoscopic systems lack depth perception during navigation and positioning, which limits the accuracy and efficiency of surgery. Traditional computer-aided navigation systems rely on preoperative images and additional hardware, cannot adapt to lung movement in real time, and vision-based technologies are not effective in endoscopic surgery.
By using a computational model to train and generate depth images and confidence maps, and combining them with monocular endoscopic images to generate a three-dimensional anatomical model, the performance on real patient images is optimized using supervised training and domain adversarial training techniques, achieving efficient three-dimensional reconstruction and navigation.
It enables efficient generation of 3D anatomical models based on monocular endoscopic images, supporting minimally invasive surgeries such as bronchoscopy, reducing reliance on additional hardware, and improving the accuracy and efficiency of surgery.
Smart Images

Figure CN116997928B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 138,186, filed January 15, 2021, entitled “Technique for Determining Depth Estimation of Tissue Images”, 35 USC §119, the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] The present invention generally relates to a process for examining the physical properties of a part of a patient based on an image, and more particularly to a technique for generating a multidimensional model of the part based on a monocular image. Background Technology
[0004] Endoscopy provides a minimally invasive surgical procedure for visually examining internal body cavities. For example, bronchoscopy is an endoscopic diagnostic technique used to directly examine the airways of the lungs by inserting a long, thin endoscope (or bronchoscope) through the patient's trachea and down into the lung pathway. Other types of endoscopes can include colonoscopes for colonoscopy, cystoscopes for the urethra, and colonoscopes for the small intestine. Endoscopes typically include a lighting system for illuminating the patient's internal cavities, a sample retrieval system for collecting samples from within the patient, and an imaging system for capturing images of the patient's internal interior for transmission to a database and / or the operator.
[0005] A major challenge during endoscopic surgery using traditional systems is positioning the endoscope within the patient's luminosity (e.g., the bronchi and bronchioles used for bronchoscopy) for accurate and efficient navigation. Computer vision systems have been developed to provide navigational assistance to the operator in guiding the endoscope to its target. However, existing endoscopes typically have imaging devices that provide a limited two-dimensional or monocular field of view and lack sufficient depth perception. Therefore, due to insufficient visual information, it is difficult for the operator to orient and navigate the endoscope within the patient's luminosity, especially during live diagnostic procedures. The limitations of computer-aided systems further complicate the achievement of successful surgeries, and the overall outcome of endoscopic surgery becomes highly dependent on the operator's experience and skill.
[0006] One navigation solution involves a tracking system (e.g., an electromagnetic (EM) tracking system) that relies on preoperative diagnostic images of the patient (e.g., computed tomography (CT) images). This solution has several drawbacks, including, but not limited to, the need for extensive preoperative patient analysis and the inability to compensate for lung movement during surgery, leading to inaccurate positioning. Other navigation systems use vision-based technologies, which require additional hardware such as sensors and / or complex data manipulation dependent on patient images, limiting their effectiveness, particularly in real-time surgical settings.
[0007] Regarding these and other considerations, improvements to the present invention may be useful. Summary of the Invention
[0008] An overview of the invention is provided to aid understanding, and those skilled in the art will understand that each of the various features of the invention may be used advantageously alone in some cases, or in combination with other features of the invention in others. The inclusion or exclusion of elements, components, etc., in this overview is not intended to limit the scope of the claimed subject matter. In one embodiment, the invention relates to...
[0009] According to various features of the embodiments described, an apparatus includes at least one processor and a memory coupled to the at least one processor. The memory may include instructions that, when executed by the at least one processor, cause the at least one processor to: access a plurality of endoscopic training images comprising a plurality of synthetic images and a plurality of real images; access a plurality of depth real information associated with the plurality of synthetic images; perform supervised training of at least one computational model using the plurality of synthetic images and the plurality of depth real information to generate a synthetic encoder and a synthetic decoder; and perform domain adversarial training on the synthetic encoder using real images to generate a real image encoder for the at least one computational model.
[0010] In some embodiments of the device, the instructions, when executed by at least one processor, can cause the processor to perform an inference process on multiple real images using a real image encoder and a synthesis decoder to generate a depth image and a confidence map. In various embodiments of the device, the real image encoder may include at least one coordinate convolutional layer.
[0011] In some embodiments of the device, the multiple endoscopic training images may include bronchoscope images. In various embodiments of the device, the multiple endoscopic training images include images generated via bronchoscope imaging through a phantom device.
[0012] In an exemplary embodiment of the device, when executed by at least one processor, the instructions may cause the at least one processor to: provide a patient image as input to a trained computational model, and generate at least one anatomical model corresponding to the patient image. In various embodiments of the device, the instructions, when executed by at least one processor, may cause the at least one processor to generate a depth image and a confidence map for the patient image.
[0013] In some embodiments of the device, the anatomical model may include a three-dimensional point cloud. In various embodiments of the device, instructions, when executed by at least one processor, may cause at least one processor to present the anatomical model on a display device to facilitate navigation of the endoscopic apparatus.
[0014] According to various features of the embodiments described, a computer-implemented method may include, via at least one processor of a computing device: accessing a plurality of endoscopic training images comprising a plurality of synthetic images and a plurality of real images; accessing a plurality of depth real information associated with the plurality of synthetic images; performing supervised training of at least one computational model using the plurality of synthetic images and the plurality of depth real information to generate a synthetic encoder and a synthetic decoder; and performing domain adversarial training on the synthetic encoder using real images to generate a real image encoder for at least one computational model.
[0015] In some embodiments of the method, the method may include performing an inference process on multiple real images using a real image encoder and a synthetic decoder to generate a depth image and a confidence map. In various embodiments of the method, the real image encoder may include at least one coordinate convolutional layer.
[0016] In some embodiments of the method, the multiple endoscopic training images may include bronchoscope images. In various embodiments of the method, the multiple endoscopic training images may include images generated via bronchoscope imaging through a phantom device.
[0017] In an exemplary embodiment of the method, the method may include: providing a patient image as input to a trained computational model, and generating at least one anatomical model corresponding to the patient image. In various embodiments of the method, the method may include generating a depth image and a confidence map of the patient image.
[0018] In some embodiments of the method, the anatomical model may include a three-dimensional point cloud. In various embodiments of the method, the method may include presenting the anatomical model on a display device to facilitate navigation of the endoscopic device. In some embodiments of the method, the method may include examining a patient portion represented by the anatomical model using an endoscopic device.
[0019] According to various features of the embodiments described, an endoscopic imaging system may include an endoscope and a computing device operatively coupled to the endoscope. The computing device may include at least one processor and a memory coupled to the at least one processor. The memory may include instructions that, when executed by the at least one processor, may cause the at least one processor to: access a plurality of endoscopic training images comprising a plurality of synthetic images and a plurality of real images; access a plurality of depth real information associated with the plurality of synthetic images; perform supervised training of at least one computational model using the plurality of synthetic images and the plurality of depth real information to generate a synthetic encoder and a synthetic decoder; and perform domain adversarial training on the synthetic encoder using real images to generate a real image encoder for the at least one computational model.
[0020] In some embodiments of the system, the endoscope may include a bronchoscope.
[0021] In some embodiments of the system, the instructions, when executed by at least one processor, may cause the at least one processor to provide a patient image as input to a trained computational model, the patient image being captured via an endoscope; and to generate at least one anatomical model corresponding to the patient image. In some embodiments of the system, the instructions, when executed by at least one processor, may cause the at least one processor to present the anatomical model on a display device to facilitate navigation of the endoscope within the patient portion represented by the anatomical model. Attached Figure Description
[0022] Specific embodiments of the disclosed machine will now be described by way of example with reference to the accompanying drawings, in which:
[0023] Figure 1 A first exemplary operating environment according to the present invention is shown;
[0024] Figure 2 The computational model training process according to the present invention is illustrated;
[0025] Figure 3 An exemplary synthetic image and corresponding depth image according to the present invention are shown;
[0026] Figure 4 An exemplary real image according to the present invention is shown;
[0027] Figure 5 An exemplary depth image based on a real image input is shown according to the present invention;
[0028] Figure 6 An exemplary depth image and anatomical model based on real image input are shown according to the present invention;
[0029] Figure 7An exemplary depth image based on a real image input is shown according to the present invention;
[0030] Figure 8 An exemplary depth image based on a real image input is shown according to the present invention;
[0031] Figure 9 An exemplary anatomical model according to the present invention is shown;
[0032] Figure 10 An exemplary anatomical model according to the present invention is shown;
[0033] Figure 11 An exemplary anatomical model according to the present invention is shown;
[0034] Figure 12 A second exemplary operating environment according to the present invention is shown; and
[0035] Figure 13 An embodiment of the computing architecture according to the present invention is shown. Detailed Implementation
[0036] This embodiment will now be described more fully below with reference to the accompanying drawings, in which several exemplary embodiments are illustrated. However, the subject matter of the invention can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will be thorough and complete, and are intended to convey the scope of the subject matter to those skilled in the art. In the drawings, the same numerals consistently denote the same elements.
[0037] Various features of diagnostic imaging apparatuses and processes will now be described more fully below with reference to the accompanying drawings, in which one or more features of the diagnostic imaging process will be shown and described. It should be understood that the various features described below, etc., can be used independently or in combination with each other. It should be understood that the diagnostic imaging processes, methods, techniques, apparatuses, systems, components, and / or parts thereof disclosed herein can be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will convey certain features of the diagnostic imaging apparatus and / or processes to those skilled in the art.
[0038] This document discloses a diagnostic imaging process operable to generate an anatomical model from a source image. In some embodiments, the anatomical model may be or may include a three-dimensional (3D) image configured to provide a 3D visualization of a patient's anatomical scene, a graphical user interface (GUI) object, a model, a 3D model, etc. In various embodiments, the source image may be or may include a non-3D source image (e.g., a two-dimensional or monocular image). In one example, the source image may include a monocular image from an endoscope. In various embodiments, the source image may include at least one color monocular image from an endoscope.
[0039] Although endoscopes, and especially bronchoscopes, are used as illustrative diagnostic imaging devices, the embodiments are not limited thereto, as images from any type of image capture system (including other types of diagnostic imaging systems) capable of operating according to some embodiments can be contemplated in the present invention.
[0040] This invention describes monocular endoscopic images as examples of diagnostic images, source images, and / or as the basis for synthetic images used in this invention; however, embodiments are not limited thereto. More specifically, any type of source image (including real or synthetic images) capable of operating in a diagnostic imaging process configured according to some embodiments is contemplated in this invention.
[0041] In various embodiments, the diagnostic imaging process may include a computational model training process operable to train a computational model to generate an anatomical model based on source image input. Illustrative and non-limiting examples of computational models may include machine learning (ML) models, artificial intelligence (AI) models, neural networks (NN), artificial neural networks (ANN), convolutional neural networks (CNN), deep learning (DL) networks, deep neural networks (DNN), recurrent neural networks (RNN), encoder-decoder networks, residual networks (ResNet), U-Net, fully convolutional networks (FCN), combinations thereof, variations thereof, etc. In exemplary embodiments, the computational model training process may include training the computational model using simulated images in a first training process and training the computational model using actual source images in a second training process.
[0042] Depth estimation from monocular diagnostic images is a core task in the localization and 3D reconstruction pipeline of anatomical scenes (e.g., bronchoscopic scenes). Traditional procedures have attempted to utilize various supervised and self-supervised ML- and DL-based approaches using real patient images. However, the lack of labeled data and the scarcity of texture features in endoscopic images (e.g., lungs, colon, intestines, etc.) render these methods ineffective.
[0043] Attempts have been made to register electromagnetic (EM) tracking data captured by bronchoscopy into segmented airway trees from preoperative CT scan data. Besides sensory errors caused by electromagnetic distortion, anatomical deformation is a major challenge for EM-based methods, making their use impractical in live surgical settings. Inspired by successes in natural scenes, vision-based methods have also been proposed. For example, direct and feature-based video CT registration techniques, as well as Simultaneous Localization and Mapping (SLAM) pipelines, have been investigated in various studies. However, feature scarcity, texture, and photometric inconsistencies caused by specular reflections found in anatomical scenes (e.g., internal lung and intestinal cavities) render these techniques insufficient for medical diagnosis, particularly endoscopic surgery.
[0044] The limitations of direct and feature-based methods have led researchers to focus on leveraging depth information to develop direct relationships with scene geometry. With advancements in learning-based techniques, supervised learning has emerged as a method for monocular depth estimation in natural scenes. However, its application to endoscopic tasks is challenging due to the difficulty in obtaining realistic data. An alternative approach is to train networks on synthetic images with their rendered depth information. However, these models tend to experience performance degradation during inference due to the domain gap between real and synthetic images, thus hindering their use in real-world medical diagnostic settings.
[0045] Therefore, some embodiments can provide image processing methods that include alternative domain adaptation methods using a two-step structure, which first trains a depth estimation network in a supervised manner with labeled synthetic images, and then employs an unsupervised adversarial domain feature adaptation process to improve and optimize performance on real patient images.
[0046] Some embodiments can provide diagnostic image processing methods that can be operated to significantly improve the performance of computing devices on real patient images, and can be employed in 3D diagnostic imaging reconstruction pipelines. In various embodiments, for example, depth estimation methods based on deep learning can be used for 3D reconstruction of endoscopic scenes, etc. Due to the lack of labeled data, the computational model can be trained on synthetic images. Various embodiments can provide methods and systems configured to use adversarial domain feature adaptive elements. When applied at the feature level, adversarial domain feature adaptive elements can compensate for the network's low generality on real patient images, etc.
[0047] The devices and methods operating according to some embodiments can provide a variety of technical advantages and features superior to conventional systems. One non-limiting example of a technical advantage may include training a computational model to efficiently generate realistic and accurate 3D anatomical models based on non-3D (e.g., monocular) diagnostic images (e.g., endoscopic images). Another non-limiting example of a technical advantage may include using a monocular imaging device, such as an endoscope, to generate 3D anatomical models without requiring additional hardware, such as additional cameras, sensors, etc. In yet another non-limiting example of a technical advantage, a monocular endoscope may be used in conjunction with a model generated according to some embodiments to navigate within a patient's cavity (e.g., the lung) to efficiently and effectively perform diagnostic tests and sample collection within the patient's cavity without requiring invasive surgery, such as biopsies (e.g., open-chest lung biopsies) or needle aspiration.
[0048] Systems and methods according to some embodiments can be integrated into a variety of practical applications, including diagnosing medical conditions, providing treatment recommendations, performing medical procedures, and delivering treatment to patients. In a particular example, a diagnostic imaging process according to some embodiments can be used to provide minimally invasive bronchoscopy to provide pathological examination of lung tissue for screening for lung cancer. Routine pathological examinations for lung cancer include invasive surgical procedures such as open-chest lung biopsy, transthoracic needle aspiration (TTNA), or transbronchial needle aspiration (TBNA). Existing bronchoscopy cannot efficiently or effectively guide the bronchoscopy through the lungs using monocular images provided by the bronchoscopic camera sensor. However, some embodiments provide a diagnostic imaging process capable of generating 3D anatomical models using images from existing monocular bronchoscopic camera sensors, which healthcare professionals can use to guide the bronchoscopy through the lungs for examination of lung cancer and / or to obtain samples from target areas.
[0049] Some embodiments may include software, hardware, and / or combinations thereof that can be used as part of and / or operatively accessible by a medical diagnostic system or tool. For example, some embodiments may include software, hardware, and / or combinations thereof that can be used as part of an endoscope system, such as a bronchoscopy system to be used during bronchoscopy, and / or operatively accessible by the endoscope system (e.g., elements for providing depth estimation to guide the bronchoscopy system).
[0050] Figure 1 An example of an operational environment 100, which may represent some embodiments, is shown. For example... Figure 1As shown, the operating environment 100 may include a diagnostic imaging system 105. In various embodiments, the diagnostic imaging system 105 may include a computing device 110 communicatively coupled to a network 180 via a transceiver 170. In some embodiments, the computing device 110 may be a server computer, a personal computer (PC), a workstation, and / or other types of computing devices.
[0051] According to some embodiments, the computing device 110 can be configured to manage operational aspects of the diagnostic imaging process, etc. Although Figure 1 Only one computing device 110 is depicted in this drawing, but the embodiments are not limited to this, as the computing device 110 may be, may include, multiple computing platforms, and / or may be distributed across multiple computing platforms. In various embodiments, the functions, operations, configurations, data storage functions, applications, logic, etc., described with respect to the computing device 110 may be executed and / or stored in one or more other computing devices (not shown), for example, connected to the computing device 110 via a network 180 (e.g., one or more client devices 184a to n). A single computing device 110 is depicted for illustrative purposes only to simplify the drawings. The embodiments are not limited to this context.
[0052] Computing device 110 may include processor circuitry 120, which may include and / or have access to various logic for performing processes according to some embodiments. For example, processor circuitry 120 may include and / or have access to diagnostic imaging logic 122. Processing circuitry 120, diagnostic imaging logic 122, and / or portions thereof may be implemented in hardware, software, or a combination thereof. As used herein, the terms “logic,” “component,” “layer,” “system,” “circuit,” “decoder,” “encoder,” “control loop,” and / or “module” refer to computer-related entities that may be hardware, a combination of hardware and software, software, or software in execution, examples of which are provided by exemplary computing architecture 1300. For example, logic, circuitry, or modules can be and / or may include, but are not limited to, processes running on a processor, processors, hard disk drives, multiple storage drives (of optical and / or magnetic storage media), objects, executable files, execution threads, programs, computers, hardware circuits, integrated circuits, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), memory cells, logic gates, registers, semiconductor devices, chipsets, microchips, chipsets, software components, programs, application programs, firmware, software modules, computer code, control loops, computational models or application programs, AI models or application programs, ML models or application programs, DL models or application programs, proportional-integral-derivative (PID) controllers, variations thereof, and any combination thereof.
[0053] although Figure 1 The diagnostic imaging logic 122 is depicted as being within the processor circuitry 120, but the embodiments are not limited thereto. For example, the diagnostic imaging logic 122 and / or any of its components may be located within an accelerator, processor core, interface, separate processor chip, or fully implemented as a software application (e.g., diagnostic imaging application 150), etc.
[0054] The memory cell 130 may include various types of computer-readable storage media and / or systems employing one or more high-speed memory cell forms, such as read-only memory (ROM), random access memory (RAM), dynamic RAM (DRAM), double data rate DRAM (DDRAM), synchronous DRAM (SDRAM), static RAM (SRAM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, polymer memory such as ferroelectric polymer memory, austenite memory, phase change or ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, magnetic cards or optical cards, arrays of devices such as redundant array of independent disk drives (RAID), solid-state memory devices (e.g., USB memory, solid-state drive (SSD), and any other type of storage medium suitable for storing information). In addition, memory unit 130 may include various types of computer-readable storage media in the form of one or more lower-speed memory units, including internal (or external) hard disk drives (HDDs), magnetic floppy disk drives (FDDs), and optical disc drives, solid-state drives (SSDs), etc., for reading from or writing to removable optical discs (e.g., CD-ROMs or DVDs).
[0055] Memory unit 130 may store various types of information and / or applications used in diagnostic imaging processes according to some embodiments. For example, memory unit 130 may store monocular images 132, computational models 134, computational model training information 136, depth images 138, anatomical models 140, and / or diagnostic imaging applications 150. In some embodiments, some or all of monocular images 132, computational models 134, computational model training information 136, depth images 138, anatomical models 140, and / or diagnostic imaging applications 150 may be stored in one or more data stores 182a to n accessible by computing device 110 via network 180.
[0056] Monocular image 132 may include any non-3D image captured via an endoscope, such as endoscope system 160, of a diagnostic tool. In some embodiments, endoscope system 160 may include a bronchoscope. Illustrative and non-limiting examples of endoscope system 160 may include EXALT, supplied by Boston Scientific, Marlborough, Massachusetts, USA. TM Type B disposable bronchoscope. Monocular image 132 may include images of the patient and / or phantom (i.e., a simulated human anatomy device). In some embodiments, endoscopy system 160 may be communicatively coupled to computing device 110 directly or via network 180 and / or client device 184a via wired and / or wireless communication protocols. In various embodiments, computing device 110 may be part of endoscopy system 160, for example, operating as a monitor and / or control device.
[0057] Computational model 134 may include any computational model, algorithm, application, process, etc., used in diagnostic imaging applications according to some embodiments. Illustrative and non-limiting examples of computational models may include machine learning (ML) models, artificial intelligence (AI) models, neural networks (NN), artificial neural networks (ANN), convolutional neural networks (CNN), deep learning (DL) networks, deep neural networks (DNN), recurrent neural networks (RNN), encoder-decoder networks, residual networks (ResNet), U-Net, fully convolutional networks (FCN), combinations thereof, variations thereof, etc.
[0058] In one embodiment, computational model 134 may include a monocular depth image and confidence map estimation model for processing real endoscopic (e.g., bronchoscopy) images. Some embodiments may include, for example, an encoder-decoder model trained on labeled synthetic images. Various embodiments may include computational models trained using domain adversarial training configurations, such as those trained on real images.
[0059] Non-limiting examples of adversarial methods are described in Goodfellow et al.'s "Generative Adversarial Networks" ("Goodfellow et al.") on pages 2672-2680 (2014) in *Advances in Neural Information Processing Systems*, the contents of which are incorporated herein by reference as if fully described herein. For example, one method for reducing domain disparity is based on the adversarial method of Goodfellow et al. Essentially, image processing methods and systems according to some embodiments can operate to optimize a generative model F by developing adversarial signals from a quadratic discriminative network A. During training in each iteration, the generator F attempts to improve by developing a distribution similar to the training set to deceive the discriminator A, while A attempts to guess whether its input is generated by F or sampled from the training set. The input to the generator is from p...z The random noise sampled from the distribution of (z). Putting them together to describe the value function V, the pipeline is similar to a binary minimax game, which was originally defined as follows:
[0060] Where E x ~p data (x) and E z ~p z (z) is the expected value of instances across real data and generated data.
[0061] The application of the above framework to the current problem is called domain adversarial training. In this specific adoption, the generator F acts as a feature extractor and the discriminator A acts as a domain discriminator. The goal of joint training is to assume that the feature extractor F will generate output vectors with the same statistical properties, regardless of the input domain. Depending on the purpose of the application, adversarial loss can be applied at the output level, as in the case of CycleGAN, or at one or more feature levels, as provided in some embodiments.
[0062] In various embodiments, computational model training information 136 may include information for training computational model 134. In some embodiments, training information 136 may include datasets from a synthetic domain and datasets from a non-synthetic or real domain. The synthetic domain may include synthetic images of anatomical regions of interest.
[0063] For example, in bronchoscopy applications, the synthesized image may include a synthesized image of the lung cavity. In one embodiment, the synthesized domain may include multiple rendered color and depth image pairs (see, for example, Figure 3 Multiple composite images may include approximately 1,000 images, approximately 5,000 images, approximately 10,000 images, approximately 20,000 images, approximately 40,000 images, approximately 500,000 images, approximately 100,000 images, and any value or range (including endpoints) between any of the above values.
[0064] In various embodiments, the real domain may include real images captured via an endoscope within a living organism or phantom. For example, for bronchoscopic applications, real images may include real monocular bronchoscopic images, including lung phantoms of human or animal (e.g., dogs) and / or in vivo recordings (see, for example, Figure 4 ).
[0065] In some embodiments, a depth image 138 may be generated using a computational model 136 based on real and / or synthetic monocular images 132. In various embodiments, the depth image 138 may include a corresponding confidence map or may be associated with a corresponding confidence map (see, for example, Figures 5 to 9In applications where depth images cannot be captured by additional sensors, such as particularly in endoscopy and bronchoscopy, according to some embodiments, DL-based methods can be used to estimate or otherwise determine depth images from color images.
[0066] In some embodiments, the anatomical model 140 may be generated based on depth images and / or confidence maps 138. The anatomical model 140 may include point clouds or other 3D representations of anatomical structures captured in source images, such as monocular images 132 (see, for example, ...). Figure 6 and Figures 10 to 12 In various embodiments, the anatomical model 140 can be generated via a 3D reconstruction pipeline configured according to some embodiments.
[0067] In various embodiments, diagnostic imaging logic 122, for example, is operable via a diagnostic imaging application 150 to train a computational model 134 to analyze a patient's monocular images 132 to determine depth images and confidence maps 138 and to generate an anatomical model 140 based on the depth images and confidence maps 138. In some embodiments, the computational model 134 used to determine the depth images and confidence maps 138 may differ from the computational model 134 used to determine the anatomical model 140.
[0068] Diagnostic imaging processes and systems according to some embodiments may or may include a two-step training structure. In one embodiment, the image processing method and system may operate to provide a monocular depth image and confidence map estimation model for processing real endoscopic (e.g., bronchoscopy) images. A first computational model training step may include supervised training of the encoder-decoder model on synthetic images. A second computational model training step may include domain adversarial training, wherein real images are incorporated into the training.
[0069] Figure 2 An illustrative computational model training process according to various features of the present invention is described. In some embodiments, processes 271 to 273 may be or may include part of an image processing pipeline. For example, process 271 may include using a synthetic color image I S and its corresponding depth D S Real-world information is used to supervise the training of the encoder-decoder architecture. In another example, process 272 may include an adversarial training scheme to target the F-value trained in the previous step. S Real Domain I R Image training of a new encoder F R F R You can use F SThe weights are initialized. A is a set of discriminators used at different feature levels of the encoder. In some embodiments, the weights are updated only on process 202 during optimization. In another example, process 273 may include inference of the ground truth domain. R Connected to process 271(G) S The decoder trained in the image is used to estimate the depth image D. R And confidence plot C R .
[0070] In some embodiments, data from the synthetic domain can be used to train the computational model. In one example, the synthetic dataset used in some embodiments may include a large number (e.g., 43,758) of rendering colors. S and depth D S Image pairs. Figure 3 Exemplary synthetic images 310a to d and corresponding depth images 320a to d according to the invention are shown. The lung volume used for data generation can be segmented from one or more computed tomography (CT) scans of a static lung phantom. For segmentation, a computational model 134 can be used to configure the simulator, wherein images are rendered inside the lung tree using a virtual camera with suitable image processing software. Non-limiting examples of computational model 134 may include CNNs, such as 3D CNNs. Illustrative and non-limiting examples of CNNs may include CNNs as described in “3D Convolutional Neural Networks with Graphical Refinement for Airway Segmentation Using Incomplete Data Labels” by Jin et al., pp. 141-149 (2017), in the International Symposium on Machine Learning in Medical Imaging, the contents of which are incorporated herein by reference as if fully set forth herein. Illustrative and non-limiting examples of image processing software may include medical image processing software provided by ImFusion GmbH, Munich, Germany. In some embodiments, the virtual camera may be modeled according to the known inherent properties of a bronchoscope used to acquire realistic images. The simulator may generate ambient lighting using Phong shading and local illumination models, etc.
[0071] In some embodiments, a virtual camera can be positioned equidistantly along an airway segment, for example, starting from the trachea and moving to a target location within the airway tree, thereby simulating images acquired during a typical bronchoscopy procedure. This set of simulated images can be a combination of multiple such simulated paths. The orientation of the virtual camera can be adjusted within reasonable limits to simulate different viewing directions of the camera along the airway segment. The position of the virtual camera within the airway segment can be offset from the centerline of the airway segment to simulate the camera's position during a real bronchoscopy procedure.
[0072] In some embodiments, a real monocular bronchoscopic image I RVarious datasets can be included, such as lung phantoms and in vivo recordings of animal patients (e.g., dogs). However, embodiments are not limited to these types of datasets because any type of real-world image dataset is contemplated in this invention and can be operated according to some embodiments. In one example, the phantom dataset can be recorded inside a lung phantom and corrected using the known lens properties of a bronchoscope. For training and evaluation, the dataset can be split into two distinct subsets, such as a training set of color images without corresponding depth data (e.g., seven video sequences, totaling 12,720 undistorted frames), and a second set of test frames (e.g., a set of 62 frames). The corresponding 3D tracking information can be registered to the centerline of the volume of the airway tree segmented from the phantom CT and used to render synthetic depth images used as real-world information. In one example, in vivo animal frames are recorded within the lung system of a dog patient using an unknown bronchoscope without detailed information about the camera properties. The resulting set of color images (e.g., 11,348 images) can be randomly split into two distinct sets for training and evaluation. Real-world depth images are not available; therefore, this dataset can be used for qualitative analysis. Figure 4 An exemplary real-world image according to the present invention is shown. More specifically, Figure 4 Images of real phantoms 410a to d and animal images 420a to d are depicted.
[0073] In some embodiments, the diagnostic imaging process can use supervised depth image and confidence map estimation. In some embodiments, a U-net variant can be used for the task of supervised depth image and confidence map estimation. For the optimal balance between accuracy and runtime performance, the ResNet-18 network skeleton can be used in the encoder portion of the model to act as a feature extractor. On the decoder side, a series of bilinear upsampling and convolutional layers can be configured to recover the original size of the input. After each upsampling operation, corresponding feature vectors from the encoder level can be concatenated to complete a skip connection structure. The outputs of the last four of these levels form scaled versions of the estimated depth image and confidence map. This output set can be used for multi-scale loss computation.
[0074] In some embodiments, the intermediate activation function may include an exponential linear unit (ELU). A non-limiting example of an ELU is described in “Fast and accurate deep network learning by exponential linear units” by Clevert et al., in the proceedings of the 4th International Conference on Learning Representations in San Juan, Puerto Rico, May 2–4, 2016, which is incorporated herein by reference as fully illustrated herein. In various embodiments, the final activation function may be set to a rectified linear unit (ReLU) and a sigmoid, respectively, for depth image and confidence map estimation. Specific variations in the encoder architecture may include adding coordinate convolutional layers, etc. A non-limiting example of a coordinate convolutional layer is described in “Interesting failures of convolutional neural networks and coordconv solutions” by Liu et al., pp. 9605–9616 (2018), in Advances in Neural Information Processing Systems, which is incorporated herein by reference as fully illustrated herein. In some embodiments, five coordinate convolutional layers are set at skip connections and bottlenecks just before being connected to the decoder or used for adversarial training, discriminator. Detailed configurations of complete models according to some embodiments are provided in Table 1.
[0075]
[0076]
[0077] Table 1
[0078] Typically, Table 1 depicts network architectures for depth image and confidence map estimation according to some embodiments, where k is the kernel size, s is the stride, H is the height, and W is the width of the input image, ↑ is the bilinear upsampling operation, and D... h and C h h∈{1,2,4,8} is the output depth image and confidence map with scaling factor h.
[0079] In various embodiments, a depth estimation loss process can be used. For example, for the estimation of depth values at the original scale as input, the estimated depth loss process can be used. A regression loss is used between the depth image and the ground truth depth image D. The BerHu loss B is used as the pixel-level error.
[0080]
[0081] in, D(i,j) is the predicted true information depth value of pixel index (i,j). The threshold components of c and B are calculated batch-wise as follows:
[0082]
[0083] Where t is an instance of the depth image within the batch.
[0084] A non-restricted example of the BerHu loss B is described in Laina et al.’s “Deeper Depth Prediction with Fully Convolutional Residual Networks”, pages 239–248 of the IEEE 4th International Conference on 3D Vision (3DV) 2016, which is incorporated by reference as fully elaborated in this paper.
[0085] Various implementations can provide scale-invariant gradient loss smoothness, a desired property in the output depth image. To ensure this, a scale-invariant gradient loss L is employed. gradient As follows:
[0086]
[0087] The gradient calculation is performed using the discrete scale-invariant finite difference operator g, which has a step size h, as shown in Equation 4.4:
[0088]
[0089] Some implementations can provide confidence loss. For example, to provide a supervisory signal, the true information confidence map is calculated as follows:
[0090]
[0091] Based on this, the confidence loss is defined as the L1 norm between the prediction and the true information, as follows:
[0092]
[0093] Table 2 below depicts the data augmentation used for the synthetic and real domains. Random values are selected from a uniform distribution, and the color augmentation results in saturation at 0 (minimum) and 1 (maximum).
[0094]
[0095] Table 2
[0096] In the multi-scale total supervision loss process, three factors are combined with four different scale spans to form the total loss:
[0097]
[0098] λ is a hyperparameter that weights each factor, h is the ratio of the predicted to the true depth image size, and u h It is a bilinear upsampling operator that upsamples the input image by a scale h.
[0099] Data augmentation plays a significant role in increasing the size and variability of the training set. Two main criteria should be considered when selecting data augmentations to apply: the function should maintain its geometry and the model should be augmented to prevent overfitting to the domain. Table 3 below describes data augmentations according to some examples:
[0100]
[0101] Table 3
[0102] Various hardware and software configurations can be used to implement and train the network. In one example, the network is implemented on PyTorch 1.5 and training is performed on a single... This was performed on an RTX 2080 graphics card. The supervised training scheme can use the Adam optimizer, for example, which uses a dynamic learning rate schedule to halve the number of training epochs at the midpoint. Further details are provided in Table 3.
[0103] In some embodiments, unsupervised adversarial domain features can be used to, for example, adapt a network previously trained on synthetic rendering to increase its generality on real (bronchoscope) images.
[0104] In some embodiments, the encoder F trained in the synthesis domain according to various embodiments S Used to adversarially train the new encoder F R For this task, three discriminators A were used at the last two jump connections and the bottleneck of the encoder. i , where i is empirically determined to be i∈{1,2,3} to reduce the domain gap at the feature level. During the prioritization period, only F is updated. R The weights. During inference, the new encoder F... R Connect to the previously trained decoder G S For use in depth image and confidence map estimation (see, for example, Figure 2 ).
[0105] Like other neural network models, Generative Adversarial Networks (GANs) have limited learning capabilities. Trained without direct supervision of the task at hand, GANs often inevitably get trapped in local minima, which is not optimal and can be very far from it. Given the small number of semantic and geometric feature differences between the two domains, the new encoder F... R Using and previously trained F SThe same weights are initialized. By doing so, adversarial training is expected to avoid a large number of possible mode collapses on geometrically irrelevant features. Typically, coordinate convolutional layers further improve the robustness of GAN models against mode collapse. In bronchoscopic scenes, deeper regions tend to exist in certain geometries, such as deformed circles and ellipses composed of darker color information. Unlike synthetic rendering, real images have uneven lighting throughout the scene, creating rather blurry dark patches that can be misinterpreted by the model as higher depth values. In some embodiments, employing coordinate convolutional layers to provide the model with supplemental spatial awareness not only reduces the chance of possible mode collapses but also simplifies the adversarial training process and provides guidance to avoid regressing large depth values to the aforementioned blurriness.
[0106] The discriminator used is based on the principles proposed for PatchGAN, as described by Isola et al., “Image-to-Image Transformation Using Conditional Adversarial Networks,” pp. 1125–1134 (2017), Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Goodfellow et al. describe a non-limiting example of a traditional GAN discriminator. Unlike traditional GAN discriminators, models according to some embodiments can generate an output grid instead of a single element. Combined with a fully convolutional structure, this produces a more comprehensive evaluation of local features.
[0107] Each element in the discriminator can employ a GAN loss. It is split into two parts, the Ldiscriminator and the Lencoder, and is employed as follows:
[0108] Ldiscriminator(A, FS, FR, IS, IR) = -EfS~FS(IS)[logA(FS(IS)]
[0109] -EfR~FR(IR)[log(1-A(F R (I R Equation (8),
[0110]
[0111] Among them, I S and I R These are color images from both the synthetic and real domains. The total adversarial loss, Ladversarial, is the sum of the Ldiscriminator and Lencoder across all i, where i∈{1,2,3} is the exponent of the feature tensor: L adversarial (A, F) S F R I S I R )=∑ i∈{1,2,3}(Ldiscriminator(A i F i S,F i R, IS, IR)+
[0112] L encoder (A i F i Equation (10) (R, IR)
[0113] Table 4 below describes the network architecture of a discriminator for adversarial domain feature adaptation according to some embodiments, where k is the kernel size, s is the stride, C is the number of channels, H is the height, and W is the width of the input feature tensor.
[0114]
[0115]
[0116] Table 4
[0117] For adversarial domain feature adaptation, a different set of data augmentations is used for each data source. Regarding the synthetic domain, the augmentations used for supervised training are retained, as detailed in Table 2. For the ground-domain images, color augmentations are skipped to prevent introducing additional complexity into adversarial training.
[0118] Table 5 below describes the details of the adversarial training scheme used. Hyperparameters for other classes and functions not mentioned are set to the library's default values.
[0119]
[0120] Table 5
[0121] The implemented 3D reconstruction pipeline can include a combination of processes derived from subsequent publications such as Park et al.'s "Revisiting Color Point Cloud Registration" (pages 143-152, IEEE International Conference on Computer Vision, 2017) and Choi et al.'s "Robust Reconstruction of Indoor Scenes" (pages 5556-5565, 2015), IEEE Conference on Computer Vision and Pattern Recognition. Essentially, this method employs a pose map formed from feature-based tracking information and color ICP for multi-scale point cloud alignment. It is best to view the pipeline as a series of steps that address the current problem from both local and global perspectives.
[0122] In the first step, the RGB-D sequence is segmented into blocks to construct local geometric surfaces, called segments. This process employs a pose map for each segment for local alignment. The edges of the pose map are formed by the estimated transformation matrix, thereby optimizing the joint photometric and geometric energy functions between adjacent frames of the subsequence. Additionally, loop closure is considered using a 5-point RANSAC algorithm on ORB-based feature matching between keyframes. Finally, the pose map is optimized using a robust nonlinear optimization method, and a point cloud is generated.
[0123] The second step uses a global-scale pose map to register the point clouds. Similar to the previous step, the edges of the pose maps between adjacent nodes are formed by the estimated transformation matrix. For this, photometric and geometric energy functions are used to match the last RGB-D frame of the former with the first RGB-D frame of the latter's point cloud, as shown in the first step. Additionally, a similar approach is used to consider loop closure, employing Fast Point Feature Histogram (FPFH) features of non-adjacent point cloud pairs. Finally, a robust nonlinear method is used to optimize the pose map.
[0124] The third step uses multi-scale point cloud alignment to refine the previously developed global pose map. The aim of this step is to reduce the chance of the optimization getting trapped in local minima by considering a smoother optimization surface for functions at coarser levels. The point cloud pyramid is constructed by increasing the voxel size to downsample the point cloud at each level. Color ICP (Joint Photometric and Geometric Purpose) is used as the optimization function to address alignment along both the normal direction and the tangent plane.
[0125] In the final step, local and global pose maps are combined to assign poses for the RGB-D frames. Ultimately, each of them is integrated into a single truncated symbolic distance function (TSDF) volume to create the final mesh.
[0126] Observations suggest that the fundamental source of the domain gap between synthetic and real images lies in the differences in the lighting and reflective properties of tissues. Furthermore, in both in vivo and in vitro scenes, obstacles arise from mucus and other natural elements, which can adhere to the camera. All these visual features and artifacts are frequently misunderstood by networks trained solely on synthetic images.
[0127] The main focus of this method is to improve the robustness of networks trained on synthetic images to visual variations when running on real data. In these experiments, we quantitatively and qualitatively evaluate the performance of the proposed method. Experiment I: Performance Analysis on the Lung Phantom Dataset
[0128] For this experiment, a two-step approach was used to train the network according to some embodiments. Supervised training was performed for 30 training epochs using the full synthetic domain set (43,758 color and depth image pairs) with the hyperparameters shown in Table 3. The second step, adversarial domain feature adaptation, was performed on 12,720 frames of training segmentation in the phantom scene. Training was conducted for 12,000 iterations, and the hyperparameters used are given in Table 5. Data augmentations applied to the synthetic and real domains are described in Table 2.
[0129]
[0130] Three evaluation metrics are used to consider quantitative analysis: where N is the total number of pixels across all instances t in the test set, and σ is the threshold.
[0131] The test images were a subset of 188 frames with their EM tracking information. For a more accurate evaluation, alignment between the rendered and original images was analyzed by visually assessing overlap at prominent edges. As a result, 62 of the better renders were selected for testing.
[0132] The method is evaluated in Table 6 below, comparing the model before and after applying adversarial domain feature adaptation. For simplicity, the former is named "original" and the latter "domain-adapted". The results show that the adopted adversarial domain adaptation step improves the basic and original models across all metrics. Table 6 describes a quantitative analysis of the impact of the domain adaptation step on depth estimation. Real-world data was rendered from the lung phantom using ImFusion Suite software based on EM tracking signals from bronchoscopy registered to preoperative CT volumes. Sufficient registrations were manually optimized to reduce the test set to 62 pairs of synthetic depth images in true color. Color images were not distorted using known lens properties. Depth values are in millimeters. Optimal values for each metric are indicated in bold.
[0133]
[0134] Table 6
[0135] During the first step of supervised training, it was observed that the learned confidence was lower at deeper locations and at high-frequency components of the image (such as edges and corners). The first property was explained as being caused by fuzziness in darker regions, while the latter was supported by the scale-invariant gradient loss introduced in equation (3). Figure 5 An exemplary depth image based on a synthetic image input is shown in Experiment I according to the present invention. Specifically, Figure 5The input images 505a to d, ground-based depth images 510a to d, raw depth images 515a to d, raw confidence maps 520a to d, domain-adaptive depth images 525a to d, and domain-adaptive confidence maps 530a to d are depicted. (As via...) Figure 5 The original network was observed to have difficulty generalizing to the relatively smooth image characteristics of bronchoscopy. Furthermore, it exhibited poor performance at darker color patches at image edges. This experiment specifically demonstrates that adaptive readjustment of the encoder to suit these two characteristics of real images is achieved through adversarial domain feature adaptation.
[0136] Experiment II: Performance Analysis of the Arterial Patient Dataset
[0137] In this experiment, the training strategy (first step) of the original network was the same as described in Table 3. The second step, adversarial domain feature adaptation, was performed on 9,078 frames of training segmentation of in vivo scenes captured from canine patients. This particular anatomical structure has more bronchi to bifurcate, resulting in finer visual details. In this experiment, 6,000 training iterations were a good balance to fine-tune the domain-specific features while preserving finer details. The remaining hyperparameters are shown in Table 5, and the data augmentations are shown in Table 2.
[0138] Qualitative analysis of the original network performed on 2,270 test frames revealed that it was misled by high spatial frequency features such as blood vessels and masses. Combined with darker textures, these regions showed a tendency to incorrectly regress to larger depth values. However, the domain-adaptive network performed more stably against these deceptive cues. Furthermore, it revealed improvements in capturing the topology around bifurcations, which was significantly more refined in detail compared to synthetic and lung phantom data. Another difference between this particular tissue and those mentioned above was the higher non-Lambertian reflectivity. While the original network was frequently fooled by contours generated by specular reflections and interpreted as having greater depth, the results show that the domain-adaptive step teaches the model to be more robust to them.
[0139] In this evaluation, some of the shortcomings of the adversarial domain feature adaptation method were also revealed. Figure 6 An exemplary depth image based on image input is shown in Experiment II according to the present invention. Specifically, Figure 6 The input images 605a to d, the original depth images 615a to d, the original confidence maps 620a to d, the domain-adaptive depth images 625a to d, the domain-adaptive confidence maps 630a to d, and the computational models 640a to d in the form of point clouds are depicted.
[0140] Experiment III: The Impact of Coordinate Convolution on Adaptive Adversarial Domain Features
[0141] Typically, coordinate convolution can influence the adaptation of adversarial domain features. For example, neural networks may have limited capacity due to their architecture. This property becomes particularly important when it comes to generative adversarial networks (GANs) because unsupervised training schemes may adapt to the properties of the data at different feature levels than intended, leading to mode collapse. Coordinate convolution can provide additional signals about the spatial properties of the feature vectors. In the following experiments, the effect of coordinate convolution on the adaptation of adversarial domain features is evaluated on previously used lung phantom and animal patient datasets. For fairness in the comparison, models were trained with the same hyperparameters introduced in Experiments I and II, respectively.
[0142] Figure 7 Exemplary depth images and confidence maps based on image input, with and without coordinate convolutional layers, according to the present invention are shown. In particular, Figure 7 The following are descriptions: input images 705a to d, ground reality depth images 720a to d, domain adaptation using coordinate convolution depth images 745a to d, domain adaptation using coordinate convolution confidence maps 750a to d, domain adaptation without coordinate convolution depth images 755a to n, and domain adaptation without coordinate convolution confidence maps 760a to d. Figure 7 The results show that, compared to models with coordinate convolutional layers, models without coordinate convolutional layers exhibit decreased robustness to specular reflection and certain high spatial frequency features.
[0143] like Figure 7 As shown, during adversarial training, the model without coordinate convolutional layers is overfitted to a certain pattern, exhibiting a tendency to estimate deeper regions with low confidence at arbitrary locations in the image. Furthermore, Table 7 below quantitatively confirms that the model may experience performance degradation before and after adversarial training without coordinate convolutions.
[0144]
[0145]
[0146] Table 7
[0147] Table 7 typically presents a quantitative analysis of the impact of coordinate convolutional layers on depth estimation. Ground-based data from a lung phantom was rendered using ImFusion Suite software, based on EM tracking signals from bronchoscopy images registered to preoperative CT volumes. Sufficient registrations were manually optimized to reduce the test set to 62 pairs of synthetic depth images in true color. Color images were not distorted using known lens properties. Depth values are in millimeters. Optimal values for each metric are indicated in bold.
[0148] Evaluations of the synthetic animal patient dataset with a large domain gap qualitatively reflect a similar performance degradation.
[0149] Figure 8 Exemplary depth images and confidence maps based on image input, with and without coordinate convolutional layers, according to the present invention are shown. In particular, Figure 8 Domain adaptation is described using input images 805a to d, domain adaptation using coordinate convolution depth images 845a to d, domain adaptation using coordinate convolution confidence maps 850a to d, domain adaptation without coordinate convolution depth images 855a to n, and domain adaptation without coordinate convolution confidence maps 860a to d. Figure 8 The results show that, compared to models with coordinate convolutional layers, models without coordinate convolutional layers exhibit decreased robustness to specular reflection and certain high spatial frequency features.
[0150] Experiment IV: 3D Reconstruction
[0151] In this experiment, two different sequences are used to qualitatively evaluate the adoption of the proposed depth estimation network in the 3D reconstruction pipeline. The model trained for the depth estimation in Experiment I is used to predict depth images.
[0152] Figure 9 An exemplary anatomical model according to the present invention is shown. More specifically, Figure 9 Anatomical models 915a to 915c, generated based on images 905a to 905c and depth images 910a to 915c, are shown. Typically, anatomical models 915a to 915c can be generated based on reconstructions of a short sequence of 55 frames. The reconstructed point cloud 971 is manually overlaid and scaled onto a segmented airway tree 970 using ImFusion Suite software. One method for aligning the reconstructed point cloud onto the segmented airway tree includes using an Iterative Closest Point (ICP) algorithm. The initial, intermediate, and final frames derived from the sequence (1005a to 905c) are shown with their corresponding estimated depth images (1010a to 905c). When pivoting forward, the mirror follows a downward tilting motion from the beginning of the sequence, creating occlusions at deeper locations in the final frame for the superior bifurcation and bronchial base. This results in inaccurate reconstructions of these points.
[0153] Figure 10 An exemplary anatomical model according to the present invention is shown. Figure 10In this example, a sequence of 300 frames was used to reconstruct a 3D anatomical model in the form of point clouds 1015a, 1015b and images 1020a, 1020b overlaid on a segmented airway tree generated based on images 1005a, 1005b and associated depth images 1010a, 1010b. The sequence was initialized at the midpoint of the bronchus, and the bronchoscope was driven to the next bifurcation point and the end. The point clouds were displayed using color information obtained from the input, and the initial, intermediate, and final frames from the sequence were displayed with their corresponding estimated depth images. In another example, Figure 11 An exemplary anatomical model according to the present invention is shown. More specifically, Figure 11 An anatomical model for a sequence of 1,581 frames is shown in the form of a point cloud 1115a generated based on images 1105a to c and associated depth images 1110a to c.
[0154] In some embodiments, the image processing method and system can be operated to perform monocular depth estimation in a bronchoscopic scene, for example, in a two-step deep learning pipeline. In the first step, a U-Net-based model with a ResNet-18 feature extractor as an encoder is used for supervised learning of depth and corresponding confidence information from the rendered synthetic image.
[0155] A non-restricted example of ResNet is provided by He et al. in “Deep Residual Learning for Image Recognition”, pages 770-778 of the 2016 IEEE Conference on Computer Vision and Pattern Recognition, and a non-restricted example of U-net is provided by Ronneberger et al. in “U-net: A Convolutional Network for Biomedical Image Segmentation”, pages 234-241 of the 2015 International Conference on Computational Medical Imaging and Computer-Aided Interventions, published by Springer. Both are incorporated by reference as if fully described in this paper.
[0156] In the second step, the network is refined using a domain adaptation process to improve its generalizability on real images. This part employs adversarial training at multiple feature levels of the encoder to ultimately reduce the domain gap between the two image sources. Training is performed in an unsupervised manner, with the training set for the second step consisting of unpaired frames. A non-restrictive example of the domain self-application process is described in Vankadari et al., “Unsupervised Monocular Depth Estimation of Night Images Using Adversarial Domain Feature Adaptation,” pp. 443–459 (2020), published in the European Conference on Computer Vision, which is incorporated herein by reference as fully elaborated in this paper.
[0157] Training Generative Adversarial Networks (GANs) is often difficult because they tend to get stuck in pattern collapse. To increase resilience against this phenomenon, coordinate convolutional layers are employed on the connection nodes between the encoder and decoder during the training procedure. Since geometrically relevant features in the images are observed to be similar across domains, the new encoder trained in adversarial training is initialized with the weights of the source encoder previously trained on the synthetic images. The discriminator is modeled after PatchGAN for a more comprehensive evaluation of local features. Synthetic images are rendered using ImFusion Suite software based on segmented airway trees from CT scans of a lung phantom. Realistic images are obtained from two sources: lung phantoms and animal patients.
[0158] In this invention, methods and systems according to some embodiments are evaluated in quantitative and qualitative tests on the aforementioned datasets. Furthermore, models according to some embodiments can be integrated into a 3D reconstruction pipeline for feasibility testing. Quantitative analysis of the lung phantom dataset shows that adversarial domain feature adaptation significantly improves the performance of the base model (also known as the original model) across all metrics. The improvements are more visually apparent in the generated results. Compared to the original model, the domain-adaptive model exhibits better depth perception in image features with smoother target domains. Furthermore, results indicate that the domain-adaptive model according to some embodiments performs better on darker patches around image boundaries. Among other things, this suggests that the employed domain-adaptive approach is able to refine the original model to accommodate a portion of cross-domain lighting and sensory variations. On the animal patient dataset, in addition to the differences described above, there are more significant texture variations in the visible domain gap. Visual examination of the results shows that the domain-adaptive model is more robust to deceptive cues such as high spatial frequency components, like blood vessels and masses, and intensity gradients generated by strong specular reflection. Moreover, it more descriptively captures the topology at bifurcation points, exhibiting finer structures and a greater number of bifurcation branches than in synthetic and lung phantom datasets. Overall, the experiment demonstrates that adversarial domain feature adaptation can also refine the model across different anatomical structures. Testing on two datasets confirms the role of coordinate convolutional layers in avoiding pattern collapse and improving network accuracy. The 3D reconstruction pipeline shows improved results compared to conventional methods and systems, demonstrating the ability to integrate it according to several embodiments for localization and reconstruction applications of endoscopic procedures, including bronchoscopy.
[0159] In exemplary embodiments, the quantity and variation in quantitative tests can be increased for use in generating, training, testing, and / or similar image processing methods and systems. While acquiring ground-based depth data from real bronchoscopic sequences is challenging, more realistic rendering methods, such as Siemens VRT technology, can be employed to generate color and depth image pairs for evaluation.
[0160] Some implementations can be configured to use Bayesian methods and heteroscedastic arbitrary uncertainty to improve accuracy and precision in confidence regression.
[0161] Similar to coordinate convolutional layers, camera convolutions incorporate camera parameters into the convolutional layer. Methods and systems according to some embodiments can, among other things, effectively improve the model's versatility for depth prediction across different sensors, achieving higher accuracy and facilitating deployment with bronchoscopes from various brands. Furthermore, combined with the non-Lambertian surface characteristics of lung tissue, the joint motion of the strongly illuminated probe on the bronchoscope disrupts illumination consistency, often leading to instability in depth prediction at a certain location across frames. In some embodiments, the loss function can be integrated with an affine model of light, from which bronchoscopic SLAM and 3D reconstruction applications can benefit.
[0162] A two-step adversarial domain feature adaptation method, according to some embodiments, enables depth image and confidence map estimation in bronchoscopy scenarios. The method has an efficient second step: adapting a network trained on a synthetic dataset under supervised conditions to generalize on real bronchoscopic images in an unsupervised adversarial training manner. Image processing methods and systems according to some embodiments can use domain adaptation schemes operable to improve the base model, thereby adapting to various sources of domain discrepancies, such as illumination, sensory, and anatomical differences. In some embodiments, integrating the methods and systems into a 3D reconstruction pipeline can allow for applications in medical imaging procedures, such as localization and reconstruction during bronchoscopy.
[0163] Figure 12 An example of an operational environment 1200, which may represent some embodiments, is shown. For example... Figure 12 As shown, the operating environment 1200 may include a bronchoscope 1260 having a camera sensor 1261 configured to be inserted into a lung pathway 1251 of a patient 1250. In some embodiments, the camera sensor 1261 may be configured to capture a monocular color image 1232. A computing device 1210 may be configured to execute a diagnostic imaging application 1250 operable to perform a diagnostic imaging process according to some embodiments.
[0164] Monocular color image 1232 from bronchoscopy 1260 can be received at computing device 1210 for processing by diagnostic imaging application 1250. In various embodiments, diagnostic imaging application 1250 may include and / or have access to a computational model trained on the bronchoscopic images according to various embodiments. Diagnostic imaging application 1250 can provide monocular color image 1232 as input to the trained computational model. Depth image and / or confidence map 1238 can be generated by the trained computational model. In some embodiments, the trained computational model can be used as part of a 3D reconstruction pipeline to generate a 3D anatomical model 1240a, such as a point cloud model. In various embodiments, an updated anatomical model 1240b depicting the location of camera sensor 1261 in 3D bronchial scene can be presented on display device 1270 to the surgeon or other medical professional performing bronchoscopy. In this way, medical professionals can have an accurate 3D visualization of the lungs 1251 for navigation and / or examination purposes.
[0165] Figure 13 Embodiments of an exemplary computing architecture 1300 suitable for implementing the various embodiments described above are illustrated. In various embodiments, the computing architecture 1300 may include or be implemented as part of an electronic device. In some embodiments, the computing architecture 1300 may represent, for example, computing devices 110 and / or 1310. The embodiments are not limited to this context.
[0166] As used herein, the terms “system,” “component,” and “module” refer to computer-related entities that can be hardware, a combination of hardware and software, software, or software in execution, examples of which are provided by the exemplary computing architecture 1300. For example, a component can be, but is not limited to, a process running on a processor, a processor, a hard disk drive, multiple storage drives (optical and / or magnetic storage media), an object, an executable file, an execution thread, a program, and / or a computer. For instance, an application running on a server and the server itself can both be a component. One or more components can reside in a process and / or an execution thread, and components can reside on a single computer and / or be distributed across two or more computers. Furthermore, components can communicatively couple with each other to coordinate operation through various types of communication media. Coordination may involve one-way or two-way exchange of information. For example, a component can convey information in the form of signals communicated through a communication medium. This information can be implemented as signals assigned to various signal lines. In such an assignment, each message is a signal. However, alternative embodiments may employ data messages. Such data messages can be sent across various connections. Exemplary connections include parallel interfaces, serial interfaces, and bus interfaces.
[0167] The computing architecture 1300 includes various common computing elements, such as one or more processors, multi-core processors, coprocessors, memory units, chipsets, controllers, peripherals, interfaces, oscillators, timing devices, video cards, audio cards, multimedia input / output (I / O) components, power supplies, etc. However, embodiments are not limited to those implemented by the computing architecture 1300.
[0168] like Figure 13 As shown, the computing architecture 1300 includes a processing unit 1304, a system memory 1306, and a system bus 1308. The processing unit 1304 may be a commercially available processor and may include dual-microprocessor, multi-core processor, and other multiprocessor architectures.
[0169] System bus 1308 provides interfaces for system components, including but not limited to system memory 1306 to processing unit 1304. System bus 1308 can be any of several types of bus architectures, and it can also interconnect to memory buses (with or without memory controllers), peripheral buses, and local buses using any of a variety of commercially available bus architectures. Interface adapters can be connected to system bus 1308 via slot architectures. Example slot architectures can include, but are not limited to, Accelerated Graphics Port (AGP), card bus, (Extended) Industry Standard Architecture ((E)ISA), Micro Channel Architecture (MCA), Network User Bus, Peripheral Component Interconnect (Extended) (PCI(X)), PCI Express, PCMCIA, etc.
[0170] System memory 1306 may include various types of computer-readable storage media employing one or more high-speed memory cell forms, such as read-only memory (ROM), random access memory (RAM), dynamic RAM (DRAM), double data rate DRAM (DDRAM), synchronous DRAM (SDRAM), static RAM (SRAM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, polymer memory such as ferroelectric polymer memory, austenite memory, phase change or ferroelectric memory, silicon-oxide-nitride-oxide-silicon (SONOS) memory, magnetic cards or optical cards, arrays of devices such as redundant arrays of independent disk drives (RAID), solid-state memory devices (e.g., USB storage, solid-state drives (SSDs), and any other type of storage media suitable for storing information). Figure 13 In the illustrated embodiment, system memory 1306 may include non-volatile memory 1310 and / or volatile memory 1312. The basic input / output system (BIOS) may be stored in non-volatile memory 1310.
[0171] Computer 1302 may include various types of computer-readable storage media in the form of one or more low-speed memory cells, including internal (or external) hard disk drive (HDD) 1314, magnetic floppy disk drive (FDD) 1316 for reading from or writing to removable disk 1311, and optical disc drive 1320 for reading from or writing to removable optical disc 1322 (e.g., CD-ROM or DVD). HDD 1314, FDD 1316, and optical disc drive 1320 may be connected to system bus 1308 via HDD interface 1324, FDD interface 1326, and optical disc drive interface 1328, respectively. HDD interface 1324 for external drive implementation may include at least one or both of Universal Serial Bus (USB) and IEEE 1114 interface technologies.
[0172] Drives and associated computer-readable media provide volatile and / or non-volatile storage of data, data structures, computer-executable instructions, etc. For example, multiple program modules may be stored in drive and memory units 1310, 1312, including operating system 1330, one or more application programs 1332, other program modules 1334, and program data 1336. In one embodiment, one or more application programs 1332, other program modules 1334, and program data 1336 may include, for example, various application programs and / or components of computing device 110.
[0173] Users can input commands and information into computer 1302 through one or more wired / wireless input devices (e.g., keyboard 1338 and pointing devices such as mouse 1340). These and other input devices are typically connected to processing unit 1304 via input device interface 1342 coupled to system bus 1308, but may also be connected via other interfaces.
[0174] Monitor 1344 or other types of display devices are also connected to system bus 1308 via an interface, such as video adapter 1346. Monitor 1344 can be internal or external to computer 1302. In addition to monitor 1344, computers typically include other peripheral output devices, such as speakers, printers, etc.
[0175] Computer 1302 can operate in a networked environment that uses logical connections via wired and / or wireless communications to one or more remote computers, such as remote computer 1348. Remote computer 1348 can be a workstation, server computer, router, personal computer, laptop computer, microprocessor-based entertainment device, peer-to-peer device, or other public network node, and typically includes many or all of the elements described relative to computer 1302; however, for brevity, only memory / storage device 1350 is shown. The depicted logical connections include wired / wireless connections to a local area network (LAN) 1352 and / or a larger network, such as a wide area network (WAN) 1354. Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks, such as corporate intranets, all of which can connect to global communication networks, such as the Internet.
[0176] Computer 1302 is operable to communicate with wired and wireless devices or entities using IEEE 802 series standards, such as wireless devices operable in wireless communication (e.g., IEEE 802.16 air modulation technology). This includes at least Wi-Fi (or Wireless Fidelity), WiMax, and Bluetooth™ wireless technologies. Therefore, communication can be a predefined structure like a traditional network, or simply self-organizing communication between at least two devices. Wi-Fi networks use radio technology known as IEEE 802.11x (a, b, g, n, etc.) to provide secure, reliable, and fast wireless connectivity. Wi-Fi networks can be used to interconnect computers, connect to the Internet, and connect to wired networks (using IEEE 802.3 related media and functions).
[0177] This document has set forth numerous specific details to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. In other instances, known operations, components, and circuits have not been described in detail to avoid obscuring the embodiments. It is understood that the specific structural and functional details disclosed herein may be representative and do not necessarily limit the scope of the embodiments.
[0178] Some embodiments may be described using the terms “connection” and “linkage” and their derivatives. These terms are not intended to be synonymous with each other. For example, some embodiments may be described using the terms “connection” and / or “linkage” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “linkage” may also indicate that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0179] Unless otherwise specifically stated, it should be understood that terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data represented as physical quantities (e.g., electronic) within the registers and / or memory of the computing system into other data similarly represented as physical quantities within the memory, registers, or other such information storage, transmission, or display devices of the computing system. Embodiments are not limited to this context.
[0180] It should be noted that the methods described herein need not be performed in the order stated or in any particular order. Furthermore, various activities described with respect to the methods identified herein can be performed serially or in parallel.
[0181] Although specific embodiments have been shown and described herein, it should be understood that any arrangement calculated for achieving the same purpose may replace the specific embodiments shown. The invention is intended to cover any and all adaptive variations or modifications of the various embodiments. It should be understood that the above description is illustrative and not restrictive. After reviewing the above description, combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art. Therefore, the scope of the various embodiments includes any other application in which the above-described components, structures, and methods are used.
[0182] Although the subject matter has been described in language specific to structural features and / or methodological behavior, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or behaviors described above. Rather, the specific features and behaviors described above are disclosed as exemplary forms for implementing the claims.
[0183] As used herein, elements or operations described in the singular and defined with the words “a” or “an” should be understood to not exclude plural elements or operations unless such exclusion is expressly stated. Furthermore, references to “one embodiment” of the invention are not intended to exclude the existence of additional embodiments that also include the described features.
[0184] The scope of this invention is not limited to the specific embodiments described herein. In fact, various embodiments and modifications of the invention, other than those described herein, will be apparent to those skilled in the art from the foregoing description and drawings. Therefore, such other embodiments and modifications are intended to fall within the scope of this invention. Furthermore, although the invention has been described in the context of a specific environment for a particular purpose, those skilled in the art will recognize that its usefulness is not limited thereto, and that the invention can be advantageously practiced in any number of environments for any number of purposes. Therefore, the claims set forth below should be interpreted in accordance with the full breadth and spirit of the invention as described herein.
Claims
1. An electronic device comprising: At least one processor; A memory coupled to the at least one processor, the memory including instructions that, when executed by the at least one processor, cause the at least one processor to: Access multiple endoscopic training images, including multiple synthetic images and multiple real images; Access multiple depth real-world information associated with the multiple synthetic images; Supervised training of at least one computational model is performed using the plurality of synthetic images and the plurality of depth real information to generate a synthetic encoder and a synthetic decoder; The synthetic encoder is subjected to domain adversarial training using the real images to generate a real image encoder for the at least one computational model; The at least one processor performs an inference process on the plurality of real images using the real image encoder and the synthetic decoder to generate a depth image and a confidence map; as well as A patient image is provided as input to a trained computational model to generate at least one anatomical model corresponding to the patient image.
2. The electronic device according to claim 1, wherein the real image encoder comprises at least one coordinate convolutional layer.
3. The electronic device according to claim 1, wherein the plurality of endoscopic training images include bronchoscope images.
4. The electronic device of claim 3, wherein the plurality of endoscopic training images include images generated via bronchoscope imaging through a phantom device.
5. The electronic device of claim 1, wherein, when executed by the at least one processor, the instructions cause the at least one processor to generate a depth image and a confidence map for the patient image.
6. The electronic device of claim 1, wherein, when executed by the at least one processor, the instructions cause the at least one processor to present the anatomical model on a display device to facilitate navigation of the endoscope.
7. A computer-implemented method comprising, via at least one processor of a computing device: Access multiple endoscopic training images, including multiple synthetic images and multiple real images; Access multiple depth real-world information associated with the multiple synthetic images; Supervised training of at least one computational model is performed using the plurality of synthetic images and the plurality of depth real information to generate a synthetic encoder and a synthetic decoder; The synthetic encoder is subjected to domain adversarial training using the real images to generate a real image encoder for the at least one computational model; The real image encoder and the synthetic decoder are used to perform an inference process on the plurality of real images to generate depth images and confidence maps; as well as A patient image is provided as input to a trained computational model to generate at least one anatomical model corresponding to the patient image.
8. The method according to claim 7, wherein the real image encoder comprises at least one coordinate convolutional layer.
9. The method according to claim 7, wherein the plurality of endoscopic training images include bronchoscope images.
10. The method of claim 7, further comprising generating a depth image and a confidence map for the patient image.
11. The method of claim 7, further comprising presenting the anatomical model on a display device to facilitate navigation of the endoscopic device.