Systems and methods for three-dimensional reconstruction of endoscopic and cystoscopic information
Neural Radiance Fields with Instant-NGP and controlled endoscope motion address texture loss and interference in cystoscopy, enabling fast and accurate 3D bladder reconstructions.
Patent Information
- Application Number
- PCT/US2025/016082
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2025-02-14
- Publication Date
- 2025-08-21
AI Technical Summary
Existing 3D reconstruction methods for cystoscopy, such as mesh-based techniques, suffer from texture loss, slow computational speed, and susceptibility to interference, making it difficult to generate accurate and complete 3D models of the bladder.
Utilizing Neural Radiance Fields (NeRF) with Instant-NGP for accelerated computations and controlled endoscope motion to reduce motion blur, combined with sparse Structure from Motion (SfM) for pose estimation, to train a NeRF model for rapid and accurate 3D reconstruction of the bladder interior.
Achieves rapid, accurate, and complete 3D reconstruction of the bladder interior in minutes, reducing computational time by a hundredfold and enhancing resistance to interference, allowing for more reliable and efficient bladder evaluations.
Smart Images

Figure US2025016082_21082025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR THREE-DIMENSIONAL RECONSTRUCTION OF ENDOSCOPIC AND CYSTOSCOPIC INFORMATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 554,439, filed on Feb. 16, 2024, the contents of which are hereby incorporated by reference in their entirety.TECHNICAL FIELD
[0002] The technical field of this disclosure pertains to systems and methods of 3D reconstruction of endoscopic imagery using neural radiance fields. The technical field also pertains to reconstructing images obtained from manual and robot-assisted endoscopies and cystoscopy.BACKGROUND
[0003] Three-dimensional (3D) reconstruction of cystoscopy holds substantial value in the observation and guided treatment of urological conditions. 3D models of the bladder, obtained through such reconstruction, can aid physicians in performing rapid and comprehensive assessment of the condition of the bladder, such as detection of bladder cancer and follow-up surveillance for the life of the patient. However, 3D reconstruction from cystoscopy applications, such as complete bladder 3D reconstruction of the bladder urothelium, have been hampered by issues such as texture loss, slow computational speed, and susceptibility to interference.SUMMARY
[0004] In a first aspect, a method is provided that includes: (i) via an aperture of an endoscope, obtaining a plurality of images of an interior space of a bladder from respective different poses within the bladder; (ii) obtaining pose information about a location and an orientation of the aperture of the endoscope when each image of the plurality of imageswas obtained; (iii) based on the plurality of images and the pose information, training a neural radiance fields (NeRF) model to represent the bladder; and (iv) for a target pose, using the trained NeRF model to render an image of the interior space of the bladder from the target pose.
[0005] In a second aspect, a non-transitory computer readable medium is provided having stored thereon program instructions executable by at least one processor to cause the at least one processor to perform the above method.
[0006] In a third aspect a system is provided that includes: (i) at least one processor; and (ii) a non-transitory computer-readable medium, having stored therein instructions executable by the at least one processor to cause the system to perform the above method.
[0007] These as well as other aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art by reading the following detailed description with reference where appropriate to the accompanying drawings. Further, it should be understood that the description provided in this summary section and elsewhere in this document is intended to illustrate the claimed subject matter by way of example and not by way of limitation.BRIEF DESCRIPTION OF THE FIGURES
[0008] Figure 1 A illustrates aspects of a method for imaging the interior space of a bladder, according to an example embodiment.
[0009] Figure IB illustrates aspects of a method for imaging the interior space of a bladder, according to an example embodiment.
[0010] Figure 1C illustrates aspects of a method for imaging the interior space of a bladder, according to an example embodiment.
[0011] Figure 2A depicts aspects of an experimental apparatus, according to an example embodiment.
[0012] Figure 2B depicts aspects of an experimental apparatus, according to an example embodiment.
[0013] Figure 2C depicts aspects of an experimental apparatus, according to an example embodiment.
[0014] Figure 3A depicts experimental results.
[0015] Figure 3B depicts experimental results.
[0016] Figure 3C depicts experimental results.
[0017] Figure 3D depicts aspects of an experimental method, according to an example embodiment.
[0018] Figure 4A depicts experimental results.
[0019] Figure 4B depicts experimental results.
[0020] Figure 4C depicts experimental results.
[0021] Figure 5A depicts experimental results.
[0022] Figure 5B depicts experimental results.
[0023] Figure 5C depicts experimental results.
[0024] Figure 5D depicts experimental results.
[0025] Figure 6 depicts elements of an example system, according to an example embodiment.
[0026] Figure 7 is a flowchart of an example method.DETAILED DESCRIPTION
[0027] Examples of methods and systems are described herein. It should be understood that the words "exemplary," “example,” and “illustrative,” are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "exemplary," “example,” or “illustrative,” is not necessarily to be construed as preferred or advantageous over other embodiments or features. Further, the exemplary embodiments described herein are not meant to be limiting. It will be readily understood that certain aspects of the disclosed systems and methods can be arranged and combined in a wide variety of different configurations.I. Overview
[0028] Cystoscopy to evaluate the inside of the bladder is indicated for a variety of conditions, including follow-up evaluations following detection and / or treatment of bladder cancer. Typical bladder examination by a urologist takes about 3-6 minutes with the urologist operating a flexible diagnostic cystoscope to manually scan the interior of the bladder. If therapy or other follow-up is required then a larger therapeutic cystoscope maybe used that is typically rigid and requires more sedation. The diagnostic bladder examination of the urothelium (lining of the urethra and bladder) includes looking at the inner surface of both the bladder and urethra. To see backwards toward the opening of the urethra, a maneuver called retroflexion is performed which the scope tip is deflected in a U-turn and, e.g., the semi-flexible shaft of a conventional flexible cystoscope is pressed against the bladder wall to allow for imaging backwards toward the urethra opening. More slender and flexible imaging catheters can make such a U-turn in the bladder without pressing against the distal bladder wall. While such manual evaluation of the interior of the bladder is useful, it is dependent upon the availability of highly trained and experienced healthcare professionals (e.g., urologists) to ensure that the entire interior of the bladder is adequately evaluated. Additionally, it is difficult to quantitatively compare results from different sessions / points in time, since manually-captured images do not necessarily correspond with respect to area, orientation, and perspective.
[0029] It would be beneficial to apply automated or semi-automated techniques to generate, from video / image data acquired using commodity endo / cystoscopies with minimal modification, full 3D reconstructions of the geometry and appearance (e.g., texture) of the interior of the bladder. This would allow such reconstructions to be compared over time, and allow the entirety of the bladder to be evaluated. However, previous attempts at this goal have exhibited significant shortcomings, including low- quality or inaccurate textures or reconstructed geometry, significant computational cost (and thus time to complete, making it difficult to know until after a patient has left whether the acquired image data is sufficient to render a complete reconstruction of the entire internal surface of the bladder), high dependence on multiple views of every region of the interior of the bladder, and other limitations.
[0030] The embodiments described herein achieve, for the first time, reconstruction of the full cystoscopy scene based on cystoscopic video data using Neural Radiance Fields (NeRF). Unlike mesh-based 3D reconstruction techniques, NeRF can restore scenes under conditions where number of views, and even distinguishable features within the available views, are limited. This can address the problem of texture loss that can arise from the use of traditional (e.g., mesh-based) 3D reconstruction algorithms. The implementation of NeRF that was experimentally assessed employed Instant-NGP, anopen-source library that accelerates NeRF computations using hash encoding. These experimental embodiments exhibited a reduction in computational time by approximately lOOx, several tens of times faster than traditional Structure from Motion (SfM)-based methods. Experimental comparison of traditional SfM methods and the NeRF-based embodiments described herein also showed that the NeRF-based embodiments exhibit stronger resistance to interference, underscoring the potential of the NeRF-based methods described herein to enhance applications in endoscopy.
[0031] The NeRF-based embodiments described herein are able to overcome the limitations of prior methods for a variety of reasons. For example, by obtaining the underlying video / image data while the bladder (or other endoscopically-imaged volume) is filled with water or other aqueous solutions, specular reflections at the tissue surface can be reduced (e.g., due to a closer match between the refractive indices of tissue and water when compared to, e.g., the refractive indices of tissue and air or other gaseous volume fill fluids). The presence of such specular reflections at the tissue surface can manifest as a time-varying light emission profile of the volume at the tissue surface, negatively affecting the convergence of NeRF models trained therefrom and negatively affecting the ability of such trained models to accurately represent the color / light emissivity of the tissue surface. This improvement can be obtained by, e.g., introducing an amount of saline or other aqueous fluid into the bladder (or other imaged volume). This could be done via one or more channels of a cystoscope / endoscope, or via another catheter or other device prior to insertion of the scope into the volume. The fluid could be refreshed during the imaging, e.g., continuously and / or at one or more discrete points in time to maintain the volume of the bladder / imaged volume. This could be done, e.g., based on a pressure sensor, a volume / distance sensor, manual feedback, or some other information.
[0032] The embodiments described herein can also overcome the limitations of prior methods by controlling the rate of motion of the imaging aperture at the end of the cystoscope / endoscope in order to reduce or eliminate motion blur in the images / video taken therewith. Such motion blur can significantly degrade the ability of a NeRF trained using imagery that includes motion blur and further can increase the amount of time needed to train such a NeRF (e.g., by reducing the rate of convergence of the trained model). To reduce or eliminate motion blur, the motion of the imaging aperture could be maintainedbelow a set rate (e.g., a set rate of linear motion, a set rate of angular motion) at which blurring occurs (e.g., related to a frame rate, exposure time, lens properties, or other properties of the imaging apparatus). For example, a set rate such that the distance of motion of the imaged tissue within the image is less than half a pixel width from one image to the next for all of the pixels of the image or for a subset of the pixels (e.g., for pixels in a central area of the image). The rate of motion / rotation of the aperture could be kept below such a rate by attaching the endoscope / cystoscope to a linear actuator or other actuator(s) that are operated to advance the scope, and thus the imaging aperture at the end of the scope, at a set, constant rate. Additional or alternative actuators could be used, e.g., control a rate of flexure of the end of the scope, and such actuators could be operated to maintain the linear and / or angular rate of speed of the aperture below a rate at which motion blur can occur.
[0033] In one example, for a camera operating at a 30 Hz frame rate having 1000 pixels spanning a field of view of 90 degrees whose aperture is approximately 15 millimeters from the wall of the bladder, motion blur could be avoided by linearly translating the aperture by less than 1 millimeter per second. In another example, for an endoscopic imaging system with resolution of at least 720x720 pixels and field of view of approximately 75 degrees whose aperture is 10-20 millimeters from the surface of an imaged tissue, a movement rate range between 0.2 mm / second and 1.0 mm / second during lateral tissue inspection can reduce or avoid motion blur artifacts.
[0034] Another manner in which the present embodiments provide for improved modeling and 3D reconstruction of the interior of the bladder (or other internal tissue volume) based on scope video / image data is by obtaining a high-quality initial estimate of the relative (or absolute) location and orientation of the imaging aperture of the endoscope / cystoscope for each frame of video or other image data used to generate the NeRF model. This can include using a sensor (e.g., a 6DOF, 9DOF, magnetic sensor, or other inertial measurement sensor) at or near the aperture to measure the location of the aperture as image data (e.g., frames of video) are captured therethrough. Additionally or alternatively, an encoder or other sensor(s) of a system (e.g., an actuated robotic system) used to control the advance (and optionally flexure or other operations) of the endoscope / cystoscope into the bladder could be operated to determine therefrom thelocation / orientation of the aperture over time. Additionally or alternatively, a known rate of actuation of such an actuator could be used to determine the location / orientation of the aperture over time. For example, if an actuator is operated to advance a cystoscope linearly into a bladder at a set rate (e.g., related to a rate of actuation of a stepper motor of the actuator), the location of the aperture over time during such a steady actuation could be determined as respective evenly-spaced locations along the direction of actuation of the actuator.
[0035] Figs. 1A-C illustrate aspects of such operations. In Fig. 1A, a cystoscope 110 is depicted having a probe tip with an aperture 120 via which the cystoscope can image the interior of a bladder 101 using, e.g., a camera 130. The location and orientation of the aperture 120 at this first point in time is depicted by first view 125a. The cystoscope is then, over time, actuated along a linear trajectory into the bladder; its advancement at a second point in time, such that the aperture 120 can observe the bladder 101 at second view 125b is depicted in Fig. IB. The distance “d” along the linear trajectory can be determined as described above, e.g., using a sensor at or near the aperture 120, using an encoder coupled to the cystoscope 120, and / or based on a pattern of drive applied to the actuator. This distance “d” can then be used to determine the relative or absolute change in the location and orientation of the aperture 120 from the first view 125a to the second view 125b, and this information can then be used to improve (e.g., to significantly speed) the training of a NeRF as described elsewhere herein to represent the entirety of the interior surface of the bladder 101. The tip of the cystoscope 110 could then be operated (e.g., by one or more additional actuators) to change its location and orientation to image the entirety of the interior surface of the bladder 101. An example of the location and orientation of such additional aperture positions is indicated by the additional views of Fig. 1C, which include the first view 125a, second view 125b, and a final view 125 c. The location and orientation of the aperture 120 could be determined from the pattern of drive of such actuators and / or encoders coupled thereto using dynamic models of the flexion of the cystoscope 120 because of such actuation. Additionally or alternatively, flex sensors along the cystoscope, an inertial measurement unit or other sensor(s) disposed at or near the aperture 120, or some other sensors could be used to directly or indirectly measure the location and orientation of the aperture 120 over time.
[0036] As noted above, such improved location and orientation information could be obtained for an actuated and / or manually operated cystoscope or endoscope as part of the embodiments described herein. The user of an endoscope or cystoscope could manually position the scope within the body and then navigate the urinary tract passageways manually or use a robotic interface that remotely performs the operations of scope insertion and retraction by axial translation, rotation of the scope, bending the distal end (tip bending) by mechanically moving the tip bending lever, and navigational maneuvers such as maintaining the distal tip at optimal separation distance from the bladder wall, maintaining a high degree of orthogonality of the scope pose to the bladder wall, and performing retroflexion by making a U-turn and optionally pressing the flexible shaft of the scope against the bladder wall to view the bladder wall around the urethra opening. The robotic interface can include motors that covert electrical signals from the user interface into force generating actions that physically move the scope and tip-bending lever that mimic manual operations or produce smoother and more stable scope motions than typical manual operation (e.g., to maintain the rate of change of the location and / or orientation of the aperture at a level less than a level that would result in motion blur in the obtain images / video data). Motorized manipulations of the scope can include sensors that send electrical signals to the computer that record these physical movements to a high degree of accuracy, such as electro- optical encoders that measure displacement or rotation. Additionally or alternatively, a multi- degree- of-freedom displacement sensor can be placed inside the working channel of the scope to produce a spatial mapping of the scope tip in real-time, such as a 6-degree-of-freedom electro-magnetic tracking sensor.
[0037] Additionally or alternatively, abbreviated SfM techniques can be applied to generate and / or refine estimates of the location and orientation at which each image / video frame was taken. Traditional SfM uses the available image data (e.g., images, video frames) to exhaustively determine a highly dense point cloud based on correspondences between many features in each image / frame. This computationally expensive process results in a sufficiently dense point cloud that can then be used to generate a mesh that represents the geometry of the bladder (or other target volume), onto which texture data from the individual images / frames can be painted. Instead, some of the embodiments described herein apply SfM techniques sparsely, identifying a sufficient number of featuresin the images / frames to determine the relative locations / orientations of the images / frames but not to generate a sufficient number of features / points to fully render the geometry of the volume. By doing so, the computational cost and time to determine the relative location / orientation of the images / frames (which is then used to train the NeRF) is substantially decreased relative to traditional SfM (where the additional features / points are needed to generate the complete mesh of the geometry of the volume). The sparse method could be adaptive, e.g., determining an initial set of features from a set of images and then using detected correspondences between features in different images to generate an initial estimate of the relative locations / orientations of the images. If the confidence / error between a pair of images is not high enough, additional features, and correspondences therebetween, could be determined from the pair of images and used to iteratively refine the estimated relative locations / orientations of the pair of images until the confidence / error is sufficiently improved.
[0038] For some or all of the images / frames used to train a NeRF model, an initial estimate of the relative or absolute location / orientation of the images / frames (e.g., of the aperture 120 when the image / frame was taken) determined from an encoder, inertial measurement unit, or other sensor could be used to ‘bootstrap’ the SfM or other method used to further refine the location / orientation. For example, an estimate of the location / orientation of the aperture 120 determined using a dynamic model of flexion of the cystoscope 110 by one or more actuators could be used as an initial estimate for images / frames taken while flexing the cystoscope 110, with SfM or other techniques used to further refine the initial estimate. In some example, additional refinement technique(s) could be applied to further refine an SfM or otherwise-derived estimate of the location / orientation of an image / frame. For example, bundling or other techniques could be applied to further refine an estimated orientation / location while also training the NeRF.
[0039] Once a NeRF has been trained using the available images / video frame depicting the inside of the bladder (or other target volume), it can be used to render a plurality of images of the inside of the bladder from various different perspectives. Such a process can include, for a target perspective, repeatedly inferencing the trained NeRF at a plurality of points along a plurality of different rays from the target perspective. A pixel of the rendered image could then be determined based on an integration of the inferencedoutputs along one or more of the rays. Such a plurality of images could be used to, e.g., reconstruct the view of the entirety of the interior surface of a bladder or other target volume. This could include.
[0040] As noted above, the embodiments described herein can allow a NeRF to be trained to represent the entirety of the interior surface of a bladder or other tissue volume using reduced computational resources (e.g., processor cycles, power, memory) and / or in less time than alternative methods (e.g., SfM-based methods that include generating a mesh from a highly dense point cloud and then painting image-based textures thereon). Such reductions in cost / latency are benefits in their own right. Additionally, the significant reductions in latency made possible by the embodiments described herein allow a complete NeRF-based reconstruction of the interior of the bladder (or other target tissue volume) to be generated in minutes or less, thereby allowing a healthcare technician to verify, before the end of an imaging session, whether the entirety of the interior surface of the bladder (or other target) has been completely imaged sufficient to allow the it to be reconstructed. Thus, if any portion of the interior surface has not been sufficiently imaged, the technician can operate the cystoscope / endoscope (e.g., manually, or by operating the controls or a robotic or otherwise actuated imaging system) to acquire additional image data of the nonrepresented surface(s), allowing the NeRF to be updated, e.g., by re-training the NeRF on a set of images and poses that includes the original set of images / poses as well as the additional images / poses acquired subsequently. This reduces the need for follow-on sessions to re-image the target, and can also allow for more direct comparison over time, since complete image data of the entire interior of the bladder can be generated more easily and more consistently within each imaging session.
[0041] Determining whether the entirety of the interior surface of the bladder (or other target) has been completely imaged could include using the trained NeRF to generate a 3D reconstruction of the bladder and then detecting whether the 3D reconstruction fully encloses a single volume. If not (i.e., if some area of the partially-enclosed volume is not enclosed by the 3D reconstruction, or is enclosed by a portion of the 3D reconstruction that exhibits insufficient textural fidelity or resolution), the computer system used to generate the reconstruction could provide an indication to a technician to operate the endoscope / cystoscope to further image the bladder. Such an indication could include adisplay of the results of the 3D reconstruction, an indication of what area of the bladder to further image, or some other indication that the technician can use to accurately direct the view of the scope toward the area(s) in need of additional image data. Additionally or alternatively, a NeRF-based 3D reconstruction of the interior surface of the bladder could be generated and displayed to a healthcare technician, and the healthcare technician could then, via visual inspection, determine whether additional image data is needed (e.g., due to portions of the reconstruction being missing, or exhibiting insufficient texture or other features).
[0042] The methods described herein may be carried out on any subject that may benefit from bladder imaging. In one embodiment, the subject is a living mammalian subject. In another embodiment, the subject is a living human subject. In a further embodiment, the subject is one that is at risk of bladder cancer, has bladder cancer, and / or is being treated for bladder cancer. Based on bladder imaging data generated via the methods herein, a diagnosis, treatment, or other course of action could be determined for and provided to the subject. This could be done based on of analysis of imaging data generated, via the methods described herein, at multiple points in time, as the more complete and high-quality image data generated thereby makes it possible to perform more accurate and complete longitudinal comparisons over time, wherein the bladder is a bladder of a living human subject. The living human subject could have been diagnosed with or suspected of having bladder cancer, and the methods described herein could be performed as part of, an initial screening for bladder cancer, a follow-up evaluation for bladder cancer, an evaluation of the progression of bladder cancer, and / or an evaluation of the effect of a treatment for bladder cancer. Based image data generated as described herein of the interior space of the bladder, a treatment could be provided to the human or animal subject or a treatment provided to the human or animal subject could be modified (e.g., to adjust a dose of a chemotherapy or other drug, to change a type of chemotherapy or other drug provided to the subject).IL Experimental Results
[0043] The embodiments described herein were implemented and experimentally validated. The particulars and results of this experimental validation are provided in this section. Note that the particulars described in this section are intended only as illustrative,non-limiting examples of the embodiments described herein.
[0044] 3D reconstruction of cystoscopy holds substantial value in the observation and guided treatment of urological conditions. 3D models of the bladder, obtained through such reconstruction, can aid physicians in performing a rapid and comprehensive assessment of the condition of the bladder, such as detection of bladder cancer and followup surveillance over the life of the patient. Prior methods for 3D reconstruction based on cystoscopy are hampered by various shortcomings, including as texture loss, slow computational speed, and susceptibility to interference. By implementing the embodiments described herein, a reconstruction of the dynamic cystoscopy scene based on Neural Radiance Fields (NeRF) was achieved for the first time. In contrast with mesh-based 3D reconstruction, NeRF can restore scenes under conditions where the number of views and features are more limited, addressing the problem of texture loss that can arise from alternative 3D reconstruction algorithms. The embodiments described herein, which in some examples include the use of Instant-NGP to accelerate NeRF computations using hash encoding, can reduce computation time by a hundredfold, making the methods described herein several tens of times faster than prior Structure from Motion (SfM) techniques. This is accomplished, in part, by providing high-quality camera pose values into the NeRF reconstruction. The NeRF-based methods described herein also exhibit stronger resistance to interference than prior SfM methods, demonstrating the potential of the NeRF=based methods herein in the field of endoscopy.
[0045] Endoscopy videos inherently include a large amount of information, but the complexity of manual post-session video reviews often results in an oversimplified and condensed form of data, reduced to a manual selection of images and succinct notes or illustrations of potential lesions and scarring. This restricted information inhibits in-depth and ongoing studies into cancer physiology, disease progression or reoccurrence, and constrains the influence these data could have on clinical decisions. The impact of complete organ reconstructions has been profound in other imaging modalities, contributing to clinical breakthroughs in fields like intravenous injection, sleep apnea assessment, and cervical and gastric cancer. Various methods can be applied to accomplish 3D reconstruction based on monocular endoscopic (e.g., cystoscopic) image data. A first approach employs Structure from Motion (SfM), which uses feature points and featurematching, two-views, and multiview geometry to reconstruct structures. A second approach uses deep learning for depth estimation, with a depth map for each image being generated through depth estimation, forming a pseudo-RGB-D image. This pseudo-RGB- D image is then used for 3D reconstruction.
[0046] However, depth estimation-based methods can struggle with inadequate generalizability. SfM-based methods also exhibit several limitations, including long processing periods, occurrences of texture missing in areas where feature points or views are limited, and a deficiency in anti-interference capability.
[0047] To compensate for these shortcomings, the embodiments described herein apply Neural Radiance Fields (NeRF) to reconstruct 3D scenes from cystoscopic or other endoscopic image data. NeRF is an inverse rendering method that is independent of feature detection and is capable of recovering photorealistic 3D scenes by inputting the camera’s pose and processing it through a Multilayer Perceptron (MLP). Previous attempts to apply NeRF to endoscopic video data indicated that this method is not well-suited for dynamic endoscopy, primarily due to the variable light field in dynamic endoscopy, which contravenes NeRF’s static light assumption. For example, one previous attempt concluded that their results “suggest that NeRF may not work well for 3D reconstruction of most if not all real-world captured endoscopy images.”
[0048] The embodiments described herein overcome these limitations. By employing these embodiments, NeRF was demonstrated to function properly in underwater conditions within an endoscopic modality, in particular cystoscopy, which is conducted in a fully submerged liquid environment. An experimental comparison was made between the 3D reconstruction of bladder phantoms using SfM and the NeRF-based embodiments described herein by using the same set of initial pose positions derived from SfM. The results depicted herein demonstrate the potential of more rapid, accurate, and complete 3D reconstructions of the bladder, facilitating remote cystoscopy for telemedical invasive procedures.
[0049] The prior state of the art method for cystoscopic 3D reconstruction was a Structure from Motion technique designed for white-light cystoscopy. In brief, this method mainly consisted of the following four steps:
[0050] Image preprocessing: A specific number of frames were chosen as“keyframes” from the video and rectified using a calibrated camera model. A color correction algorithm was used to process each keyframe twice, creating unique input images for subsequent structure-from-motion and texture-reconstruction stages.
[0051] Structure-from-motion: Appropriate keyframes were identified and used to detect points of interest, with corresponding feature descriptors being matched across the images. An initial sparse point cloud was created, representing a set of 3D points on the bladder surface, alongside the computation of camera poses to depict the location and alignment of the cystoscope in each keyframe.
[0052] Mesh generation: A comprehensive bladder surface was created using the 3D point cloud derived from the preceding step. This bladder surface was depicted via a triangular mesh. Such a mesh or other representation of the bladder surface could be generated via a variety of processes, including but not limited to marching cubes or the Poisson method.
[0053] Texture reconstruction: Texture images, camera poses, and the triangular mesh were used to project a surface texture that consisted of selected regions from the input images onto the triangular mesh. This process bestowed the 3D reconstruction with a realistic visual representation of the bladder wall. The quality of this output texture is heavily dependent on the prior steps of image preprocessing.
[0054] NeRF is a 3D scene generation technique that can synthesize novel views from input 2D images. It includes a fully-connected neural network and utilizes a rendering loss function to reconstruct input views of a scene, making it a potent tool for synthetic image generation. NeRF acts in some examples to interpolate between input images, thereby rendering a comprehensive scene.
[0055] NeRF generally operates by training a network to predict, from a five-dimensional input that includes spatial location and viewing direction, a four-dimensional output that includes color and opacity at that location and direction. The trained NeRF model can then be applied to generate novel views volume rendering. NeRF employs a multilayer perceptron (MLP), a type of artificial neural network with multiple layers of nodes, to learn and model complex, non-linear relationships in the set of input data to map a 3D point xk in space and view direction d to density c k and color ck, e.g., as
[0057] Given a pixel in an image, its color can be computed by integrating the color of the points sampled along its visual ray r through volume rendering, e.g., as
[0058]
[0059] where 8k = tk+1 - tk is the distance between adjacent sampled points. Typically, a NeRF is fitted to a scene by minimizing the reconstruction error between the rendered image C and the input image I, e.g., as
[0060]
[0061] Instant-NGP (Neural Graphics Primitives) is an improvement to NeRF computations that uses hash table encoding to allow smaller networks to be used with negligible compromise on quality, leading to a substantial reduction in floating point computations and memory access operations. Leveraging the parallel computing capabilities of modem GPUs, Instant-NGP can be deployed on fully-fused Compute Unified Device Architecture (CUD A) kernels, with a dedicated focus on minimizing inefficient computation operations and bandwidth consumption. As a result, it achieves a significant acceleration, enabling the training of high-quality neural graphics primitives in mere seconds and facilitating the rendering at a resolution of 1920x1080 in tens of milliseconds.
[0062] In the experimental study described herein, two experiments were conducted. The first one compared the performance of the NeRF-based embodiments described herein with traditional SfM algorithms in regions of low contrast and sparse views. The second experiment compared reconstruction results obtained in a realistic phantom post-injection of water. Two phantoms were used, both derived from the same subject’s CT scan. One was created using 3D printing technology to produce an outer hard shell that could be split into two halves, and onto each inner surface images from real cystoscopy were affixed (shown in Fig. 2A). The other was a silicone monolithic phantom, exhibiting a more realistic material quality, shown in Fig. 2B. The experiments and comparative analysis of NeRF and traditional SfM’s were performed on both phantoms.
[0063] To ensure the optimal performance of both methodologies, it was confirmed that the captured videos exhibited high coherence and stability. Additionally, efforts were made to mitigate the effects of motion blur. Hence, a robotic system was used for video capture, which is demonstrated in Fig. 2C.
[0064] Figure 1A-C: Different phantoms and robot setup. (A) 3D printed phantom, adorned with textures printed from genuine cystoscopy images. (B) Silicone monolithic phantom, the internal texture of which is depicted in Figs. 4A-C. (c) The custom robot-assisted flexible cystoscopy platform which was used to navigate within a phantom filled with water.
[0065] The overall structure of the robot included a Karl Storz flexible cystoscope as the main imaging component, which was connected to a 3D printed fixture and transmission mechanism with advancement controlled by a stepper motor. This robot could be remotely operated via a wired or wireless handpiece, and exhibited excellent stability in the workingenvironment of a cystoscope. Operating this robot to gather cystoscope data facilitated increased coherence between consecutive frames in the video while minimizing motion blur.
[0066] In the NeRF workflow, the camera pose is an input, making the acquisition of the camera pose one of the initial steps of the experiment. The same camera poses were used in both the SfM and NeRF methods so that the camera pose wouldn’t be a variable that influenced the reconstruction result. All experimental camera poses originated from the sparse reconstruction results of the SfM method under test.
[0067] The camera poses were optimized by bundle adjustment, which is a non-linear optimization technique to refine the estimated 3D point locations, camera positions, and orientations. The objective was to minimize the re-projection error, which measures the discrepancy between 3D points re-projected into the camera’s view and the corresponding 2D interest points observed in the image, according to
[0069] In the equation above, the indices i and j represent the ith camera and the jth 3D point, respectively. The term xj1refers to the 2D image point that corresponds to the jth 3D point as viewed by the ith camera, and Q is the set of linear correspondences. The function II : R 3 R 2 denotes the perspective projection operation, with Ri and ti being the rotation matrix (orientation) and translation vector (position) of the ith camera, respectively. The optimal adjustment of the rotation matrices Ri is typically achieved through multiplicative updates, Ri := AiRi, where incremental rotations Ai are computed on the tangent plane to the manifold of the special orthogonal group via the exponential map. For improved robustness, it is common to replace the L2-norm with a more robust cost function such as the Huber function.
[0070] The scarcity of captured views during endoscopy is predominantly attributed to the obstruction of the endoscope’s field of view by certain anatomical structures around the urethra orifice, as viewed during retroflexion. Thus, the endoscope was placed into a retroflexed position for the experiments that assessed performance under sparse view conditions. Compared to the silicone monolithic phantom, the split design of the 3D printed phantom permitted easier tracking of the camera trajectory during cystoscope retroflexion. The opacity of the monolithic phantom made it challenging to ascertain whether the endoscope’s shaft had come into contact with the bladder wall during retroflexion. Hence, the 3D printed phantom was used to evaluate the performance under conditions of sparse views.
[0071] The cystoscope was inserted into the 3D printed phantom and retroflexed back toward the ureteric orifice. This approach ensured that the bladder structure near the urethraorifice would cause a certain degree of obstruction in the captured video, effectively simulating the issues of field-of-view obstruction commonly encountered in clinical scenarios. Subsequently, an endoscopic robot was employed to sweep over the entire phantom at a constant velocity. The resulting videos were processed through both the traditional SfM algorithm and the NeRF-based algorithms described herein. The outcomes obtained from each of these procedures are shown in Figs. 3A-C.
[0072] Figure 3: Robustness test results. (A) 3D scene generated via NeRF. (B) 3D reconstructed mesh generated by SfM. (C) Photograph of the opened half of the 3D printed phantom with yellow highlight indicating location of obscured region. (D) Diagram of the retroflexed flexible cystoscope viewing back toward the urethral orifice with a region obscured from the highlighted visual field.
[0073] Fig. 3A is the 3D scene generated by NeRF. This scene presents a view following the retroflexion of the cystoscope, with the catheter visibly extending into the scene in the NeRF output, Fig. 3B is the result created by SfM. Fig. 3D illustrates how the protruding section of the bladder wall impedes and obscures the view of the cystoscope. The reconstruction of these obscured portions can be achieved by capturing images and features from different angles, however, protrusion from the ureter significantly reduced the number of clear views for the obstructed and obscured region. This type of obstruction can cause dropouts in the SfM method, as shown in Fig. 2B. It can be seen from Fig. 3A that the results produced by the NeRF-based methods described herein are more complete compared to those by the SfM method. The latter fails to effectively reconstruct the field of view obscured by the protrusion of the bladder wall by the urethral orifice, the location of which is highlighted in Fig. 3C in yellow.
[0074] As demonstrated by the results in Figs. 3a-C, NeRF can generate a comprehensive 3D reconstruction even when the number of views is limited. This robustness stems from the fact that NeRF’s generation process does not rely on feature extraction and matching. It circumvents the need for triangulation and multi -view geometry, which typically require multiple images of any particular region of the bladder to generate usable 3D information. Instead, NeRF directly predicts the 3D scene through an MLP. As a result, as long as the scene appears in the captured images, a 3D scene can be generated, even with a limited number of views or sparse feature points, given the appropriate viewing angle. Thus, NeRF can achieve good results even in regions with low features, such as walls.
[0075] In the realistic phantom experiment, injected water was injected into themonolithic silicone phantom, after which the endoscopic robot was maneuvered at a constant velocity to scan the entirety of the phantom dome. The silicone phantom’s interior contained some floating and buoyant high-contrast particles and aggregates, which for endoscopic imaging can become artifacts within the 3D bladder surface reconstruction, shown in Fig. 4A. However, given that the presence of particles or debris in the bladder is also encountered during clinical cystoscopy, these floating particles were note removed. After the acquisition of the video, it was processed via both the traditional SfM algorithm and the NeRF-based algorithms described herein. The results obtained from each of these procedures are detailed in Figs. 4A- C.
[0076] Figure 4: (A) An example of floating particles, in an acquired video frame acquired from the cystoscope video. (B) 3D reconstructed mesh generated by traditional SfM. (C) 3D surface reconstruction using NeRF of the bladder wall dome of the silicone phantom via the embodiments described herein, which is the same section as shown in (B) but without artifacts
[0077] It is evident that the phantom reconstructed via the SfM method represents the floating particles in the water as features of the bladder wall. SfM algorithms commonly misinterpret such suspended particles as features on the bladder wall, leading to their inclusion in the reconstructed bladder surface, shown in Fig. 4B. This issue arises due to the nature of the mesh representing the three-dimensional surface, since texture mapping involves projecting images onto this surface in a specific direction. Consequently, most alternative 3D reconstruction algorithms lack the capability to capture more intricate spatial structures, such as vertical layering, on the mesh. Moreover, this characteristic makes it difficult to accurately represent the realistic uneven and irregular surface of the bladder. In Fig. 4C, NeRF did not misinterpret the floating particles, which can be attributed to the fact that NeRF renders the entire 3D space and the complete light field, and is thus capable of representing structures with depth. In contrast, the SfM method first generates a mesh, transforms the corresponding images into textures, and subsequently overlays these textures onto the surface of the mesh. Consequently, if there are interfering factors (such as these high-contrast particles) present within the images that are to be textured, these factors will also appear on the textures of the final mesh.
[0078] The computational resources utilized in these experiments included an NVidia GeForce RTX TITAN GPU with 24GB of GDDR6 memory, an AMD 3800x processor, and 32GB of system RAM. Under this setup, the NeRF and SfM bladder 3D reconstructions werecompared to the bladder phantom by an experienced urologist (MPP). Qualitatively the NeRF reconstruction provided increased contrast to color and topography compared to SfM, while also insuring more complete coverage. These are useful qualities of a digital record of a remote bladder examination by a remote urologist or using Al computer-aided diagnostic tools. Although the spatial resolution was not measured, the NeRF reconstruction did not show any noticeable reduction in resolution or distortion, while the SfM reconstruction was susceptible to dropouts and artifacts. To assess the increase in computational speed, five videos were reconstructed using the NeRF approach that required an average of 20 seconds for 3D reconstruction compared to 8 - 40 minutes (200 - 500 frames) for the SfM pipeline, a 24x - 120x difference. An example of rapid reconstruction using NeRF based on instant-NGP is presented in Figs 5A-D. Photorealistic scene reconstruction was accomplished in just 20 seconds
[0079] Figure 5: The rapid generation of photorealistic 3D scenes using Instant-NGP based NeRF from t = 0s to t = 20s
[0080] The original cystoscopic videos in this preliminary study did not cover the entire bladder phantom. In addition, the full pipeline of the NeRF requires the camera pose to be known. Since this can be done in real-time using optical or electromagnetic sensing, the reported time for the NeRF algorithm does not include the input of cystoscopic camera poses. Endeavoring to provide a more objective comparison of the reconstruction results in this comparative work, we only utilized camera poses simulated through the SfM pipeline in this study.
[0081] The NeRF -based reconstruction techniques described herein were successfully implemented experimentally within the application of cystoscopy, achieving high-quality reconstruction results and effectively overcoming the limitations associated with the traditional SfM method such as slow speed, frequent dropouts, and susceptibility to interference. Through this research, it was confirmed that NeRF can be employed in cystoscopy using realistic phantoms within a liquid environment that mimics clinical bladder examination.III. Example Systems
[0082] Figure 6 illustrates an example system 600 that may be used to implement the methods described herein. By way of example and without limitation, computing system 600 may be an imaging system (e.g., an imaging system that includes and / or is adapted to interface with a cystoscope or other variety of endoscope and / or to receive image or other informationelectronically, optically, or otherwise therefrom), a computer (such as a desktop, notebook, tablet, or handheld computer, a server), elements of a cloud computing system, or some other type of device. It should be understood that computing system 600 may represent a physical computing device such as a server, a particular physical hardware platform on which an application operates in software, or other combinations of hardware and software that are configured to carry out image processing, model training, feature detection and alignment, image rendering or reconstruction, or other functions as described herein. The computing system 600 could be a central system (e.g., a server, elements of a cloud computing system) that is configured to receive image data (e.g., individual images, frames of a video) and / or other data (actuator, encoder, and / or sensor data relating to the absolute and / or relative location and / or orientation of an aperture of an endoscope or other imaging system when images of the image data were acquired) from a remote system and to responsively transmit, to that remote system or to some other system, the results of model training, inference, rendering, reconstruction, or other computations based on the image and other data. Additionally or alternatively, the computing system 600 could be such a remote system, configured to transmit image or other data to a central system, receive output models, image data, 3D models, or other information in response, and / or to take some other actions as described herein.
[0083] As shown in Figure 6, computing system 600 may include a communication interface 602, a user interface 604, a processor 606, a camera 605, an actuator / encoder 603, an endoscope 620 (e.g., a cystoscope), and data storage 608, all of which may be communicatively linked together by a system bus, network, or other connection mechanism 610.
[0084] Communication interface 602 may function to allow computing system 600 to communicate, using analog or digital modulation of electric, magnetic, electromagnetic, optical, or other signals, with other devices, access networks, and / or transport networks. Thus, communication interface 602 may facilitate circuit-switched and / or packet-switched communication, such as Internet protocol (IP) or other packetized communication. For instance, communication interface 602 may include a chipset and antenna arranged for wireless communication with a radio access network or an access point. Also, communication interface 602 may take the form of or include a wireline interface, such as an Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) port. Communication interface 602 may also take the form of or include a wireless interface, such as a Wifi, BEUETOOTH®, global positioning system (GPS), or wide-area wireless interface (e.g., WiMAX or 3GPP Long- Term Evolution (LTE)). However, other forms of physical layer interfaces and other types ofstandard or proprietary communication protocols may be used over communication interface 602. Furthermore, communication interface 602 may comprise multiple physical communication interfaces (e.g., a Wifi interface, a BLUETOOTH® interface, and a wide-area wireless interface).
[0085] In some embodiments, communication interface 602 may function to allow computing system 600 to communicate with other devices, remote servers, access networks, and / or transport networks.
[0086] User interface 604 may function to allow computing system 600 to interact with a user or other entity, for example to receive input from and / or to provide output to the user. Thus, user interface 604 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, and so on. User interface 604 may also include one or more output components such as a display screen which, for example, may be combined with a presence-sensitive panel. The display screen may be based on CRT, ECD, and / or LED technologies, or other technologies now known or later developed. User interface 604 may also be configured to generate audible output(s), via a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface 604 may be configured to receive user inputs relating to the operation of the endoscope 620 (e.g., to advance / retract the endoscope along an axis of the endoscope, to flex an end of the endoscope in one or more directions and / or at one or more locations along the endoscope) and such inputs may then be used to operate the actuator 603 to implement the inputs with respect to the endoscope.
[0087] Processor 606 may comprise one or more general purpose processors - e.g., microprocessors - and / or one or more special purpose processors - e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, tensor processing units (TPUs), or application-specific integrated circuits (ASICs). Data storage 608 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated in whole or in part with processor 606. Data storage 608 may include removable and / or non-removable components.
[0088] Processor 606 may be capable of executing program instructions 618 (e.g., compiled or non-compiled program logic and / or machine code) stored in data storage 608 to carry out the various functions described herein. Therefore, data storage 608 may include a non-transitory computer-readable medium, having stored thereon program instructions that, upon executionby computing system 600, cause computing system 600 to carry out any of the methods, processes, or functions disclosed in this specification and / or the accompanying drawings. The execution of program instructions 618 by processor 606 may result in processor 606 using data that is, e.g., organized in a database 612.
[0089] By way of example, program instructions 618 may include an operating system 622 (e.g., an operating system kernel, device driver(s), and / or other modules) and one or more application programs 620 (e.g., functions for performing one or more of the methods described herein) installed on computing system 600. Database 612 may include image data (e.g., individual images and / or frames of video data obtained via the camera 605) 614 and / or NeRF model(s) 616 that may be trained therefrom or obtained in some other manner.
[0090] Application programs 620 may communicate with operating system 622 through one or more application programming interfaces (APIs). These APIs may facilitate, for instance, application programs 620 transmitting or receiving information via communication interface 602, receiving and / or displaying information on user interface 604, and so on.
[0091] Application programs 620 may take the form of “apps” that could be downloadable to computing system 600 through one or more online application stores or application markets (via, e.g., the communication interface 602). However, application programs can also be installed on computing system 600 in other ways, such as via a web browser or through a physical interface (e.g., a USB port) of the computing system 600.IV. Example Methods
[0092] Figure 7 is a flowchart of a method 700 as described herein. The method 700 includes, via an aperture of an endoscope, obtaining a plurality of images of an interior space of a bladder from respective different poses within the bladder (710). The method 700 additionally includes obtaining pose information about a location and an orientation of the aperture of the endoscope when each image of the plurality of images was obtained (720). The method 700 additionally includes, based on the plurality of images and the pose information, training a neural radiance fields (NeRF) model to represent the bladder (730). The method 700 yet further includes, for a target pose, using the trained NeRF model to render an image of the interior space of the bladder from the target poses (740). The method 700 could include additional or alternative steps or features.VI. Conclusion
[0093] The above detailed description describes various features and functions of thedisclosed systems, devices, and methods with reference to the accompanying figures. In the figures, similar symbols typically identify similar components, unless the context indicates otherwise. The illustrative embodiments described in the detailed description, figures, and claims are not meant to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0094] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the figures and as discussed herein, each step, block and / or communication may represent a processing of information and / or a transmission of information in accordance with example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, functions described as steps, blocks, transmissions, communications, requests, responses, and / or messages may be executed out of order from that shown or discussed, including in substantially concurrent or in reverse order, depending on the functionality involved. Further, more or fewer steps, blocks and / or functions may be used with any of the message flow diagrams, scenarios, and flow charts discussed herein, and these message flow diagrams, scenarios, and flow charts may be combined with one another, in part or in whole.
[0095] A step or block that represents a processing of information may correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a step or block that represents a processing of information may correspond to a module, a segment, or a portion of program code (including related data). The program code may include one or more instructions executable by a processor for implementing specific logical functions or actions in the method or technique. The program code and / or related data may be stored on any type of computer-readable medium, such as a storage device, including a disk drive, a hard drive, or other storage media.
[0096] The computer-readable medium may also include non-transitory computer- readable media such as computer-readable media that stores data for short periods of time like register memory, processor cache, and / or random access memory (RAM). The computer- readable media may also include non-transitory computer-readable media that stores program code and / or data for longer periods of time, such as secondary or persistent long term storage, like read only memory (ROM), optical or magnetic disks, and / or compact-disc read onlymemory (CD-ROM), for example. The computer-readable media may also be any other volatile or non-volatile storage systems. A computer-readable medium may be considered a computer- readable storage medium, for example, or a tangible storage device.
[0097] Moreover, a step or block that represents one or more information transmissions may correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.
[0098] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.VII. Enumerated Example Embodiments
[0099] Embodiments of the present disclosure may thus relate to one of the enumerated example embodiments (EEEs) listed below. It will be appreciated that features indicated with respect to one EEE can be combined with other EEEs.
[0100] EEE 1 is a method including: (i) via an aperture of an endoscope, obtaining a plurality of images of an interior space of a bladder from respective different poses within the bladder; (ii) obtaining pose information about a location and an orientation of the aperture of the endoscope when each image of the plurality of images was obtained; (iii) based on the plurality of images and the pose information, training a neural radiance fields (NeRF) model to represent the bladder; and (iv) for a target pose, using the trained NeRF model to render an image of the interior space of the bladder from the target pose.
[0101] EEE 2 is the method of EEE 1, wherein obtaining the plurality of images is performed while the bladder is full of liquid.
[0102] EEE 3 is the method of EEE 2, further comprising, while obtaining the plurality of images, introducing liquid into the bladder to maintain a volume of the bladder as fluid exits from the bladder.
[0103] EEE 4 is the method of any preceding EEE, further comprising, while obtaining at least a portion of the plurality of images, operating an actuator to advance the endoscope into the bladder such that the aperture progresses along a straight trajectory.
[0104] EEE 5 is the method of EEE 4, wherein operating the actuator to advance the endoscope comprises advancing the endoscope at a speed that is lower than a blur speed at which the at least a portion of the plurality of images would exhibit motion blur.
[0105] EEE 6 is the method of any of EEEs 4 or 5, wherein obtaining pose information about the location and the orientation of the aperture of the endoscope when each image of the plurality of images was obtained comprises at least one of (i) operating the actuator to advance the endoscope at a specified speed such that a distance between respective poses of subsequent images of the plurality of images are related to the specified speed, or (ii) operating an encoder of the actuator to detect a distance between respective poses of subsequent images of the plurality of images.
[0106] EEE 7 is the method of EEE 6, further comprising, subsequent to advancing the endoscope into the bladder, operating an additional actuator to flex an end of the endoscope, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained additionally comprises determining respective poses of images of the plurality of images obtained while flexing the end of the endoscope using a model the flexion of the end of the endoscope as a result of operation of the additional actuator.
[0107] EEE 8 is the method of any preceding EEE, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained comprises operating a sensor to detect the location and orientation of the aperture.
[0108] EEE 9 is the method of any preceding EEE, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained comprises applying a sparse structure from motion (SfM) method to determine the pose information based on the plurality of images.
[0109] EEE 10 is the method of EEE 9, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained additionally comprises obtaining initial estimates of the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained, and wherein applying the sparse SfM method comprises using the plurality of images to refine the initial estimates.
[0110] EEE 11 is the method of EEE 10, wherein wherein obtaining the initial estimates comprises operating an actuator to advance the endoscope at a specified speed such that a distance between respective estimated locations of subsequent images of the plurality of images are related to the specified speed.
[0111] EEE 12 is the method of EEE 10, wherein obtaining the initial estimatescomprises operating an encoder of an actuator to detect a distance between respective locations of subsequent images of the plurality of images while the actuator advances the endoscope into the bladder.
[0112] EEE 13 is the method of EEE 10, wherein obtaining the initial estimates comprises operating a sensor to detect the location and orientation of the aperture.
[0113] EEE 14 is the method of any preceding EEE, further comprising using the trained NeRF model to render a plurality of additional images of the interior space of the bladder from respective different poses to render an entire interior surface of the interior space of the bladder.
[0114] EEE 15 is the method of EEE 14, further comprising: (i) based on the rendered plurality of additional images, determining that the rendering of the interior surface of the interior space of the bladder is incomplete; and (ii) responsively: (a) obtaining additional images of the interior space of the bladder from respective additional poses within the bladder, and (b) based on the plurality of images, the additional images, the pose information, and additional pose information about the location and orientation of the aperture of the endoscope when each of the additional images was obtained, training an updated neural NeRF model to represent the bladder.
[0115] EEE 16 is the method of EEE 15, wherein determining that the rendering of the interior surface of the interior space of the bladder is incomplete comprises generating, from the rendered plurality of additional images, a textured model of the bladder and determining that the textured model is incomplete.
[0116] EEE 17 is the method of EEE 15, wherein determining that the rendering of the interior surface of the interior space of the bladder is incomplete comprises displaying, to a user, the rendering of the interior surface of the interior space of the bladder and receiving, from the user, a user input that indicates that the rendering of the interior surface of the interior space of the bladder is incomplete.
[0117] EEE 18 is the method of any preceding EEE, wherein training the NeRF model to represent the bladder and using the trained NeRF model to render the image of the interior surface of the interior space of the bladder takes less than 40 seconds using fewer than 600 processor cores.
[0118] EEE 19 is the method of any preceding EEE, wherein the endoscope is a cystoscope.
[0119] EEE 20 is the method of any preceding EEE, wherein the bladder is a bladderof a living human subject.
[0120] EEE 21 is the method of EEE 20, wherein the living human subject has been diagnosed with or suspected of having bladder cancer, and wherein the method is performed as part of at least one of an initial screening for bladder cancer, a follow-up evaluation for bladder cancer, an evaluation of the progression of bladder cancer, or an evaluation of the effect of a treatment for bladder cancer.
[0121] EEE 22 is the method of EEE 21, wherein the method further comprises, based on the image of the interior space of the bladder from the target pose, providing a treatment to the living human subject or modifying a treatment provided to the living human subject.
[0122] EEE 23 is a non-transitory computer readable medium having stored thereon program instructions executable by at least one processor to cause the at least one processor to perform the method of any preceding EEE.
[0123] EEE 24 is a system including: (i) at least one processor; and (ii) a non-transitory computer-readable medium, having stored therein instructions executable by the at least one processor to cause the system to perform the method of any of EEEs 1-22.
[0124] EEE 25 is the system of EEE 24, further comprising an endoscope.
[0125] EEE 26 is the system of EEE 25, further comprising an actuator configured to advance an aperture of the endoscope along a straight trajectory.
[0126] EEE 27 is the system of EEE 26, further comprising an additional actuator configured to flex an end of the endoscope.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: via an aperture of an endoscope, obtaining a plurality of images of an interior space of a bladder from respective different poses within the bladder; obtaining pose information about a location and an orientation of the aperture of the endoscope when each image of the plurality of images was obtained; based on the plurality of images and the pose information, training a neural radiance fields (NeRF) model to represent the bladder; and for a target pose, using the trained NeRF model to render an image of the interior space of the bladder from the target pose.
2. The method of claim 1, wherein obtaining the plurality of images is performed while the bladder is full of liquid.
3. The method of claim 2, further comprising, while obtaining the plurality of images, introducing liquid into the bladder to maintain a volume of the bladder as fluid exits from the bladder.
4. The method of claim 1, further comprising, while obtaining at least a portion of the plurality of images, operating an actuator to advance the endoscope into the bladder such that the aperture progresses along a straight trajectory.
5. The method of claim 4, wherein operating the actuator to advance the endoscope comprises advancing the endoscope at a speed that is lower than a blur speed at which the at least a portion of the plurality of images would exhibit motion blur.
6. The method of claim 4, wherein obtaining pose information about the location and the orientation of the aperture of the endoscope when each image of the plurality of imageswas obtained comprises at least one of (i) operating the actuator to advance the endoscope at a specified speed such that a distance between respective poses of subsequent images of the plurality of images are related to the specified speed, or (ii) operating an encoder of the actuator to detect a distance between respective poses of subsequent images of the plurality of images.
7. The method of claim 6, further comprising, subsequent to advancing the endoscope into the bladder, operating an additional actuator to flex an end of the endoscope, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained additionally comprises determining respective poses of images of the plurality of images obtained while flexing the end of the endoscope using a model the flexion of the end of the endoscope as a result of operation of the additional actuator.
8. The method of any of claims 1-7, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained comprises operating a sensor to detect the location and orientation of the aperture.
9. The method of any of claims 1-7, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained comprises applying a sparse structure from motion (SfM) method to determine the pose information based on the plurality of images.
10. The method of claim 9, wherein obtaining pose information about the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained additionally comprises obtaining initial estimates of the location and orientation of the aperture of the endoscope when each image of the plurality of images was obtained, and wherein applying the sparse SfM method comprises using the plurality of images to refine the initial estimates.
11. The method of claim 10, wherein obtaining the initial estimates comprises operating an actuator to advance the endoscope at a specified speed such that a distance between respective estimated locations of subsequent images of the plurality of images are related to the specified speed.
12. The method of claim 10, wherein obtaining the initial estimates comprises operating an encoder of an actuator to detect a distance between respective locations of subsequent images of the plurality of images while the actuator advances the endoscope into the bladder.
13. The method of claim 10, wherein obtaining the initial estimates comprises operating a sensor to detect the location and orientation of the aperture.
14. The method of any of claims 1-7, further comprising using the trained NeRF model to render a plurality of additional images of the interior space of the bladder from respective different poses to render an entire interior surface of the interior space of the bladder.
15. The method of claim 14, further comprising: based on the rendered plurality of additional images, determining that the rendering of the interior surface of the interior space of the bladder is incomplete; and responsively: obtaining additional images of the interior space of the bladder from respective additional poses within the bladder, and based on the plurality of images, the additional images, the pose information, and additional pose information about the location and orientation of the aperture of the endoscope when each of the additional images was obtained, training an updated neural NeRF model to represent the bladder.
16. The method of claim 15, wherein determining that the rendering of the interior surface of the interior space of the bladder is incomplete comprises generating, from the rendered plurality of additional images, a textured model of the bladder and determining thatthe textured model is incomplete.
17. The method of claim 15, wherein determining that the rendering of the interior surface of the interior space of the bladder is incomplete comprises displaying, to a user, the rendering of the interior surface of the interior space of the bladder and receiving, from the user, a user input that indicates that the rendering of the interior surface of the interior space of the bladder is incomplete.
18. The method of any of claims 1 -7, wherein training the NeRF model to represent the bladder and using the trained NeRF model to render the image of the interior surface of the interior space of the bladder takes less than 40 seconds using fewer than 600 processor cores.
19. The method of any of claims 1-7, wherein the endoscope is a cystoscope.
20. The method of any of claims 1-7, wherein the bladder is a bladder of a living human subject.
21. The method of claim 20, wherein the living human subject has been diagnosed with or suspected of having bladder cancer, and wherein the method is performed as part of at least one of an initial screening for bladder cancer, a follow-up evaluation for bladder cancer, an evaluation of the progression of bladder cancer, or an evaluation of the effect of a treatment for bladder cancer.
22. The method of claim 21, wherein the method further comprises, based on the image of the interior space of the bladder from the target pose, providing a treatment to the living human subject or modifying a treatment provided to the living human subject.
23. A non-transitory computer readable medium having stored thereon program instructions executable by at least one processor to cause the at least one processor to perform the method of any preceding claim.
24. A system comprising: at least one processor; anda non-transitory computer-readable medium, having stored therein instructions executable by the at least one processor to cause the system to perform the method of any of claims 1-22.
25. The system of claim 24, further comprising an endoscope.
26. The system of claim 25, further comprising an actuator configured to advance an aperture of the endoscope along a straight trajectory.
27. The system of claim 26, further comprising an additional actuator configured to flex an end of the endoscope.
Citation Information
Patent Citations
Repositioning and reorientation of master / slave relationship in minimally invasive telesuregery
US20060241414A1
Image-based feedback endoscopy system
US20150208904A1
3D Reconstruction and Registration of Endoscopic Data
US20170046833A1
Systems and methods for localization based on machine learning
US20200297444A1