Fusing deep learning with geometric constraints for image-based localization, computer-implemented method, program, and computer-implemented system

By integrating deep learning with geometric constraints, the method effectively addresses the challenges of image-based localization in endoscopic systems, achieving improved accuracy and precision in tracking and localization within gastrointestinal and bronchial tracts.

JP7673392B2Active Publication Date: 2025-05-09FUJIFILM BUSINESS INNOVATION CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2020203545
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2020-12-08
Publication Date
2025-05-09
Estimated Expiration
2040-12-08

AI Technical Summary

Technical Problem

Existing image-based localization techniques for endoscopic systems in gastrointestinal and bronchial tracts face challenges due to insufficient annotated training data, lack of discernible textures, and difficulties in depth perception, leading to inaccurate tracking and localization.

Method used

A computer-implemented method that combines deep learning with geometric constraints to improve image-based localization. This involves applying a training image to a neural network, performing classification, extracting features, and using bundle adjustments to minimize reprojection errors, thereby determining the optimal pose of the test image.

Benefits of technology

The proposed solution achieves more accurate localization and tracking of endoscopic systems within the gastrointestinal and bronchial tracts, even with limited training data and challenging environments, thereby enhancing the precision of therapeutic interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673392000003
    Figure 0007673392000003
  • Figure 0007673392000004
    Figure 0007673392000004
  • Figure 0007673392000005
    Figure 0007673392000005
Patent Text Reader

Abstract

To obtain an accurate location of a test image at an optimal pose within an environment divided into zones.SOLUTION: A computer-implemented method comprises: applying training images of an environment divided into zones to a neural network and performing classification to label a test image on the basis of the closest zone of the zones; extracting a feature from retrieved training images and pose information of the test image that match the closest zone; performing bundle adjustment on the extracted feature by triangulating map points for the closest zone to generate a reprojection error and minimizing the reprojection error to determine an optimal pose of the test image; and providing an output indicating a location or probability of the location of the test image at the optimal pose within the environment for the optimal pose.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Aspects of example implementations relate to methods, systems, and user experiences associated with image-based localization in an environment, and more particularly, to an approach that fuses deep learning with geometric constraints for image-based localization. [Background technology]

[0002] Related art endoscopic systems can provide a minimally invasive method for inspecting internal body structures. More specifically, related art minimally invasive surgical (MIS) approaches can provide physicians with tools to inspect internal body structures that can be used for precise therapeutic intervention.

[0003] For example, a scope, such as an endoscope or bronchoscope, can be placed in a patient's environment, such as the intestines or lungs, to inspect the structures. Devices on the scope, such as sensors or cameras, can sense information and provide the information to a user, such as an image, video, or other image of the environment. A medical professional, such as a surgeon, may analyze the video. Based on the analysis, the surgeon can provide recommendations or perform a procedure.

[0004] Using related art robotics and sensor technology, various related art gastrointestinal (GI) tract scope solutions have been developed. In such related art GI tract approaches, precise localization and tracking enable physicians to locate and track the progression of various pathological findings such as polyps, cancerous tissue, lesions, etc. Such related art endoscopic systems can address the need for precision intervention and therefore must be able to precisely localize and track in a given gastrointestinal (GI) and / or bronchial tract.

[0005] Related art approaches to tracking the GI tract may include comparing image similarity, such as using related art image descriptors to compare image similarity, also referred to as image classification. Additionally, related art techniques may use geometry-based pose regression, such as related art geometric techniques such as SLAM or shape from shading for image-to-model registration, also referred to as geometric optimization. Related art techniques may also use image based deep learning for pose regression.

[0006] The deep learning approach of the related art has various problems and shortcomings specific to tracking for applications such as colonoscopy or bronchoscopy, such as small annotated training datasets and lack of distinguishable textures that are different from other indoor or outdoor environments in which deep learning is used in the related art. For example, there are no corner points that define textures, and the nature of body tissues is that they have blood flow, smooth curves and tubular structures, and no corner points, so there are similar corners of volumetric surfaces, and mixtures of solids and liquids.

[0007] For example, but not by way of limitation, related art deep learning and regression approaches suffer from the problems of having insufficient data sets and lack of corners and textures as described above. In these aspects, surgical scenarios are distinct and distinguishable from related art approaches used in other environments, such as autonomous driving. For example, the GI tract and / or bronchial tubes of the lungs have unique physiological characteristics, resulting in many tube-like structures without corners.

[0008] Furthermore, as related art approaches to deep learning and regression attempt to locate the scope, additional problems and / or shortcomings may arise. For example, but not by way of limitation, another problem is associated with outliers that are located completely outside the environment due to a lack of sufficient quality and quantity of datasets for training. The results of such outliers are of critical importance in the medical field, where if a scope is determined to be completely outside the environment, such as the lungs and GI tract, it may be difficult for medical personnel to rely on that information to perform appropriate analysis and treatment.

[0009] Related approaches to localization in the GI tract use monocular imaging with related art computer vision techniques (e.g., SIFT and SURF). However, such related art approaches may have various problems and shortcomings, such as deformation, intensity, and various obstructions. For example, related art systems may have a lack of depth perception or poor localization within the limited field of view provided by related art RGB / monocular images. For example, the field of view of the scope is narrow due to the close proximity of soft tissue within the patient's body environment.

[0010] Related art depth / stereo-based scope positioning systems cannot be directly adapted to monocular endoscopes because no 3D depth information is provided and the only data available is RGB video.

[0011] Furthermore, there is an unmet need to use large amounts of data to generalize deep learning-based localization and tracking. Such data is difficult to obtain, especially in the medical field, due to privacy concerns. Furthermore, the geometry-based methods of the related art are not applicable to GI tract scope tracking because the number of features is minimal and registration may be lost. Also, it is neither practical nor healthy to increase the number of datasets by actively continuing to insert a scope into the patient.

[0012] Thus, physicians may have difficulty determining the location of the scope in a body environment, such as the location of the endoscope in the GI tract. This problem is exacerbated in certain tissues, such as the lungs, due to the branching physiology of the lungs. [Prior art documents] [Non-patent literature]

[0013] [Non-Patent Document 1] ARANDJELOVIC, R., et al., NetVLAD; CNN Architecture for Weakly Supervised Place Recognition, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 5297-5307. [Non-Patent Document 2] DEGUCHI, D., et al., Selective Image Similarity Measure for Bronchoscope Tracking Based on Image Registration, Medical Image Analysis, 2009, 13, pp. 621-633. [Non-Patent Document 3] DlMAS, G., et al., Visual Localization of Wireless Capsule Endoscope Aided by Artificial Neural Networks, 2017 IEEE 30th International Symposium on Computer-Based Medical Systems, 2017, pp. 734-738. [Non-Patent Document 4] ENGEL, J., et al., Direct Sparse Odometry, IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(3), March 2018, pp. 611-625. [Non-Patent Document 5] HE, K., et al., Deep Residual Learning for Image Recognition, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [Non-Patent Document 6] KENDALL, A.., et al., PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization, Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 2938-2946. [Non-Patent Document 7] LUO, X., et al., Development and Comparison of New Hybrid Motion Tracking for Bronchoscopic Navigation, Medical image Analysis, 2012, 16, pp. 577-596. [Non-Patent Document 8] LUO, X., et al., A Discriminative Structural Similarity Measure and its Application to Video-Volume Registration for Endoscope Three-Dimensional Motion Tracking, IEEE Transactions on Medical Imaging 33(6), June 2014, pp. 1248-1261. [Non-Patent Document 9] MAHMOUD, N., et al., ORBSLAM-based Endoscope Tracking and 3D Reconstruction, arXiv: 1608.08149 [cs.CV], August 29, 2016,13 pgs. [Non-Patent Document 10] MERRITT, S. A., et al., Interactive CT-Video Registration for the Continuous Guidance of Bronchoscopy, IEEE TransMed Imaging August 2013, 32(8), pp. 1376-1396. [Non-Patent Document 11] DEGUCHI, D., et al., New Image Similarity Measure for Bronchoscope Tracking Based on Image Registration, Medical imaging 2004; Physiology, Function, and Structure from Medical Images, 5369, 2004, pp. 165-176. [Non-Patent Document 12] MUR-ARTAL, R., et al., ORB-SLAM2: an Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras, IEEE Transactions on Robotics, July 19, 2017, 33(5), pp. 1255-1262. [Non-Patent Document 13] PATEL, M., et al., ContextualNet: Exploiting Contextual Information Using LSTMs to Improve Image-Based Localization, 2018 IEEE International Conference on Robotics and Automation (ICRA), May 21-25, 2018, Brisbane, Australia, pp. 5890-5896. [Non-Patent Document 14] SGANGA, J., et al., OffsetNet Deep Learning for Localization in the Lung Using Rendered Images, arXiv: 1809.05645 [cs.CV], September 15, 2018, 7 pgs. [Non-Patent Document 15] SHEN, M., et al., Robust Camera Localization with Depth Reconstruction for Bronchoscopic Navigation, International Journal of Computer Assisted Radiology and Surgery, 10(6), 2015, 16 pgs. [Non-Patent Document 16] TURAN, M., et al., Deep EndoVO: A Recurrent Convolutional Neural Network (RCNN) Based Visual Odometry Approach for Endoscopic Capsule Robots, Neurocomputing, 275, 2018, pp. 1861-1870. [Non-Patent Document 17] WEYAND, T., et al., PlaNet - Photo Geolocation with Convolutional Neural Networks, European Conference on Computer Vision, February 17, 2016, 10 pgs. Summary of the Invention [Problem to be solved by the invention]

[0014] The disclosed technique aims to provide a computer-implemented method, program, and computer-implemented system capable of obtaining an accurate position of a test image in an optimal pose within an environment divided into zones. [Means for solving the problem]

[0015] According to a first aspect of an exemplary implementation, a computer-implemented method is provided that includes applying training images of an environment divided into zones to a neural network and performing classification to label a test image based on a closest one of the zones; extracting features from pose information of the retrieved training images and a test image that matches the closest zone; performing bundle adjustment on the extracted features by triangulating map points of the closest zone to generate a reprojection error, minimizing the reprojection error to determine an optimal pose for the test image; and providing an output indicative of a location or a probability of a location of the test image in the optimal pose in the environment, relative to the optimal pose. In a second aspect, in the first aspect, applying the training images includes receiving the training images associated with poses within a zone of the environment as historical data or simulation data, and providing the received training images to a neural network. In a third aspect, in the second aspect, the neural network is a deep learning neural network that learns zones associated with the pose and determines the closest zone for the test image. A fourth aspect is the first aspect, wherein the bundle adjustment includes reprojecting 3D points associated with the measured pose and the triangulated map points into 2D image space to generate a result, and comparing the result to a registered 2D observation to determine the reprojection error. In a fifth aspect, the pose of the test image is confirmed to be the optimal pose if a reprojection error is below a threshold. In a sixth aspect, in the fourth aspect, if a reprojection error exceeds a threshold, the pose of the test image is determined to be incorrect, and the calculation of the pose of the test image is determined to be correct. A seventh aspect is the first aspect, wherein minimizing the reprojection error includes adjusting the pose of the test images to minimize the reprojection error. The program of the eighth aspect causes a processor to perform processes including applying training images of an environment divided into zones to a neural network and performing classification to label a test image based on the closest one of the zones; extracting features from the retrieved training images and pose information of the test image that matches the closest zone; performing bundle adjustment on the extracted features by triangulating map points of the closest zone to generate a reprojection error and minimizing the reprojection error to determine an optimal pose for the test image; and, for the optimal pose, providing an output indicative of the position or probability of the test image's position in the optimal pose within the environment. A ninth aspect is the eighth aspect, wherein applying the training images includes receiving the training images associated with poses within a zone of the environment as historical data or simulation data, and providing the received training images to a neural network. A tenth aspect is the ninth aspect, wherein the neural network is a deep learning neural network that learns zones associated with the pose and determines the closest zone for the test image. An eleventh aspect is the eighth aspect, wherein the bundle adjustment includes reprojecting 3D points associated with the measured pose and the triangulated map points into 2D image space to generate a result, and comparing the result to a registered 2D observation to determine the reprojection error. In a twelfth aspect, in the eleventh aspect, the pose of the test image is confirmed to be the optimal pose if a reprojection error is below a threshold. A thirteenth aspect is the eleventh aspect, wherein if a reprojection error exceeds a threshold, the pose of the test image is determined to be incorrect, and the calculation of the pose of the test image is determined to be correct. A fourteenth aspect of the present invention relates to the eighth aspect, and further relates to the minimizing the reprojection error including adjusting the pose of the test images to minimize the reprojection error. A fifteenth aspect is a computer-implemented system for locating and tracking a scope within an environment to identify targets, configured to: apply training images of the environment associated with the scope and divided into zones to a neural network; perform classification to label a test image generated by the scope based on a closest one of the zones of the environment associated with the scope; extract features from the retrieved training images and pose information of the test image that matches the closest zone; perform bundle adjustment on the extracted features by triangulating map points of the closest zone to generate a reprojection error, and minimize the reprojection error to determine an optimal pose for the test image; and provide an output indicative of a location or probability of a location of the test image generated by the scope at the optimal pose in the environment, given the optimal pose. A sixteenth aspect is the fifteenth aspect, wherein the environment comprises the gastrointestinal tract or the bronchial tract of one or more lungs. A seventeenth aspect is the fifteenth aspect, wherein the scope is configured to provide the location of one or more targets including at least one of a polyp, a lesion, and cancerous tissue. An eighteenth aspect is the fifteenth aspect, wherein the scope comprises one or more sensors configured to receive the test image associated with the environment, and the test image is a visual image. In a nineteenth aspect, in the fifteenth aspect, the scope is an endoscope or a bronchoscope. A twentieth aspect is the fifteenth aspect, wherein the environment is a piping system, an underground environment, or an industrial facility.

[0016] Exemplary implementations can also include a non-transitory computer-readable medium having a storage device and a processor, the processor capable of executing instructions for image-based localization in a target tissue that fuses deep learning and geometric constraints for image-based localization. [Brief description of the drawings]

[0017] [Figure 1] FIG. 1 illustrates various aspects of a framework for training and testing in accordance with an example implementation. [Diagram 2] FIG. 2 illustrates an example representation and data generated by a simulator in accordance with an example implementation. [Diagram 3] FIG. 3 illustrates a training process according to an example implementation. [Figure 4] FIG. 4 illustrates a training approach in accordance with an example implementation. [Diagram 5] FIG. 5 illustrates a prediction approach according to an example implementation. [Figure 6] FIG. 6 is a diagram illustrating bundle adjustment according to an example implementation. [Figure 7] FIG. 7 illustrates the results of an exemplary implementation. [Figure 8] FIG. 8 illustrates results from an exemplary implementation. [Figure 9] FIG. 9 illustrates an example process for some example implementations. [Figure 10] FIG. 10 illustrates an example computing environment having an example computer device suitable for use with some example implementations. [Figure 11] FIG. 11 illustrates an example environment suitable for some example implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0018] The following detailed description provides further details of the drawings and exemplary implementations of the present application. Reference numbers and descriptions of redundant elements between figures are omitted for clarity. Terms used throughout this specification are provided by way of example and are not intended to be limiting.

[0019] Aspects of example implementations are directed to combining deep learning methods with geometric constraints for application in various fields, including, but not limited to, minimally invasive surgical (MIS) approaches (e.g., endoscopic approaches).

[0020] In contrast to open surgery, MIS reduces the surgical field of view. Therefore, the surgeon may have less information than in an open surgical approach. Therefore, MIS approaches must perform the procedure in a small space with long and thin tools without direct 3D vision. Moreover, the training dataset may be small and limited.

[0021] Exemplary implementations are directed to providing image-based localization for MIS techniques using tissues (eg, GI tract, lungs, etc.) by constraining localization within the tissue.

[0022] More specifically, the exemplary implementation classifies the test image to one of the training images based on similarity. The closest training image and its neighboring images are used with their pose information to generate an optimal pose (e.g., position and orientation) for the test image using feature registration and bundle adjustment. By locating the position and orientation of the scope, the surgeon can recognize the location of the scope within the body. Although the exemplary implementation refers to a scope, the exemplary implementation is not limited thereto and other MIS structures, devices, systems and / or methods may be substituted without departing from the scope of the present invention. For example, but not limited to, a probe may be substituted for a scope.

[0023] For example, but not by way of limitation, exemplary implementations are directed to a hybrid system that fuses deep learning with traditional geometry-based techniques. With this fusion approach, the system can be trained using smaller data sets. Thus, exemplary implementations can optionally provide a solution for localization using monocular RGB images with fewer samples of training data and textures.

[0024] Furthermore, the exemplary implementation uses a geometric method with deep learning techniques that can provide robustness of the estimated pose. More specifically, in the reprojection error minimization process, poses with large reprojection errors can be directly rejected.

[0025] Training images may be acquired and labeled to assign at least one image to each zone. The labeled images are used to train a neural network. Once the neural network is trained, test images are provided and classified into zones. Additionally, a training dataset tree and images from the training dataset are obtained. Key features are acquired and adjusted to recover the gaze point and minimize the projection error.

[0026] The exemplary implementations described above are directed to a hybrid system that fuses deep learning with geometry-based localization and tracking. More specifically, the deep learning component of the exemplary implementations provides high-level zone classification that can be used by geometry-based refinement to optimize the pose for a given test image.

[0027] The application of geometry to refine in example implementations can help constrain the predictions of deep learning models, optionally resulting in better pose estimation. Furthermore, by fusing deep learning and geometric techniques as described herein, accurate results can be achieved using small training datasets, and problems with related techniques such as outliers can be avoided.

[0028] This exemplary implementation provides a simulated dataset that provides a ground truth. In the training aspect, images are input to a neural network and the output is provided as image labels associated with zones of the environment. More specifically, the environment is subdivided into zones. This subdivision may be done automatically, such as by dividing the zones into equal lengths, or may be done using expertise in the medical domain, such as based on surgeon input regarding appropriate subdivision of the zones. Thus, each image is labeled with a zone and the images are classified into zones.

[0029] After the training phase, a test image is input and classified into zones. The test image is fed to the neural network and compared with the training dataset to extract corners from the training and test images and build the global locations of the map points. In other words, a comparison is made with the training dataset and key features are obtained, judged and identified as corner points.

[0030] For training images, the 3D points are projected onto the 2D images and an operation is performed to minimize the distance between the projected 3D points and the 2D images. As a result, the corner points are recovered in a way that minimizes the reprojection error.

[0031] 1 to 5 show various aspects of an exemplary implementation. Fig. 1 shows an overview of an exemplary implementation, including training and inference.

[0032] An exemplary implementation can be divided into two main blocks: PoseNet 107 (e.g., prediction) and pose refinement 109. For the prediction phase at 107, an exemplary implementation utilizes PoseNet, a deep learning framework (e.g., GoogLeNet). The system consists of a given number (e.g., 23) of convolutional layers and one or more fully connected layers (e.g., 1). At 107, the model does not learn actual poses, but rather learns zone-level classifications. During inference, PoseNet may classify the closest zone to which a given test image matches.

[0033] In the refinement phase at 109, the zones classified by PoseNet at 107 and the retrieved image and pose information from the training images are applied to determine the closest match. For pose optimization, a stream of neighboring poses is employed for the closest matching training image. The image stream and its corresponding pose information are used to estimate the pose.

[0034] More specifically, according to one exemplary implementation, Unity3D may be used to generate image-pose pairs from a phantom. A posenet model 101 is trained using these training sets from 101. For example, but not by way of limitation, pose regression may be replaced with zone classification. Thus, images of adjacent poses are classified as zones and labeled at 105.

[0035] With respect to training data, at 101, training images are provided to a deep learning neural network 103. As shown in Figure 2, at 200, the large intestine 201 may be divided into a number of zones identified by lines that process the image of the large intestine 201. For example, and without limitation, a first image 203 may represent a first image of the zones and a second image 205 may represent a second image of the zones.

[0036] 3 illustrates the aforementioned exemplary implementation as applied to a training phase at 300. As discussed above, training images 301 are provided to a deep learning neural network 303 to generate image labels 305 associated with zone classifications of the image's location, which are further represented as images at 313. A number of images 307 are correspondingly used for training at 309 and labeled at 311.

[0037] In PoseNet 107, test images 111 are provided to a deep learning neural network 113 to generate labels 115. This is also represented in FIG. 4 as 401. More specifically, for a test image, the most similar zone in the training set is predicted using the deep neural network.

[0038] For pose refinement at 109, the training database 117 receives input from PoseNet 107, which is also represented in FIG. 5 as 501. For example, and without limitation, the training database may provide image IDs associated with poses and labels, where the pose indicates the image state and the label indicates the classification associated with the pose.

[0039] This information is fed to a feature extractor, which receives output images 119, 121, and 123 associated with pose n-k 133, pose n 129, and pose n+k 125, respectively. For example, but not limited to, the zone and adjacent zones are included to avoid potential misclassification risks prior to bundle adjustment and reprojection error minimization.

[0040] Thus, in 135, 131 and 127, features are extracted from each of the images 123, 121 and 119, respectively. More specifically, a feature extractor is used to extract (e.g., SURF) features from the stream of images. These extracted features are further used in bundle adjustment, where features from each image are registered based on their characteristics.

[0041] More specifically, as shown in Figure 6, the feature extraction unit involves the use of an output image 601, which are images 119, 121 and 123. The feature extraction operation described above is performed on the output image 601 with respect to a number of adjacent poses n-k 603, n 605 and n+k 607, which may represent different zones. As shown at 609 and 611, triangulation of the map points may be performed based on the predicted zones.

[0042] At 139, the bundle adjustment uses features extracted from images 123, 121 and 119 (e.g., 135, 131 and 127) and the pose information of these images 133, 129 and 125 to perform a local bundle adjustment to map the pose. Since the poses of the relevant images are the ground truth, it is a multi-image triangulation process to map several corner feature points.

[0043] At 141, and also represented in FIG. 5 as 503, the reprojection error may be reoptimized, as may be defined in equation (1).

[0044]

number

[0045] P (position) and R (orientation) are the pose of the scope, and v i are the triangulated map points. Π() reprojects the 3D points into the 2D image space, and O i are the registered 2D observations. At 137, key features of the test image 111 may also be fed into a reprojection error minimization at 141.

[0046] If the optimized average reprojection level is below the threshold, the optimal global pose is found at 143. Otherwise, the initial pose is deemed incorrect and a failure of Posenet is attributed. Since the output of Posenet is fully measurable, the exemplary implementation provides a robust approach to identifying the validity of the output.

[0047] Furthermore, the reprojection error is minimized. More specifically, a registration is constructed between the key features and the test image, which is further used to optimize the pose of the test image by minimizing the reprojection error of the registered key feature points.

[0048] The above exemplary implementations may be implemented in a variety of applications. For example, the scope may be used in a medical setting to provide information associated with temporal changes related to characteristics. In one exemplary application, the growth of polyps may be tracked over time, and by allowing the scope to pinpoint the exact location and the ability to properly identify polyps and their size, medical personnel may perform more accurate tracking of polyps. As a result, medical personnel may provide more accurate risk analysis, as well as relevant recommendations and courses of action and more accurate methods.

[0049] Additionally, the scope may include devices or tools for performing procedures within the body's environment. For example, the scope may include tools that can change targets within the environment. In one exemplary implementation, the tool may be a cutting tool such as a laser, heat, or blade, or other cutting structure as would be understood by one of ordinary skill in the art. The cutting tool may perform procedures such as removing polyps in real time if the polyps are larger than a certain size.

[0050] In principle, polyps are only removed depending on the medical approach or if the medical practitioner judges them to be too large or harmful to the patient, and a more conservative approach may follow the growth of the target in this environment. Furthermore, the scope may be used according to an exemplary implementation to more accurately perform follow-up screening after a procedure performed by a device or tool.

[0051] Although the example of a polyp is shown herein, the present exemplary implementations are not limited to that example and other environments or targets may be substituted without departing from the scope of the present invention. For example, but not limited to, the environment may be the bronchi of the lungs rather than the GI tract. Similarly, the target may be a lesion or tumor rather than a polyp.

[0052] Further, an exemplary practical implementation may feed the results into a predictive tool. According to such an exemplary approach, analysis may be performed to generate a predictive risk assessment based on demographics, organizational growth rate, and historical data. The predictive risk assessment may be reviewed by a medical professional, who may verify or validate the results of the predictive tool. The validation or verification by the medical professional may be fed back into the predictive tool to improve its accuracy. Alternatively, the predictive risk assessment may be input into a decision support system with or without validation or verification by the medical professional.

[0053] In such a situation, the decision support system may provide recommendations to the medical practitioner in real time or after the scope is removed. In the option where recommendations are provided to the medical practitioner in real time, the scope may also carry a cutting tool so that real time action may be taken based on the decision support system's recommendations.

[0054] Further, while the exemplary implementations described above may define this environment as an environment within the human body that does not have clearly defined corner points, such as the lungs or intestines, the exemplary implementations are not limited thereto and other environments having similar characteristics may be within the scope.

[0055] For example, but not by way of limitation, piping systems such as sewer or water pipes can be difficult to inspect for damage, wear or replacement due to the difficulty of being able to determine which exact pipe segment an inspection tool is located on. Using the present exemplary implementations, sewer and water pipes can be inspected more accurately over time, and pipe maintenance or replacement, etc. can be performed with less precision. Similar approaches can be addressed in industrial safety, such as awareness in factory environments, underwater, underground environments such as caves, or other similar environments that meet the conditions associated with the present exemplary implementations.

[0056] Figure 7 shows at 700 the results associated with an exemplary implementation. At 701, a related art approach is shown that involves only deep learning. More specifically, an approach that uses regression is shown, and as can be seen, the number of outliers outside the ground truth is significant in both magnitude and number. As mentioned above, this is due to the related art problem of the narrow field of view of the camera and the associated risk of misclassification.

[0057] In 703, an approach is presented that uses test image information using classification only, but with this approach the data is limited to what is strictly available from the video.

[0058] An example implementation approach including classification and bundle adjustment is shown at 705. Although the errors are small, these errors are mainly due to the texture of the images.

[0059] Figure 8 shows a validation of the exemplary implementation, showing the difference in error over time. The X-axis shows key frames over time, and the Y-axis shows the error. At 801, the position error is shown, and at 803, the angle error is shown. The blue line shows the error using the exemplary implementation technique, and the red line shows the error calculated using only the classification technique, which corresponds to 703, as described above and shown in Figure 7.

[0060] More specifically, we generate a simulated dataset based on an off-the-shelf model of the male digestive system. A virtual colonoscope is placed in the colon and observations are simulated. Unity3D (https: / / unity.com / ) is used to generate a series of simulated 2D RGB images using a rigorous pinhole camera model. The frame rate and size of the simulated in vivo digestion (e.g., as shown in Figure 2) are 30 frames / s and 640 × 480. The global pose of the colonoscope is recorded simultaneously.

[0061] As shown and discussed above, the red plots are for results with classification only (e.g., related art), and the blue plots are for results where pose refinement according to the exemplary implementation has been performed. As can be seen, generally speaking, these results show better accuracy for the pose refinement performed, both in terms of positional and angular differences.

[0062] More specifically, as discussed above, Figure 8 shows a comparison 800 of position difference error versus keyframe ID at 801, and a comparison 803 of angle difference error versus keyframe ID. Table 1 shows an error comparison between the related art (i.e., ContextualNet) and the example implementation described herein.

[0063] [Table 1]

[0064] Exemplary implementations may be integrated with other sensors or approaches, for example, other sensors may be integrated on the scope, such as, but not limited to, an inertial measurement unit, temperature sensors, acidity sensors, or other sensors associated with sensing surroundings associated with the environment.

[0065] Similarly, multiple sensors of a given sensor type may be used, and related art approaches may not use such multiple sensors, and the related art focuses on providing accurate location as opposed to using the labeling, feature extraction, bundle adjustment and reprojection error minimization approaches described herein.

[0066] Because the present exemplary implementations do not require higher accuracy of sensors or cameras or additional training data sets, existing equipment may be used with the exemplary implementations to achieve more accurate results, and thus the need to upgrade hardware to obtain more accurate cameras or sensors may be reduced.

[0067] Additionally, increased accuracy also allows for interchangeability of different types of cameras and scopes, allowing different medical facilities to more easily exchange results and data, involve more and different medical personnel, and analyze, make recommendations, and take action appropriately, without sacrificing accuracy.

[0068] 9 illustrates an example process 900 according to an example implementation. The example process 900 may be performed on one or more devices as described herein.

[0069] At 901, a neural network receives input and labels training images. For example, but not by way of limitation, as described above, the training images may be generated from a simulation. Alternatively, historical data associated with one or more patients may be provided. The training data is used in the model to replace pose regression with zone classification. For example, but not by way of limitation, images of adjacent poses may be classified as zones.

[0070] At 903, feature extraction occurs. More specifically, the images are provided to a training database 117. Based on the key features, a classification decision is provided as to whether the image features can be classified as being in a particular pose.

[0071] At 905, adjustments are made. More specifically, the prediction zones are used to triangulate the map points as described above.

[0072] In 907, an operation is performed to minimize the reprojection error of the map points on the test image by adjusting the pose, and based on the result of this operation, an optimal pose is determined.

[0073] At 909, an output is provided. For example, but not limited to, this output may be an indication of the zone of the image or a location within the zone, or a scope associated with the image. Thus, a medical professional may be aided in determining the location of the image in a target tissue, such as the GI tract, lungs, or other tissue.

[0074] 10 illustrates an exemplary computing environment 1000 having an exemplary computer device 1005 suitable for use with some exemplary implementations. The computing device 1005 in the computing environment 1000 can include one or more processing units, cores, or processors 1010, memory 1015 (e.g., RAM, ROM, etc.), internal storage 1020 (e.g., magnetic storage, optical storage, solid-state storage, and / or organic storage), and / or input / output interfaces 1025, any of which can be coupled to or incorporated into the computing device 1005 over a communication mechanism or bus 1030 for communicating information.

[0075] According to an exemplary implementation, processing associated with neural activity may be performed on a processor 1010 that is a central processing unit (CPU). Alternatively, other processors may be substituted without departing from the concept of the present invention. For example, but not limited to, a graphics processing unit (GPU) and / or a neural processing unit (NPU) may be substituted in place of or used in combination with the CPU to perform the processing of the exemplary implementations described above.

[0076] The computing device 1005 may be communicatively coupled to an input / interface 1035 and an output device / interface 1040. One or both of the input / interface 1035 and the output device / interface 1040 may be a wired or wireless interface and may be detachable. The input / interface 1035 may include any device, component, sensor, or physical or virtual interface that may be used to provide input (e.g., buttons, touch screen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, optical reader, etc.).

[0077] Output devices / interfaces 1040 may include displays, televisions, monitors, printers, speakers, Braille, etc. In some exemplary implementations, input / interfaces 1035 (e.g., user interfaces) and output devices / interfaces 1040 may be embedded in or physically coupled to computing device 1005. In other exemplary implementations, other computing devices may function as or provide functionality for input / interfaces 1035 and output devices / interfaces 1040 for computing device 1005.

[0078] Examples of computing devices 1005 may include, but are not limited to, highly mobile devices (e.g., smart phones, devices in vehicles and other machines, devices carried by humans and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, server devices, other computers, information kiosks, televisions, radios with one or more processors embedded therein and / or coupled to the processors, etc.).

[0079] Computing device 1005 may be communicatively coupled (e.g., via input / output interface 1025) to any number of networked components, devices, including one or more computing devices of the same or different configurations, as well as external storage 1045 and a network 1050 for communicating with the system. Computing device 1005 or any connected computing device may function as, provide services to, or be considered as a server, client, thin server, general purpose machine, special purpose machine, or another label. For example, but not by way of limitation, network 1050 may include a blockchain network and / or a cloud.

[0080] The I / O interface 1025 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11xs, Universal System Bus, WiMAX, modem, cellular network protocols, etc.) for communicating information to and from at least all connected components, devices, and networks in the computing environment 1000. The network 1050 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, etc.).

[0081] The computing device 1005 can use and / or communicate using computer usable or computer readable media, including transitory media and non-transitory media. Transitory media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, etc. Non-transitory media include magnetic media (e.g., disks and tapes), optical media (e.g., CD ROM, digital video disks, Blu-ray disks), solid media (e.g., RAM, ROM, flash memory, solid state storage), and other non-volatile storage or memory.

[0082] The computing device 1005 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be retrieved from a transitory medium and can be stored in and retrieved from a non-transitory medium. The executable instructions can originate from one or more of any programming, scripting, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0083] The processor 1010 may execute under any operating system (OS) (not shown), in a native or virtual environment. One or more applications may be deployed, including a logic unit 1055, an application programming interface (API) unit 1060, an input unit 1065, an output unit 1070, a training unit 1075, a feature extraction unit 1080, a bundle adjustment unit 1085, and an inter-unit communication mechanism 1095 for different units to communicate with each other, with the OS, and with other applications (not shown).

[0084] For example, the training unit 1075, the feature extraction unit 1080, and the bundle adjustment unit 1085 may perform one or more processes described above for the structures described above. The described units and elements may be varied in design, function, configuration, or implementation and are not limited to the provided description.

[0085] In some example implementations, when information or execution instructions are received by the API unit 1060, it may be communicated to one or more other units (e.g., the logic unit 1055, the input unit 1065, the training unit 1075, the feature extraction unit 1080, and the bundle adjustment unit 1085).

[0086] For example, the training unit 1075 may receive and process simulated data, historical data, or information from one or more sensors, as described above. The output of the training unit 1075 is provided to a feature extraction unit 1080, which performs the necessary operations, for example based on the application of a neural network as described above and illustrated in Figures 1-5. Furthermore, a bundle adjustment unit 1085 may perform operations based on the output of the training unit 1075 and the feature extraction unit 1080 to minimize the reprojection error and provide an output signal.

[0087] In some cases, the logic unit 1055 may be configured to control the flow of information between units and direct the services provided by the API unit 1060, the input unit 1065, the training unit 1075, the feature extraction unit 1080, and the bundle adjustment unit 1085 in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled by the logic unit 1055 alone or together with the API unit 1060.

[0088] 11 illustrates an example environment suitable for some example implementations. Environment 1100 includes devices 1105-1145, each communicatively connected to at least one other device (e.g., by wired and / or wireless connections), for example, via network 1160. Some devices may be communicatively connected to one or more storage devices 1130 and 1145.

[0089] An example of one or more devices 1105-1145 may each be the computing device 1005 of Figure 10. The devices 1105-1145 may include, but are not limited to, a computer 1105 (e.g., a laptop computing device) having a monitor and associated webcam, a mobile device 1110 (e.g., a smartphone or tablet), a television 1115, a device associated with a vehicle 1120, a server computer 1125, computing devices 1135-1140, storage devices 1130 and 1145, as described above.

[0090] In some implementations, the devices 1105-1120 can be considered user devices associated with a user, who can remotely obtain sensed inputs that are used as inputs for the exemplary implementations described above. In the present exemplary implementations, one or more of these user devices 1105-1120 can be associated with one or more sensors, such as a camera, embedded in the user's body, temporarily or permanently, away from the patient care facility, that can sense information necessary for the present exemplary implementations, as described above.

[0091] While the foregoing exemplary implementations are provided to illustrate the scope of the invention, these implementations are not intended to be limiting, and other approaches or implementations may be substituted or added without departing from the scope of the invention. For example, and not by way of limitation, imaging techniques other than those disclosed herein may be used.

[0092] According to one exemplary implementation, algorithms such as SuperPoint may be used to train image point detection and determination. Additionally, exemplary implementations may employ alternative image classification algorithms and / or use other neural network structures (e.g., Siamese network). Additional approaches integrate expertise in zone classification, apply two-image enhancement through the use of techniques such as shaping, daylighting and illumination, and / or use a single image for depth methods.

[0093] Exemplary implementations may have various advantages and benefits, but are not required. For example, but not limited to, exemplary implementations are operable on small data sets. Additionally, exemplary implementations provide location constraints within a target tissue, such as the colon or lungs. Thus, a surgeon can more accurately localize the position of the scope by using video. Additionally, exemplary implementations provide much higher accuracy than related art approaches.

[0094] Although several exemplary implementations have been shown and described, these exemplary implementations are provided to convey the subject matter described herein to those skilled in the art. It should be understood that the subject matter described herein may be embodied in various forms, without being limited to the exemplary implementations described. The subject matter described herein may be practiced without the subject matter specifically defined or described, or with other or various elements or subject matter not described. Those skilled in the art will understand that changes can be made in these exemplary implementations without departing from the subject matter described herein as defined in the appended claims and their equivalents.

[0095] Aspects of certain non-limiting embodiments of the present disclosure address the features described above and / or other features not described above, but aspects of non-limiting embodiments need not address the features described above, and aspects of non-limiting embodiments of the present disclosure may not address the features described above.

Claims

1. 1. A computer-implemented method, comprising: applying training images of the environment divided into zones to a neural network and performing classification using the neural network to label test images based on the closest one of the zones; extracting features from the training image that matches the closest zone and its neighboring images and obtaining corresponding pose information for each of them; performing a bundle adjustment on the extracted features by triangulating the map points of the nearest zones to generate a reprojection error, and minimizing the reprojection error to determine an optimal pose for the test image; providing an output indicative of a location or a probability of a location of the test image in the optimal pose within the environment, relative to the optimal pose; 4. A computer-implemented method comprising:

2. 2. The computer-implemented method of claim 1, wherein applying the training images comprises receiving the training images associated with poses within a zone of the environment as historical or simulated data, and providing the received training images to a neural network.

3. 3. The computer-implemented method of claim 2, wherein the neural network is a deep learning neural network that learns zones associated with the pose and determines the closest zone for the test image.

4. 2. The computer-implemented method of claim 1 , wherein the bundle adjustment comprises reprojecting 3D points associated with a measured pose and the triangulated map points into a 2D image space to generate a result, and comparing the result to a registered 2D observation to determine the reprojection error.

5. The computer-implemented method of claim 4 , wherein if a reprojection error is below a threshold, the pose of the test image is confirmed to be the optimal pose.

6. The computer-implemented method of claim 4 , wherein if a reprojection error exceeds a threshold, the pose of the test image is determined to be incorrect and the calculation of the pose of the test image is determined to be correct.

7. The computer-implemented method of claim 1 , wherein minimizing the reprojection error comprises adjusting the pose of the test images to minimize the reprojection error.

8. The processor: applying training images of the environment divided into zones to a neural network and performing classification using the neural network to label test images based on the closest one of the zones; extracting features from the training image that matches the closest zone and its neighboring images and obtaining corresponding pose information for each of them; performing a bundle adjustment on the extracted features by triangulating the map points of the nearest zones to generate a reprojection error, and minimizing the reprojection error to determine an optimal pose for the test image; providing an output indicative of a location or a probability of a location of the test image in the optimal pose within the environment, relative to the optimal pose; A program that executes a process including:

9. 9. The program of claim 8, wherein applying the training images comprises receiving the training images associated with poses within a zone of the environment as historical or simulated data, and providing the received training images to a neural network.

10. 10. The method of claim 9, wherein the neural network is a deep learning neural network that learns zones associated with the pose and determines the closest zone for the test image.

11. 9. The program of claim 8, wherein the bundle adjustment comprises reprojecting 3D points associated with a measured pose and the triangulated map points into a 2D image space to generate a result, and comparing the result to a registered 2D observation to determine the reprojection error.

12. The program of claim 11 , wherein the pose of the test image is confirmed to be the optimal pose if the reprojection error is below a threshold.

13. The program of claim 11 , wherein if a reprojection error exceeds a threshold, the pose of the test image is determined to be incorrect and the calculation of the pose of the test image is determined to be correct.

14. The method of claim 8 , wherein minimizing the reprojection error comprises adjusting the pose of the test images to minimize the reprojection error.

15. 1. A computer-implemented system for locating and tracking a scope within an environment to identify a target, comprising: applying training images of the environment associated with the scope and divided into zones to a neural network and performing classification using the neural network to label test images generated by the scope based on the closest one of the zones of the environment associated with the scope; extracting features from the training image that matches the closest zone and its neighboring images and obtaining corresponding pose information for each of them; performing a bundle adjustment on the extracted features by triangulating the map points of the nearest zones to generate a reprojection error, and minimizing the reprojection error to determine an optimal pose for the test image; providing an output indicative of a location or a probability of a location of the test image generated by the scope at the optimal pose within the environment, relative to the optimal pose; 23. A computer-implemented system configured to:

16. The computer-implemented system of claim 15 , wherein the environment comprises a gastrointestinal tract or a bronchial tract of one or more lungs.

17. 16. The computer-implemented system of claim 15, wherein the scope is configured to provide a location of one or more targets including at least one of a polyp, a lesion, and cancerous tissue.

18. 16. The computer-implemented system of claim 15, wherein the scope comprises one or more sensors configured to receive the test image associated with the environment, the test image being a visual image.

19. The computer-implemented system of claim 15 , wherein the scope is an endoscope or a bronchoscope.

20. The computer implemented system of claim 15 , wherein the environment is a plumbing system, an underground environment, or an industrial facility.

Citation Information

Patent Citations

  • Endoscopic image reproducing apparatus

    JP2012139456A

  • Medical image display apparatus

    JP2013085593A

  • Systems and methods for guiding laparoscopic surgical procedures through augmentation of anatomical models

    JP2018522610A

  • Medical image processing device, medical image processing system, medical image processing method, and program

    WO2020012872A1