Methods, systems, and media for image-based localization

By combining deep learning and geometric constraints, and utilizing neural networks and bundle adjustment, the problem of inaccurate positioning of observation devices in minimally invasive surgery was solved. This enabled high-precision positioning and tracking in environments such as the GI tract and lungs, supporting more precise medical operations and risk assessments.

CN113409386BActive Publication Date: 2026-03-17FUJIFILM BUSINESS INNOVATION CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-19
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately locate and track internal structures, such as the GI tract and lungs, during minimally invasive surgery, particularly due to issues such as a lack of corner points and textures, insufficient datasets, and outliers, leading to inaccurate localization and unreliable information.

Method used

By combining deep learning and geometric constraints, the method classifies environmental images into regions using neural networks and determines the optimal pose of the observation device through bundle adjustment and reprojection error minimization, thus integrating deep learning and geometric methods to improve positioning accuracy.

Benefits of technology

It achieves higher positioning accuracy and robustness under small dataset conditions, reduces the impact of outliers, provides more accurate observation device location information, and supports more precise medical operations and risk analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113409386B_ABST
    Figure CN113409386B_ABST
Patent Text Reader

Abstract

The present application provides methods, systems, and media for image-based localization. A computer-implemented method includes the steps of: applying a training image of an environment divided into regions to a neural network and performing a classification to label a test image based on a closest region in the region; extracting features from pose information of the retrieved training image and the test image that match the closest region; bundle adjusting the extracted features by triangulating map points of the closest region to generate re-projection errors and minimizing the re-projection errors to determine a best pose of the test image; and for the best pose, providing an output indicating a position or a probability of a position of the test image in the environment in the best pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The example implementations cover various aspects of methods, systems, and user experiences associated with image-based localization in an environment, and more specifically, schemes that integrate deep learning and geometric constraints for image-based localization. Background Technology

[0002] Endoscopic systems based on related technologies can provide minimally invasive methods for examining internal structures. More specifically, minimally invasive surgical (MIS) protocols based on related technologies can provide medical practitioners with tools to examine internal structures and can be used for precise therapeutic interventions.

[0003] For example, observational devices such as endoscopes or bronchoscopes can be placed in a patient's environment, such as the intestines or lungs, to examine their structures. Devices on the observational device, such as sensors or cameras, can sense information and provide it to the user, such as images and videos of the environment. Medical professionals, such as surgeons, can analyze the video. Based on the analysis, surgeons can provide recommendations or perform actions.

[0004] Leveraging robotics and sensor technologies in related fields, various gastrointestinal (GI) tract observation device solutions have been developed. For such GI tract solutions, accurate localization and tracking enable medical practitioners to locate and monitor the progression of various pathological findings, such as polyps, cancerous tissue, and lesions. Endoscopic systems in this field can meet the need for precise therapeutic interventions and therefore must be able to accurately locate and track within a given gastrointestinal (GI) tract and / or bronchial tract.

[0005] Related techniques for tracking GI channels can include image similarity comparison, such as using image descriptors to compare image similarities, also known as image classification. Furthermore, related techniques can utilize geometry-based pose regression, such as geometric techniques like SLAM or shape recovery from shadows for image-to-model registration, also known as geometry optimization. Related techniques can also use deep learning-based images for pose regression.

[0006] Deep learning approaches in related technologies have various problems and drawbacks unique to tracking applications such as colonoscopy or bronchoscopy, such as small annotated training datasets and a lack of recognizable textures, unlike other indoor or outdoor environments where deep learning has been used in related technologies. For example, there are no corners to define textures, while body tissues are characterized by blood flow, smooth curves, and tubular structures without corners. Therefore, there are angular-like features on volumetric surfaces, as well as mixtures of solids and liquids.

[0007] For example, but not as a limitation, deep learning and regression schemes in related technologies suffer from problems such as insufficient datasets and lack of angles and textures, as described above; in these respects, surgical scenarios differ from and are distinguishable from related technological schemes used in other environments (such as autonomous driving). For example, due to the unique physiological characteristics of the GI tracts in the lungs, there are many tubular structures without angles.

[0008] Furthermore, because deep learning and regression techniques attempt to locate the observation device, additional problems and / or drawbacks arise. For example, but not as a limitation, there is another problem associated with outliers that are completely outside the environment, due to a lack of sufficiently high-quality and large datasets for training. The consequences of these outliers are quite significant in the medical field, where identifying an observation device completely outside the environment, such as the lungs or GI tract, makes it difficult for medical professionals to rely on the information and make appropriate analyses and decisions.

[0009] Correlation schemes for localization in the GI channel use monocular images and leverage computer vision techniques (e.g., SIFT and SURF). However, such correlation schemes have various problems and drawbacks, such as distortion, intensity, and varying obstacles. For example, correlation systems may lack depth perception or perform poor localization within the limited field of view provided by the RGB / monocular images of the correlation technology. For instance, the observation device may have a small field of view due to the proximity of soft tissues in the patient's environment.

[0010] Since 3D depth information is not provided and the only available data is RGB video, the depth / stereoscopic observation device positioning system of the relevant technology cannot be directly applied to monocular endoscopes.

[0011] Furthermore, the vast amount of data required to generalize deep learning-based localization and tracking remains unmet. Such data is difficult to obtain, especially in the medical field due to privacy concerns. Additionally, the geometry-based methods of related technologies are unsuitable for tracking with GI (Genomic Anatomy) devices due to the limited number of features and the potential for registration loss. Increasing the dataset size by forcefully inserting observational devices into patients is also impractical or unhealthy.

[0012] Therefore, practitioners may find it difficult to determine the location of observation equipment within the human body's environment, such as the position of an endoscope in the GI tract. This problem is exacerbated in certain tissues, such as the lungs, due to the branching physiology of the lungs. Summary of the Invention

[0013] According to an example implementation, a computer-implemented method is provided, comprising the steps of: applying training images of an environment to be divided into regions to a neural network, and performing classification to label a test image based on the nearest region in the regions; extracting features based on pose information of the retrieved training and test images that match the nearest region; performing bundle lead adjustment on the extracted features by triangulating map points of the nearest region to generate a reprojection error, and minimizing the reprojection error to determine the optimal pose of the test image; and, for the optimal pose, providing an output indicating the position or position probability of the test image in the environment at the optimal pose.

[0014] Example implementations may also include a non-transitory computer-readable medium with memory and a processor capable of executing instructions for image-based localization in a target tissue, which incorporates deep learning and geometric constraints for image-based localization. Attached Figure Description

[0015] This patent or application document contains at least one color drawing. The Patent Office will provide a copy of this patent or application publication with color drawings upon request and payment of the necessary fees.

[0016] Figure 1 This illustrates various aspects of the framework for training and testing, based on the example implementation.

[0017] Figure 2 Examples of representations and data generated by the simulator based on the example implementation are shown.

[0018] Figure 3 The training process is illustrated based on the example implementation.

[0019] Figure 4 The training scheme based on the example implementation is illustrated.

[0020] Figure 5 An example prediction scheme based on the example implementation is shown.

[0021] Figure 6 This illustrates bundle adjustment based on the example implementation.

[0022] Figure 7 The results are shown based on the example implementation.

[0023] Figure 8 The results are shown based on the example implementation.

[0024] Figure 9 Example processing is shown for some example implementations.

[0025] Figure 10 An example computing environment with example computer devices suitable for use in some example implementations is illustrated.

[0026] Figure 11 An example environment is shown that applies to some of the example implementations. Detailed Implementation

[0027] The following detailed description provides further details of the accompanying drawings and exemplary implementations of this application. For clarity, reference numerals and descriptions of redundant elements between the drawings have been omitted. The terminology used throughout the specification is provided by way of example only and is not intended to be limiting.

[0028] The aspects of the example implementations are designed to combine deep learning methods with geometric constraints for use in a variety of fields, including but not limited to minimally invasive surgical (MIS) procedures (e.g., endoscopic procedures).

[0029] In contrast to open surgery, MIS narrows the surgical field of view. Therefore, surgeons receive far less information compared to open surgery. Consequently, MIS procedures require the use of slender tools in confined spaces without direct 3D vision. Furthermore, the training dataset is small and limited.

[0030] The example implementation aims to leverage the image-based localization provided by an organization (e.g., the gastrointestinal tract, lungs, etc.) for MIS technologies by constraining localization within the organization.

[0031] More specifically, the example implementation classifies the test image into one of the training images based on similarity. The closest training image and its neighboring images, along with their pose information, are used to generate the optimal pose (e.g., position and orientation) of the test image using feature registration and bundle adjustment. By locating the position and orientation of the observation device, the surgeon can determine the device's location within the body. While this example implementation involves an observation device, it is not limited thereto and can be substituted with other MIS structures, devices, systems, and / or methods without departing from the scope of the invention. For example, but not as a limitation, a probe can be used instead of the observation device.

[0032] For example, but not as a limitation, the example implementations involve hybrid systems that fuse deep learning with traditional geometry-based techniques. Using this fusion approach, the system can be trained with a smaller dataset. Therefore, the example implementations can optionally provide a solution for localization using monocular RGB images when training data and texture samples are limited.

[0033] Furthermore, the example implementation uses a geometric approach with deep learning techniques, which can provide robustness to the estimated pose. More specifically, poses with large reprojection errors can be directly rejected during the reprojection error minimization process.

[0034] Training images can be obtained and labeled, thus assigning at least one image to each region. The labeled images are used to train the neural network. While the neural network is trained, test images can be provided and classified into the regions. Furthermore, a training dataset tree and images from the training dataset are obtained. Key features are obtained and adjusted to recover points of interest and minimize any projection errors.

[0035] The above example implementation involves a hybrid system that integrates deep learning with geometry-based localization and tracking. More specifically, the deep learning component according to the example implementation provides high-level region classification, which can be used by geometry-based thinning to optimize the pose of a given test image.

[0036] In the example implementation, applying geometry to perform refinement can help constrain the predictions of the deep learning model and optionally improve pose estimation. Furthermore, by fusing deep learning and geometric techniques as described in this paper, accurate results can be obtained using a small training dataset, and problems associated with these techniques, such as outliers, can be avoided.

[0037] This example implementation provides a simulated dataset that offers ground truth. For training, images are input into a neural network, and image labels, as associated with regions of the environment, are provided as output. More specifically, the environment is segmented into regions. This segmentation can be performed automatically, for example, by dividing regions into equal lengths, or it can be performed using expert knowledge in the medical field, such as input based on a surgeon's assessment of appropriate region segmentation. Therefore, each image is labeled against a region, and the image is classified into regions.

[0038] After the training phase, a test image is input and classified into a region. The test image is fed into the neural network and compared with the training dataset to extract corners based on both the training and test images, thus establishing the global location of map points. In other words, it is compared with the training dataset, and key features are obtained and determined, and identified as corner points.

[0039] For the training image, 3D points are projected onto the 2D image, and an operation is performed to minimize the distance between the projected 3D points and the 2D image. Therefore, corner points are recovered in a manner that minimizes reprojection error.

[0040] Figures 1 to 5 Various aspects of the example implementation are illustrated. Figure 1 This illustrates an overall view of the example implementation, including training and derivation.

[0041] The example implementation can be divided into two main parts: PoseNet107 (e.g., prediction) and pose refinement 109. In the prediction phase of 107, the example implementation utilizes PoseNet, a deep learning framework (e.g.,...).

[0042] (GoogLeNet). The system consists of a specified number (e.g., 2^3) of convolutional layers and one or more fully connected layers (e.g., 1). At 10^7, the model learns region-level classifications rather than actual poses. During derivation, PoseNet can classify the closest regions that have a matching degree to a given test image.

[0043] In the refinement stage (step 109), the regions classified by PoseNet in step 107, along with the image and pose information retrieved from the training images, are applied to determine the closest match. For pose optimization, a stream of adjacent poses is applied to the training image that provides the closest match. The stream of images and their corresponding pose information is used for pose estimation.

[0044] More specifically, according to one example implementation, Unity3D can be used to generate image-pose pairs from a phantom. The PoseNet model 101 is trained using these training sets from 101. For example, but not limited to, pose regression can be replaced with region classification. Thus, images in adjacent poses are classified as regions, and labeling is performed at 105.

[0045] Regarding the training data, at position 101, the training image is provided to the deep learning neural network 103. For example... Figure 2 As shown, at position 200, the large intestine 201 can be divided into multiple regions identified by lines that process the image of the large intestine 201. For example, but not as a limitation, the first image 203 can represent the first region within the region, and the second image 205 can represent the second region within the region.

[0046] Figure 3 The aforementioned example implementation is illustrated at 300 during the training phase. As described above, training image 301 is provided to deep learning neural network 303 to generate image labels 305 associated with region classifications of the image locations. This is further represented as the image at 313. Multiple images 307 are correspondingly used for training at 309 and labeled at 311.

[0047] In PoseNet 107, the test image 111 is fed to the deep learning neural network 113, and labels 115 are generated. This... Figure 4This is also represented as 401. More specifically, for the test image, a deep neural network is used to predict the most similar region in the training set.

[0048] In the pose refinement at position 109, training database 117 receives input from PoseNet 107. This is still... Figure 5 The value is represented as 501. For example, but not as a limitation, the training database can provide image IDs associated with pose and label. The pose indicates the state of the image, and the label indicates the classification associated with the pose.

[0049] This information is fed to a feature extractor, which receives output images associated with poses n-k133, n 129, and n+k 125 at points 119, 121, and 123, respectively. For example, but not as a limitation, the region and its adjacent regions are included to avoid potential misclassification risks before minimizing bundle adjustment and reprojection errors.

[0050] Therefore, features are extracted from each of images 123, 121, and 119 at 135, 131, and 127, respectively. More specifically, a feature extractor (e.g., SURF) is used to extract features from the image stream. These extracted features will be further used for bundle adjustment, and the features from each image are registered based on their properties.

[0051] More specifically, and as Figure 6 As shown, the feature extractor involves using output image 601 (which is images 119, 121, and 123). The aforementioned feature extraction operation is performed on output image 601 for multiple adjacent poses nk 603, n 605, and n+k 607, which can indicate various regions. As shown at 609 and 611, map point triangulation can be performed based on the predicted regions.

[0052] In step 139, during bundle adjustment, local bundle adjustment is performed using features extracted from images 123, 121, and 119 (e.g., 135, 131, and 127) and pose information 133, 129, and 125 for these images to map the pose. Since the pose of the images involved is the ground truth, mapping multiple corner feature points is a multi-image triangulation process.

[0053] At 141, and still Figure 5 The reprojection error, which can be defined in equation (1), can be re-optimized, as indicated by 503.

[0054]

[0055] P (position) and R (orientation) are the orientation of the observation device, and v iThese are triangulated map points. ∏() reprojects 3D points onto the 2D image space, and O i These are registered 2D observations. At position 137, key features of test image 111 can also be fed into the reprojection error minimization at position 141.

[0056] If the optimized average reprojection level is below or equal to a threshold, the optimal global pose is found at 143. Otherwise, the initial pose is assumed to be incorrect and caused by a fault in PoseNet. Since the output of PoseNet can be fully measured, the example implementation provides a robust way to identify the validity of the output.

[0057] Furthermore, reprojection error is minimized. More specifically, registration is established between key features and the test image, which is further used to optimize the pose of the test image by minimizing the reprojection error of the registered key feature points.

[0058] The aforementioned example implementations can be implemented in a variety of applications. For instance, observation devices can be used in medical settings to provide information associated with temporary changes in characteristics. In one example application, polyp growth over time can be tracked, and by being able to pinpoint the exact location of the observation device and correctly identify the polyp and its size, medical professionals can track polyps more precisely. As a result, medical professionals can be able to provide more accurate risk analysis and offer relevant advice and action plans in a more precise manner.

[0059] Furthermore, the observation device may also include means or tools for performing actions within the human body environment. For example, the observation device may include tools capable of modifying a target within the environment. In one example implementation, the tool may be a cutting tool, such as a laser, heat, or a knife, or other cutting structures understood by those skilled in the art. The cutting tool can perform actions, such as cutting the polyp in real time if it is larger than a certain size.

[0060] Depending on the medical protocol, polyps are typically removed, or only when medical professionals determine that the polyp is too large or harmful to the patient; according to a more conservative approach, the growth of the target in the environment can be tracked. Additionally, observational devices can be used, as illustrated in the example implementation, to perform subsequent screening more accurately after action has been taken by the device or tool.

[0061] While this document illustrates an example of a polyp, the implementation of this example is not limited thereto, and it can be substituted with other environments or targets without departing from the scope of the invention. For example, but not by limitation, the environment could be the bronchi of the lungs instead of the GI tract. Similarly, the target could be a lesion or tumor instead of a polyp.

[0062] Furthermore, the example implementation can feed the results into a predictive tool. Based on this example scheme, analysis can be performed using demographic information, organizational growth rates, and historical data to generate a predictive risk assessment. This predictive risk assessment can be reviewed by medical professionals, who can then verify or validate the results of the predictive tool. This expert validation or verification can be fed back into the predictive tool to improve its accuracy. Alternatively, with or without expert validation or verification, the predictive risk assessment can be input into a decision support system.

[0063] In this scenario, the decision support system can provide recommendations to healthcare professionals either in real time or after the observation device has been removed. Among the options where recommendations are provided to healthcare professionals in real time, since the observation device can also carry cutting tools, real-time actions can be performed based on the recommendations of the decision support system.

[0064] Furthermore, although the aforementioned example implementation can define the environment as an environment within the human body without clearly defined corners, such as the lungs or intestines, the example implementation is not limited to this, and other environments with similar characteristics can also be within the scope of this invention.

[0065] For example, but not as a limitation, piping systems such as wastewater or water supply pipes may be difficult to inspect for damage, wear, and cuts, or require replacement, due to the difficulty in accurately determining which pipe section the inspection tool is located in. By employing the implementation method described in this example, wastewater and water supply pipes can be inspected more accurately over time, and pipe maintenance, replacement, etc., can be performed with fewer accuracy issues. Similar solutions can be adopted in industrial safety settings, such as in factory environments, underwater, underground environments (e.g., caves), or other similar environments that meet the conditions associated with the implementation method described in this example.

[0066] Figure 7 The results associated with the example implementation are illustrated at 700. At 701, a related technical solution involving only deep learning is shown. More specifically, a solution employing regression is shown, and it can be seen that outliers outside the ground truth are significant in both size and number. As explained above, this is due to the technical problem of the small field of view of the camera and the associated risk of misclassification.

[0067] At point 703, a scheme for employing test image information using only classification is shown. However, according to this scheme, the data is limited to the strictly available data from the video.

[0068] In section 705, a scheme based on an example implementation is shown, which includes classification and bundle adjustment. Although a small number of errors exist, these errors are mainly due to image texture.

[0069] Figure 8 The verification of the example implementation is illustrated, showing the difference in error over time. The X-axis represents keyframes changing over time, and the Y-axis represents the error. At 801, the positional error is represented, and at 803, the angle error is represented. The blue line represents the error using the techniques of the example implementation, and the red line represents the error calculated using only classification techniques, which is consistent with the error described above and... Figure 7 This corresponds to 703 shown.

[0070] More specifically, a simulation dataset is generated based on an off-the-shelf model of the male digestive system. A virtual colonoscope is placed inside the colon, and the observations are simulated. Unity3D (https: / / unity.com / ) is used to simulate and generate continuous 2D RGB images using a rigorous pinhole camera model. The frame rate and size of the simulated in vivo digestive tract (e.g., as shown in the image) are also considered. Figure 2 (As shown) is 30 frames per second and 640*480. The overall orientation of the colonoscope was also recorded.

[0071] As illustrated and as stated above, the red graph represents results with only classification (e.g., related technologies), while the blue graph represents results with pose refinement performed according to the example implementation. It can be seen that, generally speaking, the results with pose refinement have better accuracy in terms of both positional and angular differences.

[0072] More specifically, and as explained above, Figure 8 A comparison of positional difference errors relative to the keyframe ID is illustrated at 801, and a comparison of angular difference errors relative to the keyframe ID is illustrated at 803. Table 1 shows an error comparison between the related technology (i.e., ContextualNet) and the example implementation described herein.

[0073]

[0074] Table 1

[0075] The example implementation can be integrated with other sensors or solutions. For example, but not as a limitation, other sensors, such as inertial measurement units, temperature sensors, acidity sensors, or other sensors associated with sensing parameters related to the environment, can be integrated into the observation device.

[0076] Similarly, multiple sensors of a given type can be used; related technical solutions may not employ such multiple sensors; related techniques focus on providing precise location, which is the opposite of the approach described in this paper that uses labeling, feature extraction, bundle adjustment, and minimizing reprojection errors.

[0077] Because this example implementation does not require a high level of sensor or camera accuracy or additional training datasets, existing devices can be used with this example implementation to achieve more accurate results. Therefore, the need to upgrade hardware to obtain a more accurate camera or sensor can be reduced.

[0078] Furthermore, improved accuracy allows for the interchangeability of different types of cameras and observation devices, and enables different medical institutions to more easily exchange results and data, and to involve more and more diverse medical professionals without sacrificing accuracy, with the ability to properly analyze, make recommendations, and take action.

[0079] Figure 9 Example processing 900 is illustrated according to the example implementation. As described herein, example processing 900 can be performed on one or more devices.

[0080] At position 901, the neural network receives input and labels the training images. For example, but not limited to, as explained above, the training images can be generated from a simulation. Alternatively, historical data associated with one or more patients can be provided. The training data is used in the model, and pose regression is replaced with region classification. For example, but not limited to, images in adjacent poses can be classified as regions.

[0081] At position 903, feature extraction is performed. More specifically, the image is fed into the training database 117. Based on key features, a classification determination is made regarding whether the image's features can be classified as being in a specific pose.

[0082] At position 905, bundle adjustment is performed. More specifically, as described above, the predicted regions are used to triangulate the map points.

[0083] At position 907, an operation is performed to minimize the reprojection error of map points on the test image by adjusting the pose. Based on the result of this operation, the optimal pose is determined.

[0084] At position 909, an output is provided. For example, but not limited to, the output can be an image or an indication of a region or location within a region of an observational device associated with the image. Therefore, it can assist medical professionals in determining the location of an image within a target tissue such as the GI tract, lungs, or other tissues.

[0085] Figure 10An example computing environment 1000 is illustrated having an example computing device 1005 suitable for use in some example implementations. The computing device 1005 in the computing environment 1000 may include one or more processing units, cores, or processors 1010, memory 1015 (e.g., RAM, ROM, etc.), internal storage 1020 (e.g., magnetic storage, optical storage, solid-state storage, and / or organic storage), and / or I / O interfaces 1025, any of which may be coupled to a communication mechanism or bus 1030 to transmit information, or embedded in the computing device 1005.

[0086] According to this exemplary embodiment, the processing associated with neural activity can occur on the processor 1010, which serves as a central processing unit (CPU). Alternatively, other processors may be used instead without departing from the inventive concept. For example, but not as a limitation, a graphics processing unit (GPU) and / or a neural processing unit (NPU) may be used in place of or in combination with a CPU to perform the processing for the foregoing exemplary implementation.

[0087] The computing device 1005 can be communicatively coupled to the input / interface 1035 and the output device / interface 1040. One or both of the input / interface 1035 and the output device / interface 1040 can be wired or wireless and can be detachable. The input / interface 1035 may include any device, component, sensor, or physical or virtual interface that can be used to provide input (e.g., button, touchscreen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, optical reader, etc.).

[0088] Output device / interface 1040 may include a display, television, monitor, printer, speaker, Braille, etc. In some example implementations, input / interface 1035 (e.g., a user interface) and output device / interface 1040 may be embedded in or physically coupled to computing device 1005. In other example implementations, other computing devices may be used as or provide the functionality of input / interface 1035 and output device / interface 1040 of computing device 1005.

[0089] Examples of computing device 1005 may include, but are not limited to, highly mobile devices (e.g., devices in smartphones, vehicles and other machines, devices carried by people and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices designed for non-mobility (e.g., desktop computers, server devices, other computers, kiosks, televisions and / or radios, etc., in which one or more processors are embedded).

[0090] Computing device 1005 may be communicatively coupled (e.g., via I / O interface 1025) to external storage 1045 and network 1050 for communicating with any number of networked components, devices, and systems, including the same one or more computing devices or other configurations. Computing device 1005 or any connected computing device may be used as a server, client, thin server, general-purpose machine, special-purpose machine, or other name, providing services under the names of server, client, thin server, general-purpose machine, special-purpose machine, or other name, or referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or another name. For example, but not as a limitation, network 1050 may include a blockchain network and / or a cloud.

[0091] I / O interface 1025 may, but is not limited to, use any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX, modem, cellular network protocol, etc.) for transmitting information to or from at least all connected components, devices, and networks in computing environment 1000. Network 1050 may be any network or combination of networks (e.g., Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).

[0092] The computing device 1005 can be used and / or communicate using computer-usable or computer-readable media (including transient and non-transient media). Transient media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, etc. Non-transient media include magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROMs, digital video disks, Blu-ray discs), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage devices or memories.

[0093] The computing device 1005 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in certain example computing environments. The computer-executable instructions can be retrieved from a temporary medium, stored in a non-temporary medium, and retrieved from there. The executable instructions can originate from one or more of any programming, scripting, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0094] The processor 1010 can execute under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including a logic unit 1055, an application programming interface (API) unit 1060, an input unit 1065, an output unit 1070, a training unit 1075, a feature extraction unit 1080, a bundle adjustment unit 1085, and an inter-unit communication mechanism 1095 to enable different units to communicate with each other, with the operating system, and with other applications (not shown).

[0095] For example, training unit 1075, feature extraction unit 1080, and bundle adjustment unit 1085 can implement one or more of the processes described above with respect to the above structure. The described units and elements may vary in design, function, configuration, or implementation, and are not limited to the description provided.

[0096] In some example implementations, when information or execution instructions are received by API unit 1060, they can be passed to one or more other units (e.g., logic unit 1055, input unit 1065, training unit 1075, feature extraction unit 1080, and bundle adjustment unit 1085).

[0097] For example, as described above, training unit 1075 can receive and process information from simulation data, historical data, or one or more sensors. The output of training unit 1075 is provided to feature extraction unit 1080, which, based on the above description and, for example, in… Figure 1-5 The neural network shown is used to perform the necessary operations. Additionally, the bundle adjustment unit 1085 can perform operations based on the outputs of the training unit 1075 and the feature extraction unit 1080, and minimize the reprojection error to provide the output signal.

[0098] In some cases, in some of the example implementations described above, logic unit 1055 can be configured to facilitate information flow between control units and direct services provided by API unit 1060, input unit 1065, training unit 1075, feature extraction unit 1080, and bundle adjustment unit 1085. For example, the flow of one or more processes or implementations can be controlled by logic unit 1055 alone or in combination with API unit 1060.

[0099] Figure 11 An example environment suitable for certain example implementations is shown. Environment 1100 includes devices 1105-1145, and each device is communicatively connected to at least one other device via, for example, network 1150 (e.g., via wired and / or wireless connection). Some devices may be communicatively connected to one or more storage devices 1130 and 1145.

[0100] Examples of one or more devices 1105-1145 can be respectively Figure 10 The computing device 1005 described herein. Devices 1105-1145 may include, but are not limited to, a computer 1105 (e.g., a laptop computer device) having a monitor and associated webcam as described above, a mobile device 1110 (e.g., a smartphone or tablet), a television 1115, a device associated with a vehicle 1120, a server computer 1125, computing devices 1135-1140, and storage devices 1130 and 1145.

[0101] In some implementations, devices 1105-1120 can be considered user devices associated with a user, who can remotely obtain sensing inputs used as inputs in the example implementations described above. In this example implementation, one or more of these user devices 1105-1120 can be associated with one or more sensors capable of sensing the information required for this example implementation as described above, such as cameras temporarily or permanently embedded in the user's body, away from the patient care facility.

[0102] While the foregoing example implementations have been provided to indicate the scope of the invention, they are not intended to be limiting, and other methods or implementations may be substituted or added without departing from the scope of the invention. For example, but not as a limitation, other image technologies besides those disclosed herein may be used.

[0103] According to one example implementation, algorithms such as SuperPoint can be used to train image point detection and determination. Furthermore, the example implementation can employ alternative image classification algorithms and / or use other neural network architectures (e.g., Siamese networks). Additional schemes incorporate expert knowledge into region-classing actions by applying enhancements to two images using techniques such as forming, illumination, and emission, and / or using a single image-to-depth approach.

[0104] The example implementation may have various advantages and benefits, although this is not required. For example, but not as a limitation, the example implementation can work on small datasets. Furthermore, the example implementation provides constraints on location within target tissues such as the colon or lungs. Therefore, surgeons can be able to more accurately locate the position of any person's viewing device by using video. Moreover, the example implementation provides significantly higher accuracy than related technical methods.

[0105] While some example implementations have been shown and described, these example implementations are provided to convey the subject matter described herein to those skilled in the art. It should be understood that the subject matter described herein can be implemented in various forms and is not limited to the example implementations described. The subject matter described herein can be practiced without those specifically defined or described subjects, or without other or different elements or subjects described. Those skilled in the art will understand that changes can be made to these example implementations without departing from the subject matter described herein as defined in the appended claims and their equivalents.

[0106] Some non-limiting embodiments of this disclosure address the features discussed above and / or other features not described above. However, some non-limiting embodiments do not need to address the aforementioned features, and some non-limiting embodiments of this disclosure may not address the aforementioned features.

Claims

1. A computer-implemented method, the method comprising the steps of: applying training images of an environment divided into regions to a neural network and performing a classification to label a test image based on a closest region of the regions; extracting features from pose information of a retrieved training image matching the closest region and the test image; performing bundle adjustment on the extracted features by triangulating map points of the closest region to generate a re-projection error and minimizing the re-projection error to determine a best pose of the test image; and for the best pose, providing an output indicating a position or a probability of a position of the test image in the environment in the best pose.

2. The computer-implemented method of claim 1, wherein, applying the training images comprises receiving training images associated with poses in regions of the environment as historical or simulated data and providing the received training images to a neural network.

3. The computer-implemented method of claim 2, wherein, the neural network is a deep learning neural network that learns regions associated with the poses and determines the closest region for the test image.

4. The computer-implemented method of claim 1, wherein, the bundle adjustment comprises re-projecting 3D points associated with the measured pose and triangulated map points into 2D image space to generate a result and comparing the result to a registered 2D observation to determine the re-projection error.

5. The computer-implemented method of claim 4, wherein, for a re-projection error below or equal to a threshold, a pose of the test image is confirmed as the best pose.

6. The computer-implemented method of claim 4, wherein, for a re-projection error above a threshold, a pose of the test image is determined to be incorrect and a computation of the pose of the test image is determined to be correct.

7. The computer-implemented method of claim 1, wherein, minimizing the re-projection error comprises adjusting a pose of the test image to minimize the re-projection error.

8. A non-transitory computer readable medium having a storage device storing instructions for execution by a processor, the instructions comprising: applying training images of an environment divided into regions to a neural network and performing a classification to label a test image based on a closest region of the regions; extracting features from pose information of a retrieved training image matching the closest region and the test image; performing bundle adjustment on the extracted features by triangulating map points of the closest region to generate a re-projection error and minimizing the re-projection error to determine a best pose of the test image; and for the best pose, providing an output indicating a position or a probability of a position of the test image in the environment in the best pose.

9. The non-transitory computer-readable medium of claim 8, wherein, applying the training images comprises receiving training images associated with poses in regions of the environment as historical or simulated data and providing the received training images to a neural network.

10. The non-transitory computer-readable medium of claim 9, wherein, the neural network is a deep learning neural network that learns regions associated with the poses and determines the closest region for the test image.

11. The non-transitory computer-readable medium of claim 8, wherein, The bundle adjustment includes reprojecting 3D points associated with the measured pose and triangulated map points into 2D image space to generate results, and comparing the results to registered 2D observations to determine the re-projection error.

12. The non-transitory computer-readable medium of claim 11, wherein, For re-projection errors below or equal to a threshold, the pose of the test image is confirmed as the best pose.

13. The non-transitory computer-readable medium of claim 11, wherein, For re-projection errors above a threshold, the pose of the test image is determined to be incorrect and the computation of the pose of the test image is determined to be correct.

14. The non-transitory computer-readable medium of claim 8, wherein, Minimizing the re-projection error includes adjusting the pose of the test image to minimize the re-projection error.

15. A computer-implemented system for localizing and tracking a viewing device in an environment to identify a target, the computer-implemented system comprising: a memory configured to store a program; and a processor communicatively coupled to the memory and configured to execute the program to perform the following operations: applying a training image of the environment associated with the viewing device divided into regions to a neural network and performing a classification to label a test image generated by the viewing device based on a closest region of the regions of the environment associated with the viewing device; extracting features from pose information of the test image and a retrieved training image matching the closest region; performing bundle adjustment on the extracted features by triangulating map points of the closest region to generate a re-projection error, and minimizing the re-projection error to determine a best pose of the test image; and for the best pose, providing an output indicating a location or a probability of location within the environment of the test image generated by the viewing device in the best pose.

16. The computer-implemented system of claim 15, wherein, The environment includes a gastrointestinal tract, or a bronchial tract of one or more lungs.

17. The computer-implemented system of claim 15, further comprising the observation device, wherein, The viewing device is configured to provide a location of one or more targets including at least one of a polyp, a lesion, and a cancerous tissue.

18. The computer-implemented system of claim 15, further comprising the observation device, wherein, The viewing device includes one or more sensors configured to receive the test image associated with the environment, and the test image is a visual image.

19. The computer-implemented system of claim 15, further comprising the observation device, wherein, The viewing device is an endoscope or a bronchoscope.

20. The computer-implemented system of claim 15, wherein, The environment is a piping system, an underground environment, or an industrial facility.

Citation Information

Patent Citations

  • Network training method, incremental mapping method, positioning method, device and equipment

    CN109658445A

  • A weak texture three-dimensional object attitude estimation method and device

    CN109934847A