Markerless pose estimation of medical devices using a single camera
A single-camera system for medical device tracking segments and extracts features to determine the orientation of ultrasonic transducers, addressing cost and precision issues in existing methods by using depth information and occlusion modeling.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2026-03-25
AI Technical Summary
Existing medical device tracking technologies, such as stereo-based optical tracking and electromagnetic tracking, are costly, prone to errors due to occlusions, and lack precision, especially in medical environments where markers can be blocked or the device is shielded.
A markerless pose estimation system using a single camera to detect and determine the orientation of medical devices like ultrasonic transducers by segmenting and extracting features from images, utilizing a single camera to capture depth information and comparing with pre-defined templates, while accounting for occlusions by modeling the shielding objects.
Reduces hardware and calibration costs, enhances tracking accuracy by eliminating the need for additional markers and cameras, and effectively handles occlusions, enabling precise posture estimation of medical devices during use.
Smart Images

Figure 2026053264000001_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the posture estimation of medical devices such as ultrasonic transducers.
Background Art
[0002] The tracking of medical devices is a major component of modern scanning and / or surgical procedures. Tracking the device synchronizes the position of the device with respect to the anatomical structure and reduces the possibility of errors.
[0003] Stereo-based optical tracking can be used. Two or more cameras observe the operating scene. To identify the device in the stereo view, specially designed markers (e.g., infrared reflectors or light-emitting diodes) are attached to the body of the device, such as a removable housing. Adding markers or obtaining devices with pre-designed markers can be expensive. The markers can be blocked by the user's hand, arm, etc. intervening between the marker and one or both cameras.
[0004] Electromagnetic tracking uses Faraday's law to detect the movement of the device, but a receiver module must be added to the device being tracked. The device can be tracked even if it is shielded, but the tracking is not as precise as in the optical solution.
[0005] Markerless optical tracking has been proposed. No prior information is used for tracking. General objects (objects) such as a pitcher or a drill are tracked from the video stream by segmenting the moving object. The posture is based on the extracted features using graph-based alignment. This approach can function with general objects, while rare shapes and occlusions seen in the medical environment can lead to incorrect tracking, especially when significant rotation of the device occurs.
[0006] Another common approach involves estimating the pose of previously unseen objects, as used in robotics. Prior object information, in the form of a CAD model or other 3D object representation, is used in combination with image properties, and a view of the model is rendered in the scene. The rendered object is then compared to another detected object to see if there is a sufficiently close match. Based on this match, the object's pose can be estimated. This method may not adequately handle occlusion. [Overview of the Initiative]
[0007] Regarding posture estimation, a non-temporary computer-readable medium containing a system, method, and stored instructions (program code) is provided. The posture of a medical device in use, which may be occluded, is estimated using a single camera. Multiple cameras and / or markers are not required. The medical device is detected in the camera image during use to determine the posture. Various further approaches may be used to address occlusion. Comparison with templates for the applicable type of medical device may be used to address occlusion. Modeling occlusion, such as hands, may be used to address occlusion.
[0008] In a first embodiment, a method for estimating the pose of an ultrasonic transducer is provided. A single camera acquires an image of the ultrasonic transducer. The ultrasonic transducer captured in the image is held and occluded by a user. The ultrasonic transducer is detected in the image. The pose of the ultrasonic transducer detected in the image from the single camera is determined. Ultrasonic imaging using the ultrasonic transducer has scan data aligned by the determined pose.
[0009] In a second embodiment, an ultrasonic system is provided. A single camera is configured to image an acoustic transducer that is partially shielded by the user. An image processor is configured to determine the orientation of the acoustic transducer from the image taken by the single camera. This determination is based on a transducer template.
[0010] In a third embodiment, a method for estimating the posture of a medical device is provided. A single camera captures an image of the medical device while it is being used on or in a patient. In the image, the medical device is being held and obscured by the user. The medical device is detected in the image, along with a portion of the user. The posture of the medical device detected in the image by the single camera is determined using the portion of the user detected in the image.
[0011] Any one or more of the embodiments or concepts outlined above or in the embodiments illustrated below may be used individually or in combination. An embodiment or concept described in one exemplary embodiment or aspect may also be used in other embodiments or aspects. An embodiment or concept described in a method or system may also be used in other systems, methods, or non-temporary computer-readable storage media.
[0012] The present invention is defined by the claims, and nothing in this section should be construed as a limitation to the claims. Further aspects and advantages of the present invention are described below in connection with preferred embodiments and may be claimed individually or in combination thereafter. [Brief explanation of the drawing]
[0013] The components and drawings are not necessarily to scale, but rather emphasized to illustrate the principles of the embodiments. Furthermore, the same reference numerals in the drawings indicate corresponding parts throughout the drawings. [Figure 1] A flowchart illustrating one embodiment of a method for estimating the orientation of a medical device. [Figure 2] An example of an ultrasonic transducer is shown. [Figure 3] Figure 2 shows an example of an ultrasonic transducer shielded by the user's finger or hand. [Figure 4] A block diagram of one embodiment of a system for determining the orientation of medical devices. [Modes for carrying out the invention]
[0014] A single camera is used for pose estimation without markers on the imaged object, avoiding the cost and complexity associated with additional (extra) cameras and markers. In one embodiment, a color depth camera (e.g., an RGBD camera) observes a target medical object in a medical environment. A representation of the target object is generated from the image (e.g., RGBD data). The object itself is not augmented with any additional identification markers, and the object is identified as it appears (as it looks) and used in the surgical environment. This approach is directed to a set of known objects limited to a set of tools typically used for medical procedures. Once an object is identified in the scene, key features are extracted from the observed object. These features represent important properties of the object that can be used for matching against a ground truth template. Since the poses of the set of known objects are detected, prior ground truth geometric information about the objects is known. This may be in the form of a CAD file or other object representation. Key features are extracted from the ground truth object representation, and matching is determined against the key features from the detected object. Once the matching determination is complete, the object's orientation relative to the camera is calculated.
[0015] A single camera performs object tracking without additional markers. Using a single camera reduces the amount of physical hardware required to build the tracking system and the need for complex calibration techniques to calibrate multiple cameras. The relaxed hardware requirements lower the overall system cost. Markerless detection eliminates the burden of attaching dedicated tracking hardware to all tools the operator may need during a procedure. Adding markers increases procedure time and can lead to errors if the markers are not properly attached. A markerless system allows the operator to use the tools as they are. A single camera can determine the pose with only a one-time load of an object library (template collection), which can be performed during system manufacturing or a software update.
[0016] Shielding is reliably controlled by considering the shape geometry of the medical device along with other objects that may shield it. In the case of ultrasound scanning, the hand is most likely to shield a significant portion of the probe body. To accurately determine the posture, the hand and its interaction with the probe are modeled. Shielding can result from the hand, one or more other objects, or the overlap of multiple objects. These objects can be modeled.
[0017] Figure 1 is a flowchart of one embodiment of a method for estimating the orientation of a medical device, such as an ultrasonic transducer. Landmark detection is used to identify the orientation captured by a single camera. Template matching and / or modeling of occluding objects may be included. Object segmentation may be used to extract features (feature quantities) from a mask (segmentation).
[0018] The orientation of any medical device can be specified. For example, the orientation can be specified for surgical instruments, endoscopes, or other medical tools held in the hand of a physician or nurse. In the example shown here, the medical device is an ultrasonic transducer, such as a handheld probe, used externally or internally on a patient. The user holds the handle of the ultrasonic transducer, and the camera captures images of the handle and / or other parts of the ultrasonic transducer.
[0019] The method in Figure 1 is performed by a medical imaging system, such as an image processor, camera, and / or scanner. The camera captures the scene. The image processor identifies the posture from the captured scene. Memory may store one or more templates. The image processor may be part of a medical scanner, such as an ultrasound system, or it may be a separate computer, workstation, or processor. The image processor, medical scanner, physician, or procedure may use the identified posture. In one embodiment, the system in Figure 4 performs the method, but other systems may also be used.
[0020] Additional or different processes may be provided, or processes may be reduced. For example, processes 110 and 120 may be merged or performed as a single process. In another example, processes 130, 140, and / or 150 may not be performed. In yet another example, a process for configuring a scanner to scan or a process for other uses of a medical device may be provided.
[0021] The processes are executed in the order shown in the diagram (from top to bottom or in numerical order) or in any other order. For example, process 140 may be executed as part of process 110, process 120, process 130, or independently before or after process 120.
[0022] In process 100, a single camera captures an image of a medical device (e.g., an ultrasonic transducer). The medical device is shielded in the image by being held by a user or an arm (e.g., a robotic arm). As an example, FIG. 2 shows a handheld transducer 200 when not being held. FIG. 3 shows a handheld transducer 200 with at least a part of the transducer 200 being shielded by a user's hand 310 (e.g., a finger). The image captured by the camera is of the ultrasonic transducer 200 while the ultrasonic transducer 200 is being used on or within a patient (e.g., in FIG. 3). The ultrasonic transducer 200 is being moved by the user while being held in order to scan the patient. A hand, finger, palm, arm, torso, head, and / or other parts of the user shield a part of the ultrasonic transducer 200 from the camera 300. As a result, a part of the ultrasonic transducer 200 captured within the image is shielded by the user.
[0023] The camera captures an image of a scene within the camera's field of view. The camera is positioned to capture an ultrasonic transducer within the scene while being used on or towards a patient. One or more images are captured as still images or videos, etc.
[0024] The camera captures an image as a two-dimensional distribution of pixels, for example, by obtaining red, green, blue (RGB) values. Other information such as depth information can be captured as an image or as part of an image. Thermal, infrared, and other images can also be captured.
[0025] In one embodiment, the camera is a depth camera such as a 2.5D or RGBD (RGB + depth) camera (e.g., Microsoft Kinect 2 or ASUS Xtion Pro). One sensor can capture a color image and another sensor of the depth camera can capture depth. The depth camera can directly measure depth, such as by using time-of-flight, interferometry, or coded aperture. A depth camera (e.g., an RGBD camera) outputs a color point cloud where each pixel has color and depth. Other object representations can also be used. While one way is to construct a point cloud, the color and depth images can be directly used to avoid reconstructing the point cloud. Other optical or non-ionizing sensors such as LIDAR cameras can be used.
[0026] Only one camera is used. For example, a single depth camera is used. If the depth camera uses separate color and depth sensors within one camera device, these separate data are combined to generate a representative point cloud or representation of the target device.
[0027] The camera can be placed in various indoor locations, including attachment or connection to a medical scanner, wall, ceiling, or headset. The camera is placed on a wall, ceiling, or in an imaging suite (dedicated imaging room) or other locations in an operating room, such as on a boom that is generally above the patient.
[0028] The camera is directed at the patient and / or the area where an ultrasonic transducer is used on or within the patient. The field of view of the camera covers the area where the ultrasonic transducer is used. The camera captures the ultrasonic transducer, the patient, and / or the outer surface of the user from one viewpoint. The camera is mounted to limit or avoid occlusion and / or to maximize the visualization of features used to identify the pose.
[0029] The image is used at the sensor's resolution. For example, the image or point cloud is 256 x 256 pixels. Other sizes may also be used, including a rectangular field of view. The image may be filtered and / or processed. For example, the image may be changed to a predetermined resolution. Another example is downsampling, such as reducing 256 x 256 to 64 x 64 pixels. The image may be cropped, such as limiting the field of view.
[0030] Images are captured within a predetermined time frame. By using continuous or predefined frequency capture, a stream of images (i.e., video) from different points in time can be captured. The predefined frequency may be preset, adjustable, or variable. By capturing images over time, the images constitute a video of the ultrasound transducer in use on or in the patient. In an alternative embodiment, images are captured at only one point in time.
[0031] The image shows an ultrasound transducer without any additional optical markers. No shapes other than the target, color code, or the shape of the probe itself are added. The ultrasound transducer itself is not extended with any additional identification markers for orientation identification. Trademarks and brand markers may be included, but no markers are added for orientation identification. The ultrasound transducer should be identified as it appears (as is) or as an ultrasound transducer shipped from the manufacturer and used in the surgical environment. Since triangulation for stereoscopic vision is not used, additional markers are unnecessary. Markers may be added for other purposes or may remain after use in other orientation identification approaches, but such markers are not necessary.
[0032] In process 110, the image processor detects ultrasonic transducers in the image. For example, a portion or pixel of a color point cloud belonging to or representing an ultrasonic transducer is detected. The ultrasonic transducer is segmented in or from the image. The location or representation of the ultrasonic transducer is distinguished from other objects in the image. Pixels belonging solely to ultrasonic transducers are separated or labeled.
[0033] Any segmentation method may be used. For example, random walker, thresholding, domain expansion, or other hand-coded segmentation methods may be applied. Another example is a machine learning classifier, such as a deep learning neural network or other machine learning model, which segments (outputs segments in response to image input). This machine learning classifier can be trained using image or video examples captured during a patient scan, and ground truth is manually identified for that image or video or identified by another process. Synthetic data may also be used as training examples. During model training, a CAD model is collected along with RGBD data of transducers that have been collected in the real world beforehand. In contrast to this data, simulated RGBD data for training is created by importing the CAD model into a simulation environment and generating a synthetic scene in which the transducer is scanning a patient. This allows for the generation of a large amount of synthetic data with minimal effort. The model machine training process may use real-world data and / or synthetic data to learn object shapes.
[0034] Other objects may also be segmented. For example, a user or a part of a user (e.g., a hand) may be segmented. A user represented in the image is detected, and locations belonging to the user are distinguished from other locations. The same or different algorithms or machine learning classifiers may be used to segment different objects, such as objects that obstruct the ultrasound transducer. Other medical devices used in a procedure, such as surgical tools, may be detected in addition to the ultrasound transducer. A predetermined number of segmentations may be performed on a given image, such as segmenting for any medical device within a possible set of medical devices.
[0035] Some filtering and / or smoothing can remove noise data before or after segmentation. Segmentation generates a representation of the ultrasonic transducer, with or without generating representations of other objects.
[0036] Three-dimensional reconstruction may be used. Camera characteristics or parameters are used to reconstruct a three-dimensional object or surface for the ultrasonic transducer and / or other objects (e.g., a hand). Detection by segmentation and 3D reconstruction generates a representation of the ultrasonic transducer from the camera viewpoint for further analysis. Similarly, a representation of the hand or other occluding object from the camera viewpoint is generated. Part of the user (e.g., occluding part) is detected in the image. Segmentation may be used without further three-dimensional reconstruction. By using a depth camera, segmentation represents the object as a three-dimensional surface.
[0037] In process 120, the image processor extracts multiple features of the detected ultrasonic transducer and / or other objects in the image. Landmarks are detected. For example, keypoints or landmarks of the ultrasonic transducer are detected from the segmented representation of the ultrasonic transducer. Once the ultrasonic transducer is identified in the scene, key features are extracted from the observed object. These features represent important characteristics of the object that can be used to determine a match against a ground truth template.
[0038] Feature extraction can be performed by processes such as template matching or edge detection, or by applying a machine learning model. In the case of a machine learning model, the segmented or detected representations are input to the model. Alternatively or additionally, information derived from the segmented representations is input. The machine learning model is trained to output one or more landmarks in response to the input. Neural networks, support vector machines, or Bayesian inference models may be used.
[0039] Any landmark can be extracted as a feature. This extraction detects the location of the landmark within the representation. This extraction may be the extraction of semantic features, such as by using a semantic feature extractor. A semantic feature extractor discovers features that humans would inherently interpret as important about an object. For example, the corners of a cube and the centers of each face of a cube can be considered semantic features. Figure 2 illustrates an ultrasonic transducer 200. The transducer 200 includes an elongated region 210 that houses an array and a gripping region 220 having a grip 230. A cable 240 connects to the gripping region 220. Examples of semantic features include each edge of the grip 230, and each corner and / or edge of region 210 and / or region 220. The contact points between the cable 240 and region 220 can also be another semantic feature. Additional or different semantic features may be provided, or there may be fewer semantic features.
[0040] The extraction may involve the extraction of general or pre-trained features, such as by using a general-purpose feature extractor. In one approach, a general-purpose feature extractor is a neural network, such as an encoder, that outputs feature vectors based on machine training. Feature vectors are abstract information generated by deep learning. General-purpose feature extractors can capture semantic features, but they can also capture less obvious features (e.g., general (approximate) cues). This may include features arising from lighting or surface texture, as well as subtle cues about the shape of an object. The detected features are invariant with respect to the viewpoint from which the camera observes the object. Either one or both types of feature extractors may be used. Other feature extractors or landmark detection may also be used. The same feature extractor or class of feature extractors is applied to the observed or detected ultrasonic transducer and ground truth template.
[0041] In other approaches, an image processor extracts features as part of the detection process. Processes 110 and 120 are merged.
[0042] It is possible that some features may not be extracted. For example, the user may occlude one or more features. In the example in Figure 3, the user's hand 310 occludes several semantic or generic features associated with a large portion of the grip 230. Similarly, any features on the opposite side of the ultrasonic transducer 200 from the camera 300's perspective will not appear in the image and therefore will not be extracted.
[0043] In process 130, the image processor determines the orientation of the ultrasonic transducer. The image processor determines the position, orientation, scale (size), and / or combination of two or more of these of the ultrasonic transducer relative to the camera.
[0044] Segmented ultrasonic transducers provide a point cloud or spatial position of the ultrasonic transducer relative to the camera. Extracted features based on image location provide the spatial position of the ultrasonic transducer relative to the camera. Ultrasonic transducers detected within an image from a single camera are positioned relative to the camera.
[0045] Due to occlusion, errors can occur if the orientation is determined solely by segmented objects or extracted features. For more accurate orientation determination, the image processor compares the detected ultrasonic transducer (e.g., segmentation and / or extracted features) to an ultrasonic transducer template. Various relative orientations of the template are compared to the detected ultrasonic transducer. The best-matching orientation is identified as the orientation of the detected ultrasonic transducer. Since the orientations of known types or known objects are identified, prior ground truth geometric information about the object may be used. A set of known objects and corresponding templates may be provided. The set of objects is limited to a set of tools used in a list of typical procedures. If the specific ultrasonic transducer being used is unknown, various templates corresponding to various types of transducers may be examined to find the one with the best correlation (e.g., minimizing the difference in the best orientation or minimizing the square of the difference). Category-based models of transducers may also be used. For example, a set of 10 radial transducers is provided that represents a sample of possible variations in transducer shape. The model understands possible shape variations and knows what radial transducers generally look like, so it can accurately estimate the pose of a radial transducer that is not identical to one of the 10 transducers in the training set. A “basic model” is provided because the model includes fundamental knowledge of transducers and their possible shape variations.
[0046] One or more templates are computer-aided design (CAD) files, lists of spatially related features, encoded feature vectors, or other object representations. The object in the template may have its orientation modified for comparison. Alternatively, different templates are provided for different orientations.
[0047] In one embodiment, key features (e.g., semantic and / or generalized) are extracted from the ground truth object representation and matched against key features from the detected ultrasonic transducer. Once the match determination is complete, the object's orientation relative to the camera can be calculated or identified. The image processor examines various orientations of the ultrasonic transducer template against the detected ultrasonic transducer. This examination involves a comparison of segmentation and / or extracted features.
[0048] In another embodiment, feature vectors output by a machine learning network (e.g., an encoder) are compared. Detected ultrasonic transducers and / or extracted features are input to a machine learning network, which outputs feature vectors in response to the input. A template of the pose to be examined and / or extracted features are input to a machine learning network, which outputs feature vectors in response to the input. Alternatively, a pre-processed table of feature vectors serves as the template. Feature vectors from the detected ultrasonic transducers and the template are compared. Various poses are examined to identify the pose with the best match in the feature vectors.
[0049] If there is a match, the posture is identified. Investigating different postures leads to the posture with the smallest difference in the investigation. Based on calibration, the posture relative to the camera is converted to the posture relative to the patient, ultrasound scanner, global or intra-intra
[0050] A template may be specific to the detected ultrasound transducer, such as being of the same model or type. In another embodiment, a template is generalized or standardized to represent a range of ultrasound transducers of multiple types or models. Multiple templates may be provided for multiple styles or common transducer types, e.g., templates for transducers of different sizes, different array configurations, different purposes (e.g., via esophageal vs. via abdominal scanning), and / or other differences. Comparison for determining posture uses a generalized template or transducer that is not specific to the detected ultrasound transducer. Sufficient similarity in the extracted features enables posture determination.
[0051] Generally, since the orientation is specified for each known object, the feature extractor in process 120 generates a large number (e.g., tens, hundreds, or thousands) of prominent features with respect to the ground truth object template. This large number provides a sufficient number of features for a match determination. A detected ultrasonic transducer may have less than half or even less than one-tenth of the features compared to the template due to occlusion and the field of view from a single camera. When observing an object in a real-world scenario, it is impossible to see the entire appearance from a single camera. The features extracted from the view of the object are a subset of the entire ground truth features. The endpoint of the match determination process is to match this subset with the correct features in the template.
[0052] In one embodiment, outlier removal is used for comparison. Gaussian filtering, principal component analysis, RANSAC, or other outlier removal is used to match a large number of extracted features from the template to a smaller number of extracted features from the detected ultrasonic transducer. Minimization for determining pose also selects the extracted features from the template to use for comparison. Having many features in both sets helps to generate a more accurate solution for performing a match determination. An outlier removal method such as RANSAC is used by the image processor to compute a geometric transformation between the object and ground truth in order to estimate the pose.
[0053] In process 140, the image processor models one or more occluding objects. These occluding objects may be other medical devices, beds, patients, robotic arms, support arms, sterile barriers, or other objects in the medical environment. For example, the user's hand is modeled. As shown in Figure 3, the user's hand 310 occludes a portion of the ultrasonic transducer 200 during use.
[0054] Handling occlusion of the target object (e.g., an ultrasound transducer) is a significant challenge in pose estimation. Occlusion can cause the pose to be determined without complete information about the object. In the case of medical devices and tools, occlusion is widespread in many procedures, ranging from partial to severe. Occlusion can result from the object being partially inserted into the patient or from the object's body being obscured by the operator's hand. The latter scenario is very common in scanning procedures such as ultrasound, where the camera observing the scan may not obtain an unobstructed view of the probe due to the operator's hand obstructing the object. This type of occlusion is considered in pose estimation by utilizing information about probe occlusion (e.g., a hand). Since a known object (e.g., an ultrasound transducer) is detected while being occluded by a known general shape (e.g., a hand), information from both can be used to correctly estimate the object's pose.
[0055] The model or information may be a template, a statistical geometry model, a physically based model, or a bioengineering model. The model is fitted to the occluding object. Any fitting is available, for example, performing process 110 detection, process 120 feature extraction, and process 130 pose determination for the occluding object. The occluding object is detected, and the model is fitted to the detected object. Scaling processes using rigid body transformations or affine transformations may also be provided.
[0056] This fitting may indicate the orientation of the shielding object, the position of the shielded ultrasonic transducer, and / or how the shielding object is held or positioned in relation to the ultrasonic transducer. This information can be used to determine the orientation of the ultrasonic transducer. Modeling of the shielding object may be used in the detection of the ultrasonic transducer (i.e., process 110), the extraction of ultrasonic transducer features (i.e., process 120), and / or the determination of the ultrasonic transducer's orientation (i.e., process 130). By incorporating a model of the shielding object in any of processes 110, 120, and / or 130, the detected shielding object is used in determining the orientation of the ultrasonic transducer. In the example in Figure 3, a part of the user (i.e., the hand) detected from the image is used in orientation determination.
[0057] In one embodiment, the image processor performs detection of process 110 in the individual or combined classification of both objects. Combined classification is superior in distinguishing the ultrasonic transducer from the hand and can improve the accuracy of segmentation and final orientation determination.
[0058] In another embodiment, features extracted in process 120 that may belong to the hand rather than the ultrasonic transducer are excluded. For example, any features located at the hand's position are excluded even if they could also be identified as features of the ultrasonic transducer.
[0059] In yet another embodiment, information from a modeled hand, including its orientation, gripping style, and other details, is used to indicate the orientation of the ultrasonic transducer in process 130. For example, the gripping hand 310, as shown in Figure 3, indicates the orientation of the ultrasonic transducer 200. The hand orientation may be used to initialize orientation search or identification of the ultrasonic transducer. The identified transducer orientation can be confirmed by comparison with the hand orientation. The ultrasonic transducer and hand orientations can be averaged. As another example, determining the orientation of the ultrasonic transducer may consider both the ultrasonic transducer in the image and the user's portion in the image. Minimization is based on errors or differences in fitting or matching both the ultrasonic transducer and the hand. Other uses of occluding object modeling may also be used.
[0060] In process 150, the image processor uses a specified posture. This posture may be used to position the graphical representation of the ultrasound transducer (or other medical device) for preoperative imaging to assist the physician in guiding the device to the patient. The posture may be used to adjust the position relative to the patient during scanning, treatment, biopsy, or diagnosis. The posture may be used by a robot or a user (e.g., a physician or ultrasound technician). Other applications of the specified posture in a medical setting may also be provided.
[0061] In one embodiment, posture is used for ultrasound imaging. A one-dimensional array scans a region within the patient. Different regions can be scanned to scan a volume by moving the ultrasound transducer. Posture is used to determine the alignment of each region (i.e., the alignment of the ultrasound transducer for each two-dimensional scan) to assemble a volume representation. Scan data is aligned according to the specified posture. A three-dimensional representation is rendered from the data of the three-dimensional representation aligned based on posture. This alignment may be used alternatively or additionally to indicate the scanning position relative to the patient and / or preoperative imaging.
[0062] Processes 100-150 can be performed once per image. In other approaches, process 100 captures a video or a series of images. Processes 110, 120, and 130 are performed repeatedly over time, for each image in a series, for example.
[0063] This iteration identifies the pose independently of each other. Alternatively, poses from earlier or different points in time are used to assist in identifying poses at later or different points in time. For example, an earlier pose is used to initialize comparisons for identifying poses later. Minimization (e.g., minimization using RANSAC) can reduce the number of iterations if the initial pose is closer to the actual pose. Collecting video of the object being tracked helps refine the pose estimation of the object. For example, after feature extraction and matching (template matching for pose identification) are performed on the first frame, the generated information is cached for use on one or more subsequent frames. Based on the camera's frame rate, the object may move only slightly between frames. Utilizing pose estimations from one or more preceding frames provides the matching algorithm with an initial pose estimate, which then refines that initial estimate into a new estimate. Time-based pose estimation and refinement help maintain a stable pose estimate throughout the tracking period and accelerate the processing time for generating pose estimates.
[0064] This process can be repeated or performed for one or more objects, such as any object from the expected object and the corresponding template set. The orientation is specified for a single object (e.g., an ultrasonic transducer) or for each of multiple objects. Any number of shielding objects can be modeled to address shielding.
[0065] Figure 4 shows one embodiment of an ultrasonic system. This ultrasonic system is either part of an ultrasonic scanner or a separate computer from the ultrasonic scanner. In other embodiments, the system is for other medical imaging (non-ultrasonic), therapeutic, or surgical procedures. A computer, either without a scanner or as part of a scanner, is provided to determine the orientation of the medical device.
[0066] The ultrasound system includes an image processor 400, memory 410, and a display 420. The ultrasound system also includes a camera 300 for detecting (imaging) a medical device (e.g., a transducer 200) and / or a transducer 200 for scanning a patient 430. The display 420, image processor 400, and memory 410 may be part of a medical imaging or treatment system or may be a computer, server, workstation, or other system communicatively connected to a medical imaging or treatment system for image processing.
[0067] Additional or different components may be provided, or there may be fewer components. For example, a computer network may be included for remote image processing and / or display. In another example, one or more machine learning models are stored in memory 410 and applied by processor 400 for detection, segmentation, extraction, and / or pose determination. In yet another example, a beamformer, scan converter, and / or one or more filters are provided for ultrasound imaging.
[0068] Camera 300 is a single camera configured to image the acoustic transducer 200. While other cameras may be present in the room for other purposes, camera 300 is the only camera used for attitude determination.
[0069] Camera 300 is a depth sensor, optical camera, 3D camera, infrared camera, thermal camera, and / or other type of camera. LiDAR, 2.5D, color depth (RGBD), or other depth cameras may be used. Camera 300 may include a separate processor for determining depth measurements from the image and / or detecting objects represented in the image, or an image processor 400 may determine depth measurements and / or detect objects from the image captured by camera 300. Camera 300 may directly measure the depth from camera 300 to the patient. The depth may be relative to camera 300 and / or the bed or table 440. Alternatively, a camera without depth sensing may be used. An optical projector may be provided.
[0070] The camera 300 is directed towards the patient 430 and / or the acoustic transducer 200. The camera 300 may be part of or connected to an ultrasonic scanner. In one embodiment, the camera 300 is positioned on a boom, robotic arm, ceiling, and / or wall. The field of view of the positioned camera 300 includes the area of the patient 430 in which the transducer 200 is to be used.
[0071] Camera 300 is calibrated to transducer 200, patient 430, bed 440, ultrasound system, or other coordinate system. Known relationships of camera 300 to other devices or rooms, and known camera parameters, allow the orientation specified for the camera to be translated to the orientation for other equipment or rooms.
[0072] The acoustic transducer 200 is a probe for scanning. The acoustic transducer 200 is part of the scanner that scans the patient. For example, a beamformer uses the acoustic transducer 200 to scan patient 430 with ultrasound for treatment and / or imaging. Other medical devices that interact with the patient in ways other than imaging or scanning may be used instead.
[0073] The image processor 400 is a control processor (e.g., a controller), a general-purpose processor, a digital signal processor, a three-dimensional data processor, a graphics processing unit, an application-specific integrated circuit, a field-programmable gate array, an artificial intelligence processor, a digital circuit, an analog circuit, a combination thereof, or any other currently known or future-developed device capable of performing image processing to determine pose. The image processor 400 is a single device, multiple devices, or a network. In the case of multiple devices, parallel or sequential partitioning of processing may be used. Each device constituting the image processor 400 may perform different functions, such as detecting an object in an image by one device and determining its pose by another device. In one embodiment, the image processor 400 is a control processor or other processor of a medical scanner. The image processor 400 operates according to stored instructions, hardware, and / or firmware and is configured to perform the various processes described herein by the stored instructions, hardware, and / or firmware.
[0074] In one embodiment, the image processor 400 is configured to identify the orientation of an acoustic transducer 200 from an image taken by a single camera 300. The acoustic transducer 200 represented in the image from the camera 300 is detected, for example, by applying a machine learning classifier. The orientation is identified based on a transducer template. By determining whether the transducer 200 detected in the image matches the template at different orientations, scales, and / or positions, the orientation of the detected transducer is identified at the orientation, scale, and / or position that best matches (e.g., has the smallest difference).
[0075] In one approach, the image processor 400 is configured to extract features of an acoustic transducer from an image or segmentation. A machine learning model, template (e.g., a statistical shape model), or other image processing is applied to identify the location of the semantic and / or generalized features of the acoustic transducer 200 in the image. The image processor 400 then determines the pose based on the extracted features relative to the template features.
[0076] The template used may be specific to the acoustic transducer 200 (e.g., a template for a transducer of the same model). Alternatively, the template may be a generalized transducer that can be applied to different transducer models. The generalized transducer of the template is not specific to the acoustic transducer 200 captured in the image. The image processor 400 is configured to use extracted features and the generalized template to determine the pose of the acoustic transducer 400.
[0077] Display 420 is a CRT, LCD, projector, plasma, printer, tablet, smartphone, or other display device currently known or to be developed in the future for displaying preoperative images with a graphical representation of captured images, postures, medical images, or an acoustic transducer (or other medical device) in a posture conforming to a specified posture. Display 420 may also display scanning information such as medical images or treatment progress.
[0078] Camera data (images), segmentation, extracted features (landmarks), one or more machine learning models, one or more templates, rendered ultrasound images, and / or other information are stored in non-temporary computer-readable memory such as memory 410. Memory 410 is an external storage device, RAM, ROM, database, and / or local memory (e.g., a solid-state drive or hard drive). The same or a different non-temporary computer-readable medium may be used for instructions and other data. Memory 410 may reside in memory such as a hard disk, RAM, or removable medium using a database management system (DBMS). Alternatively, memory 410 is built into the processor 400 (e.g., a cache).
[0079] Instructions for performing the methods, processes, and / or techniques described herein are provided in non-temporary computer-readable storage media or memory, such as caches, buffers, RAM, removable media, hard drives, or other computer-readable storage media (e.g., memory 410). Computer-readable storage media include various types of volatile and non-volatile storage media. The functions, processes, or tasks illustrated or described herein are performed in response to one or more instruction sets stored in computer-readable storage media. The functions, processes, or tasks may be performed by software, hardware, integrated circuits, firmware, microcode, etc., acting alone or in combination, regardless of any particular type of instruction set, storage medium, processor, or processing strategy.
[0080] In one embodiment, the instructions are stored on a removable media device so that they can be read by a local or remote system. In another embodiment, the instructions are stored in a remote location so that they can be transferred over a computer network. In yet another embodiment, the instructions are stored in a given computer, CPU, GPU, or system. Since some of the components and method steps of the illustrated system configuration can be implemented in software, the actual connections between the components of the system (or between process steps) may vary depending on how this embodiment is programmed.
[0081] Various exemplary embodiments are listed below. The following exemplary embodiments outline various combinations of embodiments. Other combinations of any embodiment with any one or more other embodiments may be provided. An embodiment of one type (e.g., a method or system) may also be used in another type (a system or method).
[0082] Exemplary Embodiment 1: A method for estimating the orientation of an ultrasonic transducer, To capture an image of the ultrasonic transducer using a single camera, The ultrasonic transducer is obscured by the user's gripping in the image. To detect the ultrasonic transducer in the aforementioned image, To identify the orientation of the ultrasonic transducer detected in the image from the single camera, A method comprising performing ultrasonic imaging using the ultrasonic transducer and aligning the scanning data according to the identified orientation.
[0083] Exemplary Embodiment 2: The capturing described above includes capturing with a color depth camera, wherein the image includes a color point cloud. The detection method of exemplary embodiment 1 includes, at least in part, detection based on the color point cloud.
[0084] Exemplary Embodiment 3: The capturing method of exemplary embodiment 1 or 2 includes capturing an image of the ultrasonic transducer without an optical marker attached.
[0085] Exemplary Embodiment 4: The detection is performed by any of the exemplary embodiments 1 to 3, which includes segmenting the ultrasonic transducer in the image.
[0086] Exemplary Embodiment 5: A method of any of the exemplary embodiments 1 to 4, further comprising extracting a plurality of features of the ultrasonic transducer detected in the image.
[0087] Exemplary Embodiment 6: The extraction method of exemplary embodiment 5 includes extracting the features as semantic features and general clues.
[0088] Exemplary Embodiment 7: Identifying the orientation is a method of any of the exemplary embodiments 1 to 6, which includes comparing the detected ultrasonic transducer with an ultrasonic transducer template.
[0089] Exemplary Embodiment 8: The comparison described above includes comparing the detected ultrasonic transducer with the ultrasonic transducer template, the ultrasonic transducer template representing a generalized transducer that is not specific to the detected ultrasonic transducer, according to the method of exemplary embodiment 7.
[0090] Exemplary Embodiment 9: The comparison described above includes investigating various orientations of the ultrasonic transducer template with respect to the detected ultrasonic transducer, thereby obtaining the orientation in which the difference is minimized, according to exemplary embodiment 7 or 8.
[0091] Exemplary Embodiment 10: The comparison is performed by any of the exemplary embodiments 7 to 9, which includes comparing feature vectors output by an encoder in response to the detected ultrasonic transducer and the ultrasonic transducer template inputs.
[0092] Exemplary Embodiment 11: The method further includes extracting multiple features of the ultrasonic transducer detected in the image, The ultrasonic transducer template has more template features than the number of the aforementioned features, The comparison described above is any of the methods in the exemplary embodiments 7 to 10, which include comparing using outlier removal.
[0093] Exemplary Embodiment 12: The capturing, detecting, and identifying are performed repeatedly over time. A method in any of the exemplary embodiments 1 to 11, wherein the posture at a previous point in time is used for initialization when determining the posture at a later point in time.
[0094] Exemplary Embodiment 13: This further includes modeling the user's hands, Detecting the ultrasonic transducer and / or determining the orientation is done in any of the exemplary embodiments 1 to 12, taking into account the modeled hand.
[0095] Exemplary Embodiment 14: The ultrasonic imaging method is any of the exemplary embodiments 1 to 13, which includes aligning the data of the three-dimensional representation within the three-dimensional representation based on the orientation and rendering from the three-dimensional representation.
[0096] Exemplary Embodiment 15: It is an ultrasonic system, A single camera configured to image an acoustic transducer partially shielded by the user, The system includes an image processor configured to determine the orientation of the acoustic transducer from an image taken by the single camera, An ultrasonic system in which the orientation is determined based on the template of the acoustic transducer.
[0097] Exemplary Embodiment 16: The single camera comprises an ultrasonic system of exemplary embodiment 15, which includes a color depth sensor.
[0098] Exemplary Embodiment 17: An exemplary ultrasonic system of embodiment 15 or 16, wherein the image processor is configured to extract features of the acoustic transducer from the image, compare them with template features, and determine the orientation based on the extracted features.
[0099] Exemplary Embodiment 18: The template includes a generalized representation of a transducer that is not specific to the acoustic transducer in the image, representing an ultrasonic system of any of the exemplary embodiments 15 to 17.
[0100] Exemplary Embodiment 19: A method for estimating the orientation of a medical device, A single camera captures images of the medical device while it is being used on or within the patient. The medical device is obscured by the user's grasp in the image. To detect the medical device and the user portion in the image, A method comprising determining the posture of the medical device detected in the image from the single camera, using the portion of the user detected in the image.
[0101] Exemplary Embodiment 20: Identifying the aforementioned posture is Based on the user portion, remove the features detected with respect to the medical device. To indicate the posture using the orientation of the user's part, and / or, A method of exemplary embodiment 19, which includes determining the posture by considering both the medical device and the user's portion in the image.
[0102] The various improvements described herein may be used together or individually. While exemplary embodiments of the present invention are described herein with reference to the drawings, the present invention is not limited to those embodiments, and a person with ordinary skill in the art will understand that various other changes and modifications may be made without departing from the scope or spirit of the invention.
Claims
1. A method for estimating the orientation of an ultrasonic transducer, To capture an image of the ultrasonic transducer using a single camera, The ultrasonic transducer is obscured by the user's gripping in the image. To detect the ultrasonic transducer in the aforementioned image, To identify the orientation of the ultrasonic transducer detected in the image from the single camera, A method comprising performing ultrasonic imaging using the ultrasonic transducer and aligning the scanning data according to the identified orientation.
2. The capturing described above includes capturing with a color depth camera, wherein the image includes a color point cloud. The method according to claim 1, wherein the detection includes, at least in part, detection based on the color point cloud.
3. The method according to claim 1, wherein the capturing includes capturing an image of the ultrasonic transducer without an optical marker attached.
4. The method according to claim 1, wherein the detection includes segmenting the ultrasonic transducer in the image.
5. The method according to claim 1, further comprising extracting a plurality of features of the ultrasonic transducer detected in the image.
6. The method according to claim 5, wherein the extraction includes extracting the features as semantic features and general clues.
7. The method according to claim 1, wherein determining the orientation includes comparing the detected ultrasonic transducer with an ultrasonic transducer template.
8. The method according to claim 7, wherein the comparison includes comparing the detected ultrasonic transducer with the ultrasonic transducer template, the ultrasonic transducer template representing a generalized transducer that is not specific to the detected ultrasonic transducer.
9. The method according to claim 7, wherein the comparison includes investigating various orientations of the ultrasonic transducer template with respect to the detected ultrasonic transducer, and by investigating, the orientation in which the difference is minimized.
10. The method according to claim 7, wherein the comparison includes comparing feature vectors output by an encoder in response to the inputs of the detected ultrasonic transducer and the ultrasonic transducer template.
11. The method further includes extracting multiple features of the ultrasonic transducer detected in the image, The ultrasonic transducer template has more template features than the number of the aforementioned features, The method according to claim 7, wherein the comparison includes comparing using outlier removal.
12. The capturing, detecting, and identifying processes are repeatedly performed over time. The method according to claim 1, wherein the posture at a previous point in time is used for initialization when determining the posture at a later point in time.
13. This further includes modeling the user's hands, The method according to claim 1, wherein detecting the ultrasonic transducer and / or determining the posture takes into account the modeled hand.
14. The method according to claim 1, wherein the ultrasonic imaging includes aligning the data of the three-dimensional representation within the three-dimensional representation based on the orientation and rendering from the three-dimensional representation.
15. It is an ultrasonic system, A single camera configured to image an acoustic transducer partially shielded by the user, The system includes an image processor configured to determine the orientation of the acoustic transducer from an image taken by the single camera, An ultrasonic system in which the orientation is determined based on the template of the acoustic transducer.
16. The ultrasonic system according to claim 15, wherein the single camera comprises a color depth sensor.
17. The ultrasonic system according to claim 15, wherein the image processor is configured to extract features of the acoustic transducer from the image, compare them with template features, and determine the orientation based on the extracted features.
18. The ultrasonic system according to claim 15, wherein the template includes a generalized representation of a transducer that is not specific to the acoustic transducer in the image.
19. A method for estimating the orientation of a medical device, A single camera captures images of the medical device while it is being used on or inside the patient. The medical device is obscured by the user's grasp in the image. To detect the medical device and the user portion in the image, A method comprising determining the posture of the medical device detected in the image from the single camera, using the portion of the user detected in the image.
20. Identifying the aforementioned posture is Based on the user portion, remove the features detected with respect to the medical device. To indicate the posture using the orientation of the user's part, and / or, The method according to claim 19, comprising determining the posture by considering both the medical device and the user's portion in the image.