System and method for real-time multi-modality image alignment
The system achieves real-time, sub-millimeter accurate image alignment of 3D images without markers, addressing precision challenges in surgical navigation and instrument tracking.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ZETA SURGICAL INC
- Filing Date
- 2020-08-14
- Publication Date
- 2026-05-13
AI Technical Summary
Existing image registration systems face challenges in achieving sub-millimeter accuracy and real-time alignment without the use of markers, particularly in environments sensitive to lighting, shading, occlusion, and sensor noise, making precise surgical navigation and instrument tracking difficult.
A system and method for real-time multi-modality image alignment that transforms 3D point clouds into different reference frames using processors to align 3D images with sub-millimeter accuracy, employing techniques like downsampling, feature matching, and reference frame selection without markers, enabling precise alignment of medical scans and instrument tracking.
Enables precise, real-time alignment of 3D images with sub-millimeter accuracy, facilitating surgical navigation and instrument tracking, and reducing processing hardware requirements, improving surgical precision and efficiency.
Smart Images

Figure 0007857861000005 
Figure 0007857861000006 
Figure 0007857861000007
Abstract
Description
Cross-reference of related applications
[0001] This application claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 888,099, the disclosure of which is incorporated herein by reference in its entirety. [Technical Field]
[0002] Image registration (or image alignment) can be used for a variety of purposes. For example, it can be used to align image data from a camera with a 3D model and to establish a correlation between the image data and the stored 3D information.
[0003] This disclosure generally relates to the field of image detection and registration. More specifically, this disclosure relates to a system and method for real-time multi-modality image alignment. The system and method of this technical solution can be used for real-time 3D point cloud and image alignment, such as for medical analysis or surgical applications. [Overview of the Initiative]
[0004] Various embodiments generally relate to systems and methods for real-time multi-modality image alignment using three-dimensional (3D) image data, which can be performed without markers and with sub-millimeter accuracy. 3D images, including scans such as CT or MRI, can be directly aligned to a subject, such as a patient's body, captured in real time using one or more capture devices. This allows for the real-time display of specific scan information, such as information about internal tissues, along with a point cloud representation of the subject. This is beneficial for surgical procedures where a manual process is used to position instruments in the same reference frame as in a CT scan. Instruments can be tracked, their trajectories drawn, and targets highlighted on the scan. This solution can provide real-time registration with sub-millimeter accuracy for various applications such as aligning depth capture information to medical scans (e.g., for surgical navigation), aligning depth capture information to CAD models (e.g., for manufacturing and troubleshooting), aligning and merging multiple medical imaging modalities (e.g., MRI and CT, CT and 3D ultrasound, MRI and 3D ultrasound), aligning multiple CAD models (e.g., to find differences between models), and merging depth capture data from multiple image capture devices. [Problems that the invention aims to solve]
[0005] This solution can be used for image-guided procedures in various settings, including operating rooms, outpatient settings, CT suites, ICUs, and emergency rooms. It can be used in neurosurgical applications such as CSF diversion procedures including external ventricle placement and VP shunt placement, brain tumor resection and biopsy, and electrode placement. It can also be used in interventional radiology procedures such as abdominal and lung biopsy, resection, aspiration, and drainage. The solution of this invention can be used in orthopedic surgeries such as spinal fusion procedures. [Means for solving the problem]
[0006] At least one aspect of this disclosure relates to a method for transforming a three-dimensional point cloud into a different reference frame. This method can be implemented, for example, by one or more processors of a data processing system. The method may include one or more processors accessing a first set of data points of a first point cloud captured by a first capture device having a first pose, and a second set of data points of a second point cloud captured by a second capture device having a second pose different from the first pose. The method may include a step of selecting a reference frame based on the first set of data points. The method may include a step of using the reference frame and the first set of data points to determine a transformation data structure for the second set of data points. The method may include a step of using the transformation data structure and the second set of data points to transform the second set of data points into a transformed set of data points.
[0007] In some implementations (embodiments) of this method, the step of accessing the first set of data points of the first point cloud may include the step of receiving three-dimensional image data from the first capture device. In some implementations of this method, the step of accessing the first set of data points of the first point cloud may include the step of generating the first point cloud having the first set of data points using the three-dimensional image data from the first capture device.
[0008] In some implementations of this method, the second capture device may be identical to the first capture device. In some implementations of this method, the step of selecting the reference frame may include the step of selecting a first reference frame of the first point cloud as the first reference frame. In some implementations of this method, the step of selecting the reference frame may include the step of obtaining color data assigned to one or more of the first set of data points of the first point cloud. In some implementations of this method, the step of selecting the reference frame may include the step of determining the reference frame based on the color data. In some implementations of this method, the step of determining the transformed data structure may include the step of generating the transformed data structure to include a change in position or rotation of at least one point in the second set of data points.
[0009] In some implementations of this method, the step of transforming the second set of data points may include a step of applying the position change or the rotation change to at least one point in the second set of data points to generate a transformed set of data points. In some implementations of this method, the step of generating a combined set of data points including the first set of data points and the transformed set of data points may further include a step of accessing the first set of data points and the second set of data points may include a step of downsampling at least one of the first set of data points or the second set of data points. In some implementations of this method, the step of transforming the second set of data points may include a step of matching at least one first point in the first set of data points to at least one second point in the second set of data points.
[0010] At least one other aspect of the present disclosure relates to a system configured to transform a three-dimensional point cloud into different reference frames. The system may include one or more processors configured with machine-readable instructions. The system may have access by one or more processors to a first set of data points of a first point cloud captured by a first capture device having a first pose, and a second set of data points of a second point cloud captured by a second capture device having a second pose different from the first pose. The system may have one or more processors select a reference frame based on the first set of data points. The system may have one or more processors use the reference frame and the first set of data points to determine a transformation data structure for the second set of data points. The system may have one or more processors use the transformation data structure and the second set of data points to transform the second set of data points into a transformed set of data points.
[0011] In some implementations, the system can access a first set of data points in a first point cloud by receiving three-dimensional image data from the first capture device. In some implementations, the system can access a first set of data points in a first point cloud by generating the first point cloud using the three-dimensional image data from the first capture device so that it has a first set of data points. In some implementations of the system, the second capture device can be identical to the first capture device. In some implementations, the system can select a reference frame by selecting a first reference frame of the first point cloud as the first reference frame.
[0012] In some implementations, the system can obtain color data assigned to one or more of the first set of data points in the first point cloud. In some implementations, the system can determine the reference frame based on the color data. In some implementations, the system can determine the transformed data structure by generating the transformed data structure to include a change in position or rotation of at least one point in the second set of data points. In some implementations, the system can transform the second set of data points by applying the change in position or rotation to at least one point in the second set of data points to generate a transformed set of data points.
[0013] In some implementations, the system can generate a combined set of data points, which includes a first set of data points and a transformed set of data points. In some implementations, the system can downsample at least one of the first set of data points or the second set of data points. In some implementations, the system can match at least one first point in the first set of data points with at least one second point in the second set of data points.
[0014] At least one other aspect of this disclosure relates to a method for downsampling three-dimensional point cloud data. The method may include the step of accessing a set of data points corresponding to a point cloud representing the surface of an object. The method may include the step of applying a response function to the set of data points to assign a set of response values to the set of data points. The method may include the step of selecting a subset of the set of data points using a selection policy and the set of response values. The method may include the step of generating a data structure containing the subset of the set of data points.
[0015] In some implementations of this method, the step of accessing the set of data points may include the step of receiving three-dimensional image data from at least one capture device. In some implementations of this method, the step of accessing the set of data points may include the step of using the three-dimensional image data to generate the point cloud representing the surface of the object. In some implementations of this method, the step of applying the response function to the set of data points may include the step of generating a graph data structure using the set of data points corresponding to the point cloud representing the surface of the object. In some implementations of this method, the step of applying the response function to the set of data points may include the step of determining a graph filter as part of the response function.
[0016] In some implementations of this method, the step of determining the graph filter may include the step of generating a k-dimensional binary tree using the set of data points. In some implementations of this method, the step of determining the graph filter may include the step of generating the graph filter using the k-dimensional binary tree. In some implementations of this method, the step of determining the graph filter may include the step of generating the graph filter using the Euclidean distance between pairs of data points in the set of data points. In some implementations of this method, the step of generating the graph filter using the Euclidean distance between pairs of data points in the set of data points may further be based on at least one color channel of the pairs of data points in the set of data points.
[0017] In some implementations of this method, the step of generating a graph filter may include a step of identifying the brightness parameter of the set of data points. In some implementations of this method, the step of generating a graph filter may include a step of deciding whether or not to generate the graph filter based on at least one color channel of the pairs of data points in the set of data points. In some implementations of this method, the step of selecting the subset of the set of data points may include a step of performing a weighted random selection using each response value of the set of response values as a weight. In some implementations of this method, the selection policy may be configured to select a subset of the data points corresponding to one or more contours on the surface of the object. In some implementations of this method, the step of storing the data structure containing the subset of the set of data points in memory may further include a step of storing the data structure containing the subset of the set of data points in memory.
[0018] At least one other aspect of this disclosure relates to a system configured for downsampling three-dimensional point cloud data. The system may include one or more processors configured with machine-readable instructions. The system may have access to a set of data points corresponding to a point cloud representing the surface of an object. The system may apply a response function to the set of data points to assign a set of response values to the set of data points. The system may use a selection policy and the set of response values to select a subset of the set of data points. The system may generate a data structure containing the subset of the set of data points.
[0019] In some implementations, the system can receive three-dimensional image data from at least one capture device. In some implementations, the system can use the three-dimensional image data to generate the point cloud representing the surface of the object. In some implementations, the system can use the set of data points corresponding to the point cloud representing the surface of the object to generate a graph data structure. In some implementations, the system can determine a graph filter as part of the response function. In some implementations, the system can use the set of data points to generate a k-dimensional binary tree. In some implementations, the system can use the k-dimensional binary tree to generate the graph filter.
[0020] In some implementations, the system can generate the graph filter using the Euclidean distance between pairs of data points in the set of data points. In some implementations, the system can further generate the graph filter using the Euclidean distance between pairs of data points in the set of data points based on at least one color channel of the pairs of data points in the set of data points. In some implementations, the system can identify the brightness parameter of the set of data points. In some implementations, the system can determine whether to generate the graph filter based on the at least one color channel of the pairs of data points in the set of data points.
[0021] In some implementations, the system can perform weighted random selection using each response value in the set of response values as a weight. In some implementations, the selection policy can be configured to cause the system to select a subset of the data points corresponding to one or more contours on the surface of the object. In some implementations, the system can store in memory the data structure including the subset of the set of data points.
[0022] At least one other aspect of the present disclosure relates to a method for scheduling processing jobs for point cloud and image alignment. The method can include the step of identifying a first processing device having a first memory and a second multi-processor device having a second memory. The method can include the step of identifying a first processing job for a first point cloud. The method can include the step of determining to allocate the first processing job to the second multi-processor device. The method can include the step of allocating information for the first processing job including the first point cloud to the second memory. The method can include the step of receiving a second processing job for a second point cloud. The method can include the step of determining to allocate the second processing job to the first processing device. The method can include the step of allocating information for the second processing job to the first memory. The method can include the step of transferring, by the one or more processors, a first instruction for causing the second multi-processor device to execute the first processing job, to the second multi-processor device. The first instruction can be specific to the second multi-processing device. The method can include the step of transferring, by the one or more processors, a second instruction for causing the first processing device to execute the second processing job, to the first processing device. The second instruction can be specific to the first processing device.
[0023] In some implementations of this method, the step of deciding to assign the first processing job to the second multiprocessor device may include the step of deciding that the first processing job includes operations for feature detection in the first point cloud. In some implementations of this method, the step of deciding to assign the first processing job to the second multiprocessor device may further be based on at least one of the following: the number of data points in the first point cloud is less than a predetermined threshold; the utilization of the second multiprocessor device is less than a predetermined threshold; or the processing complexity of the first job exceeds a complexity threshold. In some implementations of this method, the step of identifying the first processing job for the first point cloud may include the step of receiving the first point cloud from a first capture device.
[0024] In some implementations of this method, the step of identifying a predecessor second processing job for the second point cloud may correspond to the step of identifying a first processing job for the first point cloud. In some implementations of this method, the first or second processing job may include at least one of downsampling, normalization, feature detection, contour detection, registration, or rendering operations. In some implementations of this method, the second point cloud may be extracted from at least one of an image captured from a capture device or a 3D medical image. In some implementations of this method, the step of deciding to assign the second processing job to the first processing device may correspond to the step of deciding to extract the second point cloud from a 3D medical image.
[0025] In some implementations of the Method, the first processing job may be associated with a first priority value. In some implementations of the Method, the second processing job may be associated with a second priority value. In some implementations of the Method, the step of deciding to assign the second processing job to the first processing device may include a step of deciding that the first priority value is greater than the second priority value, and may further include a step of deciding to assign the second processing job to the first processing device. In some implementations of the Method, each of the first priority values may be based on a first frequency at which the first processing job is performed. In some implementations of the Method, the second priority value may be based on a second frequency at which the second processing job is performed.
[0026] At least one other aspect of this disclosure relates to a system for scheduling processing jobs for point cloud and image alignment. The system may include one or more processors configured with machine-readable instructions. The system may identify a first processing device having a first memory and a second multiprocessor device having a second memory. The system may identify a first processing job for a first point cloud. The system may decide to assign the first processing job to the second multiprocessor device. The system may allocate information for the first processing job, including the first point cloud, to the second memory. The system may receive a second processing job for a second point cloud. The system may decide to assign the second processing job to the first processing device. The system may allocate information for the second processing job to the first memory. The system may use one or more processors to transfer a first instruction to the second multiprocessor device causing the second multiprocessor device to perform the first processing job. The first instruction may be specific to the second multiprocessor device. This system can transfer a second instruction to the first processing device, via one or more processors, to cause the first processing device to execute the second processing job. The second instruction may be specific to the first processing device.
[0027] In some implementations, the system can determine that the first processing job includes operations for feature detection in the first point cloud. In some implementations, the system can decide to assign the first processing job to the second multiprocessor device based on at least one of the following: the number of data points in the first point cloud is below a predetermined threshold; the utilization of the second multiprocessor device is below a predetermined threshold; or the processing complexity of the first job exceeds a complexity threshold. In some implementations, the system can receive the first point cloud from a first capture device.
[0028] In some implementations, the system can identify a second processing job for a second point cloud in response to identifying a first processing job for the first point cloud. In some implementations of the system, the first or second processing job may include at least one of downsampling, normalization, feature detection, contour detection, alignment, or rendering operations. In some implementations, the system can extract the second point cloud from at least one of an image captured from a capture device or a 3D medical image. In some implementations, the system can decide to assign the second processing job to the first processing device in response to determining that the second point cloud is extracted from a 3D medical image.
[0029] In some implementations of this system, the first processing job may be associated with a first priority value. In some implementations of this system, the second processing job may be associated with a second priority value. In some implementations, the system may decide to assign the second processing job to the first processing device in response to determining that the first priority value is greater than the second priority value. In some implementations of this system, the first priority value may be based on a first frequency at which the first processing job is performed, and the second priority value may be based on a second frequency at which the second processing job is performed.
[0030] At least one other aspect of the present disclosure relates to a method for aligning a three-dimensional medical image to a point cloud. The method may include the step of accessing a first set of data points of a first point cloud representing a global scene having a first reference frame. The method may include the step of identifying a set of feature data points of a feature of a three-dimensional medical image having a second reference frame different from the first reference frame. The method may include the step of using the first reference frame, the first set of data points, and the set of feature data points to determine a transformation data structure for the three-dimensional medical image. The method may include the step of using the transformation data structure to align the 3D medical image to the first point cloud representing the global scene so that the 3D medical image is positioned relative to the first reference frame.
[0031] In some implementations of the Method, the step of determining the transformation data structure may include a step of downsampling the first set of data points to generate a reduced set of first data points. In some implementations of the Method, the step of determining the transformation data structure may include a step of using the reduced set of first data points to determine the transformation data structure for the 3D medical image. In some implementations of the Method, the step of determining the transformation data structure may further include a step of generating the transformation data structure to include a change in the position or rotation of the 3D medical image. In some implementations of the Method, the step of aligning the 3D medical image to the first point cloud representing the global scene may further include a step of applying the change in position or rotation to the 3D medical image to align the features of the 3D medical image with corresponding points in the first point cloud.
[0032] In some implementations of this method, the step of identifying the set of feature data points for the features of the 3D medical image may include the step of assigning weight values to each data point in the medical image to generate a set of weight values. In some implementations of this method, the step of identifying the set of feature data points for the features of the 3D medical image may include the step of selecting the data points of the 3D medical image that correspond to weight values that satisfy a weight value threshold as the set of feature data points. In some implementations of this method, the step of displaying a rendering of the first point cloud and the 3D medical image in response to aligning the 3D medical image with the first point cloud may be included. In some embodiments of this method, the step of receiving tracking data from a surgical instrument may be included. In some embodiments of this method, the step of converting the tracking data forming the surgical instrument into a first reference frame to generate converted tracking data may be included. In some embodiments of this method, the step of rendering the converted tracking data in the rendering of the first point cloud and the 3D medical image may be included.
[0033] Some embodiments of this method may include a step of determining a position of interest within a first reference frame relating to the first point cloud and the 3D medical image. Some implementations of this method may include a step of generating an action command for a surgical instrument based on the first point cloud, the 3D medical image, and the position of interest. Some embodiments of this method may include a step of transmitting the action command to the surgical instrument. Some embodiments of this method may further include a step of displaying a highlighted region corresponding to the position of interest within the rendering of the 3D medical image and the first point cloud. Some embodiments of this method may further include a step of determining the distance of the patient represented in the 3D medical image from a capture device that is at least partially involved in generating the first point cloud.
[0034] At least one other aspect of the present disclosure relates to a system configured for aligning a three-dimensional medical image to a point cloud. The system may include one or more processors configured with machine-readable instructions. The system may have access to a first set of data points of a first point cloud representing a global scene having a first reference frame. The system may identify a set of feature data points for features of a three-dimensional medical image having a second reference frame different from the first reference frame. The system may use the first reference frame, the first set of data points, and the set of feature data points to determine a transformation data structure for the 3D medical image. The system may use the transformation data structure to align the 3D medical image to the first point cloud representing the global scene so that the 3D medical image is positioned relative to the first reference frame.
[0035] In some implementations, the system can downsample the first set of data points to generate a reduced set of first data points. In some implementations, the system can use the reduced set of first data points to determine the transformation data structure for the 3D medical image. In some implementations, the system can generate the transformation data structure to include a change in the position or rotation of the 3D medical image. In some implementations, the system can apply the change in position or rotation to the 3D medical image to align the features of the 3D medical image with corresponding points in the first point cloud.
[0036] In some implementations, the system can assign weight values to each data point in the medical image to generate sets of weight values. In some implementations, the system can select the data points in the 3D medical image corresponding to weight values that satisfy a weight value threshold as a set of feature data points. In some implementations, the system can display renderings of the first point cloud and the 3D medical image in response to aligning the 3D medical image with the first point cloud. In some implementations, the system can receive tracking data from surgical instruments. In some implementations, the system can convert the tracking data forming the surgical instruments into a first reference frame to generate converted tracking data. In some implementations, the system can render the converted tracking data within renderings of the first point cloud and the 3D medical image.
[0037] In some implementations, the system can determine a position of interest within a first reference frame relating to the first point cloud and the 3D medical image. In some implementations, the system can generate action commands for a surgical instrument based on the first point cloud, the 3D medical image, and the position of interest. In some embodiments, the system can transmit the action commands to the surgical instrument. In some embodiments, the system can display highlighted regions corresponding to the position of interest within the rendering of the 3D medical image and the first point cloud. In some embodiments, the system can determine the distance of the patient represented in the 3D medical image from a capture device that is at least partially involved in generating the first point cloud.
[0038] These and other embodiments and implementations are described in detail below. The information set forth herein and the detailed description below include exemplary examples of various embodiments and implementations and provide an overview or framework for understanding the nature and features of the claimed embodiments and implementations. The drawings, which provide examples and further understanding of various embodiments and embodiments, are incorporated herein and constitute part of this specification. The embodiments can be combined and it will be readily apparent that features described in the context of one embodiment of the invention can be combined with other embodiments. Embodiments can be carried out in any convenient form, for example, by a suitable computer program that can be carried on a suitable carrier medium (computer-readable medium) which may be a tangible carrier medium (e.g., disk) or an intangible carrier medium (e.g., communication signals). Embodiments can also be carried out using a suitable device, which may take the form of a programmable computer that runs a computer program arranged to carry out the corresponding embodiment. Where used herein and in the claims, the singular forms "a," "an," and "the" include plural references unless otherwise specified by the context. [Brief explanation of the drawing]
[0039] The attached drawings are not intended to be drawn to scale. Similar reference numbers and designations in various drawings refer to similar elements. For clarity, not all components are shown in all drawings. [Figure 1] This is a perspective view of an image processing system according to one embodiment of the present disclosure. [Figure 2] This is a block diagram of an image processing system according to an embodiment of the present disclosure. [Figure 3] This is a flowchart illustrating a method for aligning image data from multiple modalities according to embodiments of the present disclosure. [Figure 4] This is a flowchart of a method for aligning multiple depth cameras in an environment based on image data, according to an embodiment of the present disclosure. [Figure 5] This is a flowchart of a method for segmenting the surface of a medical image according to embodiments of the present disclosure. [Figure 6] This is a flowchart of a method for generating a 3D surface model from a medical image based on segmentation, according to an embodiment of the present disclosure. [Figure 7] This is a flowchart of a method for generating a 3D surface model from a medical image based on segmentation, according to an embodiment of the present disclosure. [Figure 8] This is a flowchart of a method for downsampling a point cloud generated from a 3D surface model to improve image alignment efficiency, according to embodiments of the present disclosure. [Figure 9] This is a flowchart of a method for detecting contour points from a downsampled point cloud and performing a prioritization analysis of the contour points, according to an embodiment of the present disclosure. [Figure 10] This is a flowchart of a method for aligning a point cloud of a medical image with a global scene point cloud according to an embodiment of the present disclosure. [Figure 11] This is a flowchart of a method for performing real-time surgical planning visualization using pre-captured medical images and global scene images according to embodiments of the present disclosure. [Figure 12] This is a flowchart of a method for dynamically tracking the movement of an instrument in a 3D image environment according to embodiments of the present disclosure. [Figure 13A] This is a block diagram of a computing environment according to one embodiment of the present disclosure. [Figure 13B] This is a block diagram of a computing environment according to one embodiment of the present disclosure. [Figure 14] This figure shows a resampled image according to an embodiment of the present disclosure. [Figure 15] This is a flowchart of a method for dynamically allocating processing resources to different computational objects within a processing circuit, according to an embodiment of the present disclosure. [Modes for carrying out the invention]
[0040] The following describes in detail various concepts and implementations related to techniques, approaches, methods, apparatus, and systems for real-time multimodality image alignment. The various concepts introduced above and described in more detail below can be implemented in one of many ways, as the described concepts are not limited to a specific form of implementation. Examples of specific implementations and applications are presented primarily for illustrative purposes.
[0041] I. Overview
[0042] The systems and methods according to this solution can be used to perform real-time alignment of image data from multiple modalities, such as aligning or registering 3D image data to medical scan data. Some systems can use markers for registration, which can be numerous, require attachment to the subject, or may interfere with one or more image capture devices. Due to processing requirements in the image processing pipeline, it may be difficult to compute such systems with high precision, such as sub-millimeter accuracy, and in real time. Furthermore, various image processing calculations may be highly sensitive to factors affecting image data, such as lighting, shading, occlusion, sensor noise, and camera pose.
[0043] The system and method according to this solution can improve the speed at which image data from multiple sources is processed and aligned by applying various image processing solutions, thereby improving performance to achieve desired performance benchmarks without the use of markers and reducing processing hardware requirements. This solution can enable a precise, responsive, and user-friendly surgical navigation platform. For example, this solution can directly align 3D scans, such as CT scans or MRI scans, to the subject (or image data representing the subject), and further enable instrument tracking, instrument trajectory mapping, and target highlighting on the scan.
[0044] Figures 1-2 show an image processing system 100. The image processing system 100 may include a plurality of image capture devices 104, such as a 3D camera. The cameras may be visible light cameras (e.g., color or black and white), infrared cameras, or a combination thereof. Each image capture device 104 may include one or more lenses 204. In some embodiments, the image capture device 104 may include a camera for each lens 204. The image capture device 104 may be selected or designed to have a predetermined resolution and / or a predetermined field of view. The image capture device 104 may have a resolution and field of view for detecting and tracking objects. The image capture device 104 may have a pan, tilt, or zoom mechanism. The image capture device 104 may have a pose (orientation) corresponding to the position and orientation of the image capture device 104. The image capture device 104 may be a depth camera. The image capture device 104 may be a Microsoft Kinect.
[0045] The light of the image captured by the image capture device 104 is received through one or more lenses 204. The image capture device 104 may include, but is not limited to, a sensor circuit including a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) circuit, which can detect the light received through one or more lenses 204 and generate an image 208 based on the received light.
[0046] The image capture device 104 can provide the image 208 to the processing circuit 212, for example, via a communication bus. The image capture device 104 can provide a timestamp corresponding to the image 208, which facilitates synchronization of the image 208 when image processing is performed on it. The image capture device 104 can output a 3D image (e.g., an image with depth information). The image 208 may contain multiple pixels, each of which is assigned spatial position data (e.g., horizontal, vertical, and depth data), brightness or luminance data, and / or color data.
[0047] Each image capture device 104 can be coupled to the respective end of one or more arms 108 that can be coupled to a platform 112. The platform 112 can be a cart that includes wheels for movement and various support surfaces for supporting devices used with the platform 112.
[0048] The arm 108 can change its position and orientation by rotating, extending, retracting, or inserting and removing it, and can control the pose of the image capture device 104. The platform 112 can support processing hardware 116, which includes at least a portion of the processing circuit 212, and the user interface 120. The image 208 can be processed by the processing circuit 212 for presentation via the user interface 120.
[0049] The processing circuit 212 can incorporate features of the computing device 1300 described with reference to Figures 13A and 13B. For example, the processing circuit 212 may include a processor and memory. The processor may be implemented as a purpose-specific processor, an application-specific integrated circuit (ASIC), one or more field-programmable gate arrays (FPGAs), a set of processing components, or other suitable electronic processing components. The memory is one or more devices (e.g., RAM, ROM, flash memory, hard disk storage) for storing data and computer code to complete or facilitate the various user or client processes, layers, and modules described herein. The memory may be or include volatile memory or non-volatile memory and may include database components, object code components, script components, or any other type of information structure to support various activities and information structures of the inventive concepts disclosed herein. The memory is communicably connected to the processor and includes computer code or instruction modules for performing one or more processes described herein. The memory includes various circuits, software engines, and / or modules that cause the processor to perform the systems and methods described herein.
[0050] Some parts of the processing circuit 212 can be provided by one or more devices located away from the platform 112. For example, one or more servers, cloud computing systems, or mobile devices (such as those described with reference to Figures 13A and 13B) can be used to implement various parts of the image processing pipeline described herein.
[0051] The image processing system 100 may include a communication circuit 216. The communication circuit 216 can implement features of the computing device 1300 described with reference to Figures 13A and 13B, such as a network interface 1318.
[0052] The image processing system 100 may include one or more infrared (IR) sensors 220. The IR sensors 220 can detect IR signals from various devices in the environment surrounding the image processing system 100. For example, the IR sensors 220 can be used to detect IR signals from an IR emitter that can be coupled to equipment for tracking. The IR sensors 220 can be communicatively coupled to other components of the image processing system 100 so that other components of the image processing system 100 can utilize the IR signals in appropriate operation in the image processing pipeline, as described later herein.
[0053] Figure 3 shows an image processing pipeline 300 that the image processing system 100 can perform using image data from one or more image modalities. Various features of the image processing pipeline 300 that enable the image processing system 100 to perform high-precision real-time image alignment will be further described herein.
[0054] In some embodiments, setup procedures can be performed to enable the image processing system 100 to perform various functions described herein. For example, the platform 112 can be positioned in close proximity to a subject, such as a patient. The image capture device 104 can be positioned and oriented in various poses to detect image data relating to the subject. The image capture device 104 can be positioned in different poses so as to face the subject from multiple directions, thereby improving the quality of the image data generated by fusing the image data from the image capture device 104.
[0055] In 305, first image data can be received. The first image data may be model data such as CT, MRI, ultrasound, or CAD data (e.g., medical scan data, DICOM data). The model data may be received from a remote source (e.g., a cloud server) via a healthcare facility network, such as a network connected to an image archiving and communication system (PACS), or it may be in the memory of the processing circuit 216. The model data may be intraoperative data (e.g., detected while a procedure is being performed on the subject) or preoperative data.
[0056] At 310, a second image data can be received. The second image data may be of a different modality than the first image data. For example, the second image data may be 3D image data from a 3D camera.
[0057] In 315, the first image data may be resampled, such as through downsampling, which is performed in a manner that preserves the key features of the first image data, reduces the data complexity of the first image data, and improves the efficiency of further calculations performed on the first image data. Similarly, in 320, the second image data may be resampled. Resampling may include identifying features in the image data that are not relevant to image alignment and removing them to generate a reduced or downsampled image.
[0058] In 325, one or more first feature descriptors can be determined with respect to the resampled first image data. The feature descriptors can be determined to relate to contours or other features of the resampled first image data that correspond to a 3D surface represented by the resampled first image data. Similarly, in 330, one or more second feature descriptors can be determined with respect to the second image data.
[0059] In 335, feature matching can be performed between one or more first feature descriptors and one or more second feature descriptors. For example, feature matching can be performed by comparing each first and second feature descriptor to determine a match score, and by identifying a match in response to the match score satisfying a match threshold.
[0060] In 340, in response to feature matching, one or more alignments can be performed between the first image data and the second image data. One or more alignments can be performed to transform at least one of the first image data or the second image data into a common reference frame.
[0061] II. System and method for aligning multiple depth cameras in an environment based on image data.
[0062] By using multiple depth cameras, such as 3D cameras, the quality of 3D image data collected about the subject and its surrounding environment can be improved. However, aligning image data from multiple depth cameras with different orientations can be difficult. This solution can transform various point cloud data points and effectively determine a reference frame for aligning point cloud data points to a reference frame in order to generate aligned image data.
[0063] Referring back to Figures 1 and 2, the image processing system 100 can use the image capture device 104 as a 3D camera to capture real-time 3D image data of a subject. For example, each image capture device 104 can capture at least one 3D image of a subject, object, or environment. This environment may include other features that would not medically be present in the environment. The 3D image can consist of a number or set of points within a reference frame provided by the image capture device. The set of points constituting the 3D image may have color information. In some implementations, this color information is discarded and not used in further processing steps. Each set of points captured by each image capture device 104 can be called a "point cloud". When multiple image capture devices 104 are used to capture images of a subject, each image capture device 104 may have a different reference frame. In some implementations, the 3D images captured by the image capture devices 104 can be recorded in real time. In such an implementation, a single image capture device 104 is used to capture a first 3D image in a first pose, and then it is repositioned to a second pose to capture a second 3D image of the subject.
[0064] In response to capturing at least one 3D image (e.g., as at least one of image 208), the image capture device 104 can provide image 208 to the processing circuit 212, for example, via a communication bus. The image capture device 104 can provide a timestamp corresponding to image 208, which facilitates synchronization of image 208 when image processing is performed on it. The image capture device 104 can output a 3D image (e.g., an image with depth information). Image 208 may contain multiple pixels, each of which is assigned spatial position data (e.g., horizontal, vertical, and depth data), brightness or luminance data, or color data. In some implementations, the processing circuit 212 can store image 208 in its memory. For example, storing image 208 may include indexing image 208 in one or more data structures in the memory of the processing circuit 212.
[0065] The processing circuit 212 accesses a first set of data points of a first point cloud captured by a first capture device 104 having a first pose, and a second set of data points of a second point cloud captured by a second capture device 104 having a second pose different from the first pose. For example, each 3D image (e.g., image 208) can contain one or more 3D dimensional data points that make up the point cloud. A data point can correspond to a single pixel captured by a 3D camera and can be at least a 3D data point (e.g., each containing at least three coordinates corresponding to a dimension). A 3D data point can contain at least three coordinates within a reference frame shown in each image 208. In this way, different image capture devices 104 with different poses can generate 3D images with different reference frames. To improve the overall accuracy and feature density of the system when 3D images of a subject are captured, the system can align the point clouds of the 3D images captured by the image capture devices 104 to generate a single composite 3D image. The 3D data points that make up one of the images 208 can be considered together as a single "point cloud".
[0066] The processing circuit 212 can extract 3D data from each data point of the image 208 received from the image capture device 104 to generate a first point cloud corresponding to the first image capture device 104 and a second point cloud corresponding to the second image capture device 104. Extracting 3D data from the point cloud may include accessing and extracting only the three coordinates of the data points in the 3D image (e.g., x-axis, y-axis, and z-axis, etc.) (for example, by copying them to a different area of memory in the processing circuit 212). Such processing may, in a further processing step, remove or discard color information or other irrelevant information.
[0067] In some implementations, to improve the overall computational efficiency of the system, the processing circuit can downsample or selectively discard specific data points that make up the 3D image to generate a downsampled set of data points. The processing circuit 212 can selectively remove data points uniformly, for example, by discarding one of four data points in an image (e.g., not extracting any data points from the image) (e.g., 75% of the points are uniformly extracted). In some implementations, the processing circuit 212 can extract different percentages of points (e.g., 5%, 10%, 15%, 20%, or any other percentage). Therefore, when extracting or accessing data points in a point cloud of a 3D image, the processing circuit 212 can downsample the point cloud to reduce its overall size and improve image processing without significantly affecting the accuracy of further processing steps.
[0068] In response to the 3D image data from each of the image capture devices 104 being converted into two or more point clouds, or otherwise accessed, the processing circuit 212 can select one of the point clouds to serve as a baseline reference frame for the alignment of any of the other point clouds. To improve the accuracy and overall resolution of the point clouds representing the surface of an object in the environment, two or more image capture devices 104 can capture 3D images of the object. The processing circuit 212 can synthesize the images so that they reside within a single reference frame. For example, the processing circuit 212 can select one of the point clouds corresponding to a single 3D image captured by the first image capture device 104 as the reference frame. Selecting a point cloud as the reference frame may include copying the selected point cloud (e.g., data points and coordinates that make up the point cloud) to different areas of memory. In some implementations, selecting a point cloud may include allocating memory points to at least a portion of the memory of the processing circuit 212 where the selected point cloud is stored.
[0069] Selecting a reference frame may involve obtaining color data assigned to one or more of the first sets of data points in the first point cloud. For example, the processing circuit 212 can extract color data (e.g., red / green / blue (RGB) values, cyan / yellow / magenta / luminance (CMYK) values, etc.) from pixels or data points in the 3D image 208 received from the image capture device 104 and store the color data in the data points of each point cloud. The processing circuit 212 can determine whether a reference frame is more uniformly illuminated by comparing the color data of each data point with a luminance value (e.g., a threshold for the average color value). The processing circuit 212 can perform this comparison for a certain number of data points in each point cloud, for example, by looping through every N data points and comparing the color threshold with the color data of each data point. In some implementations, the processing circuit 212 can average the color data across the data points in each point cloud to calculate an average color luminance value. In response to the average color brightness value being greater than a predetermined threshold, the processing circuit 212 can determine that the point cloud is uniformly illuminated.
[0070] In some implementations, the processing circuit 212 can select a reference frame by determining the most illuminated (e.g., most uniformly illuminated) point cloud. The most uniformly illuminated (e.g., thus having a high-quality image) point cloud can be selected as the reference frame for further alignment calculations. In some implementations, the processing circuit can select the reference frame as the reference frame for the least uniformly illuminated point cloud. In some implementations, the processing circuit 212 can arbitrarily select a reference frame for a point cloud (e.g., using pseudorandom numbers) as the reference frame.
[0071] The processing circuit 212 can determine a transformation data structure for a second set of data points using a reference frame and a first set of data points. The transformation data structure can include one or more transformation matrices. The transformation matrices can be, for example, 4x4 rigid body transformation matrices. To generate the transformation matrices for the transformation data structure, the processing circuit 212 can identify one or more feature vectors by performing one or more steps in method 900, which will be described herein in relation to Figure 9. The result of this processing can include a set of feature vectors for each point cloud, where one point cloud is used as the reference frame (for example, the points in that point cloud are not transformed). The processing circuit 212 can generate transformation matrices such that when each matrix is applied to its respective point cloud (for example, used for transformation), the features of the transformed point cloud are aligned with similar features in the reference frame point cloud.
[0072] To generate the transformation matrix (for example, as part of the transformation data structure, or as a transformation data structure), the processing circuit 212 can access or otherwise retrieve from the memory of the processing circuit 212 the features corresponding to each point cloud. To find the points in the reference frame point cloud that correspond to the points in the point cloud to be transformed, the processing circuit 212 uses L between the feature vectors in each point cloud. 2 The distance can be calculated. The L of the feature points in each point cloud. 2 Calculating the distance returns a list of initial (and potentially inaccurate) correspondences for each point. These correspondences indicate that a given data point corresponds to the same location on the object surface represented in each point cloud. After these initial correspondences are enumerated, the processing circuit 212 can apply a Random Sample Consensus (RANSAC) algorithm to identify and reject inaccurate correspondences. The RANSAC algorithm can be used to iteratively identify and fit correspondences between each point cloud using the list of initial correspondences.
[0073] The RANSAC algorithm can be used to determine which correspondences between features in both point clouds are relevant to the alignment process and which are incorrect correspondences (e.g., features in one point cloud that are incorrectly identified as corresponding to features in the point cloud being transformed or aligned). The RANSAC algorithm is iterative and can reject incorrect correspondences between the two point clouds until a satisfactory model is fitted. The resulting satisfactory model can identify each data point in the reference point cloud that has data points corresponding to the point cloud being transformed, and vice versa.
[0074] When implementing the RANSAC algorithm, the processing circuit 212 performs L between feature vectors. 2From the full set of initial correspondences identified using distance, a sample subset of feature correspondences containing the minimum correspondence can be randomly selected (e.g., pseudo-randomly). The processing circuit 212 can use the elements of this sample subset to compute the fitted model and its corresponding model parameters. The cardinality (concentration) of the sample subset can be kept to the minimum necessary to determine the model parameters. The processing circuit 212 can check which elements of the full set of correspondences match the model instantiated by the estimated model parameters. A correspondence can be considered an outlier if it does not fit the fitted model instantiated by the set of estimated model parameters within a certain error threshold (e.g., 1%, 5%, 10%) that defines the maximum deviation due to noise. The set of inliers obtained for the fitted model can be called the consensus set of correspondences. The processing circuit 212 can repeat the steps of the RANSAC algorithm until the consensus set obtained in a given iteration has sufficient inliers (e.g., above a predetermined threshold). The consensus set can be an exact list of correspondences between data points in each point cloud that fit the parameters of the RANSAC algorithm. The parameters for the RANSAC algorithm can be predetermined parameters. The consensus set can then be used in an iterative nearest nearest (ICP) algorithm to determine the transformed data structure.
[0075] The processing circuit 212 can implement the ICP algorithm using a consensus set of corresponding features generated by the RANSAC algorithm. Each corresponding feature in the consensus set can contain one or more data points in each point cloud. When implementing the ICP algorithm, the processing circuit 212 can match the nearest point in the reference point cloud (or a selected set) to a point closet point in the point cloud being transformed. The processing circuit 212 can then estimate the rotation and translation combination using the root mean square point distance metric minimization technique that best transforms each point in the point cloud to its match in the reference point cloud. The processing circuit 212 can transform the points in the point cloud to determine the amount of error between features in the point cloud, and iterate through this process to determine the optimal transformation values for the position and rotation of the point cloud being transformed. These output values can be combined into a transformation matrix, such as a 4x4 rigid body transformation matrix. This output transformation matrix can be used as the transformation data structure.
[0076] The processing circuit 212 uses a transformation data structure to transform a set of data points in the point cloud to be transformed, creating a transformed set of data points. The transformation data structure can include one or more transformation matrices. The transformation matrices contain transformation values that indicate changes in the position or rotation of points in the point cloud to be transformed. The processing circuit 212 can apply the values in the transformation data structure to each point in the point cloud to be transformed (for example, by offsetting its position or applying a rotation around a reference point) to change its respective position or rotation. This transformed point cloud can then be made into the same reference frame as the point cloud selected as the reference frame.
[0077] The processing circuit 212 can generate a combined set of data points, which includes a first set of data points and a transformed set of data points. In some implementations, the combined set of data points can include all data points in the reference point cloud and the transformed point cloud. Since data points represent captures of the same subject from two different angles, the processing circuit 212 can use the combined set of data points to assemble a more complete 3D point-based image. For example, each of the combined set of data points can represent a 3D image of the subject being analyzed. This 3D image can be converted into display data (e.g., a 3D point-based mesh rendered using 3D rendering techniques) and provided to the user interface 120 for display. For further processing, the processing circuit 212 can store the combined set of data points in its memory.
[0078] Referring now to Figure 4, a flowchart of Method 400 for aligning multiple depth cameras in an environment based on image data is shown. Method 400 can be implemented using the image processing system 100 or various devices and systems described herein, such as its components or modules mentioned herein in conjunction with Figures 1 and 2.
[0079] In 405, the first and second point cloud data points are accessed. The point cloud data points can include spatial coordinates and various parameters assigned to the spatial coordinates. For example, the point cloud data points can include spatial coordinates in a specific reference frame (e.g., Cartesian coordinates, cylindrical coordinates, spherical coordinates). The point cloud data points can represent information such as luminance or brightness, grayscale data, color data (e.g., RGB, CYMK), density, or various combinations thereof. In some embodiments, image data is processed to generate point cloud data points.
[0080] The first point cloud data points can correspond to a first pose relative to the subject, and the second point cloud data points can correspond to a second pose relative to the subject. The pose can represent the position and orientation of a device that detected the image data corresponding to the point cloud data points, such as an image capture device (e.g., a camera), an MRI machine, or a CT machine.
[0081] At least one of the first or second pose can be determined based on pose data received from each image capture device. For example, the pose can be determined based on data received from a position sensor (e.g., an accelerometer) coupled to the image capture device.
[0082] At least one of the first or second pose can be determined based on image data captured by one or more image capture devices in the environment surrounding the subject. For example, image data captured by the first image capture device can be processed to identify the second image capture device if the second image capture device is within the field of view of the first image capture device. The pose of the second image capture device can be determined from the image data in which the second image capture device has been identified.
[0083] In 410, a reference frame for the image data is determined based on at least one of the first pose or the second pose. The reference frame can be determined by comparing the points of the first point cloud data with the points of the second point cloud data. For example, the point cloud data, or features extracted from the point cloud data, can be compared, a match score can be generated based on the comparison (e.g., similarity can be determined), and an alignment transformation can be determined based on the match score.
[0084] In some embodiments, color data from point cloud data points can be used to determine a reference frame. For example, in addition to lightness or luminance values, data from one or more color channels assigned to each point cloud data point can be used when comparing point cloud data points, which can increase the accuracy of the match score generated based on the comparison.
[0085] In 415, at least one of the first point cloud data points or the second point cloud data points is transformed to align with a reference frame. For example, if the reference frame corresponds to a first pose (or a second pose), the second point cloud data points (or the first point cloud data points) can be transformed to align with the reference frame. In some embodiments, the reference frame is different from the first and second poses, and the first and second point cloud data points can be transformed to align with the reference frame, respectively.
[0086] In some embodiments, color data is not used when transforming point cloud data points. For example, the transformation may be applied to the spatial coordinates of the point cloud data points but not to the color data. The color data can be discarded before transforming the point cloud data points, or it can be retained in the point cloud data structure for later retrieval.
[0087] III. Systems and methods for segmenting the surface of medical images
[0088] Segmenting images, such as medical images or 3D images, can improve the efficiency of further calculations performed on the images in the image processing pipeline. However, it can be difficult to segment images in a computationally efficient way while retaining information relevant to the intended use of the image processing. This solution can implement a segmentation model that effectively retains anatomically relevant information and improves computational efficiency. For example, this solution can implement a segmentation model that effectively distinguishes the surface of an object, including its anatomical features, from its surrounding environment by identifying the density difference between the object's surface and its surrounding environment, or by using a machine learning model trained to classify point cloud data points as corresponding to the object's surface.
[0089] Figure 5 shows a method 500 for segmenting the surface of a medical image. Method 500 can be carried out using various devices and systems described herein, such as the image processing system 100.
[0090] In step 505, point cloud data points are accessed. These point cloud data points can correspond to the surface of a subject. For example, the point cloud data points can correspond to medical images or 3D images of a subject detected by a 3D camera, MRI device, or CT device.
[0091] In 510, multiple point cloud data points are applied as input to a segmentation model. The segmentation model may include one or more functions that generate segments based on density data indicated by the image data of the point cloud data points (for example, segmenting a surface from air based on the brightness data of the point cloud data points). A segment can be a group of pixels or data points that share similar characteristics, such as a region of a group of pixels or data points that represent the surface of an object from the air surrounding the object.
[0092] In some embodiments, the segmentation model includes a machine learning model. The machine learning model can be trained using training data that includes a given image (e.g., a given image data) and labeled segments associated with the image (e.g., a given segment). For example, the machine learning model may include a neural network trained to generate output data containing one or more segments based on input data that includes image data such as point cloud data points.
[0093] In 515, multiple segments are generated. These segments can be generated using a segmentation model. The segments can correspond to the surface of an object in response to input to the segmentation model.
[0094] In step 520, multiple segments are output. These segments can be used to generate a 3D surface model of the subject's surface.
[0095] IV. System and method for generating 3D surface models from medical images based on segmentation.
[0096] 3D surface models based on medical images such as MRI or CT scans can be used for a variety of applications, including surgical navigation, planning, and instrument tracking. To effectively render and perform image processing calculations using 3D surface models, it may be useful to generate the 3D surface model from the segmentation of the underlying 3D image data. For example, triangular or quadrilateral 3D surface models can be generated from segments produced using segmentation models as described herein. Point cloud data can be generated from 3D surface models or directly from segmentation data (e.g., from segments generated by the segmentation model).
[0097] Figure 6 shows a method 600 for generating a 3D surface model from a medical image based on segmentation. Method 600 can be carried out using various devices and systems described herein, such as the image processing system 100.
[0098] In 605, multiple segments are accessed. A segment can correspond to the three-dimensional surface of an object. For example, a segment can be generated from image data representative of the surface, such as 3D point cloud data points corresponding to detected images of the surface.
[0099] In step 610, a 3D model of multiple segments is generated. The 3D model can represent the 3D surface of the subject. The 3D model can be generated as a triangular or quadrilateral 3D surface model, for example, by connecting the points of the segments to form a surface portion with three or four sides.
[0100] In step 615, a point cloud data structure is generated. This point cloud data structure can represent the three-dimensional surface of the subject. For example, a point cloud data structure can be generated by sampling points that form a surface portion of a surface generated using segments. The point cloud data structure can contain multiple point cloud data points corresponding to a surface portion.
[0101] In step 620, a point cloud data structure is output. This point cloud data structure can be output to represent a 3D surface model, for example, to match the 3D surface model with image data from other modalities (e.g., 3D image data).
[0102] Figure 7 shows a method 700 for generating a 3D surface model from a medical image based on segmentation. Method 700 can be carried out using various devices and systems described herein, such as the image processing system 100. Method 700 can be similar to Method 600, but generates a point cloud representation of the surface of the subject directly from surface segmentation (e.g., not via a 3D surface model).
[0103] In 705, multiple segments are accessed. A segment can correspond to the three-dimensional surface of an object. For example, a segment can be generated from image data representative of the surface, such as 3D point cloud data points corresponding to detected images of the surface.
[0104] In step 710, a point cloud data structure is generated. This point cloud data structure can represent the three-dimensional surface of the subject. For example, a point cloud data structure can be generated by sampling points using segments. The point cloud data structure can contain multiple point cloud data points corresponding to segments.
[0105] In step 715, a point cloud data structure is output. This point cloud data structure can be output to represent the surface of an object, such as for matching a 3D surface with image data from other modalities (e.g., 3D image data).
[0106] V. System and method for downsampling point clouds generated from 3D surface models to improve image alignment efficiency
[0107] By comparing and matching point cloud data from multiple modalities, point cloud data can be aligned and used for various purposes. However, point cloud data, 3D image data, and 3D models representing the 3D surface of an object are often large and complex, making them difficult to process efficiently in image processing pipelines. For example, Microsoft Kinect can generate 9 million point cloud data points per second. The execution time of image processing operations using 3D point cloud data can be directly related to the density of the 3D point cloud data (including being slower than linear time execution). Furthermore, various image processing operations can be highly sensitive to environmental factors that affect image data, such as lighting, shading, occlusion, and pose.
[0108] This solution enables effective point cloud resampling (e.g., downsampling) in a way that effectively preserves anatomical or other physically relevant information, such as object contours and edges, and the relationships between point cloud data points, while reducing the computational complexity involved in further image processing. The solution can reduce point cloud density to enable faster image processing while preserving relevant information. Thus, this solution can enable real-time image processing that meets target performance criteria such as sub-millimeter precision (e.g., preserving image data where the distance between point cloud data points is less than 1 millimeter).
[0109] Returning to Figures 1 and 2, the image processing system 100 can resample or downsample the point cloud to improve computational efficiency in the image processing pipeline without significantly reducing the accuracy of image registration. By selectively reducing the total number of points that need to be processed to achieve the desired 3D image registration, the image processing system 100 can improve the speed of the 3D image registration technique while reducing the overall computational requirements. The image processing system 100 can perform contour-based resampling of the point cloud data, thereby reducing the density of the point cloud in a 3D image (e.g., captured by an image capture device, or extracted from a 3D medical image such as a CT scan or MRI image) while preserving relevant points and relationships between points, such as contours and edges. Relevant parts of the point cloud are those that have a significant impact or importance on the image registration process described herein.
[0110] The processing circuit 212 can access a set of data points corresponding to a point cloud representing the surface of an object. The 3D data points constituting the point cloud (e.g., extracted from a 3D medical image) can be multidimensional data points describing a set of coordinates in a particular reference frame. For example, the 3D data points of the point cloud (e.g., a collection of data points) can include three coordinates (e.g., Cartesian coordinates, cylindrical coordinates, etc.). A data point can correspond to a single pixel captured by a 3D camera and can be at least a three-dimensional data point (e.g., containing at least three coordinates, each corresponding to a spatial dimension). In some implementations, a data point can correspond to a point or vertex in a 3D medical image, such as a CT scan model or an MRI image. In some implementations, the data points accessed (e.g., retrieved from one or more data structures in memory by the processing circuit 212) can be a combined set of data points generated from point clouds captured from two or more image capture devices 104. In some implementations, the processing circuit 212 accesses or receives a set of data points from at least one of the image capture devices 104, for example in real time, when the image capture device 104 captures a 3D image of a subject or environment.
[0111] Accessing a set of data points may include generating a point cloud representing the surface of an object using 3D image data. In some implementations, the processing circuit 212 may receive or acquire a 3D image or model containing a set of 3D data points. These data points, along with other relevant point data (e.g., color, temperature at each point, and other factors), may be extracted from the 3D image or model. For example, in the case of a 3D model (e.g., a CT scan mesh or an MRI image model), the processing circuit 212 may use the data present in the 3D model to extract one or more slices or vertices. In some implementations, the processing circuit 212 may generate a point cloud from a 3D model using the steps of method 600 or method 700 described herein in relation to Figures 6 and 7.
[0112] The processing circuit 212 can apply a response function to a set of data points to assign a set of response values to each set of data points. The response function can be a function that takes a set of data points (e.g., from a point cloud) and generates response values based on the relationships between the input points. The response function can generate response values for each input data point by applying one or more matrix operations to the points in the point cloud. The response values can be weighted values indicating whether each data point is part of a feature of interest, such as a contour. For example, contours are more complex structures and can therefore be considered more relevant to image alignment or alignment of two different point clouds. To determine the response value for each data point in the point cloud, the processing circuit 212 can perform an analysis of each data point with respect to one or more adjacent data points. For example, the processing circuit 212 can apply the response function to generate response values with greater weights based on data points that have a greater relevance to features of interest, such as contours, edges, segments, or other image features that are likely to represent anatomical features or surgical instruments.
[0113] For example, the processing circuit 212 can apply a response function that includes a graph filter that can be applied to a graph data structure generated from data points in a point cloud. The graph filter can be a function that takes a graph data structure as input and produces an output that is indexed by the same graph data. The graph data structure can be an adjacency matrix that shows the relationships between different nodes in the graph (e.g., point cloud data points). The adjacency matrix can be a square matrix having values corresponding to the edge weights between nodes in the graph.
[0114] To generate a graph data structure, the processing circuit 212 can generate an adjacency matrix with edge weights W as follows:
number
[0115] Here, W can be the adjacency matrix between points in the point cloud, xi and xj correspond to the i-th and j-th data points in the point cloud, respectively, and sigma is an adjustable parameter for the graph filter response function. In some implementations, the processing circuit 212 can generate the adjacency matrix such that the edge weights for data points are set to zero if the distance between points connected by edges is greater than a predetermined threshold. Using the above adjacency matrix, the processing circuit 212 can utilize a graph filter function.
number
[0116] Here, h(A) is a graph filter, and A is a graph shift operator as follows:
number
[0117] Here, W is the adjacency matrix outlined above, and D is a diagonal matrix, D i,j This can be the sum of all elements in row i of W. Using the above graph filter, the processing circuit 212 can define the response function as follows:
number
[0118] The response function operates across the entire set of data points X, and it is possible to assign a weight value to each data point X that indicates the likelihood that each point is part of the contour.
[0119] To improve computational efficiency, the processing circuit 212 can use the set of data points to generate a k-dimensional binary tree (sometimes referred to herein as a "kd tree") to generate the adjacency matrix W. In some implementations, the processing circuit 212 does not create a kd tree, but instead generates an adjacency matrix that contains non-zero edges from each data point in the point cloud to each other data point in the point cloud.
[0120] The processing circuit 212 uses point cloud data points to generate a k-dimensional tree as a binary tree, sorting the point cloud data points into nodes based on parameters of the point cloud data points such as spatial coordinates. Furthermore, it is possible to generate a k-dimensional tree using brightness, luminance, color, density, or other parameters assigned to each point cloud data point. The adjacency matrix W of a particular point cloud data point gives weights W to each point cloud data point that is k or more away from the particular point cloud data point in the k-dimensional tree. i,j It can be generated based on a k-dimensional tree such that is set to zero. The number of dimensions k can correspond to the number of different parameters used to generate the k-dimensional tree (e.g., three spatial coordinate dimensions or three color dimensions). The number of dimensions k can be a predetermined parameter that can be used to control the computational demands associated with generating a response function and applying the response function to point cloud data points.
[0121] In some implementations, when assembling the adjacency matrix, the processing circuit 212 can determine the Euclidean distance based on at least one color channel in a pair of data points in the set of data points (for example, based only on position data independent of any color channel data). i -x j(Instead of determining the color). In some implementations, the processing circuit 212 can calculate the Euclidean distance based on additional red, green, or blue color channel data (or, optionally, cyan, yellow, magenta, and lightness color channel data) contained in each data point. For example, the processing circuit 212 can determine or calculate the Euclidean distance using each color channel as three additional independent distances, which can range, for example, from 0 to 1. The color values can be stored as 8-bit color data (one 8-bit number for each color channel), for example, ranging from 0 to 255, either as part of the data point or associated with the data point.
[0122] Generating a graph filter may involve identifying a brightness parameter for a set of data points. The brightness parameter could be, for example, a parameter describing a channel included in or calculated from the color value of a data point in the point cloud. For instance, if a data point stores the CYMK color channel, the data point could use K as its brightness value. In some implementations, such as when a data point stores the RGB color channel, the processing circuit 212 can calculate the brightness value for each data point by calculating a weighted average of the color channels at each data point. The processing circuit 212 can compare the brightness values generated or accessed for data points in the point cloud and determine whether the variation between a significant number of brightness values (e.g., greater than 10%, 15%, 40%, 50%, or any other predetermined amount) is greater than a predetermined threshold. If this variation is greater than the predetermined threshold, the surface represented by the point cloud cannot be uniformly illuminated, and the processing circuit 212 can use a non-color-based graph filter variant. In contrast, if the fluctuation is not greater than a predetermined threshold, the surface can be illuminated uniformly and clearly, and the processing circuit 212 can utilize variations of a color-based graph filter.
[0123] The processing circuit 212 can select a subset of data points using a selection policy and a set of response values corresponding to each data point in the point cloud. The selection policy can, for example, indicate which points are relevant to further processing operations and which can be removed from the entire point cloud without sacrificing image alignment accuracy. The selection policy executed by the processing circuit 212 can, for example, compare the response value of each data point with a predetermined threshold. If the response value is greater than or equal to the predetermined threshold, the selection policy can indicate that the data point should not be culled (downsampled) from the point cloud. If the response value is less than the predetermined threshold, the selection policy can indicate that the point should be removed from the point cloud or downsampled. Thus, the selection policy can be configured to select a subset of data points in the point cloud that sufficiently correspond to one or more contours on the surface of an object represented by the point cloud.
[0124] In some implementations, the selection policy can remove points on a pseudo-random basis. In addition to removing data points based on a calculated response value, the processing circuit 212 can further improve performance by uniformly removing data points from the point cloud. For example, to remove data points from the entire point cloud in a pseudo-random and uniform manner, the selection policy may include an instruction to generate a pseudo-random number for each data point in the point cloud. The pseudo-random number can be a value within a range of values, such as between 0 and 100. The processing circuit 212 can determine whether the pseudo-random value of a data point is less than a predetermined threshold. Following the previous example, if the selection policy indicates that approximately 25% of the data points should be uniformly culled from the point cloud, the selection policy may include an instruction to remove or cull data points that have been assigned a pseudo-random value less than 25.
[0125] Once the selection policy indicates which points can be removed from the point cloud without sacrificing system accuracy, the processing circuit 212 can generate a data structure containing a selected subset of the data points that were not culled by the selection policy. This data structure can be smaller than the data structure containing the entire set of data points in the point cloud. In some implementations, the data points in the subset can be assigned an index value corresponding to their respective positions in the data structure. The processing circuit 212 can then store the generated data structure of the data point subjects in its memory.
[0126] Figure 8 shows a method 800 for downsampling a point cloud generated from a 3D surface model to improve image alignment efficiency. Method 800 can be carried out using various devices and systems described herein, such as the image processing system 100. Method 800 can be used to perform contour-based resampling of point cloud data, which can reduce the density of the point cloud data while preserving related points and inter-point relationships such as contours and edges.
[0127] In 805, multiple point cloud data points are accessed. These point cloud data points can correspond to the surface of a subject. For example, the point cloud data points can correspond to medical images or 3D images of a subject detected by a 3D camera, MRI device, or CT device.
[0128] In 810, a response function based on a graph filter is applied to each point cloud data point of multiple data points. The response function can be applied to assign a response value to each point cloud data point of multiple point cloud data points.
[0129] For example, the graph filter can be Equation 2 ([Equation 2]), where A is the graph shift operator Equation 3 ([Equation 3]). W can be the adjacency matrix of the point cloud data points such that W has the edge weight Equation 1 ([Equation 1]), and D is D i,i is a diagonal matrix that is the sum of all elements in the i-th row of W. The response function can be defined as Equation 4 ([Equation 4]).
[0130] In some embodiments, the graph filter is generated using a k-dimensional tree, thereby reducing the computational requirements. For example, the k-dimensional tree can be generated using the point cloud data points as a binary tree that sorts the point cloud data points into nodes based on parameters of the point cloud data points such as spatial coordinates, and can further be generated using the brightness, luminance, color (hue), density, or other parameters assigned to each point cloud data point. The adjacency matrix W of a particular point cloud data point can be generated based on the k-dimensional tree such that the weight W i,j is set to zero for each point cloud data point more than k-nearest neighbors away from the particular point cloud data point in the k-dimensional tree. The number of dimensions k can correspond to the number of different parameters used to generate the k-dimensional tree (e.g., 3 spatial coordinate dimensions or 3 color dimensions). The number of dimensions k can be a predetermined parameter used to control the computational requirements associated with generating the response function and applying the response function to the point cloud data points. The parameter σ 2 can also be a predetermined parameter. The response function is applied to each point cloud data point to generate a response value corresponding to each respective point cloud data point.
[0131] In 815, a subset of multiple point cloud data points is selected. The subset can be selected using a selection policy and multiple response values. For example, the subset can be selected based on response values. The selection policy can perform weighted selection, in which each point is selected for the subset based on the response value assigned to the point (for example, based on whether the response value satisfies a threshold). In some embodiments, the selection policy performs random weighted selection, such as by randomly generating one or more thresholds for comparing response values.
[0132] In 820, a subset of multiple point cloud data points is output. This subset can be output for further image processing operations such as feature matching and point cloud alignment, which can be improved due to the reduced density of the point cloud data point subset. Figure 14 shows k=10 and σ 2 This figure shows an example of image 1400, which was resampled according to this solution at =0.0005, resulting in 19.31% of the points retained in 1405 and 5.30% of the points retained in 1410. As shown in Figure 13, this solution can reduce the density of point cloud data points by approximately 4 to 20 times while retaining relevant features such as edges and contours of the subject.
[0133] As described above, the point cloud data points x used to determine the edge weights of the adjacency matrix W are i and x j The distance between is the Euclidean distance (e.g., x) based on the spatial coordinates of the point cloud data points. i -x j L to determine 2 (norm) can be used. Thus, color data is not used to generate the adjacency matrix W. In some embodiments, point cloud data point x i and x jThe distance between points can be further determined based on color data from one or more color channels assigned to the point cloud data points (in addition to spatial coordinates). For example, when determining the Euclidean distance between point cloud data points, one or more color data values from one or more color channels (e.g., red, green, and blue channels) can be used as an additional dimension, in addition to the spatial dimension. The color data can be normalized to a specific scale (e.g., a scale from 0 to 1), which can be the same scale as the spatial dimension being compared, or a different scale so that different weights are applied to the spatial distance and color distance.
[0134] Using color data (e.g., implementing color recognition filters) can be effective in various situations, although it may increase the computational complexity of resampling point cloud data points. For example, text (e.g., colored text) can be sampled more frequently when using color data because it can be detected that the text forms part of the plane on which it lies, rather than forming the edges or contours of the subject. Furthermore, color data captured by the image capture device used to generate point cloud data can be highly dependent on the lighting present when the image data was captured. Therefore, factors such as lighting, shading, and occlusion can affect the effectiveness of using color data. If the point clouds being resampled downstream for comparison have similar or uniform lighting, using color data can improve the estimation of the transformation. Also, using color data can mitigate image capture device issues such as flying pixels (where pixels at the edges of the subject may take the same color).
[0135] In some embodiments, the response function can be applied in a first operating mode that uses color data or a second operating mode that does not use color data. The operating mode can be selected based on information from the image data, such as the brightness parameter of the point cloud data points. For example, the first operating mode can be selected if the brightness parameter indicates that the illumination of the point cloud data points is greater than a uniformity threshold measure (e.g., based on a statistical measure of brightness such as the median, mean, or standard deviation of brightness).
[0136] In some embodiments, one or more preliminary filters are applied to the point cloud data points before resampling using a response function and a graph filter. For example, a voxel grid filter can be applied to the point cloud data points, and a response function can be applied to the output of the voxel grid filter, which can improve the overall effectiveness of resampling. Applying a voxel grid filter may include generating a grid of side length l (e.g., a 3D grid where each voxel acts as a container for a specific spatial coordinate) on top of the point cloud data points, assigning each point cloud data point to its respective voxel based on the spatial coordinate of the point cloud data points, and then generating updated point cloud data points with the centroid (center of mass) position and centroid color of each point cloud data point assigned to each respective voxel. The voxel grid filter can enable a uniform density of point cloud data points (e.g., compared to a density that decreases as the distance from the camera increases by a method of detecting image data that the camera can retain in some random resampling method) and can smooth out local noise variations.
[0137] VI. System and method for detecting contour points from downsampled point clouds and prioritizing the analysis of contour points.
[0138] To perform further image processing operations, features of the resampled point cloud, such as contour points, can be identified. For example, identifying features can enable feature matching and point cloud alignment based on feature matching. Effective feature selection (e.g., preserving physically relevant features) can reduce the computational requirements for performing alignment while maintaining the target performance and quality of the alignment. In some embodiments, features are selected using keypoint detection methods such as scale-invariant feature transformation (SIFT) or speed-up robust feature (SURF) algorithms.
[0139] Figure 9 shows a method 900 for detecting contour points from a downsampled point cloud and prioritizing the analysis of contour points. Method 900 can be carried out using various devices and systems described herein, such as the image processing system 100. In particular, the processing circuit 212 can perform any of the operations described herein.
[0140] In step 905, multiple point cloud data points are accessed. These point cloud data points can correspond to the surface of a subject. For example, the point cloud data points can correspond to medical images or 3D images of a subject detected by a 3D camera, MRI device, or CT device.
[0141] In step 910, a feature vector is generated for a point cloud data point. A feature vector can be generated for each point cloud data point in at least one subset of multiple point cloud data points. The feature vector can be based on the point cloud data point and multiple adjacent point cloud data points.
[0142] In some embodiments, feature vectors are generated by assigning a plurality of rotation values between a point cloud data point and a plurality of adjacent point cloud data points to each of a plurality of spatial containers representing the feature vector. For example, feature vectors can be generated using a fast point feature histogram (FPFH). Rotation values on each spatial axis (e.g., theta, phi, and alpha angles) can be determined for each spatial container. Each spatial container can be assigned adjacent point cloud data points within a predetermined radius of the point cloud data point; for example, 11 spatial containers can be used, resulting in a vector of length 33 (each of the 11 spatial containers is assigned 3 rotation angles). Feature vector generation takes O(n*k) time. 2 The order can be as follows: where n is the number of point cloud data points and k is the number of neighboring points within the radius of a point cloud data point.
[0143] In some embodiments, feature vectors are generated by determining a reference frame for a point cloud data point using adjacent point cloud data points within a given radius of the point cloud data point, and then generating feature vectors based on the reference frame and multiple spatial containers. For example, feature vectors can be generated using the sign number (SHOT) of the histogram orientation. The reference frame can be a 9-dimensional reference frame determined using adjacent point cloud data points. A grid (e.g., an isotropic grid) can be generated to identify multiple spatial containers, such as a grid with 32 containers and 10 angles assigned to each container, which can result in a feature vector of length 329 (320 dimensions describing the containers and 9 dimensions for the reference frame). In some embodiments, the feature vectors include color data, which can be assigned to additional containers. The generation of feature vectors can be of order O(n*k).
[0144] The processing performed to generate feature vectors can be selected based on factors such as accuracy and computation time. For example, generating feature vectors using a reference frame can reduce computation time while maintaining similar performance in terms of accuracy and fidelity in retaining relevant information about the subject. As mentioned above regarding resampling, the use of color data can be affected by the uniformity of ambient lighting, so feature vector generation can be performed in various operating modes in which color data responsive to lighting can be used or not. For example, using color data may be useful when operating in a scene alignment mode that aligns point cloud data to a scene (e.g., medical scan data aligned to 3D image data of a subject).
[0145] In 915, each feature vector is output. For example, the feature vectors can be output to perform feature matching, and computational efficiency can be improved due to the way in which this solution generates feature vectors. The feature vectors can be stored in one or more data structures in memory, such as the memory of the processing circuit 212.
[0146] VII. System and method for dynamically allocating processing resources in a parallel processing environment for image alignment and point cloud generation calculations.
[0147] The image processing pipeline described herein can improve the computation time of image alignment and point cloud generation by using parallel processing operations. For example, the processing circuit 212 can allocate processing resources such as separate threads, separate processing cores, or separate virtual machines (e.g., controlled by a hypervisor) and can be used to perform parallel processing such as point cloud resampling and feature vector determination (e.g., resampling two different point clouds in parallel, or generating feature vectors from point clouds resampled in parallel). The processing circuit 212 may include other computing devices or computing machines such as graphics processing units, field-programmable gate arrays, computing clusters with multiple processors or computing nodes, or other parallel processing devices. Based on the current demand for processing resources and the size of the processing jobs, the processing circuit 212 can dynamically allocate certain processing jobs to processing machines specialized in parallel processing and assign other processing jobs to processing machines specialized in sequential processing.
[0148] In some implementations, parallel processing can be performed based on the type of image data or image stream capture source. For example, DICOM data (e.g., CT data, MRI data) may have certain characteristics that differ from 3D image data detected by a depth camera. This solution can assign different processing threads to each point cloud received from each source, and can allocate larger or smaller processing resources to different image source modalities based on expected computational demands to maintain synchronization between processing performed in each modality. By performing point cloud calculations in a tightly scheduled manner across parallel and sequential computing devices as appropriate, the processing circuit 212 can perform accurate image alignment in real time.
[0149] Referring here to Figure 15, an exemplary flowchart of Method 1500 for allocating processing resources to different computing machines in order to improve the computational performance of point cloud alignment calculations is shown. Method 1500 can be implemented, executed, or otherwise carried out by the image processing system 100, in particular at least the processing circuit 212, the computer system 1300 described herein in relation to Figures 13A and 13B, or any other computing device described herein.
[0150] In 1505, the processing circuit 212 can identify a first processing device having a first memory and a second multiprocessor device having a second memory. The processing circuit 212 can include different processing machines, such as processing machines specialized for parallel computation (e.g., clusters of computing nodes, graphics processing units (GPUs), field-programmable gate arrays (FPGAs), etc.) and computing machines specialized for sequential computation (e.g., high-frequency single-core or multi-core devices, etc.). Each of these devices can include a memory bank or computer-readable memory for processing operations. In some implementations, the memory bank or other computer-readable medium can be shared among different processing devices. Furthermore, the memory of some processing devices can have a higher bandwidth than the memory of other processing devices.
[0151] In some implementations, the processing circuit 212 of the image processing system 100 can be modularized. For example, specific processing devices and memories can be added to or removed from the system via one or more system buses or communication buses. Communication buses may include PCI Express and Ethernet, etc. Further descriptions of different system buses and their operations will be given later in relation to Figures 13A and 13B. The processing circuit 212 can query one or more communication or system buses to identify and enumerate the processing resources available for processing point cloud data. For example, the processing circuit 212 can identify one or more parallel processing units (e.g., a cluster of computing nodes, a GPU, an FPGA, etc.) or sequential processing units (e.g., a high-frequency single-core or multi-core device, etc.). Once the devices are identified, the processing circuit 212 can communicate with each processing device to determine the parameters and memory banks, maps, or regions associated with each device. Parameters may include processing power, cores, memory maps, configuration information, and other processing-related information.
[0152] In 1510, the processing circuit 212 can identify a first processing job for the first point cloud. After the processing device of the processing circuit 212 is identified, the processing circuit 212 can begin performing the processing tasks detailed herein. The processing circuit 212 can execute instructions for processing points related to point cloud information, composite point clouds (e.g., global scene point clouds), or 3D image data from the image capture device 104. For example, one such processing job is to compute a k-dimensional tree for the point cloud, as described herein. Other processing jobs that can be identified by the processing circuit 212 may include graph filter generation, calculation of Euclidean distance, determination of overall brightness value, downsampling of point cloud, calculation of normal map, generation of features for point cloud, and conversion of 3D medical image data (e.g., segmented) into a 3D image that can be represented as a point cloud. In some implementations, the jobs or operations described herein have a specific order that the processing circuit 212 can identify.
[0153] A processing job identified by the processing circuit 212 may include job information, such as metadata about the information to be processed by the processing circuit 212. The job information may also include one or more data structures or areas in computer memory that contain the information to be processed when the job is executed. In some implementations, the job information may include a pointer to an area of memory containing that information to be processed when the job is executed. Other job information identified by the processing circuit 212 may include instructions that, when executed by the identified processing device, cause the processing device to perform a computational task on the information to be processed in order to perform the processing job. In some implementations, a first processing job may be identified in response to receiving point cloud data from at least one image processing device 104.
[0154] In step 1515, the processing circuit 212 may decide to assign the first processing job to a second multiprocessor device. A particular processing job can be performed more quickly on different, higher-performance processing hardware. For example, if a processing job involves calculations for feature detection in a point cloud, and these calculations involve many calculations that can be performed in parallel, the processing circuit 212 may decide to perform the processing job on a parallel computing device. A processing device(s) may be selected for a particular job based on information about the job to be performed. Such information may include the number of data points in the point cloud to be processed by the job, the utilization of processing devices that are part of the processing circuit 212, or the overall processing complexity of the job.
[0155] If the number of points to be processed in a particular job exceeds a threshold, and the processing job is not a sequential-based algorithm, the processing circuit 212 may decide to process the job on a multiprocessor device such as a GPU, cluster, or FPGA, if any of the devices described below are available. Otherwise, if the number of points to be processed is below a predetermined threshold (e.g., much smaller or fewer points), such as when the point cloud is a 3D medical image, the processing circuit 212 may decide to process the job on a processing device that specializes in sequential processing. In another example, if one of the processing devices is overutilized (e.g., utilization is greater than a predetermined threshold), the processing circuit 212 may decide to assign the job to a different computing device. Conversely, if the processing circuit 212 determines that a processing device is underutilized (e.g., utilization is greater or less than a predetermined threshold) and suitable for performing the processing job in a reasonable amount of time, the processing circuit may decide to assign the processing job to that processing device. If the complexity of the job processing exceeds a predetermined threshold (for example, a very high computation sequence), the processing circuit 212 can assign the job to a computing device suitable for the complexity, such as a multiprocessor device. If the processing circuit 212 determines that the processing job should be performed by a processing device specialized in sequential computation operations, the processing device can perform step 1520A. If the processing circuit 212 determines that the processing job should be performed by a second multiprocessor device specialized in parallel computing operations, the processing device can perform step 1520B.
[0156] In 1520A and 1520B, the processing circuit 212 can allocate information for a first processing job, including a first point cloud, to a second memory or a first memory. Once a processing device is determined for a particular job, the processing circuit 212 can allocate job-specific resources for performing the job to the appropriate processing device. If the processing circuit 212 decides to perform the job using one or more parallel processing devices, the processing circuit 212 can send job-specific data, such as point clouds or other related data structures, to the memory of the parallel processing device, or allocate it in other ways. If the job-specific resources reside in memory at a location shared with the parallel processing device, the processing circuit 212 can provide the processing device with a pointer to the location of the job-related data. Otherwise, the processing circuit 212 can prepare the job for execution by transmitting the processing-specific data to the memory of the parallel processing device, or otherwise copying it (e.g., via direct memory access (DMA)). The processing circuit 212 can communicate with any number of parallel processing devices or resources, or perform any of the operations disclosed herein, using one or more application programming interfaces (APIs) such as NVIDIA CUDA or OpenMP.
[0157] If the processing circuit 212 decides to execute a job using a sequential processing device, it may send or otherwise allocate job-specific data, such as a point cloud or any other relevant data structure, to the memory of the sequential processing device. If the job-specific resources reside in memory at a location shared with the parallel processing device, the processing circuit 212 may provide the processing device with a pointer to the location of the job-related data. Otherwise, the processing circuit 212 may prepare the job for execution by transmitting the processing-specific data to the memory of the sequential processing device or by other means (e.g., via direct memory access (DMA)). The processing circuit 212 may communicate with any number of sequential processing devices or resources or perform any of the operations disclosed herein using one or more application programming interfaces (APIs), such as OpenMP.
[0158] In 1525, the processing circuit 212 can identify a second processing job for the second point cloud. While another job is assigned or executed by another computing device, the processing circuit 212 can process points related to the point cloud information, composite point cloud (e.g., global scene point cloud), or 3D image data of the image capture device 104 and execute instructions to assign these jobs to other computing devices. For example, one such processing job is to compute a k-dimensional tree for the point cloud, as described herein. Other processing jobs that can be identified by the processing circuit 212 may include graph filter generation, calculation of Euclidean distance, determination of overall brightness value, downsampling of the point cloud, calculation of normal map, generation of features for the point cloud, conversion of 3D medical image data (e.g., segmented) into a 3D image that can be represented as a point cloud, or any of the other processing operations described herein. In some implementations, the jobs or operations described herein have a specific order in which the processing circuit 212 can identify them. In some implementations, the processing circuit 212 can stop processing a job if a previous job on which the current job depends is still being processed by the computing machine of the processing circuit 212.
[0159] A processing job identified by the processing circuit 212 may include job information, such as metadata about the information to be processed by the processing circuit 212. The job information may also include one or more data structures or areas in computer memory that contain the information to be processed when the job is executed. In some implementations, the job information may include a pointer to an area of memory containing that information to be processed when the job is executed. Other job information identified by the processing circuit 212 may include instructions that, when executed by the identified processing device, cause the processing device to perform a computational task on the information to be processed in order to perform the processing job. In some implementations, a processing job may be identified in response to receiving point cloud data from at least one image processing device 104, or in response to the completion of another job.
[0160] At 1530, the processing circuit 212 may decide to assign the second processing job to the first processing device. If the identified processing job has high-order complexity and does not have many operations that can be performed in parallel, the processing circuit 212 may decide to assign the processing job to a sequential computing device with a high clock frequency. This decision may also be made based on information about the job to be performed. Such information may include the number of data points in the point cloud to be processed by the job, the utilization of processing devices that are part of the processing circuit 212, or the overall processing complexity of the job. In some implementations, jobs may be assigned to specific computing devices based on priority. For example, a job with a high priority for sequential computing devices will be assigned to sequential computing devices before being assigned to non-sequential computing devices.
[0161] If the number of points to be processed in a particular job exceeds a threshold, and the processing job is not a sequential-based algorithm, the processing circuit 212 may decide to process the job on a multiprocessor device such as a GPU, cluster, or FPGA, if any of the devices described below are available. Otherwise, if the number of points to be processed is below a predetermined threshold (e.g., much smaller or fewer points), such as when the point cloud is a 3D medical image, the processing circuit 212 may decide to process the job on a processing device that specializes in sequential processing. In another example, if one of the processing devices is overutilized (e.g., utilization is greater than a predetermined threshold), the processing circuit 212 may decide to assign the job to a different computing device. Conversely, if the processing circuit 212 determines that a processing device is underutilized (e.g., utilization is greater or less than a predetermined threshold) and suitable for performing the processing job in a reasonable amount of time, the processing circuit may decide to assign the processing job to that processing device. If the complexity of a job is greater than a predetermined threshold (for example, a very high computation sequence), the processing circuit 212 can assign the job to a computing device suitable for the complexity, such as a multiprocessor device. Furthermore, if a particular algorithm or process includes a large number of operations that cannot be performed in parallel, the processing circuit 212 can assign the job to a sequential processing device. If the processing circuit 212 determines that the processing job should be performed by a processing device specialized in sequential computation, the processing device can perform step 1535A. If the processing circuit 212 determines that the processing job should be performed by a second multiprocessor device specialized in parallel computation, the processing device can perform step 1535B.
[0162] In steps 1535A and 1535B, the processing circuit 212 can allocate information for a second processing job to the first or second memory (steps 1535A and 1535B). Once a processing device is determined for a particular job, the processing circuit 212 can allocate job-specific resources for performing the job to the appropriate processing device. If the processing circuit 212 determines that the job will be performed using a sequential processing device, the processing circuit 212 can send or otherwise allocate job-specific data, such as point clouds or other relevant data structures, to the memory of the sequential processing device. If the job-specific resources reside in memory at a location shared with the parallel processing device, the processing circuit 212 can provide the processing device with a pointer to the location of the job-related data. Otherwise, the processing circuit 212 can prepare the job for execution by transmitting or otherwise copying the processing-specific data to the memory of the sequential processing device (e.g., via direct memory access (DMA)). The processing circuit 212 can communicate with any number of sequential processing devices or resources, or perform any of the operations disclosed herein, using one or more application programming interfaces (APIs) such as OpenMP.
[0163] If the processing circuit 212 determines that a job will be executed using one or more parallel processing devices, the processing circuit 212 may send or otherwise allocate job-specific data, such as a point cloud or any other relevant data structure, to the memory of the parallel processing device. If job-specific resources reside in memory at a location shared with the parallel processing device, the processing circuit 212 may provide the processing device with a pointer to the location of the job-related data. Otherwise, the processing circuit 212 may prepare the job for execution by transmitting or otherwise copying the processing-specific data to the memory of the parallel processing device (e.g., via direct memory access (DMA)). The processing circuit 212 may communicate with any number of parallel processing devices or resources, or perform any of the operations disclosed herein, using one or more application programming interfaces (APIs), such as NVIDIA CUDA or OpenMP.
[0164] In 1540, the processing circuit 212 can transfer instructions to the first processing device and the second multi-processing device to perform their assigned processing jobs. To have the appropriate processing device perform processing for a particular job, the processing circuit 212 can transfer instructions related to each job to the appropriate computing device (e.g., via one or more system buses or communication buses, etc.). In some implementations, the instructions transferred to the computing device may include device-specific instructions. For example, if a GPU device is selected for a processing job, the processing circuit 212 can identify and send GPU-specific instructions (e.g., CUDA instructions, etc.) to execute the job. Similarly, if a standard CPU device (e.g., a sequential processing device, etc.) is selected, the processing circuit 212 can identify and send CPU-specific instructions to perform the processing job. Processing instructions for each computing device may be included in the job information identified by the processing circuit 212. When a processing job is completed on a particular processing device, the processing circuit 212 can identify a signal from that device indicating that the job has been completed. Subsequently, the processing circuit 212 can identify the memory region containing the results of calculations performed as part of the processing job and copy them to another region of working memory for further processing.
[0165] VIII. System and method for aligning medical image point clouds with global scene point clouds
[0166] By downsampling the image 208 captured from the capture device 104 to improve computational efficiency and extracting feature vectors from the point cloud, the processing system can align 3D medical images, such as 3D models generated from CT scan images or MRI images, to the point cloud. By aligning the CT scan image to the point cloud in real time, the medical image is rendered on the same reference frame as the real-time subject information, allowing medical professionals to more easily align surgical instruments to the features shown in the medical image. Furthermore, the reference frame data can be used in combination with positional information from surgical instruments. Tracking data is converted to the same reference frame as the subject's point cloud data and the converted medical image, which can improve the accurate application of surgical treatment.
[0167] Returning to Figures 1 and 2, the processing circuit 212 of the image processing system 100 can align point clouds from one or more image capture devices 104 to a 3D medical image of a subject. The 3D medical image can be, for example, a 3D model generated from a CT scan image or an MRI image. To improve the processing speed of the alignment process, the processing circuit 212 can be used to identify feature vectors of the medical image in offline processing. Therefore, when aligning a medical image to a point cloud, the processing circuit 212 only needs to calculate the feature vectors of the point cloud captured in real time, thus improving the overall system performance.
[0168] The processing circuit 212 can access a set of data points in a first point cloud representing a global scene having a first reference frame. The global scene can be, for example, a scene represented by a set of data points in a point cloud. For example, when capturing an image using a 3D camera such as the capture device 104, features other than the object of analysis, such as the surroundings or room in which the subject is located, can also be captured. Therefore, the point cloud cannot represent only the surface of the area of interest, such as the subject, but can also include the surface of the environment which is not very relevant to the image alignment process. The global scene point cloud can be a combined point cloud generated from the image capture device 104, as described above herein.
[0169] The processing circuit 212 can identify a set of feature data points for a 3D medical image having a second reference frame different from a first reference frame. Feature data points can be, for example, one or more feature vectors extracted from a 3D medical image acquired in an offline process. Feature vectors can be generated, for example, by the processing circuit 212 performing one or more steps of the method 900 described herein in relation to Figure 9. Accessing feature vectors of a 3D medical image may include retrieving feature vectors from one or more data structures in the memory of the processing circuit 212.
[0170] In some implementations, before determining the feature vectors present in the 3D medical image, the processing circuit 212 may downsample the point cloud generated from the 3D medical image according to the embodiments described herein. For example, the processing circuit 212 may extract one or more data points from the 3D medical image to generate a point cloud representing the 3D medical image. The point cloud extracted from the 3D medical image may have a different reference frame than that of the point cloud generated by the image capture device 104. In some implementations, the point cloud captured from the 3D medical image is not downsampled; instead, the feature vectors are determined based on the entire point of the 3D medical image.
[0171] The processing circuit 212 can determine a transformation data structure for a 3D medical image using a first reference frame from feature vectors, a first set of data points, and a set of feature data points. In implementations where at least one of the point clouds representing the global scene or the point clouds representing the 3D medical image is downsampled, the processing circuit 212 can generate a transformation data structure using the reduced or downsampled point cloud(s). The transformation data structure may include one or more transformation matrices. The transformation matrices may be, for example, 4x4 rigid body transformation matrices. To generate the transformation matrices for the transformation data structure, the processing circuit 212 can identify one or more feature vectors of the global scene point cloud by, for example, performing one or more steps of method 900 described herein in relation to Figure 9. The result of this processing may include a set of feature vectors for each point cloud, where the global scene point cloud can be used as a reference frame (for example, the points of that point cloud are not transformed). The processing circuit 212 is capable of generating transformation matrices so that when each matrix is applied to the point cloud of a medical image (for example, used for transformation), the features of the medical image are aligned with similar features in the global scene point cloud.
[0172] To generate the transformation matrix (for example, as part of the transformation data structure, or as a transformation data structure), the processing circuit 212 can access or otherwise obtain the features corresponding to each point cloud from the memory of the processing circuit 212. To find the points in the reference frame point cloud that correspond to those in the point cloud to be transformed, the processing circuit uses L between the feature vectors in each point cloud. 2 The distance can be calculated. After these correspondences are enumerated, the processing circuit 212 can apply the Random Sample Consensus (RANSAC) algorithm to identify and reject incorrect correspondences.
[0173] The RANSAC algorithm can be used to determine which correspondences in each point cloud feature are relevant to the alignment process and which are false correspondences (e.g., features in one point cloud that are incorrectly exceptional for corresponding to features in the point cloud being transformed or aligned). The RANSAC algorithm is iterative and can reject false correspondences between two point clouds until a satisfactory model is fitted. The resulting satisfactory model can identify each data point in the reference point cloud that has data points corresponding to the point cloud being transformed, and vice versa.
[0174] When implementing the RANSAC algorithm, the processing circuit 212 performs L between feature vectors. 2 From the full set of correspondences identified using distance, a sample subset of feature correspondences containing the minimum correspondence can be randomly selected (e.g., pseudo-randomly). The processing circuit 212 can use the elements of this sample subset to compute a fitted model and its corresponding model parameters. The cardinality of the sample subset can be the minimum sufficient to determine the model parameters. The processing circuit 212 can check which elements of the full set of correspondences match the model instantiated by the estimated model parameters. A correspondence can be considered an outlier if it does not fit the fitted model instantiated by the set of estimated model parameters within a certain error threshold (e.g., 1%, 5%, 10%) that defines the maximum deviation due to noise. The set of inliers obtained for the fitted model can be called the consensus set of correspondences. The processing circuit 212 can repeat the steps of the RANSAC algorithm until the consensus set obtained in a given iteration has a sufficient number of inliers (e.g., above a predetermined threshold). The consensus set can then be used with an iterative closest approach (ICP) algorithm to determine the transformed data structure.
[0175] The processing circuit 212 can implement the ICP algorithm using a consensus set of corresponding features generated by the RANSAC algorithm. Each corresponding feature in the consensus set can contain one or more data points in each point cloud. When implementing the ICP algorithm, the processing circuit 212 can match the nearest point in the reference point cloud (or a selected set) to a point closet point in the point cloud to be transformed. The processing circuit 212 can then estimate the rotation and translation combination using a root mean square point distance metric minimization technique that optimally aligns each point in the point cloud to be transformed with its matching point in the reference point cloud. The processing circuit 212 can transform points in the point cloud to determine the amount of error between features in the point cloud, and iterate using this process to determine the optimal transformation values for the position and rotation of the point cloud to be transformed. These output values can be assembled into a transformation matrix, such as a 4x4 rigid body transformation matrix that includes the position or rotation changes of the 3D medical image. This output transformation matrix can be used as the transformation data structure.
[0176] The transformation matrix for a 3D medical image can be adapted to changes in the position or rotation of the 3D medical image. To align the 3D medical image with a point cloud representing the global scene, the processing circuit 212 can apply positional or rotational changes to the transformation matrix to points within the 3D medical image. By applying the transformation matrix, the 3D medical image is transformed into the same reference frame as the point cloud captured by the capture device 104. Therefore, by applying the transformation matrix, the 3D medical image is aligned with the point cloud of the global scene. The calculation of the transformation matrix and the alignment of the 3D medical image can be performed in real time. In some implementations, the data points of the global scene and the data points of the transformed 3D medical image can be placed within a single reference frame so that the 3D medical image is positioned relative to the reference frame of the global scene together with the data points of the global scene point cloud.
[0177] The processing circuit 212 can provide display information to the user interface 120 to display the rendering of the first point cloud and the 3D medical image in response to aligning the 3D medical image with the first point cloud. The global scene reference frame can be used with the converted 3D medical image to generate display data using one or more 3D rendering processes. The display data can be displayed, for example, on the user interface 120 of the image processing system 100.
[0178] The image processing system 100 can include information in addition to the converted 3D image in the reference frame of the global scene point cloud. For example, the processing circuit 212 can receive tracking data from surgical instruments and provide a display of surgical instructions in the reference frame of the global scene point cloud. For example, image data of surgical instructions (e.g., position, motion, tracking data, etc.) can be received using one of the image capture devices 104 that generate the reference frame of the global scene point cloud. Since the instruments are in the same reference frame as the global scene point cloud, the position and tracking data of the surgical instruments can be displayed together with the converted 3D medical image, including the global scene point cloud.
[0179] In some implementations, the processing circuit 212 can convert tracking data from surgical instruments to a first reference frame to generate converted tracking data. The converted tracking data may include changes in position, rotation, or other information received from the surgical instruments. For example, if there is an offset detected between the position of the surgical instrument and the reference frame of the global scene, the processing circuit 212 can rearrange or deform the tracking data to compensate for the offset. The offset can be manually corrected by user input. For example, if a user observing the user interface 120 notices an offset, they can manually input a conversion value to convert and correct the tracking data of the surgical instrument. In some implementations, this process can be performed automatically by the processing circuit 212. The processing circuit 212 can then use the converted tracking data to create display information that renders the converted surgical instrument tracking data in the global scene reference frame along with the global scene point cloud and the converted 3D medical image.
[0180] Using information from the global scene, the processing circuit 212 can determine a location of interest within a first reference frame related to the first point cloud and the 3D medical image. For example, the location of interest may include areas where the 3D medical image cannot be properly aligned with the global scene (e.g., outside the tolerance range). Under certain circumstances, the 3D medical image may become outdated and not properly aligned with the global scene. From the output values from the ICP process detailed above in this specification, the processing circuit can identify some locations where feature correspondence pairs were not aligned within the tolerance range. If a location of interest is detected, the processing circuit 212 can generate a highlighted area (e.g., highlighted in some way, flashing red, etc.) in the display data rendered to the user interface 120. The highlighted area may correspond to a location of interest on the 3D medical image or the global scene. In some implementations, determining the location of interest can be obtained from patent data such as lesions, fractures, or other medical problems that can be treated using surgical information. This location can be entered by a medical professional or automatically detected using other processing. A medical professional may, for example, enter or identify the location using one or more inputs on a user interface.
[0181] If the location of interest is a location related to a surgical procedure or other medical procedure that can be partially automated by a robotic device, the processing circuit 212 can generate motion instructions for a surgical instrument or other robotic device based on the global scene point cloud, 3D medical images, and the location of interest. Using the global scene point cloud data and the tracked location of the surgical instrument, the processing circuit can identify a path or set of locations that do not interfere with the global scene point cloud (e.g., such as causing the surgical instrument to collide with the subject in an undesirable way). Since the global scene point cloud can be calculated in real time and the surgical instrument can be tracked in real time, the processing circuit 212 can calculate and provide up-to-date motion instructions to move the surgical instrument to the location of interest within a given period of time. The generated motion instructions include instructions to move the surgical instrument to a specific location, or instructions to move the surgical instrument along a path calculated by the processing circuit 212 that allows the surgical instrument to reach the location of interest without undesiring interference with the patient. After generating a movement command, the processing circuit 212 can transmit the movement command to the surgical instrument using a communication circuit 216 that can be communicatively coupled to the surgical instrument. The command can be transmitted, for example, in one or more messages or data packets.
[0182] The processing circuit 212 can be configured to determine the distance of the patient represented in the 3D medical image from the capture device which is at least partially responsible for generating the first point cloud. For example, the processing circuit 212 can use a reference marker or object in the global scene to determine the actual distance between the capture device 104 which captures the global scene point cloud and the subject being imaged. If there is a reference object or field in the global scene point cloud with a known distance or length, the processing circuit 212 can use the known distance or length to determine or calculate different dimensions or parameters of the global scene point cloud, such as the distance from the image capture device 104 to other features in the global scene. Using the features of the subject in the global point cloud corresponding to features in the 3D medical image, the processing circuit 212 can determine the average position of the subject. Using this average position and the reference length or distance, the processing circuit 212 can determine the distance of the subject from the image capture device 104.
[0183] Figure 10 shows a method 1000 for aligning a medical image point cloud to a global scene point cloud. Method 1000 can be carried out using various devices and systems described herein, such as the image processing system 100.
[0184] In step 1005, multiple first feature vectors are accessed. The first feature vectors can correspond to a first point cloud representing first image data of the subject. For example, the first feature vectors can be generated from first point cloud data of the subject that can be resampled before feature detection. The first image data can be medical images (e.g., CT, MRI).
[0185] In step 1010, multiple second feature vectors are accessed. The second feature vectors can correspond to a second point cloud representing a second image data of the subject. For example, the second feature vectors can be generated from second point cloud data of the subject that can be resampled before feature detection.
[0186] Multiple second feature vectors can be mapped to a reference frame. For example, the first image data may be a global scene point cloud (which can be generated and updated over time) corresponding to the reference frame.
[0187] In 1015, the transformations of multiple first feature vectors are determined. The transformations can be determined to align the multiple first feature vectors with a reference frame. In some embodiments, a correspondence is generated between one or more first feature vectors and one or more second feature vectors (for example, based on the L2 distance between the feature vectors). The transformations can be determined by applying one or more alignment algorithms to the feature vectors or the correspondences between feature vectors, such as Random Sample Consensus (RANSAC) and Iterative Closest Approach (ICP). In some embodiments, the first pass is performed using RANSAC and the second pass is performed using ICP, which can improve the accuracy of the transformations identified. The transformations can be determined using the alignment algorithm(s) as a transformation matrix that can be applied to the first point cloud data points.
[0188] In 1020, multiple first feature vectors (or first point cloud data points) are aligned with a second point cloud (e.g., with a reference frame of the global scene). Alignment can be performed by applying a transformation (e.g., a transformation matrix) to the first feature vectors or the first point cloud data points associated with the first feature vectors.
[0189] IX. System and method for real-time surgical planning visualization using pre-captured medical images and global scene images.
[0190] The image processing pipeline described herein can enable improved surgical planning visualization, such as visualizing 3D images along with medical images and models, and along with planned trajectories for instrument navigation.
[0191] Figure 11 illustrates Method 1100 for real-time visualization of surgical planning using pre-captured medical images and global scene images. Method 1100 can be implemented using various devices and systems described herein, such as the image processing system 100.
[0192] In 1105, medical images and 3D image data relating to the subject are accessed. Medical images may include various types of medical images such as CT images or MRI images. 3D image data can be received from one or more 3D cameras, such as depth cameras.
[0193] In 1110, medical images are aligned to three-dimensional image data. Alignment can be performed using various processes described herein, including resampling the medical image data and three-dimensional image data, determining features from the resampled data, identifying transformations to align the medical image data and three-dimensional image data (e.g., to each other or to a global reference frame), and applying transformations to one or both of the medical image data or three-dimensional image data.
[0194] In 1115, a visual indicator is received via the user interface. The visual indicator can show a trajectory or path in an environment presented using medical image data and 3D image data. For example, the visual indicator can show the path through which an instrument is introduced to a subject.
[0195] In 1120, the visual indicator is mapped to a medical image. For example, a reference frame receiving the visual indicator can be identified, and the conversion of the visual indicator to the medical image can be determined in order to map the visual indicator to the medical image.
[0196] In 1125, medical images, 3D image data, and visual indicators are presented. These medical images, 3D image data, and visual indicators can be presented using a display device. For example, visual indicators can be presented as overlays on 3D images of the subject and on CT or MRI images of the subject. Presenting medical images may include presenting display data corresponding to visual indicators that include at least one of the highlighting of target features of the subject or the trajectory of an instrument.
[0197] X. System and method for dynamically tracking the movement of instruments in a 3D imaging environment
[0198] As mentioned above, IR sensors can be used not only to track instruments in the environment surrounding a subject, but also while the instrument is being operated on the subject. This solution can use tracking data to display a representation of the tracked instrument along with 3D image data and medical image data (e.g., CT or MRI), enabling users to effectively visualize how the instrument interacts with the subject.
[0199] In 1205, it is possible to access 3D image data relating to the environment of the subject. 3D image data can be received from one or more 3D cameras, such as depth cameras.
[0200] In 1210, medical images related to the subject can be accessed. These medical images may include various types of medical images, such as CT or MRI images.
[0201] In 1215, medical images can be aligned to three-dimensional image data. Alignment can be performed using various processes described in this document, including resampling the medical image data and three-dimensional image data, determining features from the resampled data, identifying transformations to align the medical image data and three-dimensional image data (e.g., to each other or to a global reference frame), and applying transformations to one or both of the medical image data or three-dimensional image data.
[0202] In 1220, an instrument can be identified from 3D image data. An instrument can be identified by performing one of various object recognition processes using 3D image data, such as obtaining template features of an object and comparing the template features with features extracted from the 3D image data. An instrument can also be identified based on an identifier (e.g., a visual indicator) attached to the instrument, which can reduce the computational requirements for identifying the instrument by reducing the search space of 3D image data in which the instrument is identified.
[0203] In 1225, access the model of the device. The model can show the shape, contour, edges, or other features of the device. The model may include template features used to identify the device.
[0204] In 1230, positional data related to the device is tracked by matching a portion of the 3D image data representing the device to a model of the device. For example, by matching features extracted from the image data to a model of the device, it is possible to identify the location of the device in the 3D image data and track the device by monitoring it across the image (e.g., a stream of images from a 3D camera).
[0205] XI. Computing Environment for Real-Time Multi-Modality Image Alignment
[0206] Figures 13A and 13B show block diagrams of computing devices 1300. As shown in Figures 13A and 13B, each computing device 1300 includes a central processing unit (CPU) 1321 and a main memory unit 1322. As shown in Figure 13A, computing devices 1300 may include a storage device 1328, an installation device 1316, a network interface 1318, an I / O controller 1323, display devices 1324a to 1324n, a keyboard 1326, and a pointing device 1327, such as a mouse. The storage device 1328 may include, but is not limited to, an operating system, software, and software for system 200. As shown in Figure 13B, each computing device 1300 may also include additional optional elements, such as a memory port 1303, a bridge 1370, one or more input / output devices 1330a to 1330n (generally referred to by reference numeral 1330), and a cache memory 1340 that communicates with the CPU 1321.
[0207] The CPU 1321 is any logic circuit that processes instructions fetched from the main memory unit 1322. In many embodiments, the CPU 1321 is provided by a microprocessor unit, such as those manufactured by Intel Corporation in Mountain View, California; Motorola Corporation in Schaumburg, Illinois; ARM processors (e.g., those manufactured by ST, TI, ATMEL, etc. from ARM Holdings); TEGRA system-on-chip (SoC) manufactured by Nvidia in Santa Clara, California; POWER7 processors manufactured by International Business Machines in White Plains, New York or Advanced Microdevices in Sunnyvale, California; or field-programmable arrays ("FPGAs") from Altera, Intel Corporation in San José, California, Xlinix in San José, California, or Microsemi in Aliso Viejo, California. The computing device 1300 may be based on any of these processors or other processors capable of operating as described herein. The CPU 1321 may utilize instruction-level parallelism, thread-level parallelism, different levels of cache, and multi-core processors. A multicore processor can contain two or more processing units on a single computing component. Examples of multicore processors include the AMD PHENOM II X2, Intel Core i5, and Intel Core i7.
[0208] The main memory unit 1322 may include one or more memory chips that can store data and allow direct access to any storage location by the microprocessor 1321. The main memory unit 1322 may be more volatile and faster than the memory of the storage 1328. The main memory unit 1322 may be any variant of dynamic random access memory (DRAM) or static random access memory (SRAM), burst SRAM or sync-burst SRAM (BSRAM), fast page mode DRAM (FPM DRAM), extended DRAM (EDRAM), extended data output RAM (EDO RAM), extended data output DRAM (EDO DRAM), burst extended data output DRAM (BEDO DRAM), single data rate synchronous DRAM (SDR SDRAM), double data rate SDRAM (DDR SDRAM), direct Rambus DRAM (DRDRAM), or extreme data rate DRAM (XDR DRAM). In some embodiments, the main memory 1322 or storage 1328 can be non-volatile, such as non-volatile read-access memory (NVRAM), flash memory non-volatile static RAM (nvSRAM), ferroelectric RAM (FeRAM), magnetoresistive RAM (MRAM), phase-change memory (PRAM), conductive bridge RAM (CBRAM), silicon oxide-nitride-silicon (SONOS), resistive RAM (RRAM), racetrack, nanoRAM (NRAM), or millipade memory. The main memory 1322 can be based on any of the memory chips described above, or on any other available memory chip that can operate as described herein. In the embodiment shown in Figure 13A, the processor 1321 communicates with the main memory 1322 via the system bus 1350 (described in more detail below). Figure 13B depicts one embodiment of a computing device 1300 in which the processor communicates directly with the main memory 1322 via a memory port 1303. For example, in Figure 13B, the main memory 1322 can be DRDRAM.
[0209] Figure 13B shows an embodiment in which the main processor 1321 communicates directly with the cache memory 1340 via a secondary bus sometimes called the backside bus. In other embodiments, the main processor 1321 communicates with the cache memory 1340 using the system bus 1350. The cache memory 1340 has a faster response time than the main memory 1322 and is basically provided by SRAM, BSRAM, or EDRAM. In the embodiment shown in Figure 13B, the processor 1321 communicates with various I / O devices 1330 via the local system bus 1350. Various buses can be used to connect the CPU 1321 to any of the I / O devices 1330, including the PCI bus, PCI-X bus, PCI-Express bus, or NuBus. In embodiments where the I / O device is a video display 1324, the processor 1321 can use an Advanced Graphics Port (AGP) to communicate with the display 1324 or the I / O controller 1323 for the display 1324. Figure 13B shows one embodiment of a computer 1300 in which the main processor 1321 communicates directly with I / O device 1330b or other processors 1321' via HYPERTRANSPORT, RAPIDIO, or INFINIBAND communication technology. Figure 13B also shows a mixed embodiment in which processor 1321 communicates with I / O device 1330a using a local interconnect bus while also communicating directly with I / O device 1330b. In some embodiments, processor 1321 can communicate with other processing devices such as other processors 1321', GPUs, and FPGAs via various buses connected to the processing unit 1321. For example, processor 1321 can communicate with a GPU via one or more communication buses such as a PCI bus, PCI-X bus, PCI-Express bus, or NuBus.
[0210] The computing device 1300 can contain a wide variety of I / O devices 1330a to 1330n. Input devices may include keyboards, mice, trackpads, trackballs, touchpads, touch mice, multi-touch touchpads and touch mice, microphones (analog or MEMS), multi-array microphones, drawing tablets, cameras, single-lens reflex (SLR) cameras, digital single-lens reflex (DSLR) cameras, CMOS sensors, CCDs, accelerometers, inertial measurement units, infrared optical sensors, pressure sensors, geomagnetic sensors, angular velocity sensors, depth sensors, proximity sensors, ambient light sensors, gyroscopes, or other sensors. Output devices may include video displays, graphical displays, speakers, headphones, inkjet printers, laser printers, and 3D printers.
[0211] Devices 1330a to 1330n may include combinations of multiple input or output devices, such as Microsoft Kinect, Nintendo Wiimote for Wii, Nintendo Wii U GamePad, or Apple iPhone. Some devices 1330a to 1330n enable gesture recognition input by combining some of the inputs and outputs. Some devices 1330a to 1330n provide facial recognition, which can be used as input for different purposes, including authentication and other commands. Some devices 1330a to 1330n provide voice recognition and input, such as Microsoft Kinect, Siri for iPhone by Apple, Google Now, or Google Voice Search.
[0212] Additional devices 1330a–1330n have both input and output capabilities, including, for example, haptic feedback devices, touchscreen displays, or multitouch displays. Touchscreens, multitouch displays, touchpads, touch mice, or other touch-sensing devices can use different technologies for touch sensing, such as capacitive, surface capacitive, projected capacitive touch (PCT), in-cell capacitive, resistive, infrared, waveguide, dispersed signal touch (DST), in-cell optics, surface acoustic wave (SAW), bent wave touch (BWT), or force-based sensing technologies. Some multitouch devices have two or more contact points with the surface, enabling advanced functions including gestures such as pinch, spread, rotate, and scroll. For example, some touchscreen devices, including Microsoft PIXELSENSE or the Multitouch Collaboration Wall, can have larger surfaces, such as on a tabletop or a wall, and can also interact with other electronic devices. Several I / O devices 1330a-1330n, display devices 1324a-1324n, or groups of devices can constitute an augmented reality device. The I / O devices can be controlled by an I / O controller 1321, as shown in Figure 13A. The I / O controller 1321 can control one or more I / O devices, such as a keyboard 126 and a pointing device 1327, e.g., a mouse or optical pen. Furthermore, the I / O devices can also provide storage and / or installation media 116 for the computing device 1300. In yet another embodiment, the computing device 1300 can provide a USB connection (not shown) to receive a handheld USB storage device. In a further embodiment, the I / O device 1330 can act as a bridge between the system bus 1350 and an external communication bus, such as a USB bus, SCSI bus, FireWire bus, Ethernet bus, Gigabit Ethernet bus, Fibre Channel bus, or Thunderbolt bus.
[0213] In some embodiments, the display devices 1324a to 1324n can be connected to the I / O controller 1321. The display devices may include, for example, liquid crystal displays (LCDs), thin-film transistor LCDs (TFT-LCDs), blue phase LCDs, electronic paper (E-ink) displays, flexil displays, light-emitting diode displays (LEDs), digital light-processing (DLP) displays, liquid crystal displays on silicon (LCOS) displays, organic light-emitting diode (OLED) displays, active-matrix organic light-emitting diode (AMOLED) displays, liquid crystal laser displays, time-multiplexed optical shutter (TMOS) displays, or 3D displays. Examples of 3D displays may use, for example, stereoscopy, polarizing filters, active shutters, or auto-stereoscopy. The display devices 1324a to 1324n may also be head-mounted displays (HMDs). In some embodiments, the display devices 1324a to 1324n or the corresponding I / O controller 1323 may be controlled through the OPENGL or DIRECTX API or other graphics libraries, or have hardware support for them.
[0214] In some embodiments, the computing device 1300 may include or be connected to a plurality of display devices 1324a-1324n, each of which may be the same or different type and / or form. Thus, either the I / O devices 1330a-1330n and / or the I / O controller 1323 may include any type and / or form of appropriate hardware, software, or combination of hardware and software to support, enable, or provide the connection and use of the plurality of display devices 1324a-1324n by the computing device 1300. For example, the computing device 1300 may include any type and / or form of video adapters, video cards, drivers, and / or libraries to interface, communicate, connect, or otherwise use the display devices 1324a-1324n. In one embodiment, the video adapter may include multiple connectors to interface with the plurality of display devices 1324a-1324n. In other embodiments, the computing device 1300 may include multiple video adapters, each video adapter connected to one or more display devices 1324a-1324n. In some embodiments, any part of the operating system of the computing device 1300 may be configured to use the multiple displays 1324a-1324n. In other embodiments, one or more of the display devices 1324a-1324n may be provided by one or more other computing devices 1300a or 1300b connected to the computing device 1300 via the network 1340. In some embodiments, software may be designed and constructed to use a display device of another computer as a second display device 1324a for the computing device 1300.For example, in one embodiment, an Apple iPad can be connected to a computing device 1300 and used as an additional display screen, allowing the device 1300's display to be used as an extended desktop. Those skilled in the art will recognize and understand various ways and embodiments by which the computing device 1300 can be configured to have multiple display devices 1324a to 1324n.
[0215] Referring again to Figure 13A, the computing device 1300 can be configured with a storage device 1328 (e.g., one or more hard disk drives or a redundant array of independent disks) for storing the operating system or other related software, and for storing application software programs such as any programs related to the software for system 200. Examples of storage devices 1328 include, for example, hard disk drives (HDDs), optical drives including CD drives, DVD drives, or Blu-ray drives, solid-state drives (SSDs), USB flash drives, or any other devices suitable for storing data. Some storage devices may include multiple volatile and non-volatile memories, for example, a solid-state hybrid drive that combines a hard disk and a solid-state cache. Some storage devices 1328 may be non-volatile, modifiable, or read-only. Some storage devices 1328 may be internal and connected to the computing device 1300 via bus 1350. Some storage devices 1328 may be external and connected to the computing device 1300 via I / O device 1330 that provides an external bus. Some storage devices 1328 can be connected to computing devices 1300 via a network interface 1318 over a network, including, for example, a remote disk for Apple's MacBook Air. Some client devices 1300 do not require non-volatile storage devices 1328 and can be thin clients or zero clients 202. Some storage devices 1328 can also be used as installation devices 1316, making them suitable for installing software and programs.Furthermore, the operating system and software can be run from a bootable medium, such as a bootable CD, for example, KNOPPIX, a bootable CD for GNU / Linux available as a GNU / Linux distribution from knoppix.net.
[0216] Computing device 1300 can also install software or applications from application distribution platforms. Examples of application distribution platforms include the Apple App Store for iOS, the Apple Mac App Store, Google Play for Android OS, the Google Chrome Webstore for Chrome OS, and the Amazon Appstore for Android OS and Kindle Fire provided by Amazon.com.
[0217] Furthermore, the computing device 1300 may include a network interface 1318 that interfaces to the network 1340 via a variety of connections, including but not limited to, standard telephone line LAN or WAN links (e.g., 802.11, T1, T3, Gigabit Ethernet, Infiniband), broadband connections (e.g., optical fiber such as ISDN, Frame Relay, ATM, Gigabit Ethernet, Ethernet-over-Sonet, ADSL, VDSL, BPON, GPON, FiOS), wireless connections, or any combination of the above. The connection can be established using various communication protocols (e.g., TCP / IP, Ethernet, ARCNET, SONET, SDH, Fiber Distributed Data Interface (FDDI), IEEE 802.11a / b / g / n / ac CDMA, GSM, WiMAX, and direct asynchronous connections). In one embodiment, computing device 1300 communicates with other computing device 1300' via any type and / or form of gateway or tunneling protocol, such as SSL (Secure Socket Layer) or TLS (Transport Layer Security), or Citrix Gateway Protocol from Citrix Systems, Fort Lauderdale, Florida. The network interface 1318 may include an internal network adapter, a network interface card, a PCMCIA network card, an EXPRESSCARD network card, a CardBus network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for interface connecting computing device 1300 to any type of network capable of communication and performing the operations described herein.
[0218] The type of computing device 1300 depicted in Figure 13A can operate under the control of an operating system that controls task scheduling and access to system resources. The computing device 1300 can run any operating system, including any version of the MICROSOFT WINDOWS operating system, different releases of the Unix and Linux operating systems, any version of MAC OS for Macintosh computers, any embedded operating system, any real-time operating system, any open-source operating system, any proprietary operating system, any operating system for mobile computing devices, or any other operating system that can run on the computing device and perform the operations described herein. Typical operating systems include, but are not limited to, the following: These include Windows 7000, Windows Server 2012, Windows CE, Windows Phone, Windows XP, Windows Vista, and Windows 7, Windows RT, and Windows 8, all manufactured by Microsoft in Redmond, Washington; Mac OS and iOS, manufactured by Apple in Cupertino, California; and Linux, Unix, and other Unix-like operating systems, such as the Linux Mint distribution ("distro") or Ubuntu, distributed by Canonical in London, UK; and Android, designed by Google in Mountain View, California. For example, some operating systems, including Google's Chrome OS, can be used on zero clients or thin clients, such as Chromebooks.
[0219] The computer system 1300 can be any workstation, telephone, desktop computer, laptop or notebook computer, netbook, ultrabook, tablet, server, handheld computer, mobile phone, smartphone or other mobile communication device, media playback device, game system, mobile computing device, or any other type and / or form of computer, telecommunications or media device capable of communication. The computer system 1300 has sufficient processor power and memory capacity to perform the operations described herein. In some embodiments, the computing device 1300 may have different processors, operating systems, and input devices that match the device. For example, the Samsung GALAXY smartphone operates under the control of the Android operating system developed by Google. The GALAXY smartphone receives input via a touch interface.
[0220] In some embodiments, the computing device 1300 is a game system. For example, the computer system 1300 may include a PLAYSTATION 3, PSP (PERSONAL PLAYSTATION PORTABLE), or PLAYSTATION VITA device manufactured by Sony Corporation in Tokyo, Japan; a Nintendo DS, Nintendo 3DS, Nintendo Wii, or Nintendo Wii U device manufactured by Nintendo Co., Ltd. in Kyoto, Japan; an XBOX 360 manufactured by Microsoft Corporation in Redmond, Washington; or an OCULUS RIFT or OCULUS VR manufactured by OCULUS VR Inc. in Menlo Park, California.
[0221] In some embodiments, the computing device 1300 is a digital audio player such as the Apple iPod, iPod Touch, and iPod Nano line of devices manufactured by Apple Inc. in Cupertino, California. Some digital audio players may have other functionalities, including, for example, a game system or any functions made available by applications from a digital application distribution platform. For example, the iPod Touch may be able to access the Apple App Store. In some embodiments, the computing device 1300 is a portable media player or digital audio player that supports file formats including, but not limited to, MP3, WAV, M4A / AAC, WMA protected AAC, AIFF, Audible audiobook, Apple Lossless audio file format, and mov, m4v, mp4 MPEG-4 (H.264 / MPEG-4 AVC) video file format.
[0222] In some embodiments, the computing device 1300 is a tablet, such as Apple's iPad line of devices, Samsung's GALAXY TAB family of devices, or Amazon's Kindle Fire from Seattle, Washington. In other embodiments, the computing device 1300 is an e-book reader, such as Amazon's Kindle family of devices, or Barnes & Noble's NOOK family of devices from New York City, New York.
[0223] In some embodiments, the communication device 1300 includes a smartphone combined with a combination of devices, such as a digital audio player or a portable media player. For example, one of these embodiments is a smartphone, such as Apple's iPhone family of smartphones, Samsung's Samsung Galaxy family of smartphones, or Motorola's Motorola DROID family of smartphones. In yet another embodiment, the communication device 1300 is a laptop or desktop computer equipped with a web browser and a microphone and speaker system, such as a telephony headset. In these embodiments, the communication device 1300 is web-enabled and can receive and initiate phone calls. In some embodiments, the laptop or desktop computer also includes a webcam or other video capture device that enables video chat and video calls.
[0224] In some embodiments, the status of one or more machines 1300 in the network is monitored, generally as part of network management. In one of these embodiments, the machine status may include identification of load information (e.g., the number of processes on the machine, CPU and memory usage), port information (e.g., the number of available communication ports and port addresses), or session status (e.g., the duration and type of the process, and whether the process is active or idle). In another of these embodiments, this information may be identified by multiple metrics, which can be applied at least partially to decisions in load balancing, network traffic management, and network fault recovery, as in any aspect of the operation of the Solutions described herein. The aspects of the operating environment and components described above will become apparent in the context of the Systems and Methods disclosed herein.
[0225] Implementations of the subject matter and operations described herein can be implemented in digital electronic circuits or in computer software embodied on tangible media, firmware, or hardware, including, or in combination with, the structures disclosed herein and their structural equivalents. Implementations of the subject matter described herein can be implemented as one or more computer programs encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device, for example, as one or more components of computer program instructions. Program instructions can be encoded in artificially generated propagating signals, for example, mechanically generated electrical signals, optical signals, or electromagnetic signals generated to encode information for transmission to a suitable receiving device for execution by a data processing device. Computer storage mediums can be, or include, computer-readable storage devices, computer-readable storage boards, random or serial access memory arrays or devices, or one or more combinations thereof. Furthermore, although computer storage mediums are not propagating signals, computer storage mediums can contain the source or destination of computer program instructions encoded in artificially generated propagating signals. Furthermore, computer storage media may consist of one or more separate physical components or media (for example, multiple CDs, disks, or other storage devices), or be contained within them.
[0226] The operations described herein can be implemented as operations performed by a data processing device on data stored in one or more computer-readable storage devices or received from other sources.
[0227] The terms "data processing device," "data processing system," "client device," "computing platform," "computing device," or "device" encompass, for example, all types of devices, machines, and equipment that process data, including programmable processors, computers, systems on chips, or combinations of the aforementioned. A device may include special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, a device may also include code that constitutes the execution environment of the computer program, such as processor firmware, protocol stacks, database management systems, operating systems, cross-platform execution environments, virtual machines, or one or more combinations thereof. This device and execution environment can realize various different computing model foundations, such as web services, distributed computing, and grid computing infrastructure.
[0228] Computer programs (also called programs, software, software applications, scripts, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, such as a standalone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. Computer programs can, but do not have to, correspond to files in a file system. A program can be part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program, or multiple collaborative files (e.g., a file that holds one or more modules, subprograms, or parts of code). Computer programs can be deployed to run on a single computer or on multiple computers distributed across one or more sites and interconnected by a communication network.
[0229] The processes and logic flows described herein are implemented by one or more programmable processors running one or more computer programs, which can perform actions by acting on input data to produce outputs. The processes and logic flows can also be implemented by special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the apparatus can also be implemented in this manner.
[0230] Processors suitable for executing computer programs include, as an example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory or both. The elements of a computer include a processor for performing operations according to instructions, and one or more memory devices for storing instructions and data. Generally, a computer will include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, which are operablely coupled to receive data from or transfer data to both, or both. However, a computer does not need to have such devices. Furthermore, a computer can be incorporated into another device, such as a mobile phone, a personal digital assistant (PDA), a portable audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive). Devices suitable for storing computer program instructions and data include, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and all forms of non-volatile memory, media, and memory devices, including CD-ROMs and DVD-ROMs. Processors and memory may be supplemented or incorporated by logic circuits for special purposes.
[0231] To provide user interaction, implementations of the subject matter described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), plasma, or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, to which the user can provide input to the computer. It is also possible to provide user interaction using other types of devices, for example, the feedback provided to the user may include any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, such as acoustic, voice, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from a web browser.
[0232] Implementations of the subject matter described herein can be implemented as a backend component, e.g., a data server; a middleware component, e.g., a computing system including an application server; or a frontend component, e.g., a client computer having a graphical user interface or a web browser that allows a user to interact with the implementation of the subject matter described herein; or in any combination of one or more such backend, middleware, or frontend configurations. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad-hoc peer-to-peer networks).
[0233] This specification includes many specific implementation details, which should not be construed as limitations on the scope of any invention or claims, but rather as descriptions of features specific to particular implementations of the systems and methods described herein. Certain features described herein in the context of individual implementations may also be implemented in combination within a single implementation. Conversely, various features described within the context of a single implementation may be implemented separately or in any suitable subcombination within multiple implementations. Furthermore, features described above as acting in a particular combination, and may even be initially claimed to act in this way, but one or more features from the claimed combination may, in some cases, be removed from the combination, and the claimed combination may be directed towards a subcombination or a variation of a subcombination.
[0234] Similarly, while operations are depicted in a specific order in the drawings, this should not be interpreted as requiring that such operations be performed in the specific order shown, or sequentially, or that all illustrated operations be performed, in order to achieve the desired result. In some cases, the operations described in the claims may be performed in a different order, and the desired result may still be achieved. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown, or a sequential order, to achieve the desired result.
[0235] Under certain circumstances, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.
[0236] Now, while several exemplary implementations and embodiments have been described, it is clear that these are presented illustratively and are not limiting. In particular, many of the examples presented herein involve specific combinations of method actions or system elements, but these actions and their elements can be combined in other ways to achieve the same objective. Actions, elements, and features discussed in relation to only one implementation are not intended to preclude similar roles in other implementations or embodiments.
[0237] The expressions and terms used herein are for illustrative purposes only and should not be considered limiting. The use herein of “including,” “comprising,” “having,” “containing,” “involving,” “characterized by,” “characterized in that,” and their variations is intended to include the items listed thereafter, their equivalents, and additional items, as well as alternative implementations consisting only of the items listed thereafter. In a single implementation, the systems and methods described herein consist of one, any combination of, or all of the elements, actions, or components described herein.
[0238] Any singular reference to an implementation or element or act of a system or method in this specification may include implementations that include multiple such elements, and any plural reference to an implementation or element or act in this specification may include implementations that include only a single element. Singular or plural references are not intended to limit the currently disclosed systems or methods, their components, acts, or elements to one or more configurations. A reference that any act or element is based on any information, act, or element may include implementations in which the act or element is at least partially based on any information, act, or element.
[0239] Any implementation disclosed herein can be combined with any other implementation, and references to “an implementation,” “some implementation,” “an alternate implementation,” “various implementation,” “one implementation,” etc., are not necessarily mutually exclusive and are intended to indicate that certain features, structures, or characteristics described in relation to an implementation may be present in at least one implementation. Such terms used herein do not necessarily refer to the same implementation. Any implementation can be combined with any other implementation, inclusively or exclusively, in any manner consistent with the embodiments and implementations disclosed herein.
[0240] A reference to "or" can be interpreted comprehensively, so that any term described using "or" can refer to a single term, a group of terms, or all of the terms being described.
[0241] Where reference numerals follow drawings, detailed descriptions, or technical features in claims, the reference numerals are included solely for the purpose of enhancing the understanding of the drawings, detailed descriptions, and claims. Therefore, the reference numerals, or their absence, do not limit the scope of any element of a claim.
[0242] The systems and methods described herein can be embodied in other specific forms without departing from their features. While the provided examples are useful for transforming three-dimensional point clouds into different reference frames, the systems and methods described herein can also be applied to other environments. The aforementioned embodiments are illustrative, not limiting, the systems and methods described herein. Accordingly, the scope of the systems and methods described herein can be expressed rather than from the foregoing description by the appended claims, and any changes that fall within the meaning and equivalence of the claims are encompassed therein.
Claims
1. It is a method, The steps of accessing, by one or more processors, a first set of data points of a first point cloud captured in a first capture session by a first capture device having a first modality and a first pose, and a second set of data points of a second point cloud captured in a second capture session by a second capture device having a second modality different from the first modality and a second pose different from the first pose, Here, the second capture session is temporally separated from the first capture session. The steps of applying a response function to the first set of data points and the second set of data points in order to assign a set of response values corresponding to the first set of data points and the second set of data points, respectively, by one or more processors, wherein each response value corresponds to a weight value, and a larger weight value indicates that the associated data points have a higher relevance to the feature of interest, A step of resampling a first set of data points and a second set of data points to generate a second set of downsampled data points using one or more processors, wherein the resampling includes removing data points having associated response values below a threshold. The steps include: selecting a reference frame based on a first set of downsampled data points using one or more processors; The steps include: determining a transformed data structure for a second set of downsampled data points using the reference frame and a first set of downsampled data points by one or more processors; The steps include: using one or more processors, converting the second set of downsampled data points into a converted set of data points using the converted data structure and the second set of downsampled data points; Equipped with, The step of applying the response function includes applying the graph filter to the graph data structure generated from each of the first set of data points and the second set of data points, The graph data structure comprises an adjacency matrix, the values of which correspond to edge weights between point cloud data points.
2. The method according to claim 1, wherein the step of accessing a first set of data points of the first point cloud is: The steps include: receiving three-dimensional (3D) image data from the first capture device using one or more processors, A step of generating a first point cloud having a first set of data points using the 3D image data by one or more processors, Methods that include...
3. A method according to claim 1, wherein the step of selecting the reference frame includes the step of selecting the first reference frame of the first point cloud as the first reference frame.
4. The method according to claim 1, wherein the step of selecting a reference frame is: The steps include: obtaining color data assigned to one or more of the first sets of data points of the first point cloud using one or more processors, and The step of determining the reference frame based on the color data using one or more processors, Methods that include...
5. A method according to claim 1, wherein the step of determining the transformation data structure includes the step of generating the transformation data structure to include a change in position or a change in rotation.
6. A method according to claim 5, wherein the step of transforming a second set of data points includes the step of applying the change in position or the change in rotation to at least one data point in the second set of data points by one or more processors to generate a transformed set of data points.
7. The method according to claim 1, further, The steps include downsampling at least one of the first set of data points or the second set of data points using one or more processors, The steps include determining the transformed data structure in response to downsampling at least one of the first set of data points or the second set of data points by one or more processors, A method that includes [a certain feature].
8. A method according to claim 1, wherein the step of transforming a second set of data points includes the step of matching at least one first point from the first set of data points with at least one second point from the second set of data points using one or more processors.
9. It is a system, One or more processors configured to perform the following steps using machine-readable instructions, namely, The steps of accessing a first set of data points of a first point cloud captured in a first capture session by a first capture device having a first modality and a first pose, and a second set of data points of a second point cloud captured in a second capture session by a second capture device having a second modality different from the first modality and a second pose different from the first pose, Here, the second capture session is temporally separated from the first capture session. The steps of applying a response function to the first set of data points and the second set of data points in order to assign a set of response values corresponding to the first set of data points and the second set of data points, respectively, by one or more processors, wherein each response value corresponds to a weight value, and a larger weight value indicates that the associated data points have a higher relevance to the feature of interest, A step of resampling a first set of data points and a second set of data points to generate a second set of downsampled data points using one or more processors, wherein the resampling includes removing data points having associated response values below a threshold. The steps include selecting a reference frame based on a first set of downsampled data points, A step of determining a transformed data structure for a second set of downsampled data points using the reference frame and a first set of downsampled data points, A step of converting the second set of data points into a converted set of data points using the converted data structure and the second set of downsampled data points, The system comprises one or more processors configured to perform the following: The step of applying the response function includes applying the graph filter to the graph data structure generated from each of the first set of data points and the second set of data points, The aforementioned graph data structure comprises an adjacency matrix, the values of which correspond to edge weights between point cloud data points.
10. The system according to claim 9, wherein one or more processors further perform the following steps by machine-readable instructions, namely: The steps include receiving three-dimensional (3D) image data from the first capture device, and A step of generating a first point cloud having a first set of data points using the 3D image data from the first capture device, A system configured to perform the following actions.
11. A system according to claim 9, wherein one or more processors are further configured to perform the step of selecting a first reference frame of the first point cloud as the first reference frame by machine-readable instructions.
12. The system according to claim 9, wherein one or more processors further process machine-readable instructions: A step of acquiring color data assigned to one or more of the first sets of data points of the first point cloud, and A step of determining the reference frame based on the color data, A system configured to perform the following actions.
13. The system according to claim 9, wherein one or more processors are further configured to perform the step of generating the transformation data structure, which includes a change in position or a change in rotation, by machine-readable instructions.
14. A system according to claim 13, wherein one or more processors are further configured to perform the step of applying the change in position or the change in rotation to at least one data point of a second set of data points by machine-readable instructions to generate a set of transformed data points.
15. The system according to claim 9, wherein one or more processors further perform the following steps by machine-readable instructions, namely: The steps include downsampling at least one of the first set of data points or the second set of data points, and A step of determining the transformed data structure in response to a step of downsampling at least one of the first set of data points or the second set of data points, A system configured to perform the following actions.
16. A system according to claim 9, wherein one or more processors are further configured to perform the step of matching at least one first point in a first set of data points with at least one second point in a second set of data points by machine-readable instructions.