Data processing method, data processing system, data processing program and training method
The AI-driven data processing method addresses the limitations of conventional point cloud registration by rapidly and accurately aligning endoscopic camera data with anatomical information from other imaging modalities, enhancing the efficiency and reliability of medical imaging in real-time surgical procedures.
Patent Information
- Application Number
- JP2024188406
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-10-25
- Publication Date
- 2025-05-26
AI Technical Summary
Conventional point cloud registration methods for aligning endoscopic camera images with anatomical information from other imaging modalities are slow, unreliable, and often require user interaction, limiting their effectiveness in real-time surgical procedures.
A data processing method utilizing an AI system with an artificial neural network to identify a first correlation between source data from an endoscopic camera and target data from another imaging modality, enabling rapid and accurate alignment of the data for improved integration and display.
The proposed method significantly enhances the speed and accuracy of data registration, enabling real-time integration and display of endoscopic and anatomical data, thus improving the efficiency and reliability of medical imaging applications.
Smart Images

Figure 2025080754000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a data processing method, a data processing system, a data processing program, and a training method for processing image data. In particular, the present invention relates to a data processing method and a data processing system for processing source data obtained from an image captured by a camera that captures a target and target data related to the target so as to obtain a correlation between the source data and the target data.
Background Art
[0002] What is usually required in endoscopic imaging is to enhance the value of an image captured by an endoscopic camera with anatomical information of a target (e.g., organs such as the liver, lung, colon, rectum, kidney, pancreas, etc.) being examined. The anatomical information may be provided in the form of another image of the target obtained using another imaging modality such as computed tomography (CT) or magnetic resonance imaging (MRI) before the endoscopic examination. Subsequently, an image obtained from another imaging modality may be overlaid on an image captured by the endoscopic camera so that high-value-added information that facilitates medical analysis is added to the camera image. In another conventional application example, an augmented reality (AR) device is used to display a video captured by the endoscopic camera, and anatomical information is overlaid on the endoscopic video.
[0003] To properly align relative to each other an image captured by a camera and additional information obtained through another imaging modality for display together, a positional correlation needs to be explored between points (source points) within the image captured by the camera and corresponding points (target points) within the image obtained from another imaging modality. Such a correlation is typically explored by a point cloud registration (PCR) process, in which the source point cloud containing the source points and the target point cloud containing the target points are processed such that a transformation matrix that enables aligning the source point cloud with the target point cloud is explored. As a result, PCR enables integration of information obtained by an endoscopic camera and information obtained from another imaging modality.
[0004] FIG. 1 shows a conventional PCR pipeline that uses as inputs a source point cloud obtained from a camera image and a target point cloud obtained from a CT scan, an MRI scan, or any imaging modality. The source point cloud and the target point cloud are input into a point cloud registration section, which aims to identify the position of the target point cloud within the source point cloud. To achieve this, in a conventionally known PCR algorithm, feature descriptors are applied to both the target point cloud and the source point cloud to assign a fingerprint to each individual point within these point clouds, thereby obtaining a unique description of each point such that, in the best case, each point is distinguishable from all other points. Thereafter, in the feature matching step of the algorithm, using the feature descriptor of each point in the target point cloud, the corresponding point with the most similar feature descriptor in the source point cloud is identified. Thereafter, a transformation (e.g., a transformation matrix) between the source point cloud and the target point cloud can be estimated using the corresponding points. Thereafter, based on the transformation, a data fusion section may perform transformation, alignment, and superimposition of the source point cloud and the target point cloud to generate an integrated image used for display on a display section.
[0005] Feature matching components are an important part of any registration pipeline and aim to find the proper corresponding points between a source point cloud and a target point cloud. As an example, to facilitate high accuracy in the results of rigid registration, it is necessary to identify at least three sets of corresponding points in both point clouds. Based on these three points, a transformation matrix can be calculated and then used to align the source point cloud and the target point cloud. The most frequently used algorithms for feature matching are the Random Sample Consensus (RANSAC) method and, for example, derivatives of the RANSAC method disclosed by Fischler et al. (Martin A. Fischler and Robert C. Bolles, 1981, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography”, Commun. ACM 24, 6 (June 1981), 381-395). This method is an iterative approach that aims to randomly sample points from the source and target point clouds and perform a match between these points and the points in the target point cloud that are most likely to correspond. Subsequently, the transformation matrix is estimated using the correspondence. Then, using an accuracy metric, it is determined whether the estimated transformation matrix provides a higher accuracy than that estimated from previous iterations. The iterative process is repeated until an end criterion is reached.
[0006] The accuracy of the results is high by the method using RANSAC, but since it is an iterative process, a lot of time is required to execute this method. Furthermore, based on the selection result of the wrong corresponding point pair that brings about the wrong registration result, a low-accuracy result may be obtained. As a result, the use of this method is restricted in examples applied in real time in surgical procedures and endoscopic procedures. Another drawback of the conventional PCR algorithm occurs because the field of view obtained by the endoscopic camera during the procedure is limited. In an endoscope, usually only a small area that is part of the organ surface can be seen, and this may be reconstructed as a source point cloud. In many cases, this means that the smaller the size of the area, the higher the possibility of obtaining an ambiguous solution for the selection result of the corresponding points. As shown in Figure 2, in some cases, it can be seen that the accuracy of the result is clearly high even though it is based on something that mimics the wrong registration result following the selection of the wrong corresponding points. To avoid the above problems, commercially implemented PCR and feature matching algorithms require interaction with the user to select the corresponding point pairs from both the source point cloud and the target point cloud. However, even in this case, the accuracy of this method is extremely low, especially when used for registration in application examples of soft tissues such as abdominal surgery.
Prior Art Documents
Non-Patent Documents
[0007]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] In view of the above circumstances, an object of the present invention is to address the problems of the prior art described above and enable high-speed and reliable registration of source data and target data so as to improve the efficiency of the registration process, and to provide a data processing method, a data processing system, a data processing program, and a training method for processing image data.
Means for Solving the Problems
[0009] According to a first aspect of the present invention, the above object is achieved by a data processing method for processing image data, the method including: providing source data from an image captured by a camera that captures an object; providing target data related to the object; operating an AI system including at least one artificial neural network trained to identify a first correlation between the source data and the target data, inputting the source data into the AI system, and identifying the first correlation; and processing the source data and the target data based on the first correlation to obtain a second correlation between the source data and the target data.
[0010] According to an important idea of the present invention, an artificial intelligence system (AI system) is used to search for a first correlation between source data and target data and use this first correlation to more efficiently and / or more accurately search for a second correlation between the source data and the target data. Therefore, the target data and the source data can be aligned with each other more quickly and more accurately, thereby enabling effective integration of information obtainable from the source data and information obtainable from the target data.
[0011] In particular, according to the present invention, the AI system is trained to identify (predict) a first correlation between source data and target data. The training of the AI system can be performed using a plurality of different training data sets, each training data set including a source data set obtained from at least one image captured by a training camera that photographs a target, a target data set related to the target, and a first correlation between the source data set and the target data set. Then, using the trained AI system as described above, a first correlation can be identified based on the first-seen source data obtained from an image captured by a camera that photographs a first-seen target, and in particular, the first correlation can be identified without requiring target data in this step of the present data processing method.
[0012] In one embodiment of the present invention, the second correlation defines a spatial relationship between the source data and the target data, the spatial relationship describes at least one spatial transformation between at least a part of the source data and at least a part of the target data, and the spatial transformation is at least one of translation, rotation, reflection, and deformation or a combination thereof. Therefore, the position information can be aligned between the source data and the target data using the second correlation, and in particular, registration can be performed between a source point group described by the source data and a target point group described by the target data. Therefore, the position information of the source data and the position information of the target data can be merged.
[0013] When a second correlation is searched, one or more points or areas in the image captured by the camera (providing the basis for the source data) may be annotated, interpreted, supplemented, or otherwise combined with the information obtained from the target data acquired for the point or area of the object (through another imaging). As a specific example, such a combination may be implemented by superimposing an image corresponding to at least a part of the source data (particularly, at least a part of the image captured by the camera) and an image obtained from a spatial transformation of at least a part of the target data. In other words, the images of the same part of the same object, obtained from the camera on the one hand and from another imaging modality on the other hand, may be aligned and displayed by superimposing them.
[0014] Since the source data and the target data are acquired from different measurements using different imaging modalities, different viewpoints, shooting angles, etc., spatial transformation may be required, particularly for registering the source point cloud and the target point cloud, in order to align the source data and the target data. In addition, for example, due to deformation of soft tissue organs caused by intraoperative inflation, patient alignment, mechanical shock, etc., the spatial transformation between the source data and the target data may include deformation based on the actual deformation of the object from preoperative imaging to intraoperative imaging, and thus the shape of the object itself, particularly the shape of the organ, may change. Furthermore, due to the difference between the measurement area and the field of view of one camera and those of another imaging modality on the other hand, deformation components may occur in the spatial transformation.
[0015] In another embodiment of the present invention, the source data relates to a portion of an object observed by a camera, the target data relates to a target area that is larger than the observed portion and includes the observed portion, and the AI system is trained to identify a first correlation as something that describes the position of the observed portion within the target area (particularly position estimation). Thus, the camera may observe only a small portion of the object (e.g., a small surface section) within the field of view of the camera, and the target data may be obtained through another imaging modality for a large area of the object or the entire object, and the AI system may be trained to predict the approximate position of the observed small portion of the source data within the large target area of the target data. Thereafter, the search for the second correlation can be accelerated using the predicted position of the source data within the target data, for example, by reducing the solution space by cropping the target data. Cropping of the target data will be described in more detail later. Thus, the possibility of finding an appropriate transformation in a short time can be increased, and in particular, the possibility of finding a pair of corresponding points between the target point cloud and the source point cloud can be increased.
[0016] The first correlation identified by the AI system may include specific position information such as the coordinates of the observed portion within the target area. However, the effect of the present invention of increasing the speed and accuracy of searching for the second correlation has already been achieved by using an AI system trained to classify the source data with respect to the position within the target area, and has already been achieved when the AI system is trained to perform only approximate position estimation. In a simple embodiment, two classes are sufficient for effective classification (e.g., left or right, up or down). Further, three or more classes can be defined based on the segmentation of the object, i.e., the AI system can identify the source data as corresponding to a specific segment X within the target data of the organ. Further, such combinations of classifications can also be considered.
[0017] When obtaining a second correlation by processing source data and target data, particularly in the case based on the classification of the source data determined by an AI system, the step of obtaining the second correlation preferably includes determining a cropping area that extends around the position of the observed portion within the target area (or, when the observed portion is aligned within the target area according to the first correlation, the cropping area extends around the area corresponding to the observed portion), and processing the target data within the cropping area and the source data to obtain the second correlation. Therefore, based on the first correlation, the target data is cropped by removing data corresponding to positions outside the cropping area that includes the observed portion of the source data. Therefore, the size of the target data is effectively reduced, and the search space, i.e., the solution space, used to search for the second correlation is reduced.
[0018] To actively search for the first correlation, the AI system may be trained to recognize the presence of target features in at least one of the source data and the target data and determine the positions of the target features in both the coordinate system of the source data and the coordinate system of the target data. Visible features (points, groups of points, or areas) in the images corresponding to the source data and the target data respectively, i.e., visible features suitable for enabling the AI system to determine that the visible features belong to the target features of the object being inspected, may correspond to the target features. For example, an area where various blood vessels of an organ intersect can form a target feature in the sense of the present invention, thereby making it possible to generate a visible feature in the image corresponding to the source data and another visible feature in the image corresponding to the target data. The AI system may be trained to recognize both visible features in the images of the source data and the target data respectively and identify that both visual features correspond to the same target feature of the object based on the similarity between the two visual features.
[0019] The target feature may be a visible structural feature of the target that enables visual features to be recognizable (distinguishable from other visual features) within the images of the source data and the target data. Alternatively, the target feature may be a dye marking artificially applied to the target to mark points, areas, or structures within the organ and make the feature or area of the organ visible or more visible within the image corresponding to the source data and the image corresponding to the target data.
[0020] In another embodiment of the present invention, the source data may include the spatial coordinates of a plurality of source points forming a source point cloud, the target data may include the spatial coordinates of a plurality of target points forming a target point cloud, and the step of processing the source data and the target data to obtain a second correlation may include point cloud registration for estimating a transformation matrix between at least a part of the source point cloud and at least a part of the target point cloud. Preferably, a feature matching algorithm is used for point cloud registration. Therefore, by the method of this embodiment, it is possible to align the three-dimensional source data and the three-dimensional target data and realize the integration of the information obtained from the target data and the information obtained from the source data. The transformation matrix may be a mathematical matrix defined to calculate the coordinates of the corresponding points in the target point cloud for each point in the source point cloud, or vice versa, that is, to calculate the coordinates of the corresponding points in the source point cloud for each point in the target point cloud, or any other algorithm or formula.
[0021] For point cloud registration, it is preferable to use a feature matching algorithm that recognizes visible features in the images of target data and source data corresponding to the same target feature and performs matching of the features, for example, an algorithm such as the one described above. The point cloud registration itself, particularly the feature matching algorithm itself, may be of a conventional type as described, for example, by Robu et al. (Robu MR, Ramalhinho J, Thompson S, Gurusamy K, Davidson B, Hawkes D, Stoyanov D, Clarkson MJ., “Global rigid registration of CT to video in laparoscopic liver surgery”, lnt J Comput Assist Radiol Surg. June 2018;13(6):947-956).
[0022] In yet another embodiment of the present invention, the target data is obtained from an imaging modality applied to the subject, particularly computed tomography (CT) and / or magnetic resonance imaging (MRI). Such imaging modalities are constructed to provide comprehensive three-dimensional data (target point cloud) through non-invasive measurements. The measurement principle is different from visual observation by a camera, and thus, it is possible to provide high-added-value additional information regarding the subject.
[0023] In yet another embodiment of the present invention, the source data and the target data are preferably processed using an augmented reality device for two-dimensional or three-dimensional display as visually recognized by the user, and the source data and the target data are preferably displayed so as to be superimposed on each other based on a second correlation. By the two-dimensional or three-dimensional display of the source data and the target data, the user viewing the display can directly observe the object, and the information obtained from both the camera image and another imaging modality can be displayed simultaneously. More comprehensive information can be displayed using a three-dimensional display such as a head-mounted stereoscopic display (3D glasses) that realizes a real or nearly real immersive viewing experience for the user.
[0024] In yet another embodiment, the source data may be processed for two-dimensional or three-dimensional display as visually recognized by the user as a source image, and the target data may be displayed to the user based on a second correlation, provided that it is not in the form of an image, but in the form of one or more markings (annotations, highlighted areas, and other visible features) generated by a computer and placed at positions within the source image corresponding to the second correlation. For example, computer-generated labels may be attached to some points or areas of the source image, or a part of an organ may be highlighted within the source image. It should be noted that the source image may be the same as the image captured by the camera, and processing the source data for display may include sending or transferring the image captured by the camera without modification.
[0025] According to a second aspect of the present invention, the above object is achieved by a data processing system for processing image data, comprising a camera adapted to capture an image of an object, a source data processing unit adapted to provide source data based on the image captured by the camera, a target data processing unit adapted to provide target data related to the object, at least one artificial neural network trained to identify a first correlation between the source data and the target data, an AI system adapted to receive the source data as an input, and a registration unit adapted to process the source data and the target data based on the first correlation to obtain a second correlation between the source data and the target data.
[0026] By using the system according to the second aspect of the present invention, the technical effects described above for the data processing method according to the first aspect of the present invention (including the embodiments described above) can be realized. In particular, the data processing system according to the second aspect of the present invention may be configured to execute the data processing method according to the first aspect of the present invention, including one or more of the above embodiments. In particular, some parts and components of the embodiments of the data processing system according to the second aspect of the present invention are described in the dependent claims, and have the same or similar effects as the corresponding method steps described above for the embodiments of the first aspect of the present invention.
[0027] In yet another embodiment of the present invention, the display device may be an augmented reality device adapted to perform real-time display of source data according to an image captured by a camera. The augmented reality device may be further adapted to display target data information obtained from target data so as to be superimposed on the source data. By displaying the source data in real time, the user can track the changes in the source data without interruption, and thus can recognize time-dependent information in real time. In particular, when the source data displayed by the augmented reality device is equivalent to or corresponds to the image captured by the camera, the user can observe the live stream of the target image, and thus can observe the movement of the target and the movement of the camera relative to the target without interruption. Further, when the augmented reality device is adapted to process the source data in real time, i.e., to search for the first correlation and the second correlation in real time, target data information can be obtained from the target data related to the current source data, and can be displayed in real time on the same display, and in particular, can be displayed so as to be superimposed on the current source data.
[0028] The term "real time" in the present disclosure refers to a time delay or latency that is small enough for the user to be able to use the data processing system continuously without interruption, i.e., a time delay or latency that either allows continuous operation such that the user is unaware or, if aware, there is no time delay or latency. For example, "real time" refers to a time delay or latency of less than 5 seconds, preferably less than 1 second, and more preferably less than 500 milliseconds.
[0029] In particular, in the data processing system according to the second aspect of the present invention, the time delay between the first time when the camera captures a specific image of the object, and the second time when the registration unit acquires the second correlation between the source data provided based on the specific image and the target data is less than 5 seconds, preferably less than 1 second, more preferably less than 500 milliseconds. Such a data processing system enables not only real-time display of the camera image but also real-time display of relevant target information superimposed on the current source data. Therefore, the user can directly recognize the target information related to the live camera image, and the target data information is displayed at an appropriate position within the camera image.
[0030] In yet another embodiment of the second aspect of the present invention, the data processing system may be an endoscopic imaging system, and the camera may be an endoscopic camera adapted to be attached to the endoscope. By implementing the data processing system as an endoscopic imaging system, a wide range of applications for observing complex objects or objects that are difficult to access, such as cavities, internal regions of the object, especially tissues, are opened up. At that time, the data processing system may be further configured to be used for gastrointestinal intervention or surgical intervention in the human body.
[0031] The data processing system according to the second aspect of the present invention may be configured to execute the method according to the first aspect of the present invention, in particular the data processing method according to at least one of claims 1 to 8, so as to achieve the same effects or corresponding effects as those described above for the first aspect of the present invention.
[0032] According to a third aspect of the present invention, the above object is achieved by a data processing program configured to execute, when executed by a computer, a data processing method according to the first aspect of the present invention, particularly an embodiment of the first aspect of the present invention as described above. The data processing program may particularly provide source data from an image captured by a camera that captures a target, provide target data related to the target, and provide and / or operate an AI system comprising at least one artificial neural network. The source data is used as an input to the AI system, and the AI system is trained to identify a first correlation between the source data and the target data. The data processing program may further comprise an instruction for processing the source data and the target data based on the first correlation to obtain a second correlation between the source data and the target data. Such a data processing program exhibits the actions and effects described above for the data processing method of the first aspect of the present invention. The data processing program of the third aspect of the present invention is particularly suitable for being executed by a computer that forms part of the data processing system of the second aspect of the present invention so as to exhibit the actions and effects described above for the second aspect of the present invention.
[0033] According to a fourth aspect of the present invention, there is provided a training method for training an AI system including at least one artificial neural network, the method comprising: providing a plurality of different training data sets; and training the AI system by inputting the training data sets into the AI system, each training data set comprising a source data set obtained from at least one image captured by a camera photographing a target, a target data set related to the target, and a first correlation between the source data set and the target data set. Such a training method may be implemented as a supervised machine learning algorithm that uses the first correlation as a classification of source data with respect to related target data. Preferably, this classification is a positional classification. In particular, the source data preferably corresponds to a source point cloud obtained from a camera image, the target data preferably corresponds to a target point cloud obtained from another imaging modality of the same target, and the first correlation preferably corresponds to a classification of the position of the source point cloud within the target point cloud.
[0034] In one embodiment of the present invention, the classification is performed for a finite number of classes (e.g., less than 20 classes, preferably less than 8 classes) so as to specify an approximate position (e.g., left, right, bottom, top) and / or the type of the target (e.g., one of a predetermined number of organs). To prepare the training data set, a specific first correlation between a specific source data set and a specific target data set, i.e., an appropriate positional classification of the source point cloud within the target point cloud, may be manually given by the user training the AI system (manual classification), or may be given using historical data including search steps and measured results constructed for the correlation between the source data and the target data. Thereafter, the AI system thus trained can predict the first correlation based on the source data of the first sight.
[0035] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the accompanying drawings.
Brief Description of the Drawings
[0036]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0037] FIG. 3 shows a data processing system 10 (hereinafter referred to as system 10) for processing image data according to a first embodiment of the present invention. The system 10 includes a camera 12, and the camera 12 is connected or connectable to a computing device 14 configured to process an image captured by a camera 12 that captures an object (not shown). The system 10 may further include a display unit 16 that displays the result of the processing of the computing device 14.
[0038] The camera 12 may be any suitable image capturing means capable of capturing a two-dimensional or three-dimensional image of the object. For example, a white light, narrow-band visible light, or infrared camera equipped with a CCD array that detects reflected light or emitted light from the object may be used.
[0039] The display unit 16 may be any means for visually presenting information, particularly two-dimensional or three-dimensional images, to the user. The display unit 16 may be, for example, a computer screen, a tablet screen, or a head-mounted screen (head-mounted device). In particular, the head-mounted device may be implemented in the form of 3D glasses equipped with two displays respectively visible to the user's left and right eyes, enabling the presentation of three-dimensional images (stereoscopic images).
[0040] The computing device 14 may be implemented by one or more computers comprising at least a processor, volatile and / or non-volatile data storage means, input means for receiving user input through, for example, a keyboard and / or a mouse (not shown), and connection means for connecting the computing device 14 to peripheral devices such as the camera 12 and the display unit 16 and / or for connecting the computing device 14 to a local network and / or a remote network such as the Internet. The computing device 14 is preferably configured and optimized to host and operate one or more neural networks so as to achieve high efficiency, particularly real-time AI processing, within the computing device 14. For example, it is preferably equipped with at least one deep learning processor (DLP) specialized for executing learning algorithms and / or at least one AI accelerator.
[0041] Computing device 14 may comprise a source data processing unit 18 adapted to process data, particularly images, obtained from camera 12 so as to provide source data corresponding to the images captured by camera 12. The source data may simply be the camera image itself (in which case the camera image passes through source data processing unit 18 unchanged), or may be a source point group including the spatial coordinates of a plurality of source points obtained from the camera image, preferably a three-dimensional source point group. For example, source data processing unit 18 may combine depth information captured by camera 12 for the object being photographed with a two-dimensional image simultaneously captured by camera 12 to generate a three-dimensional source point group.
[0042] Source data processing unit 18 may in particular be configured to process a live image (video) from camera 12, i.e., may be configured to update the source point group in real time according to the camera image. Accordingly, the source points may correspond to what camera 12 is currently photographing.
[0043] Computing device 14 is further configured to receive or store target data of the same object photographed by camera 12, and the target data is preferably obtained through another imaging modality 19, such as computed tomography (CT) and / or magnetic resonance imaging (MRI). Alternatively, the target data may be a model of the object, particularly a three-dimensional model, pre-stored in computing device 14.
[0044] Target data processing unit 20 of computing device 14 may be configured to process the target data and then obtain a target point group defined by the spatial coordinates of a plurality of target points. In addition to or instead of the above, target data processing unit 20 may output other data related to the object in the form of target information, such as in the form of labels, annotations, metadata, and other information.
[0045] In this embodiment, the computing device 14 has an AI system 22, and the AI system 22 may be any suitable machine learning module, and in particular, may include at least one neural network to be trained. The AI system 22 may include at least one dedicated processor, for example, a deep learning processor (DLP) specialized for executing learning algorithms or a general-purpose microprocessor, and the processor may host at least one neural network. Alternatively, the AI system 22 may be implemented as a physical or virtual part of any other processing unit of the computing device 14, and this part may also execute other processing, such as image processing.
[0046] In this embodiment, the AI system 22 includes a deep learning network trained to classify the position of the source point cloud in the related target point cloud for the source point cloud input to the AI system 22. Such a classification corresponding to the first correlation in the meaning of the present invention may be the approximate position of the source point cloud in the target point cloud. The approximate position in the present disclosure means that the number of classes of such a classification is significantly less than the total number of possible positions of the source point cloud in the target point cloud so as to achieve a reasonable balance between the prediction accuracy and the processing time of the AI system. The number of classes is preferably less than 20, and more preferably less than 8. It has been found that a classification with only two classes, for example, a classification with only "the left part of the object" and "the right part of the object", can also bring very good results, which will be further described later. Other classification methods for the position of the source point cloud with respect to the target point cloud may include, for example, four classes of "upper left", "lower left", "upper right", and "lower right".
[0047] Therefore, the AI system 22 provides a predicted classification of the position of the source point cloud in the target point cloud. The predicted classification result output by the AI system 22 may be provided as a first correlation to the target point cloud cropping unit 24, and the target point cloud cropping unit 24 receives the target point cloud from the target data processing unit 20 as a second input. Based on the first correlation, the target point cloud cropping unit 24 may crop the target point cloud to generate a cropped target point cloud that includes only the points of the original target point cloud within or near the predicted position of the source point cloud. In particular, for the target area defined by the outer periphery of the target point cloud, the target point cloud cropping unit 24 may determine a cropping area corresponding to the outer periphery of the source point cloud according to the first correlation, that is, the approximate position classified by the AI system 22, when the source point cloud is arranged in the target point cloud.
[0048] Thereafter, the cropped target point cloud may be passed from the target point cloud cropping unit 24 to the point cloud registration unit 26. The point cloud registration unit 26 may further receive source data, particularly a copy of the source point cloud, from the source data processing unit 18. Thereafter, the point cloud registration unit 26 may use any point cloud registration algorithm including a conventional point cloud registration algorithm, preferably, an algorithm based on feature matching known in the prior art as such an algorithm (see, for example, Robu et al. above). Further, a random sample consensus (RANSAC (Fischler et al. above)) method, which is an iterative method aimed at randomly sampling points from the source point cloud and the target point cloud, performing matching of these points, and searching for the points with the highest probability of correspondence between the two point clouds, may be used as the feature matching algorithm.
[0049] The result of the point cloud registration and the output of the point cloud registration unit 26 may describe a spatial transformation between at least a part of the source point cloud and a part of the target point cloud, for example, a transformation matrix. Such a transformation forms a second correlation in the sense of the present invention.
[0050] The computing device 14 of the present embodiment further includes a data fusion unit 28 that receives the second correlation as a first input. Further, the data fusion unit 28 receives source data (for example, a copy of the source point cloud) from the source data processing unit 18 as a second input, and receives target information (for example, a copy of the target point cloud or a part thereof) from the target data processing unit 20 as a third input. Then, the data fusion unit 28 may generate an integrated image in which the source data and the target information are superimposed on each other based on the second correlation. In particular, in order to properly align the source data and the target information with each other in the integrated image, the data fusion unit 28 can apply a transformation according to the second correlation to either the target information or the source data. In one example, the source data may be the same as or substantially the same as the camera image provided by the camera 12, the target information may be an image obtained from another imaging modality, and the positions, orientations, sizes, and shapes of both images are matched so that the positions of the corresponding points corresponding to the specific target points of the original object are aligned.
[0051] It is preferable that the data fusion unit 28 provides the integrated image as a moving image based on the live image (raw video) captured by the camera 12, and the integrated image is provided in real time, or at least with a time delay that enables continuous observation of the target and relative movement between the camera 12 and the target. For example, the time delay between the first time when the camera 12 captures a specific image of the target and the second time when the data fusion unit 28 integrates the source data of the specific image and the target information based on the second correlation obtained from the source data of the specific image may preferably be less than 5 seconds, more preferably less than 1 second, and most preferably less than 500 milliseconds. Such a delay is either unrecognizable by the user or does not significantly interfere with the continuous operation of the system 10.
[0052] FIG. 4 shows an example of a source point cloud 30 and a target point cloud 32 that can be processed by the system 10 according to the first embodiment of the present invention. In this example, the source point cloud 30 is a two-dimensional white light image of the right part of the liver captured by the camera 12. The target point cloud 32 is a two-dimensional image of the entire liver obtained from a whole-body CT scan or a whole-body MRI scan of the liver.
[0053] The AI system 22 of this example is trained to recognize several parts of the liver. For example, the training data used to train the AI system 22 may include a plurality of training data sets, and each training data set includes an image of a part of the liver as a source data set, an image of the entire liver as a related target data set, and an appropriate classification (first correlation) indicating the approximate position and area of the source point cloud in the target point cloud. After sufficiently training the AI system 22 using a plurality of different training data sets, the AI system 22 can predict the first correlation of an unseen source point cloud. From this, in the example shown in FIG. 4, the AI system 22 predicts that the source point cloud 30 is aligned with the right part of the liver. As a result, the target point cloud cropping unit 24 crops the target point cloud 32 and provides a cropped target point cloud 34 that includes only the points belonging to the right part of the liver. As a result, the size of the target point cloud is reduced to approximately half, which means that the target area that needs to be analyzed by the point cloud registration unit 26 for registration with the source point cloud 30 is also reduced to approximately half.
[0054] The point cloud registration unit 26 receives the source point cloud 30 and the cropped target point cloud 34, and performs point cloud registration to search for the exact transformation between the two point clouds 30, 34. Since the cropped target point cloud 34 is significantly smaller in size than the original target point cloud 32, the processing speed of the point cloud registration unit 26 can be increased accordingly.
[0055] After that, an integrated image 36 may be generated by transforming both the source point cloud 30 and the target point cloud 32 into a common coordinate system based on the second correlation so that the points of each group corresponding to the same target point are superimposed at the same coordinates. In the integrated image 36, the pixels obtained from the source point cloud may be displayed in a color different from the pixels obtained from the target point cloud so that the two groups are visually distinguishable.
[0056] FIG. 5 shows experimental data obtained by training and testing the system 10 in an example of an operation mode for observing the liver. In this example, PointNet was used as an artificial intelligence network that uses a binary classification method, that is, predicts the binary position (left lobe of the liver or right lobe of the liver) of the source point cloud in the target point cloud. The AI system 22 was trained on samples (n = 19,000) of the source point cloud obtained from randomly sampled surface sections (observed target parts) located in the left lobe of the liver and the right lobe of the liver, respectively. After training, the trained AI system was tested using new source point clouds (n = 4,000) obtained from images of the liver that were not used to train the network, and the targets of this image had different anatomical shapes. The results obtained from the test show that the true positive detection rate is 94% and the false negative detection rate is only 6%.
[0057] Furthermore, the performance of point cloud registration according to an embodiment of the present invention using the cropped target point cloud was measured and compared with the performance of point cloud registration using the same source point cloud and the original (uncropped) target point cloud as inputs. FIG. 5 shows the results of this comparison for both the left part and the right part of the liver. For both liver parts, when using the cropped target point cloud instead of the original target point cloud, the root mean square error (RMSE) of the point cloud registration algorithm is significantly reduced, and as a result, an average RMSE improvement rate of about 61.09% for the left liver part and about 64.56% for the right liver part is obtained.
[0058] It has been shown that the AI system 22 has been successfully trained to predict the position of a smaller source point cloud in a larger target point cloud, and that the first correlation predicted to appropriately crop the target point cloud can be used, whereby the search space (solution space) of the point cloud registration algorithm is reduced, resulting in a significant increase in the accuracy of registration regarding the alignment of both point clouds, an increase in its robustness, and / or a significant acceleration of the point cloud registration algorithm.
[0059] FIG. 6 shows a system 110 according to a second embodiment of the present invention, which is a modified example of the system 10 according to the first embodiment of the present invention described above with reference to FIGS. 3 and 4. Only the points of modification and differences from the first embodiment will be described in more detail, and for all other features and operations, refer to the above description of the first embodiment.
[0060] The system 110 may include a camera 112 that illuminates an object to be observed with light of different wavelengths. When illuminating with white light, the camera 112 may capture a white light image (WLI). When illuminating the object with light of a predetermined color (light within a narrow wavelength band of visible light) corresponding to the excitation wavelength of a predetermined first dye, the camera 112 may further capture another image (dye 1). Further, the camera 112 may have the ability to illuminate the object with light of a second color or yet another color (preferably another wavelength band that is completely different from each other) corresponding to the excitation wavelength of the second dye and / or yet another dye. At this time, the camera 112 may capture another camera image (dye 2,...).
[0061] Camera images (WLI, Dye 1, Dye 2) may be input into the computing device 114 as source data. The computing device 114 is preferably a type of computing device as described above for the computing device 14 of the first embodiment. In the second embodiment, the computing device 14 may include a first AI system AI1 that receives an image (WLI and Dye 1) as input and is trained to recognize a source structure in the source data based on the light reflected by the first dye. Further, a second AI system AI2 may receive an image (WLI and Dye 2) as input, and the second AI system AI2 may be trained to recognize the same source structure or another source structure in the source data based on the light reflected by the second dye. By using different dyes due to the different reflection / absorption characteristics of different dyes and the various interactions between the dyes and the object (such as the interaction with the part of the object passing through the object), information about one or more source structures within the object, such as blood vessels and regions where blood vessels intersect, can be added. Therefore, the accuracy and robustness of structure recognition are higher compared to structure recognition based only on white light images.
[0062] Thereafter, the source structure identified in the source data may be passed to a third artificial intelligence system AI3, which includes a machine learning network trained to recognize the source structure as corresponding to a specific target structure within the target point cloud. To do this, AI3 may be trained using a supervised training algorithm with a plurality of training data sets, each training data set including a definition of a sample source structure and an indication (classification) of the presence and / or position of the sample source structure within the target point cloud. After training, AI3 can predict whether a source structure of interest exists at a predetermined position within the target point cloud or within a predetermined region of the target point cloud.
[0063] Such information regarding the presence and / or position of the source structure forms a first correlation in the sense of the present invention and is provided to the point cloud registration unit 126 of the computing device 114. In particular, the point cloud registration unit 126 receives as input source data, in particular a white light image captured by the camera 112, target data, in particular a target point cloud obtained from another imaging modality 119, and the first correlation, and calculates a second correlation between the source data and the target data, which corresponds to the transformation of the points of the source data and the points of the target data. The calculation of the transformation may be based on any suitable PCR algorithm as described above for the first embodiment.
[0064] According to the second embodiment, the PCR calculation is assisted by a first correlation that provides information about the structure present in the source data. Further, the target data may also have the same structure or corresponding features. In any case, by the structure recognition as described above, the solution space of the point cloud registration algorithm, particularly the number of point pairs considered within the feature matching algorithm, can be significantly reduced. For example, the number can be significantly reduced by considering only the points belonging to the detected structure or by preferentially considering such points. In other words, based on the prediction made by the AI system, promising sample point candidates within the source point cloud and / or target point cloud can be estimated, and feature matching (e.g., Random Sample Consensus) can be restricted to these point candidates of the samples. As a result, the effectiveness of the point cloud registration unit can be enhanced, the feature matching result can be improved, and a more robust and efficient registration can be realized. In particular, by increasing the processing speed of the point cloud registration unit 126, real-time processing of the source data becomes possible. For example, the display unit 116 can display a continuous video corresponding to the live image captured by the camera 112, and continuously enhance the value of the video with the relevant target information. In this case, the processing delay is small enough not to be noticed by the user or to prevent continuous operations by the user (the delay time is preferably less than 5 seconds, more preferably less than 1 second, and even more preferably less than 500 milliseconds).
[0065] The display unit 116 may be an extended reality device as described above with respect to the display unit 16 of the first embodiment.
Description of Reference Numerals
[0066] 10 System 12 Camera 14 Computing Device 16 Display Unit 18 Source Data Processing Unit 19 Imaging Modality 20 Target Data Processing Unit 22 AI System 24 Target Point Cloud Cropping Unit 26 Point Cloud Registration Unit 28 Data Fusion Unit 30 Source Point Cloud 32 Target Point Cloud 34 Target Point Cloud 36 Integrated Image 110 System 112 Camera 114 Computing Device 116 Display Unit 119 Imaging Modality 126 Point Cloud Registration Unit
Claims
1. providing source data from an image captured by a camera photographing an object; providing target data relating to said subject; operating an AI system comprising at least one artificial neural network trained to identify a first correlation between the source data and the target data, inputting the source data into the AI system and identifying the first correlation; processing the source data and the target data based on the first correlation to obtain a second correlation between the source data and the target data; A data processing method for processing image data, comprising:
2. 2. The data processing method of claim 1, wherein the second correlation defines a spatial relationship between the source data and the target data, the spatial relationship describing at least one spatial transformation between at least a portion of the source data and at least a portion of the target data, the spatial transformation being at least one of a translation, a rotation, a reflection and a deformation or a combination thereof.
3. the source data pertains to a portion of the object viewed by the camera; the target data relates to a target area that is larger than and includes the observed portion; The data processing method of claim 1 , wherein the AI system is trained to identify the first correlation as describing a location of the observed portion within the target area.
4. The step of processing the source data and the target data to obtain the second correlation comprises: defining a cropping region extending around the location of the observed portion within the target area; 4. The data processing method of claim 3, further comprising processing the target data of the cropping region and the source data to obtain the second correlation.
5. 2. The data processing method of claim 1, wherein the AI system is trained to recognize the presence of a feature of interest in at least one of the source data and the target data and to determine the position of the feature of interest in both the coordinate system of the source data and the coordinate system of the target data, the feature of interest being a pigment marking or a visible structural feature of the object.
6. the source data includes spatial coordinates of a plurality of source points forming a source point cloud; the target data includes spatial coordinates of a plurality of target points forming a target point cloud; the step of processing the source data and the target data to obtain the second correlation includes point cloud registration to estimate a transformation matrix between at least a portion of the source point cloud and at least a portion of the target point cloud, the point cloud registration using a feature matching algorithm; 2. The data processing method according to claim 1.
7. 2. The data processing method of claim 1, wherein the target data is obtained from imaging modalities applied to the subject: computed tomography (CT) and / or magnetic resonance imaging (MRI).
8. The data processing method of claim 1 , wherein the source data and the target data are processed using an augmented reality device for two-dimensional or three-dimensional display to be viewed by a user, and the source data and the target data are displayed superimposed on each other based on the second correlation.
9. a camera adapted to capture an image of the object; a source data processing unit adapted to provide source data based on the image captured by the camera; a target data processing unit adapted to provide target data relating to said object; an AI system adapted to accept the source data as an input, the AI system comprising at least one artificial neural network trained to identify a first correlation between the source data and the target data; and a registration unit adapted to process the source data and the target data based on the first correlation to obtain a second correlation between the source data and the target data.
10. 10. The data processing system of claim 9, wherein the registration unit is adapted to determine a transformation between at least a portion of the source data and at least a portion of the target data, the transformation being at least one of a translation, a rotation, a reflection and a deformation, or a combination thereof.
11. the source data pertains to a portion of the object viewed by the camera; the target data relates to a target area that is larger than and includes the observed portion; the AI system is trained to identify the first correlation as describing a location of the observed portion within the target area; 10. The data processing system of claim 9, further comprising a cropping unit adapted to define a cropping region within the target area, the cropping region extending around the location of the observed portion, and the registration unit adapted to process the source data and the target data of the cropping region to obtain the second correlation.
12. 10. The data processing system of claim 9, further comprising a display device adapted to provide a two-dimensional or three-dimensional display of the source data and the target data to be viewed by a user, the display device adapted to display the source data and the target data so as to be superimposed on one another based on the second correlation, the display device comprising an augmented reality device adapted to provide a real-time display of the source data according to the image captured by the camera, the augmented reality device further adapted to provide a real-time display of target information obtained from the target data so as to be superimposed on the source data.
13. 10. The data processing system of claim 9, wherein the data processing system is an endoscopic imaging system and the camera is an endoscopic camera adapted to be attached to an endoscope.
14. A data processing program that causes a computer to execute the data processing method according to claim 1.
15. 1. A method for training an AI system comprising at least one artificial neural network, comprising: providing a plurality of different training data sets; training the AI system by inputting the training data set into the AI system; 11. A training method, wherein each of the training datasets comprises a source dataset obtained from at least one image captured by a camera photographing an object, a target dataset related to the object, and a first correlation between the source dataset and the target dataset.
Citation Information
Patent Citations
Image processing method and device, electronic device, storage medium, and program product
JP2022543531A
Computer-implemented method for executing image registration, system and program (deformable registration of medical images)
JP2023026400A