Apparatus and method for processing image data of a medical imaging device, and method for training an artificial intelligence entity
The device and method for processing medical imaging data address the challenge of determining actual sizes of scaleless 3D structures by using a scaling module to calculate a scale factor, thereby enabling accurate size determination across different medical imaging scales.
Patent Information
- Application Number
- PCT/EP2024/084665
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-04
- Publication Date
- 2025-06-26
AI Technical Summary
Current medical imaging technologies struggle to determine the actual size of scaleless 3D structures in image data, lacking a reliable scale factor to relate these structures to real-world sizes.
A device and method that process image data from medical imaging devices by using a structure recognition module to identify scale-free 3D structures and a scaling module to determine a scale factor, allowing for the calculation of actual sizes of components within these structures.
Enables the accurate determination of actual sizes of 3D structures in medical imaging, facilitating applications across various medical scenarios regardless of scale, from microscopic to macroscopic.
Smart Images

Figure EP2024084665_26062025_PF_FP_ABST
Abstract
Description
[0001] Apparatus and method for processing image data of a medical imaging device and method for training an artificial intelligence entity
[0002] The present invention relates to a device and a method for processing image data from a medical imaging device, in particular to improve the determination of dimensions or size ratios in 3D structures captured in the image data. The invention also relates to a method for training an artificial intelligence entity to determine a scale factor for a scaleless 3D structure in image data from a medical imaging device.
[0003] Imaging techniques are frequently used in modern medicine. A medical imaging device is used to capture images of a medical scene. These images can provide a user, such as a surgeon, with information that the user could not visually perceive without aids, either because the images originate from a location not directly visible (e.g., inside a patient) or because the information is based on electromagnetic radiation with non-visible wavelengths.
[0004] The use of such imaging devices as aids thus brings many advantages, but also a certain indirectness. For example, it becomes more difficult for the user to develop an intuitive understanding of the captured medical scene. This is particularly true for scale. By moving an input lens of the imaging device closer or closer, the structures of the medical scene appear larger or smaller, respectively.
[0005] Methods are known for determining the relative dimensions or sizes of structures or sections of a medical scene, for example, in the context of generating a depth map. However, the state of the art lacks a scale factor that relates the structure—represented in a size-consistent manner—to real-world sizes.
[0006] For example, the scientific publication “Detecting Deficient Coverage in Colonoscopies” by D. Freedman et al., arXiv:2001,08589v3, from March 29, 2020 (hereinafter referred to as “Freedman et al.”) describes a network structure that is trained using unsupervised learning to generate a depth map based on a standard RGB image of intestinal structures. This is used to determine whether there are any areas in a series of sequential RGB images that were not visible in the RGB images. If this is found, the corresponding areas can be deliberately examined again. Freedman et al. explain that the generated depth maps can only be determined up to a scale factor, which is random, but this is effectively compensated for by the algorithm of the described network structure. Furthermore, for the purpose of Freedman et al.It is also clearly unimportant to know the true scale factor, since the goal is to determine whether there are any areas that were not seen, regardless of their actual size. It should also be noted that in Freedman et al.'s only application, namely colonoscopy, both the training and application image data all contain structures with very similar proportions, and thus the random scale factor will be within a narrow range of values anyway.
[0007] It is therefore an object of the present invention to provide an apparatus and a method that process image data in such a way that the actual size of scaleless 3D structures captured in the image data can be determined. A further object is to train an artificial intelligence entity suitable for use in such processing.
[0008] This object is solved by the subject matter of the independent patent claims of the present invention.
[0009] According to a first aspect, a device for processing image data from a medical imaging device is provided. This device comprises at least: an input interface configured to receive image data acquired by a medical imaging device; a computing device configured to implement at least one structure recognition module and a scaling module; wherein the structure recognition module is configured to determine a scale-free 3D structure in the acquired image data; wherein the scaling module is configured to determine a scale factor for the scale-free 3D structure, based on which scale factor an actual size of at least one component of the scale-free 3D structure (preferably the entire 3D structure) can be determined; and an output interface configured to generate an output signal that includes or indicates (or indicates) the determined scale factor.
[0010] A fundamental idea of the present invention is that prior art methods, which can partially provide useful reconstructions of scaleless 3D structures (e.g. depth maps), can be supplemented with the determination of a scale factor so that the actual size of 3D structures in the image data can later be determined using the scale factor.
[0011] Scaleless 3D structures are understood in particular to be 3D structures which specify accurate positional relationships and distances between points of the 3D structure, but do not relate these to an actual absolute value for the distances, in particular not to an actual distance exhibited by the points in the physical 3D structure from which the image data in which the scaleless 3D structure was determined was acquired. According to one aspect of the invention, a scaleless 3D structure is thus first determined, i.e. a 3D structure is determined up to a scale factor, and then, or simultaneously but separately, the associated scale factor is determined. This device can advantageously be used for all types of medical scenarios, in particular also those which have completely different scales or sizes.In contrast to the known state of the art, the same device can therefore be used for both microscopic and macroscopic, endoscopic and external use.
[0012] The image data acquired by the imaging device can, in particular, be one or more images, in particular a temporal series of images, of a medical scene. These are, in particular, images from a medical video, such as those recorded with endoscopes or microscopes. The images of the scene are acquired in the optical or near-infrared range, in particular with a conventional frame rate and image resolution, for example, HD or 4K resolution. The term "medical scene" is defined broadly here: It can refer to an external or even internal view of a patient who is currently undergoing or is about to undergo a medical procedure.In other words, the medical scene can also be a scene in which an organic, particularly human, tissue can be seen, for example in a laboratory or an operating room, either in vitro and / or in vivo.
[0013] Although some functions are described herein, above, and below as being performed by "devices," "interfaces," or "modules," it is understood that this does not necessarily mean that such devices, interfaces, or modules are provided as separate entities. In cases where one or more devices, interfaces, or modules are provided, in whole or in part, as software, the devices, interfaces, or modules may be implemented by sections or snippets of program code that are distinct from one another but may also be intertwined.
[0014] Similarly, in the case where one or more devices, interfaces, or modules are provided as hardware, the functions of one or more devices, interfaces, or modules may be provided by one and the same hardware component, or the functions of one device, interface, or module, or the functions of several devices, interfaces, or modules may be distributed among several hardware components that do not necessarily have to correspond one-to-one to the devices, interfaces, or modules. Therefore, any device, system, method, etc. that has all the features and functions attributed to a particular device, interface, and / or module is to be understood as constituting, comprising, or implementing the device, interface, and / or module.
[0015] In particular, it is possible for all modules to be implemented by program code that is executed by the computing device. The computing device can be realized as any device or means for computing, in particular for executing software, an app, or an algorithm. For example, the computing device can comprise at least one processor, such as at least one central processor (CPU), and / or at least one graphics processor (GPU), and / or at least one field-programmable gate array (FPGA), and / or at least one application-specific integrated circuit (ASIC), and / or any combination of the foregoing. The computing device can further comprise a main memory operatively connected to the at least one processor, and / or a non-volatile memory operatively connected to the at least one processor and / or the main memory.The computing device may be implemented partially and / or entirely in a local device and / or partially and / or entirely in a remote system, such as through a cloud computing platform.
[0016] The input interface can be configured to receive the image data directly from the imaging device, in particular in real time, or alternatively from a picture archiving and communications system (PACS), whereby the latter can also occur in real time, i.e. in particular as soon as the image data is received in the PACS. The device can also comprise an imaging device, so that receiving the image data can also comprise acquiring the image data by the imaging device. The device can comprise several different imaging devices, which can be designed, for example, to acquire medical scenes of different sizes (microscopic, macroscopic) or from different perspectives (endoscopic, open surgical, etc.). However, processing can always be carried out using the same computing device.
[0017] According to some preferred embodiments, variants, or refinements of embodiments, the computing device is additionally configured to implement an object recognition module configured to recognize at least one object with at least one known dimension in the obtained image data. For this purpose, for example, a database of objects of known size can be configured in the device, along with the respective known dimension(s).
[0018] For example, all dimensions of at least one object (or of all objects in the database) can be known, so that one can speak of a digital twin of the object(s). Alternatively, only some, or just a single, dimension can be known. Objects that can either be easily added to the medical scene from which the image data is acquired, or objects that are frequently visible in such medical scenes anyway, are particularly suitable for the database. The latter include, for example, medical instruments such as scalpels, scissors, holding instruments, laser fibers and / or the like, or parts or sections thereof, or markings on them.Especially for objects with a clearly recognizable elongated body, such as a laser fiber, it may be advantageous to know the diameter of this elongated body as the only dimension.
[0019] In some variants, one of the objects with at least one known dimension can also be a component of a tissue to be examined, for example, an implant or a tissue interval with a known dimension. The information about the dimension can be obtained, for example, from a database, such as a medical database, an implant database, or a patient database.
[0020] The scaling module may be configured to determine the scale factor based on the position and / or orientation of the detected object and its at least one known dimension, or at least based on the position and / or orientation of the at least one known dimension of the detected object.
[0021] The position of the object or its dimension can be a position of one or more points of the object or its dimension in three-dimensional space. The orientation of the object or its dimension can be a location of the object or its dimension in three-dimensional space, specified by the positions of two or more points of the object or its dimension that are in a known relationship to each other, for example, specified by a vector.
[0022] The object recognition module can be configured to first determine the position and / or orientation of the object in the acquired image data. The scaling module can be configured to relate the at least one known dimension of the object to the scaleless 3D structure based on the determined position and / or orientation and to determine the scale factor therefrom.
[0023] According to some preferred embodiments, variants, or refinements of embodiments, the structure recognition module and / or the scaling module comprise and utilize at least one artificial intelligence entity (AI). The artificial intelligence entity may, for example, be an artificial neural network (ANN).
[0024] According to some preferred embodiments, variants, or refinements of embodiments, the structure recognition module and the scaling module comprise a common artificial intelligence entity (KIE), which is configured to receive the obtained image data as input data and, based thereon, to generate an output indicating (or: indicating) the scaleless 3D structure and the scale factor. In other words, in this artificial intelligence entity (KIE), the 3D structure and the scale factor are generated jointly. A method for training an artificial intelligence entity suitable for this purpose is provided below as a further aspect of the present invention.According to some preferred embodiments, variants, or refinements of embodiments, the structure recognition module comprises a first artificial intelligence entity, KIE, which is configured to receive the obtained image data as input data and, based thereon, to generate a first output indicating the scaleless 3D structure, for example, as a point cloud, as a voxel structure, or as a vertex or polygon structure. The computing device can advantageously also be configured to implement a second artificial intelligence entity, KIE, which is configured to receive at least the first output of the first artificial intelligence entity, KIE, or data based thereon as input and, based thereon, to generate a second output.In other words, this variant provides a pipeline in which the scaleless 3D structure is first generated, followed by further processing using the second artificial intelligence entity (KIE). For example, the first KIE can be configured for 3D modeling, followed by a second KIE for instrument segmentation. This can contain either the obtained image data and / or the output of the 3D modeling as input images, in order to then calculate the scale factor from the width of, for example, an instrument shaft.
[0025] According to some preferred embodiments, variants, or refinements of embodiments, the computing device implements an object recognition module comprising the second artificial intelligence entity (KIE). The second KIE is advantageously configured to recognize the at least one object with the at least one known dimension in the obtained image data.
[0026] The second output, i.e., the output of the second artificial intelligence entity, KIE, can display (or index or encode) the recognized object, for example, selecting it from a list of known objects (in a database). For this purpose, this second artificial intelligence entity, KIE, can comprise or be an artificial neural network, KNN, which in particular has a softmax layer as an output layer. The scaling module can be configured to obtain the known size of the recognized object using the database and to determine the scale factor based thereon. For this purpose, the scaling module can comprise, for example, a third artificial intelligence entity, KIE. The first, the second, and / or the third artificial intelligence entity, KIE, can thus be integrated into an artificial intelligence pipeline.
[0027] According to some preferred embodiments, variants, or refinements of embodiments, the computing device is further configured to generate a scaled depth map of at least the determined scaleless 3D structure based on the determined scale factor and to store or output it. Preferably, the determined scaleless 3D structure comprises all of the acquired image data, or, in other words, a depth map of all of the acquired image data is generated as the scaleless 3D structure. This depth map, provided with the determined scale factor, can then be stored or output as a scaled depth map. The output can be sent, for example, to a display device, such as a screen or a touchscreen. Storage can occur either within the device or, for example, in a picture archiving and communications system (PACS).The generation of a depth map may be unnecessary if a stereo reconstruction is carried out. For this purpose, the imaging device can comprise a stereo camera or two separate lenses. For example, the imaging device can be a stereo endoscope. In other variants, however, a depth map can also be created in addition to a stereo reconstruction, and the method according to the invention can be carried out on this basis or the device according to the invention can process it. In these cases, dimensions within the image data can be determined on the one hand via the stereo reconstruction and on the other hand via the method according to the invention. A comparison of the results can then take place, and a warning can be issued to a user, for example on a display device, if the results are not within a predeterminable tolerance range of one another.Furthermore, the method according to the invention can also be used with a stereo camera, for example, if one of its lenses is unusable (e.g. dirty).
[0028] According to some preferred embodiments, variants, or refinements of embodiments, the structure recognition module is configured to determine the scale-free 3D structure substantially in real time. Advantageously, the structure recognition module is also configured to determine the scale factor in real time. Thus, a display device can advantageously always display (i.e., in real time) the currently acquired image data, together with the scale factor or at least one dimension based on the scale factor. For example, two instrument tips can be determined in the acquired image data (e.g., by segmenting the image data by an artificial intelligence entity), and the actual distance between the two instrument tips can be displayed automatically, or upon request from the user, by the display device. In this way, the user can perform length measurements using the instrument tips.
[0029] For the described real-time application, it is advantageous if obtaining the image data involves capturing the image data in real time using the imaging device. The imaging device can advantageously be a camera operating on the basis of visible or near-infrared light, such as an RGB camera, e.g., in an endoscope or microscope, since these generate image data that can be processed particularly quickly.
[0030] According to some preferred embodiments, variants, or refinements of embodiments, the device comprises the medical imaging device, in particular an endoscope or a microscope. The input interface can thus be configured to receive the image data acquired by the imaging device directly from the imaging device, for example, in real time.
[0031] Processing the image data in real time has the advantage that instructions can be provided to a user via a display device of the device. For example, the user can be prompted to introduce one of the objects with at least one known dimension into the detection area of the imaging device. This can occur periodically or rule-based, for example, when an error value determined for determining the scale factor exceeds a threshold, or when the depth map changes more than a predetermined threshold within a predetermined time.
[0032] On the other hand, processing the image data in real time has the advantage that a user of the imaging device, who often also operates an instrument at the same time, can always be informed about the actual dimensions of the 3D structure via the display device.
[0033] According to some preferred embodiments, variants, or refinements of embodiments, the medical imaging device comprises a camera control unit. Advantageously, at least the computing device is integrated into the camera control unit. Camera control units, particularly those of endoscopes, already have image processing means. Thus, the tasks of the computing device can be advantageously integrated into the camera control unit according to the present invention.
[0034] According to some preferred embodiments, variants or refinements of embodiments, the acquired image data comprises an image data set comprising first image data of a medical scene at a first time and at least second image data of the (same) medical scene at a second time.
[0035] The structure recognition module can also be configured to determine the scale-free 3D structure at least for the second image data based on both the first image data and the second image data. Alternatively or additionally, the scaling module can be configured to determine the scale factor for the scale-free 3D structure at least for the second image data based on both the first image data and the second image data. In other words, the determination of the scale-free 3D structure and / or the determination of the scale factor can be based on a temporal progression, i.e., a temporal change, of the image data. For this purpose, the image data sets can optionally also comprise third or even further image data at correspondingly even further points in time.
[0036] According to some preferred embodiments, variants, or refinements of embodiments, the device comprises a user interface by means of which a user can define points in an image, in particular in a real-time image displayed by a display device, as measurement points. The points can in particular be defined dynamically, i.e. in such a way that they move along with moving objects on which they were defined. Thus, a user can, for example, define the tips of known objects, markings on objects, and / or freely selectable points on an object (such as, for example, on a medical instrument) as measurement points. As a further example, a user can also define a point on a tissue as a measurement point, e.g., a point that an instrument should not or must not touch, while the tip of the corresponding instrument is defined as a further measurement point.The display device can thus be configured to display to the user - preferably in real time - the distance maintained, always based on the scale factor displayed by the output signal.
[0037] Such a user interface can be implemented, for example, as a graphical user interface by a display device of the device, in particular by a touchscreen.
[0038] According to a second aspect, the present invention provides a computer-implemented method for training an artificial intelligence entity (AI) to determine a scale factor for a scaleless 3D structure in image data from a medical imaging device. The method comprises at least the steps:
[0039] Providing a plurality of image data sets, each image data set comprising first image data of a medical scene at a first time and at least second image data of the medical scene at a second time;
[0040] Providing a randomly initialized or pre-trained artificial intelligence entity, KIE, which is configured to receive the first image data and the second image data as input, to determine a first scale factor for the first image data, to determine a second scale factor for the second image data, to determine a first pose for the first image data, to determine a second pose for the second image data, and to calculate an estimated first scale factor for the first image data based on the first pose, the second pose, and the second scale factor; and training the artificial intelligence entity, KIE, wherein parameters of the artificial intelligence entity, KIE, are iteratively changed to minimize a loss function which, among other things, penalizes differences between the determined first scale factor and the estimated first scale factor.
[0041] In this context, a pose refers to the orientation of an imaging device relative to the medical scene to be captured. The pose is typically specified with an orientation matrix (which specifies rotation angles relative to an orthonormal coordinate system) and a translation vector. In other words, a pose recognition module, for example, determines how (or whether) the imaging device moved between the acquisition of the first image data and the acquisition of the second image data.
[0042] According to a third aspect, the present invention provides a computer-implemented method for processing image data from a medical imaging device. The method comprises at least the steps:
[0043] Providing image data acquired by a medical imaging facility;
[0044] Determining a scale-free 3D structure in the acquired image data; and
[0045] Determining a scale factor for the scaleless 3D structure, based on which an actual size of a component of the scaleless 3D structure can be determined.
[0046] According to some preferred embodiments, variants, or refinements of embodiments, determining the scale factor comprises at least the step of detecting an object with at least one known dimension in the obtained image data. Determining the scale factor can be based on the position and / or orientation of the detected object and its known dimension, or at least based on the position and / or orientation of the known dimension of the detected object. The at least one known dimension can also be read from a database after the object has been detected, for example, using an artificial intelligence entity (AI).
[0047] According to some alternatives, the size of the detected object can also be determined directly by the artificial intelligence entity (AI), for example, based on a classification of the instrument. For this purpose, the instruments to be used can be assigned codes, e.g., color coding, where each color indicates a specific shaft diameter of the instrument's shaft.
[0048] According to some preferred embodiments, variants or refinements of embodiments, the method further comprises the steps:
[0049] Detecting two measurement points within the captured image data;
[0050] Determining, based on the scale factor, the actual distance between the detected measuring points; and
[0051] Display, by means of a display device, the determined actual distance.
[0052] The two measuring points can be located on one and the same object within the captured image data, with the two measuring points being rigidly arranged relative to each other, so that the object can be used, for example, as a kind of ruler. The two measuring points can also be located on the same object, but on parts of the object that are movable relative to each other, for example, on two wings of a pair of scissors, on a fixed part of a jaw that is movable relative to the fixed part, and / or the like. Thus, the object can be used in a variety of ways to measure distances.
[0053] Finally, measurement points can also be defined on different objects, which are thus also movable relative to each other. It is also possible to define more than two measurement points, and the distances between all detected defined measurement points are determined and displayed accordingly.
[0054] In the cases mentioned, the method can also be referred to as a method for performing a length measurement; it can in particular be carried out in real time, so that a user can move the objects and arrange the measuring points so that they limit the length to be measured on both sides. The measuring points to be detected can be predetermined, for example the tips of known objects (e.g. the tips of medical instruments), or specially marked points (i.e., markings known to the device or recognizable by the device) on known or unknown objects, in particular instruments. Particularly preferably, in one method step, a user can define points in an image, in particular in a real-time image displayed by a display device, as measuring points, for example by means of a user interface implemented by a display device. The points can in particular be defined dynamically, i.e.such that they move with the moving objects on which they were defined. Thus, a user can define, for example, the tips of known objects, markers on objects, and / or freely selectable points on an object (such as a medical instrument) as measurement points. As another example, a user can also define a point on a tissue as a measurement point, e.g., a point that an instrument must not touch, while defining the tip of the corresponding instrument as another measurement point.
[0055] This is particularly advantageous when the method is performed in real time, which is preferred. In this case, a user can use the objects shown in the image data, for example, instruments or objects guided by the user, to measure length. The options and variants mentioned are of course also applicable to the aforementioned user interface of the device according to the invention, and vice versa.
[0056] According to a fourth aspect, the invention provides a computer program product comprising executable program code which, when executed by a computing device, is configured to perform the method according to an embodiment of the second and / or the third aspect of the present invention.
[0057] According to a fifth aspect, the invention provides a non-transitory, computer-readable data storage medium comprising executable program code which, when executed by a computing device, is adapted to perform the method according to an embodiment of the second and / or third aspect of the present invention.
[0058] The non-volatile, computer-readable data storage medium may comprise or consist of any type of computer memory, in particular semiconductor memory, such as solid-state memory. The data carrier may also comprise or consist of a CD, a DVD, a Blu-ray disc, a USB memory stick, or the like.
[0059] According to a sixth aspect, the invention provides a data stream comprising executable program code or configured to generate executable program code which, when executed by a computing device, is configured to perform the method according to an embodiment of the second and / or third aspect of the present invention.
[0060] Further advantageous variants, options, embodiments, and modifications will become apparent from the following figures, the detailed description, and the claims. It should be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given for illustrative purposes only, since various changes and modifications within the scope of the invention will become apparent to those skilled in the art.
[0061] Individual embodiments of the present disclosure will be explained in detail with reference to the following figures. The components in the drawings are not necessarily to scale, but serve to illustrate the principles of the present invention. Parts in the various figures that correspond to the same elements or method steps have been provided with the same reference numerals in the figures. The numbering of method steps initially serves only to distinguish them and does not necessarily imply a corresponding order; however, it is a variant to perform the steps in the order of their numbering. Multiple steps can also be performed overlappingly or simultaneously. The figures show:
[0062] Fig. 1 is a schematic block diagram for explaining an apparatus according to an embodiment of the present invention;
[0063] Fig. 2 is a schematic representation of image data of a medical scene;
[0064] Fig. 3 shows an exemplary representation of image data of a medical scene with two instruments;
[0065] Fig. 4 is a schematic diagram illustrating the artificial intelligence entity during its training according to another embodiment of the present invention;
[0066] Fig. 5 is a schematic flow chart for explaining another embodiment of the present invention;
[0067] Fig. 6 is a schematic flow chart for explaining yet another embodiment of the present invention;
[0068] Fig. 7 is a schematic block diagram of a computer program product according to another embodiment of the present invention; and
[0069] Fig. 8 is a schematic block diagram of a non-transitory computer-readable data storage medium according to another embodiment of the present invention.
[0070] Fig. 1 shows a schematic block diagram for explaining a device 100 according to a first embodiment of the present invention, ie a device 100 for processing image data of a medical imaging device 50.
[0071] The imaging device 50 can be or include, for example, an endoscope or a microscope. It is configured to capture image data 71 from a medical scene 1. Typically, such image data 71 is displayed on a display device 200, often in real time. Overlays, evaluations, visualizations of radiation invisible to the human eye, etc., can be added to the visual display. This is often done by a camera control unit of the endoscope or microscope. The endoscope can be rigid, flexible, or semi-rigid.
[0072] The device 100 according to the present invention may, in some variants, also comprise the imaging device 50 and / or the display device 200, and may comprise a camera control unit or be fully or partially integrated into a camera control unit.
[0073] The device 100 comprises an input interface 110 which is configured to receive image data 71 acquired by the medical imaging device 50, either directly, via a PACS, or from another type of database.
[0074] The device 100 also includes a computing device 150 configured to perform various functions, which are explained below using various modules. As already explained above, these modules do not necessarily have to be implemented as clearly distinguishable units.
[0075] First, the computing device 150 provides a structure recognition module 152 configured to determine a 3D structure in the acquired image data 71, preferably to convert the entire acquired image data 71 into such a 3D structure, for example, in the form of a depth map. For this purpose, the structure recognition module 152 may comprise an artificial intelligence entity (AI), such as that described in Freedman et al.
[0076] The computing device further provides a scaling module 154, which is configured to determine a scale factor for the 3D structure, based on which an actual size of a component (or, preferably, all components) of the 3D structure can be determined. For example, a depth map can be provided with a scale factor that relates distances within the depth map to the actual distances in the physical basis for the 3D structure. The display device 200 can then, for example, automatically or at the request of a user, display dimensions of components of interest in the 3D structure. Alternatively or additionally, a scale can also be displayed, which always shows the actual size of the displayed 3D structure.
[0077] Finally, an output interface 190 of the device 100 is configured to output an output signal 79 that includes or indicates the determined scale factor. The output signal 79 can be a signal to the display device 200 to display a virtual representation of the 3D structure based on the scale factor, a scale display based on the scale factor, or the like.
[0078] As already briefly explained above, there are numerous different variants for determining the scale factor and for what purposes it can be used. The following describes a specific example of a variant with which, for example, the size of stones (e.g., kidney stones, gallstones, bladder stones) can be measured. This makes it possible to estimate whether stone fragments can be removed through a working channel. For illustration, reference is also made to Fig. 2 in the following description.
[0079] Fig. 2 shows a schematic representation of image data 71 of a medical scene 1. Visible are organic tissue and an instrument, namely a laser fiber 2 (or laser fiber, from the English "laser fiber"). Using such laser fibers 2, laser light, e.g., from a holmium laser or thulium laser, can be transported to the target location, for example, to fragment stones (lithotripsy). Such laser fibers 2 have a constant diameter, which thus represents a known dimension 4 for a known laser fiber 2.
[0080] To enable the user to guide the laser fiber 2 to the correct location and trigger the laser, image data 71 is acquired in real time by an imaging device 50, which at least partially (ideally always) contains, i.e., depicts, the tip of the laser fiber 2. The acquired image data 71 can, as shown in Fig. 2, be displayed in real time by a display device 200. Fig. 2 also clearly illustrates how the tissue structure itself hardly allows any conclusions to be drawn about the scale factor of the medical scene 1 shown.
[0081] The image data 71 acquired in real time can thus advantageously comprise an image dataset, namely a series of temporally consecutive images, for example, with a refresh rate between 10 and 100 Hertz, i.e., with images acquired sequentially between 10 and 100 milliseconds apart. These can be easily processed by the human eye and simultaneously provide a good data basis for subsequent processing.
[0082] In the described variant, the 3D structure of the image data 71 is first determined (or: recognized, or: detected) by the structure recognition module 152. This advantageously includes both the tissue sections contained in the image data 71 and any objects visible therein, such as the laser fiber 2.
[0083] The structure recognition module 152 can use various algorithms to recognize the 3D structure, for example, classic image processing algorithms such as SfM (Structure from Motion), for example using SLAM or vSLAM algorithms, where SLAM stands for simultaneous localization and mapping and vSLAM for visual simultaneous localization and mapping. SfM refers to the detection of the 3D structure of the medical scene 1 from an image dataset of consecutive 2D images. With vSLAM, the position and orientation of the imaging device 50 (specifically its input optics) with respect to its surroundings are calculated while simultaneously mapping this surroundings.Alternatively, or additionally, the structure recognition module 152 may also comprise a first artificial intelligence entity, KIE, which is trained and configured to receive the obtained image data 71 as input data and, based thereon, to generate a first output indicating the 3D structure.
[0084] In the present embodiment, the computing device further comprises an object recognition module 156, which also recognizes an object with at least one known dimension 4 in the image data 71. In the example of Fig. 2, this is the known laser fiber 2 with its known diameter. In fact, however, not just "one" diameter is known for the laser fiber 2, but rather a plurality of diameters at a plurality of locations along the laser fiber 2, since the diameter is constant across the elongated body of the laser fiber 2. With this additional information, the position of the laser fiber 2 relative to the remaining 3D structure can optionally be determined particularly precisely from the image shown in Fig. 2.
[0085] For example, the object recognition module 156 may comprise a second artificial intelligence entity, KIE, which is trained and configured to receive the obtained image data 71 as input data and, based thereon, to generate a second output which indicates (or: indexes, or: indicates) at least one object recognized in the image data 71 (here: the laser fiber 2) with at least one known dimension 4, preferably together with additional information such as its exact position (i.e., the object is segmented in the image data 71) and / or orientation, or at least the position and / or orientation of the known dimension 4 in the image data 71.
[0086] This object recognition, specifically segmentation, can be performed on the obtained (typically 2D) image data 71, whereupon the thus determined boundaries of the object, or at least the known dimension 4, can be transferred into the determined 3D structure. The object recognition, or segmentation, can be performed using a typical state-of-the-art artificial neural network (ANN), for example, a ResNetöO trained with segmented object images, in particular images of medical instruments.
[0087] The scaling module 154 can now relate the detected 3D structure and the at least one known dimension 4 to one another and thus determine the scale factor. The scale factor can be adjusted so that the diameter of the laser fiber 2 in the determined 3D structure corresponds as closely as possible to the known dimension 4. The scaling module 154 can also be implemented entirely or partially by the second artificial intelligence entity, KIE, in that it is configured to output an output that also displays the scale factor. The output can thus be the scale factor itself, or a control signal for a display device 200 to display a size or scale information, or a scaled 3D structure (e.g., a scaled depth map), or the like.
[0088] Similar to the laser fiber 2, other instruments can also be used which have elongated sections or shafts, the diameter of which can be known to the computing device 150, in particular the object recognition module 156. The device 100 can also comprise an input interface by means of which a user can enter a dimension. If the object recognition module 156 detects, for example, a suitable tool with an elongated shaft, a user can be prompted to enter the diameter of the shaft via the user interface. If the user complies, the diameter is henceforth treated as a known dimension 4. The user interface can be implemented, for example, using the display device 200, wherein user input can be made via conventional means (keyboard, mouse, trackpad, touchscreen, voice recognition, and the like).
[0089] Fig. 3 shows an exemplary representation of image data 71 of a medical scene 1 with two instruments 3, whose shaft diameters each have known dimensions 4. Just as with the laser fiber 2, the respective diameter is a known dimension 4 at a plurality of points along the shaft. For example, the shaft diameter may be 5 mm and mapped in the image (here an endoscopic image) to a width of 100 pixels. Using the scaling module 154, the previously determined depth map can now be scaled such that the corresponding points in the 3D structure at the same depth as the shaft diameter are spaced 5 mm apart.
[0090] Fig. 3 also shows, by way of example, that a distance d can be calculated between two detected objects 3 and displayed by the display device 200, here in the top left of the image. For this purpose, the display device 200 can display additional auxiliary lines based on the detected (or segmented) objects 3, as indicated in Fig. 3. Small "+" symbols indicate fictitious instrument tips between which the distance is calculated. The user can thus guide these to the tissue or object of interest and thereby perform a length measurement. Further small "+" symbols indicate the points at which the radii, and thus the diameters, of the shafts of the instruments 3 are calculated.
[0091] Computing device 150 can also include an artificial intelligence entity (AII), which is configured to determine the 3D structure, including the scale factor. This artificial intelligence entity (AII), can thus function both as a structure recognition module 152 and as a scaling module 154, and possibly even as an object recognition module 156 (e.g., in an intermediate stage).
[0092] In the following, a method for training such an artificial intelligence entity, KIE, is described with reference to Fig. 4 and Fig. 5.
[0093] Fig. 4 shows a schematic diagram illustrating the artificial intelligence entity, KIE 80, during its training. Fig. 5 shows a schematic flowchart according to an aspect of the third aspect of the present invention, i.e., a method for training an artificial intelligence entity, KIE, to determine a scale factor for a 3D structure in image data 71 of a medical imaging device 50.
[0094] Referring to Fig. 4 and Fig. 5, in a step S10, a plurality of image data sets are provided (as training data), wherein each image data set comprises first image data 71 (or: a first frame) of a medical scene 1 at a first time t=0 and at least second image data 72 (or: a second frame) of the medical scene 1 at a second time t'= 1.
[0095] The medical scenery 1 can advantageously vary between the individual image datasets, or at least between some of them, so that the artificial intelligence entity, KIE 80, is not excessively trained on a specific medical scene 1. In contrast to the known methods in the prior art, an implementation of the artificial intelligence entity, KIE 80, described here, is applicable to all types of medical scenes 1, regardless of their size scale, since the determination of the associated size scale (i.e., the scale factor) is included. By having a large number of different medical sceneries 1 in the training data, i.e., by having as many image datasets of different medical sceneries 1 as possible, the KIE is trained particularly robustly.In this context, different medical scenes 1 are understood to mean in particular those which show different organs, for example a blood vessel, an abdominal cavity, an intestine, a heart, a kidney, a gall bladder, a bladder and the like.
[0096] In a step S20, a randomly initialized or pre-trained artificial intelligence entity, KIE 80, is provided, with an architecture or structures that are explained in more detail below:
[0097] The KIE 80 is designed such that the provided first image data 71, optionally also the second image data 72, are input into a structure recognition module 152 of the artificial intelligence entity, KIE 80, which is configured to determine a 3D structure, for example a depth map 91, of the image data 71, 72 input into it.
[0098] The KIE 80 is further configured such that the first image data 71 and the second image data 72 are input into a pose recognition module 81 of the KIE 80. The pose recognition module 81 is configured to recognize a respective pose 93, 94 of the image data 71, 72, at least in relation to one another, based on the image data 71, 72 input into it. In this context, a pose is understood to mean the orientation of an imaging device 50 relative to the medical scene 1 to be captured. The pose is typically specified with an orientation matrix (which specifies rotation angles with respect to an orthonormal coordinate system) and a translation vector. In other words, the pose recognition module 81 determines how (or whether) the imaging device 50 has moved between the capture of the first image data 71 and the second image data 72.The difference between the determined first pose 93 of the first image data 71 and the determined second pose 94 of the second image data 72 can be used to relate a depth map 91 for the first image data 71 to a second depth map 92 for the second image data 72, or to convert one into the other, or to estimate one using the other.
[0099] For example, the structure recognition module 152 can determine the first depth map 91 for the first image data 71 based on the first image data 71, and can determine the second depth map 92 for the second image data 72 based on the second image data 72. The depth map 92 determined for one of the image data (here, without loss of generality, the second image data 72) is then converted (or adjusted, or evolved) into an estimated first depth map for the first image data 71 using the two determined poses 93, 94. Thus, the estimated first depth map ideally corresponds to the determined first depth map 91.
[0100] A loss function 99 of the training procedure contains a term that penalizes deviations between the determined first depth map 91 and the estimated first depth map. As is common in machine learning, parameters (e.g., node weights, thresholds for activation functions, etc.) are iteratively adjusted during training to minimize the overall loss function.
[0101] The term mentioned thus contributes to training the artificial intelligence entity, KIE 80, in such a way that the structure recognition module 152 becomes increasingly better at generating a depth map 91, 92 in such a way that it is as consistent as possible with an estimated depth map which is generated based on a pose difference to preceding or subsequent image data 71, 72.
[0102] According to the present invention, the artificial intelligence entity, KIE 80, additionally includes a scaling module 154 configured to determine a first scale factor 95 for the first image data 71 and a second scale factor 96 for the second image data 72. As already mentioned above, such a scaling module 154 can be implemented together with other modules or separately. For example, it can be implemented together with an object recognition module 156 and determine the scale factor based on any objects 2, 3 visible in the image data 71, 72 with at least one known dimension 4.
[0103] Accordingly, the loss function 99 can have a further term, or an existing term can be modified such that deviations between the scale factors 95, 96 for the first image data 71 and the second image data 72, respectively, are penalized and thus minimized during training. Similar to what was described with regard to the depth maps 91, 92, an estimated scale factor for the first image data 71 can also be determined from the second scale factor 96 for the second image data 72 using the determined poses 93, 94, which in turn can be compared with the first scale factor 95 actually determined for the first image data 71, wherein the difference is positively incorporated into the loss function 99, i.e., a higher difference, ceteris paribus, leads to a higher value of the loss function 99 to be minimized.
[0104] In step S30, the artificial intelligence entity, KIE 80, is iteratively trained to minimize the loss function 99. This can be done, for example, using unsupervised learning, such as that described in Freedman et al. All mechanisms and methods known in the art can be used for training, such as gradient descent, training over epochs, and the like.
[0105] In this way, the artificial intelligence entity, KIE 80, is trained to obtain temporal sequences of (at least or exactly 2) image data 71, 72 (or: frames) and to find a solution for the depth maps 91, 92, the poses 93, 94, and the scale factors 95, 96 that is consistent across the image data 71, 72. In this way, the artificial intelligence entity, KIE 80, is also able to determine a scale factor 95, 96 for current image data 71, 72, even if it currently does not contain an object 2, 3 with at least one known dimension 4.
[0106] The artificial intelligence entity, KIE 80, trained to a desired level of accuracy can then be implemented, for example, by the computing device 150 of the device 100.
[0107] For its deployment (in production mode, deployment stage), the artificial intelligence entity, KIE 80, can be deployed as it was trained. Alternatively, the KIE 80 can also be deployed with only the structure recognition module 152 and the scaling module 154. The pose recognition module 81, however, can be removed or deactivated, depending on the application, or its output can be ignored. Depending on the application and training, the KIE 80 can also be configured in production mode to receive individual images as input, or an image dataset (i.e., an image sequence). It is also conceivable that the structure recognition module 152 receives and uses individual images as input, while the scaling module 154 receives image datasets (i.e., image sequences), or vice versa.
[0108] Although the description herein has always focused on first image data 71 and second image data 72, it is understood that the image data sets can also additionally comprise further image data, for example, third image data. In this case, the described mechanisms can be extended to a triple, wherein, for example, for the second image data 72 located in the middle in time, estimates are generated both based on the temporally preceding first image data 71 and based on the temporally succeeding third image data. The loss function 99 can thus correspondingly comprise deviations of the second depth map 92 determined for the second image data 72 and / or the second scale factor 96 both from the corresponding estimates based on the quantities determined for the first image data 71 and from the corresponding estimates based on the quantities determined for the third image data. Fig.Fig. 6 shows a schematic flow diagram for explaining a method according to an embodiment of the third aspect of the present invention, ie, a computer-implemented method for processing image data 71, 72 of a medical imaging device 50. The method according to Fig. 6 can be carried out in particular by means of the device 100 from Fig. 1 and can therefore be adapted according to all options or variants described with regard to the device 100 according to the invention and vice versa.
[0109] In a step S100, image data 71, 72 acquired by a medical imaging device 50 are provided. This provision S10 may include acquiring the image data 71, 72 by the medical imaging device 50. Further options have already been explained above by way of example with reference to the input interface 110.
[0110] In a step S200, a 3D structure is determined in the acquired image data 71, 72, for example, a depth map 91, 92. The determination S200 of the 3D structure can be carried out in particular as described above with reference to the structure recognition module 152, for example, using an artificial intelligence entity, KIE 80.
[0111] In an optional step S300, an object 2, 3 with at least one known dimension 4 is recognized in the acquired image data 71, for example as described above with reference to the object recognition module 156.
[0112] In a step S400, a scale factor 95, 96 is determined for the 3D structure, based on which an actual size of a component (or each component) of the 3D structure can be determined, for example as described above with reference to the scaling module 154 and / or the object recognition module 156 and / or the artificial intelligence entity, KIE 80.
[0113] Fig. 7 shows a schematic block diagram of a computer program product 300 according to an embodiment of the third aspect of the present invention. The computer program product 300 comprises executable program code 350 which, when executed (e.g., by a computing device), is configured to perform the method according to an embodiment of the present invention, for example, according to Fig. 5 or Fig. 6.
[0114] Fig. 8 shows a schematic block diagram of a non-volatile computer-readable data storage medium 400 according to an embodiment of the present invention. The data storage medium 400 comprises executable program code 450 which, when executed (e.g., by a computing device), is configured to perform the method according to an embodiment of the present invention, for example, according to Fig. 5 or Fig. 6. The non-volatile computer-readable data storage medium 400 may, for example, be designed as or comprise a semiconductor memory, e.g., an SSD memory chip. The data storage medium 400 may also comprise or comprise a CD, DVD, Blu-ray, or a magnetic storage device.
[0115] The above description of the disclosed embodiments merely contains examples of possible implementations described to enable a person skilled in the art to make or use the present invention. Various variations and modifications of these embodiments will be readily apparent to those skilled in the art, given knowledge of the present invention, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure.
[0116] Thus, the present invention is not intended to be limited to the specific embodiments shown herein, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0117] List of reference symbols
[0118] 1 Medical scenery
[0119] 2 laser fibers
[0120] 3 instruments
[0121] 4 Known dimensions
[0122] 5 Tip of the laser fiber
[0123] 50 imaging device
[0124] 71 image data
[0125] 79 Output signal
[0126] 80 Artificial Intelligence Entity
[0127] 81 Pose detection module
[0128] 91 Depth map
[0129] 92 Depth map
[0130] 93 Pose
[0131] 94 Pose
[0132] 95 scale factor
[0133] 96 scale factor
[0134] 99 Loss function
[0135] 100 device
[0136] 110 Input interface
[0137] 150 computing device
[0138] 152 Structure recognition module
[0139] 154 Scaling module
[0140] 156 Object recognition module
[0141] 190 Output interface
[0142] 200 display device
[0143] 300 computer program product
[0144] 350 program code
[0145] 400 data storage medium
[0146] 450 program code
[0147] S10..S30; S100..S400
[0148] Procedural steps
Claims
Patent claims 1. A device (100) for processing image data (71, 72) of a medical imaging device (50), comprising: an input interface (110) configured to receive image data (71, 72) acquired by a medical imaging device (50); a computing device (150) configured to implement at least one structure recognition module (152) and one scaling module (154); wherein the structure recognition module (152) is configured to determine a scale-free 3D structure (91, 92) in the acquired image data (71, 72); wherein the scaling module (154) is configured to determine a scale factor (95, 96) for the scaleless 3D structure (91, 92), on the basis of which an actual size of at least one component of the scaleless 3D structure (91, 92) can be determined;and an output interface (190) configured to generate an output signal (79) comprising or indicating the determined scale factor (95, 96); 2. Device (100) according to claim 1, wherein the computing device (150) is additionally configured to implement an object recognition module (156) which is configured to recognize at least one object (2, 3) with at least one known dimension (4) in the obtained image data (71, 72); and wherein the scaling module (154) is configured to determine the scale factor (95, 96) based on the position and / or orientation of the recognized object (2, 3) and its known dimension (4), or at least based on the position and / or orientation of the at least one known dimension (4) of the recognized object (2, 3).
3. Device (100) according to claim 2, wherein the object (2, 3) with the at least one known dimension (4) comprises at least one medical instrument (2, 3), in particular a laser fiber (3), and / or at least one component of a medical imaging device.
4. Device (100) according to one of claims 1 to 3, wherein the structure recognition module (152) and the scaling module (154) have a common artificial intelligence entity (80) which is configured to receive the obtained image data (71, 72) as input data and, based thereon, to generate an output which indicates the scale-free 3D structure (91, 92) and the scale factor (95, 96).
5. The device (100) according to any one of claims 1 to 3, wherein the structure recognition module (152) comprises a first artificial intelligence entity, KIE, which is configured to receive the obtained image data (71, 72) as input data and, based thereon, to generate a first output indicating the scaleless 3D structure (91, 92); and wherein the computing device (150) implements a second artificial intelligence entity, KIE, which is configured to receive at least the first output of the first artificial intelligence entity or data based thereon as input and to generate a second output based thereon.
6. The device (100) according to claim 2 or 3 in conjunction with claim 5, wherein the object recognition module (156) comprises the second artificial intelligence entity, KIE, and this second KIE is configured to recognize the at least one object (2, 3) with the at least one known dimension (4) in the obtained image data (71, 72), wherein the second output indicates the recognized object (2, 3), and wherein the scaling module (154) is configured to obtain the at least one known dimension (4) of the recognized object (2, 3) using a database and to determine the scale factor (95, 96) based thereon.
7. Device (100) according to one of claims 1 to 6, wherein the computing device (150) is further configured to generate a scaled depth map of at least the determined scaleless 3D structure (91, 92) based on the determined scale factor (95, 96) and to store or output this.
8. Device (100) according to one of claims 1 to 8, wherein the structure recognition module (152) is configured to determine the scale-free 3D structure (91, 92) substantially in real time and / or wherein the scaling module (154) is configured to determine the scale factor (95, 96) substantially in real time.
9. Device (100) according to one of claims 1 to 8, wherein the device (100) comprises the medical imaging device (50), in particular an endoscope or a microscope.
10. The device (100) according to claim 9, wherein the medical imaging device (50) comprises a camera control unit, and at least the computing device (150) is integrated into the camera control unit.
11. The device (100) according to one of claims 1 to 10, wherein the acquired image data (71) comprise an image data set comprising first image data (71) of a medical scene (1) at a first point in time and second image data (72) of the medical scene (1) at a second point in time; and wherein the structure recognition module (152) is further configured to determine the scale-free 3D structure (91, 92) at least for the second image data (72) based on both the first image data (71) and the second image data (72) and / or the scaling module (154) is configured to determine the scale factor (95, 96) for the scale-free 3D structure (91, 92) at least for the second to determine image data (72) based on both the first image data (71) and the second image data (72).
12. A computer-implemented method for training an artificial intelligence entity, KIE (80), for determining a scale factor (95, 96) for a scaleless 3D structure (91, 92) in image data (71, 72) of a medical imaging device (50), comprising: Providing (S10) a plurality of image data sets, each image data set comprising first image data (71) of a medical scene (1) at a first time and second image data (72) of the medical scene at a second time (1); Providing a randomly initialized or pre-trained artificial intelligence entity, KIE (80), which is designed to receive the first image data (71) and the second image data (72) as input, to determine a first scale factor (95) for the first image data (71), to determine a second scale factor (96) for the second image data (72), to determine a first pose (93) for the first image data (71), to determine a second pose (94) for the second image data (72), and to calculate an estimated first scale factor for the first image data (71) based on the first pose (93), the second pose (94), and the second scale factor (96); and Training the artificial intelligence entity, KIE (80), whereby parameters of the KIE (80) are iteratively changed to minimize a loss function (99) which, among other things, penalizes differences between the determined first scale factor (95) and the estimated first scale factor.
13. Computer-implemented method for processing image data (71, 72) of a medical Imaging device (50) comprising: Providing (S100) image data (71, 72) acquired by a medical imaging device (50); Determining (S200) a scale-free 3D structure (91, 92) in the acquired image data (71, 72); and determining (S300) a scale factor (95, 96) for the scale-free 3D structure (91, 92), based on which an actual size of a component of the scale-free 3D structure (91, 92) can be determined.
14. A computer program product (300) comprising executable program code (350) which, when executed, is configured to perform the method according to claim 12 or 13.
15. A non-transitory, computer-readable data storage medium (400) comprising executable program code (450) which, when executed, is configured to perform the method according to claim 12 or 13.
Citation Information
Patent Citations
Automated monitoring of medical imaging procedures
WO2019145951A1