Apparatus and method for processing image data of a medical imaging device and method for training an artificial intelligence entity
The apparatus and method for processing medical imaging data address the challenge of determining actual sizes of scaleless 3D structures by using a scaling module to calculate a scale factor, achieving accurate and scalable size determination for diverse medical scenes.
Patent Information
- Application Number
- DE102023136115
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Existing medical imaging technologies struggle to determine the actual size of scaleless 3D structures in image data, lacking a reliable scale factor to relate the structures to real-world dimensions.
An apparatus and method that process image data from medical imaging devices by using a computing device with a structure recognition module to identify scaleless 3D structures and a scaling module to determine a scale factor, allowing for the calculation of actual sizes of the structures.
Enables accurate determination of the actual size of 3D structures in medical imaging, applicable to various medical scenes regardless of scale, improving size ratio determination and facilitating precise measurements in real-time applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to an apparatus and a method for processing image data of a medical imaging device, in particular in order to improve the determination of dimensions or size ratios in 3D structures which are captured in the image data. The invention also relates to a method for training an artificial intelligence entity for determining a scale factor for a scaleless 3D structure in image data of a medical imaging device.Imaging methods are frequently used in modern medicine. A medical imaging device is used to acquire images of a medical scene. These images may provide a user, for example an operator, with information that the user cannot visually grasp without assistance, whether because the images originate from a location that is not directly visible (e.g., from the interior of a patient), or whether because the information is based on electromagnetic radiation having invisible wavelengths.The use of such imaging devices as auxiliary means thus entails many advantages, but also a certain mediumability. For example, it becomes more difficult for the user to develop an intuitive understanding of the acquired medical scene. This applies in particular to the size relationships. By closer approach or removal of an input optics of the imaging device, the structures of the medical scene appear larger or smaller.Methods are known for determining relative dimensions or sizes of structures or sections of a medical scene with respect to one another, for example in the context of generating a depth map. In the prior art, however, there is a lack of a scale factor which links the structure--shown in a structurally consistent manner--to variables in reality.For example, the scientific publication "Detecting Deficient Coverage in Colonoscopies" by D. Fredman et al, arXiv:2001.08589v 3, dated March 29, 2020 (hereafter cited as "Fredman et al.") describes a network structure that is trained with unsupervised learning ("unsupported learning") to generate a depth map based on an ordinary RGB image of intestinal structures. This is used to determine in a series of time-sequential RGB images whether there are areas that were not visible on the RGB images. If this is determined, the corresponding regions can be deliberately examined again. Freedman et al. state that the generated depth maps can only be determined to a scale factor that is random, but which is effectively compensated by the algorithm of the described network structure. Moreover, for the purpose of Freedman et al., it is also noticeably unimportant to know the true scale factor, since it is important to determine whether there are areas that were not seen, regardless of how large they are in reality. It should also be noted that in the single application of Freedman et al., namely intestinal mirroring (colonoscopy), both training and application image data all have structures with very similar size ratios and thus the random scale factor will be in a narrow range of values anyway.It is therefore an object of the present invention to provide an apparatus and a method which process image data such that the actual size of scaleless 3D structures detected in the image data can be determined. It is a further object to train an artificial intelligence entity suitable for use in such processing.This object is achieved by the subject matters of the independent claims of the present invention.According to a first aspect, an apparatus for processing image data of a medical imaging device is accordingly provided. This comprises at least:an input interface configured to receive image data acquired from a medical imaging device;a computing device configured to implement at least one structure recognition module and a scaling module;wherein the structure recognition module is configured to determine a scaleless 3D structure in the captured image data;wherein the scaling module is configured to determine a scale factor for the scaleless 3D structure, on the basis of which an actual size of at least one component of the scaleless 3D structure (preferably of the entire 3D structure) can be determined; andan output interface configured to generate an output signal that includes or indicates (or: indicates) the determined scale factor.A basic idea of the present invention is that prior art methods, which can in some cases provide well-usable reconstructions of scaleless 3D structures (e.g. depth maps), can be supplemented with the determination of a scale factor, so that the actual size of 3D structures in the image data can be determined later on the basis of the scale factor.Scaleless 3D structures are to be understood in particular as 3D structures which indicate accurate positional relationships and distances of points of the 3D structure to one another, but do not relate these to an actual absolute value for the distances, in particular not to an actual distance which the points in the physical 3D structure from which the image data in which the scaleless 3D structure was determined have been captured. According to one aspect of the invention, a scaleless 3D structure is thus initially determined, i.e. a 3D structure is determined apart from a scale factor, and the associated scale factor is then determined, or simultaneously but separately.This device can be used advantageously for all types of medical scenes, in particular also those which have completely different scales or sizes. In contrast to the known prior art, the same device can therefore be used for both microscopic and macroscopic, endoscopic and external use.The image data acquired by the imaging device can be, in particular, an image or a plurality of images, in particular a temporal series of images, of a medical scene. The term "medical scene" is broadly used herein: it may refer to an exterior or even interior view of a patient currently undergoing or about to undergo a medical procedure. In other words, the medical scene can also be a scene in which an organic, in particular human, tissue can be seen, for example in a laboratory or an operating room, either in vitro and / or in vivo.Although some functions are described herein as being performed by "devices", "interfaces" or "modules", it is understood that this does not necessarily mean that such devices, interfaces, or modules are provided as separate units from one another. In cases where one or more devices, interfaces, or modules are provided in whole or in part as software, the devices, interfaces, or modules may be implemented by program code portions or program code switches that are different from each other, but may also be interwoven with each other.Similarly, in the case where one or more devices, interfaces, or modules are provided as hardware, the functions of one or more devices, interfaces, or modules may be provided from one and the same hardware component, or the functions of one device, interface, or module, or the functions of multiple devices, interfaces, or modules may be distributed among multiple hardware components that do not necessarily need to correspond to the devices, interfaces, or modules one-to-one. Thus, any apparatus, system, method, etc., having all features and functions attributed to a particular device, interface, and / or module, is to be understood to include, or implement, the device, interface, and / or module.In particular, it is possible for all modules to be implemented by program code which is executed by the computing device.The computing device can be realized as any device or any means for computing, in particular for executing software, an app or an algorithm. For example, the computing device may include at least one processor, such as at least one central processor, CPU and / or at least one graphics processor, GPU, and / or at least one field programmable gate array, FPGA, and / or at least one application specific integrated circuit, ASIC, and / or any combination of the foregoing. The computing device may further include a memory operatively connected to the at least one processor and / or a non-volatile memory operatively connected to the at least one processor and / or the memory. The computing device may be partially and / or fully implemented in a local device and / or partially and / or fully implemented in a remote system, such as a cloud computing platform.The input interface can be configured to receive the image data directly from the imaging device, in particular in real time, or alternatively also from a picture archiving and communications system (PACS), wherein the latter can also take place in real time, i.e. in particular as soon as the image data enter the PACS. The apparatus can also comprise an imaging device, so that obtaining the image data can also comprise capturing the image data by the imaging device. The apparatus can comprise a plurality of different imaging devices which can be designed, for example, for the acquisition of medical scenes of different sizes (microscopic, macroscopic) or different perspectives (endoscopic, open-surgical,... ). However, the processing can always be carried out using the same computing device.According to some preferred embodiments, variants or refinements of embodiments, the computing device is additionally configured to implement an object recognition module which is configured to recognize at least one object having at least one known dimension in the image data obtained. For this purpose, for example, a database of objects of known size can be formed in the device, together with the respective known dimension / s.For example, all dimensions of the at least one object (or all objects in the database) can be known, so that one can speak of a digital twin of the object or objects. Alternatively, only some, or only a single, dimension may also be known. Objects are particularly suitable for the database which can either be introduced into the medical scene additionally without great effort, from which the image data are captured, or else objects which are frequently visible in such medical scenes in any case. The latter include, for example, medical instruments such as scalpels, scissors, holding cutlery, laser fibers and / or the like, or parts or portions thereof, respectively, or markings thereon.It is precisely in the case of objects with a clearly recognizable elongated body, such as a laser fiber, for example, that it may already be sufficient to know as the single dimension about a diameter of this elongated body.In some variants, one of the objects with at least one known dimension can also be a component of a tissue to be examined, for example an implant or a tissue distance with known dimension. The information about the dimension can be taken from a database, for example, such as a medical database, an implant database or a patient database.The scaling module may be configured to determine the scale factor based on the position and / or orientation of the detected object and its at least one known dimension, or at least based on the position and / or orientation of the at least one known dimension of the detected object.The position of the object or of the dimension thereof can be a position of one or more points of the object or of the dimension of the object in three-dimensional space. The orientation of the object or of its dimension can be a position of the object or of its dimension in three-dimensional space, indicated by the positions of two or more points of the object or of the dimension which are in a known relationship to one another, for example indicated by a vector.The object recognition module can be configured to first determine the position and / or orientation of the object in the captured image data. The scaling module may be configured to relate the at least one known dimension of the object to the scaleless 3D structure based on the determined position and / or orientation and to determine the scale factor therefrom.According to some preferred embodiments, variations or refinements of embodiments, the structure recognition module and / or the scaling module include and use at least one artificial intelligence entity, KIE. The artificial intelligence entity may be, for example, an artificial neural network, ANN.According to some preferred embodiments, variants or refinements of embodiments, the structure recognition module and the scaling module comprise a common artificial intelligence entity, KIE, which is configured to obtain the obtained image data as input data and, based thereon, to generate an output which indicates (or: indicates) the scaleless 3D structure and the scale factor. In other words, in this artificial intelligence entity, KIE, 3D structure and scale factor are generated together. A method for training an artificial intelligence entity suitable for this is provided below as a further aspect of the present invention.According to some preferred embodiments, variants or refinements of embodiments, the structure recognition module comprises a first artificial intelligence entity, KIE, which is designed to obtain the obtained image data as input data and based thereon to generate a first output which indicates the scaleless 3D structure, for example as a point cloud, as a voxel structure, or as a vertex or polygon structure. The computing device can advantageously also be configured to implement a second artificial intelligence entity, KIE, which is configured to receive at least the first output of the first artificial intelligence entity, KIE, or data based thereon as input and to generate a second output based thereon. In other words, in this variant a pipeline is provided in which firstly the scaleless 3D structure is generated and based thereon a further processing by means of the second artificial intelligence entity, KIE, then takes place. For example, the first KIE can be configured for 3D modeling, and then a second KIE can be provided downstream for instrument segmentation. This can contain as input images either the image data obtained and / or the output of the 3D modeling in order to then calculate the scale factor from the width of, for example, an instrument shaft.According to some preferred embodiments, variants or refinements of embodiments, an object recognition module is implemented by the computing device, which comprises the second artificial intelligence entity, KIE. The second KIE is advantageously configured to recognize the at least one object having the at least one known dimension in the obtained image data.The second output, i.e. the output of the second artificial intelligence entity, KIE, can display (or: index or code) the recognized object, for example from a list of known objects (in a database). For this purpose, this second artificial intelligence entity, KIE, can comprise or be an artificial neural network, KNN, which in particular has a softmax layer as output layer. The scaling module may be configured to obtain the known size of the detected object using the database and determine the scale factor based thereon. For this purpose, the scaling module can comprise, for example, a third artificial intelligence entity, KIE. The first, second, and / or third artificial intelligence, KIE, entities may thus be integrated into an artificial intelligence pipeline.According to some preferred embodiments, variants or refinements of embodiments, the computing device is also configured to generate a scaled depth map of at least the determined scaleless 3D structure on the basis of the determined scale factor and to store or output said scaled depth map. Preferably, the determined scaleless 3D structure includes the entire captured image data, or in other words, a depth map of the entire captured image data is generated as the scaleless 3D structure. This depth map, provided with the determined scale factor, can then be stored or output as a scaled depth map. The output can be effected, for example, to a display device, for example a screen or a touchscreen. The storage can be effected either within the apparatus or, for example, also in a picture archiving and communications system (PACS).The generation of a depth map can be dispensed with if a stereo reconstruction is carried out. For this purpose, the imaging device can comprise a stereo camera or two separate objectives. For example, the imaging device can be a stereo endoscope. In other variants, however, a depth map can also be created in addition to a stereo reconstruction, and the method according to the invention can be carried out based thereon or the device according to the invention can process the latter. In these cases, dimensions within the image data can be determined on the one hand via the stereo reconstruction and on the other hand via the method according to the invention. A comparison of the results can then take place, wherein a warning can be output to a user, for example on a display device, if the results are not within a predeterminable tolerance range of one another. In addition, the method according to the invention can also be used in a stereo camera, for example if one of its lenses is unusable (e.g. soiled).According to some preferred embodiments, variants or refinements of embodiments, the structure recognition module is configured to determine the scaleless 3D structure substantially in real time. Advantageously, the structure recognition module is also configured to determine the scale factor in real time. Thus, advantageously, the currently captured image data can always be displayed (i.e. in real time) by a display device, together with the scale factor or at least one dimension based on the scale factor. For example, two instrument tips can be determined in the captured image data (for example by segmenting the image data by an artificial intelligence entity), and the actual distance between the two instrument tips can be displayed by the display device automatically or upon request from the user. In this way, the user can make length measurements using the instrument tips.For the described real-time application, it is advantageous if obtaining the image data comprises capturing the image data by means of the imaging device in real time. The imaging device can advantageously be an RGB camera (i.e. a camera operating on the basis of visible light), for example in an endoscope or microscope, since these generate image data that can be processed particularly quickly.According to some preferred embodiments, variants or refinements of embodiments, the apparatus comprises the medical imaging device, in particular an endoscope or a microscope. The input interface can thus be configured to receive the image data captured by the imaging device directly from the imaging device, for example in real time.The processing of the image data in real time has the advantage, on the one hand, that indications can be output to a user by means of a display device of the apparatus. For example, the user may be prompted to insert one of the objects having at least one known dimension into the imaging device's detection range. This can be done periodically or on a rule basis, for example if an error value determined for determining the scale factor exceeds a threshold value, or the depth map changes more strongly than a predetermined threshold value within a predetermined time.The processing of the image data in real time, on the other hand, has the advantage that a user of the imaging device, who frequently simultaneously also operates an instrument, can always be informed about the actual size ratios of the 3D structure by means of the display device.According to some preferred embodiments, variations or refinements of embodiments, the medical imaging device comprises a camera control unit. At least the computing device is advantageously integrated into the camera control unit. Namely, camera control units, in particular of endoscopes, already have means for image processing. Thus, the tasks of the computing device according to the present invention can be advantageously integrated into the camera control unit.According to some preferred embodiments, variants or refinements of embodiments, the captured image data comprise an image data set which comprises first image data of a medical scene at a first point in time and at least second image data of the (same) medical scene at a second point in time.The structure recognition module may also be configured to determine the scaleless 3D structure at least for the second image data based on both the first image data and the second image data. Alternatively or additionally, the scaling module may be configured to determine the scale factor for the scaleless 3D structure at least for the second image data based on both the first image data and the second image data. In other words, the determination of the scaleless 3D structure and / or the determination of the scale factor can take place on a temporal profile, i.e. a temporal change, of the image data.For this purpose, the image data sets can also optionally comprise third or even further image data at correspondingly even further times.According to some preferred embodiments, variants or refinements of embodiments, the apparatus has a user interface, by means of which a user can define points in an image, in particular in a real-time image displayed by a display device, as measurement points. The points can be defined dynamically, i.e. in such a way that they move with moving objects on which they were defined. Thus, for example, a user can define the tips of known objects, markings on objects, and / or freely selectable points on an object (such as on a medical instrument) as measurement points. As another example, a user may also define a point on a tissue as a measurement point, e.g., a point that an instrument should or may not touch, while the tip of the corresponding instrument is defined as another measurement point. The display device can thus be configured to display the respectively observed distance to the user-preferably in real time-always on the basis of the scale factor displayed by the output signal.Such a user interface can be implemented, for example, as a graphical user interface by a display device of the apparatus, in particular by a touchscreen.In a second aspect, the present invention provides a computer-implemented method for training an artificial intelligence, KIE, entity to determine a scale factor for a scaleless 3D structure in image data of a medical imaging device. The method comprises at least the steps of:providing a plurality of image data sets, each image data set comprising first image data of a medical scene at a first time and at least second image data of the medical scene at a second time;providing a randomly initialized or pre-trained artificial intelligence, KIE, entity configured to receive the first image data and the second image data as input, determine a first scale factor for the first image data, determine a second scale factor for the second image data, determine a first pose for the first image data, determine a second pose for the second image data, and calculate an estimated first scale factor for the first image data based on the first pose, the second pose, and the second scale factor; andtraining the artificial intelligence entity, KIE, iteratively changing parameters of the artificial intelligence entity, KIE, to minimize a loss function that penalizes, among other things, differences between the determined first scale factor and the estimated first scale factor.In this context, a pose is to be understood as meaning the alignment of an imaging device with respect to the medical scene to be captured. The pose is typically indicated with an orientation matrix (which indicates rotational angles with respect to an orthonormal coordinate system) and a translation vector. In other words, for example, a pose recognition module determines how (or whether) the imaging device has moved between the acquisition of the first image data and the second image data.According to a third aspect, the present invention provides a computer-implemented method for processing image data of a medical imaging device. The method comprises at least the steps of:providing image data acquired by a medical imaging device;determining a scaleless 3D structure in the captured image data; anddetermining a scale factor for the scaleless 3D structure, from which an actual size of a component of the scaleless 3D structure can be determined.According to some preferred embodiments, variants or refinements of embodiments, determining the scale factor comprises at least the step:detecting an object having at least one known dimension in the obtained image data. The scale factor can be determined based on the position and / or orientation of the detected object and its known dimension, or at least based on the position and / or orientation of the known dimension of the detected object. The at least one known dimension may also be read from a database after the object is detected, for example using an artificial intelligence entity, KIE.According to some alternatives, the size of the detected object may also be directly determined by the artificial intelligence entity, KIE, for example based on a classification of the instrument. For this purpose, instruments to be used can be provided with codes, for example color codes, wherein each color, for example, specifies a specific shaft diameter of a shaft of the instrument.According to some preferred embodiments, variants or refinements of embodiments, the method further comprises the steps of:detecting two measurement points within the captured image data;determining, based on the scale factor, the actual distance between the detected measurement points; anddisplaying, by means of a display device, the determined actual distance.The two measurement points can be arranged on one and the same object within the captured image data, wherein the two measurement points are arranged rigidly with respect to one another, such that the object can be used, for example, as a kind of ruler. The two measurement points can also be arranged on the same object, but on parts of the object that are movable relative to one another, for example on two wings of scissors, on a fixed part and a part of a jaw part that is movable relative thereto and / or the like. Thus, the object can be used in a variety of ways to measure distances.Finally, measurement points on different objects can also be defined, which are thus likewise movable relative to one another. It is also possible for more than two measurement points to be defined, and for distances between all defined measurement points detected to be determined and displayed accordingly.In the cases mentioned, the method can also be referred to as a method for carrying out a length measurement; it can be carried out in particular in real time, so that a user can move the objects and arrange the measurement points in such a way that they delimit the length to be measured on both sides. The measurement points to be identified can be predetermined, for example the tips of known objects (e.g. the tips of medical instruments), or specially marked points (i.e. markings known to the device or recognizable by the device) on known or unknown objects, in particular instruments.Particularly preferably, in a method step, a user can define points in an image, in particular in a real-time image displayed by a display device, as measurement points, for example by means of a user interface implemented by a display device. The points can be defined dynamically, i.e. in such a way that they move with moving objects on which they were defined. Thus, for example, a user can define the tips of known objects, markings on objects, and / or freely selectable points on an object (such as on a medical instrument) as measurement points. As another example, a user may also define a point on a tissue as a measurement point, e.g., a point that an instrument may not contact while the tip of the corresponding instrument is defined as another measurement point.This is particularly advantageous if the method is carried out in real time, which is preferred. In this case, a user can namely use the objects shown in the image data, for example instruments or objects guided by the user, for length measurement. Said options and variants are of course also applicable to the aforementioned user interface of the device according to the invention, and vice versa.According to a fourth aspect, the invention provides a computer program product comprising executable program code which, when executed by a computing device, is configured to perform methods according to an embodiment of the second and / or the third aspect of the present invention.According to a fifth aspect, the invention provides a non-transitory computer readable storage medium comprising executable program code which, when executed by a computing device, is arranged to perform methods according to an embodiment of the second and / or third aspect of the present invention.The non-transitory computer readable data storage medium may comprise or consist of any type of computer memory, in particular semiconductor memory, such as a solid state memory. The data carrier can also comprise or consist of a CD, a DVD, a Blu-ray disc, a USB memory stick or the like.According to a sixth aspect, the invention provides a data stream comprising executable program code or configured to generate executable program code which, when executed by a computing device, is configured to perform methods according to an embodiment of the second and / or third aspect of the present invention.Further advantageous variants, options, embodiments and modifications are evident from the following figures, the detailed description, and from the claims. It is to be understood, however, that the detailed description and specific examples, while indicating preferred embodiments of the invention, are given by way of illustration only, since various changes and modifications within the scope of the invention will be apparent to those skilled in the art.Individual embodiments of the present disclosure will be explained in detail with reference to the following figures. The elements in the drawings are not necessarily to scale, presenting a further understanding of the principles of the present invention. Parts in the various figures corresponding to the same elements or steps have been given the same reference numerals throughout the figures. The numbering of method steps is used initially only for distinguishing them and does not necessarily imply a corresponding sequence, wherein however it represents a variant to carry out the steps in the sequence of their numbering. Several steps can also be carried out overlapping or simultaneously. Of the figures, FIG. 1 is a schematic block diagram for explaining an apparatus according to an embodiment of the present invention; FIG. 2 shows a schematic representation of image data of a medical scene; FIG. 3 shows an exemplary representation of image data of a medical scene with two instruments; FIG. 4 is a schematic diagram illustrating the artificial intelligence entity during its training in accordance with another embodiment of the present invention; FIG. 5 is a schematic flow chart for explaining another embodiment of the present invention; FIG. 6 is a schematic flow chart for explaining still another embodiment of the present invention; FIG. 7 is a schematic block diagram of a computer program product according to another embodiment of the present invention; and FIG. 8 is a schematic block diagram of a non-transitory computer readable storage medium according to another embodiment of the present invention.FIG. 1 shows a schematic block diagram for explaining an apparatus 100 according to a first embodiment of the present invention, i.e. an apparatus 100 for processing image data of a medical imaging device 50.The imaging device 50 can be or comprise, for example, an endoscope or a microscope. It is designed for capturing image data 71 from a medical scene 1. Typically, such image data 71 are displayed by a display device 200, frequently in real time, wherein overlays, evaluations, visualizations of radiations invisible to the human eye, etc., can be added to the visual display. This is often done by a camera control unit of the endoscope or microscope. The endoscope can be a rigid, a flexible or a semi-rigid endoscope.The apparatus 100 according to the present invention may also include the imaging device 50 and / or the display device 200 in some variants, and may include a camera control unit or may be wholly or partly integrated into a camera control unit.The apparatus 100 comprises an input interface 110 which is configured to obtain image data 71 recorded by the medical imaging device 50, whether directly, whether via a PACS, or from another type of database.The apparatus 100 also comprises a computing device 150 which is configured to carry out various functions which are explained below with reference to various modules. As has already been explained in the preceding, these modules do not necessarily have to be implemented as clearly distinguishable units.First, the computing device 150 provides a structure recognition module 152 which is configured to determine a 3D structure in the captured image data 71, preferably to convert the entire captured image data 71 into such a 3D structure, for example in the form of a depth map. The structure recognition module 152 may include an artificial intelligence, KIE, entity for this purpose, such as described in Freedman et al.The computing device further provides a scaling module 154, which is configured to determine a scale factor for the 3D structure, on the basis of which an actual size of a constituent (or, preferably, of all constituents) of the 3D structure can be determined. For example, a depth map may be provided with a scale factor that relates distances within the depth map to the actual distances in the physical basis for the 3D structure. The display device 200 can then, for example, automatically or upon request from a user, superimpose dimensions of components of interest of the 3D structure. Alternatively or additionally, a scale can also be blended in, which always indicates the actual size of the 3D structure shown.An output interface 190 of the device 100 is finally configured to output an output signal 79 which comprises or indicates the determined scale factor. The output signal 79 may be a signal to the display device 200 to present a scale factor-based virtual representation of the 3D structure, a scale factor-based scale display, or the like.As has already been explained in the foregoing in an overview, there are a large number of different variants of how the scale factor can be determined and for which it can be used. In the following, a variant is described as a specific exemplary embodiment, with which, for example, the size of stones (e.g. kidney stones, gall stones, bladder stones) can be measured. Thereby, it can be estimated whether fragments of stones can be removed by a working channel. For illustrative purposes, reference is also made to FIG. 2 in the following description.FIG. 2 shows a schematic representation of image data 71 of a medical scene 1. an organic tissue and an instrument, namely a laser fiber 2 (or also: laser fiber, from Engl. "laser fiber") can be seen. Laser light, e.g. a holmium laser or thulium laser, can be transported by means of such laser fibers 2 in order to fragment stones at the target site, for example (lithotripsy). Such laser fibers 2 have a constant diameter, which thus represents a known dimension 4 in the case of known laser fiber 2.In order for the user to guide the laser fiber 2 to the correct location and to be able to trigger the laser, image data 71 are acquired in real time by means of an imaging device 50, which image data at least partially (ideally always) contain, i.e. image, the tip of the laser fiber 2. The captured image data 71 can be displayed in real time by a display device 200, as shown in FIG. 2. FIG. 2 also illustrates quite clearly how the tissue structure per se hardly allows any conclusions about the scale factor of the medical scene 1 shown.The image data 71 acquired in real time can thus advantageously comprise an image data set, namely a series of images following one another in time, for example with an image repetition rate of between 10 and 100 hertz, that is to say with images recorded successively between 10 and 100 milliseconds. These can be processed well by the human eye and at the same time offer a good database for the subsequent processing.In the described variant, the 3D structure of the image data 71 is first determined (or: recognized, or: acquired) by the structure recognition module 152. This advantageously comprises both the tissue sections contained in the image data 71 and any objects visible therein, such as the laser fiber 2.The structure detection module 152 may use various algorithms for detecting the 3D structure, for example, classic image processing algorithms such as SfM (Structure from Motion), for example, using SLAM or vLAM algorithms, where SLAM stands for simultaneous localization and mapping (SLAM) and vS stands for visual simultaneous localization and mapping (SLAM). SfM denotes the acquisition of the 3D structure of the medical scene 1 from an image data set of successive 2D images. In vLAM, the position and orientation of the imaging device 50 (specifically its input optics) with respect to its environment is calculated with simultaneous mapping of this environment.Alternatively, or additionally, the structure recognition module 152 may also include a first artificial intelligence, KIE, entity trained and configured to obtain the obtained image data 71 as input data and, based thereon, generate a first output indicative of the 3D structure.In the present embodiment, the computing device also comprises an object recognition module 156, by means of which an object having at least one known dimension 4 is also recognized in the image data 71. In the example of FIG. 2, this is the known laser fiber 2 with its known diameter. In fact, however, not only "one" diameter is known in the laser fiber 2, but a plurality of diameters are known at a plurality of locations along the laser fiber 2, since the diameter is constant across the elongated body of the laser fiber 2. With this additional information, the position of the laser fiber 2 with respect to the remaining 3D structure can optionally be determined particularly accurately from the illustration shown in FIG. 2.For example, the object recognition module 156 may include a second artificial intelligence, KIE, entity trained and configured to obtain the obtained image data 71 as input data and, based thereon, generate a second output that displays (or: indexes, or: indicates) at least one object (here: the laser fiber 2) recognized in the image data 71 having at least one known dimension 4, preferably together with additional information such as its exact position (i.e., the object is segmented in the image data 71) and / or orientation, or at least the position and / or orientation of the known dimension 4 in the image data 71.This object recognition, especially segmentation, can be performed in the obtained (typically 2D) image data 71 whereupon the boundaries of the object thus determined, or at least the known dimension 4, can be transferred into the determined 3D structure. Object recognition, or segmentation, can be performed using a typical artificial neural network, ANN, from the prior art, for example using a ResNet50 trained with segmented object images, in particular images of medical instruments.The scaling module 154 may now associate the captured 3D structure and the at least one known dimension 4 to each other, thereby determining the scale factor. The scale factor can be adapted in such a way that the diameter of the laser fiber 2 in the specific 3D structure corresponds as well as possible to the known dimension 4. The scaling module 154 may also be implemented in whole or in part by the second artificial intelligence entity, KIE, by being configured to output an output indicative of the scale factor. The output may thus be the scale factor itself, or a control signal for a display device 200 to display a size or scale indication, or a scaled 3D structure (e.g., a scaled depth map), or the like.Similar to the laser fiber 2, it is also possible to move with further instruments which have elongate sections or shafts, the diameter of which can be known to the computing device 150, in particular to the object recognition module 156. The device 100 may also include an input interface by means of which a user may input a dimension. For example, if the object detection module 156 detects a suitable tool having an elongated shank, a user may be prompted by the user interface to enter the diameter of the shank. If the user subsequently arrives, the diameter is then treated as known dimension 4. The user interface can be realized, for example, using the display device 200, wherein the user can be input via conventional means (keyboard, mouse, trackpad, touchscreen, speech recognition and the like).FIG. 3 shows an exemplary representation of image data 71 of a medical scene 1 with two instruments 3, the shaft diameters of which are each known dimensions 4. As in the case of the laser fiber 2, the respective diameter is thus a known dimension 4 at a plurality of locations along the shaft. For example, the shaft diameter can be 5 mm and can be mapped in the image (here an endoscopic image) to a width of 100 pixels. The scaling module 154 may now scale the previously determined depth map such that the corresponding points in the 3D structure are spaced 5 mm apart at the same depth at which the shank diameter is located.In FIG. 3, it is also shown by way of example that a distance d between two detected objects 3 can be calculated and displayed by the display device 200, here on the upper left in the image. For this purpose, the display device 200 can display additional auxiliary lines based on the detected (or segmented) objects 3, as indicated in FIG. 3. Small "+" symbols show, on the one hand, fictitious instrument tips between which the distance is calculated. The user can thus guide this information to tissue of interest or an object of interest and thus carry out a length measurement. Further small "+" symbols indicate the points at which the radii, and thus the diameters, of the shafts of the instruments 3 are calculated.The computing device 150 may also include an artificial intelligence entity, KIE, configured to determine the 3D structure identically including the scale factor. This artificial intelligence entity, KIE, can thus function as both a structure detection module 152 and a scaling module 154, and optionally even as an object detection module 156 (e.g., in an intermediate stage).In the following, with reference to FIGS. 4 and 5, a method for training such an artificial intelligence entity, KIE, is described.FIG. 4 is a schematic diagram illustrating the artificial intelligence entity, KIE 80, during its training.FIG. 5 shows a schematic flow diagram according to an aspect of the third aspect of the present invention, i.e. a method for training an artificial intelligence entity, KIE, for determining a scale factor for a 3D structure in image data 71 of a medical imaging device 50.Referring to FIGS. 4 and 5, in a step S 10, a plurality of image data sets is provided (as training data), each image data set comprising first image data 71 (or: a first frame) of a medical scene 1 at a first time t=0 and at least second image data 72 (or: a second frame) of the medical scene 1 at a second time t'=1.Between the individual image data sets, or at least between some, the medical scene 1 may advantageously vary so that the artificial intelligence entity, KIE 80, is not excessively trained on a particular medical scene 1. In contrast to the known methods in the prior art, an implementation of the artificial intelligence entity, KIE 80, described herein is applicable to all types of medical scenes 1, regardless of their size scale, since the determination of the associated size scale (i.e. the scale factor) is included. By a plurality of different medical scenes 1 in the training data, i.e. by as many image data sets as possible of different medical scenes 1, the KIE is trained particularly robust. In this context, various medical scenes 1 are to be understood as meaning, in particular, those which show different organs, that is to say, for example, a blood vessel, an abdominal cavity, an intestine, a heart, a kidney, a bile, a bladder and the like.In a step S 20, a randomly initialized or pre-trained ("pre-trained") artificial intelligence entity, KIE 80, is provided, having an architecture or structures which are explained in more detail below:The KIE 80 is designed such that the first image data 71 provided, optionally additionally also the second image data 72, are input into a structure recognition module 152 of the artificial intelligence entity, KIE 80, which is designed to determine a 3D structure, for example a depth map 91, of the image data 71, 72 input into it.The KIE 80 is further configured such that the first image data 71 and the second image data 72 are input into a pose recognition module 81 of the KIE 80. The pose recognition module 81 is configured to recognize a respective pose 93, 94 of the image data 71, 72, at least in relation to one another, on the basis of the image data 71, 72 input thereto. In this context, a pose is to be understood as the alignment of an imaging device 50 with respect to the medical scene 1 to be captured. The pose is typically indicated with an orientation matrix (which indicates rotational angles with respect to an orthonormal coordinate system) and a translation vector. In other words, it is determined by the pose recognition module 81 how (or whether) the imaging device 50 has moved between the acquisition of the first image data 71 and the second image data 72.The difference between the determined first pose 93 of the first image data 71 and the determined second pose 94 of the second image data 72 can be used to associate a depth map 91 for the first image data 71 with a second depth map 92 for the second image data 72, or to convert one into the other, or to estimate one by means of the other.For example, the structure recognition module 152 may determine the first depth map 91 for the first image data 71 based on the first image data 71 and determine the second depth map 92 for the second image data 72 based on the second image data 72. The depth map 92 determined for one of the image data (here, without limiting generality, the second image data 72) is then converted (or: adapted, or: evolved) into an estimated first depth map for the first image data 71 by means of the two determined poses 93, 94. Thus, the estimated first depth map ideally corresponds to the determined first depth map 91.A loss function 99 of the training method contains a term that abstracts deviations between the determined first depth map 91 and the estimated first depth map. As is common in machine learning, during training, parameters (e.g., weights for nodes, thresholds for activation functions, etc.) are iteratively adjusted to minimize the overall loss function.The term thus contributes to training the artificial intelligence entity, KIE 80, such that the structure recognition module 152 becomes more and more effective in generating a depth map 91, 92 to be as consistent as possible with an estimated depth map generated based on a pose difference to previous or subsequent image data 71, 72.According to the present invention, the artificial intelligence entity, KIE 80, additionally includes a scaling module 154 configured to determine a first scale factor 95 for the first image data 71 and a second scale factor 96 for the second image data 72. As already mentioned in the foregoing, such a scaling module 154 may be implemented together with other modules, or separately therefrom. For example, it can be realized together with an object recognition module 156 and determine the scale factor based on any objects 2, 3 with at least one known dimension 4 that can be seen in the image data 71, 72.Accordingly, the loss function 99 may have a further term, or an existing term may be modified such that deviations between the scale factors 95, 96 are penalised for the first image data 71 and the second image data 72, respectively, and thus minimized during training. Similar to the depth maps 91, 92, an estimated scale factor for the first image data 71 can also be determined from the second scale factor 96 for the second image data 72 using the determined poses 93, 94, which scale factor can in turn be compared with the first scale factor 95 actually determined for the first image data 71, wherein the difference is positively incorporated into the loss function 99, i.e. a higher difference ceteris paribus leads to a higher value of the loss function 99 to be minimized.In a step S 30, the artificial intelligence entity, KIE 80, is thus now iteratively trained in such a way that the loss function 99 is minimized. This can be done, for example, with unsupervised learning, for example in the manner as described in Freedman et al. For training, all mechanisms and methods known in the prior art can be used, such as a gradient descent method ("gradient descent"), training across epochs, and so forth.In this way, the artificial intelligence entity, KIE 80, is thus trained to obtain temporal sequences of (at least or exactly 2) image data 71, 72 (or: frames) and to find a consistent solution across the image data 71, 72 for the depth maps 91, 92, the poses 93, 94, and the scale factors 95, 96. In this way, the artificial intelligence entity, KIE 80, is also able to determine a scale factor 95, 96 for current image data 71, 72, even if no object 2, 3 with at least one known dimension 4 is currently contained therein.The artificial intelligence entity, KIE 80, trained to a desired level of accuracy may then be implemented, for example, by the computing device 150 of the apparatus 100.For its use (in the productively mode, "deployment stage"), the artificial intelligence entity, KIE 80, can be provided as it was trained. Alternatively, the KIE 80 may be provided with only the structure recognition module 152 and the scaling module 154. The pose detection module 81, on the other hand, may be removed or deactivated depending on the application, or its output may be ignored. Depending on the application and training, the KIE 80 in productive mode can also be configured to receive individual images as input in each case, or else an image data set (i.e. an image sequence). It is also conceivable that the structure recognition module 152 receives and uses individual images as input, while the scaling module 154 receives image data sets (i.e. image sequences), or vice versa.Although the description herein has always focused on first image data 71 and second image data 72, it is understood that the image data sets may also additionally include further image data, for example third image data. The described mechanisms can be extended to a triple in this case, wherein, for example, estimates are generated for the second image data 72 situated in the middle in time based both on the temporally preceding first image data 71 and on the temporally following third image data. The loss function 99 can thus comprise deviations of the second depth map 92 and / or of the second scale factor 96 determined for the second image data 72 both from the corresponding estimates on the basis of the quantities determined for the first image data 71 and from the corresponding estimates on the basis of the quantities determined for the third image data.FIG. 6 shows a schematic flow diagram for explaining a method according to an embodiment of the third aspect of the present invention, i.e. a computer-implemented method for processing image data 71, 72 of a medical imaging device 50. The method according to FIG. 6 can be carried out in particular by means of the apparatus 100 from FIG. 1 and can therefore be adapted according to all options or variants described with respect to the apparatus 100 according to the invention and vice versa.In a step S 100, image data 71, 72 are provided, which have been acquired by a medical imaging device 50. This providing S 10 can comprise capturing of the image data 71, 72 by the medical imaging device 50. Further options have already been explained above by way of example with reference to the input interface 110.In a step S 200, a 3D structure is determined in the captured image data 71, 72, for example a depth map 91, 92 The determination S 200 of the 3D structure can in particular take place as described above with reference to the structure recognition module 152, for example using an artificial intelligence entity, KIE 80.In an optional step S 300, an object 2, 3 having at least one known dimension 4 is detected in the captured image data 71, for example as described above with reference to the object detection module 156.In a step S 400, a scale factor 95, 96 is determined for the 3D structure, from which an actual size of a constituent (or each constituent) of the 3D structure can be determined, for example as described above with reference to the scaling module 154 and / or the object recognition module 156 and / or the artificial intelligence entity, KIE 80.FIG. 7 shows a schematic block diagram of a computer program product 300 according to an embodiment of the third aspect of the present invention. The computer program product 300 comprises executable program code 350 which, when executed (e.g. by a computing device), is configured to perform the method according to an embodiment of the present invention, for example according to FIG. 5 or 6.FIG. 8 shows a schematic block diagram of a non-transitory computer readable storage medium 400 according to an embodiment of the present invention. The data storage medium 400 comprises executable program code 450 configured, when executed (e.g. by a computing device), to perform the method according to an embodiment of the present invention, for example according to FIG. 5 or 6.The non-transitory computer-readable data storage medium 400 may be embodied as or comprise a semiconductor memory, e.g. an SSD memory brick, for example. The data storage medium 400 may also include or comprise a CD, DVD, Blu-ray or magnetic storage device.The foregoing description of the disclosed embodiments includes merely examples of possible implementations described to enable one skilled in the art to make or use the present invention. Various variations and modifications of these embodiments will be readily apparent to those skilled in the art, having the benefit of the present invention, and the general principles defined herein may be applied to other embodiments without departing from the scope of the present disclosure.Thus, the present invention is not intended to be limited to the specific embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and features disclosed herein.List of reference characters1 Medical scene 2 Laser fiber 3 Instruments 4 Known dimension 5 Tip of laser fiber 50 Imaging device 71 Image data 79 Output signal 80 Artificial intelligence entity 81 Pose detection module 91 Depth map 92 Depth map 93 Pose 94 Pose 95 Scale factor 96 Scale factor 99 Loss function 100 Apparatus 110 Input interface 150 Computing device 152 Structure detection module 154 Scaling module 156 Object detection module 190 Output interface 200 Display device 300 Computer program product 350 Program code 400 Data storage medium 450 Program code S 10..S 30; S 100..S 400 Method stepsReferences included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Cited Non-Patent LiteratureDetecting Deficient Coverage in Colonoscopes" by D. Freedman et al, arXiv:2001.08589v 3, March 29, 2020
[0005]
Claims
Apparatus (100) for processing image data (71, 72) of a medical imaging device (50), comprising: an input interface (110) which is configured to obtain image data (71, 72) acquired by a medical imaging device (50); a computing device (150) which is configured to implement at least one structure recognition module (152) and a scaling module (154); wherein the structure recognition module (152) is configured to determine a scaleless 3D structure (91, 92) in the acquired image data (71, 72); wherein the scaling module (154) is configured to determine a scale factor (95, 96) for the scaleless 3D structure (91, 92), on the basis of which an actual size of at least one component of the scaleless 3D structure (91, 92) can be determined; and an output interface (190) which is configured to generate an output signal (79) which comprises or displays the determined scale factor (95, 96).The apparatus (100) according to claim 1, wherein the computing device (150) is additionally configured to implement an object detection module (156) configured to detect, in the obtained image data (71, 72), at least one object (2, 3) having at least one known dimension (4); and wherein the scaling module (154) is configured to determine the scale factor (95, 96) based on the position and / or orientation of the detected object (2, 3) and its known dimension (4), or at least based on the position and / or orientation of the at least one known dimension (4) of the detected object (2, 3).The device (100) according to claim 2, wherein the object (2, 3) with the at least one known dimension (4) comprises at least one medical instrument (2, 3), in particular a laser fiber (3), and / or at least one component of a medical imaging device.The apparatus (100) according to any one of claims 1 to 3, wherein the structure recognition module (152) and the scaling module (154) comprise a common artificial intelligence entity (80) which is configured to obtain the obtained image data (71, 72) as input data and based thereon generate an output which indicates the scaleless 3D structure (91, 92) and the scale factor (95, 96).The apparatus (100) according to any one of claims 1 to 3, wherein the structure recognition module (152) comprises a first artificial intelligence entity, KIE, configured to obtain the obtained image data (71, 72) as input data and based thereon generate a first output indicative of the scaleless 3D structure (91, 92); and wherein the computing device (150) implements a second artificial intelligence entity, KIE, configured to obtain at least the first output of the first artificial intelligence entity or based thereon data as input and based thereon generate a second output.The apparatus (100) of claim 2 or 3 in combination with claim 5, wherein the object detection module (156) comprises the second artificial intelligence entity, KIE, and this second KIE is configured to detect the at least one object (2, 3) having the at least one known dimension (4) in the obtained image data (71, 72), wherein the second output indicates the detected object (2, 3), and wherein the scaling module (154) is configured to obtain the at least one known dimension (4) of the detected object (2, 3) using a database and to determine the scale factor (95, 96) based thereon.The apparatus (100) according to any one of claims 1 to 6, wherein the computing device (150) is further configured to generate a scaled depth map of at least the determined scaleless 3D structure (91, 92) based on the determined scale factor (95, 96) and to store or output said scaled depth map.The apparatus (100) according to any one of claims 1 to 8, wherein the structure recognition module (152) is configured to determine the scaleless 3D structure (91, 92) substantially in real time and / or wherein the scaling module (154) is configured to determine the scale factor (95, 96) substantially in real time.The apparatus (100) according to any one of claims 1 to 8, wherein the apparatus (100) comprises the medical imaging device (50), in particular an endoscope or a microscope.The apparatus (100) according to claim 9, wherein the medical imaging device (50) comprises a camera control unit, and at least the computing device (150) is integrated into the camera control unit.The apparatus (100) according to any one of claims 1 to 10, wherein the captured image data (71) comprises an image data set, which comprises first image data (71) of a medical scene (1) at a first point in time and second image data (72) of the medical scene (1) at a second point in time; and wherein the structure recognition module (152) is further configured to determine the scaleless 3D structure (91, 92) at least for the second image data (72) based on both the first image data (71) and the second image data (72), and / or the scaling module (154) is configured to determine the scale factor (95, 96) for the scaleless 3D structure (91, 92) at least for the second image data (72) based on both the first image data (71) and the second image data (72).A computer-implemented method for training an artificial intelligence, KIE, entity (80) to determine a scale factor (95, 96) for a scaleless 3D structure (91, 92) in image data (71, 72) of a medical imaging device (50), comprising: providing (S10) a plurality of image data sets, each image data set comprising first image data (71) of a medical scene (1) at a first time and second image data (72) of the medical scene at a second time (1); providing a randomly initialized or pre-trained artificial intelligence, KIE, entity (80) configured to receive the first image data (71) and the second image data (72) as input, determine a first scale factor (95) for the first image data (71), determine a second scale factor (96) for the second image data (72), determine a first pose (93) for the first image data (71), determine a second pose (94) for the second image data (72), and calculate an estimated first scale factor for the first image data (71) based on the first pose (93), the second pose (94) and the second scale factor (96); and training the artificial intelligence entity, KIE (80), iteratively changing parameters of the KIE (80) to minimize a loss function (99) that penalizes, among other things, differences between the determined first scale factor (95) and the estimated first scale factor.A computer-implemented method for processing image data (71, 72) of a medical imaging device (50), comprising: providing (S100) image data (71, 72) acquired by a medical imaging device (50); determining (S200) a scaleless 3D structure (91, 92) in the acquired image data (71, 72); and determining (S300) a scale factor (95, 96) for the scaleless 3D structure (91, 92), on the basis of which an actual size of a constituent part of the scaleless 3D structure (91, 92) can be determined.A computer program product (300) comprising executable program code (350) which, when executed, is configured to perform the method of claim 12 or 13.A non-transitory computer readable storage medium (400) comprising executable program code (450) which when executed is configured to perform the method of claim 12 or 13.
Citation Information
Patent Citations
Surgical microscope
DE102014007909A1
Cell counting or cell confluence with rescaled input images
US20220284719A1