Co-design of optical filters and fluorescence applications using artificial intelligence

JP2024526099A5Pending Publication Date: 2025-06-18CARL ZEISS MEDITEC AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577468
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-16
Filing Date
2022-06-15
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

Current surgical methods for differentiating between healthy and diseased tissue during brain tumor resection are inefficient, requiring high computational power and time, leading to potential removal of excessive healthy or diseased tissue, and necessitate frequent lighting adjustments, increasing surgical time and cost.

Method used

A computer-implemented method using a microsurgical optical system with a machine learning model that predicts digital fluorescence images by capturing tissue samples with a first digital image capture unit and an optical filter, training the model with hyperspectral data to reduce color channels, and optimizing filter parameters for real-time differentiation.

Benefits of technology

Enables real-time, minimally invasive tissue differentiation without additional biopsies, reducing computational power requirements and surgical time, while improving precision in tissue removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method for predicting a digital fluorescence image, comprising the steps of capturing a first digital image of a tissue sample by a microsurgical optical system using a first digital image capture unit having a first number of color channel information with white light and at least one optical filter, and predicting a second digital image in the form of a digital fluorescence representation of the captured first digital image using a trained machine learning system having a trained learning model for predicting a corresponding digital fluorescence representation of the input image, The first captured digital image is used as an input image for the trained machine learning system, and parameter values ​​of the at least one optical filter are determined during the training process of the machine learning system.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a prediction method with a machine learning model, in particular to a computer implemented method for predicting digital fluorescence images.The present invention further relates to a prediction system and a computer program product for predicting digital fluorescence images. [Background technology]

[0002] Brain tumors include a type of cancer that is relatively aggressive and often has a relatively low success rate of treatment, with a survival probability of about 1 / 3. The treatment of such diseases typically requires surgical intervention for removal, radiation therapy, and / or subsequent chemotherapy, usually over a long period of time. Biopsy often forms the basis of the respective treatment, and molecular tests are also used. Of course, such treatments are associated with medical risks. The potential for analyzing images recorded by radiometers has recently been very advanced, so that such tumor tests can at least complement biopsies. Such complementation may be the case when biopsies are not possible or are not desirable. Recently, such image-based diagnostics have even become available during surgery. However, the required computing power is currently very high, which is why real real-time assistance has not been possible so far.

[0003] It is not only in the field of brain tumors that diseased tissue areas must be completely and as accurately as possible removed in order to prevent diseased tissue, i.e. tumor-containing tissue, from growing back into healthy tissue and to leave as much of the healthy tissue as possible. This partial resection (resection) of tissue is usually performed by a surgeon in an operating room equipped with special instruments. For this purpose, an operating microscope is also generally used. In this case, the exact border between healthy and cancerous tissue can only be recognized with difficulty under the typical white operating light. As a result, there is an obvious risk of removing too much healthy tissue or too little cancerous tissue after resection. Either result can be designated as suboptimal.

[0004] Up until now, it has generally been necessary to further inject contrast agents to relatively clearly distinguish healthy and diseased tissue under optimized illumination (e.g., BLUE 400 or YELLOW 560). As a result, it is also necessary to have to switch between illumination presets relatively frequently during surgery, which in principle leads to longer surgery times, higher overall surgical costs, and less patient-friendly.

[0005] The prior art does disclose early approaches for using artificial intelligence to verify fluorescent images from tissue recordings recorded under white light, but still requires high-resolution cameras with a large number of color channels.

[0006] It is therefore desirable to have minimally invasive surgical assistance that can help surgeons clearly distinguish between healthy and cancerous tissue in real time (i.e., without significant time delay) and without significant distraction, so that patients only need to undergo minor additional treatments (e.g., chemotherapy). Summary of the Invention [Means for solving the problem]

[0007] This object is achieved by the method, a corresponding system and an associated computer program product proposed herein according to the independent claims. Further configurations are described by the respective dependent claims.

[0008] According to one aspect of the present invention, a computer-implemented method for predicting a digital fluorescence image is presented. In this case, the method can include capturing a first digital image of a tissue sample by a microsurgery optical system having a first digital image capture unit having a first number of color channel information (or number of color channels) using white light and at least one optical filter. The method can further include predicting a second digital image in the form of a digital fluorescence representation of the captured first digital image by a trained machine learning system including a trained learning model for predicting a corresponding digital fluorescence representation of the input image. In this case, the first captured digital image can be used as an input image for the trained machine learning system, and parameter values ​​of the at least one optical filter can have been determined during training of the machine learning system.

[0009] According to another aspect of the present invention, a prediction system for predicting a digital fluorescence image is presented. The prediction system may include a microsurgery optical system having a first digital image capture unit with a first number of color channel information (or number of color channels) and an optical filter for capturing a first digital image of a tissue sample using white light, and at least one trained machine learning system. In this case, the machine learning system may include a trained learning model for predicting a corresponding digital fluorescence representation of an input image. The machine learning system may be adapted to predict a second digital image in the form of a digital fluorescence representation of the captured first digital image. In this case, the first captured digital image may serve as an input image of the trained machine learning system, and parameter values ​​of the at least one optical filter may have been determined during training of the machine learning system.

[0010] The proposed computer-implemented method for predicting digital fluorescence images has several advantages and technical effects that may be suitably applied to related systems.

[0011] The practice in which underlying computer systems of different performance levels are used to create machine learning models in a training phase and to use them in a prediction phase is further developed by the method proposed herein. As a rule, the time available for developing machine learning models is longer during the training phase than during the prediction phase. During the prediction phase, i.e. during productive use, the output of the machine learning system must be possible with as short a time delay as possible. The shortest possible time delay is essential mainly when the machine learning system is used in real-time situations, for example to support surgical interventions.

[0012] According to the concept proposed herein, a hyperspectral camera can be conveniently used to create the training data, which makes available a large number of different color channel information. From this starting point, during the training process in the integrated approach, the number of color channels is significantly reduced by digitally simulated optical filters in order to anticipate the available resources during the prediction phase as early as during the training phase. To that end, both the machine learning system with the integrated machine learning model and the parameters of the digitally simulated optical filters are adapted or trained simultaneously in an integrated optimization process. Furthermore, additional constraints can be taken into account, such as the spectrum of the illumination light used for the tissue sample illuminated by the digital capture system.

[0013] After the training is finished, the ascertained parameters of the digitally simulated optical filter can be adopted to the physical filter present in reality. Furthermore, a digital capture unit (e.g. RGB camera) can be used, the number of color channels and its type (e.g. in terms of wavelength sensitivity) being adapted as far as possible to the reduced number of color channels during the training process. In this regard, the constraints of the digital capture unit (e.g. RGB camera) available during the later productive use can also be already predefined during training. These measures serve to significantly reduce the required computational power during the prediction phase, making the proposed method and the corresponding system usable in real time to support surgical interventions. Previously proposed methods have been frustrated by this obstacle, either because no integrated optimization was provided or because the required computational power was insufficient for real-time support.

[0014] In this way, a representation of the tissue to be operated on can be provided to the surgeon by a fluorescent representation, which allows the surgeon to distinguish between healthy and diseased tissue clearly and without further distracting the surgeon's attention. This is true even if in productive operation, i.e. during the prediction phase (inference), only relatively simple cameras (as opposed to hyperspectral cameras for the training phase) can be used.

[0015] In this way, an optical representation of the surgical target tissue is generated in real time, without the need for time-consuming multidimensional biopsies, which is of great advantage to the surgeon.

[0016] From a technical point of view, the juxtaposition of differently optimized individual machine learning systems, the data transfer between the learning systems, as well as additional high expenditures on time and optimization of the individual machine learning systems are also avoided.

[0017] The system presented here is finally very advantageously based on hardware in the form of optical filters that are physically manufactured for the prediction phase, the parameters of which are ascertained during a joint optimization during the training phase, and on the optimization of the hardware and software of the machine learning system used.

[0018] Further exemplary embodiments are presented below, which may have utility both in connection with the methods and in connection with corresponding systems.

[0019] According to an advantageous embodiment of the method, the trained learning model of the trained machine learning system may include providing a plurality of first digital training images of the tissue sample captured under white light by a microsurgical optics system having a second image capture unit, where a second number of color channel information in different spectral ranges (e.g., 64 color channels of a hyperspectral camera) may be available for each first digital training image.

[0020] The method may further include providing a plurality of second digital training images each representing the same tissue sample as the first set of digital training images, the second digital training images may have an indication of a pathological element of the tissue sample, and training the machine learning system to form a trained machine learning model for predicting digital images of the type of the plurality of second digital training images. In this case, as input values ​​for the machine learning system, the plurality of first digital training images in the form of a second number of color channel information (i.e., for example, 64 color channels), the plurality of second digital training images as "ground truth", and parameter values ​​for reducing the second number of color channel information by at least one digitally simulated optical filter to form the first number of color channel information may be used. In this case, the term "ground truth" as a learning target of training will be recognized by those skilled in the art. This includes generating an annotated version of the training images intended to predict their appearance in the case of unknown input for the machine learning system, or it as an output value / output mapping.

[0021] In this case, after reducing the second number of color channel information to the first number of color channel information by the digitally simulated optical filter, the plurality of first digital training images can be used as training data for predicting digital images of the plurality of second digital training images type. In addition, after the training of the machine learning system is completed, at least a portion of the parameter values ​​of the at least one optical filter can be output as output values ​​of the machine learning system. In this context, the portion of the parameter values ​​of the at least one optical filter can also mean that all the confirmed parameter values ​​are output.

[0022] It is noted that the tissue sample may be diseased tissue, for example in the form of cancer tissue, specifically diseased brain tissue. It should be further noted that the second training images were generated as captured images under visible light and / or under specific lighting conditions (e.g. ultraviolet light).

[0023] According to a further advantageous embodiment of the method, the parameter values ​​of the at least one optical filter may include the number of first color channel information and / or the filter shape of the digitally simulated optical filter. It should be noted in this case that the filter can also consist of a plurality of digitally simulated optical filters. In this way, different parameter values ​​of the optical filter can be varied. During the manufacture of the optical filter, a plurality of these parameter values ​​can be realized simultaneously.

[0024] According to a sophisticated embodiment of the method, the second number of color channel information (i.e. the number of color channels) may be greater than the first number of color channel information. By way of example, the second number may constitute the number of color channels of the hyperspectral camera, for example 64 color channels (more broadly, for example 30 to 130 color channels), whereas the first number constitutes the number of color channels of the first digital image capture unit. The latter may comprise less than 10, or preferably 3 to 4 color channels. In this way, the number of color channels can be significantly reduced, which in turn clearly reduces the required computational power. This has a favorable effect on the direct use of the method during ongoing operation.

[0025] According to an extended embodiment of the method, the parameter value for reducing the second number of color channels may relate to the filter shape of the filter or the respective center frequency of the first number of color channel information. According to further embodiments based on this embodiment, the parameter value for reducing the second number of color channels may also relate to one or more camera sensitivity curves as envelopes (e.g. one per color channel) or their number and / or the shape of the geometrically described structure, which may relate to a Gaussian distribution, a rectangular shape, etc.

[0026] According to an extended embodiment of the method, parameter values ​​for controlling the source of white light during capture of the first digital image can be generated as additional output values ​​of the machine learning system after training of the machine learning system has been completed. This also allows the environmental variables used during training of the machine learning system with the machine learning model to be made available for actual productive use of the method.

[0027] According to one useful embodiment of the method, the digital fluorescence representation may correspond to a representation as produced using a light source in the UV range, e.g. BLUE 400. Furthermore, the spectral range may also be represented using preferences. Nevertheless, in the context of the present application, the designation "fluorescence representation" should be usable for these other spectral ranges as well. Conveniently, the fluorescence representation uses (depending on the light used) a contrast agent in the tissue sample, e.g. 5-ALA (aminolevulinic acid) for excitation in the wavelength range of 430-440 nm or sodium fluorescein for excitation in the wavelength range around 560 nm.

[0028] According to a further advantageous embodiment of the method, the learning model may correspond in its setup to an encoder-decoder model.

[0029] Furthermore, embodiments may relate to a computer program product accessible from a computer usable or computer readable medium including program code for use by or in association with a computer or other instruction processing system. In the context of this specification, a computer usable or computer readable medium may be any device suitable for storing, communicating, transferring or carrying program code.

[0030] It is noted that exemplary embodiments of the present invention may be described with respect to different implementation categories. In particular, some exemplary embodiments may be described with respect to a method, whereas other exemplary embodiments may be described in the context of a corresponding device. Nevertheless, a person skilled in the art may identify and combine possible combinations of method features from the above and below descriptions, as well as possible combinations of features from the corresponding system, even if they belong to different claim categories, unless otherwise specified.

[0031] The above-mentioned aspects and further aspects of the present invention will become apparent from the additional and further illustrative configurations described, inter alia, in relation to the exemplary embodiments and drawings described below.

[0032] Preferred exemplary embodiments of the present invention will now be described by way of example and with reference to the following drawings, in which: [Brief description of the drawings]

[0033] [Figure 1] 1 shows a flow chart-like representation of an exemplary embodiment of a computer-implemented method according to the present invention for predicting digital fluorescence images. [Diagram 2] 2 shows an extension of the exemplary embodiment according to FIG. [Diagram 3] 1 shows a diagram of an example embodiment of the training and prediction stages of a machine learning system and the use of filters. [Figure 4] 1 shows a diagram of an example embodiment of the number of color channels during training. [Diagram 5] Two possible wavelength distributions are shown as a result of the joint optimization of the machine learning system and the digitally simulated filter. [Figure 6] 1 shows a prediction system for predicting digital fluorescence images. [Figure 7] 7 illustrates an exemplary embodiment of a computer system including a system according to FIG. 6. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0034] In the context of this specification, the conventions, terms and / or expressions should be understood as follows.

[0035] The term "machine learning system" as used herein may refer to a system or method used to generate output values ​​in a non-procedurally programmed manner. To that end, in the case of supervised learning, a machine learning model present in the machine learning system is trained using training data and associated desired output values ​​(annotated data or ground truth data). The training phase may be followed by a production phase, i.e. a prediction phase ("inference phase"), while output values ​​are generated / predicted from previously unknown input values ​​in a non-procedural manner. Many different architectures for machine learning systems are known to those skilled in the art. They also include neural networks, which may be trained and used, for example, as classifiers. During the training phase, desired output values ​​given predefined input values ​​are learned, typically by a method called "backpropagation", and parameter values ​​of the nodes of the neural network or connections between the nodes are automatically adapted. In this way, an inherently existing machine learning model is adjusted or trained to form a trained machine learning system having the trained machine learning model.

[0036] In light of the above discussion, the term "prediction" may refer to a stage of productive use of a machine learning system. During the prediction stage of a machine learning system, output values ​​are generated or predicted based on a trained machine learning model provided with previously unknown input data.

[0037] The term "digital fluorescence image" as used herein refers to a digital image of a tissue sample having the appearance of an image observed under a particular wavelength of light (e.g., UV light) with the aid of a contrast agent previously added to the tissue sample.

[0038] The term "first digital image" as used herein refers to the capture recording of a digital image by a digital image capture unit (e.g., an RGB camera) having a small number of color channels (e.g., monochromatic or 3 to 4 color channels, generally less than 10 color channels). This first digital image is captured during a prediction stage in order to predict a digital fluorescence image therefrom. Generally, the terms "number of color channels" and "number of color channel information" can be understood as synonyms in the context of this application.

[0039] The term "tissue sample" as used herein may refer to biological material. The tissue sample may be advantageously available for testing in the present method. The tissue sample may include "normal" human tissue or brain tissue.

[0040] The term "microsurgical optics" as used herein may refer to a surgical microscope. The latter may be used for surgical procedures, e.g. minimally invasive interventions. The microsurgical optics may include an illumination system, a digital capture unit (e.g. a digital camera with one or more color channels), an image processing and operator control unit, and an output screen.

[0041] The term "digital image capture unit" as used herein may refer to a camera having a digital image converter. The camera may capture multiple color channel information simultaneously. In the case of a hyperspectral camera, for example, 64 color channels may be captured, although other numbers of color channels (e.g., 30 to 130) are also contemplated. While such cameras can be used to capture training images, cameras having a significantly more conventional number of color channels, for example, on the order of 3 to 4 (or monochromatic) color channels, may be used for productive use during the prediction phase.

[0042] As used herein, the term "first number of color channel information" may refer to the number of color channels of an image capture unit that are used in a predictive mode of operation.

[0043] As used herein, the term "optical filter" may refer to a device that allows selective passage of an incident light beam according to certain criteria. This criteria may be wavelength selective (or polarization state dependent), such that the filter is more transparent to certain distinct wavelength ranges and less transparent (up to essentially no transmission) to other wavelength ranges.

[0044] The term "second digital image" may refer to a predicted image generated by the machine learning system during productive operations.

[0045] As used herein, the term "digital fluorescence representation" may refer to a particular representation of a digital image that appears as if illuminated with light of a particular wavelength (e.g., UV light), such that, for example, fluorescent effects associated with contrast agents present in biological tissue become visible as a result of the light of the particular wavelength.

[0046] The term "white light" refers to light from a light source that emits primarily in the visible wavelength spectrum (eg, white light LEDs, xenon, etc.).

[0047] The term "parameter value of at least one optical filter" refers to substantially the wavelength (or range thereof) over which the filter is transparent. The envelope in the transmittance versus wavelength representation can be a further parameter, such as the average value of the transmission range.

[0048] The term "first digital training image" as used herein refers to a digital image with a large number of color channels available (eg, 64, or more broadly, 30 to 130 color channels).

[0049] The term "second image capture unit" refers to an image capture unit that is adapted to provide multiple color channel data for recording, as is for example the case for a hyperspectral camera.

[0050] The term "second digital training images" refers to digital images of ground truth data that are essential for training a machine learning system in the case of supervised learning. Ground truth data refers to digital images of the kind that are expected as predictions, i.e., that are intended to be "learned from".

[0051] The term "indication of lesion elements of a tissue sample" can be identified by various annotations in the digital image, which may include pixel-by-pixel annotation of a particular image excerpt, areas marked by color in some other way, or a particular form of segmentation of the digital image.

[0052] The term "parameter value for reducing the second number of color channel information" may refer to a characteristic of a digitally simulated filter, which may relate to the number of filters or to the transmission characteristics.

[0053] The term "digitally simulated optical filter" refers to a unit of a machine learning system that modifies color channel information. This may include a digital color transmission filter.

[0054] The term "digitally simulated filter shape of an optical filter" may refer, for example, to the envelope of the transmission spectral line (or wavelength range) of the filter, its mean value, or the associated standard deviation.

[0055] The term "parameter values ​​for controlling a source of white light" may refer to characteristics of the white light used to create a digital image, which may essentially include the spectral range and intensity values ​​of the emitted light of an illumination source.

[0056] The term "encoder-decoder model" as used herein refers to the architecture of a machine learning system that encodes or codes input data so that it is decoded again immediately afterwards. Between the encoder and the decoder, the necessary data is present as a kind of feature vector. During decoding, certain features in the input data can be given special emphasis depending on the training of the machine learning model.

[0057] The term "U-Net architecture" as used herein refers to the architecture of a machine learning model that has contraction and expansion paths internally. A more detailed definition is provided in conjunction with FIG.

[0058] A detailed description of the drawings is given below, where it is understood that all the details and indications in the figures are given in schematic form. First, a flow-chart-like representation of one exemplary embodiment of a computer-implemented method according to the present invention for predicting digital fluorescence images is given. Further exemplary embodiments, or exemplary embodiments of the corresponding system, are described below.

[0059] 1 shows a flow chart-like representation of one preferred exemplary embodiment of a computer-implemented method 100 for predicting a digital fluorescence image. In this case, the method includes capturing 102 a first digital image of a tissue sample (e.g., diseased brain tissue) with a microsurgery optics system having a first digital image capture unit, e.g., an RGB camera, having a first number of color channel information, e.g., 3 to 4 color channels, using white light and at least one optical filter. The filter is typically located between the illuminated tissue sample and the camera.

[0060] The method 100 further includes predicting (104) a second digital image in the form of a digital fluorescence representation of the captured first digital image by a trained machine learning system including a trained learning model for predicting a corresponding digital fluorescence representation of the input image.

[0061] The first captured digital image is used as input data or an input image for a trained machine learning system, and further, parameter values ​​of the at least one optical filter have been determined during training of the machine learning system.

[0062] Figure 2 shows an extension of the exemplary embodiment according to Figure 1, in particular the training phase 200 of the above method shown in Figure 1 during its production or prediction phase. In this case, training the composite machine learning model of the trained composite machine learning system comprises providing (202) a plurality of first digital training images of a tissue sample (i.e. again for example brain tissue or cancer tissue), which images are captured under white light by a microsurgery optical system having a second image capture unit, i.e. for example a surgical microscope, and for each first digital training image a second number of color channel information in different spectral ranges is available. This second number is typically higher than the first number of color channels in the productive prediction operation mode according to Figure 1. A hyperspectral camera with for example 64 color channels is advantageously used.

[0063] The training stage 200 further includes providing (204) a plurality of second digital training images, each of which represents the same tissue sample as the first set of digital training images. In this case, the second digital training images include indications of diseased elements of the tissue sample. These indications can take a variety of forms, such as pixel-by-pixel annotations, optically enhanced boxed regions, and / or boundaries or other image segmentations of diseased tissue regions. The representations can be under normal visible light, e.g., under white light, or under special light, and enhanced by contrast agents (fluorescence representations).

[0064] The training stage 200 further comprises the actual training 206 (under so-called supervised learning) of the machine learning system to form a trained machine learning model for predicting digital images of the type of a plurality of second digital training images, i.e. digital images with visible markings of diseased tissue. In this case, the following are used as input values ​​of the combined machine learning system: (i) a plurality of first digital training images in the form of a second number of color channel information, i.e. having a larger number of color channel information, (ii) a plurality of second digital training images as ground truth, i.e. the desired result image to be trained, and (iii) parameter values ​​for reducing the second number of color channel information by at least one digitally simulated optical filter for forming the first number of color channel information. In this case, the parameter values ​​may comprise instructions regarding the number of filters, the color spectrum or the filter shape.

[0065] Further, after reducing the second number of color channel information to the first number of color channel information by a digitally simulated optical filter, the plurality of first digital training images are used as training data for predicting digital images of the type of the plurality of second digital training images at one time without different systems or different machine learning systems being used.

[0066] In addition, after the training of the machine learning system is completed, the parameter values ​​of the at least one optical filter are output as output values ​​of the machine learning system. That is, the learning system learns such parameter values ​​as being properties of the simulated optical filters during the training of the machine learning system to form the machine learning model. The parameter values ​​of the at least one optical filter may relate, for example, to the number of filters or the associated frequency band. As will be explained in more detail in FIG. 3, the properties of the digitally simulated filters can be directly used to generate such optical filters for use in the prediction stage.

[0067] 3 shows a simultaneous representation 300 of the training and prediction stages of a machine learning system as well as the use of filters. The left side of the figure relates to the training stage, whereas the right side of the figure relates to the prediction stage. A number of tissue samples 302 are captured by a digital capture system 304 with multiple color channels 306 to generate training data from the tissues 302. These digital images (first training data) are directed through a digitally simulated filter 308, and downstream of the digitally simulated filter 308 these digital images 310 are made available in a representation with a reduced number of color channels 310.

[0068] These digital images 310 thus reduced in terms of the number of color channels are used as training data for the machine learning system 312. As annotated data for training, marked versions of the digital images of the first training data are provided to the machine learning system 312 as second training data.

[0069] In addition, the properties of the optical filter 308 are altered, thereby influencing the learning speed and learning success (faster or slower convergence) of the machine learning system 312. This results in a joint optimization 314 of the parameters of the digitally simulated optical filter 308 and the machine learning system 312, so that at the end of the training phase a machine learning model is available that can be productively used in the prediction phase (right side of FIG. 3).

[0070] The use phase of the trained machine learning system 324, shown as an example on the right side of FIG. 3, also starts with capturing a biological tissue 302a. This can be done, for example, during surgery or examination, by an image capture unit 320, i.e., for example, an RGB camera. The potentially diseased tissue 302a is illuminated, for example, by white light (illumination source not shown) and converted into a digital image by the capture unit / camera 320 after filtering of the beam by an optical filter 318. This digital image is composed of color information from, for example, 3-4 color channels 322. Ideally, the number of color channels 322 corresponds to the number of color channels 312 generated from the color channels 306 of the hyperspectral camera 304 by the digitally simulated optical filter 310 during training of the machine learning system 312. The equivalent elements of the training and prediction phases are indicated by dashed arrows from the left side of FIG. 3 to the right side of FIG. 3.

[0071] Digital images captured by a camera 320 having a small number of color channels 322 (e.g., an RGB camera) are fed to a trained machine learning system 324, i.e. used as input data, to output a digital output image 326 in the form of a fluorescence representation.

[0072] In this context, it should again be clearly pointed out that while the filter 308 during training of the machine learning system 312 is digitally simulated, the filter 318 used during the prediction phase of the trained machine learning system 324 is a filter that exists in reality, a filter having optical properties or parameters that the trained machine learning system 314 actually outputs at the end of training. As a result, the real optical filter 318 is manufactured exactly according to the specifications ascertained during the joint optimization 314 of the digitally simulated filter 308 and the parameters of the machine learning system 312 in training. Ground truth data 316 is also required as additional input for training.

[0073] 4 shows a diagram of an example embodiment 400 regarding the number of color channels during training. The N color channels 402, which are the result of digital recording (of a digital image), for example of a hyperspectral camera, can be neutralized (i.e., made independent) with respect to the characteristics 404 of the light source, so that these characteristics can also be used as additional constraints during training of the machine learning system 414. Examples of such constraints can be the number, as well as the availability of light at different wavelengths, its density, and the respective intensities.

[0074] The color information available by the same number of color channels 406 for the captured image is then abstracted from the characteristics of the light source used. The constraints of the simulated optical filters 408, which are optimized together with the characteristics of the machine learning system 414 (see correspondingly the machine learning systems 312, 324 in FIG. 3), can then be taken into account. Furthermore, the camera sensitivity parameters (particularly with respect to certain wavelengths) can also be taken into account as additional constraints 410. As a result, a digital image is obtained with a reduced number of color channels 412 that is less (or significantly less) than the number of color channels 402 of the digital input image.

[0075] The double arrows between the symbolically represented constraints 404, 408, 410 represent their influence on the machine learning system 414 during training. Digital training images with a reduced number of color channels 412 are provided as input data to the machine learning system. In supervised learning, reference data (ground truth data) 416 representing the expected results for each digital training image are also provided to the machine learning system 414 during training. These ground truth data 416 ultimately represent the expected output images in a fluorescent representation, which are intended (predicted) in the presence of digital input images with a large number of color channels 402.

[0076] Furthermore, further parameters 418 can be output, which mainly relate to the properties of the digitally simulated optical filter 408, but also for example the spectral range properties of the illumination source for the prediction step, or the spectral range sensitivity parameters of the digital capture unit (i.e. RGB camera, monochrome camera, etc.) used in the prediction step.

[0077] As an example, U-Net can be used as the machine learning system 414. U-Net is composed of a contraction path and an augmentation path, as known, indicated by the symbol 414. U-Net is composed of, for example, two 3×3 convolutions (convolutions without padding) repeatedly applied, each preceded by a ReLU and a stride 2 2×2 pooling operation for downsampling. In this case, the number of feature channels doubles for each downsampling step. On the other hand, the augmentation path is composed of an upsampling of the feature maps, followed by a 2×2 convolution ("upsampling") that halves the number of feature channels in each case, and a concatenation of the feature maps that are "cropped" accordingly by the contraction path, and two 3×3 convolutions, each followed by a ReLU.

[0078] If there are under-represented fluorescent groups, data augmentation techniques (L / R mirror images, top / bottom mirror images, random rotation, random image crop, etc.) can be applied in order to use as many training images in all possible groups with the same intensity as possible. Alternatively, images of anatomical variants can be used.

[0079] FIG. 5 shows two possible wavelength (or frequency) distributions as a result of the joint optimization of a machine learning system (see 312 in FIG. 3) and a digitally simulated filter. Diagram 502 shows essentially three mutually separated frequency / wavelength ranges that can be easily converted into an RGB representation. Diagram 504 has four filter ranges with different attenuation for the different frequency ranges. It should be pointed out that this actually contains only two possible examples where each spectral range can represent a sensor channel of the digital capture unit. The height of the vertical lines corresponds to the weighted sensitivity of the respective filter at the respective wavelength. In this case, an additional constraint of representation 502 concerns that the three channels each have a Gaussian sensitivity distribution with a relatively small standard deviation. Representation 504 has four maxima, of which one can be, for example, in the blue range, one in the green range and two in the red range. Again, the standard deviation is relatively small. If no optimization were performed during the training phase of the machine learning system, the intensity lines of the potential filters would, so to speak, be distributed mathematically more chaotically on the x-axis (not shown), with no discernible distinct spectral ranges that could be translated into actual physical optical filters.

[0080] Ideally, the actual optical filters (see 318 in FIG. 3) would have exactly the same sensitivity curves.

[0081] These three possible physical constraints presented for the filters represent good examples here. In this case, representation 502 represents the fully represented sensitivity specification (i.e. no changes in the filter parameters are made during learning). In this case, representation 504 corresponds to the characteristics (or parameters) of the sensitivity curve of the camera to be used later (during the prediction stage), in other words a Gaussian sensitivity state curve with constraints such as a small standard deviation and a specific value for the highest sensitivity value for a given wavelength. The last case described in the two previous paragraphs can probably only be realized with difficulty in a 1:1 way from a hardware point of view with respect to real physical filters.

[0082] Similarly, in further embodiments, the properties of the light source used can be used as additional constraints during the training phase.

[0083] Further embodiments can provide a temporal sequence of digitally captured images. The aim here is to avoid digital specking ("flicker"), which changes significantly when consecutive images are used. A simple countermeasure that can be provided is post-processing of the individual captured frames (individual images) with softer transitions over time. In further developed forms, the generation of the 3D model can be performed in such a way that in further embodiments all 2D operations of the 2D model are performed in 3D space (i.e. 3D convolution, 3D pooling, etc.). A combination of a 2D model with a time-limited model can also be used as an additional further developed exemplary embodiment.

[0084] 6 shows an exemplary embodiment of a prediction system 600 for predicting a digital fluorescence image. In this case, the prediction system includes a memory 604 for storing a program code and one or more processors 602 connected to the memory and which, when executing the program code, cause the prediction system 600 to control the following units: a first digital image capture unit 606 having a first number of color channel information for capturing a first digital image of a tissue sample by microsurgery optics using white light and at least one optical filter, and a trained machine learning system 608 for predicting a second digital image in the form of a digital fluorescence representation of the captured first digital image. In this case, the trained machine learning system includes a trained learning model for predicting a corresponding digital fluorescence representation of an input image, the first captured digital image being used as an input image of the trained machine learning system, and parameter values ​​of the at least one optical filter having been determined during training of the machine learning system.

[0085] Furthermore, the prediction system 600 may explicitly include, or function as, a second image capture unit 610 for a plurality of first digital training images of the tissue sample captured under white light by the microsurgical optics, where a second number of color channel information in different spectral ranges is available for each first digital training image.

[0086] Further, there may be a providing unit 612 for providing a plurality of second digital training images, the second digital training images in each case representing the same tissue sample as the first set of digital training images, the second digital training images having indications of pathological elements of the tissue sample. The providing unit 612 may be implemented as an image memory.

[0087] Further, the present invention may further comprise a training system 614 for training the machine learning system to form a trained machine learning model for predicting digital images of the type of the plurality of second digital training images. In this case, the plurality of first digital training images in the form of a second number of color channel information, the plurality of second digital training images as ground truth (or ground truth data), and a parameter value for reducing the second number of color channel information by at least one digitally simulated optical filter for forming the first number of color channel information are used as input values ​​of the machine learning system, and the plurality of first digital training images are used as training data for predicting digital images of the type of the plurality of second digital training images after reducing the second number of color channel information to the first number of color channel information by the digitally simulated optical filter. After the training of the machine learning system is completed, the parameter value of the at least one optical filter is further output as an output value of the machine learning system.

[0088] It is explicitly mentioned that the modules and / or units, specifically the processor 602, the memory 604, the first digital image capture unit 606, the trained machine learning system 608, the second image capture unit 610, the providing unit 612, and the training system 614, may be connected by electrical signal lines or via a system internal bus system 616 to exchange signals and / or data and for coordinated behavior.

[0089] Fig. 7 shows a diagram of a computer system 700 that may include at least a part of the prediction system. The embodiments of the concepts proposed herein can in principle be used with virtually any kind of computer, regardless of the platform used therein to store and / or execute the program code. Fig. 6 shows by way of example a computer system 700 suitable for executing the program code according to the method presented herein. A computer system already present in a surgical microscope may also serve as a computer system for implementing the concepts presented herein, possibly with corresponding extensions.

[0090] The computer system 700 has multiple general-purpose functions, where the computer system may be a tablet computer, a laptop / notebook computer, another portable or mobile electronic device, a microprocessor system, a microprocessor-based system, a computer system with specially configured dedicated functions, or a component part of a computer-controlled microscope system. The computer system 700 may be configured to execute computer system executable instructions, such as program modules, executable to implement the functionality of the concepts proposed herein. To that end, program modules may include routines, programs, objects, components, logic, data structures, etc., for implementing particular tasks or particular abstract data types.

[0091] The components of the computer system may include one or more processors or processing units 702, a storage system 704, and a bus system 706 that connects various system components including the storage system 704 to the processor 702. The computer system 700 typically includes multiple accessible volatile or non-volatile storage media accessible by the computer system 700. The storage system 704 may store data and / or instructions (commands) on the storage media in a volatile format, such as a RAM (random access memory) 708, for execution by the processor 702. These data and instructions implement one or more functions and / or steps of the concepts presented herein. Further components of the storage system 704 may be a persistent memory (ROM) 710 and a long-term memory 712 that may also store program modules and data (reference numeral 716) as well as workflows.

[0092] The computer system includes several dedicated devices for communication purposes, such as a keyboard 718, a mouse pointing device (not shown), a screen 720, etc. These dedicated devices can also be combined with a touch-sensitive display. A separately provided I / O controller 714 ensures smooth data exchange with external devices. A network adapter 722 can be used for communication over a local or global network (LAN, WAN, e.g. via the Internet). The network adapter can be accessed by other components of the computer system 700 via the bus system 706. In this case, it will be understood that other devices, not shown in the drawings, can also be connected to the computer system 700.

[0093] Additionally, at least a part of a prediction system 600 (see FIG. 6) for predicting a digital fluorescence image can be connected to the bus system 706. The prediction system 600 and the computer system 700 can optionally share a memory and / or a processor.

[0094] The description of various exemplary embodiments of the present invention is provided for the purpose of improving understanding, but does not directly limit the concept of the present invention to these exemplary embodiments. Further modifications and alterations may be developed by those skilled in the art. The terms used in this specification are selected to best explain the basic principles of the exemplary embodiments and to allow those skilled in the art to easily understand them.

[0095] The principles presented herein may be embodied as any of a system, method, combination thereof, and / or computer program product, which may include one or more computer readable storage media containing computer readable program instructions for causing a processor or control system to implement various aspects of the invention.

[0096] As a medium, electronic, magnetic, optical, electromagnetic, or infrared media, or semiconductor systems, such as SSD (Solid State Device / Drive as solid state memory), RAM (Random Access Memory) and / or ROM (Read Only Memory), EEPROM (Electrically Erasable Read Only Memory), or any combination thereof, are used as a transmission medium. Suitable transmission media also include propagating electromagnetic waves, electromagnetic waves in a wave guide or other transmission medium (e.g. light pulses in an optical cable), or electrical signals transmitted in an electrical wire.

[0097] A computer readable storage medium may be an embodied device that holds or stores instructions for use by an instruction execution device. The computer readable program instructions described herein may also be downloaded onto a compatible computer system as an app (for example, on a smartphone) from a service provider via a cable connection or a mobile wireless network.

[0098] The computer readable program instructions for carrying out the operations of the invention described herein may be machine dependent or machine independent instructions, microcode, firmware, status definition data, or any source or object code written in a conventional procedural programming language such as C++, Java, or the like, or in a programming language such as the programming language "C". The computer readable program instructions may be fully executed by a computer system. In some exemplary embodiments, there may also be electronic circuitry, such as a programmable logic device, field programmable gate array (FPGA), or programmable logic array (PLA), that executes the computer readable program instructions by using status information of the computer readable program instructions to configure or individualize the electronic circuitry in accordance with aspects of the invention.

[0099] The invention presented herein is further illustrated with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to exemplary embodiments of the invention. It should be noted that substantially any block of the flowcharts and / or block diagrams may be embodied as computer readable program instructions.

[0100] The computer readable program instructions may be provided to a general purpose computer, a special purpose computer, or an otherwise programmable data processing system to create a machine whose instructions, executed by a processor or computer or other programmable data processing apparatus, generate means for implementing the functions or procedures shown in the flow charts and / or block diagrams. These computer readable program instructions may also be stored on a computer readable storage medium corresponding thereto.

[0101] In this sense, any block in the illustrated flow charts or block diagrams may represent a module, segment, or portion of instructions that represent multiple executable instructions for implementing certain logical functions. In some exemplary embodiments, the functions represented in individual blocks may be implemented in different orders, optionally even in parallel.

[0102] The illustrated structures, materials, sequences, and equivalents of all of the means and / or steps having the relevant functions in the appended claims are intended to cover all of the structures, materials, or sequences expressed by the claims. [Explanation of symbols]

[0103] 100 ways 102 Method 100 Steps 104 Method 100 Steps 200 Extended Methods 100 202 Method 200 Steps 204 Method 200 Steps 206 Steps of Method 200 300 Comparison of training and prediction stages 302 Organization 304 Hyperspectral camera, capture system 306 Color Channels 310 Reduced number of color channels 312 Machine Learning Systems in Training 314 Simultaneous Optimization 316 Ground Truth Data 318 Physical Filters 320 Image capture unit, camera 322 Color Channels 324 Learning System 326 Output Images 400 Exemplary Embodiments 402 Color Channels 404 Constraints, characteristics 406 Color Channels 408 Filters, Constraints 410 Filters, Constraints 412 Reduced number of color channels 414 Learning System 416 Ground Truth Data 418 More output parameters 502 Three mutually separated frequency / wavelength ranges 504 Representation with four wavelength maxima 600 Prediction System 602 Processor 604 Memory 606 Image capture unit 608 Learning System 610 Image capture unit 612 units offered 614 Training System 616 Bus System 700 Computer Systems 702 Processor 704 Storage System 706 Bus System 708 RAM 710 ROM 712 Long-term Memory 714 I / O Controller 716 Program Modules and Data 718 Keyboard 720 screen 722 Network Adapter

Claims

1. A computer-implemented method (100) for predicting a digital fluorescence image (326), comprising: - Capturing a first digital image of a tissue sample (302a) by a microsurgical optical system having a first digital image capture unit (304) with a first number (322) of color channel information using white light and at least one optical filter (318) (step 102); - Predicting a second digital image (326) in the form of a digital fluorescence representation of the captured first digital image by a trained machine learning system (324) including a trained learning model for predicting the corresponding digital fluorescence representation of the input image (step 104); comprising: - The first captured digital image is used as the input image of the trained machine learning system (324); - The parameter values of the at least one optical filter are determined during the training of the machine learning system (324). A computer-implemented method (100).

2. Training the learning model of the trained machine learning system comprises: - Providing a plurality of first digital training images of a tissue sample (302) captured under white light by a microsurgical optical system having a second image capture unit (304), wherein a second number (306) of color channel information in a different spectral range is available for each first digital training image (step 202); - Providing a plurality of second digital training images (316) each representing the same tissue sample (302) as the first set of digital training images, wherein the second digital training images have an indication of a lesion element of the tissue sample (step 204); - Training the machine learning system (312) to form the trained machine learning model for predicting the type of digital image of the plurality of second digital training images, wherein as the input value of the machine learning system, - the plurality of first digital training images in the form of the second number (306) of the color channel information, - the plurality of second digital training images as ground truth (316), - a parameter value for reducing the second number (306) of the color channel information by at least one digitally simulated optical filter (308) for forming the first number of the color channel information is used, a step (206) of training including, After reducing the second number (306) of the color channel information to the first number (310, 322) of the color channel information by the digitally simulated optical filter (308), the plurality of first digital training images are used as training data for predicting the digital images of the type of the plurality of second digital training images, After the training of the machine learning system is completed, at least a part of the parameter values of the at least one optical filter is output as an output value of the machine learning system, The method (100) according to claim 1.

3. The method (100) according to claim 1, wherein the parameter value of the at least one optical filter (318) includes the number (310, 322) of the first color channel information and / or the filter shape of the digitally simulated optical filter (308).

4. The method (100) according to claim 1, wherein the second number of the color channel information exceeds the first number of the color channel information.

5. The method (100) according to claim 1, wherein the parameter value for reducing the second number (306) of the color channel is at least one selected from the group consisting of a filter shape and a center frequency of each of the first numbers (310, 322) of the color channel information.

6. The parameter value for controlling the light source of the white light during the capture of the first digital image is generated as an additional output value of the machine learning system (312) after the training of the machine learning system (312) has ended, according to the method (100) of claim 2.

7. The digital fluorescence representation corresponds to a representation generated using a light source within the UV range, according to the method (100) of claim 1.

8. The learning model corresponds to an encoder-decoder model in terms of its setup, according to the method (100) of claim 1.

9. The encoder-decoder model is a convolutional network in the form of a U-Net architecture, according to the method (100) of claim 1.

10. A prediction system (600) for predicting a digital fluorescence image (326), - A memory (604) for storing program code and one or more processors (602) connected to the memory (604) comprising, When the one or more processors (602) execute the program code, the prediction system (600) is caused to have the following units, namely - A first digital image capture unit (606) having a first number (306) of color channel information for capturing a first digital image of a tissue sample (302a) by means of a microsurgical optical system using white light and at least one optical filter (318), - A trained machine learning system (326) for predicting a second digital image (326) in the form of a digital fluorescence representation of the captured first digital image, the trained machine learning system (326) comprising a trained learning model for predicting the corresponding digital fluorescence representation of the input image to be controlled, - The first captured digital image is used as the input image of the trained machine learning system (326), - The parameter values of the at least one optical filter (318) are determined during the training of the machine learning system (312), Prediction system (600).

11. A computer program product for predicting a digital fluorescence image, comprising a computer-readable storage medium including program instructions stored thereon, the program instructions being executable by one or more computers or control units, and causing the one or more computers or control units to execute the method according to any one of claims 1 to 9.