Cell type classification based on cell location information

Through a two-stage processing method, the microscopic image data with nonspecific and specific contrast is solved, and the problem of time-consuming and poor generalization of training data set generation in the prior art is achieved, and a cell type classification with low computational density and good generalization is achieved.

CN120071335APending Publication Date: 2025-05-30CARL ZEISS MICROSCOPY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411672222.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Prior Art In the technology of automatically evaluating cell arrangement in microscopic images, generating training data sets for training machine learning algorithms is time-consuming and requires a lot of annotation work, and is poorly generalized, making it difficult to adapt to different imaging parameters and unknown cell types.

Method used

Through a two-stage processing method, cell instances are firstly located using image data with nonspecific contrast, and then cell types are determined using image data with specific contrast, and distribution complexity is processed in two independent working steps.

Benefits of technology

A low computational density and good generalization of cell type classification is achieved, enabling robust determination of cell types and adapting to different imaging parameters and emerging cell types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071335A_ABST
    Figure CN120071335A_ABST
Patent Text Reader

Abstract

The present disclosure relates to techniques based on a combination of image data (210) having a cell type non-specific contrast, optionally with coloration, and additional image data (230) having a cell type specific contrast. Individual cells may be localized (220) based on image data having a cell type non-specific contrast. Based on the additional image data, distinguishing between different cell types can be made, i.e., making it possible to classify (245) cell types and in particular cell models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Aspects of the present disclosure relate to techniques for classifying cells of different cell types based on microscopic images. Background Art

[0002] In the fields of biology and medicine, microscopes are used to study biological samples. For example, microscopic image analysis is used to identify and classify individual cells to obtain relevant information about their structure, function, and potential abnormalities.

[0003] Techniques for automatically evaluating cell arrangements depicted in microscopic images are known. In particular, techniques for distinguishing between different types of cells are known.

[0004] In a variant of the prior art, cell type classification information is generated based on image data that does not have cell type-specific contrast. For example, see Meng, N., Lam, E.Y., Tsia, K.K., & So, H.K.H. (2018), Large-scale multi-class image-based cell classification with deep learning, IEEE Journal of Biomedical and Health Informatics, 23(5), 2091-2098. It uses a machine learning algorithm that takes image data as input and outputs semantic cell instance segmentation. Thus, the machine learning algorithm can infer cell types only based on the morphological characteristics of different types of cells, which are reflected in the different appearances of different types of cells in the image data. The disadvantage of this technique is that it is very time-consuming to generate the corresponding training data set for training the machine learning algorithm. Since individual cells must be manually located and classified, a particularly large amount of annotation work is required. It has also been found that the generalization of such machine learning algorithms is very poor, i.e., limited to specific imaging methods with specific imaging parameters. For example, if the magnification of the microscopic image changes or cells of another previously unknown cell type appear, the corresponding machine learning algorithm will produce poor results. In addition, the corresponding machine learning algorithm is complex, i.e., has many parameters. Therefore, the derivation is computationally intensive and requires a large amount of running time in edge applications. Such techniques generally cannot be implemented in real time. Summary of the Invention

[0005] There is a need for improved techniques to determine cell type classification information for cells in microscopic images. Techniques are needed that at least solve or mitigate some of the above disadvantages and limitations. Techniques are needed for determining cell type classification information with low computational intensity and good generalization so that cell type classification information can be robustly determined.

[0006] The above object is achieved by the features of the independent claims. The features of the dependent claims define embodiments.

[0007] The following describes a technique for classifying cells in a cell arrangement. The method optionally includes obtaining a microscopic image suitable for estimating cell size. Then, size estimation can be performed based on the microscopic image. The method also includes obtaining first image data that shows cells using a non-specific contrast method. The method includes localizing cell instances, for example, the size information obtained from the size estimation can be used to localize cell instances. The method also includes obtaining second image data that shows cells using a specific contrast method. The method includes classifying the cell type of each cell position determined based on the localization using the size information obtained from the size estimation. The method includes using the obtained information to, for example, generate a histogram, generate feedback for a workflow, or issue a warning to a user.

[0008] A computer-executed method is disclosed. The method includes acquiring first image data. The first image data microscopically depicts an arrangement of cells with a non-cell-type-specific contrast. The method also includes obtaining second image data. The second image data microscopically depicts an arrangement of cells with a multi-cell-type-specific contrast. The second image data is obtained by cell-type-specific coloring of the cells. The method also includes determining position information of the cells based on the first image data. The method also includes determining cell-type classification information of the cells based on the second image data and the position information.

[0009] A data processing device is disclosed. It is provided to execute such a program. For example, a processor can load and execute program code from a memory.

[0010] A program code is also disclosed. The program code can be loaded and executed by a processor. This causes the processor to execute the computer-executed method described above.

[0011] A system is disclosed, including a data processing device and a microscope, which is configured to obtain image data such as first image data and second image data.

[0012] A training data set is disclosed. The training data set includes a plurality of input-output data pairs, where the output data of the input-output data pairs is determined based on cell-type classification information.

[0013] Without departing from the scope of the present disclosure, the features set forth above and the features described below can be used not only in the explicitly described corresponding combinations, but also in further combinations or individually.

[0014] The above features and advantages of the present disclosure, as well as the ways and types of achieving these features and advantages, will be more clearly understood in connection with the following description of exemplary embodiments that will be explained in more detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A flowchart of an exemplary method according to an embodiment of the present disclosure is shown.

[0016] Figure 2 An exemplary data processing flow according to an embodiment of the present disclosure is shown.

[0017] Figure 3 An exemplary data processing apparatus according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0018] The above features and advantages of the present disclosure, as well as the ways and types of achieving these features and advantages, will be more clearly understood in connection with the following description of exemplary embodiments that will be explained in more detail with reference to the accompanying drawings.

[0019] The present disclosure is explained in more detail below with reference to preferred embodiments using the accompanying drawings. In the drawings, the same reference numerals denote the same or similar elements. The drawings are schematic diagrams of various embodiments of the present disclosure. The elements shown in the drawings are not necessarily drawn to scale. Instead, the various elements shown in the drawings are reproduced in a manner that those skilled in the art can understand their functions and general purposes. The connections and couplings between the functional units and elements shown in the drawings can also be implemented as indirect connections or couplings. The connections or couplings can be implemented in a wired or wireless manner. The functional units can be implemented as hardware, software, or a combination of hardware and software.

[0020] Techniques for automatically evaluating microscope images depicting cell arrangements are disclosed below. Specifically, techniques for determining cell type classification information of imaged cells are disclosed.

[0021] The cell type classification information indicates the cell type of each individual cell in the arrangement. This means that each imaged cell in the arrangement is assigned an associated cell type.

[0022] Different cell types can be represented, for example, by different types of cells such as nerve cells, muscle cells, blood cells, suspension cells, sperm, stem cells, etc.

[0023] Different cell types can also represent different cell lines such as LOVO (human colon cancer cell line), COS7 (monkey kidney-derived cell line), HeLa Kyoto (human cancer cell line), LCPK (porcine kidney cell line), U2OS (human osteosarcoma cell line), etc. A cell line is a group of cells cultured (grown) under laboratory conditions that can reproduce by continuous division.

[0024] However, different cell types can also be represented by cells of the same type but in different life stages or states, such as live cells and dead cells. Different cell states include, for example: dividing, dead, surviving, about to die, mutated, detached and adherent, transfected and non-transfected, etc.

[0025] Different cell types can have different sizes.

[0026] Different cell types can also be related to each other in a predefined hierarchy. For example, at the highest level of a predefined hierarchy, cell types "normal" and "detached" can be distinguished; at the next lower level, "dead" and "mitotic" of the "detached" cell type can be distinguished, or different cell cycle stages of live cells can be distinguished; at the next lower level, "necrotic" and "apoptotic" of "dead" can be distinguished. This is just an example, and other cell types and other hierarchies can also be envisioned. Different cell types can be, for example: detached and non-detached; alive and dead; mitosis and apoptosis.

[0027] This hierarchy can also be reflected by cell type classification information. The cell type classification information can be output, for example, as a tree data structure to represent the hierarchy.

[0028] All kinds of examples are based on the understanding that different cell types have significant differences in size, morphology, staining behavior, etc. In principle, these characteristics are very suitable for classifying cells into different cell types; however, it has been found that these different characteristics are usually reflected in different appearances in microscope images, so that it may be difficult to detect cells (without cell type classification).

[0029] To solve this problem, according to the present disclosure, cells are classified (i.e., cell type classification information is determined) in the following manner: (i) First, all cells in the first image data are located (but not classified) based on cell type non-specific contrast; then (ii) based on the second image data, the cell type classification information of each located cell is determined, where the second image data has cell type specific contrast. The second image data is usually obtained by cell type specific staining to achieve cell type specific contrast.

[0030] This two-stage process (first, determining the position information of cells based on the first image data; second, determining the cell type classification information based on the second image data and the position information) distributes the complexity of cell separation and cell type assignment over two independent work steps, which is more robust than a solution with only one work step.

[0031] Figure 1 A flowchart of an exemplary method of an embodiment of the present disclosure is shown. Figure 1The method can determine the cell type classification information of cells in a cell arrangement. Figure 1 The method can be executed, for example, by a processor that loads and executes corresponding program code from a memory.

[0032] In block 905, first image data is received. The first image data microscopically depicts the arrangement of cells with cell type non-specific contrast.

[0033] For example, the first image data can include one or more microscopic images. One or more microscopic images can include one or more channels. One microscopic image can have multiple channels; or multiple channels can be assigned to different microscopic images.

[0034] Block 905 can include loading the first image data from a database or a memory.

[0035] Block 905 can include controlling a microscope to obtain the first image data and obtaining the first image data from the microscope.

[0036] One or more of the following imaging methods can be used to obtain the first image data: autofluorescence, transmitted light imaging, wide-field imaging, DIC phase contrast, Zernike phase contrast, digital phase contrast, overview contrast, for example using a dedicated overview camera. Digital phase contrast includes, for example, using multiple illumination directions and combining the corresponding images.

[0037] The first image data can be two-dimensional (2-D) or three-dimensional (3-D).

[0038] The first image data can include a time series of multiple microscopic images. For example, a time series of 2-D microscopic images or a time series of 3-D microscopic images can be obtained.

[0039] The first image data can be obtained without staining the cells. However, the first image data can also be obtained by staining the cells, for example, with cell type non-specific staining of the cells.

[0040] Generally, stained cells can be observed using fluorescence imaging. For staining, the cells are first grown on a suitable substrate such as a coverslip. Then they are fixed with a fixative (usually formaldehyde) to preserve their structure. After fixation, the cells are permeabilized to make the cell membrane permeable to the dye. A blocking solution is usually used to block non-specific binding sites, thus minimizing background fluorescence. Then the cells are incubated with a specific primary antibody that binds to the desired target structure. After a washing step, a fluorescently labeled secondary antibody is added, which specifically binds to the primary antibody. This is the dye. Again, washing removes the excess dye and unbound antibody.

[0041] For example, the DAPI dye can be used as a dye for non-cell-type-specific staining. DAPI (4',6-diamidino-2-phenylindole) is a fluorescent dye that binds to double-stranded DNA. The DAPI dye emits almost no fluorescence when unbound. However, once bound to DNA, it produces strong fluorescence with a maximum wavelength of approximately 461 nm. This enables the observation of cell nuclei under a fluorescence microscope. Since the DAPI dye mainly attaches to the cell nuclei, different types of cells can be well localized with the help of the DAPI channel, but can only be very poorly distinguished or not at all.

[0042] There are also other dyes for non-cell-type-specific cell nucleus staining. Examples include, for example, Hoechst 33342 and Hoechst 33258.

[0043] In block 910, second image data is received. The second image data is microscopically depicted in a plurality of cell-type-specific contrasts.

[0044] For example, the second image data can include one or more microscopic images. One or more microscopic images can include one or more channels. The second image data can be a multi-channel image (e.g., different wavelengths for different labels) or a collection of single images (if each cell type is stained separately, one wavelength is sufficient). The second image data can include a time series.

[0045] Block 910 can include loading the second image data from a database or memory.

[0046] Block 910 can also include controlling the microscope to obtain the second image data and obtaining the second image data from the microscope.

[0047] For example, the second image data can be obtained by non-cell-type-specific cell staining. Fluorescent staining can be used. For example, special fluorescent labels that specifically bind to different cell types (or their cell organelles) can be used.

[0048] This cell-type-specific staining enables cells of different cell types to be imaged with distinguishing features, thereby differentiating the individual cell types.

[0049] Alternatively, as an alternative or supplement to this cell-type-specific staining, it is also conceivable to use special imaging methods, for example, which make the differences in cell morphology particularly obvious.

[0050] In principle, it is conceivable to obtain the first image data in block 905 and the second image data in block 910 simultaneously. For example, the first image data may correspond to one or more first channels in image acquisition, and the second image data may correspond to one or more second channels in image acquisition. Thus, it is conceivable to control the microscope to obtain the first image data and the second image data simultaneously. If the first image data and the second image data are obtained using the same imaging process, there may be an inherent registration between the first image data and the second image data.

[0051] In block 915, third image data may optionally be obtained. For example, the third image data may be obtained without staining. The third image data may serve as a reference.

[0052] The image data of block 905, block 910, and block 915 all represent the same cell arrangement. This means that one imaging region of the sample is the same (different image data may also image other edge regions, but at least image the common region). In some embodiments, registration may be performed between the different image data of block 905, block 910, and block 915. In this way, a pixel-to-pixel mapping between the image data can be obtained. This is shown in block 916. However, such registration is usually optional. In some cases, the image data may already be (inherently) registered with each other, for example because they are obtained using the same imaging optics.

[0053] In block 920, size information may optionally be determined and / or the size information (e.g., based on the size information) may be scaled to a standard imaging scale.

[0054] For example, the size information may be determined as the imaging scale of the cells.

[0055] Then, the size information can be used as an input to a scaling algorithm (or used later in the method). The scaling algorithm can also be trained as an image-to-image model.

[0056] Block 920 may include: scaling the first image data and / or the second image data based on an analysis of at least one of the first image data of block 905, the second image data of block 910, or additional image data depicting the cell arrangement (e.g., the third image data of block 915), such that they depict the cells according to a reference scale. The first image data of block 905 may be scaled individually or together with the second image data of block 910. Then subsequent blocks use the scaled image data.

[0057] Details related to the size information are described below. The size information may indicate the size of the cells in the image data (e.g., "average diameter of 10 pixels"). The size information may indicate the imaging scale of the cells.

[0058] In one example, for the first image data of box 905 and the second image data of box 910, the size information in box 920 can be determined respectively. This means that the size information is determined twice, once for the first image data of box 905 and once for the second image data of box 910. In another example, for the first image data of box 905 and the second image data of box 910, the size information can also be determined only once and together.

[0059] If the size information is determined in box 920, the size information can be determined in a spatially resolved manner. This means that, for example, different values of the size information will be determined at that location in the microscope image. This technique is particularly useful when the cell arrangement includes multiple cell types with significantly different sizes. Then, different image regions can be associated with different size information.

[0060] To effectively handle variations in different cell sizes, scaling the image data to a uniform cell size has proven helpful. This means that the complexity of the subsequent localization and / or classification algorithms (boxes 925 and 930) can be significantly reduced, and their robustness can be significantly improved. By reducing the complexity of the algorithms used, the techniques described herein can achieve real-time performance.

[0061] In box 920, there are different ways to determine the size information and / or perform the scaling.

[0062] The first variant is based on heuristic analysis. For example, it can be envisioned to use a corresponding algorithm to find contours in the image data used (which can be the first image data, the second image data, or other image data depicting the cell arrangement). Then a predefined shape, such as an ellipse, can be fitted to the found contour. Then the evaluation of the size distribution of the predefined shape can be analyzed. For example, the distribution of the ellipse diameter or radius can be analyzed. For example, the diameters of the ellipse can be averaged. In this way, the imaging scale of the cells can be determined, and then the scaling factor for scaling to the nominal imaging scale can be determined. In the block heuristic analysis, the size information can be determined in a spatially resolved manner, as described above.

[0063] In another variant, as an alternative or addition to such heuristic analysis of the corresponding image data, a machine learning algorithm is applied to the corresponding image data. The machine learning algorithm can be, for example, a classifier or an (ordinal) regression model that determines size information (such as a scaling factor). For example, such a machine learning algorithm can make inferences on a block-by-block basis, whereby size information can be determined in a spatially resolved manner. However, it is also conceivable that the machine learning algorithm uses an image-to-image transformation to directly output a scaled image (without explicitly outputting the scaling factor of the imaging ratio or other size information as an intermediate result). This can also be referred to as domain transformation.

[0064] In block 920, as an alternative or complement to determining size information or scaling, other normalizations can also be performed on the first image data and / or the second image data. For example, normalization can be performed to promote a uniform appearance of the first image data and the second image data or their channels. For example, normalization regarding noise, contrast, and / or brightness, etc., can be performed. Edge regions can be trimmed.

[0065] In block 925, the position information of the arranged cells is determined based on the first image data.

[0066] Determining the position information can include applying one or more image processing algorithms to the first image data (which has optionally been scaled in block 920).

[0067] Such image processing algorithms can be selected from the following group: smoothing operations (such as low-pass filters); threshold operations for determining segments (such as the Otsu algorithm); morphological operations for determining segments (such as closing holes, removing smaller islands, etc.); single-segment filtering.

[0068] For example, single-segment filtering can be performed using so-called "contour finding" or "blob detection" or "ellipse fitting".

[0069] Therefore, in this technique, smoothing can be performed first, for example, by using a low-pass filter, which reduces image noise. After reducing the image noise, a threshold operation can be applied to the image contrast. An example of this is the so-called Otsu algorithm. For example, a binary mask can be created that includes all pixels with a contrast value greater than or less than a predetermined threshold. Pixel clusters can then be formed in this mask. Then, individual segments can be identified in this mask, for example, by searching for specific contours or by identifying pixel aggregations with a specific size and / or shape (such as an ellipse). Optionally, filtering can then be performed based on size and / or shape, and this filtering can optionally be supported by context information regarding the arrangement of the cells. For example, such context information can be information about the cell size (such as the size information of box 920). Such context information can alternatively or additionally relate to pixel resolution. A set of individual segments is thus obtained. Then, the center of each segment can be determined for centering. A bounding box can also be created.

[0070] Determining the position information can include applying one or more image processing algorithms to the first image data, where the one or more image processing algorithms have parameters of machine learning.

[0071] Therefore, a machine learning algorithm (such as an artificial neural network) trained to provide the position information as an output can be used. Examples include single-stage or multi-stage detectors, point segmentation models, or image-to-image models (providing a density map of cell positions). For example, an R-CNN mask can be used, as described by He, Kaiming, et al., "Mask r-cnn", Proceedings of the IEEE International Conference on Computer Vision, 2017. For example, the training data for the corresponding machine learning model can be created during a manual labeling process. For example, the training data can have corresponding reference image data as input, which is acquired corresponding to the first image data (i.e., using the same or sufficiently similar imaging method). Then, an expert can create labels, such as locating the center of a cell or a cell nucleus, or other reference structures of the cell, or the cell boundary. Then, this training data can be used to train such a machine learning model.

[0072] For example, it has been shown that when obtaining the first image data by using nuclear staining, using a (for example, heuristically parameterized) threshold operation to determine the segments (associated with cells) in the first image data is particularly effective. The cell nucleus roughly corresponds to a point position of the cell, and such a threshold operation can be used to filter it out. On the other hand, it has been shown that using an image processing algorithm of machine learning to determine the position information is particularly helpful for the first image data obtained without staining. For example, a machine learning algorithm can be used to provide a reliable determination of the position information for the first image data with phase contrast.

[0073] Depending on the information content of the first image data or on the algorithm used to determine the position information, the position information itself can take different forms. For example, the position information can be selected from the following groups: central position; cell density map; cell instance segmentation; bounding box; density map, point position, cell center or a list of coordinates of cell centers, etc.

[0074] The box 925 can use the dimension information in the box 920. For example, the dimension information can be used to determine the parameters of the positioning algorithm. For example, for the DAPI blob detector, the reference blob size can be determined according to the dimension information.

[0075] Alternatively or additionally, an appropriate algorithm can be selected based on the dimension information of the box 920. For example, there may be multiple machine learning algorithms for different cell sizes; then the algorithm to be used can be selected according to the dimension information.

[0076] In the box 930, the cell type classification information is determined. This determination is based on the second image data in the box 910 (possibly scaled in the box 920). In the box 930, the result of the box 925 is preferably used for the determination. The box 930 can also be performed based on the dimension information of the box 920. The image processing algorithm used in the box 930 can be heuristic or machine learning.

[0077] The cell type classification information can include a list of the cell types of each cell. The cell type classification information can include, for example, a semantic instance segmentation map. The cell type classification information can include, for example, the probability distribution of the softmax function. The cell type classification information can be probabilistic. However, the cell type classification information can also be non-probabilistic, for example, specifying the cell type with the highest probability as the class label.

[0078] Determining the cell type classification information can include, for example, applying a second image processing algorithm to the second image data in the box 910, where the second image processing algorithm selectively processes the image contrast values of the second image data with multiple cell type-specific contrasts in the region determined based on the position information. This is an example of heuristic image processing.

[0079] For example, the dimensions of these regions can be determined based on the dimension information of the box 920. More generally, this means that it is conceivable that the second image processing algorithm for determining the cell type classification information is parameterized based on the dimension information of the box 920.

[0080] These techniques are based on the recognition that size information can provide an indication of the typical range of cells in the second image data. Thus, by sizing regions based on the size information, it can be ensured that the corresponding image patches each depict an entire cell, but only a limited region surrounding the cell. Therefore, the pixel values in the image patches corresponding to the regions are characteristic of the corresponding cell type. For example, the second image processing algorithm can classify based on the measured intensity values or statistics of pixel values with different contrasts within the region. This is based on the recognition that by using cell type-specific contrasts in the second image data, different contrasts each have significantly different pixel values in the region, depending on the cell type present. For example, a contrast sensitive to the "type A" cell type typically has predominantly bright pixels in a region centered on the cell center of the corresponding cells of the "type A" cell type; while the same contrast typically has predominantly dark pixels in another region centered on the cell center of the cells of the "type B" cell type.

[0081] A "region of interest" (ROI) can be determined in block 930. The ROI can be centered on the point location in block 925. For example, the point location can indicate the cell nucleus or the cell center. Then, these regions represent image chunks or patches that have been processed to determine cell type classification information.

[0082] Here, the size of the region can be determined based on, for example, the size information of the cells in block 920. Such size information indicates, for example, the imaging scale of the cells in the second image data.

[0083] Machine learning algorithms can also be used to determine cell type classification information. It can, for example, receive the second image data as input, or receive image chunks or patches of the second image data as input. The image chunks or patches of the second image data correspond, for example, to the regions or ROIs determined based on the size information as described above.

[0084] The machine learning algorithm can also receive location information as a further input. Optionally, the machine learning algorithm can receive size information (e.g., of block 920) as a further input.

[0085] The machine learning algorithm can be, for example, a classifier. It can be trained using traditional training based on manually generated training data. For example, when there is a cell type hierarchy, a multi-stage or hierarchical training process can be used.

[0086] The machine learning algorithm can be trained and applied based on image patches; or based on a "fully convolutional network", which means the entire image is processed. A hybrid form can also be adopted, where the model is trained in a block form and then fully convolutional inference is performed.

[0087] The model architectures of machine learning algorithms can have various implementation methods, such as convolutional networks and transform networks.

[0088] Convolutional networks use convolutional layers to identify local features of the input image. Each convolutional layer consists of a series of learnable filters that move over the input image. The output of this layer is called the feature map, which is generated by the convolution operation between the filter and the input image. Pooling layers usually aggregate the statistical features of small adjacent regions through max-pooling or averaging, thereby reducing the dimension. The fully connected layer, usually located at the later stage of the convolutional network, allows combining features for final classification or regression. Activation functions (such as ReLU) are applied after the convolutional layer or the fully connected layer to introduce non-linearity into the network. Batch normalization and Dropout are techniques commonly used to stabilize training and avoid overfitting. Convolutional networks are usually trained using backpropagation and optimization methods such as SGD or Adam. Loss functions such as cross-entropy measure the difference between the predicted output and the actual output. The ultimate goal of the convolutional network is to learn hierarchical features from simple to complex structures and effectively extend them to new unknown data. To solve the classification output problem, the following operations are performed: after the input image passes through multiple convolutional layers, pooling layers, and fully connected layers, the output layer with the softmax function generates the probability of each class. The class with the highest probability is used as the network's prediction for the given input image.

[0089] The transform architecture is a neural network model that relies on self-attention-based mechanisms to process sequences. It contains multiple encoder blocks and decoder blocks, where each block contains multiple self-attention-based layers and feed-forward networks. Self-attention allows the model to weight different positions of the input sequence according to the meaning and context of the input sequence. Positional encoding is added to the input data to consider the ordering information because the transform is inherently position-independent.

[0090] For example, image patches can be used as the input image of the corresponding artificial neural network. These image patches can be determined based on the region determined using the position information of box 925 and the size information of box 920. Alternatively, the position information can also be passed as a further input into the convolutional network, and the convolutional network can process the entire image.

[0091] For example, based on the size information of box 920, a specific machine learning algorithm can be selected from multiple candidate algorithms, which include various different architectures.

[0092] In box 935, the information obtained from the previous boxes can be used, especially the cell type classification information and / or position information. There are various ways to use such data.

[0093] In block 935, the analysis of cell arrangements can be performed based on cell type classification information. For example, performing the analysis can include one or more of the following evaluations: determining the cell count; determining the cell density; determining the occupied area; determining the neighborhood relationships between cells.

[0094] For example, feedback can be provided to the user. For example, a list of the found cell types can be output in combination with quantities, densities, confluences, etc. A control signal can be provided. For example, an operation can be triggered in an automated workflow. A measurement report can be created.

[0095] The following information can be determined: information about the presence of specific cell types; statistics (such as the frequency distribution of cell types); detecting local density changes of all or individual cell types (which can also be combined with the results of adjacent images); detecting "migratory movements" when classification is repeatedly applied at multiple time points; etc.

[0096] Thus, typical results are both single-image statistics (where in the local position of the sample are which cell types located?) and time-series statistics (in a time series, where do certain cell types locally move / aggregate in the sample and what happens to them?).

[0097] A special form of application is to use the results to generate a training dataset to train another machine learning algorithm. This is shown in block 940.

[0098] The training dataset includes various input-output data pairs. The training dataset is used to train a machine learning algorithm, such as a neural network. The input data of these data pairs is input into the network to generate predictions. The difference between the predictions of the network and the output data (labels or ground truth) of the training dataset is measured as an error or loss. To minimize this error, optimization techniques are applied. A common method is gradient optimization, where the error gradient is calculated based on the network parameters. Using backpropagation, this gradient is propagated backward through the network to update the individual weights. The optimization process continuously adjusts the network parameters to reduce the overall error of the training dataset. After multiple iterations or epochs, the network is continuously optimized until it exhibits satisfactory performance.

[0099] Next, the details related to constructing the training dataset are described.

[0100] First variant: For example, output data of the input-output data pairs of the training dataset can be generated based on cell type classification information. The input data of the data pairs can include, for example, first image data and optional position information of the box 925. Thus, a machine learning algorithm can be trained to predict cell type classification information based on the first image data. Then, such a machine learning algorithm can determine cell type classification information based on relatively simple first image data (without any cell type-specific coloring). In the future, second image data will no longer be required.

[0101] In another variant, the training dataset can include third image data of the box 915 as input data. The output data is formed by the position information of the box 925. The third image data is obtained without any coloring of the cells. This imaging method without using coloring is particularly easy and fast to perform compared to imaging methods using coloring. For example, phase contrast (such as digital phase contrast or Zernike phase contrast) can be utilized to obtain the third image data. Then, a machine learning algorithm can be trained to predict the position information based on the image data corresponding to the third image data. On the other hand, the position information of the box 925 can be determined based on the first image data of the box 905, which is obtained using an imaging method with cell type-nonspecific coloring. This means that the output data of the training dataset is based on an imaging method with coloring, which means that labels can be created particularly reliably. However, in a further iteration of this method, the newly trained machine learning algorithm can be used to determine the position information in the box 925 based on a simple imaging method without the need for coloring. There is no longer a need to use colored, cell type-nonspecific image data. This simplifies the imaging.

[0102] In yet another variant, the training dataset can include third image data of the box 915 as input data. The output data is formed by the cell type classification information of the box 930. The third image data can be obtained without coloring the cells. This imaging method without using coloring is particularly easy and fast to perform compared to imaging methods using coloring. For example, phase contrast (such as digital phase contrast or Zernike phase contrast) can be utilized to obtain the third image data. Then, a machine learning algorithm can be trained to predict the cell type classification information based on the image data corresponding to the third image data. On the other hand, the cell type classification information of the box 930 (via a detour of the position information) can be determined based on the first image data of the box 905 and the second image data of the box 910 (which is obtained by coloring). This means that the output data of the training dataset is based on an imaging method with coloring, which means that labels can be created particularly reliably. Using the newly trained machine learning algorithm, the cell type classification information can be directly predicted only based on the image data obtained without coloring.

[0103] Thus, with such training data, other machine learning algorithms can be trained. This can also be referred to as "bootstrapping" because the machine learning algorithms trained in block 940 can then be used, for example, in block 925 or block 930 to achieve better results and / or be able to generate input data more easily. These trained machine learning algorithms can even make Figure 1 the entire process in

[0104] redundant (since a large number of training data pairs can be generated, which enables even complex machine learning algorithms to be robustly trained, so that classification can be directly based on, for example, the morphological characteristics of cells).

[0105] For example, machine learning algorithms that do not require fluorescent staining (neither DAPI nuclear staining nor specific staining) can be trained and can predict position information and / or cell type classification information purely based on structure (e.g., based on phase contrast or brightfield images). Figure 1 To this end, the input data is obtained according to the

[0106] conventional workflow shown (still including fluorescent staining, at least for the second image data in block 910 and possibly also for the first image data in block 905), and further registering a contrast for it, which is non-specific and advantageous (e.g., brightfield or non-specific fluorescence; block 915). Then, as described above, the position information and cell type classification information are determined. Now, the results can be transferred to the "more favorable" image data using the bootstrapping in block 940. Thereby, a training data set is generated, which only contains the favorable contrast as input data and the results of the machine learning algorithm to be trained as output data (e.g., position information or cell type classification information). Using this training data set, machine learning models can now be trained, which can learn to localize and / or classify purely based on structure and do not require complex contrasts. For example, training data can also be generated based on partial results of blocks 925 and 930.

[0107] Figure 2 Schematically shows a data processing flow according to various examples. A microscope is shown at 205, using which the first image data 210 (see Figure 1 : block 905) and the second image data 230 (see Figure 1 : block 910) are obtained, and if necessary, other image data (see Figure 1: The frame 915) can be acquired. For example, the microscope can be an optical microscope that can work with and without fluorescence. For example, the microscope can provide an imaging method for phase contrast. For example, the microscope can have an illumination module, using which different illumination directions or illumination geometries can be activated to achieve digital phase contrast.

[0108] As Figure 2 shown, the first image data 210 is obtained using nuclear staining that is independent of cell type (cell type agnostic). The second image data 230 has multiple channels obtained with different cell type-specific stainings.

[0109] At 215, an artificial neural network (as an example of a machine learning algorithm), here a convolutional neural network (English: convolutional neural network), is used to predict the position information 220. The position information 220 corresponds here to a point position; this point position marks the stained cell nucleus. Thus 215 corresponds to Figure 1 the frame 925.

[0110] At 235, based on the position information 220 and the second image data 230, cell type classification information 245 (here distinguishing cell types "type A", "type B", and "type C") is determined. This is done using a machine learning algorithm that processes the image patches 241, 242, 243, 244 separately.

[0111] These image patches can, for example, correspond to regions centered on the point position 220. These regions can be sized based on separately determined size information (not shown in Figure 2 ; see Figure 1 : frame 920) such that each region depicts a cell. If the image data 210, 230 is scaled, regions of a fixed size can be used since cells have a specific standard size anyway.

[0112] Figure 3 Schematically shown is a data processing device 81 according to various examples. The data processing device 81 includes a processor 82, a memory 83, and a communication interface 84. The processor 82 can communicate with other devices via the communication interface 84, for example with an image acquisition device such as a microscope. The processor 82 can also communicate with a database (such as an image database) via the communication interface 84. For example, the processor 82 can send control commands to the image acquisition device via the communication interface 84. For example, the processor 82 can use such control commands to trigger the acquisition of image data, as described above in connection with Figure 1as described by box 905, box 910, and box 915. Alternatively or additionally, the processor 82 may receive image data via the communication interface 84, as described above in connection with Figure 1 box 905, box 910, and box 915 therein. The image data may be received from an image acquisition device and a database. The image data may also be loaded from the memory 83. The processor 82 may load and execute program code from the memory 83. When the processor 82 executes such program code, this causes the processor 82 to perform techniques such as those associated with the Figure 1 method in Figure 2 or the data processing flow described in

[0113] Generally speaking, the above techniques are based on a combination of image data having cell-type non-specific contrast (optionally with coloring) and additional image data having cell-type specific contrast. The individual cells may be located based on the image data having cell-type non-specific contrast. Different cell types may be distinguished based on the additional image data.

[0114] Of course, the features of the foregoing embodiments and aspects of the present disclosure may be combined with each other. Specifically, these features may not only be used in the described combinations, but also in other combinations or alone, without departing from the scope of the present disclosure.

[0115] For example, techniques for determining the size information of various image data and then using the size information to further evaluate the image data have been described above. Generally speaking, determining such size information is optional. Therefore, scaling the image data is also optional. The size information may also be determined but only for determining the position information (and not for determining cell type classification information based on the position information and the corresponding image data), or vice versa.

[0116] Techniques for using image data without temporal resolution have been described above. However, it is also possible to consider using image data having temporal resolution, such as a time series of images or contrasts. Cells (or their distribution / density) may also be observed and classified over time. By allocating cells (tracking) between time steps, individual decisions from multiple input images (or their individual results) may be combined and / or re-evaluated based on confidence levels, thus achieving more robust classification.

Claims

1. A computer-implemented method comprising: - obtaining (905) first image data (210) microscopically depicting an arrangement of cells with cell type non-specific contrast, - obtaining (910) second image data (230), the second image data microscopically depicting an arrangement of cells with a plurality of cell type specific contrasts, wherein the second image data is obtained by cell type specific staining of the cells, - determining (925) position information (220) of cells based on said first image data (210), and - determining cell type classification information (245) of the cell based on the second image data (230) and the position information (220).

2. The computer-implemented method of claim 1 , wherein: The first image data (210) is obtained by cell type non-specific cell staining.

3. A computer-implemented method according to claim 1 or 2, wherein: The first image data (210) is obtained by staining the nuclei of cells.

4. A computer-implemented method according to any one of the preceding claims, wherein: The method further comprises: performing (935) an analysis of the arrangement of cells based on the cell type classification information (245) and optionally the position information (220), Wherein, performing the analysis includes one or more of the following evaluations: determining the number of cells; determining the cell density; determining the occupied area; determining the neighborhood relationship between cells.

5. A computer-implemented method according to any one of the preceding claims, wherein: The second image data includes a multi-channel image having a plurality of channels, each channel being obtained by different cell type-specific staining.

6. A computer-implemented method according to any one of the preceding claims, wherein: Determining (925) the position information (220) includes applying one or more first image processing algorithms to the first image data (210), Wherein, the one or more first image processing algorithms are selected from the following group: threshold operation for determining fragments; morphological operation for determining fragments; single fragment filtering; filtering fragments according to shape and / or size; and fragment center determination.

7. A computer-implemented method according to any one of the preceding claims, wherein: Determining (925) the position information (220) includes applying one or more first image processing algorithms to the first image data (210), Wherein, the one or more first image processing algorithms have machine learning parameters.

8. A computer-implemented method according to any one of the preceding claims, wherein: The position information (220) is selected from the following group: center position; point position; cell density map; cell instance segmentation; bounding box; segmentation map.

9. A computer-implemented method according to any one of the preceding claims, wherein: Determining the cell type classification information includes applying a second image processing algorithm to the second image data (230), The second image processing algorithm selectively processes the second image data (230) in an area (241, 242, 243, 244) determined based on the position information (220).

10. The computer-implemented method of claim 9, wherein: The size of the regions (241, 242, 243, 244) is determined based on the size information of the cells, The size information indicates the imaging ratio of the cells in the second image data (230).

11. A computer-implemented method according to any one of the preceding claims, wherein: Determining the cell type classification information includes applying a second image processing algorithm to the second image data, Wherein, the second image processing algorithm includes machine learning parameters.

12. The computer-implemented method of claim 11, wherein: The second image processing algorithm receives position information (220) as a further input.

13. The computer-implemented method of any preceding claim, wherein the method further comprises: Image registration of the first image data (210) and the second image data (230) is performed (916).

14. A computer-implemented method according to any one of the preceding claims, wherein the method further comprises: determining (920) size information of the cells based on an analysis of at least one of the first image data (210), the second image data (230), or other image data depicting an arrangement of cells, the size information indicating an imaging scale of the cells in the first image data (210) and / or the second image data (230), Wherein, the position information (220) and / or the cell type classification information (245) are determined based on the size information.

15. The computer-implemented method of claim 14, wherein the method further comprises: Based on the size information, a first image processing algorithm for determining the position information (220) is parameterized and / or a second image processing algorithm for determining the cell type classification information (245) is parameterized.

16. The computer-implemented method of claim 14 or 15, wherein the method further comprises: Based on the size information, a first image processing algorithm is selected for determining the position information (220) and / or a second image processing algorithm is selected for determining the cell type classification information (245).

17. A computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Based on an analysis of at least one of the first image data (210), the second image data (230), or other image data depicting cell arrangements, the first image data (210) and / or the second image data (230) are scaled (920) so that they depict cells according to a reference scale.

18. A computer-implemented method according to any one of the preceding claims, wherein the method further comprises: Based on at least one of the position information (220) or the cell type classification information (245), a training data set is generated (940), and a machine learning algorithm is trained based on the training data set.

19. The computer-implemented method of claim 18, wherein: The training data set includes the first image data (210) and optionally the position information (220) as input data for a machine learning algorithm, The training data set includes the cell type classification information (245) as output data of the machine learning model.

20. The computer-implemented method of claim 18, wherein: The first image data (210) is obtained by cell type non-specific cell staining, The method further comprises: obtaining (915) third image data, the third image data microscopically imaging the arrangement of cells with a cell type non-specific contrast, wherein the third image data is obtained without staining the cells, Wherein, the training data set includes the third image data as input data of the machine learning model, Wherein, the training data set includes the location information (220) as output data of the machine learning model.

21. The computer-implemented method of claim 18, wherein: The first image data (210) is obtained by cell type non-specific cell staining, The method further comprises: obtaining third image data, wherein the third image data is microscopically imaged for the arrangement of cells with a non-specific contrast of the cell type, wherein the third image data is obtained without staining the cells, Wherein, the training data set includes the third image data as input data of the machine learning model, Wherein, the training data set includes the cell type classification information (245) as output data of the machine learning model.

22. A training data set comprising a plurality of input-output data pairs, wherein output data of the input-output data pairs are determined based on cell type classification information (245).

23. A data processing device comprising a processor and a memory, wherein the processor is configured to load and execute a program code from the memory, wherein execution of the program code causes the processor to perform the method according to any one of the preceding claims 1-21.