Method for determining sample location based on structural information
By using the selection model and recognition model in the sample image data, the sample location is determined based on structural information, and the resource-intensive and time-intensive coverslip edge recognition problem in the prior art is solved, and faster and more accurate sample positioning is achieved.
Patent Information
- Application Number
- CN202510140792.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-09
- Filing Date
- 2025-02-08
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, when identifying the edge of the coverslip, especially in low-contrast image recording, there are problems such as resource-intensive, time-consuming and data-intensive, resulting in inefficient analysis and affecting customer acceptance.
By using the structural information of the sample image data, the selection model and identification model are used to determine the location of the sample in the imaging device, the coarse image data with reduced detail is used to identify the structural area, the consumption of calculation and data resources is reduced, and the analysis speed and accuracy are improved.
Faster and more reliable sample positioning is achieved, reducing the use of computing resources and data resources, while improving analysis efficiency and accuracy.
Smart Images

Figure CN120472138A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method of controlling an imaging device based on structural information determined in sample image data, a method of training a machine learning system, a machine learning system and a computer program product. Background Art
[0002] A method for identifying a cover glass edge is known from the prior art, in which a single high-resolution recording of the sample is used in order to search for the cover glass edge or similar relevant structures in the high-resolution recording using conventional image analysis tools.
[0003] In "Learning to zoom: a saliency-based sampling layer for neural networks," Recasens, A., et al., describe a method for determining a low-resolution intermediate image for a model based on its high-resolution output image. The intermediate image retains relevant or interesting regions at sufficient resolution, while reducing the resolution of irrelevant regions. To capture regions of varying resolution in the joint intermediate image, an elastic mesh must be estimated that irregularly deforms the high-resolution original image. This irregular deformation can lead to geometrical artifacts in the intermediate image, particularly distortions of straight and circular shapes or objects in the sample. Furthermore, the deformations applied by the model cannot be anticipated or explained, which is particularly problematic for previously unseen data.
[0004] In practice, detecting the coverslip edge, especially in low-contrast image recordings, has been found to pose significant challenges for conventional methods. Image analysis is often resource-intensive, time-consuming, and data-intensive. Consequently, the analysis has been found to be often flawed. This has a particularly negative impact on customer acceptance. Summary of the Invention
[0005] The present invention relates to a method of controlling an imaging device based on structural information determined in sample image data, a method of training a machine learning system, a machine learning system and a computer program product.
[0006] The present invention is based on the object of providing an improved method for controlling an imaging device, which method is based on structural information determined in image data of a sample. In particular, the method should be faster, use fewer computing resources, use fewer data resources and be more reliable.
[0007] One aspect of the present invention relates to a method for determining a sample positioning based on structural information, wherein elements of a sample holding device that holds the sample in an imaging device form a structure in image data captured with the imaging device, comprising: providing image data, determining a structural region based on coarse image data with the aid of a selection model, determining structural information based on the structural region with the aid of a recognition model, and determining the positioning of the sample based on the structural information, characterized in that the coarse image data is reduced in detail compared to the image data, the structural region is a region in the image data where the structure is captured with a certain probability, and the total amount of data of the structural region is less than the amount of data of the image data.
[0008] In the sense of the present invention, the "positioning" of a sample is simply the position of the sample in the sample holding device or in the imaging device, or the position of the sample in the sample carrier. These different positions are treated in an equivalent manner and differ from each other only by a relative displacement.
[0009] In the sense of the present invention, a "sample" is in particular a biological sample. However, the sample can also be any other sample that can be conveniently controlled by the imaging device based on structural information.
[0010] For the purposes of the present invention, an "imaging device" is particularly a microscope. This is not limited to certain types of microscopes. In particular, a microscope includes at least one camera capable of recording image data. Alternatively, however, the imaging device may be any other device for recording samples, particularly biological samples.
[0011] In the sense of the present invention, "structural information" is information about the respective sample holder, in particular about the construction, in particular the physical construction, of the sample holder. In particular, this can be shape information, position information and / or height information. If the sample holder is, for example, a cover glass, a sample carrier or a slide, the shape information comprises in particular information whether the sample holder is round, square or rectangular. Furthermore, the structural information can comprise information about the dimensions of the sample holder and the position of the sample holder. The position can be determined in particular relative to the sample. For different sample holders from the prior art, specific standardized object shapes or object geometries are generally known, wherein, in particular, for each object shape, a set of structural information comprising shape information and size information is known and determined accordingly. If, in particular, edges are identified in the image data, the structural information can be inferred based on the shape and size or extent of the edges.
[0012] In the sense of the present invention, a "sample holding device" comprises in particular a plurality of components which together provide a sample in an imaging device. In particular, these components comprise one or more cover slips, a sample carrier, a label, a spacer and a slide, in particular the sample holding device is only partially visible, in particular the edges of the cover slip, the sample carrier and the slide are visible.
[0013] In the sense of the present invention, "image data" is data captured by an imaging device. In particular, in this application, the terms "image data" and "image" are used synonymously. In particular, image data may also include contextual information.
[0014] For the purposes of the present invention, a "camera" can be any camera of an imaging device. If the imaging device is a microscope, the camera is particularly an overview camera. Alternatively, the camera can also be a microscope camera. A microscope camera can use an objective lens with the smallest possible magnification to record an overview image of the sample, based on which structural information can then be determined. Alternatively, it can also record multiple adjacent images, so-called tiled images, and then combine them to form an overview image. Possible magnifications include, for example, 1x, 2x, 3x, or 5x magnification.
[0015] For the purposes of the present invention, an "image" is a recording made by an imaging device. In particular, an image for the purposes of the present invention may also include multiple images, in particular image stacks and time sequences of images or image stacks. In particular, an image may include depth information in addition to color and brightness information. Images may be recorded in brightfield, darkfield, or as phase contrast images, for example.
[0016] For the purposes of the present invention, a "structure region" is a region in the image data where the structure of the sample holder is typically captured or visible with a certain probability. In particular, a structure region refers to a region of both the image data and the coarse image data. Specifically, the structure region can be first determined based on the coarse image data and then transferred or applied to the image data. Exactly how this conversion is performed depends on how the coarse image data is calculated from the image data. However, in particular, the structure region can also be determined with reference to the sample holder and thus be directly applied to both the coarse and image data. In particular, the structure in the structure region may be better represented in the image data, visible due to greater detail, or only visible in the image data, compared to the structure region in the coarse image data. If the image data is a simple image, the structure region is an image portion. If the image data is a stack of images, the structure region is an image portion that is arranged one above the other. If the image data is a time series of images, the structure region is an image portion of the same point in the time series of images, where the same point of the mapped sample holder portion is captured at the same image point in the image portion of the image.
[0017] In the sense of the present invention, "coarse image data" are data with reduced details compared to the image data. In particular, the coarse image data comprise the same area of the sample that is also captured by the image data.
[0018] In the sense of the present invention, a "selection model" is a machine learning model configured to identify structural regions in the coarse image data. Thus, the selection model finds partial regions in the coarse image data that at least partially capture the sample holding device with a certain probability.
[0019] In the sense of the present invention, a "recognition model" is a processing model, in particular a machine learning model, that is configured to recognize structural information. The recognition model can be configured as a classical processing model for recognizing edges, wherein, for example, the edge shape is determined based on the recognized edges. Alternatively, the recognition model can also be configured as a machine learning model.
[0020] In the sense of the present invention, a “machine learning model” is a processing model, in particular a neural network, which is configured to process input data and output output data or result data by supervised or unsupervised learning, and which can in particular be trained to perform a specific mapping.
[0021] In the context of the present invention, a "processing model" is a model configured to process input data and output output data, in particular, result data. A processing model can be, for example, a classical model created for the application of classical optimization or analytical methods. Similarly, a processing model can be a model trained using a learning method, which is then referred to as a machine learning model. In the case of a processing model constructed with multiple layers, a distinction is made between the output data of intermediate layers and the output data of the last layer. The output data of the last layer (the so-called output layer) is also referred to as result data.
[0022] In the sense of the present invention, "input data" is data that is input into a processing model and processed by the processing model. In the present invention, the input data is particularly image data. The input data may particularly include a single image or multiple images, a stack of images, or a time series of images or image stacks.
[0023] In the sense of the present invention, "result data" is data output by the processing model, which is calculated and output by processing the input data by the processing model. In a multi-layer model, the last layer of the model, the so-called output layer, also called the result layer, outputs the result data.
[0024] In the sense of the present invention, "output data" are data output by a process model, wherein a process model can in particular output a plurality of output data, in particular result data. In addition to the result data, a process model can also output one or more intermediate data.
[0025] In the sense of the present invention, a "convolutional network" is a neural network with convolutional layers. In particular, in addition to convolutional layers, the network may also include pooling layers, nonlinear layers, and other known layers. The arrangement of layers is defined in the network architecture.
[0026] In the sense of the present invention, an “intermediate layer” is a layer in a machine learning model (in particular a network or neural network) that receives input data from the previous layer and forwards output data to subsequent layers of the network.
[0027] In the sense of the present invention, “intermediate data”, also called intermediate output or intermediate layer output, is the output of a layer of a multi-layer processing model that is not the last layer of the processing model, i.e., an intermediate layer.
[0028] Prior art methods for determining sample location based on identified structural information (particularly the glass edge and its shape) either require specialized lighting equipment to reliably identify the glass edge or require analyzing images recorded from the glass edge, which incurs significant computational and memory overhead. Consequently, image analysis is very time-consuming. Despite sophisticated analysis, the results are often flawed. This has had a significant negative impact on customer acceptance of such automated systems for sample location. The inventors have recognized that a large portion of image data is irrelevant to the discovery of structural information and therefore propose first using a selection model to identify structural regions relevant to sample discovery based on coarse image data, which has significantly reduced detail compared to the image data. Using this coarse image data with reduced detail significantly reduces the amount of data the selection model must analyze, significantly speeding up the analysis and consuming less computing resources. Furthermore, using this coarse image data with reduced detail significantly improves data efficiency during the training of the selection model. Because the coarse image data exhibits less detail, the selection model must learn fewer details during training, which is why less training data is required for complete training. Furthermore, the reduction in detail ensures that the selection model is not trained on barely visible small details. Instead, the semantics of the image (also known as context) allows for the identification of large structures within the image, and structural regions can be identified based on these large structures within the image data. In fact, in the sample under consideration, the sample holder is always in a similar position relative to these large structures, which may be found at different locations in the image data, depending on the imaging device. To actually locate the structures for determining structural information, image data—that is, actual image data—is again required. However, only those portions of the image data identified as structural regions are processed by the recognition model, so the recognition model must also process a significantly smaller amount of image data for structural regions, as structural regions only occupy a small portion of the image data. Consequently, the described method achieves a further reduction in required computing and storage resources, as well as reduced user waiting time. Furthermore, structural regions can be analyzed with full detail fidelity, thereby improving accuracy in determining structural information. Image data not identified as structural regions does not need to be processed.
[0029] The image data preferably comprises in particular images captured by a camera of the imaging device, in particular one or more of: time series images, image stacks, stereo images, images with depth information, low contrast images or images captured with a low magnification objective.
[0030] Because the method can be applied to many different types of image data, the disclosed method can be applied to image data from many different imaging devices.
[0031] Preferably, the coarse image data exhibit one or more of the following compared to the image data: a lower sampling depth, a lower image resolution, a lower temporal resolution, a lower height resolution, in particular a larger distance between adjacent images of the stack.
[0032] Since the reduction of details can be performed in different ways, a suitable detail reduction can be selected in particular depending on the context of the respective sample indicating or reproducing the large structure to be recognized or to be identified.
[0033] Preferably, the sample holding device comprises one or more elements, in particular a slide, a cover slip, a spacer, a sample carrier, a holding frame, an inscription, a mark or a label, wherein the structure of the sample holding device captured in the image data is in particular visible as light or dark lines, light or dark arcs, arcs or circles, so-called spots, in particular light or dark image areas, so-called dots, distortions, mirror images, doublings, textures or characters.
[0034] In the sense of the present invention, an "element" is one or more components of the sample holding device that form one or more structures in the image data during capture with an imaging device. In particular, the element may depend on the position in the imaging device, the different elements of the sample holding device, and the illumination of the imaging device, depending on the imaging device used.
[0035] Preferably, the structural information comprises in particular information about the geometry, orientation and / or position of the structure, in particular whether the structure is straight or round.
[0036] Since very different sample holding devices and their respective shapes and sizes can be determined, the method can be applied to very different imaging devices.
[0037] The method preferably comprises determining coarse image data based on the image data, and in particular determining structure regions in the coarse image data, and selecting structure regions of the image data corresponding to the structure regions of the coarse image data, wherein structure regions corresponding to each other capture the same element of the sample holding device.
[0038] Since the coarse image data and the image data each capture the same image region, the structural region between the coarse image data and the image data can be determined particularly easily. Since the structural region is determined first on the coarse image data, much less data must be processed by the selection model, which improves data efficiency.
[0039] The selection model is preferably a machine learning model implemented as a classifier, a detector, a segmentation model or an image-to-image model, and the determination of the structural region comprises: inputting at least a partial region of the coarse image data as input data into the selection model, outputting output data, and in particular selecting the structural region from the coarse image data based on the output data.
[0040] In the sense of the present invention, a "classifier", also called a classification model, is a processing model that assigns categories to input data or can be trained to assign categories to input data. The classifier can in particular be a machine learning model. The result data can in particular be the categories assigned to the input data, wherein the format can in particular be a vector, wherein each entry of the vector corresponds exactly to one of the possible categories to be assigned, and in particular a "1" entry in the vector indicates the category of the input data. Alternatively, a category number can be output. However, as a further alternative, the classifier can also be trained such that the result data is a vector, wherein the entries in the vector respectively indicate the probability that the respective input data belongs to the category corresponding to the entry of the result data. Depending on the implementation, the respective formats of the result data of the classifier and correspondingly the formats of the target data in the annotation dataset used to train the classifier also differ.
[0041] In the context of the present invention, a "detector," also called a detection model, is a machine learning model that has been trained to recognize predetermined detection patterns in input data and output a list. In particular, this list is a list of localizations, such as localizations in the input data. The input data can be, in particular, an image, an image stack, or an input tensor. The exact format of the localizations depends, in particular, on the format of the input data and the detection patterns to be identified.
[0042] In the sense of the present invention, a "segmentation model", also called a semantic segmentation model or semantic segmenter, is a processing model that assigns an output value to each item of input data; the resulting data is also called a segmentation mask or semantic segmentation mask. If the input data is an image, the segmentation model performs an image-to-image model, where an output value corresponding to a semantic meaning is assigned to each image point of the input data.
[0043] In the sense of the present invention, an "image-to-image model" is a processing model configured to perform image-to-image mapping. An image-to-image model assigns a value in the output data to each entry of the input data.
[0044] In the sense of the present invention, an "annotated dataset" includes input data and target data, where annotations or identifications of target data (referred to as target data) correspond to each input data element. The target data are typically generated through complex processing or manual labeling. The processing model used to perform the desired mapping is trained based on the target data.
[0045] In the sense of the present invention, "target data" are data used in training a process model for performing process mapping, to which result data output by the process model based on input data is to be adjusted. Approximation is performed with the aid of an objective function.
[0046] In the sense of the present invention, an “objective function” is in particular a gain function or a loss function, which specifies how the difference between the result data of the processing model and the target data is evaluated. In the training of a machine learning model, the training is performed by optimizing the objective function, wherein the model parameters of the trained machine learning model are adjusted during the training so that the objective function is optimized.
[0047] In the sense of the present invention, a "gain function" is an objective function that captures the match and is maximized during training, compared to a loss function that captures the difference between result data and target data and is minimized during training.
[0048] In the sense of the present invention, a "loss function" is a function that captures the difference between the result data and predefined target data. For example, if the result data and the target data are images, the comparison can be performed pixel by pixel. For example, if the result data and the target data are vectors or tensors, the difference can be performed entry by entry. The differences can be added in absolute value (as absolute values) in the L1 loss function. The sum of the squares of the differences is formed in the L2 loss function. In order to minimize the loss function, the values of the model parameters of the processing model are changed, which can be calculated, for example, by gradient descent and backpropagation. Further possible loss functions are in particular the cross entropy loss, the hinge loss, the logistic loss, the log-likelihood loss, the Gaussian negative log-likelihood loss or the Kullback-Leibler loss.
[0049] For the purposes of this invention, "model parameters" are parameters of a machine learning model that determine how the model's output values are calculated from its input values. During machine learning model training, the model parameters of the machine learning model are adjusted so that the model's output matches the desired output as closely as possible, i.e., the result data matches the target data as closely as possible. The machine learning model learns the desired mapping by appropriately adjusting the model parameters.
[0050] Since the selection model is implemented as a machine learning model, the selection model can be used for image analysis particularly efficiently, for example on special graphics cards provided for this purpose or special processors for the calculation of neural networks or similar machine learning models.
[0051] The selection model is preferably a classifier, and the result data includes a class assignment of the input data to a result class, and the selection of the structure region is performed based on the result class, the result class including one or more of the following classes: a structure class, a non-structure class, a label class, and a non-sample class, and the structure region is exactly the image region assigned to the structure class. In particular, the classifier can be a binary classifier having a structure class and a non-structure class.
[0052] Since the selection model can be implemented in various ways, the method can be optimized for specific purposes depending on the specification. If a particularly precise determination of structural regions is required, it is preferable to use an image-to-image model that outputs a probability map based on which structural regions can be determined very accurately. While generating the target data is certainly more complex when training such an image-to-image model, it allows for very accurate determination of structural regions. If, instead, a simple classifier is used, annotation is correspondingly less complex.
[0053] The classifier is preferably a conventional classifier or a patch classifier, wherein the input data of the conventional classifier is coarse image data, the result data includes a classification map, wherein each result category is respectively assigned to a partial area of the coarse image data in the classification map, and the input data of the patch classifier are respectively partial areas of the coarse image data, which are continuously selected by a sliding window function, and the individual result categories are output as result data of each partial area.
[0054] In the context of the present invention, a "sliding window function" is a function that continuously selects partial regions (also referred to as partial data sets) from a data set corresponding to the size of the sliding window and provides these partial regions for further processing. In this case, the sliding window can particularly be a one-dimensional sliding window, a two-dimensional sliding window, or a three-dimensional sliding window. The sliding window function selects spatially and / or temporally coherent data from each data set. If a first partial data set is selected and provided, the sliding window is shifted further in the data set by a number of entries (step size) before the next sliding window is provided for further processing.
[0055] Since the coarse image data is divided into an actual result class and a non-result class, the structure region can be selected from the coarse image data in a simple manner based on the result class.
[0056] The conventional classifier is preferably configured to omit the final pooling layer, so that the conventional classifier outputs a classification map as result data.
[0057] According to the present invention, a conventional classifier in which the final pooling layer is omitted is also referred to as a graph classifier. Since the selection model is implemented as a graph classifier, the coarse image data only needs to be input into the selection model once, and the result data then provides the result class for all image regions of the input data.
[0058] The selected model is preferably a neural network, in particular one or more of: a fully convolutional network, in particular a DenseNet, a Resnet or a ResNext.
[0059] The selection model is preferably implemented as a detector, and the result data comprises a list with data locations by which the structure region was selected from the image data.
[0060] Since the selection model is implemented as a detector, the structural regions can be determined in a particularly simple manner via the list output by the detector.
[0061] The selection model is preferably implemented as a segmentation model, and the output data is a segmentation mask, wherein a resulting class is assigned to each entry of the input data, indicating whether the respective entry belongs to a structural region or not.
[0062] In the sense of the present invention, a "segmentation mask" is output data of a machine learning model, wherein an output value is assigned to each entry of the input data. The output value can in particular correspond to an assigned class.
[0063] Since the selection model is implemented as a segmentation model and outputs a segmentation mask, the results can be checked in a particularly simple way, e.g. by overlaying the segmentation mask and the image data. An operator can quickly and simply check the consistency of the segmentation mask, e.g. thereby quickly creating annotated datasets.
[0064] The selection model is preferably configured as an image-to-image model, wherein the output data is a probability map, wherein probability values (in particular a probability distribution) are assigned to entries of the input data, which indicate the probability of capturing a structure at the entry location, in particular a probability value is assigned to each entry or respectively to a group of entries, and the determination of the structure region comprises grouping the entries of the image data based on the probability, wherein in particular consecutive entries of the image data with probability values above a certain probability are combined to form a structure region or to form a structure region.
[0065] In the sense of the present invention, a "probability map" indicates the probability of capturing the sample holding device at various points in the image data. In particular, a confidence map can be determined from the probability map.
[0066] Since the selection model is implemented as an image-to-image model which outputs a probability map, the image data to be further processed can be changed in a particularly simple manner during further processing, for example by changing the specific probabilities.
[0067] If the selection model is implemented as an image-to-image model or a segmentation model, the selection model is, for example, a Unet or an encoder-decoder network.
[0068] Preferably, the determination of the structural information comprises inputting structural regions of the image data into the recognition model according to a sequence, wherein the sequence is determined based on the probability values of the structural regions, in particular structural regions with higher probability values are classified in the sequence before structural regions with lower probability values, and the input of the structural regions is in particular terminated when a certain amount of structural information has been determined, or only a predetermined number of structural regions with the highest corresponding probability values are input into the recognition model, and no further structural regions with lower probability values are input into the recognition model.
[0069] Since the structural regions are processed according to their respective probability values, the computational efficiency can be further improved, since structural regions with higher probability values can be expected to provide better results during further processing with the recognition model.
[0070] The recognition model is preferably implemented as a classifier, a segmentation model, a detector or an image-to-image model.
[0071] The recognition model is preferably implemented as a classifier, the result data of the classifier particularly including shape categories, and the structural information is determined based on the shape categories, wherein the shape categories particularly include one or more of the following: no structure, structure, circular structure, straight structure, polygonal structure, straight coverslip edge, polygonal coverslip edge, circular coverslip edge, sample carrier edge, gasket structure, retaining frame structure, microtiter plate edge structure, microtiter plate well structure, sample chamber edge structure, sample chamber structure.
[0072] Since the classifier is used as the recognition model, creating annotated datasets is particularly simple.
[0073] The recognition model is preferably implemented as a classifier, wherein the result data of the classifier include component classes, and the structural information is determined based on the assigned component classes, wherein a type of possible component types of the sample holding device is specifically assigned to each component class.
[0074] Since the recognition model is implemented as a classifier which outputs a component class, the localization of the sample can be determined particularly easily based on the respectively recognized component.
[0075] The recognition model is preferably implemented as a segmentation model, and the result data comprises a segmentation mask for the respective structural region, wherein a shape class is assigned to each entry of the input data, the structural information is determined based on the segmentation mask, and the shape class in particular comprises one or more of the following classes: no structure, structure, circular structure, straight structure, polygonal structure, straight coverslip edge, polygonal coverslip edge, circular coverslip edge, sample carrier edge, gasket structure, retaining frame structure, microtiter plate edge structure, microtiter plate well structure, sample chamber edge structure, sample chamber structure.
[0076] Because the recognition model is implemented as a segmentation model, the granularity of shape categories in the image data is significantly improved, allowing for more accurate determination of structural information based on shape categories. This is particularly useful when incorporating structural information.
[0077] The image data preferably comprise a plurality of overview images, a plurality of structure regions being determined in the overview images, and the positioning of the sample is determined based on a plurality of segmentation masks determined based on the structure regions.
[0078] In the sense of the present invention, an "overview image" is an image recorded with an overview camera or an image recorded with a microscope camera having an objective lens with a small magnification. In this case, the overview image particularly captures one or more of the sample, the sample carrier and the sample holding device.
[0079] Since multiple overview images are determined, the sample can be better localized. In particular, gaps in the segmentation mask can be filled by merging them, which improves the accuracy of localization.
[0080] Preferably, the result data output for different structural areas are merged with the remaining image data to form detail-reduced result data, the individual result data are assigned to the structural areas, and the values of the non-structural shape category are assigned to the entries of the remaining image data, and the positioning is determined based on the detail-reduced result data.
[0081] By combining multiple segmentation masks, localization can be verified particularly simply.
[0082] The recognition model is preferably configured as an image-to-image model, wherein each entry of the input data is assigned a value in the result data, which value indicates whether the respective entry captures a structure, wherein the value is in particular a probability, and in particular the result data is a probability map, in particular the probability value can also be a probability distribution over multiple shape classes.
[0083] Since the recognition model outputs a probability map, the resulting data can be verified particularly well and is explanatory. Furthermore, the information content in the probability map is higher. In particular, during further processing of the probability distribution, the result directly provides information about the quality of the distribution, which can also be taken into account during further processing, thus further increasing the reliability of the correct result, i.e., the final location of the sample.
[0084] Preferably, the determination of the localization comprises merging structure information of a plurality of source identical structure regions from different overview images, wherein one or more structures of each structure caused by the same element of the sample holding device are captured in the source identical structure regions.
[0085] Since different structural information of the same structural region of the source is merged, the reliability of the entire method can be further improved.
[0086] The recognition model is preferably a neural network, in particular one or more of: a fully convolutional network, in particular a DenseNet, a ResNet or a ResNext.
[0087] Preferably, the selection model and / or the recognition model is selected from a list of machine learning models based on the imaging device, the sample holding device used, the sample carrier 106 used and in particular on the overview camera used, respectively.
[0088] Since different machine learning models are selected based on the configuration of the image data evaluation system, a specific machine learning model can be provided that functions very well in each case, which significantly improves the quality of the results.
[0089] Preferably, context information is used when determining the structure region, determining the structure information or determining the positioning.
[0090] In the sense of the present invention, "context information" includes one or more of the following:
[0091] information on available and / or used light sources and / or their exposure spectra,
[0092] information on available and / or used light source filters and / or their chromaticity characteristics during spectral filtering of the illumination spectrum of the light source,
[0093] information about available and / or used fluorescence filters for spectrally filtering the fluorescence spectrum emitted by the sample,
[0094] Information on available and / or used dichroic mirrors and their chromaticity characteristics,
[0095] The illumination of the light source,
[0096] The lighting time of the light source,
[0097] The type of sample, sample holding device, or imaging device being recorded,
[0098] the type of sample support used, e.g. whether chamber slides, microtiter plates, slides with coverslips, or culture dishes were used,
[0099] Image recording parameters, such as information about illumination intensity, illumination time, filter settings, fluorescence excitation, contrast method or sample stage setup,
[0100] information about the objects contained in the individual overview images,
[0101] Application information indicating the type of application that recorded the overview image,
[0102] Information about the user who recorded the sample image.
[0103] Another aspect of the invention relates to a method of controlling an imaging device for capturing a sample based on a positioning of the sample, wherein the positioning has been determined according to the above method, the method further comprising controlling the imaging device based on the determined positioning.
[0104] Since the positioning of the sample is automatically determined, the sample can be analyzed fully automatically thereafter, thus increasing throughput.
[0105] Another aspect of the present invention relates to a method for training a selection model for determining object areas based on coarse image data, wherein the selection model is specifically trained for performing the above method, comprising: providing image data, determining structural areas in the image data, determining coarse image data with reduced details compared to the image data, determining object areas corresponding to the structural areas, providing the coarse image data as input data and target data, based on which the structural areas can be identified as an annotation data set for training the selection model.
[0106] Since the annotated dataset can be determined automatically, the selection model can be trained quickly and reliably in a simple manner even for new image data evaluation systems.
[0107] A further aspect of the invention relates to a control device for controlling an image data evaluation system, in particular designed as a microscope, comprising means for carrying out the above-described method.
[0108] Another aspect of the invention relates to a computer program product comprising instructions which, when the program is executed by one or more computers, cause them to carry out the method described above.
[0109] Another aspect of the present invention relates to an imaging device, in particular designed as a microscope, comprising the above-mentioned control device.
[0110] Another aspect of the invention relates to an image data evaluation system comprising at least the imaging device described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0111] The invention is explained in more detail below based on the examples shown in the accompanying drawings. The drawings show:
[0112] Figure 1 A system for a method according to an embodiment is schematically shown;
[0113] Figure 2A schematically illustrates parts of a system for a method according to one embodiment;
[0114] Figure 2B schematically illustrates parts of a system for a method according to one embodiment;
[0115] Figure 2C schematically illustrates parts of a system for a method according to one embodiment;
[0116] Figure 3 schematically illustrates image data used for a method according to one embodiment;
[0117] Figure 4 schematically illustrates parts of a system for a method according to one embodiment;
[0118] Figure 5 schematically illustrates parts of a system for a method according to one embodiment;
[0119] Figure 6 Schematically illustrates a machine learning model used for a method according to one embodiment;
[0120] Figure 7 schematically shows an arrangement of components of a system for a method according to one embodiment,
[0121] Figure 8 Schematically illustrates a method according to one embodiment;
[0122] Figure 9A Schematically illustrates an arrangement of components of a system for a method according to one embodiment;
[0123] Figure 9B Schematically illustrates an arrangement of components of a system for a method according to one embodiment;
[0124] Figure 9C Schematically illustrates an arrangement of components of a system for a method according to one embodiment;
[0125] Figure 10schematically illustrates a method according to one embodiment; and
[0126] Figure 11 A method according to one embodiment is schematically illustrated. DETAILED DESCRIPTION
[0127] An exemplary embodiment relates to an image data evaluation system 1. The image data evaluation system 1 comprises an imaging device 100, which is in particular a microscope, and a control device 130 (also referred to as evaluation and control device). The control device 130 can in particular be connected to a monitor 120. The control device 130 is communicatively coupled to the imaging device 100 (e.g., by a wired or wireless communication link). The control device 130 can evaluate image data 500 (see FIG. 1 ) captured using the microscope 100. Figure 1 ) and controls the imaging device 100, for example based on the evaluated image data 500. If the image data evaluation system 1 comprises a machine learning model, for example a neural network, it is also referred to as a machine learning system.
[0128] The image data evaluation system 1 is in particular configured for automatic evaluation, in particular for determining the position 202 of the sample based on structures 312 contained in the overview image 300 recorded from the sample.
[0129] The imaging device 100 according to the illustrated embodiment is an optical microscope. Microscope 100 comprises a stand 101, which includes further microscope components. These further microscope components include, in particular, a nosepiece or turret 102 with a mounted objective lens 103, a sample stage 104 with a holding frame 105 for holding a sample carrier 106 (also referred to as a sample holder), and a microscope camera 107. The combination of sample stage 104, holding frame 105, and sample carrier 106 is also referred to as a sample storage device. The sample holder may also include, in particular, a cover glass and a spacer.
[0130] If a sample is clamped in the sample carrier 106 and the objective lens 103 is pivoted into the microscope optical path, the fluorescence illumination device 108 can illuminate the sample for fluorescence recording, and the microscope camera 107 can receive the fluorescence from the clamped sample as detection light and record image data 500 with fluorescence contrast. If the microscope 100 is to be used for transmitted light microscopy, the transmitted light illumination device 109 can be used to illuminate the sample. The microscope camera 107 receives the detection light after passing through the clamped sample and records image data 500. The sample can be any object, fluid, or other microstructure, in particular biological microstructures.
[0131] Microscope 100 optionally also includes an overview camera 110, with which an overview image of the sample environment can be recorded. The overview image particularly shows the sample holding device including sample carrier 106. The field of view 111 of overview camera 110 is larger than the field of view during recording of image data 500 by microscope camera 107. In particular, overview camera 110 observes the sample holding device via a mirror 112. Mirror 112 is arranged on objective turret 102 and can be selected instead of objective lens 103.
[0132] According to some embodiments of the present invention, different sample carriers 106 can be used in the microscope 100. In particular, a simple sample carrier 106 (also called a slide) can be used, in which the sample is arranged on a glass plate. In addition, there are sample carriers 106 in which the sample is arranged between the above-mentioned glass plate of the sample carrier 106 and the cover glass 204. Such a sample carrier 106 is schematically shown in FIG. Figure 2A Here, the positioning 202 of the sample is shown schematically in a shaded manner. For the sake of clarity, the positioning 202 is shown for only one of the two samples arranged under the cover glass 204. There are many different models and types of sample carriers 106, which are known to those skilled in the art, and the typical positioning 202 of the sample in each sample carrier 106 is also known to those skilled in the art.
[0133] Furthermore, the sample carrier 106 may also have a label 206, which is also schematically shown in FIG. Figure 2A The label 206 can be used in particular to identify the sample; for example, a barcode, QR code, or handwritten marking can be used for this purpose. The label 206 can also be recognized and analyzed when recording the overview image 300. Instead of the label 206, a barcode or QR code can also be printed directly on the sample carrier 106, without the need for the label 206 to be provided on the sample carrier 106.
[0134] Figure 2B Another type of sample carrier 106 is shown. Figure 2B The sample carrier 106 in FIG. 1 is a so-called microtiter plate, also known as a multi-well plate. Such a microtiter plate comprises a plurality of wells 208, which are generally arranged in a regular pattern, into which the sample is arranged, so that the locations 202 correspond approximately to the wells 208. The number of wells 208 and the size of the wells 208 can vary widely. The wells 208 are typically arranged in a regular pattern in such a microtiter plate, wherein the pattern can vary from one microtiter plate to another.
[0135] Figure 2C Also schematically shown is a sample carrier 106 having a plurality of generally closed sample chambers 210 (also referred to as chamber slides). Also for this type of sample carrier 106, the sample locations 202 approximately coincide with the individual sample chambers 210. Figures 2A to 2C The different sample carriers 106 shown in FIG. 1 are merely exemplary selections of sample carriers 106 that may be used in an experiment. A person skilled in the art is aware of the different commonly used sample carriers 106 and the various positions 202 of the sample in each sample carrier 106.
[0136] even though Figure 2B and Figure 2C The sample carriers 106 shown in FIG1 do not have labels 206; these sample carriers 106 may also have one or more labels 206, but for better clarity, labels 206 are omitted for these sample carriers 106. In particular, each well 208 or each sample chamber 210 may be provided with a label 206, respectively, so that the respective well 208 or the respective sample chamber 210 can be identified by the label 206 based on the respective label 206. In particular, the label 206 may include contextual information, for example, encoded in a barcode, a QR code or a handwritten inscription.
[0137] The image data evaluation system 1 is designed in particular to record an overview image 300 of the sample carriers 106 in order to determine the position 202 of one or more samples in the respective sample carrier 106 based on the overview image 300. In the overview image 300 of the sample held in the sample holding device, structures 312 are formed which are caused by elements of the sample holding device.
[0138] By way of example, Figure 3 Such an overview image 300 recorded with the overview camera 110 is shown. Six bright spots 314 can be identified in the overview image 300. The bright spots 314 are caused by the lighting of the overview camera 110, which are reflected on the glass surface of the sample carrier 106 and are therefore detected by the overview camera 110 in the overview image 300. Depending on the geometry of the overview camera 110 (also with respect to the lighting), the geometry of the sample holder, the relative arrangement of the sample holder with respect to the overview camera 110, or the geometry of the optics of the overview camera 110, as well as the respective arrangement of the sample holder, the sample carrier 106, and the overview camera 110 or the camera of the microscope 100 relative to one another, for example, certain elements of the sample holder can produce structures 312 in the overview image 300 given a specific arrangement with respect to the overview camera 110. These structures 312 (also called image artifacts) can produce spherical or chromatic properties of the elements by casting shadows, reflections, refractions, or diffractions. Structures 312 or image artifacts may include, in particular, light or dark lines, light or dark arcs, arcs or circles, so-called spots, in particular light or dark image areas, so-called dots, distortions, mirror images, doublings, etc.
[0139] Reference below Figure 4 The schematic diagram in FIG. 3 illustrates the occurrence of structure 312 by way of example. Figure 4Shown using its record Figure 3 Schematic side view of an arrangement of an overview image 300 is shown in FIG. In this example, the overview camera 110 has an illumination device comprising a plurality of LEDs, of which only a first LED 402 and a second LED 404 are shown in the side view shown. The overview camera 110 is arranged above the glass surfaces of the sample holding device, wherein the glass surfaces include a glass surface 414 of the sample carrier 106 and a glass surface 412 of the cover glass 204.
[0140] The overview camera 110 has an indicated field of view 111 and records an overview image 300 of the sample carrier 106 including the cover glass 204 in the field of view 111. In this example, a first LED 402 of the illumination device of the overview camera 110 is arranged relative to the overview camera 110 and the sample holding device including at least the sample carrier 106 and the cover glass 204 in such a way that the light 406 emitted by the first LED 402 illuminates the edge 410 of the cover glass 204 in such a way that structures 312 can be identified as dark lines in the overview image 300 in which the reflected light 408 is captured, see also FIG. Figure 3 The light 406 emitted by the second LED 404 is reflected by the glass surface 412 of the cover glass 204 facing the overview camera 110, so that the reflected light 408 of the second LED 404 captured by the overview camera 110 in the overview image 300 does not form the structure 312 in the overview image 300. Instead, the aforementioned reflections of the illumination device of the overview camera 110 and, by way of example, of the sample carrier 106 can be seen.
[0141] according to Figure 3 and Figure 4 In the embodiment shown in FIG, the overview camera 110 comprises an illumination device; in this geometry, the presence of the structure 312 can be explained in particular illustratively. A person skilled in the art knows many different forms of the overview camera 110, in particular Figure 1 Another alternative is shown, in which the overview camera 110 observes the sample through a mirror 112 on the objective lens turret 102. In such an embodiment, in which the overview camera 110 does not, for example, include an illumination device arranged around the overview camera 110, the overview image 300 in particular does not have a bright spot 314, or the camera cannot be identified in the overview image 300. A person skilled in the art knows different designs of the overview camera 110 and the corresponding illumination devices, and knows how the corresponding overview image 300 appears. These may be different from Figure 3 and Figure 4 The arrangement shown in is quite different.
[0142] According to this embodiment, Figure 1As schematically shown in FIG, the control device 130 is connected to the monitor 120. The control device 130 is configured to control the microscope 100 to record image data 500 using the microscope camera 107 or the overview camera 110, evaluate the recorded image data 500 using the evaluation module 131, and store the image data 500 in the memory module 132 (see FIG. Figure 5 If necessary, the recorded image data 500 can be displayed on the monitor 120. The control device 130 is configured to process or evaluate the recorded image data 500. The image data 500 particularly include the overview image 300, but can also include microscope images recorded with the microscope camera 107.
[0143] Control device 130 includes not only an evaluation module 131 and a memory module 132, but also a control module 133. The modules of control device 130 are connected to one another via channels 134, whereby they can exchange data with one another via channels 134. Channels 134 are logical data connections between the individual modules. These modules can be designed as either software modules or hardware modules.
[0144] The evaluation module 131 evaluates the input image data 500 and, based on the evaluation, forwards the information to the control module 133 or forwards the evaluation results to the memory module 132 for storage.
[0145] The memory module 132 stores the image data 500 recorded by the microscope 100 and manages the data to be evaluated in the control device 130 .
[0146] The control module 133 can read the image data 500 from the memory module 132 and forward them to the evaluation module 131 for evaluation. In addition, the control module 133 can send control or steering commands (also referred to as control information) to the microscope 100. In particular, the control module 133 can be configured to generate control information based on information obtained from the evaluation module 131.
[0147] In this case, the control information can control the entire microscope 100 or only certain parts. In particular, the control information can include information about the position of the sample in the sample holding device (also called positioning 202), which has been determined by the evaluation module 131. Alternatively, the control module 133 can send control information to the microscope 100 for determining the positioning of the sample in the sample holding device.
[0148] According to the present embodiment, the image data evaluation system 1 is designed to determine the position 202 of the sample in the sample holding device. In particular, the position 202 of the sample is determined based on the image data 500 recorded by the overview camera 110. Alternatively, the microscope camera 107 can also be used to record the overview image 300. In this case, an objective lens 103 with the smallest possible magnification is used so that as many components of the sample holding device as possible are visible in the overview image. The overview image 300 can then be processed as image data 500.
[0149] In particular, the evaluation module 131 may include one or more machine learning models 600. In particular, the machine learning model 600 is implemented as a neural network. According to one configuration, the evaluation module 131 includes one or more processing models that are not machine learning models 600.
[0150] Different processing models are particularly machine learning models 600. Machine learning models 600 (see Figure 6 ) can be a neural network having multiple layers. In particular, machine learning model 600 has an input layer 602, one or more intermediate layers 604, and an output layer 606. Input layer 602 receives input data 608, processes the input data 608 through input layer 302, intermediate layers 604, and output layer 606, and outputs result data 610. The form and scope of input data 608 and result class 310 vary depending on the type or implementation of the processing model used. For some machine learning models 600, intermediate data 612 may also be output.
[0151] In the sense of the present invention, an “input layer” is the first layer of a machine learning model having multiple layers, in particular the first layer of a neural network.
[0152] The machine learning model 600 may be implemented, in particular, as a regressor, a classifier, a segmentation model, or an image-to-image model.
[0153] In the context of the present invention, a "regressor," also called a regression model, is a processing model that performs regression, in particular a machine learning model. If the regressor is a machine learning model, it is trained to perform regression, in particular by supervised or unsupervised learning. The regressor then outputs as result data the probability that the input data is a structure region.
[0154] In the sense of the present invention, “supervised learning” is a learning process in which a machine learning model used to perform the desired mapping is trained using an annotated dataset.
[0155] In the sense of the present invention, “unsupervised learning” training or learning of a machine learning model is training based solely on a non-annotated dataset without specifying a desired objective, wherein the machine learning model automatically finds or should find specific cluster points in the dataset.
[0156] Independent of the implementation, the machine learning model 600 must be trained during training to perform the processing mapping. During the training of the machine learning model 600, the evaluation module 131 controlled by the control module 133 reads some image data 500 of the training dataset from the memory module 132, wherein the training dataset is particularly an annotation dataset, and inputs the training data into each machine learning model 600. The evaluation module 131 determines the objective function based on the output data or result data 610 of the machine learning model 600 and the target data contained in the annotation dataset, and optimizes the objective function by adjusting the model parameters of the machine learning model 600 based on the optimization of the objective function.
[0157] In particular, the optimization of the objective function is performed by the stochastic gradient descent method. In the stochastic gradient descent method, only a small portion of the training data of the annotation dataset is used each time (called a batch). For each input data 608 of the batch, based on the result data 610 output by the machine learning model 600 and the target data of the annotation dataset corresponding to the input data 608, the control module 133 determines an objective function (here a loss function) that captures the difference between the output data 240 and the target data. Thereafter, the control module 133 calculates the gradient of each objective function calculated relative to the model parameters of the machine learning model 600, sums the gradients calculated over the batch and determines an average value. Based on the average value, the control module 133 determines the updated model parameters of the machine learning model 600 by so-called backpropagation. The control module 133 reinitializes the machine learning model 600 in the evaluation module 131 with the updated model parameters and executes the next step of the stochastic gradient descent method.
[0158] Once the objective function reaches a predetermined threshold through optimization of the objective function, the training of the machine learning model 600 is terminated.
[0159] Once training is completed, the control module 133 stores the model parameters most recently used by the machine learning model 600 in the memory module 132, in particular together with the context information, so that the just trained machine learning model 600 can be recognized again later and can be initialized, for example, for further training or inference.
[0160] As an alternative to the stochastic gradient descent method, other methods can also be used. In particular, any other training method can be used.
[0161] Once the training of the machine learning model 600 is completed, the corresponding model parameters are stored in the memory module 132 and can be read out in later inference to execute the learned processing mapping.
[0162] According to some embodiments of the present invention, a machine learning model 600 implemented as a classifier, detector, or segmentation model is particularly used. Furthermore, two different machine learning models 600, each trained with a different training dataset, are used to perform the processing mapping, specifically a selection model 620 and a recognition model 630.
[0163] The selection model 620 is configured to identify the structure regions 306 in the coarse image data 502. The input data 608 of the selection model 620 is the coarse image data 502, and the result data 610 is the structure regions 306, or the structure regions 306 can be determined based on the result data 610.
[0164] Recognition model 630 is configured to determine structural information 316 in image data 500. Input data 608 to recognition model 630 is, for now, structural regions 308 of image data 500. Recognition model 630 is configured to determine structural information 316, wherein result data 610 each depend on the type of implementation.
[0165] Reference below Figures 7 to 11 A method of determining a position 202 of a sample in a sample holding device based on structural information is described.
[0166] In particular, hardware-based solutions for determining the position 202 of a sample in a sample holding device are known from the prior art. For this purpose, in particular special lighting devices are used, which are used to illuminate the structure 312 generated by certain components of the sample holding device (see, for example, Figure 3 ) is more visible in image data 500, particularly due to the improved image contrast provided by the illumination device. This type of illumination device is expensive, not suitable for every microscope type, and the results obtained with it are often flawed, resulting in low customer acceptance. All software-based solutions known to date in the prior art place high demands on computing hardware, are very time-consuming, and still suffer from quality issues.
[0167] In the described method for determining the positioning 202 of a sample based on structural information, image data 500 are first provided in step S1. According to this embodiment, the provision of the image data 500 comprises recording one or more overview images 300 with an overview camera 110. In the overview image 300, the overview camera 110 at least partially captures the sample holding device or elements thereof, in particular the sample carrier 106. These elements may in particular be one or more of the following: the holding frame 105, the sample carrier 106, the sample stage 104, but also a coverslip 204 with which the sample is covered on the sample carrier 106, a spacer arranged between the sample carrier 106 and the coverslip 204. The elements or regions of the elements of the sample holding device captured in the overview image 300 give rise to structures 312 in the image data 500. The induced structures 312 may be in particular edges (e.g. Figure 3 ), spots, textures, characters, shadows, etc. An example of image data 500 is particularly shown in Figure 3 middle.
[0168] For better understanding, Figure 7 The recording of a plurality of overview images 300 with an overview camera 110 is schematically shown. In this case, the overview camera 110 detects objects located within its field of view 111. Consequently, the field of view 111 corresponds precisely to the details captured by the overview camera 110 in the overview images 300. By way of example, a shaded area is drawn here; if elements of the sample holding device are located within the shaded area, these elements can cause structures 312 in the individual overview images 300, which is why the shaded area corresponds precisely to structure areas 308. The shaded area is drawn here merely by way of example; the actual distribution of structure areas 308 may also depend on the orientation of the elements of the sample holding device relative to one another.
[0169] Figure 7 The overview path 702 drawn in FIG3 indicates a possible relative movement of the overview camera 110 via the sample holding device. In this case, it is irrelevant whether the overview camera 110 or the sample holding device is movable, or both are movable; the overview path 702 is indicated here only for a better understanding. If one of the areas drawn in the possible structure area 308 is arranged, for example, above the cover glass 204, a structure 312 is formed in the overview image 300. As a result of the movement of the overview camera 110 relative to the sample holding device, the entire sample holding device can be scanned so that, with a suitably selected overview path 702, all edges 410 are scanned once and captured as a structure 312 in one of the overview images 300. In each case, the relative position of the sample holding device and the overview camera 110 relative to each other is also captured for each recorded overview image 300.
[0170] According to one configuration, overview images 300 of image data 500 are provided directly with coordinate information during storage; this coordinate information can be converted, in particular, into coordinates in a stationary coordinate system of the sample holder, so that identical objects in overview images 300 always have identical coordinates. Alternatively, other coordinate systems can also be selected; however, the advantage of selecting a stationary coordinate system of the sample holder is that objects in overview images 300 are always assigned identical coordinates, which makes them easier to identify and structural regions 308 of different overview images 300 easier to compare, combine, or merge. Depending on the accuracy of the drive of the movable parts of the microscope, the different overview images 300 must also be registered with one another; for this purpose, in particular, edges, spots, or similar image details visible in multiple overview images 300, particularly those stationary relative to the sample holder, can be used.
[0171] In particular, the step for registering the image data 500 of different images with one another can be performed during the entire method.
[0172] According to one configuration of the first embodiment, the overview image 300 can also be recorded using the microscope camera 107. If the microscope camera 107 is used to record the overview image 300, the microscope 100 must be set up or used so that it uses an objective lens 103 with the smallest possible magnification during the recording of the overview image 300, for example 1x, 1.5x, 2x, or 3x to 5x.
[0173] According to another configuration of the first embodiment, the image data 500 include a plurality of overview images 300 captured with a camera, wherein during the recording of the overview images 300 the position of the camera relative to the position of the sample holder was varied, or the position of the sample holder relative to the position of the respective camera was varied, so that in each case a different part of the sample holder was captured at a different position in the overview image 300. In particular, for all overview images 300 the relative position of the sample holder with respect to the respective camera used is in each case stored together with the image data 500.
[0174] According to another configuration of the first embodiment, instead of the overview image, the image data 500 may also include an image sequence, an image stack including a plurality of images offset in height relative to each other, a stereo image, an image with depth information, or one or more of the above images with low contrast.
[0175] The inventors have recognized that an image region in which a structure 312 caused by an element of the sample holding device always appears similarly in the overview image 300 to further image content surrounding the structure 312. The image content surrounding the structure 312 is also referred to as the environment of the structure 312, and the image region in which the structure 312 is typically visible is referred to below as the structure region. Figure 3 Examples of such environments and locations of structural areas 308 are described. Figure 3 In the example, the reflections of the lighting device and the overview camera 110 accurately form the environment, as shown in FIG. Figure 4 The described structures 312, for example caused by the cover glass 204, are only visible when arranged as described above relative to the first LED 402 and the overview camera 110. Furthermore, the inventors have recognized that the image information or image details required for identifying the structure region 308 in relation to its surroundings are also still contained in the coarse overview image 302 with reduced details compared to the overview image 300.
[0176] Therefore, according to a first embodiment, step S1 is followed by step S2, in which coarse image data 502 is determined based on image data 500. Coarse image data 502 has reduced detail compared to image data 500. This detail reduction can be achieved by, inter alia, reducing the bit depth for capturing color and / or intensity, reducing the number of pixels, downsampling, and, in particular, if image data 500 includes an image stack, transferring only some of the highly offset images of the image stack into coarse image data 502. According to another configuration, the image format used by each camera can be an image format with progressive image compression.
[0177] If the coarse image data 502 are stored with progressive image compression, the determination of the coarse image data 502 will correspond precisely to selecting the portion of the data set corresponding to the desired compression or desired degree of detail reduction.
[0178] According to this embodiment, the image data 500 is exactly the overview image 300, and the rough image data 502 is exactly the rough overview image 302. Figure 3 As depicted, a detail-reduced coarse overview image 302 is determined from the overview image 300 , which has fewer pixels than the overview image 300 .
[0179] According to step S3 , the coarse image data 502 are used to determine the structure region 306 based on the coarse image data 502 by means of a selection model 620 .
[0180] The inventors have recognized that the detail-reduced coarse overview image 302 still includes a sufficient amount of detail to train the selection model 620 for determining the structure region 306, provided that during the selection of the coarse image data 502, only a reduction in detail was selected such that the image content, image information, or image details of the environment constituting the structure region 306 are still recognizable in the coarse image data 502. Since the coarse image data 502 contains less detail than the image data 500, the selection model 620 does not necessarily have to learn the fine image structure present in the image data 500. Instead, the selection model learns the relative position of the structure region 306, for example, the structure region 306 relative to the bright spot 314 in the environment. As mentioned above, the bright spot 314 is mentioned here only by way of example; the environment of the structure region 306 may also have a completely different structure than that of the bright spot 314. If, in particular, a different illumination of the sample is selected, the overview image 300 changes accordingly.
[0181] If the selection model were to evaluate image data 500 instead of coarse image data 502, the selection model would have to learn more details, in particular all details of the environment of structure region 306, which is why the model must be more complex and the scope of the training data set must be significantly larger. However, since the coarse image data 502 with reduced detail already contains the required information or required image details, data efficiency during training and inference can be significantly improved by reducing the details, and therefore the computational effort in inference can also be significantly reduced.
[0182] According to the method, in inference, the rough image data 502 is input into the selection model 620 implemented in the evaluation module 131. The selection model is a machine learning model 600, such as Figure 6 As shown. According to the first embodiment, the selection model 620 is implemented as an image-to-image model and outputs a probability map 304 in which a probability is assigned to each entry of the coarse image data 502, i.e., each entry captures the structure 312. In the probability map 304 shown, the brighter areas happen to be areas where the structure 312 is captured with a high probability. In the output probability map 304, the structure region 306 of the coarse image data 502 is determined based on the probability values in the probability map 304. In this case, the determination of the structure region 306 is not necessarily performed directly by the selection model, but is performed by appropriately selecting image regions in the probability map 304 whose probability values are higher than a certain probability, as in the example shown. For example, as shown in the figure, by selecting a rectangular image region from the probability map 304 whose probability values are higher than a certain probability. Then, the structure region 306 happens to be the white border image region of the coarse overview image 302, as shown Figure 8Alternatively, however, all image regions having a probability value higher than a certain probability may simply be selected, and the corresponding image regions or data regions of the image data 500 may be processed further.
[0183] According to one configuration, the selection model 620 may alternatively be implemented as a classifier, a detector, or a segmentation model.
[0184] If the selection model 620 is implemented according to one of the above-mentioned alternatives, the input data 608 input into the selection model 620 and the result data 610 output by the selection model 620 may differ from the input data 608 and result data 610 of the image-to-image model, but in principle the respective result data 610 in turn allow the structural region 306 to be determined.
[0185] According to the first embodiment, step S3 is followed by step S4, in which structural information is determined based on the structural region 306 by means of a recognition model 630. According to the first embodiment, the structural region 308 of the image data 500 corresponding to the structural region 306 of the coarse image data 502 is input to the recognition model 630. The recognition model is particularly implemented as a classifier. Figure 8 and Figure 10 In the embodiment shown in , the recognition model 630 determines a shape category in the input structure region 308, which indicates whether the structure 312 can be recognized in the structure region 308, and if the structure 312 can be recognized, the type of the structure 312. Figure 8 and Figure 3 As shown in the example, the structure 312 is an edge 410, in particular a cover glass edge of a straight cover glass 204, i.e. Figure 8 The result class 310 for the structure area 308 shown on the right in the middle is the shape class "straight cover glass edge" or, alternatively, depending on the number of different shape classes used, only "straight structure". For the structure area 308 shown on the left, the shape class "no structure" is output. Depending on the elements of the sample holding device used, the shape classes to be identified can also include other forms of structures 312, in particular circular structures. The shape classes can include, in particular: no structure, structure, circular structure, straight structure, polygonal structure, straight cover glass edge, polygonal cover glass edge, circular cover glass edge, sample carrier edge, gasket structure, holding frame structure, microtiter plate edge structure, microtiter plate well structure, sample chamber edge structure, sample chamber structure, and alternatively only circular, polygonal, square, rectangular, straight, round, elliptical, no shape or no structure.
[0186] If the image data 500 includes a plurality of overview images 300, a shape class is determined for each of all overview images 300 and all of the structure regions 308. The position of the individual structure regions 308 in the individual overview images 300 and the position of the sample holder relative to the overview camera 110 are also stored in each case together with the specific shape class as structural information 316. In particular, as described above, coordinates in the stationary coordinate system can also be directly assigned to each image pixel of the overview image 300.
[0187] According to a first embodiment, step S4 is followed by step S5 of determining the location 202 of the sample based on the structural information. According to a first embodiment, the structural information comprises a shape class output by a recognition model implemented as a classifier and individual positions in a stationary coordinate system, wherein in each case a position (e.g. a center point) is used here as the position of the individual structure region 308. Figure 8 In the example shown, the recognition model finds exactly one edge 410 in one of the overview images 300 shown, in this example a straight cover glass edge. Figure 8 The shape class of the structure region 308 shown on the right in the middle is correspondingly “straight structure.” With appropriately selected overview path 702 , further edges 410 are found in further structure regions 308 .
[0188] In particular, in the overview image 300 a Figure 8 , the edge 410 of the coverslip 204 of the structure 312 shown in the image 300 may also cause a structure 312 in the further overview images 300. If the same element of the sample holding device (here for example the edge 410 of the coverslip 204) produces a structure 312 in a plurality of overview images 300, then according to some embodiments of the present invention, these multiple structures 312 are also referred to as source-identical structures, and the structural region in which the source-identical structures are captured is a source-identical structure region. Depending on the image quality, the recognition model 630 may determine different shape classes for structures of the same origin. Therefore, for the source-identical structure region, step S5 may in particular also include a step of merging structure information 316. In the sense of the present invention, structures in the overview images are "source-identical" if they are caused by the same element of the sample holding device.
[0189] Merging structure information 316 specifically includes determining the original identical structure regions by determining the positions of the structure regions in the stationary coordinate system of the sample holder and determining, based on these positions and the respective extents of the structure regions, whether the structure regions at least partially overlap. If this is the case, the structure regions 306 are further processed as original identical structure regions 306.
[0190] A majority decision can be made in particular when merging structure information 316. If, for example, the source identical structure region comprises structure regions 308 from six different overview images 300 and the four result data 310 of the recognition model 630 (i.e., here respectively assigned shape classes "straight structure") match, e.g. Figure 8 As shown in the example, and the two result data 310 are different from the others, then according to majority decision, the assigned shape class of the source same structure area is "straight line structure".
[0191] According to one embodiment, the orientation of the structural regions 308 can be taken into account when determining structural regions of the same origin. For example, in the case of a square cover glass 204, even if the structures 312 captured in the structural regions 308 are not caused by the same edge 410, the two structural regions may overlap at one of their corners. In particular, one edge may be aligned vertically and one edge may be aligned horizontally. The orientation of the structural regions 308 can be used in particular to determine whether the structural regions 308 have the same origin, and structural regions 308 with different orientations can then be considered to have different origins. However, according to another embodiment, corners can also be considered to have the same origin. In particular, the recognition model 630 can then output "polygonal structure" as the shape class for such corners.
[0192] In addition to regions of identical structure, structures 312 caused by different elements of the sample holding device may also appear in the overview image 300. For example, each of the four edges 410 of the cover glass 204 may cause a structure 312 in one or more overview images 300. If four edges 410 are identified, the result is a location 202 of the sample, such as Figure 7 As outlined in , to a square area within the found edge. For this type of sample carrier 106 having only one cover glass 204, the center of the sample can be found, for example by determining the center of mass from the found edges 410, when a total of four edges 410 are found. The extent of the sample, and therefore the position 202, can then be determined based on the individual distances of the edges 410 from each other or from the center.
[0193] However, in particular, the position 202 of the sample can also be determined solely based on the edge as being located somewhere in the edge region. Due to the identification of a single edge 410 of the cover glass 204, the position 202 of the sample is already limited compared to the entire overview image, so that, for example, the sample can also be found by a conventional microscope camera 107.
[0194] refer to Figure 9A Another configuration is shown. The sample carrier 106 shown corresponds to Figure 2ASpecimen carrier 106 in FIG. In the schematic diagram, the field of view 111 of overview camera 110 is depicted with dashed lines. Objects within field of view 111 are visible in overview image 300 with good illumination of field of view 111. In the example shown, structure region 306 within field of view 111 is indicated by cross hatching. According to the illustration, from the right-hand cover glass 204, the lower left tip will fall precisely within structure region 306 and should be visible in overview image 300. From the left-hand cover glass 204, a portion of the right edge will fall precisely within structure region 306 and will be visible in overview image 300. With this configuration, for example, in addition to the shape class, the angle of rotation of the cover glass 204 from the horizontal can also be determined. For specimen carriers 106 having multiple cover glasses 204, a position 202 is determined for each cover glass 204, wherein the angle of rotation of each cover glass 204, i.e., the orientation of the cover glass 204, is taken into account when determining position 202.
[0195] Another configuration of the first embodiment is shown in Figure 9B The sample carrier 106 shown corresponds to Figure 2B The sample carrier 106 in FIG. Figure 9A As described, the field of view 111 of the overview camera 110 is again depicted here. Since the sample carrier 106 used is a microtiter plate, corresponding structures 312 must be identified for a certain number of wells 208, depending on this configuration. In particular, for example, the individual wells 208 can be identified based on their edges. Alternatively, it may also be sufficient to identify an outer row of regularly arranged wells 208 along a first direction (e.g., in the x-direction) and an outer row of regularly arranged wells 208 in the y-direction, in particular the distances between the outer wells and the individual wells 208 in each case. For the sake of clarity, the structural areas 308 are not depicted in this illustration; as mentioned above, they depend on the imaging device 100 used.
[0196] As described above with reference to the first embodiment, according to one configuration, when using a sample carrier 106 comprising a microtiter plate, the overview path 702 can be adjusted after identifying the wells 208. For example, the overview path 702 is adjusted so that the regular pattern of wells 208 is first determined and then the outer edge of the pattern is determined, or vice versa, and if the outer edge is found, the extent along the pattern can be determined particularly quickly. To this end, for example, the pattern can be scanned once along the outer edge in the vertical and horizontal directions.
[0197] Another configuration of the first embodiment is shown in Figure 9C The sample carrier 106 corresponds to Figure 2C, and includes a plurality of sample chambers 210. In this configuration, the overview camera 110 differs from the previous configuration, but this schematic diagram should not be considered limiting. A person skilled in the art will be aware of different embodiments of the overview camera 110 and their respective lighting arrangements. As already described with reference to the microtiter plate, in order to identify the individual sample chambers 210, when a suitable structure 312 is found, the pattern and edge can again be appropriately determined; this allows the overview path 702 to be optimized for the most efficient positioning possible.
[0198] Figure 10 Another configuration of the first embodiment is shown. According to the configuration shown, the selection model is implemented as a classifier. The classifier outputs the assigned class for various input image regions or data regions of the rough image data 502 as result data 610 in each case. The selection model is each configured to determine the structure region 306 in the rough image data 502. Therefore, the assigned classes include at least one class assigned to the structure region 306 (i.e., the so-called structure class) and one class indicating that each data region is not a structure region (i.e., the so-called non-structure class). Other classes that the selection model 620 (implemented as a classifier) can identify will, for example, be another class of image regions including the label 206 (i.e., the so-called label class) and another class of image regions in which samples are not reliably captured (i.e., the so-called non-sample class). The classifier can, in turn, be configured in different ways.
[0199] According to one configuration, the class identified by the classifier may include the class of the tag 206. The identified tag 206 may then be read out using, in particular, a tag reading model to determine contextual information. However, the location of the tag 206 may also be used, for example, to determine the position 202.
[0200] For a so-called patch classifier, a so-called sliding window function selects data regions, in particular consecutive adjacent data regions, from the coarse image data 502 respectively and inputs the selected data regions into the patch classifier. For the embodiment described here, the sliding window function selects rectangular or square image parts from the overview image respectively and inputs them into the patch classifier. The patch classifier outputs the found result class for each input data region (here an image part), i.e. whether the image part is a structure region 306 or not. The sliding window function now continuously selects data regions so that finally all coarse image data 502 are input into the patch classifier and the patch classifier assigns or has assigned a class to each input data region. Thus, for Figure 10The data region with a white border in the middle, the data region output by the patch classifier is the structure region 306. In particular, the data regions selected by the sliding window function can be separate or overlapping. If overlapping data regions are used, the white border data region can be selected from the possible multiple overlapping data regions of the class to which the patch classifier assigns the structure region 306 by the non-maximum suppression function, so that the structure region 308 is forwarded only once to the recognition model 630.
[0201] According to a modification, the classifier can also be a modified conventional classifier, here referred to as a graph classifier. The graph classifier is implemented as a convolutional neural network (CNN), wherein the network architecture is modified so that the last pooling layer of the CNN is omitted. In the pooling layer, the outputs of the previous layer of the CNN are combined, for example, to combine the intermediate outputs of the previous layer (in particular different feature maps) to form a result output. The overview image 300 is respectively input as a whole into the graph classifier, and the graph classifier outputs a spatially resolved classification map as result data by omitting the last pooling layer. In the classification map, various partial areas of the coarse image data 502 are assigned a class, which is spatially resolved in each case, which indicates whether the respective partial area is a structural area 306 or a class corresponding to the above-mentioned class.
[0202] If selection model 620 is instead a segmentation model, the segmentation model outputs a segmentation mask as result data 610. A result class is assigned to each entry in the segmentation mask of coarse image data 502 corresponding to the aforementioned class. Thus, result data 610 comprises, for example, a bitmap in which, for example, each entry assigned by the segmentation model as belonging to a structure region is assigned a "1" entry, particularly since each entry captures a structure 312 generated in the image data by the elements of the specimen holder with a certain probability, and entries not belonging to a structure region 306 are assigned a "0" entry. Thus, relevant regions of coarse image data 502 are identified or further processed as structure regions 306. Alternatively, more than two different classes can be assigned; in particular, a class can still be assigned to a found label 206. In the case of a particular occurrence of structure 312, selection model 620 may be able to directly recognize that it is a non-sample region in overview image 300 and then assign it accordingly to the non-sample class.
[0203] According to one configuration of the first embodiment, the recognition model may also output an element class as result data 610. Here, the element class may specifically correspond to each of the possibly different elements of the sample holding device, and the result data 610 then indicates exactly to which element of the sample holding device each captured structure 312 corresponds.
[0204] In particular, the recognition model may be implemented such that it has two different output paths, where one output path outputs the component class and the other output path outputs the component shape.
[0205] In particular, a specific machine learning model 600 (so-called component-specific machine learning model 600) can be stored for each possible element of a sample holding device, different sample holding devices, different sample carriers 106, or different microscope types for use in the selection model 620 and the recognition model 630. The respective specific machine learning model 600 then outputs corresponding result data depending on the structure 312 that occurs.
[0206] According to another configuration, the recognition model 630 can also distinguish between different designs of different component classes. For example, if the element of the sample holding device that causes the structure 312 is a cover glass, the recognition model 630 may have been trained to distinguish between different cover glass shapes and cover glass sizes. For example, circular and polygonal cover glasses and different cover glass sizes for each shape, such as 2, 3, 4 or 5 different cover glass sizes. Then, the recognition model 630 can be specifically designed and trained so that it outputs the cover glass shape, i.e., circular, polygonal or no structure found, and the cover glass size for each structure area 306. For this type of recognition model 630, the shape categories include not only the pure shape categories "circular", "polygonal", "no structure", but also the size categories. Accordingly, corresponding recognition models 630 can be provided for other elements in each case, and these recognition models are correspondingly designed to recognize component shapes and component sizes.
[0207] In particular, a shape class can be provided for each combination of element shape and element size as a possible class to be assigned by the recognition model. Thus, based on the result data 610, not only the shape of the individual structures (e.g. Figure 8 and Figure 10 ), it is also possible to determine from which element of the sample holding device each classified structure 312 originates and which element size causes the classified structure 312.
[0208] However, according to one configuration of the first embodiment, the recognition model can also be designed as a segmentation model. The resulting data is then a segmentation mask for each input structure region 308, in which a class is assigned to each entry. In particular, the assigned class can be one of the classes described above with reference to the recognition model 630 implemented as a classifier. However, alternatively, the segmenter can also be configured to assign each entry to a class from "structure found" and "structure not found". The resulting segmentation mask can then be analyzed in step S5 after step S4 to determine where the sample is located based on the respectively identified structures 312.
[0209] According to another configuration of the first embodiment, the recognition model can also be implemented as an image-to-image model. The result data of the image-to-image model is then a probability map, in which a probability is assigned to each entry in the input data to find the structure 312 at each point. The output probability map can in turn be used to locate the sample. According to another modification, the result data includes a probability distribution for each entry of different result classes. Like the result classes, these are exactly the shape classes described above with reference to the recognition model 630 implemented as a classifier.
[0210] According to the first embodiment, steps S1 to S5 are performed one after another. However, according to modifications of the first embodiment, other orders of the above steps are also possible. Some modifications are outlined below, which show, for example, how steps S1 to S5 can still be combined.
[0211] For example, first, according to step S2, rough image data can be determined, in particular by recording a rough overview image 302. Thereafter, step S3 determines structure regions 306. Based on structure regions 306, image data 500 are then recorded, in particular using a microscope camera 107 with an objective 103 having the smallest possible magnification.
[0212] For example, for each overview image 300 recorded in step S1 along the overview path 702 described above, steps S2 to S5 can be respectively performed, wherein, in particular, if a structure region 306 or a structure 312 is identified in one of the steps, the overview path is appropriately adjusted to configure the overview path 702 as efficiently as possible. This can include, in particular, recording further overview images 300 after the edge 410 of the coverslip 204 is identified, in particular in the vicinity of the identified edge 410. If, for example, non-sample structure regions are identified in the overview image 300, the position of the sample holder and the overview camera 110 is adjusted so that these non-sample structure regions are not captured as much as possible or are captured as little as possible in the overview image.
[0213] According to another embodiment, the merging of the source identical structure regions 306 is performed based on the probability map output by the selection model 620, wherein the source identical structure regions 308 are sorted according to the probability values within the structure regions 308 and then processed by the recognition model 630 such that the structure regions 308 with the highest probability values are processed first, wherein the processing is terminated as soon as the recognition model 630 outputs an unambiguous result. Within the meaning of the present invention, an unambiguous result is particularly a result according to which the structure 312 has been identified, i.e., the structure 312 has been unambiguously identified according to the output of the recognition model 630, for example, by a majority decision. Alternatively, the merging can be terminated after analyzing a predetermined number k of the original identical structure regions 308, in particular if the k original identical structure regions are the structure regions with the highest probability values.
[0214] According to another embodiment, the incorporation of the structural information 316 can be performed using an aggregated statistical model. In the aggregated statistical model, the result data output by the recognition model 630 implemented as an image-to-image model (where each entry is assigned a probability distribution of possible result classes) is fed into the statistical model as input. Based on the probability distribution determined for each structural region 308, the statistical model can then determine which type of structure 312 is involved.
[0215] According to another embodiment, when merging the structural information 316 of the source identical structural regions 308, the result data (i.e., the segmentation mask) output by the recognition model 630 implemented as a segmentation model is processed by a localization model implemented as a classifier. The localization model then in turn outputs a shape classification as the result data based on the input segmentation mask.
[0216] According to another configuration, recording a single overview image 300 is sufficient for determining the position. This may be the case, in particular, if just as many structures 312 as are identifiable in the overview image 300 are sufficient to unambiguously determine the position of the specimen. This is particularly the case, for example, when using a specimen carrier 106 with a cover glass 204 of known dimensions. For example, if corners or, for example, circular arcs are identified in the structure area 308, the specimen can be unambiguously located. This is also true for other types of specimen carriers 106, in particular, if only certain types or designs of specimen carriers 106 or specimen holding devices are used in certain microscopes 100.
[0217] Figure 11 A flow chart of a method according to a first embodiment is schematically shown.
[0218] Step S0 is included here as optional. Step S0 includes training a selection model 620 and a recognition model 630. An annotation dataset is determined for training the selection model 620. In this case, the determination includes recording image data 500, in particular one or more overview images 300, wherein the image data 500 is initially not reduced in detail. In the recorded, non-detail-reduced image data 500, structures 312 are then identified, for example using classical image processing methods. If structures 312 are identified, target data are determined, the exact form of the target data (as described above) depending on the implementation of the respective selection model 620. If target data are determined, the detail of the image data 500 is reduced as described above with reference to the detail-reduced coarse image data 502. Target data are then correspondingly determined for the coarse image data 502. The coarse image data 502 can then be used together with the target data as an annotation dataset for training the selection model 620.
[0219] To train the recognition model 630 , the training data set comprises corresponding target data, in each case depending on the implementation, as described above with reference to the different configurations of the recognition model 630 . In particular, the structure regions 308 are here respectively annotated accordingly.
[0220] The training of the machine learning model 600 can be respectively specialized, for example, for different types or configurations of sample holding devices, imaging devices 100 or different configurations of the image data evaluation system 1. Therefore, when evaluating the image data 500, a respective machine learning model 600 can be selected according to the respective configuration.
[0221] According to one configuration, the overview path 702 is not a predefined path as shown, but is selected randomly or based on certain statistical criteria so that the structure 312 can be located with as high a probability as possible. In particular, after the first structure 312 is found, the next point on the overview path is selected based on the shape of the structure, for example, at the end of the first structure found.
[0222] According to the second embodiment, after step S5, the method further comprises step S6 for controlling an imaging device for capturing the sample based on the positioning of the sample, wherein capturing here particularly refers to automatically capturing the sample, the imaging device 100 being currently controlled such that an image of the entire sample is captured.
[0223] According to a third embodiment, a control device 130 is provided having means for carrying out the method according to the first to third embodiments.
[0224] According to a fourth embodiment, a computer program product is provided, which includes instructions that, when executed by one or more computers, cause the computers to perform the method according to the above-described embodiments.
[0225] According to a fifth embodiment, an image data evaluation system 1 is provided, which comprises the control device 130 according to the third embodiment. The image data evaluation system 1 comprises in particular a microscope.
[0226] According to a sixth embodiment, there is provided a microscope 100 configured to perform the methods according to the first and second embodiments.
[0227] The variants and configurations described with reference to the different figures can be combined with each other. The configurations shown and described are purely illustrative and modifications can be made within their scope.
Claims
1. A method for determining the location of a sample based on structural information, wherein: Elements of a sample holding device that holds the sample in an imaging device form structures in image data captured with the imaging device, including: Provide image data, Determine the structure areas based on the coarse image data with the aid of a selection model, determining structural information based on the structural region by means of a recognition model, and determining a location of the sample based on the structural information, Its characteristics are: The rough image data is reduced in detail compared to the image data, the structure region is a region in the image data where the structure is captured with a certain probability, and the total amount of data of the structure region is smaller than that of the image data.
2. The method according to claim 1, wherein The image data particularly comprises an image captured by a camera of the imaging device and particularly comprises one or more of the following: a plurality of images, wherein, in particular during the capture of at least one pair of said plurality of images, the relative positions of the cameras used with respect to said sample holding device differ from one another, a time series of one or more image stacks, Stereoscopic images, Images with depth information, Low contrast images, or Image captured with a low magnification objective.
3. The method according to claim 1, wherein Compared to the image data, the coarse image data exhibits one or more of the following: Lower sampling depth, Lower image resolution, Lower temporal resolution, or The lower the resolution along the height, the greater the distance between adjacent images of the stack.
4. The method according to claim 1, wherein The sample holding device comprises one or more elements, in particular a holding frame, a slide, a cover slip, a spacer, a sample carrier, a holding frame, an inscription, a mark or a label, and the structure of the sample holding device captured in the image data is in particular visible as light or dark lines, light or dark arcs, arcs or circles, so-called spots, in particular light or dark image areas, so-called dots, distortions, mirror images, doublings, textures or characters.
5. The method according to claim 1, further comprising: Based on the image data, coarse image data are determined, and in particular determining structure regions in the coarse image data, and A structural region of the image data corresponding to the structural region of the coarse image data is selected, wherein the structural regions corresponding to each other capture the same element of the sample holding device.
6. The method according to claim 1, wherein The selection model is a machine learning model implemented as a classifier, a detector, a segmentation model, or an image-to-image model, and the determination of the structural region includes: inputting at least a partial region of the rough image data as input data into a selection model, Output the resulting data, and in particular The structure region is selected from the image data based on the result data.
7. The method according to claim 6, wherein: The selection model is configured as an image-to-image model, wherein the result data is a probability map, in which probability values are assigned to entries of the input data, the probability values indicating the probability of the corresponding entry capturing a structure, in particular a probability value is assigned to each entry or respectively a group of entries, and the determination of the structure region comprises grouping the entries of the image data based on the probabilities, in particular consecutive entries of the coarse image data form a structure region having a probability value higher than a certain probability.
8. The method according to claim 7, wherein: The determination of the structural information comprises inputting structural regions of the image data into the recognition model according to a sequence, the sequence being determined based on probability values of the structural regions, in particular structural regions with higher probability values being sorted in sequence before structural regions with lower probability values, and in particular inputting of the structural regions is terminated once a certain amount of structural information has been determined, or only a predetermined number of structural regions with the highest corresponding probability values are input into the recognition model, and no further structural regions with lower probability values are input into the recognition model.
9. The method according to claim 1, wherein The recognition model is implemented as a classifier, a segmentation model, a detector, or an image-to-image model.
10. The method according to claim 9, wherein: The recognition model is implemented as a segmentation model, and the result data includes segmentation masks for the respective structure regions, wherein a shape category is assigned to each entry of the input data and structure information is determined based on the segmentation mask, the shape categories particularly including one or more of the following categories: no structure, structure, circular structure, straight structure, polygonal structure, straight coverslip edge, polygonal coverslip edge, circular coverslip edge, sample carrier edge, gasket structure, retaining frame structure, microtiter plate edge structure, microtiter plate well structure, sample chamber edge structure, sample chamber structure.
11. The method according to claim 10, wherein: The determination of the structure information is performed based on a segmentation mask, wherein a mask classifier determines a shape class based on the segmentation mask and assigns specific structure information to the shape class.
12. The method according to claim 10, wherein: The result data output for different structure areas are merged with the remaining image data to form detail-reduced result data, the corresponding result data are assigned to the structure areas, and the values of the non-structural shape category are assigned to the entries of the remaining image data, and the positioning is determined based on the detail-reduced result data.
13. The method according to claim 9, wherein: The recognition model is configured as an image-to-image model, wherein each entry of the input data is assigned a value in the result data, the value indicating whether the respective entry captures a structure, wherein the value is in particular a probability and in particular the result data is a probability map.
14. The method according to claim 1, wherein The determination of the localization comprises merging structure information of a plurality of source identical structure regions from different overview images, wherein one or more structures of each structure caused by the same element of the sample holding device are captured in the source identical structure regions.
15. A method of controlling an imaging device for capturing a sample based on the positioning of the sample, wherein: The positioning has been determined according to claim 1, the method further comprising: The imaging device is controlled based on the determined positioning.
16. A method for training a selection model for determining object regions based on coarse image data, wherein: The selection model is particularly trained to perform the method according to claim 1, comprising: Provide image data, Identify structural regions in image data, determining coarse image data having reduced detail compared to the image data, determining an object region corresponding to the structure region, Coarse image data is provided as input data and target data, based on which structural regions can be identified, as an annotation dataset for training the selection model.
17. A control device for controlling an image data evaluation system (1), in particular designed as a microscope, comprising means for carrying out the method according to claim 1.
18. An imaging device (100), in particular designed as a microscope, comprising a control device (130) according to claim 17.
19. An image data evaluation system (1) comprising at least one imaging device (100) according to claim 18.
20. A computer program product comprising instructions which, when executed by one or more computers, cause the computers to perform the method according to claim 1.