Machine learning techniques for detecting artifact pixels in images
Patent Information
- Application Number
- JP2026098352
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-10-15
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-08
Smart Images

Figure 2026143738000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application is a "Machine-Learning" application filed on October 15, 2021. Priority is claimed to be granted by U.S. Provisional Patent Application No. 63 / 256,328, entitled “Techniques For Detecting Artifact Pixels In Images,” the contents of which are incorporated herein by reference in their entirety for any purpose. [Background technology]
[0002] Background of the Invention Immunohistochemical (IHC) assays enable the visualization and quantification of biomarker locations, playing a crucial role in both cancer diagnosis and oncology research. In addition to the "gold-standard" DAB(3,3'-diaminobenzidine)-based IHC assays, recent advancements have been seen in both bright-field multiplex IHC assays and multi-fluorescence IHC assays. These multiplex IHC assays can be used, among other things, to identify multiple biomarkers within the same slide image. Such assays not only improve efficiency in identifying biomarkers on a single slide but also facilitate the identification of additional properties associated with such biomarkers (e.g., co-localized biomarkers).
[0003] Quality control of slide images can be performed to improve performance and reduce errors in digital pathology analysis. In particular, quality control enables digital pathology analysis to accurately detect diagnostic or prognostic biomarkers from slide images. Quality control may include, among other things, detecting and excluding pixels in the slide image that are expected to depict one or more image artifacts. Artifacts may include tissue folds, foreign bodies, blurred image areas, and any other distortions that interfere with the accurate display of corresponding areas of a biological specimen. For example, tissue folds present in a biological specimen may blur one or more parts of the image. These artifacts can cause errors or inaccurate results in subsequent digital pathology analysis. For example, artifacts detected in a slide image may lead to errors in counting the number of cells detected in digital pathology analysis, or a group of tumor cells may be incorrectly identified as normal. In fact, artifacts can cause inaccurate diagnoses of subjects associated with slide images. [Overview of the project]
[0004] overview In some embodiments, a method is provided for generating training data to train a machine learning model to detect predicted artifacts in an image. This method may include accessing an image that displays at least a portion of a biological sample. Furthermore, the method may include applying an image preprocessing algorithm to the image to generate a preprocessed image. In some cases, the preprocessed image includes a plurality of labeled pixels. Each of the plurality of labeled pixels can be associated with a label that predicts whether the pixel accurately displays a corresponding point or region of at least a portion of the biological sample.
[0005] Furthermore, this method may include applying a machine learning model to the preprocessed image to identify one or more labeled pixels from multiple labeled pixels. In that case, one or more labeled pixels are predicted to have been incorrectly labeled by the image preprocessing algorithm. Furthermore, the method may include correcting the label for each of the one or more labeled pixels. Furthermore, the method may further include generating a training image that includes at least one or more labeled pixels with corrected labels. Furthermore, the method may include outputting the training image.
[0006] In some embodiments, a method is provided for training a machine learning model to detect predicted artifacts in an image at a target image resolution. The method may include accessing a training image that displays at least a portion of a biological sample. In some cases, the training image includes a plurality of labeled pixels, each of which is associated with a label that predicts whether the pixel accurately displays a corresponding point or region of at least a portion of the biological sample.
[0007] Furthermore, the method may include accessing a machine learning model that includes a set of convolutional layers. In some cases, the machine learning model is configured to apply each convolutional layer of the set of convolutional layers to a feature map representing the input image. Additionally, the method may include training the machine learning model to detect one or more artifact pixels in an image at a target image resolution. In some cases, the artifact pixels among the one or more artifact pixels are predicted not to accurately represent at least a portion of a point or region of a biological sample.
[0008] In some cases, training includes, for each labeled pixel among multiple labeled pixels of a training image, (i) determining a first loss for the labeled pixel at a first image resolution by applying a first convolutional layer of a set of convolutional layers to a first feature map representing a training image at a first image resolution; (ii) determining a second loss for the labeled pixel at a second image resolution by applying a second convolutional layer of a set of convolutional layers to a second feature map representing a training image at a second image resolution, wherein the second resolution has a higher image resolution than the first; (iii) determining a total loss for the labeled pixels based on the first and second losses; and (iv) determining, based on the total loss, that the machine learning model has been trained to detect one or more artifact pixels at a target image resolution. Furthermore, the method may include outputting the trained machine learning model.
[0009] In some embodiments, methods are provided for using a machine learning model to detect artifacts predicted at a target image resolution. These methods may include accessing an image displaying at least a portion of a biological sample, wherein the image is at a first image resolution. Furthermore, these methods may include accessing a machine learning model trained to detect artifact pixels in an image at a second image resolution. In some cases, the first image resolution has a higher resolution than the second image resolution.
[0010] Furthermore, the method may include transforming the image to generate a transformed image that displays at least a portion of the biological sample at a second image resolution. The method may also include applying a machine learning model to the transformed image to identify one or more artifact pixels from the transformed image. In some cases, it is predicted that one or more artifact pixels do not accurately represent a point or region of at least a portion of the biological sample. Furthermore, the method may include output containing one or more artifact pixels. It may include generating force.
[0011] Some embodiments of this disclosure include a system comprising one or more data processors. In some embodiments, the system includes a non-temporary computer-readable storage medium containing instructions such that, when executed on one or more data processors, these instructions cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more processes. Some embodiments of this disclosure include a computer program product tangibly embodied in a non-temporary machine-readable storage medium, which includes instructions configured to cause one or more data processors to execute some or all of the methods disclosed herein and / or some or all of one or more processes.
[0012] The terms and expressions used are for illustrative purposes only, not limitation, and in using such terms and expressions, there is no intention to exclude any equivalent or part of any of the features shown and described, but it should be understood that various modifications are possible within the scope of the invention as described in the claims. Accordingly, although the invention as described in the claims is specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein can be made by those skilled in the art, and such modifications and variations will be considered within the scope of the invention as defined by the appended claims.
[0013] The file of this patent or this application contains at least one drawing prepared in color. A copy of the publication of this patent or this patent application with color drawings will be provided by the competent authority upon request and payment of the required fee.
[0014] The features, embodiments, and advantages of the present disclosure will be better understood when the following detailed description is considered with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] [Figure 1] shows an example set of images including artifact pixels. [Figure 2] shows a flowchart illustrating an example process for generating training data according to some embodiments. [Figure 3] shows a flowchart illustrating an example process for using a machine learning model to generate training data according to some embodiments. [Figure 4] shows an example set of images with generated labels according to some embodiments. [Figure 5] shows an example image including image portions each showing varying levels of blur. [Figure 6] shows an example set of images including one or more tissue folds in a corresponding biological sample. [Figure 7] shows an example schematic diagram for generating training data relating to tissue fold regions according to some embodiments. [Figure 8] comprises an example set of three-class tissue fold masks according to some embodiments. [Figure 9] shows an example image (e.g., FOV) including a tissue fold region, where the tissue fold region displays both non-blurred and blurred regions. [Figure 10] shows an example set of images displaying various artifact regions. [Figure 11] This shows an exemplary artifact mask generated using composite artifact classification techniques according to several embodiments. [Figure 12] Further artifact masks using composite artifact classification techniques in several embodiments are shown. [Figure 13] The slide images show that pixels in several embodiments are associated with two or more classification labels. [Figure 14] The process for identifying classification labels for pixels associated with both blurred and tissue fold regions according to several embodiments is shown. [Figure 15] This shows a comparison between classifications predicted by ground truth masks in several embodiments. [Figure 16] This shows the predicted region in slide images from different types of IHC assays according to several embodiments. [Figure 17] Each image shows an exemplary set of images stained using a staining protocol corresponding to a specific type of IHC assay. [Figure 18] A schematic diagram illustrates an exemplary architecture used to train a machine learning model to detect artifacts in images according to several embodiments. [Figure 19] A flowchart illustrating an exemplary process for training a machine learning model to accurately detect artifact pixels in several embodiments is shown. [Figure 20] A flowchart illustrating the process for training a machine learning model to detect artifact regions within slide images according to several embodiments is shown. [Figure 21] A flowchart illustrates an exemplary process for using a trained machine learning model to accurately detect artifact pixels in several embodiments. [Figure 22]This shows an exemplary set of graphs that identify the accuracy and recall scores of a machine learning model trained to detect artifact pixels. [Figure 23] This shows a set of exemplary image masks generated for a set of images that are not visible from the same type of assay and the same type of tissue as the training images. [Figure 24] This shows an exemplary set of image masks generated from images that reveal invisible assay patterns or tissue types. [Figure 25] Examples of computer systems for implementing some of the embodiments disclosed herein are shown. [Modes for carrying out the invention]
[0016] Detailed explanation I. Overview The following examples are provided to illustrate specific embodiments. In the following description, specific details are given for illustrative purposes to provide a complete understanding of the examples in this disclosure. However, it will be apparent that various examples may be carried out without these specific details. For example, apparatus, systems, structures, assemblies, methods, and other components may be shown as components in block diagram form so as not to obscure the examples with unnecessary details. In other cases, well-known apparatus, processes, systems, structures, and techniques may be shown without the necessary details so as not to obscure the examples. The drawings and descriptions are not intended to be limiting. The terms and expressions used in this disclosure are for illustrative purposes only, not limiting purposes, and the use of such terms and expressions is not intended to exclude equivalents of any feature or part thereof shown and described. The term “example” is used herein to mean “serving as an example, case, or illustration.” Any embodiment or design described herein as an “example” should not necessarily be construed as being preferable or advantageous to other embodiments or designs.
[0017] Several techniques for detecting artifacts have been used for quality control of slide images. Exemplary techniques include observing a given image and manually identifying sets of pixels within the image that are predicted to display one or more artifacts. However, manual identification of artifacts can be time-consuming. The manual process relies heavily on expertise to examine each image and accurately determine whether it contains artifacts. Furthermore, the classification of certain types of artifacts is subjective and can vary from expert to expert. For example, one expert might label a set of pixels in a slide image as representing a blurred tissue area, while a second expert might label the same set of pixels in the same slide image as an unblurred tissue area. Such discrepancies that can arise in artifact identification can reduce the accuracy of subsequent digital pathology analysis.
[0018] As an alternative to manual identification, machine learning models can predict which pixels display artifacts. While these machine learning models have been successful in detecting certain types of artifacts, their accuracy is limited by several factors. For example, these factors may stem from existing training techniques, as follows: (1) the inability to efficiently generate accurate training data; (2) the inability to efficiently train machine learning models to accurately detect artifacts in images with various staining patterns; and (3) the inability to integrate artifact detection machine learning models into subsequent digital pathology analyses (e.g., cell classification models, image segmentation techniques) while minimizing increased processing time and consumption of computing resources. In fact, existing machine learning models typically require considerable computing resources and processing time for training and testing. Furthermore, using existing machine learning models to detect artifacts can significantly increase processing time and consume a large amount of computing resources for subsequent digital pathology analyses (e.g., cell classification). As will be described in detail below, embodiments of this application can address each of the three factors to optimize artifact detection performance and improve efficiency.
[0019] A. Generating training data The first factor that can impair accurate artifact detection by machine learning models includes existing training techniques that cannot efficiently generate accurate training data. Existing techniques may include manually annotating sets of pixels in an image that may not accurately represent the corresponding part of a biological sample. However, manually annotating artifacts in slide images can be time-consuming. This problem can be further exacerbated when machine learning models require large amounts of training data to achieve an acceptable level of performance.
[0020] In addition to the above, manual annotation can result in inconsistent training data. Generally, several experts are involved in manually annotating images to generate training data. As mentioned above, each expert may have a different perspective on how a given set of pixels in an image should be considered blurred, especially when the image contains pixels with varying levels of blur. For a particular set of pixels in an image, the annotation from the first expert (e.g., not blurred pixels) may be the opposite of the annotation from the second expert (e.g., blurred pixels). Such differences in perspective can lead to inconsistencies in the training data. These inconsistencies can result in the machine learning model being trained to perform below the optimal level of accuracy. Therefore, it is necessary to generate consistent and accurate training data while simultaneously reducing the time required for its generation.
[0021] To address the above challenges, some embodiments have determined, for each pixel of the image, whether the pixel accurately represents the corresponding point or region of a biological sample (e.g., a stained sample). Techniques are used to generate identifying labels. These labels can then be used as training data to train a machine learning model. In some cases, the labels identify whether a pixel displays at least a portion of a blurred image. Pixels to which a “blurred” label is associated can be determined by estimating the amount of blur in the pixel and comparing the estimate to a blur threshold. As used herein, the term “blur threshold” corresponds to the level of blur that is expected to result in a degradation of the classification model's performance beyond an acceptable threshold. If the estimate exceeds the blur threshold, the label may indicate that the corresponding pixel does not accurately represent the corresponding point or region of a biological sample. In some cases, the blur threshold is determined by performing a digital pathology analysis on other images at a particular level of blur, determining that the output of the digital pathology analysis results in a result below an acceptable threshold (e.g., the amount of misclassified pixels), and setting this particular level of blur as the blur threshold.
[0022] This technique may include the use of image preprocessing algorithms to generate initial label sets and machine learning models to correct those initial label sets. For example, image blur detection can be applied to an image to generate a preprocessed image. The preprocessed image can identify the initial label of each pixel in the image. A machine learning model can be applied to the preprocessed image to correct the label of each mislabeled set of pixels. The image with the corrected pixel sets can be used as a training image to train a model for detecting artifacts in an image.
[0023] B. Training a machine learning model to accurately detect artifact pixels. A second factor that can impair the accurate artifact detection of machine learning models is the lack of existing training techniques that can efficiently train machine learning models to detect artifacts in images with various staining patterns. In particular, multiple stains corresponding to specific types of IHC assays (e.g., Ki67 IHC assay) can be applied to tissue samples to determine a specific diagnosis or prognosis of a subject. Images displaying such tissue samples may show distinct staining patterns.
[0024] Recent advances in IHC assay technology have made it easier to detect multiple biomarkers in a single image. For example, fluorescence-based IHC assays can use multispectral imaging to separate several different fluorescence spectra, which can enable the precise identification of multiple antigens on the same tissue section. However, these multiplex IHC assays can result in more complex staining patterns compared to single IHC assays (e.g., IHC assays targeting a single type of antigen). However, training a single machine learning model to detect artifacts across images with complex staining patterns can be challenging, especially when various types of IHC assays are considered. Existing techniques may include training a machine learning model with a first set of training images corresponding to a first type of assay, and then training the machine learning model with a second set of training images corresponding to a second type of assay. In some cases, the machine learning model is trained with sets of training images collected from several IHC assays under study. These techniques can lead to time-consuming labeling and training processes. Therefore, there is a need to efficiently train machine learning models to detect artifacts in images with diverse staining patterns.
[0025] To address the above challenges, some embodiments include techniques for training a machine learning model to detect artifacts in images having various staining patterns. The technique may include accessing a training image that displays at least a portion of a biological sample. The training image contains a plurality of labeled pixels, each with a label associated with it. The labels indicate that the pixels correspond to at least a portion of the biological sample. It predicts whether a point or region is accurately represented. For example, a pixel that displays an out-of-focus area of a biological sample can be labeled as not accurately representing the corresponding region.
[0026] In some cases, the training image is converted to a grayscale image. The grayscale image is used to train machine learning to detect artifact pixels. As used herein, “artifact pixel” refers to a pixel that is predicted not to accurately represent a corresponding point or region of at least a portion of a biological sample. In some cases, the artifact pixel is predicted to represent at least a portion of the artifact. For example, the artifact pixel may be predicted to represent a portion of a blurred area in a given image, or a portion of a foreign object (e.g., hair, dust particles, fingerprints) shown in the image. In addition to or instead of this, the training image is converted to a preprocessed image by converting its pixels from a first color space (e.g., RGB) to a second color space (e.g., L*a*b). The first color channel (e.g., L channel) of the second color space can be extracted and used to train a machine learning model to detect artifact pixels.
[0027] In some cases, a set of image features can be added to a training image to train a machine learning model. For example, a set of image features can include a matrix of image gradient values. The image gradient value matrix can identify the image gradient value of each pixel in the training image. The image gradient value indicates whether the corresponding pixel corresponds to an edge of an image object. In some cases, the image gradient value matrix is determined by applying a Laplacian of Gaussian (LoG) filter to the training image.
[0028] A machine learning model can include a set of convolutional layers. Each convolutional layer can be configured to contain one or more filters (or "kernels"). For each pixel, the parameters of each filter in the set of convolutional layers can be modified by backpropagating a loss based on a comparison between the output of the set of convolutional layers and a value representing the pixel's label.
[0029] In some cases, the machine learning model includes, or corresponds to, a machine learning model that includes a reduction path and an expansion path. For example, the machine learning model may include, or be, a U-Net machine learning model. The reduction path may include a first set of processing blocks, each processing block corresponding to processing a training image at a corresponding image resolution. For example, a processing block may include applying two 3x3 convolutions (convolutions without padding) to an input (e.g., a training image), each convolution followed by a normalized linear unit (ReLU). Thus, the output of the processing block may include a feature map of the training image at the corresponding image resolution. Furthermore, the processing block includes a 2x2 maximum pooling operation with a stride of 2 for downsampling the feature map of the processing block to a subsequent processing block that can repeat the above steps at a lower image resolution. In each downsampling step, the number of feature channels may be doubled.
[0030] Following the reduction pass, the expansion pass includes a second set of processing blocks, each processing block corresponding to the processing of the feature map output from the reduction pass at the corresponding image resolution. For example, the processing blocks in the second set of processing blocks receive the feature map from the preceding processing block, apply a 2x2 convolution ("de-convolution") that halves the number of feature channels, and concatenate the feature map with the cut-off feature map from the corresponding processing block of the reduction pass. The processing block can then apply two 3x3 convolutions to the concatenated feature map, each convolution followed by a ReLU. The output of the processing block is the corresponding image resolution. It includes a feature map of image resolution, which can be used as input for subsequent processing blocks at higher image resolutions. Processing blocks can be applied until the final output is generated. The final output may include an image mask. The image mask can identify sets of artifact pixels, each of which is predicted not to accurately represent at least a portion of a point or region of the biological sample.
[0031] In some cases, the loss in each processing block of the second set of processing blocks can be calculated and used to determine the total loss of the U-Net machine learning model. For example, the total loss can correspond to the sum of the losses generated from each of the second set of processing blocks. The total loss of the U-Net machine learning model can then be used to learn the parameters of the U-Net machine learning model (e.g., the parameters of one or more filters in the convolutional layer). In some cases, the loss of each processing block in the second set of processing blocks can be determined by applying a 1x1 convolutional layer to the feature map output by the processing block to generate a modified feature map, and then determining the loss from the modified feature map.
[0032] In addition to this, or alternatively, a set of machine learning models can be trained by training each model in the set at a specific image resolution. This set of machine learning models can then be used to determine a target image resolution for detecting artifact pixels in a slide image. In some cases, the output from a machine learning model trained at a lower image resolution is compared to a set of labels from a training image at a higher image resolution to determine the minimized loss. If the minimized loss indicates that the output can detect artifact pixels within an acceptable level of accuracy, the machine learning model can be deployed to detect artifact pixels in the image at a lower image resolution. For example, if a machine learning model can be trained to detect artifact pixels within an acceptable level of accuracy at 5x resolution, there is no need to deploy machine learning models at higher image resolutions (e.g., 10x, 20x, 40x). In this way, the inference time for artifact detection can be reduced to 1 / 16th compared to another machine learning model processing the image at 20x the original image resolution.
[0033] C. Implementation of a machine learning model for detecting artifact pixels in an image. A third factor that may impair the accurate artifact detection of machine learning models is the existing training techniques that cannot incorporate machine learning models for artifact detection into subsequent digital pathology analyses while minimizing increased processing time and consumption of computing resources. In particular, existing digital pathology analyses for detecting objects of interest (e.g., tissues, tumors, lymphocytes) in a given whole slide image may include dividing the whole slide image into a set of smaller image tiles. For each image tile in the set of image tiles, an analysis may be performed on the image tile to determine the classification of each image object appearing in the image tile. Thus, incorporating artifact detection into such a digital pathology analysis may include, for each image tile in the set of image tiles of the image, (i) applying a machine learning model to detect artifact pixels in the image tile, (ii) removing the detected artifact pixels from the image tile, and (iii) performing a digital pathology analysis (e.g., an image segmentation algorithm) to classify the image objects displayed in the image tile from which the artifact pixels are removed. Applying multiple algorithms to each image tile can increase processing time and consume additional computing resources, potentially making digital pathology analysis as a whole inefficient.
[0034] Furthermore, digital pathology analysis uses high image resolution to achieve accurate results. This may require scanned images. For example, machine learning models used in digital pathology analysis to detect tumor biomarkers in images may require scanning images at 20 or 40 times the original image resolution. Therefore, tumor biomarker detection can already be resource-intensive and time-consuming. If machine learning models for detecting artifact pixels require the same image resolution, the processing time for detecting tumor biomarkers may increase even further. Therefore, machine learning models for artifact detection should be integrated into digital pathology analysis to keep the increase in processing time and computing resource consumption at an acceptable level.
[0035] To address the above challenges, some embodiments include techniques for using different image resolutions to detect artifact pixels in an image. In some cases, artifact pixels are predicted to represent a portion of an artifact. Artifact pixels can be detected during the scanning of a slide and / or after the generation of a digital image of the slide. In some embodiments, a machine learning model is trained to generate an image mask containing a set of pixels from the image. The set of pixels in the image mask represents artifact pixels, which are predicted not to represent at least a portion of a point or area of a biological sample. The machine learning model is further trained to process images at specific image resolutions to generate image masks. Thus, in some cases, an image with a higher image resolution is converted to a lower image resolution, the machine learning model is applied to the converted image, and an image mask is generated. In addition to or instead of this, the machine learning model can be further trained to identify the amount of artifact pixels (e.g., the ratio of artifact pixels to the total number of pixels in the image). For example, the estimated amount may include the predicted count of artifact pixels, the cumulative area corresponding to multiple or all artifact pixels, the ratio of slide area or tissue area corresponding to the predicted artifact pixels, etc.
[0036] In some cases, an image is divided into sets of image tiles. A machine learning model can be applied to each image tile in the set to generate an image mask. The image mask identifies a subset of image tiles, and each image tile in the subset may display one or more artifact pixels. The image mask can then be applied to the image to allow the user to deselect one or more image tiles from the subset, and the deselected image tiles are excluded from further digital pathology analysis. Alternatively, the image mask can be applied to the image to select a subset of image tiles from the image without user input, and then exclude them from further digital pathology analysis.
[0037] A trained machine learning model can be applied to an image at a specific point in time to generate an image mask. For example, a machine learning model can be applied to an existing scanned image to generate an image mask. In another example, a machine learning model can be applied while the image is being captured by the scanning device. Alternatively, a preview image (e.g., a thumbnail image) can be initially captured by the scanning device. Image preprocessing algorithms, such as a blur detection algorithm, can be applied to the preview image. If a tissue region is detected in the preview image, an initial image displaying the biological sample can be scanned. The initial image can display the biological sample at the target image resolution.
[0038] By applying a machine learning model to the initial image, it is possible to generate an image mask that identifies predicted artifact pixels and to identify the amount of artifact pixels present in the image. If the amount of artifact pixels exceeds an image region threshold, a warning can be generated indicating that the image is unlikely to produce accurate results when subsequent digital pathology analysis is performed. In some cases, the artifact region threshold corresponds to a value representing the relative size of the image portion within the image (e.g., 40%, 50%, 60%, 70%, 80%, 90%). If the amount of artifact pixels exceeds the artifact region threshold, it can be predicted that one or more artifacts occupy a large portion of the image and are therefore likely to cause a decrease in the performance of subsequent digital pathology analysis (e.g., cell classification). In such cases, the image can be rejected (e.g., automatically or in response to receiving user input corresponding to an instruction to reject the image) and / or the biological sample can be rescanned to capture a different image. In some cases, an image mask is overlaid on the image to display the image with the predicted artifact pixels on the user interface. In addition to or instead of this, the application of machine learning models and the generation of warnings can be performed on each image tile of the set of image tiles that make up the image. In this way, the decision to rescan the biological sample (e.g.) can be made before the entire image is scanned, thus saving additional processing time and reducing the use of computing resources.
[0039] II. Generating training data for training a machine learning model to detect artifact pixels To improve digital pathology for the accurate detection of diagnostic or prognostic biomarkers from slide images, quality control can be implemented to detect and remove artifacts from slide images. Artifacts can include tissue folds, foreign bodies, blurred image areas, and any other image distortions. Figure 1 shows a set of exemplary images 100 containing artifact pixels. As shown in Figure 1, exemplary image 102 shows a biological sample with acceptable focus quality, and the cellular phenotyping classification results are shown superimposed as dots. Red dots correspond to positively stained cells. Black dots correspond to negatively stained cells. In contrast, exemplary image 104 shows the same image as image 102, but with a set of artifact pixels on the left. The small number of red dots in exemplary image 104 indicates that the cellular phenotyping classification model was unable to identify all the positively stained cells present in the biological sample. Due to the artifact pixels, the cellular phenotyping classification model was unable to perform its corresponding digital pathology analysis. The illustrative image shown in Figure 1 clearly shows the pixels that display the blurred portion of the image, but other images may contain pixels that display blurred portions with varying levels of blur.
[0040] To improve the efficiency of generating training data from slide images (e.g., image 104 in Figure 1), some embodiments include accelerating label collection for overall slide image quality control. In some cases, the proposed framework is applicable to label collection for other types of digital pathology analysis.
[0041] Regarding artifact identification, two options exist: (1) pixel-by-pixel classification using an image segmentation method in which a classification label is assigned to each image pixel, and (2) tile-by-tile classification using an image classification method in which a classification label is assigned to each image tile. As used herein, an image tile refers to a portion of an image (e.g., a rectangular portion, a triangular portion) that contains a set of pixels. An image tile may display a corresponding point or region of a biological sample, such as cells and / or biomarkers. In some cases, a given slide image contains multiple image tiles, and the number of image tiles may range from tens, hundreds, or thousands. The image tiles may be distributed such that they cover the entire image or a region of interest within the image.
[0042] To generate training data, pixel-by-pixel classification is used to identify whether each image pixel passes quality control or fails (e.g., is the pixel blurry). This provides greater flexibility for downstream analysis compared to tile-by-tile classification. Flexibility can stem from the pixel-level accuracy provided by the image segmentation algorithm. Alternatively, or in addition to this, tile-based classification can be used to generate pixel-based classification.
[0043] A. Framework for generating training data Figure 2 shows a flowchart illustrating an exemplary process 200 for generating training data according to several embodiments. Images for generating training data can be accessed. In some cases, the images show at least a portion of a biological sample. The images may be slide images showing tissue sections of a particular organ. In some examples, the biological sample is stained using one or more types of assays (e.g., IHC, H&E).
[0044] In step 202, a specific quality control problem is determined. This specific quality control problem may include the detection of artifact pixels. In addition, or alternatively, the quality control problem may include the detection of other types of artifacts, such as foreign objects, tissue folds, or any other image objects or distortions that result in an inaccurate representation of a portion of a biological specimen. As step 204, it is determined whether an existing deep learning model or an existing labeled dataset exists for a similar purpose with the same image configuration. If such resources are available, process 200 proceeds to step 206, where initial labels are generated by (1) performing inference on the target dataset using an existing model designed for a similar purpose, and (2) training the relevant model on an existing labeled dataset and then applying such a model to the target dataset. If such resources (i.e., models or labeled datasets) are from different image configurations or image distributions, unsupervised domain adaptation can be leveraged to adapt the existing model to the unlabeled target dataset.
[0045] If none of the aforementioned resources are available or valid, process 200 proceeds to step 208 to determine whether the quality control problem can be reduced to an image processing problem. If so (the "yes" path from step 208), an image preprocessing algorithm can be applied to predict labels (step 210). As a result, an initial set of labels can be generated. Each label can predict whether the corresponding pixel in the image accurately represents a corresponding point or region of a part of a biological sample. In some cases, the image preprocessing algorithm may include image segmentation, morphological processing, image thresholding, image filtering, image contrast enhancement, blur detection, other image preprocessing algorithms, or a combination thereof. In addition to or instead of this, the image preprocessing algorithm may include using one or more other machine learning models to preprocess the image so that a set of initial labels can be generated.
[0046] For example, an image preprocessing algorithm may include blur detection to predict artifact pixels, and blur detection includes image filtering by calculating the image gradient and then thresholding to identify low-gradient pixels. A set of pixels with a low image gradient can be defined as a group of adjacent image pixels with relatively small intensity variations. In particular, pixels with a low image gradient are considered to have more uniform pixel intensity compared to pixels with a relatively high image gradient. In another example, an image preprocessing algorithm may include tissue fold detection to predict tissue folds (i.e., where one part of the tissue overlaps with another, creating darker tissue areas). Tissue fold detection may include identifying sets of pixels with low image intensity that are significantly darker than other tissue areas. Sets of pixels can be identified by first applying an image filter (in this case, using a smoothing kernel such as a Gaussian filter) and then applying intensity thresholding.
[0047] If an image preprocessing algorithm is unavailable or invalid (the "No" path from step 208), one or more weakly supervised image processing models can be used to generate initial labels (step 212). For example, a learning-based interactive segmentation model can be used with a graphical user interface, which allows the user to provide weak annotations, such as mouse clicks, to generate an object segmentation map.
[0048] If initial labels are generated in the presence of existing resources, the initial labels can be modified to correct errors (step 214). Although not illustrated, correction of initial labels can also be performed after step 212. In some cases, a machine learning model is applied to the initial labels and a subset of the initial labels is determined to be mislabeled. For example, an initial label may indicate that a corresponding pixel accurately represents a corresponding point or region of a biological sample, but the corresponding pixel contains one or more artifacts. By applying a machine learning model, this error can be addressed by modifying the initial labels.
[0049] Once a set of labels (including the modified labels) is obtained, training images containing the set of labels can be generated. In step 216, additional labels can be iteratively generated using the training images with the set of labels to generate additional training data (step 216). The additional training data may include additional training images, each containing a corresponding set of labels. For example, if there are pre-trained models available from similar or the same image domain, transfer learning or fusion shot learning can be applied to the training images to generate an initial model. The initial model can then be used to make predictions for other unlabeled images and generate labels for the other unlabeled images to generate additional training data. In another example, active learning can be applied to the training images to select a set of images from multiple images, and the subset of images can be used to generate a corresponding set of labels. In yet another example, semi-supervised or fully supervised domain adaptation can be performed on the training images to generate additional training data. After that, process 2 ends.
[0050] Using the framework described above, various types of artifacts that affect the accurate representation of biological samples can be considered as labels. In some cases, additional types of artifacts are added to existing label types associated with the training data. For example, a new type of artifact can be merged with existing labels so that all artifacts have the same classification label as "artifact tissue." In some embodiments, a new label, distinct from any of the existing labels, is associated with the new type of artifact so that the number of label types can be increased. For example, a new classification label for "tissue folds" can be generated.
[0051] In some cases, multiple labels are assigned to the same pixel to generate training data. In this case, each of the multiple labels can predict whether the corresponding pixel displays at least some artifacts associated with a particular artifact type (e.g., blur, foreign body, tissue fold). For example, tissue folds may be interwoven with or otherwise correlated with blur artifacts. A pixel labeled as "tissue fold" may also display a blurred portion of the image. Thus, the following two labels can be associated with the above pixel: (i) "tissue fold," and (ii) "blur artifact." Machine learning techniques, such as multi-label classification techniques, can be used to predict that each image pixel is associated with one or more types of artifacts.
[0052] B. Process for generating labels for training images Figure 3 shows a flowchart illustrating an exemplary process 300 for using a machine learning model to generate training data according to several embodiments. The exemplary process 300 for generating training data may include generating labels that predict whether a corresponding pixel accurately represents a point or region of a biological sample. The exemplary process can be incorporated into the exemplary process presented in Figure 2.
[0053] In step 302, an image displaying at least a portion of the biological sample can be accessed. The image may be a slide image displaying a tissue section of a particular organ. In some examples, the biological sample has been stained using a staining protocol corresponding to a specific type of assay (e.g., IHC, H&E). For example, the image may be Ki67 Biological samples stained using a staining protocol corresponding to an IHC assay can be displayed.
[0054] In step 304, image preprocessing can be applied to the image to generate a preprocessed image. The preprocessed image may contain multiple labeled pixels. Each of the multiple labeled pixels can be associated with a label that predicts whether the pixel accurately represents a corresponding point or region of at least a portion of the biological sample. Thus, the label can indicate whether the corresponding pixel originates from an artifact, non-artifact tissue, or other type of region.
[0055] In some cases, image preprocessing algorithms include image segmentation, morphological processing, image thresholding, image filtering, image contrast enhancement, blur detection, other image preprocessing algorithms, or combinations thereof. Image preprocessing may include analyzing the image gradient of pixels across an image. For example, image preprocessing can be used to identify sets of smooth pixels (i.e., those with little to no local variation in image intensity). Smooth pixels can be identified by calculating the image gradient and applying a segmentation threshold. The segmentation threshold can represent a value that predicts whether a given pixel displays at least a portion of the edges shown in the image. The segmentation threshold can be a predetermined value. In some cases, the segmentation threshold is determined by performing Otsu's method or a balanced histogram thresholding method. Smooth pixels with an image gradient lower than the segmentation threshold can be identified as either blurred tissue or non-tissue regions with uniform image intensity. In addition to or alternatively, image preprocessing algorithms may include preprocessing images using one or more other machine learning models to enable the generation of multiple labeled labels.
[0056] In step 306, a machine learning model can be applied to the preprocessed image to identify one or more labeled pixels from multiple labeled pixels. The label of each of these one or more labeled pixels may be predicted to be incorrectly labeled by the image preprocessing algorithm. The error may stem from the image preprocessing algorithm not being sufficiently effective to identify the correct label for all pixels. For example, a segmentation threshold applied as part of the image preprocessing algorithm may accurately identify artifacts in some images, but the same segmentation threshold may be too low for the rest of the images. In another example, a segmentation threshold that correctly identifies artifacts in some parts of an image may be too low for the rest of the same image. In both examples, some artifact pixels may be incorrectly labeled as tissue regions.
[0057] In step 308, for each of the one or more labeled pixels, the label is Labels can be modified. In some cases, modifications are performed by the user via a graphical user interface. Alternatively, labels can be automatically modified using one or more executable instructions (e.g., if-else conditional statements).
[0058] In step 310, training images can be generated. The training images may include labeled pixels, such as labeled pixels with modified labels. In some cases, additional image features (e.g., image gradient values) can be associated with each labeled pixel to further facilitate training a machine learning model to identify artifact pixels.
[0059] In step 312, training images are output. These training images can be used to generate additional training data. The additional training data may include additional training images, each containing a corresponding set of labels. Various types of machine learning techniques can be used to generate the additional training data. For example, the machine learning techniques may include, but are not limited to, using active learning, transfer learning, fusion learning, or machine learning models trained via domain adaptation. After that, process 300 ends.
[0060] C. Exemplary training image with label Figure 4 shows a set of 400 exemplary images in which labels have been generated according to several embodiments. Each label indicates whether the corresponding pixel accurately represents at least a portion of the biological sample. Image 402 shows a thumbnail image corresponding to a full slide image displaying a tissue section. Image 404 shows a gradient-based map of the image generated by applying a Laplacian filter and subsequent Gaussian smoothing to the image corresponding to the thumbnail image or a full slide image at a different resolution. As shown in Image 404, pixels with an image gradient lower than the segmentation threshold can be identified as either blurred tissue or non-tissue regions with uniform image intensity.
[0061] Image 406 shows a tissue mask generated by applying a uniform filter and then thresholding a thumbnail image (or a corresponding image at a different resolution). For example, a tissue detector can be applied to an image by smoothing the image with a uniform filter and then applying a segmentation threshold to the intensities of the R, G, and B channels. As previously mentioned, the segmentation threshold can represent a value that predicts whether a given pixel displays at least a portion of the edges shown in a given image. The segmentation threshold can be a predetermined value. In some cases, the segmentation threshold is determined by performing Otsu's method or a balanced histogram thresholding method. Pixels with intensity values higher than the edge detection threshold across all three channels can be identified as tissue pixels. Using the tissue mask, a whole slide blur mask (e.g., Image 408) can be generated with three classifications: non-tissue, blurred tissue, and unblurred tissue.
[0062] Image 408 shows a preprocessed image (e.g., a blur map) generated by merging images 406 and 404. For example, the preprocessed image shown in Image 408 shows the predicted artifact pixels in the thumbnail image, with dark red identifying the predicted artifact pixels.
[0063] Image 410 shows a preprocessed image in which a set of image tiles can be identified. In some cases, image tiles with some artifact pixels are automatically selected. Preprocessed image 410 can correspond to image 408 and labels with varying amounts of blur. This may include: Image tiles 412 and 414 represent image tiles selected from the pre-processed image 410. Specifically, image tile 412 displays a region of the biological sample stained using the ER-Dabsyl IHC assay (ER: estrogen receptor). Image tile 414 identifies the initial label within the region. The initial label may include multiple classifications such as blurred tissue, non-tissue, and unblurred tissue.
[0064] Image 416 shows a screenshot of an interactive graphical user interface displaying image tiles, which the user can use to modify the initial labels through interaction (e.g., mouse clicks). In some cases, the modification of the initial labels is performed by applying a machine learning model to the preprocessed image 410 or selected image tiles from 410. The machine learning model can modify the entire blur mask with high accuracy using a limited number of annotations. The machine learning model may be a separate process (not shown) or can be integrated into the graphical user interface. The application of the machine learning model can be performed using a CPU or by leveraging highly parallelized computation using a GPU, thus ensuring efficient label correction.
[0065] D. Determining the level of blur for accurate labeling. Subjectively determining a specific threshold for detecting artifact pixels inevitably leads to discrepancies between experts' perceptions of blur and blur levels. Such discrepancies can result in a significant decrease in the performance of digital pathology algorithms. Figure 5 shows an exemplary image containing image portions exhibiting various levels of blur. In particular, a pathologist might state that image 500 is analyzable for over 70% of the image. However, in reality, a large portion of image 500 may be blurred, posing a problem for the classification model. For example, image tile 502 might be considered not blurred by a particular expert, but it may not be sufficiently focused for the classification model to accurately detect a biomarker.
[0066] To improve the consistency of identifying artifact pixels in an image, the performance changes of a classification model (e.g., cell classification) can be quantitatively evaluated at various blur levels. A blur threshold can be selected that corresponds to the blur level predicted to cause a degradation in the classification model's performance beyond an acceptable threshold. The blur threshold can be used to indicate any tile in the image that is considered more blurred (e.g., image tile 504). In some cases, any pixel in an image tile is labeled as blurred tissue if these pixels are localized in the image portion corresponding to a tissue region (e.g., within a tissue region in tissue mask 406) and the respective image gradient is lower than the blur threshold.
[0067] In some cases, the blur threshold can be determined by generating a set of sample images. Each sample image in the set of sample images can be generated by applying a Gaussian filter with a specific sigma value to display one or more regions of the sample at various levels of blur.
[0068] In addition to or instead of this, the volume scan function of a digital pathology scanner can be used to set the blur threshold. For example, the z-stack of a digital pathology scanner and / or microscope can be used to scan the z-axis of a slide and obtain a set of scans in which the distance from the nominal focal plane increases. The set of scans can correspond to an increasing level of blur. An exemplary process for using the volume scan function to determine the blur threshold may be as follows: Firstly, with respect to a fixed assay and a fixed downstream digital pathology analysis (e.g., a cell classification model), a training image with labels is rescanned in the “volume scan” mode of the scanner. This allows for the generation of volume scan images. In some cases, the volume scan setup involves scanning training images using non-nominal focal scan planes at regular intervals (e.g., 1 micron). Based on the volume scan images, sets of pixels that result in insufficient accuracy in digital pathology analysis can be detected. In some cases, the range of image gradients in the identified sets of pixels can be calculated. The maximum image gradient within the range of image gradients can be set as the blur threshold. Pixels with image gradients exceeding the blur threshold can be predicted as pixels that will cause accuracy degradation exceeding an acceptable level for subsequent digital pathology analysis.
[0069] E. Identification of tissue fold artifacts for generating training data Tissue folds typically occur during tissue processing (e.g., preparation of tissue slides) where one or more portions of a tissue section fail to adhere firmly to the slide glass and flip over onto other portions of the tissue section. Figure 6 shows a set of exemplary images 600 containing one or more tissue folds in corresponding biological specimens. As shown in Figure 6, tissue folds can have diverse appearances. For example, the first image 602 shows a tissue fold that is much darker in intensity than the surrounding non-tissue fold area. The second image 604 shows a tissue fold that is brighter in intensity, but still allows cells in the underlying tissue layer to be fairly visible. The third image 606 shows a tissue fold area with a blurred area. The blurred area may be caused by the thickness of the tissue fold exceeding the scanner's depth of field.
[0070] Different processes may be used to generate ground truth images that include tissue fold regions. Figure 7 shows an exemplary schematic diagram 700 for generating training data for tissue fold regions according to several embodiments. The training data can then be used to generate a binary tissue fold mask. In Figure 7, no existing machine learning model, existing ground truth, or effective image processing technique for generating the initial ground truth was present. Therefore, referring again to Figure 2, the answers for steps 204 and 208 were both identified as "no". As a result, step 212 was performed to generate training data for tissue fold regions, where an interactive GUI was used to generate ground truth images. Generating training data for tissue fold regions may include using an interactive segmentation GUI to generate a binary tissue fold mask where the two classifications are tissue folds and non-tissue folds.
[0071] In some cases, the tissue fold mask 702 is generated based on one of the following three methods: (1) blurred ground truth (for FOV 704 selected from blurred ground truth) (block 706), (2) regions identified by an image processing algorithm (e.g., identifying FOVs with tissue folds from additional Mosaic WSI (block 708)), and (3) regions selected by an interactive GUI (block 710). In some cases, each of the three methods is executed sequentially to generate the tissue fold mask. For example, it can be determined whether the tissue fold mask generated based on blurred ground truth is accurate (e.g., based on visual inspection). If it is not accurate, regions generated by an image processing algorithm can be used. If regions generated by an image processing algorithm do not result in an accurate tissue fold mask, regions manually selected by an interactive GUI can be used to generate the tissue fold mask (block 710).
[0072] In some cases, an interactive GUI includes one or more machine learning models to facilitate the selection of tissue fold regions. For example, an interactive GUI may include: (i) a first GUI component that allows manual writing for selecting image regions and visualizes the selected regions for iterative manual correction; and (ii) a second GUI component that allows user input such as writing and mouse clicks to guide the automatic identification of target regions. Regarding interactive GUIs, image processing methods can be designed, or machine learning models can be trained, to generate segmented masks in response to user input. For example, a machine learning model can be trained with simulated user clicks within a target image region, as well as the original image as model input, to output a segmented mask. In practice, a deep learning interactive GUI can learn to identify a target region with user input that is typically only a few pixels or a portion of the target image region. Furthermore, the user can iteratively modify existing inputs or add new inputs to refine the segmented mask until the mask is accurate and can be used as ground truth for training the machine learning model.
[0073] By combining a binary tissue fold mask with a corresponding tissue mask, a three-class tissue fold mask can be generated. Figure 8 includes an exemplary set of three-class tissue fold masks 800 according to several embodiments. The first three-class tissue fold mask 804 corresponds to the first slide image 802, and the second three-class tissue fold mask 808 corresponds to the second slide image 806. Each of the three-class tissue fold masks 804 and 808 may include, for each pixel, a first classification for non-tissue regions, a second classification for non-tissue fold tissue regions, and a third classification for tissue fold regions.
[0074] F. Integration of various types of artifacts into classification labels To train a machine learning model to detect two or more types of artifacts, artifact regions detected in training images can be distinguished between blurred ground training labels and tissue fold ground labels. For example, four types of classifications can be integrated into a segmentation mask, and these four types of classifications may include: (1) non-tissue regions, (2) blurred non-tissue fold regions, (3) unblurred tissue fold regions, and (4) analyzable tissue regions.
[0075] However, the four-category labeling system assumes that blurred tissue regions and tissue fold regions are mutually exclusive, which is not always the case. For example, as shown in images 602 and 606 of Figure 6, tissue fold regions are often accompanied by blurred regions. For instance, Figure 9 shows an exemplary image 900 (e.g., FOV) containing a tissue fold region, where the tissue fold region displays both unblurred and blurred areas. Thus, in Figure 9, assigning the blurred pixels of the tissue fold region to either the "blurred" class or the "tissue fold" class would lead to confusion when training the corresponding machine learning model.
[0076] In another example, Figure 10 shows a set of 1000 exemplary images displaying various artifact regions. For example, FOV 1002 displays blurred regions interwoven with tissue fold regions. The tissue fold binary mask 1004 can distinguish tissue folds from non-tissue fold regions. In contrast to the tissue fold binary mask 1004, the blurred ground truth mask 1006, which segments the image into four classifications, displays confusing patterns of classified regions. For example, the presence of interwoven blurred regions may divide a tissue fold region into multiple smaller tissue fold subregions, which can lead to ambiguity and significantly increase the difficulty for machine learning models to learn meaningful features with respect to tissue fold classes.
[0077] To address inaccurate classification of tissue fold regions, two types of classification strategies can be implemented. The first strategy may involve combining tissue fold regions and blurred tissue regions into a single class (e.g., a non-analyzable region class). Blurred tissue regions and tissue fold regions are classified as the "non-analyzable tissue" class. In practice, a three-class segmentation can be output, classifying each pixel into one of the following three classes: non-tissue, analyzable tissue, and non-analyzable tissue. For example, Figure 11 shows several implementations. Figure 12 shows an exemplary artifact mask 1100 generated using a morphological composite artifact classification technique. Figure 12 shows a further artifact mask 1200 using a composite artifact classification technique according to several embodiments. In Figure 12, three-class ground truth masks 1204 and 1208 are shown. The regions of the biological specimen shown in the FOV (raw RGB) stained with 1202 and QM-Dabsyl-ER / TAMRA-PR 1206 are segmented. For FOV 1202, the non-analyzed regions were hardly blurred. For FOV 1206, the non-analyzed regions included both unblurred and blurred tissue folds. Thus, Figures 11 and 12 classify the image regions into one of three classes: (i) non-tissue, (ii) analyzable tissue, and (iii) non-analyzable tissue.
[0078] A second classification strategy may involve associating pixels with two or more classification labels. Multi-label segmentation can facilitate the classification of each pixel as one or more of the following four classes: non-tissue, analyzable tissue, non-analyzable tissue, and tissue folds. To generate multiple classifications, each pixel location can be assigned a binary value (either positive or negative for that classification). For example, Figure 13 shows slide image 1300 in which pixels are associated with two or more classification labels according to several embodiments. In Figure 13, slide image 1300 shows pixels 1302 and 1306 associated with a single label. Furthermore, slide image 1300 shows pixel 1304 associated with two labels (e.g., a blurred tissue label and a tissue fold label). As shown in Figure 13, a region of a biological sample shown in slide image 1300 can be associated with multiple different labels. Therefore, the 4×1 array shown in Figure 13 identifies the label of each pixel, where 0 indicates the presence of a negative for the corresponding region (e.g., this pixel does not belong to this class), and 1 indicates the presence of a positive for the corresponding region (e.g., this pixel belongs to this class).
[0079] Figure 14 shows a process 1400 for identifying classification labels for pixels associated with both blurred and tissue fold regions according to several embodiments. In step 1402, pixels associated with multiple classifications (e.g., non-analyzable and tissue folds) can be identified. In step 1404, pixels can be associated with binary values for each classification in a set of classifications. The set of classifications may include: (a) non-tissue, (b) analyzable tissue, (c) blurred tissue, and (d) tissue folds. The binary values may indicate the presence of an object associated with the corresponding classification (e.g., tissue folds). For example, pixel 1304 contains a binary value of "1" for both blurred tissue and tissue folds, and a binary value of "0" for non-tissue and analyzable tissue. The binary value of "1" may indicate that pixel 1304 displays both blurred tissue and tissue folds. In some cases, the "blurred tissue" class can be replaced with a non-analyzable class (either blurred or tissue folds). In some cases, one of multiple classifications is selected to represent the pixel. In any step 1406, the sets of classifications can be ranked based on their respective predicted probability values at the pixel location. Specifically, the probability of how likely a pixel is to belong to each classification in the set can be generated (e.g., using a machine learning model). For example, a three-class segmentation model generates three numbers for each pixel, such as [0.1, 0.2, 0.7], which are the probabilities that this pixel belongs to each class. In this example, the sets of classifications can be ranked according to their probabilities, and class 3 is expected to be ranked with the highest value. In practice, the highest probability corresponds to the class to which the pixel is most likely to belong. In any step 1408, the classification with the highest probability value can be selected as the final predicted label for the pixel.
[0080] No additional processing is required for ground truth masks to generate labeled images for training machine learning models. Rather, two sets of ground truth masks can be used, one set corresponding to a three-category blur mask (e.g., non-tissue, analyzable tissue, blurred tissue) and another set corresponding to tissue fold regions within tissue regions (e.g., a binary tissue fold mask). In some cases, the labeling of each pixel can be performed using a 4x1 array during model training.
[0081] The machine learning techniques described above can facilitate the accurate classification of regions within slide images. For example, Figure 15 shows a comparison between classifications predicted by ground truth masks in several embodiments. For instance, the first set of images 1502 shows a comparison between predicted masks and corresponding ground truth masks for images stained using a single IHC containing estrogen receptor (ER). The second set of images 1504 shows a comparison between predicted masks and corresponding ground truth masks for images stained using a single IHC containing cytokeratin 7 (CK7). The third set of images 1506 shows a comparison between predicted masks and corresponding ground truth masks for images stained using a dual IHC containing estrogen receptor and progesterone receptor (ER / PR). Based on the comparisons, it can be seen that the predicted segmented masks are qualitatively similar to the corresponding ground truth masks.
[0082] Furthermore, Figure 16 shows the predicted regions in slide image 1600 using different types of IHC assays according to several embodiments. For example, the first image 1602 shows the image tile stained using a dual IHC assay including LIV / HER2 and the corresponding predicted segmentation mask including three classifications (e.g., unblurred tissue, blurred tissue, non-tissue). The second image 1604 shows the second image tile and the corresponding predicted segmentation mask including three classifications using a triple IHC assay including ER / Ki67 / PR. The third image 1606 shows the third image tile and the corresponding predicted segmentation mask including three classifications using a triple IHC assay including CD8 / BCL2 / CD3. Qualitative evaluation of the predicted segmentation masks shows the accurate classification of unblurred tissue, blurred tissue, and non-tissue.
[0083] III. Training machine learning models to accurately detect artifact pixels As mentioned above, since images can be stained using various types of IHC assays, training machine learning models to accurately detect artifact pixels can be complex. For example, fluorescence-based IHC assays can use multispectral imaging to separate several different fluorescence spectra, which can enable the accurate identification of multiple antigens on the same tissue section. However, these multiplexed IHC assays can result in more complex staining patterns compared to single IHC assays (e.g., IHC assays targeting a single type of antigen).
[0084] Figure 17 shows a set of 1700 exemplary images, each stained using a staining protocol corresponding to a specific type of IHC assay. Image 1702 shows a biological sample stained with hematoxylin alone. Image 1704 shows a biological sample stained using a single IHC assay. In particular, Image 1704 shows the nuclear staining pattern of a biological sample having Dabsyl-stained estrogen receptors, where Dabsyl was used as the chromogen, resulting in yellow staining. Identification of artifacts (e.g., artifact pixels) in Images 1702 and 1704 may be relatively straightforward.
[0085] The artifact detection process becomes considerably more difficult in images stained using staining protocols corresponding to dual IHC assays. For example, image 1706 shows dual IHC The images show biological samples stained using assays. In particular, Image 1706 shows the nuclear staining patterns of biological samples stained with Tamra to identify the estrogen receptor (i.e., ER) and Dabsyl to identify the progesterone receptor (i.e., PR). In Image 1706, Tamra can represent purple staining and Dabsyl can represent yellow staining. However, Image 1706 further shows a blend of both stains, exhibiting various hues that can result from a variety of factors, including the staining protocol, chromophor interference, and the relative expression levels of the biomarkers. In another example, Image 1708 shows biological samples stained using a different type of dual IHC assay. In particular, Image 1708 shows biological samples stained with Tamra-PDL1 (programmed death ligand 1) and Daybsyl-CK7 (cytokeratin 7), where the tissue area stained with PDL1 shows primarily membrane staining and the tissue area stained with CK7 shows primarily cytoplasmic staining. However, image 1708 also shows tissue regions where both stains overlap. Therefore, detecting artifacts from these types of images can be difficult.
[0086] Therefore, a machine learning model can be trained to detect artifact pixels in images with various staining patterns. This technique may involve accessing training images that display at least a portion of a biological sample. The training images may contain multiple labeled pixels, each associated with a label. The labels predict whether a pixel is an artifact pixel. The training images can be used to train a machine learning model. The machine learning model may include a set of convolutional layers, and a first loss calculated for a first convolutional layer and a second loss calculated for a second convolutional layer can be used to train the machine learning model to detect artifact pixels at a target image resolution.
[0087] A. Architecture for training machine learning models for artifact pixel detection To enhance the ability of machine learning models to effectively detect artifact pixels across various image resolutions, supervision can be added during the training phase of the machine learning model, which is trained in each of the sets of image resolutions. Figure 18 shows a schematic diagram illustrating exemplary architectures 1800 used to train a machine learning model to detect artifacts in an image according to several embodiments. Figure 18 shows an encoder-decoder model architecture for image segmentation, in which features from each of multiple image resolutions in the decoder pass can be used for pixel-by-pixel classification.
[0088] In some cases, the encoder-decoder model architecture includes a U-Net. The machine learning model includes a U-Net machine learning model trained to detect artifact pixels in an image. The U-Net machine learning model can include a reduction pass and an expansion pass. The reduction pass can include a first set of processing blocks, each processing block corresponding to processing a training image at a corresponding image resolution. For example, a processing block can include applying two 3x3 convolutions (convolutions without padding) to an input (e.g., a training image), each convolution followed by a normalized linear unit (ReLU). Thus, the output of the processing block can include a feature map of the training image at a corresponding image resolution. Furthermore, the processing block includes a 2x2 maximum pooling operation with a stride of 2 for downsampling the feature map of the processing block to a subsequent processing block that can repeat the above steps at a lower image resolution. In each downsampling step, the number of feature channels can be doubled.
[0089] Following the reduction path, the expansion path includes a second set of processing blocks, each processing block This corresponds to processing the feature map output from the reduction path at the corresponding image resolution. For example, the second set of processing blocks in the processing block receives the feature map from the preceding processing block, applies a 2x2 convolution ("de-convolution") that halves the number of feature channels, and concatenates the feature map with the cut-off feature map from the corresponding processing block in the reduction path. The processing block can then apply two 3x3 convolutions to the concatenated feature map, each convolution followed by an (optional) batch normalization layer and ReLU. The output of the processing block contains the feature map at the corresponding image resolution, which can be used as input for subsequent processing blocks at higher image resolutions. Processing blocks can be applied until the final output is generated. The final output may include an image mask. The image mask can identify sets of artifact pixels, each of which is predicted not to accurately represent at least a portion of a point or region of the biological sample.
[0090] In some cases, the loss in each or several processing blocks of a second set of processing blocks can be calculated and used to determine the total loss of the U-Net machine learning model. The total loss may be calculated based on the losses calculated from the selected processing blocks. For example, the total loss can be determined based on the sum or weighted sum of the losses generated from each of the second set of processing blocks. In the second example, the total loss can be determined based on the average or weighted average of the losses calculated for each processing block. The total loss of the U-Net machine learning model can then be used to learn the parameters of the U-Net machine learning model (e.g., the parameters of one or more filters in the convolutional layer). Using the total loss of the U-Net machine learning model enables the detection of artifact pixels across various image resolutions.
[0091] In some cases, the loss of each processing block in the second set of processing blocks can be determined by applying a 1x1 convolutional layer to the feature maps output by the processing blocks to generate one or more modified feature maps, and then determining the loss from one or more modified feature maps. In particular, a 1x1 convolution can be applied such that the number of modified feature maps corresponds to the number of class labels (e.g., three modified feature maps for three label types). In some cases, the modified feature maps are upsampled to the same resolution as the output of the machine learning model (e.g., the image mask). Alternatively, the image mask (having the same size as the training image) can be downsampled to the same resolution as the modified feature maps.
[0092] B. Integration of global information to detect large-scale artifacts in images. In some cases, a second machine learning model is trained to detect artifact pixels in an image that are predicted to correspond to larger artifacts within the image. The use of additional machine learning models can circumvent limitations on input tile size due to the constraints of computing resources (e.g., hardware memory). For this purpose, the parameters of the additional machine learning model can be learned based not only on the features of a particular image tile in an image, but also on the features of adjacent tiles in the same image, in order to incorporate information from the adjacent image regions of each image tile. Thus, the additional machine learning model can be trained using information corresponding to the dependencies between the target image tile and its adjacent image tiles. In some cases, the additional machine learning model includes recurrent neural networks (e.g., gated recurrent neural networks) and long- and short-term memory.
[0093] A second machine learning model can be trained as follows: (i) to replace a machine learning model having a set of convolutional layers (e.g., a convolutional neural network), (ii) to be used before or after the execution of the machine learning model, and / or (iii) to be integrated into the machine learning model.
[0094] A recurrent neural network includes a chain of iterative modules ("cells") of the neural network. Specifically, the operation of the recurrent neural network comprises repeating a single cell indexed by the position of a target image tile (t). To provide its recurrent behavior, the recurrent neural network maintains a hidden state s t that is provided as an input to the next iteration of the network. The hidden state may be a vector or matrix representing information from adjacent image tiles. As mentioned herein, the variables s t and h t are used interchangeably to represent the hidden state of the recurrent neural network. The recurrent neural network receives a feature representation x t of the target image tile and a hidden state value s t-1 determined using a set of input features of adjacent image tiles. In some cases, the feature representation x t of the target image tile is generated using a machine learning model having a set of convolutional layers. The following formula provides how the hidden state s t is determined. s t =φ(Ux t +Ws t-1 ) wherein U and W are weight values applied to x t and s t-1 respectively, and φ is a non-linear function such as tanh or ReLU.
[0095] As shown, the s t value generated based on the application of Ux t-1 and Ws t can be used as the hidden state value for the next iteration of the recurrent neural network that processes features corresponding to a subsequent image tile.
[0096] The output of the recurrent neural network is expressed as follows. o t=softglass(Vs t ) Here, V is the hidden state value s t This is the weight value applied to it.
[0097] Therefore, hidden states t This can be called the network's memory. In other words, the hidden state s t This relies on information related to inputs and / or outputs that is used or otherwise derived from one or more previous image tiles. t The output in this case is a set of values used to identify artifact pixels, which are calculated based at least partially on memory at the target image tile position t.
[0098] C. Process for training a machine learning model to accurately detect artifact pixels Figure 19 shows a flowchart illustrating an exemplary process 1900 for training a machine learning model to accurately detect artifact pixels according to several embodiments. In step 1902, a training image displaying at least a portion of a biological sample can be accessed. The training image contains a plurality of labeled pixels, each of which is associated with a label. The labels predict whether the corresponding pixel accurately displays a corresponding point or region of at least a portion of the biological sample. For example, a pixel displaying an out-of-focus region of a biological sample may be labeled as not accurately displaying the corresponding region (e.g., a portion of a tissue section).
[0099] In some cases, the training image is converted to a grayscale image. The grayscale image is used to train machine learning to detect artifact pixels. In addition to this, or instead, the training image is converted to a preprocessed image by converting its pixels from a first color space (e.g., RGB) to a second color space (e.g., L*a*b). The first color channel (e.g., L channel) of the second color space is extracted and can be used to train the machine learning model to detect artifact pixels. By converting to a different color space, non-informational color information can be extracted from the artifact detection model. The machine learning model is run to learn discriminative image features that can be eliminated and are unrelated to complex color changes and uneven staining patterns, which are largely irrelevant to artifacts.
[0100] In step 1904, a machine learning model containing a set of convolutional layers can be accessed. For example, the machine learning model is a U-Net architecture. In some cases, the machine learning model is configured to apply each convolutional layer of the set of convolutional layers to a feature map representing an input image.
[0101] In step 1906, the machine learning model is trained to detect one or more artifact pixels in an image at the target image resolution. The artifact pixels are predicted to not accurately represent at least a portion of a point or region of the biological sample. For example, artifact labels can predict the presence of artifacts (e.g., blur, tissue folds, foreign bodies) that may result in pixels that do not accurately represent the corresponding region of the biological sample.
[0102] In some cases, a set of image features is used along with training images to train a machine learning model. For example, a set of image features may include a matrix of image gradient values. The image gradient value matrix can identify the image gradient value of each pixel in the training image. The image gradient value indicates whether the corresponding pixel corresponds to an edge of an image object. In some cases, the image gradient value matrix is determined by applying a Laplacian of Gaussian (LoG) filter to the training image.
[0103] Training a machine learning model may involve learning the model's parameters based on loss values calculated for each pixel. For each labeled pixel among a set of multiple labeled pixels in a training image, training may involve determining a first loss for the labeled pixel at a first image resolution by applying a first convolutional layer of a set of convolutional layers to a first feature map representing the training image at a first image resolution. Then, a second loss for the labeled pixel at a second image resolution may be determined by applying a second convolutional layer of a set of convolutional layers to a second feature map representing the training image at a second image resolution. In some cases, the second image resolution has a higher image resolution than the first image resolution.
[0104] Training may further include determining the total loss for labeled pixels based on the first and second losses. The total loss can be used to determine if the machine learning model has been trained to detect one or more artifact pixels at the target image resolution.
[0105] In step 1908, the trained machine learning model is output. The trained machine learning model can then be used by another system to detect artifacts in other images with different staining patterns. After that, process 1900 ends.
[0106] D. Training a machine learning model to predict two or more types of artifacts. In some cases, a machine learning model is trained using a three-class tissue fold mask (e.g., the three-class tissue fold mask 804 in Figure 8) to identify two or more types of artifacts in at least a portion of a given slide image (e.g., FOV). For example, a machine learning model can be trained to detect a first set of pixels corresponding to blurred areas and a second set of pixels corresponding to tissue folds. To train a multi-class model, pixels in each training image are labeled as "tissue areas," "unanalyzed areas" corresponding to blurred areas, and It can be labeled as one or more "non-analytical tissues" corresponding to tissue folds within an tissue area.
[0107] A multi-classification model can be initially trained to detect and segment artifact regions. Figure 20 shows a flowchart illustrating process 2000 for training a machine learning model to detect artifact regions in slide images according to several embodiments. In Figure 20, a first training stage can be performed to train the machine learning model to detect artifact regions displayed in the slide images, and a second training stage can be performed to evaluate the performance of the machine learning model. With respect to the first training stage, a set of slides can be selected to generate a ground truth mask (step 2002). In step 2004, a first set of training images can be generated. The first set of training images may include a segmentation mask that identifies two or more types of artifacts. The process for generating the first set of training images may include process 700 in Figure 700. In step 2006, the first set of training images can be split into a model training dataset, a validation dataset, and a test dataset. In step 2008, the machine learning model can be trained using the first set of training images. The step of training a machine learning model using a first set of training images is further illustrated in process 1900 of Figure 19. In step 2010, the segmentation performance of the machine learning model can be tested based on performance scores such as accuracy and recall.
[0108] For the second training phase, an additional set of slides can be selected (step 2012). In some cases, the additional set of slides may include slides from invisible tissue types, biomarkers, and chromogens. The additional set of slides may differ from those used in the first training phase, as the objective is to have a separate set of slides to pressure-test the model performance in order to assess the generalizability to invisible images, chromogens (combinations), biomarkers, and tissue types. In some cases, FOV is selected from the additional set of slides. A second set of training images can then be generated based on the additional set of slides. In step 2014, a label can be assigned to each pixel of the training image in the second set of training images. Labels can be assigned by receiving two types of readouts from annotators. The first type of readout may include whether or not tissue folds are present within the FOV, and the second type of readout may include the percentage of non-analyzed tissue within the tissue region within each selected FOV. In step 2016, the machine learning model can be trained and tested using a second set of training images for an independent test of model generalizability. In some cases, the percentage of annotations is compared to the model prediction as a surrogate for model generalizability. As a result, the machine learning model can be trained and tested to detect artifact regions in other slide images.
[0109] IV. Implementation of a machine learning model for detecting artifact pixels in images Due to the large size of the overall slide images, automated digital pathology analysis needs to be performed as efficiently as possible without sacrificing accuracy. Typically, digital pathology analysis (e.g., cell classification modeling) can involve generating sets of image tiles from the overall slide image, where each image tile can represent a portion of the image with a specific size and dimensions (e.g., 20x20 pixels). Cell classification can then be performed on each image tile to generate corresponding prediction results, which can then be reassembled back to the overall slide image resolution.
[0110] Therefore, applying quality control to digital pathology analysis can double processing time. However, the main digital pathology analysis, which typically has a resolution of 20x or 40x, It is not necessary to perform slide quality control at the same resolution as the analysis. This is because many types of artifacts can be identified even at lower image resolutions. Furthermore, large artifacts such as large tissue folds cannot fit within image tiles at high resolutions. Therefore, in some cases, performing quality control at high image resolutions can lead to inconsistent results.
[0111] To improve efficiency when implementing artifact detection in digital pathology analysis, some embodiments include using different image resolutions to detect artifact pixels in images. A set of training images can be obtained. For each training image, labels corresponding to each pixel can be collected at high image resolutions (e.g., 40x, 20x, 10x). An artifact detection machine learning model (e.g., the U-Net machine learning model in Figure 18) is trained using the set of training images. The machine learning model is further trained and tested using images of lower image resolutions to determine the target image resolution. At the target image resolution, the machine learning model can maintain the accuracy of artifact pixel detection while increasing efficiency. For example, if a machine learning model can be trained to detect artifact pixels within an acceptable level of accuracy at 5x resolution, the target resolution can be determined to be 5x, and there is no need to apply the machine learning model to detect artifact pixels at higher image resolutions (e.g., 10x, 20x). In this way, the inference time can be reduced to 1 / 16th compared to another machine learning model that detects artifact pixels at (e.g.) 20x image resolution.
[0112] A. Process for integrating artifact detection into digital pathology analysis Figure 21 shows a flowchart illustrating an exemplary process 2100 for using a machine learning model trained to accurately detect artifact pixels according to several embodiments. In step 2102, an image displaying at least a portion of a biological sample is accessed. For example, the image may be a slide image displaying a tissue section of a particular organ. The biological sample may include tissue sections stained using a staining protocol for a particular type of IHC assay. In some cases, the image is at a first image resolution (e.g., 40x).
[0113] In step 2104, a machine learning model trained to detect artifact pixels in an image of a second image resolution is accessed. The machine learning model may be a machine learning model having a set of convolutional layers (e.g., U-Net). In some cases, the first image resolution of the image (e.g., 40x) has a higher image resolution than the second image resolution (e.g., 5x).
[0114] In step 2106, the image is transformed to generate a transformed image that displays at least a portion of the biological sample at a second image resolution. For example, one or more image resolution changing algorithms, including mipmapping, nearest neighbor interpolation, and Fourier transform, can be used to change the image resolution and generate the transformed image.
[0115] In step 2106, the machine learning model is applied to the transformed image to identify one or more artifact pixels from the transformed image. The artifact pixels are predicted to not accurately represent at least a portion of the biological sample, whether a point or region. For example, the artifact pixels may be predicted to represent a portion of a blurred area in a given image, or a portion of a foreign object shown in the image.
[0116] In step 2108, an output is generated that includes one or more artifact pixels. In some cases, the output identifies the artifact pixels with pixel-level precision. This includes artifact masking. Artifact masking can be used to identify parts of an image corresponding to different classes (e.g., unblurred tissue, blurred tissue, non-tissue). In addition, or alternatively, the output may indicate the amount of artifact pixels (e.g., the ratio of artifact pixels to the total number of pixels in the image). For example, the estimated amount may include the predicted count of artifact pixels, the cumulative area corresponding to multiple or all artifact pixels, the ratio of slide area or tissue area corresponding to the predicted artifact pixels, etc. Process 2100 then terminates.
[0117] In some cases, predicted artifacts in an image (e.g., artifacts displayed by one or more artifact pixels) fall into one of the following classifications: (a) a first artifact classification where the artifact occurs only during the slide scan, and (b) a second artifact classification where the artifact occurs at any point in time (e.g., experiment, staining). If the predicted artifact corresponds to the first artifact classification, the digital pathology analysis can proceed without further quality control operations. If the predicted artifact corresponds to the second artifact classification, a warning is generated to the user prompting them to reject the image and / or rescan the biological sample to generate a different image displaying the biological sample. In some cases, the graphical user interface is configured to allow the user to reject the image. In addition to this, or alternatively, a quality control algorithm can be designed for each type of predicted artifact. A quality control algorithm for a predicted type of artifact can output results that trigger image rejection and / or rescanning of the biological sample.
[0118] B. Configuration for implementing quality control It is not uncommon for slide images to contain some artifacts (e.g., blurred areas). From a user experience perspective, scanned slides with a large number of artifacts generated during scanning are undesirable. Furthermore, considering the large size of histology slides, digitizing all slides with obvious quality issues can increase storage space and scanning time. This problem can become more pronounced in large-scale projects when scanning speed is not optimal. Therefore, artifact detection during the scanning phase can be considered as an alternative to performing artifact detection after digitizing the slides.
[0119] 1. Preprocessing of slide images In some cases, image preprocessing is applied to an image before a machine learning model is applied to detect artifact pixels. For example, a preview image (e.g., a thumbnail image) can be initially captured by scanning a slide displaying a biological sample using a scanning device. Image preprocessing algorithms, such as a blur detection algorithm, can be applied to the preview image. If tissue regions are detected in the preview image, an initial image displaying the biological sample can be scanned. The initial image can then display the biological sample at the target image resolution.
[0120] As an exemplary example, a slide of a biological sample can be scanned at thumbnail resolution (e.g., 1.25x) or another lower resolution to generate a preview image. The lower resolution of the preview image allows the scan time to be within a predetermined time threshold. The predetermined time threshold can be selected from various time values, such as 10 seconds, 15 seconds, 20 seconds, or any larger value. Image preprocessing can be applied to the preview image to identify the portion of the image that is expected to display one or more tissue regions of the biological sample. If no tissue regions are identified, the quality control process ends. If one or more tissue regions are identified, the machine learning model is used with images captured at a relatively high resolution (e.g., 4x). It can be applied to images.
[0121] 2. Artifact detection during image scanning Scanning systems for digital pathology typically include line scanners and tile-based area scanners. In a line scanner system, a line sensor can perform image acquisition of one line / stripe at a time, and the line may be one pixel wide and have a length specified by the design of the sensor in the system. After scanning the entire slide is complete, the image data acquired from the line scan can be rearranged into image tiles according to the pixel positions corresponding to the image tiles on the slide. These image tiles can then be stitched together into an image of the entire slide. In a tile-based scanner system, an area sensor performs image acquisition of one tile at a time, and the tile corresponds to a rectangular field of view.
[0122] In both types of scanner systems, image tiles can be generated during scanning, at which point a machine learning model can be applied to detect artifact pixels. Regarding line scanners, the image data acquired from the line sensor is not image tiles. Therefore, scan data can be generated every few line sweeps. The scan data can then be rearranged into image tiles, at which point a machine learning model can be applied to the image tiles to detect artifact pixels. In some cases, the processing can be performed using hardware components (e.g., FPGA) and / or software components.
[0123] Artifact detection during scanning allows the scanner or scanner-related software to warn the user of slide quality issues (e.g., artifact type, artifact location, artifact size) during scanning, enabling the user to decide whether to save or delete a particular scan. In addition, or alternatively, artifact detection during scanning can be used by the scanner to intelligently and automatically adjust settings in response to the detection of predicted artifacts. For example, autofocus parameters can be adjusted by the scanner in response to the determination that artifact pixels are present in a portion of the image displaying a tissue area of a biological sample, or that the amount of artifact pixels exceeds an artifact area threshold.
[0124] (a) Artifact detection by initial scan at low image resolution For slides with detectable tissue regions, a machine learning model can be applied to the image to generate an image mask that identifies artifact pixels and to determine the amount of artifact pixels present in the image. In some cases, artifact detection can be performed at low image resolution. Low-resolution artifact detection can be used to detect artifact pixels that are expected to display artifacts occupying a large portion of the image, including large tissue folds, large blurred areas caused by tissue folds, etc.
[0125] For example, a machine learning model can be trained to detect artifact pixels at a target image resolution. During scanning, a first scan image of a slide displaying a biological sample can be performed at the target image resolution. The machine learning model can be applied to the first scan image to identify one or more artifact pixels. The amount of artifact pixels can be determined. A value representing the amount of artifact pixels can be compared to an artifact region threshold. In some examples, the artifact region threshold corresponds to a value representing the relative size of the image portion within the image (e.g., 40%, 50%, 60%, 70%, 80%, 90%). The artifact region threshold is user-selectable. It is possible that, if the number of artifact pixels exceeds the artifact region threshold, it can be predicted that one or more artifacts occupy a large portion of the image and are therefore likely to cause a decrease in the performance of subsequent digital pathology analysis (e.g., cell classification). If the value representing the number of artifact pixels is determined to exceed the artifact region threshold, it can be determined that there is a potential quality control failure. In some cases, a warning is generated in response to the determination of a quality control failure.
[0126] In addition to this, or alternatively, an image mask containing one or more artifact pixels (sometimes called an "artifact mask") can also be generated. The artifact mask can then be used by a graphical user interface to identify the portion of the image that is superimposed on the image and therefore expected to contain artifacts. This allows the user to decide whether to rescan the slide or reject the image (for example, the user may repeat the experiment to generate another image with better image quality).
[0127] If the value representing the amount of artifact pixels is determined to be below the artifact region threshold, the biological sample can be scanned at a higher image resolution for digital pathology analysis. In some cases, scanning at a higher image resolution involves switching magnification, such as using a different objective lens or changing the scanner's tube lens. Both operations may involve moving the optical elements.
[0128] Switching resolutions can involve scanning two passes through the slide, which requires additional scanning time. By performing the initial scan at a lower resolution, the initial scan speed can be made faster than the time required to scan the slide at the target image resolution, thus minimizing the additional scanning time. For example, scanning at 5x resolution can result in 1 / 16 the number of pixels compared to scanning at 20x resolution. Such a difference can mean that only a small portion of the time is required for scanning at 5x resolution. In another example, with a line scanner, if the length of the stripe / line is large enough to cover the width (or height) of a given slide at a lower resolution, the scan can be completed in a single sweep through the slide, thus minimizing the increase in total scanning time.
[0129] (b) Artifact detection by converting scanned images to a lower image resolution. In some cases, machine learning models are applied to slide images after scanning biological samples at high image resolution and then converting the slide images to a lower image resolution. Artifact detection can be performed on the slide images before they are further processed (e.g., stored in another database for further digital pathology analysis). Such a design can facilitate the early elimination of low-quality scans before other time-consuming processes (e.g., data transfer, long-term data storage) take place. Recent advances in computing hardware and software algorithms have made such implementations feasible, as processing a full slide image at (e.g.) 20 times higher resolution can be completed in tens of seconds.
[0130] For example, an initial image can be generated by scanning a slide displaying a biological sample at a higher image resolution. A machine learning model for detecting artifacts can be applied to the transformed image to identify one or more artifact pixels. The transformed image can be generated by converting the initial image to a lower image resolution image. The amount of artifact pixels can be determined. A value representing the amount of artifact pixels can be compared to an artifact area threshold. If the value is determined to exceed the artifact area threshold, it can be determined that there is a potential failure in quality control. In some cases, a warning is generated in response to a quality control failure. Alternatively, an artifact mask containing one or more artifact pixels can be generated, allowing the user to rescan the slide or reject the image (for example, the user may repeat the experiment to generate another image with better image quality).
[0131] If the value is determined to be below the artifact region threshold, the initial image scanned at a high image resolution can be accepted and saved directly in DICOM format and / or another file format. In some cases, information corresponding to the artifact pixels (e.g., the location of the artifact pixels in the initial image, an artifact mask at the same or lower image resolution as the initial image, etc.) is saved together with the initial image and / or in a separate file format. Furthermore, subsequent digital pathology analysis can be performed on the initial image.
[0132] (c) Artifact detection per image tile In some cases, a machine learning model can be applied to each image tile of a slide image. The slide image can be divided into sets of image tiles. The machine learning model can be applied to each image tile in a set of image tiles to generate an image mask. The image mask identifies a subset of image tiles, and each image tile in the subset can display one or more artifact pixels. The image mask can then be applied to the image to allow the user to deselect one or more image tiles from the subset, and the deselected image tiles are excluded from further digital pathology analysis. Alternatively, the image mask can be applied to the image to select a subset of image tiles from the image without user input, and then exclude them from further digital pathology analysis.
[0133] As an illustrative example, a portion of a slide of a biological sample can be scanned to obtain a corresponding part of the image (e.g., an image tile). The image tile can then be scanned at the target image resolution. After obtaining the image tile, a machine learning model can be applied to the image tile to identify one or more artifact pixels (e.g., batch size = 1). In some cases, the machine learning model can be applied to multiple image tiles to identify artifact pixels in each image tile (e.g., batch size ≥ 1). Processing of multiple image tiles can be performed based on multi-processing on a GPU or CPU.
[0134] For each image tile that has identified artifact pixels, additional processing can be performed. Additional processing of image tiles with artifact pixels may include: (i) determining the amount of artifact pixels identified in the image tile (e.g., the ratio of artifact pixels to the total number of pixels), and (ii) determining the amount of pixels that represent tissue areas (e.g., the ratio of pixels that represent tissue areas to the total number of pixels). Additional processing can be performed while additional image tiles of the image are being scanned and processed by the machine learning model. In some cases, the image tiles are first downsampled to display the biological sample at a lower image resolution, and the machine learning mode is applied to identify artifact pixels.
[0135] If the amount of artifact pixels determined from an image tile exceeds an artifact region threshold, a warning can be generated to alert the user that the image tile is not expected to accurately represent the corresponding point or region of the biological sample. In some cases, an artifact mask is generated in response to the above determination.
[0136] For the amount of artifact pixels that display tissue areas below the artifact threshold. The entire slide can be scanned at the target resolution for subsequent digital pathological analysis. In addition to this, or alternatively, the scanning system used to generate image tiles (e.g., tile-based scanner, line-based scanner) can be configured to modify its settings based on the detection of artifact pixels. In some cases, the modification of settings with respect to artifacts corresponding to blurred image portions includes: (i) comparing the focus quality of scanned / assembled image tiles in multiple z-planes, and (ii) excluding image tiles in the z-plane where artifact pixels are identified and / or adjusting the z-plane to reduce the artifacts. Such configurations can be integrated with or replace existing autofocus systems in the scanner.
[0137] 3. Additional artifact detection In some cases, in addition to artifact detection during scanning, post-scan artifact detection can be performed. Post-scan artifact detection can further improve the accuracy of artifact detection within images. For example, an algorithm for artifact detection during scanning may be designed specifically for downstream digital pathology analysis, or it may be designed to be applicable to other general algorithms. If integrating a customized artifact detection algorithm into the scanner is not practical, post-scan artifact detection can be used to maintain quality control of the entire slide image for downstream digital pathology analysis.
[0138] In another example, regarding artifact detection at low image resolution, new or different artifacts may arise in subsequent scans at higher image resolutions. In particular, because different objective or tube lenses may be used, and scans may result from separate scanning operations, the out-of-focus areas of the image may differ between these two scans. Therefore, if artifact detection is not performed after scanning, new artifacts may exist that could reduce the accuracy of downstream digital pathology analysis.
[0139] In some cases, post-scan artifact detection is more effective for certain types of machine learning models. For example, post-scan artifact detection may be more effective when certain machine learning models (e.g., recurrent neural networks) integrate features from adjacent image tiles. While image data generated during scanning can be organized during the scan to evaluate adjacent image tiles, such techniques can slow down the scan speed and / or significantly increase the computational load on hardware integrated into or associated with the scanner.
[0140] V. Experimental Results An evaluation was performed to identify the performance level of the machine learning model used to detect artifacts in slide images.
[0141] A. Dataset A set of labels for identifying pixels was collected from each of the 50 full slide images. The labels for the corresponding pixels relate to one of three types of classes: non-tissue, blurred tissue, and unblurred tissue. The full slide images display at least a portion of the biological samples obtained from two cohorts (breast cancer and lung cancer). Each biological sample was stained with one of the following: (1) hematoxylin, (2) single staining for ER, PR, PDL1, or CK7, and (3) double staining for ER / PR or PDL1 / CK7. The chromogens used in the assays were Dabsyl (yellow), Tamra (purple), SRB (red), or DAB (single IHC only). Body slide images were from various tissue types (breast, lung, liver, kidney, and colon) and from single, dual, and triple assays (additional chromogen: Teal, additional biomarkers: LIV1, HER2, CD8, and BCL1).
[0142] From 50 full slide images, 978 image tiles were selected. Each image tile has a size of 512 x 512 pixels and was scanned at 5x image resolution. From the selected image tiles, 462 were used for training, 246 for validation, and 270 for testing. An additional 100 full slide images were selected for independent testing.
[0143] B. Model Selection and Configuration Two modified U-Net machine learning models were selected for evaluation. For the first machine learning model, the number of channels in the intermediate convolutional layer was reduced by half, resulting in Model 1 (7.76 million parameters). For the second machine learning model, the number of channels in the intermediate convolutional layer was reduced by a quarter, resulting in Model 2 (1.94 million parameters).
[0144] C. Image preprocessing and training Each selected image tile was converted to grayscale and augmented with random rescaling, flipping, contrast jittering, and intensity jittering. Each augmented grayscale image tile was concatenated with its corresponding image gradient map (Laplacian filtering with kernel size 3, followed by Gaussian filtering with kernel size 25 and sigma 3). Each of the two U-Net models was trained using the augmented grayscale image tiles with the corresponding gradient features. The training of the two U-Net models was performed using the multi-resolution training technique described in Section III. In particular, the losses calculated from each of the last two processing blocks in the augmentation pass were used for pixel-level classification.
[0145] D. Results Figure 22 shows a set of exemplary graphs 2200 that identify the accuracy and recall scores of machine learning models trained to detect artifact pixels. Graph 2202 shows the accuracy scores corresponding to Model 1 and Model 2, and graph 2204 shows the accuracy scores corresponding to Model 1 and Model 2. The accuracy and recall scores for each of graphs 2202 and 2204 were calculated based on provided test images with corresponding labels. Note that the performance of Model 1 and Model 2 is similar in artifact detection. This may mean that machine learning models with relatively few parameters (e.g., 1.94 million parameters) can be sufficiently effective in artifact detection.
[0146] Figure 23 shows a set of exemplary image masks 2300 generated for sets of images not visible from the same type of assay and the same type of tissue as the training images. Each image shows a tissue section stained using a staining protocol corresponding to a specific type of IHC assay. Figure 23 presents the predicted image mask 2302 generated by Model 2, along with the ground truth image mask 2304. A comparison between the predicted image mask 2302 and the ground truth image mask 2304 demonstrates that Model 2 can accurately identify artifact pixels.
[0147] Furthermore, the trained machine learning model was applied to independent test images (i.e., the set of 100 whole slide images identified in Section VA) to identify artifact pixels. For example, Figure 24 shows a set of 2400 exemplary image masks generated from images displaying invisible assay patterns or tissue types. Figure 24 shows invisible assay patterns. It can be used to qualitatively evaluate the generalizability of machine learning models for detecting misfocus artifacts for both invisible and non-visible tissue types.
[0148] As shown in Figure 24, the exemplary image mask demonstrates accurate artifact detection from invisible assay 2402 and invisible tissue type 2404. The image mask further demonstrates that the trained machine learning model can perform accurate artifact detection for images displaying various types of biological specimens and / or biological specimens stained using staining protocols corresponding to different types of assays.
[0149] VI. Computing Environment Figure 25 shows an example of a computer system 2500 for implementing some embodiments disclosed herein. The computer system 2500 may include a distributed architecture in which some components (e.g., memory and processor) are part of an end-user device and several other similar components (e.g., memory and processor) are part of a computer server. In some cases, the computer system 2500 is a computer system for determining a target genetic feature based on the size distribution of nucleic acid molecules of a biological sample and includes at least a processor 2502, memory 2504, storage device 2506, input / output (I / O) peripherals 2508, communication peripherals 2510, and an interface bus 2512. The interface bus 2512 is configured to communicate, transmit, and transfer data, control, and commands between various components of the computer system 2500. The processor 2502 may include one or more processing units such as a CPU, GPU, TPU, systolic array, or SIMD processor. The memory 2504 and the storage device 2506 include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronically non-volatile computer storage devices such as flash memory, and other tangible storage media. Any of such computer-readable storage media may be configured to store instructions or program code that embody aspects of the present disclosure. Furthermore, the memory 2504 and the storage device 2506 include computer-readable signal media.
[0150] A computer-readable signaling medium includes a propagating data signal in which a computer-readable program code is embodied. Such a propagating signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signaling medium is not a computer-readable storage medium, but includes any computer-readable medium capable of communicating, propagating, or transmitting a program for use in connection with the computer system 2500.
[0151] Furthermore, memory 2504 includes an operating system, programs, and applications. Processor 2502 is configured to execute stored instructions and includes, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. For example, computing system 2500 can execute instructions (e.g., program code) that configure processor 2502 to perform one or more of the operations described herein. Program code includes, for example, code that performs analysis of array data, and / or any other suitable application that performs one or more of the operations described herein. Instructions include, for example, C, C++, C#, Visual It can include processor-specific instructions generated by a compiler or interpreter from code written in any suitable computer programming language, including Basic, Java, Python, Perl, JavaScript, and ActionScript.
[0152] The program code can be stored in memory 2504 or any suitable computer-readable medium and executed by processor 2502 or any other suitable processor. In some embodiments, all modules in the computer system for predicting loss of heterozygosity in HLA alleles are stored in memory 2504. In additional or alternative embodiments, one or more of these modules from the above computer system are stored in different memory devices of different computing systems.
[0153] The memory 2504 and / or processor 2502 may be virtualized and hosted, for example, in another computing system in a cloud network or data center. The I / O peripherals 2508 include user interfaces such as a keyboard, screen (e.g., a touchscreen), microphone, speaker, and other input / output devices, as well as computing components such as a graphical processing unit, serial ports, parallel ports, a universal serial bus, and other input / output peripherals. The I / O peripherals 2508 are connected to the processor 2502 via any of the ports coupled to the interface bus 2512. The communication peripherals 2510 are configured to facilitate communication between the computer system 2500 and other computing devices over a communication network and include, for example, a network interface controller, modem, wireless and wired interface cards, antenna, and other communication peripherals. For example, the computing system 2500 can communicate with one or more other computing devices via a data network using the network interface device of the communication peripheral 2510 (e.g., a computing device that determines the genetic features of a target based on the size distribution of nucleic acid molecules in a biological sample, and another computing device that generates sequence data of the biological sample in question).
[0154] While the subject matter has been described in detail with respect to its specific embodiments, it will be understood that achieving the above understanding may facilitate modifications, variations, and the creation of equivalents to such embodiments by those skilled in the art. Therefore, it should be understood that this disclosure is presented for illustrative purposes only, not limitation, and does not preclude such modifications, variations, and / or additional inclusions to the subject matter that would be readily apparent to those skilled in the art. Indeed, the methods and systems described herein may be embodied in various other forms. Furthermore, various omissions, substitutions, and modifications may be made in the forms of the methods and systems described herein without departing from the spirit of this disclosure. The appended claims and equivalents are intended to encompass forms or modifications that fall within the scope and spirit of this disclosure.
[0155] Unless otherwise specified, discussions throughout this specification using terms such as “processing,” “computing,” “calculating,” “determining,” and “identifying” are understood to refer to the operation or process of one or more computers or similar electronic computing devices or devices, such as computers, that manipulate or transform data represented as physical electronic or magnetic quantities in the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0156] The one or more systems described herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditional on one or more inputs. A suitable computing device ranges from a general-purpose computing device to a dedicated computing device implementing one or more embodiments of this subject, and may also include a multipurpose microcontroller that accesses stored software that programs or configures the computing system. This includes a processor-based computing system. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in software used for programming or configuring computing devices.
[0157] Certain embodiments of the methods disclosed herein may be implemented in the operation of such computing devices. The order of the blocks shown in the above examples can be changed, and for example, blocks can be rearranged, combined, and / or divided into subblocks. Certain blocks or processes may be executed in parallel.
[0158] The conditional language used herein, such as in particular "can," "could," "might," "may," and "eg," is generally intended to convey that certain examples include certain features, elements, and / or steps, but other examples do not, unless otherwise specified or understood in the context in which they are used. Therefore, such conditional language is not generally intended to imply that features, elements, and / or steps are somehow required in one or more examples, or that one or more examples necessarily include logic for determining whether these features, elements, and / or steps are included in or performed in any particular example, with or without input or prompting from the author.
[0159] Terms such as “comprising,” “including,” and “having” are synonyms and are used comprehensively and in an open-ended manner, without excluding additional elements, features, actions, etc. Similarly, the term “or” is used in a comprehensive sense (rather than an exclusive sense), for example, when used to connect a list of elements, “or” means one, some, or all of the elements in the list. The use of “adapted to” or “configured to” herein means an open and comprehensive language that does not exclude devices adapted or configured to perform additional tasks or steps. Furthermore, the use of “based on” means open and comprehensive in that a process, step, calculation, or other action “based” on one or more enumerated conditions or values may actually be based on additional conditions or values beyond those enumerated. Similarly, the use of “based at least in part on” means that the process, step, calculation, or other action may actually be based on additional conditions or values beyond those enumerated, “at least in part on” one or more enumerated conditions or values. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to limit the scope of the text.
[0160] The various features and processes described above may be used independently of each other or combined in various ways. All possible combinations and partial combinations are intended to fall within the scope of this disclosure. Furthermore, in some implementations, certain methods or process blocks may be omitted. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith may be executed in other appropriate sequences. For example, the described blocks or states may be executed in an order other than the specifically disclosed order, or multiple blocks or states may be combined into a single block or state. The exemplary blocks or states may be executed in series, in parallel, or in any other way. Blocks or states may be added to or removed from the disclosed examples. Similarly The exemplary systems and components described herein may be configured differently from those described. For example, elements may be added, removed, or rearranged compared to the disclosed examples.
Claims
1. Access to an image showing at least a portion of a biological sample, Applying an image preprocessing algorithm to the aforementioned image to generate a preprocessed image, wherein the preprocessed image includes a plurality of labeled pixels, and each of the plurality of labeled pixels is associated with a label that predicts whether the pixel accurately represents a corresponding point or region of at least a portion of the biological sample; Applying a machine learning model to the preprocessed image to identify one or more labeled pixels from the plurality of labeled pixels, wherein the one or more labeled pixels are predicted to have been incorrectly labeled by the image preprocessing algorithm; For each of the one or more labeled pixels, modify the label, To generate a training image that includes at least one or more labeled pixels with the label corrected, Outputting the aforementioned training image Methods that include...
2. The method according to claim 1, wherein the label further identifies the type of artifact, and the pixel is further predicted to display at least a portion of the artifact corresponding to the type of artifact.
3. Applying a blur threshold to each of the multiple labeled pixels, Based on the application of the blur threshold, it is determined that the labeling of a further labeled pixel among the plurality of labeled pixels is incorrect, To modify the labels corresponding to the further labeled pixels mentioned above. The method according to claim 1 or 2, further comprising:
4. The method according to claim 3, wherein the blur threshold is determined based on the performance of a downstream algorithm for a set of z-axis images showing at least a portion of the biological sample across the depth dimension.
5. The method according to any one of claims 1 to 4, wherein the image preprocessing algorithm includes image segmentation, morphological processing, image thresholding, image filtering, image contrast enhancement, blur detection, or a combination thereof.
6. The method according to any one of claims 1 to 5, wherein the label further predicts whether the pixel displays at least a portion of an artifact associated with a particular artifact type.
7. The method according to claim 6, wherein the specific artifact types include blurred areas, tissue folds, and foreign bodies.
8. Accessing a training image that displays at least a portion of a biological sample, wherein the training image includes a plurality of labeled pixels, and each of the plurality of labeled pixels is associated with a label that predicts whether the pixel accurately displays a corresponding point or region of the at least portion of the biological sample. This involves accessing training images, Accessing a machine learning model that includes a set of convolutional layers, wherein the machine learning model is configured to apply each of the convolutional layers in the set of convolutional layers to a feature map representing an input image. The machine learning model is trained to detect one or more artifact pixels in an image at a target image resolution, wherein the artifact pixels are predicted not to accurately represent at least a portion of a point or region of the biological sample, and the training is performed accordingly. For each of the labeled pixels among the plurality of labeled pixels in the training image, The first loss of the labeled pixels at the first image resolution is determined by applying the first convolutional layer of the set of convolutional layers to a first feature map representing the training image at the first image resolution, Determining a second loss for the labeled pixels at the second image resolution by applying a second convolutional layer of the set of convolutional layers to a second feature map representing the training image at the second image resolution, wherein the second resolution has a higher image resolution than the first image resolution. Determining the total loss for the labeled pixel based on the first loss and the second loss, Based on the total loss, it is determined that the machine learning model has been trained to detect one or more artifact pixels at the target image resolution. The machine learning model, including the above, is trained to detect one or more artifact pixels in an image at a target image resolution. Outputting the aforementioned trained machine learning model Methods that include...
9. The method according to claim 8, further comprising converting the training images into grayscale training images, wherein the machine learning model is trained using the grayscale training images.
10. The method according to claim 8, further comprising converting the plurality of labeled pixels of the training image from a first color space to a second color space to generate a modified training image, wherein the machine learning model is trained using the modified training image.
11. The method according to claim 8, wherein the total loss is determined based on the sum of the first loss and the second loss.
12. The method according to claim 8, wherein the total loss is determined based on the average of the first loss and the second loss.
13. The method according to claim 8, wherein the target image resolution is the first image resolution.
14. Accessing an image that displays at least a portion of a biological sample, wherein the image has a first image resolution, Accessing a machine learning model trained to detect artifact pixels in an image of a second image resolution, wherein the first image resolution has a higher image resolution than the second image resolution. Convert the aforementioned image to the second image resolution and at least a portion of the biological sample To generate a converted image that displays the following: Applying the machine learning model to the converted image to identify one or more artifact pixels from the converted image, wherein it is predicted that the artifact pixels among the one or more artifact pixels do not accurately represent at least a portion of the point or region of the biological sample; To generate an output that includes one or more artifact pixels. Methods that include...
15. The output is an image mask containing the one or more artifact pixels, and the method is The image mask is superimposed on the image to distinguish a set of pixels in the image from the one or more artifact pixels, Applying a cell classification model to the aforementioned set of pixels The method according to claim 14, further comprising:
16. The method according to claim 14 or 15, wherein the output identifies the amount of one or more artifact pixels.
17. The method according to any one of claims 14 to 16, further comprising adjusting one or more scan parameters of a scanning device using the output.
18. One or more data processors, A non-temporary computer-readable storage medium containing instructions, wherein, when the instructions are executed on the one or more data processors, the one or more data processors cause some or all of the methods disclosed herein to be executed. A system that includes these features.
19. A computer program product tangibly embodied in a non-temporary machine-readable storage medium, comprising instructions configured to cause one or more data processors to perform some or all of the methods disclosed herein.