Automatic Segmentation of Artifacts in Pathological Tissue Images

The method automates the segmentation of artifact regions in digital pathology images using a trained generation network, addressing the challenges of manual annotation and improving diagnostic accuracy and efficiency.

JP7699227B2Active Publication Date: 2025-06-26VENTANA MEDICAL SYSTEMS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023571834
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2022-05-20
Publication Date
2025-06-26
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

Manual annotation of digital pathology slides to exclude artifact regions is costly, time-consuming, and subjective, leading to potential misdiagnosis or delays in diagnosis due to the degradation of image quality by artifacts.

Method used

A computer-implemented method for image segmentation using a generation network trained with image pairs to automatically detect and segment artifact regions in digital pathology images, thereby generating a segmentation image that highlights the boundaries of these regions.

Benefits of technology

The method significantly reduces the burden of manual annotation, enhances image quality by excluding artifacts, and improves the accuracy and efficiency of pathological diagnosis by providing a scalable, robust, and accessible image quality control algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699227000001
    Figure 0007699227000001
  • Figure 0007699227000002
    Figure 0007699227000002
  • Figure 0007699227000003
    Figure 0007699227000003
Patent Text Reader

Abstract

A technique for image segmentation of a digital pathology image may include accessing an input image indicative of a section of tissue, and generating a segmentation image by processing the input image using a generative network, the generative network being trained using a dataset including a plurality of image pairs. The segmentation image indicates artifact region boundaries for each of a plurality of artifact regions of the input image, at least one of the plurality of artifact regions being indicative of an abnormality that is not a tissue structure. Each image pair of the plurality of image pairs includes a first image of the section of tissue, the first image including at least one artifact region, and a second image indicating artifact region boundaries for each of the at least one artifact region of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 191,567, filed May 21, 2021, which is hereby incorporated by reference in its entirety for all purposes.

[0002] The present disclosure relates to digital pathology, and more particularly, to techniques including semantic segmentation of digital pathology images.

Background Art

[0003] Histopathology can include the examination of slides made from tissue sections for various reasons such as disease diagnosis, determination of response to treatment, and / or development of drugs to fight disease. Since tissue sections and their cells are virtually transparent, slide preparation typically involves staining the tissue sections to make the relevant structures more visible. Digital pathology can include scanning stained slides to obtain digital images, which can then be examined by digital pathology image analysis and / or interpreted by a pathologist.

[0004] In addition to one or more regions to be analyzed, a digital pathology slide can include regions that should be excluded from further analysis. Such regions can include, for example, regions that may be distracting during the task of annotating a tumor region and / or regions that may produce false results if not excluded from an automated scoring operation. The task of manually annotating the slide to indicate regions to be excluded is costly, time - consuming, and subjective.

Summary of the Invention

[0005] In various embodiments, a computer-implemented method of image segmentation includes accessing an input image that shows a tissue section and includes a plurality of artifact regions, and generating a segmentation image by processing the input image using a generation network, the generation network being trained using a training data set that includes a plurality of image pairs. The segmentation image shows the boundaries of the artifact regions for each of the plurality of artifact regions of the input image. In this method, at least one of the plurality of artifact regions shows an abnormality that is not a tissue structure, and in each image pair of the plurality of image pairs, the pair includes a first image of a tissue section that includes at least one artifact region, and a second image that shows the boundaries of the at least one artifact region of the first image for each of the at least one artifact regions of the first image.

[0006] In some embodiments, the abnormality is out-of-focus, a fold in the tissue section, a pigment deposit in the tissue section, or something trapped between the tissue section and the slide cover.

[0007] In some embodiments, the segmentation image includes a binary segmentation mask.

[0008] In some embodiments, the method further includes generating an annotated image that includes the segmentation image overlaid on the input image.

[0009] In some embodiments, the method further includes estimating the quality of the input image based on the total area of the plurality of artifact regions.

[0010] In some embodiments, the input image includes a second plurality of artifact regions, and the method further includes generating a second segmentation image by processing the input image using a second generation network, the second generation network being trained using a second training data set including a second plurality of image pairs, the second segmentation image indicating boundaries of the artifact regions for each of the second plurality of artifact regions of the input image, and at least one of the second plurality of artifact regions indicating a biological structure of tissue.

[0011] In some embodiments, the computer-implemented method further includes a user determining a diagnosis of the subject based on the segmentation image.

[0012] In some embodiments, the computer-implemented method further includes a user performing treatment with a compound based on (i) the segmentation image and / or (ii) the diagnosis of the subject.

[0013] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the one or more methods disclosed herein.

[0014] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of the one or more methods disclosed herein.

[0015] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system, when executed on one or more data processors, includes a non-transitory computer-readable storage medium containing instructions that cause the one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes.

[0016] The terms and expressions used are used as terms of explanation and not of limitation, and in the use of such terms and expressions, it is not intended to exclude any equivalents or portions of the features shown and described, but it should be recognized that various modifications are possible within the scope of the claimed invention. Thus, the claimed invention is specifically disclosed by way of embodiments and optional features, but it should be understood that those skilled in the art may adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0017] The patent or application file contains at least one drawing made in color. Copies of this patent or patent application publication with color drawings will be provided by the Patent Office upon request and payment of the necessary fees.

[0018] Aspects and features of various embodiments will become more apparent by way of example with reference to the accompanying drawings.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

[0020] The systems, methods, and software disclosed herein facilitate the segmentation of artifact regions in digital pathology images (e.g., WSIs). Although certain embodiments are described, these embodiments are presented by way of example only and are not intended to limit the scope of protection. The devices, methods, and systems described herein may be embodied in various other forms. Further, various omissions, substitutions, and changes may be made to the exemplary methods and system forms described herein without departing from the scope of protection.

[0021] I. Overview Digital pathology may involve the interpretation of digital images to accurately diagnose a subject and guide treatment decisions. In digital pathology solutions, the image analysis workflow can be established to automatically detect or classify objects of biological interest (e.g., positive tumor cells, negative tumor cells, etc.). FIG. 1 shows an exemplary diagram of a workflow 100 of a digital pathology solution. The workflow 100 of the digital pathology solution includes receiving a specimen at block 110, creating a specimen at block 120 (e.g., fixing, dehydrating, washing, wax infiltration, embedding), performing microtomy at block 130 (e.g., slicing a specimen created to obtain one or more tissue sections), staining the tissue sections at block 140, digitizing the created slide at block 150 (e.g., scanning), performing quality control (QC) by an expert on the slide image (e.g., WSI) and sending it out at block 160, and reporting the analysis results (e.g., diagnosis) by a pathologist at block 240.

[0022] Figure 2A shows another exemplary diagram of workflow 200 of a digital pathology solution. Workflow 200 of the digital pathology solution includes, at block 210, obtaining a tissue slide; at block 220, scanning a preselected site or the entire tissue slide by a digital image scanner (e.g., a whole slide image (WSI) scanner) to obtain a digital image; at block 230, performing image analysis on the digital image using one or more image analysis algorithms; and at block 240, scoring an object of interest (e.g., quantitative or semi - quantitative scoring such as positive, negative, moderate, weak, etc.) based on the image analysis.

[0023] For example, the evaluation of tissue changes caused by a disease can be performed by examining thin tissue sections. A tissue sample (e.g., a tumor sample) can be sliced to obtain a series of sections, each section having a thickness of, for example, 4 - 5 microns. Since tissue sections and their cells are virtually transparent, slide preparation typically involves staining the tissue sections to make the relevant structures more visible. For example, sections of different tissues may be stained with one or more different stains to represent different characteristics of the tissues.

[0024] After each section is mounted on a slide, it is scanned to create a digital image, which can then be examined by digital pathology image analysis and / or interpreted by a pathologist (e.g., using image viewer software). A pathologist may review and manually annotate the digital image of the slide (e.g., a tumor site, necrosis, etc.) to enable the use of image analysis algorithms to extract meaningful quantitative metrics (e.g., to detect and classify biological objects of interest). Conventionally, a pathologist may manually annotate each consecutive image of multiple tissue sections from a tissue sample to identify the same aspect for each consecutive tissue section. Figure 2B shows an example of a digital pathology image annotated by drawing a boundary with a pen.

[0025] The stained sections of tissue samples on pathology slides can have various types of defects that can obscure the information to be conveyed. Such defects may be due to causes that occur during tissue preparation. For example, a tissue section may be thicker in one part than another tissue section (Figure 3 shows an example), and / or the section may contain one or more pigment deposits (e.g., precipitates) (Figure 4 shows an example). Other defects may be due to causes that occur during slide preparation. For example, a tissue section may be folded on the slide (Figure 5 shows an example), and / or unwanted substances (e.g., one or more foreign objects such as dirt or other debris, one or more air bubbles) may be trapped between the tissue section and the slide cover (Figures 6 and 7 show examples). Other defects may be due to causes that occur during the handling of the prepared slides by a person and / or during automated processing (e.g., scanning), such as pen marks (Figure 8 shows an example), pinholes (i.e., out-of-focus) (Figure 26 shows three examples), and / or false features introduced during the digital stitching of individual scan tiles to obtain a WSI (Figure 27 shows some examples where the black and white boxes in the image are enlarged in detail images c and d respectively, and the black and white boxes in image b are enlarged in detail images e and f respectively). Such defects are examples of "artifacts" that indicate something (e.g., a structure) that does not actually exist in the tissue in a part of the image. To support the accurate analysis of slide images, it may be desirable to detect such artifacts and, if possible, process the image to improve the extent to which the image conveys accurate information about the subject so that the interpretation of such information can be improved (e.g., for pathological diagnosis, prognosis, and / or treatment selection).

[0026] For this purpose, current practice may include the evaluation of digital pathology images by a pathologist to assess the quality (e.g., of the image, of the section, and / or generally) before the image of the stained sample section is analyzed (e.g., to detect and / or characterize a particular organism or a particular biomarker in the stained sample section). If the quality of the stained sample section is poor, the corresponding digital pathology image may be discarded from the digital pathology analysis performed on a given subject. However, detecting artifacts in digital pathology images can be both subjective and time-consuming.

[0027] Image artifacts such as those described above (also referred to as "anomalies") are an important issue in the adoption of the digital pathology (DP) workflow (e.g., as shown in FIGS. 1 and / or 2A). Artifacts can degrade the quality of the image of the stained sample section, which can lead to misdiagnosis or a delay in diagnosis. For example, multiple artifacts present in an image of a stained sample (e.g., out-of-focus, watermark, and tissue folds) can potentially obscure diagnostic features. These artifacts can even render the tissue entirely useless. To accurately analyze the slide image, it may be desirable to detect such artifacts in the slide image and, if possible, process the slide image so that the artifacts do not interfere with the pathological diagnosis.

[0028] In addition to artifacts that can reduce the overall quality of the scanned image, slide images may contain areas that show the biological characteristics (e.g., structure) of the tissue section and that should be excluded from subsequent analysis operations. Examples of such features that are generally excluded from the analysis (also referred to as "biological artifacts") include necrotic tissue, blood pools, and serum. Other such features may be excluded from subsequent analysis operations depending on the application. For example, it may be desirable to exclude areas showing macrophages to facilitate the programmed death ligand 1 (PD-L1) scoring of slides stained using the SP142 assay.

[0029] In current practice, the DP workflow relies on a pathologist to visually inspect digital slide images to identify artifacts and outline such areas that are to be excluded from subsequent analysis by manually annotating the images, which is a labor-intensive and costly process. FIG. 9 shows an example of a manually annotated slide image of FIG. 2B to outline the folds (black) of tissue, the necrotic regions (yellow), and the blood (cyan) to be excluded from the area of analysis (the tumor shown in green).

[0030] The generation of the segmentation image can be performed by a trained generation network, which may include parameters learned during the training of a fully convolutional network (FCN). The FCN may further include an encoding / decoding network and may be configured as a U-Net.

[0031] One exemplary embodiment of the present disclosure is a computer-implemented method for image segmentation, comprising accessing an input image showing a tissue section and including a plurality of artifact regions, and generating a segmentation image by processing the input image using a generation network, the generation network being trained using a training data set including a plurality of image pairs, the segmentation image showing the boundaries of the plurality of artifact regions of the input image. In this method, at least one of the plurality of artifact regions indicates an abnormality that is not a tissue structure, and in each image pair of the plurality of image pairs, the pair includes a first image of a tissue section including at least one artifact region, and a second image showing the boundaries of the at least one artifact region of the first image for each of the at least one artifact regions of the first image.

[0032] Advantageously, the methods of image segmentation described herein may be applied to optimize the digital pathology workflow at one or more different levels. In one example, such methods may be applied to provide a more scalable, robust, and accessible image quality control (QC) algorithm. In another example, such methods may be applied to optimize annotation and review time by pathologists. In a further example, such methods may be applied to optimize the performance of other downstream tasks (i.e., automated image analysis).

[0033] II. Definitions As used herein, when an action is “based on” something, this means that the action is at least partially based on at least a portion of that something.

[0034] As used herein, the terms “substantially,” “approximately,” and “about” are defined as being mostly (and including all of the specified thing), but not necessarily all, of the specified thing, as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms “substantially,” “approximately,” or “about” may be replaced by “within [percentage]” of the specified thing, where the percentage includes 0.1 percent, 1 percent, 5 percent, and 10 percent.

[0035] As used herein, the terms "sample", "biological sample", or "tissue sample" refer to any sample containing biomolecules (such as proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof) obtained from any organism, including viruses. Other examples of organisms include mammals (humans; domestic animals such as cats, dogs, horses, cows, and pigs; and laboratory animals such as mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and tissue needle biopsies), cell samples (such as cytological smear specimens like Pap smears, or blood smear specimens, or samples of cells obtained by microdissection), or cell fractions, fragments, or organelles (such as those obtained by lysing cells and separating their components by centrifugation or other means). Other examples of biological samples include blood, serum, urine, semen, fecal matter, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (such as obtained by surgical or needle biopsy), nipple aspirate, earwax, milk, vaginal fluid, saliva, swabs (such as buccal swabs), or any substance containing biomolecules removed from an initial biological sample. In certain embodiments, the term "biological sample" as used herein refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or a portion thereof obtained from a subject.

[0036] III. Techniques for Automatic Segmentation of Artifacts in Digital Pathology Images Applications of automated techniques for segmenting artifacts in digitized slides, as described herein, may include tools for facilitating pathological review and / or pathological scoring (e.g., immunohistochemistry (IHC) estimation scoring workflows, hematoxylin and eosin (H&E) tumor segmentation, tumor prediction workflows) by removing regions to be excluded from further analysis. Such tools may output a segmentation mask, apply the mask to the original corresponding input image, and also be implemented to input the masked image into an image analysis algorithm for further processing (e.g., to segment tumor cells, to count cells, etc.). Automated techniques for segmenting artifacts in digitized slides, as described herein, may be implemented to be scanner-independent, tissue-independent, and / or stain-independent.

[0037] Figure 10 shows a flowchart of an exemplary process 1000 for image segmentation according to some embodiments. Process 1000 may be executed using one or more computing systems, models, and networks (as described herein with respect to FIGS. 14 and 15, for example). Referring to FIG. 10, at block 1004, access an input image showing a tissue section and including a plurality of artifact regions. At least one of the plurality of artifact regions exhibits an abnormality that is not a tissue structure. At block 1008, a segmentation image is generated by processing the input image using a generation network. The segmentation image indicates the boundaries of the artifact regions for each of the plurality of artifact regions of the input image. The generation network is trained using a training data set that includes a plurality of image pairs, where each pair includes a first image of a tissue section that includes at least one artifact region, and a second image that indicates the boundaries of the at least one artifact region of the first image for each of the at least one artifact regions of the first image. In some embodiments, process 1000 also includes generating an annotated image that includes the segmentation image overlaid on the input image and / or estimating the quality of the input image based on the total area of the plurality of artifact regions.

[0038] In some embodiments of process 1000, the abnormality is a blur, a fold in the tissue section, or a pigment deposit in the tissue section.

[0039] In some embodiments of process 1000, the segmentation image includes a binary segmentation mask.

[0040] In some embodiments of process 1000, the generation network is implemented as a fully convolutional network, as a U-Net, and / or as an encoder / decoder network.

[0041] In some embodiments of process 1000, the input image includes a second plurality of artifact regions, and the process includes generating a second segmentation image by processing the input image using a second generation network, where the second generation network is trained using a second training data set that includes a second plurality of image pairs. Generating the second segmentation image also includes that at least one of the second plurality of artifact regions shows a biological structure of the tissue, and the second segmentation image shows the boundaries of the artifact regions for each of the second plurality of artifact regions of the input image.

[0042] One or more methods according to the present disclosure may be implemented to reduce the burden of manual contouring of at least some types of artifacts in whole slide images by a pathologist and to focus on viable tumors. For example, the task of manually delineating the contours of artifacts can be simplified into performing a QC review on the results of an automated artifact detection process, as described herein. In the previous DP algorithm, it may not be possible to delineate the contours of the tumor sites and the contours of their own exclusion regions, and thus such a process may enable a fully automated DP algorithm, as it results in a manual, time-consuming preprocessing step for the pathologist.

[0043] FIG. 11 shows an example of the automatic segmentation of tissue folds resulting from an implementation of process 1000. For comparison, FIG. 12 shows an example of the contouring of tissue folds performed by a pathologist on the same slide image. The segmentation resulting from the algorithm is further detailed, and it can be seen that it includes many smaller folds that the pathologist did not annotate at all. Furthermore, in the algorithm segmentation, the tissue folds are very closely followed (thus minimizing the area that would otherwise be inaccurately excluded), and the contouring of tissue folds by a pathologist includes the adjacent regions of each fold.

[0044] Manually drawing the contours of the exclusion regions is very time-consuming, and depending on the amount of exclusion parts that may exist, it may take up to 30 minutes or more to finish one whole slide image. Therefore, as the time spent on slides starts to accumulate, pathologists may experience a decline in accuracy or may miss small exclusion regions. Additionally, variability among and within observers, which can also affect the results of subsequent analysis, is another issue. In contrast, the algorithm is highly reproducible, and these contour drawings will be close to the actual boundaries of the exclusion regions. Better boundary contour drawings protect against false results due to incomplete exclusion of unwanted regions and retain more of the desired regions for analysis. For contour drawings that can be inherently subjective, such as those of blurry regions, the improved uniformity provided by an automated solution as described herein may also be highly desirable.

[0045] It may be desirable to perform process 1000 to provide a scalable image quality control (QC) algorithm. FIG. 13A shows an example of such an application 1300 of process 1000 according to some embodiments. At block 1310, a created tissue sample (e.g., a created slide) is provided. At block 1320, the created sample is digitized (scanned) to obtain a slide image (e.g., a WSI) using a DP scanner equipped with an artifact detection model. For example, the DP scanner may be configured to execute an embodiment of process 1300 as described herein. At block 1330, an expert performs a QC review of the annotated slide image (e.g., as shown in FIG. 13B). At block 1340, the slide image that passes the QC review is transferred for automated tissue analysis or visual tissue analysis (e.g., as shown in FIG. 13C). Slide images that fail the QC review are rejected, and the process may return to block 1310 to rescan the corresponding created sample (e.g., correct blurring or stitching artifacts) or perform other operations. For example, the sample may be washed and / or the slide may be recreated if possible (e.g., to correct artifacts due to unwanted substances, etc.), or the slide may be replaced using a new section of the sample. In another example of application 1300, at block 1320, it may be enhanced or replaced by a quality score calculated based on the segmentation image generated by process 1000. For example, process 1000 may be configured to calculate a quality score based on the total image area consumed by the artifact region (alternatively, the total foreground area of the image consumed by the artifact region, where this foreground area is the area occupied by the tissue section). In such a case, the process may be configured to indicate a failed quality score if the site exceeds a threshold.

[0046] It may be desirable to perform process 1000 to provide for automatic exclusion of artifact regions (e.g., to reduce the workload of a pathologist). FIG. 14 shows an example of such an application 1400 of process 1000 according to some embodiments. In the previous workflow (top), annotation by a pathologist of an input scan image (e.g., as shown in the example of FIG. 2B) includes annotation of one or more viable tumor regions (e.g., target regions) for analysis, as well as annotation of artifact regions and, optionally, other non-target regions excluded from the analysis (e.g., as shown in the example of FIG. 9). In block 1410, automatic artifact detection is performed to provide automatic annotation of the exclusion regions (e.g., according to an implementation of process 1000 as described herein). The automatic artifact detection may be performed using a neural network architecture such as that shown in FIG. 15D, for example. In block 1420, the task of annotation by a pathologist of the input scan image is simplified by reducing the types of exclusion regions that are manually annotated, including annotation of viable tumor regions, or by eliminating the task of annotating the exclusion regions.

[0047] FIG. 15A shows an example of a training workflow 1500 according to some embodiments. At block 1510, the WSI is manually annotated (e.g., as shown in FIG. 15B) to indicate the boundaries of artifact regions. At block 1520, the original (unannotated) WSI and the annotated WSI are divided into training patches (e.g., as shown in FIG. 15B) (e.g., of size 64×64 pixels, 128×128 pixels, or 256×256 pixels), and one or more training sets of matched pairs of images are created. Each matched pair includes an unannotated version of the corresponding image patch and an annotated version of the same image patch. At block 1530, one or more training sets of matched pairs are used to train a custom fully convolutional network (e.g., as shown in FIG. 15D) for performing image segmentation, and the trained network is used to automatically delineate the boundaries of artifact regions in digital pathology slides (e.g., WSIs).

[0048] Since image artifacts (e.g., anomalies) may occur regardless of the scanner brand, tissue designation, or staining used, it may be desirable to train the network to be scanner - agnostic, tissue - agnostic, and / or staining - agnostic. To ensure that the detection algorithm is robust against different scanners, tissue types, staining types, and creation protocols, it may be desirable to train the deep - learning model using images from different bright - field microscopy methods (e.g., H&E, IHC PD - L1, and epidermal growth factor receptor (EGFR)), different tissues (e.g., lung and colon), and different scanners (e.g., Ventana DP200 and Ventana Aperio).

[0049] As described above, it may be desirable to train a network to map an input image to a corresponding segmentation image using supervised learning (e.g., based on an original / annotated image patch pair of an exemplar). Supervised learning may include penalizing a model that makes an error regarding a prediction error or mismatch between a generated output segmentation mask and an available "ground truth" mask (e.g., an image patch manually annotated).

[0050] In one example, the network is trained to simultaneously handle different anomalies (e.g., tissue fold artifacts and blurring artifacts) as (e.g., as a single "excluded" class). However, in such a situation of supervised learning, it may be desirable to handle some different types of artifacts separately (e.g., as different classes). For example, it may be desirable to separate training for artifacts indicating anomalies that are not part of the tissue structure from training for artifacts indicating the biological structure of the tissue, both with respect to the output of the network and providing separate ground truths for each of the desired classes of artifacts. In one example, the prediction model 1415 is implemented to support segmentation of both anomalies and unwanted biological structures by modifying the last layer of the network to support multi-class output.

[0051] FIG. 16 shows a block diagram illustrating an exemplary computing environment 1600 (i.e., a data processing system) for segmenting an input image according to various embodiments. The computing environment 1600 can include an analysis system 1605 for training and executing a prediction model, such as a two-dimensional CNN model. More specifically, the analysis system 1605 can include training subsystems 1610a-n (where "a" and "n" represent any natural numbers) for building and training each prediction model 1615a-n (which can be individually referred to as prediction model 1615 herein or collectively referred to as prediction model 1615) used by other components of the computing environment 1600. The prediction model 1615 can be a deep convolutional neural network (CNN), such as an initial neural network, a residual neural network ("Resnet"), or a recurrent neural network, such as a long short-term memory ("LSTM") model or a gated recurrent unit ("GRU") model, etc., which are machine learning ("ML") models. The prediction model 1615 can also be any other suitable ML model trained to segment non-target regions (e.g., artifact regions), segment target regions, or provide image analysis of target regions, such as a two-dimensional CNN ("2DCNN"), Mask R-CNN, Feature Pyramid Network (FPN), Dynamic Time Warping ("DTW") technique, Hidden Markov Model ("HMM"), etc., or a combination of one or more of such techniques, such as CNN-HMM or MCNN (multi-scale convolutional neural network). The computing environment 1600 can use the same type or different types of prediction models trained to segment non-target regions, segment target regions, or provide image analysis of target regions. For example, the computing environment 1600 can include a first prediction model (e.g., U-Net) for segmenting non-target regions (e.g., artifact regions).The computing environment 1600 can also include a second prediction model (e.g., 2DCNN) for segmenting a target region (e.g., a region of tumor cells). The computing environment 1600 can also include a third model (e.g., CNN) for image analysis of the target region. The computing environment 1600 can also include a fourth model (e.g., HMM) for diagnosing a disease for treatment or prognosis of a subject such as a patient. In other examples according to the present disclosure, still other types of prediction models may be implemented.

[0052] In various embodiments, each of the respective prediction models 1615a - n corresponding to the training subsystems 1610a - n is separately trained based on one or more sets of the input image elements 1620a - n. In some embodiments, each of the input image elements 1620a - n includes image data from one or more scanned slides. Each of the input image elements 1620a - n may correspond to a single specimen from which the underlying image data for the image was collected and / or image data from one day. The image data may include the image, as well as any information regarding the imaging platform on which the image was generated. For example, a tissue section may need to be stained by the application of a staining assay that contains one or more different biomarkers associated with a chromogenic stain for brightfield imaging or a fluorophore for fluorescence imaging. The staining assay can use a chromogenic stain for brightfield imaging, an organic fluorophore, a quantum dot, or an organic fluorophore together with a quantum dot for fluorescence imaging, or any other combination of stain, biomarker, and viewing or imaging device. Further, a typical tissue section is processed on an automated staining / assay platform that applies the staining assay to the tissue section, resulting in a stained sample. There are various commercially available products suitable for use as the staining / assay platform, and as one example, there is the product VENTANA SYMPHONY of assignee Ventana Medical Systems. The stained tissue section can be provided to, for example, a microscope or an imaging system on a whole slide scanner having a microscope and / or imaging components, and as one example, there is the product VENTANA iScan Coreo of assignee Ventana Medical Systems. Multiple tissue slides can be scanned with an equivalent multiplexed slide scanner system.Additional information provided by the imaging system can include any information regarding the staining platform, such as the concentration of the chemicals used in the staining, the reaction time of the chemicals applied to the tissue in the staining, and / or the pre - analysis conditions of the tissue, e.g., the age of the tissue, the fixation method, the duration, how the sections were embedded, cut, etc.

[0053] The input image elements 1620a - n may include one or more training input image elements 1620a - d, validation input image elements 1620e - g, and unlabeled input image elements 1620h - n. It should be understood that it is not necessary to access the input image elements 1620a - n corresponding to the training, validation, and unlabeled groups simultaneously. For example, an initial set of the training and validation input image elements 1620a - n may be first accessed and used to train the prediction model 1615, and the unlabeled input image elements may then be accessed or received (e.g., at one or more subsequent times) and used by the trained prediction model 1615 to provide a desired output (e.g., segmentation of non - target regions). In some cases, the prediction models 1615a - n are trained using supervised training, and each of the training input image elements 1620a - d and optionally the validation input image elements 1620e - g are associated with one or more labels 1625 that identify the "correct" interpretation of the identification of non - target regions, target regions, and various biological substances and structures within the training input image elements 1620a - d and validation input image elements 1620e - g. The labels can alternatively or additionally be used to classify the corresponding training input image elements 1620a - d and validation input image elements 1620e - g, or the pixels therein, with respect to the presence and / or interpretation of staining associated with normal or abnormal biological structures (e.g., tumor cells). In a particular example, alternatively or additionally, the labels can be used to classify the corresponding training input image elements 1620a - d and validation input image elements 1620e - g at the time when the underlying image was captured or at a subsequent time point corresponding to (e.g., a predetermined duration following the time when the image was captured).

[0054] In some embodiments, the training subsystems 1610a - n include a feature extractor 1630, a parameter data store 1635, a classifier 1640, and a trainer 1645, which are collectively used to train the prediction model 1615 based on training data (e.g., training input image elements 1620a - d) and to optimize the parameters of the prediction model 1615 during supervised or unsupervised training. Optionally, the training process may include iterative operations to find a set of parameters of the prediction model 1615 that minimizes the loss function of the prediction model 1615. Each iteration may involve finding a set of parameters of the prediction model 1615 such that the value of the loss function using the set of parameters is less than the value of the loss function using another set of parameters in the previous iteration. The loss function may be constructed to measure the difference between the output predicted using the prediction model 1615 and the label 1625 contained in the training data. Once the set of parameters is identified, the prediction model 1615 is trained and available for segmentation and / or prediction as designed.

[0055] In some embodiments, the training subsystems 1610a - n access training data from the training input image elements 1620a - d at the input layer. The feature extractor 1630 may preprocess the training data to extract relevant features (e.g., edges) detected in specific portions of the training input image elements 1620a - d. The classifier 1640 receives the extracted features and, according to weights associated with a set of hidden layers in one or more prediction models 1615, converts the features into one or more output metrics that segment non - target or target regions, provides image analysis, provides a diagnosis of a disease for the treatment or prognosis of a subject such as a patient, or provides a combination thereof. The trainer 1645 may train the feature extractor 1630 and / or the classifier 1640 by facilitating the learning of one or more parameters using the training data corresponding to the training input image elements 1620a - d. For example, the trainer 1645 can use the backpropagation technique to facilitate the learning of the weights associated with the set of hidden layers of the prediction model 1615 used by the classifier 1640. Backpropagation may, for example, use the Stochastic Gradient Descent (SGD) algorithm to cumulatively update the parameters of the hidden layers. The learned parameters may include, for example, weights, biases, and / or other hidden - layer - related parameters, which can be stored in the parameter data store 1635.

[0056] Individually trained prediction models or a group of trained prediction models can be deployed to process unlabeled input image elements 1620h-n to segment non-target or target regions, provide image analysis, provide a diagnosis of a disease for the treatment or prognosis of a subject such as a patient, or provide a combination thereof. More specifically, a trained version of the feature extractor 1630 can generate a feature representation of the unlabeled input image elements that can then be processed by a trained version of the classifier 1640. In some embodiments, image features can be extracted from the unlabeled input image elements 1620h-n based on one or more convolutional blocks, convolutional layers, residual blocks, or pyramid layers that utilize an extension of the prediction model 1615 in the training subsystems 1610a-n. The features can be organized in a feature representation such as a feature vector of the image. The prediction model 1615 can be trained to learn the feature type based on the classification and subsequent adjustment of the parameters in the hidden layer including the fully connected layer of the prediction model 1615.

[0057] In some embodiments, the image features extracted by the convolutional block, convolutional layer, residual block, or pyramid layer include a feature map that is a matrix of values representing one or more portions of a specimen slide on which one or more image processing operations (e.g., edge detection, sharpening of image resolution) have been performed. These feature maps may be flattened for processing by a fully connected layer of a prediction model 1615 that outputs one or more metrics corresponding to a non-target region mask, a target region mask, or a current or future prediction regarding the specimen slide. For example, the input image elements may be fed to an input layer of the prediction model 1615. The input layer can include nodes corresponding to specific pixels. The first hidden layer can include a set of hidden nodes, each of which is connected to a plurality of input layer nodes. Nodes within subsequent hidden layers can likewise be configured to receive information corresponding to a plurality of pixels. Thus, the hidden layers can be configured to learn to detect features that extend across a plurality of pixels. Each of the one or more hidden layers can include a convolutional block, convolutional layer, residual block, or pyramid layer. The prediction model 1615 can further include one or more fully connected layers (e.g., a softmax layer).

[0058] At least a portion of the training input image elements 1620a - d, the validation input image elements 1620e - g, and / or the unlabeled input image elements 1620h - n may include, or may be derived from, data obtained directly or indirectly from a source that may or may not be an element of the analysis system 1605. In some embodiments, the computing environment 1600 includes an imaging device 1650 for imaging a sample to obtain image data such as a multi - channel image (e.g., a multi - channel fluorescence or bright - field image) having several (e.g., 10 - 16, etc.) channels. The imaging device 1650 may include, but is not limited to, a camera (e.g., an analog camera, a digital camera, etc.), an optical system (e.g., one or more lenses, a sensor focus lens group, a microscope objective lens, etc.), an imaging sensor (e.g., a charge - coupled device (CCD) or a complementary metal - oxide - semiconductor (CMOS) image sensor, etc.), or photographic film, etc. In digital embodiments, the image - capturing device may include a plurality of lenses that cooperate to provide on - the - fly focusing. An image sensor, e.g., a CCD sensor, can capture a digital image of a specimen. In some embodiments, the imaging device 1650 is a bright - field imaging system, a multispectral imaging (MSI) system, or a fluorescence microscopy system. The imaging device 1650 may image an image using invisible electromagnetic radiation (e.g., UV light) or other imaging techniques. For example, the imaging device 1650 may include a microscope and a camera arranged to image an image magnified by the microscope. The image data received by the image analysis system 1605 may be the same as, and / or may be derived from, the raw image data captured by the imaging device 1650.

[0059] In some cases, the labels 1625 associated with the training input image elements 1620a - d and / or the validation input image elements 1620e - g may be received, or may be derived from data received from one or more provider systems 1655 that may be associated with (e.g.) a doctor, a nurse, a hospital, a pharmacist, etc., each associated with a particular subject. The received data may include, for example, one or more medical records corresponding to a particular subject. The medical records may indicate, for example, whether the subject had a tumor with respect to a period corresponding to the time when one or more input image elements associated with the subject were collected or a defined period thereafter, and / or the stage of progression of the subject's tumor (e.g., by identifying an expert diagnosis or characterization along standard metrics and / or such metrics as total metabolic tumor volume (TMTV)). The received data may further include pixels of the location of a tumor or tumor cells within one or more input image elements associated with the subject. Thus, the medical records may include or be used to identify one or more labels for each of the training / validation input image elements 1620a - g. The medical records may further indicate each of one or more treatments (e.g., medication) the subject has received and the period during which the subject has received the treatment. In some cases, the images or scans input into one or more training subsystems are received from the provider system 1655. For example, the provider system 1655 may receive an image from the imaging device 1650 and then transmit the image or scan to the analysis system 1605 (e.g., along with a subject identifier and one or more labels).

[0060] In some embodiments, data received or collected by one or more of the imaging devices 1650 may be aggregated with data received or collected by one or more of the provider systems 1655. For example, the analysis system 1605 may identify corresponding or identical identifiers for the subject and / or period so as to associate the image data received from the imaging device 1650 with the label data received from the provider system 1655. The analysis system 1605 may further use metadata or automated image analysis to process the data and determine which training subsystem to supply a particular data component to. For example, the image data received from the imaging device 1650 may correspond to an entire slide, or multiple regions of a slide or tissue. Metadata, automated alignment, and / or image processing may indicate for each image which region of the slide or tissue the image corresponds to. For example, automated alignment and / or image processing may include detecting whether the image has image characteristics corresponding to a slide substrate, or has a biological structure and / or shape associated with a particular cell such as a white blood cell. The label-related data received from the provider system 1655 may be slide-specific, region-specific, or subject-specific. When the label-related data is slide-specific or region-specific, metadata or automated analysis (e.g., using natural language processing or text analysis) can be used to identify which region the particular label-related data corresponds to. When the label-related data is subject-specific, the same label data (for a given subject) may be supplied to each of the training subsystems 1610a - n during training.

[0061] In some embodiments, the computing environment 1600 can further include a user device 1660 that can be associated with a user who requests and / or coordinates the execution of one or more iterations of the analysis system 1605 (e.g., each iteration corresponds to one execution of the model and / or one generation of the output of the model). The user can correspond to a physician, a researcher (e.g., associated with a clinical trial), a subject, a medical professional, etc. Thus, it will be understood that in some cases, the provider system 1655 can include the user device 1660 and / or can act as the user device 1660. Each iteration may be associated with a particular subject (e.g., a person) that may be different from the user (although this is not required). The iteration request may include and / or be accompanied by information about a particular subject (e.g., the name or other identifier of the subject such as an unspecified patient identifier). The iteration request may include identifiers of one or more other systems that collect data such as input image data corresponding to the subject. In some cases, the communication from the user device 1660 includes an identifier for each subject represented in a particular set of subjects in response to a request to perform an iteration for each subject in the set.

[0062] Upon receiving a request, the analysis system 1605 can send the request for unlabeled input image elements (e.g., including an identifier of the subject) to one or more corresponding imaging systems 1650 and / or provider systems 1655. The trained prediction model 1615 can then process the unlabeled input image elements to segment non-target or target regions, provide image analysis, provide a diagnosis of a disease for treatment or prognosis of a subject such as a patient, or provide a combination thereof. The results for each identified subject may include or be based on the segmentation and / or one or more output metrics from the trained prediction model 1615 developed by the training subsystems 1610a - n. For example, the segmentation and / or one or more output metrics can include or be based on the output generated by one or more fully connected layers of a CNN. In some cases, such output may be further processed using (e.g.) a softmax function. Further, the output and / or the further processed output may then be aggregated using an aggregation technique (e.g., random forest aggregation) to generate one or more subject-specific metrics. One or more results (e.g., including plane-specific output and / or one or more subject-specific output and / or their processed versions) can be transmitted to and / or utilized by the user device 1660. In some cases, some or all of the communication between the analysis system 1605 and the user device 1660 occurs via a website. It will be understood that the CNN system 1605 can gate access to results, data, and / or processing resources based on authentication analysis.

[0063] Although not explicitly shown, it will be understood that computing environment 1600 may further include a development device associated with a developer. Communication from the development device to the components of computing environment 1600 may indicate what types of input images are used in each prediction model 1615 in analysis system 1605, the number and types of models used, the hyperparameters of each model (e.g., learning rate and number of hidden layers), how much data request is formatted, which training data is used (e.g., and how can access to the training data be obtained), and which verification techniques are used, and / or how the controller process is configured.

[0064] As described above, prediction model 1615 may be implemented using a U-Net architecture, which includes an encoder having layers that gradually downsample the input to a bottleneck layer and a decoder having layers that gradually upsample the bottleneck output to produce an output. The U-Net also includes skip connections having feature maps of equal size between the encoding and decoding layers, and these connections concatenate the channels of the feature maps of the encoding layer with the channels of the feature maps of the corresponding decoding layer. In a particular example, prediction model 1615 is updated via cross-entropy loss measured between a generated image and an expected output image (e.g., "predicted image" and "ground truth", respectively). Other examples of loss functions that may be used to update prediction model 1615 include, for example, L1 loss or L2 loss.

[0065] Figures 17 - 25 show exemplary results of the implementation of process 1000, where FIGS. 18, 21, and 24 are enlarged versions of the portions outlined by the dotted lines in FIGS. 17, 20, and 23 respectively, and FIGS. 19, 22, and 25 are enlarged versions of the portions outlined by the dashed lines in FIGS. 17, 20, and 23 respectively. In this example, the model was trained on tissue fold artifacts and blurry artifacts, and the trained model was tested on various whole slide images resulting from different scanners, different stainings, and different tissue types. As described herein, the method can similarly be extended to other common tissue - based or image - based artifacts.

[0066] The algorithm was found to identify many small regions with out - of - focus problems (see FIGS. 22 and 25 in particular), for which it would seem nearly impossible for a pathologist to similarly outline without spending hours on the same slide. Further, as discussed above with reference to FIGS. 11 and 12, it can be seen that the algorithm draws segmentation boundaries that are very close to the actual boundaries of the artifacts. It is impossible for a pathologist to achieve such precision at high magnifications such as 40x (and thus significantly increase the size of the image area being reviewed) without spending much more time trailing the boundaries of the artifacts. Such a method can provide an automated solution for the segmentation of artifacts in DP images that is more scalable, more robust, and less expensive compared to manual annotation by a pathologist.

[0067] V. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system, when executed on one or more data processors, includes a non-transitory computer-readable storage medium containing instructions that cause the one or more data processors to execute some or all of one or more of the methods and / or some or all of one or more of the processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to execute some or all of one or more of the methods and / or some or all of one or more of the processes disclosed herein.

[0068] The terms and expressions used are used as terms of description and not of limitation, and in the use of such terms and expressions, it is not intended to exclude any equivalents or portions thereof of the features shown and described, but it should be recognized that various modifications are possible within the scope of the claimed invention. Thus, the claimed invention is specifically disclosed by way of embodiments and optional features, but it should be understood that those skilled in the art may adopt modifications and variations of the concepts disclosed herein, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0069] This specification provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. More precisely, the description of the preferred exemplary embodiments will provide those skilled in the art with a practical explanation for implementing various embodiments. It should be understood that various changes may be made in the functions and arrangements of the elements without departing from the spirit and scope as set forth in the appended claims.

[0070] The following description sets forth specific details in order to provide a thorough understanding of the embodiments. It will be understood, however, that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may sometimes be shown as components in block diagram form in order not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

Claims

1. A method for image segmentation, comprising: accessing an input image showing a tissue section and including a plurality of artifact regions; generating a segmentation image by processing the input image using a generation network, wherein the generation network is trained using a training dataset including a plurality of image pairs; wherein the segmentation image shows the boundaries of the plurality of artifact regions in the input image for each of the plurality of artifact regions; at least one of the plurality of artifact regions shows an abnormality that is not part of the tissue structure; each image pair of the plurality of image pairs comprises a first image of a tissue section including at least one artifact region, and a second image showing the boundary of the at least one artifact region for each of the at least one artifact region in the first image.

2. The method according to claim 1, wherein the abnormality is blurring.

3. The method according to claim 1, wherein the abnormality is a fold in the tissue section.

4. The method according to claim 1, wherein the abnormality is a pigment deposit in the tissue section.

5. The method according to claim 1, wherein the segmentation image includes a binary segmentation mask.

6. The method according to claim 1, further comprising generating an annotated image including the segmentation image superimposed on the input image.

7. The method according to claim 1, further comprising estimating the quality of the input image based on the total area of the plurality of artifact regions.

8. The input image includes a second plurality of artifact regions, and the method further comprises generating a second segmentation image by processing the input image using a second generation network, wherein the second generation network is trained using a second training dataset including a second plurality of image pairs; wherein the second segmentation image shows the boundaries of the second plurality of artifact regions in the input image for each of the second plurality of artifact regions. The method according to claim 1, wherein at least one of the second plurality of artifact regions shows the biological structure of the tissue.

9. The method according to claim 1, wherein the generation network is implemented as a fully convolutional network.

10. The method according to claim 1, wherein the generation network is implemented as a U-Net.

11. The method according to claim 1, wherein the generation network is implemented as an encoding / decoding network.

12. The method according to claim 1, wherein the generation network is updated via a cross-entropy loss measured between an image by the generation network and a predicted output image.

13. The method according to claim 1, further comprising determining a diagnosis of a subject by a user based on the segmentation image.

14. The method according to claim 13, further comprising treating with a compound by the user based on (i) the segmentation image and / or (ii) the diagnosis of the subject.

15. One or more data processors, A non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the method according to any one of claims 1 to 14 A system comprising.

16. One or more data processors, When executed on the one or more data processors, cause the one or more data processors to Access an input image showing a tissue section and including a plurality of artifact regions, Generating a segmentation image by processing the input image using a generation network, wherein the generation network is trained using a training data set including a plurality of image pairs, generating a segmentation image A non-transitory computer-readable storage medium containing instructions for performing a method comprising A system comprising, The segmentation image shows the boundary of the artifact region for each of the plurality of artifact regions of the input image, At least one of the plurality of artifact regions shows an abnormality that is not the structure of the tissue, Each image pair of the plurality of image pairs is A first image of a tissue section, the first image including at least one artifact region, and for each of the at least one artifact region of the first image, a second image showing the boundary of the artifact region A system comprising.

17. The system according to claim 16, wherein the abnormality is blurring, a fold in the section of the tissue, or a deposit of pigment in the section of the tissue.

18. A computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to execute some or all of the method according to any one of claims 1 to 14.

19. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product including instructions configured to cause one or more data processors to access an input image showing a tissue section and including a plurality of artifact regions, and generate a segmentation image by processing the input image using a generation network, the generation network being trained using a training data set including a plurality of image pairs, generating the segmentation image including instructions configured to cause execution of a method including the segmentation image showing the boundary of each of the plurality of artifact regions of the input image, at least one of the plurality of artifact regions showing an abnormality that is not a structure of the tissue, each image pair of the plurality of image pairs a first image of a tissue section, the first image including at least one artifact region, and for each of the at least one artifact region of the first image, a second image showing the boundary of the artifact region A computer program product comprising.

20. The computer program product according to claim 19, wherein the abnormality is blurring, a fold in the section of the tissue, or a deposit of pigment in the section of the tissue.

Citation Information

Patent Citations

  • Scanning / pre-scanning quality control of slides

    EP3757872A1

  • Systems and methods for probabilistic segmentation in anatomical image processing

    JP2020503603A

  • Automated pattern recognition and scoring method for histological images.

    JP2021500691A

  • Information processing device, information processing method, and information processing system

    WO2020174862A1