Adaptive Learning Framework for Digital Pathology

The adaptive learning framework in digital pathology uses existing datasets for preconditioning and continuous learning to reduce resource requirements and maintain performance across varying data distributions, addressing the inefficiencies and high costs of existing model development and adaptation methods.

JP2025521529AActive Publication Date: 2025-07-10VENTANA MEDICAL SYSTEMS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024575040
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2023-06-22
Publication Date
2025-07-10
Estimated Expiration
2043-06-22

AI Technical Summary

Technical Problem

Existing digital pathology models are resource-intensive and time-consuming to develop, requiring large numbers of annotations, and lack generalizability to unseen data, necessitating continuous redevelopment, while current adaptation methods suffer from catastrophic forgetting and high computational costs.

Method used

An adaptive learning framework that utilizes existing annotated datasets for preconditioning and employs continuous learning techniques to reduce resource requirements and adapt models to different image domains with minimal annotations, using strategies like Elastic Weight Consolidation and Learning Without Forgetting to prevent forgetting.

Benefits of technology

The framework enables efficient initial model development with fewer annotations and rapid convergence, reducing computational resources and maintaining model performance across varying data distributions, thus addressing the challenges of resource-intensity and adaptability in digital pathology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025521529000001_ABST
    Figure 2025521529000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to techniques for efficient development of an initial model using an adaptive learning framework, as well as for efficient model updating and / or adaptation to different image domains. For efficient development of the initial model, a two-stage development strategy may be implemented as follows. That is, stage 1; pre-adjustment of the model, in which an artificial intelligence system utilizes existing annotated datasets and improves learning skills through training on these datasets, and stage 2: target model training, in which the artificial intelligence system utilizes the learning skills learned in stage 1 to expand to a different image domain (target domain) that requires a smaller number of annotations in the target domain than conventional learning methods. To efficiently perform model updating and adaptation to new datasets after the initial model development, a digital pathology scenario is identified, an adaptive learning method is selected based on the scenario, and the model is updated and adapted to the new dataset using the adaptive learning method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to digital pathology, and more particularly, to techniques for efficient development of initial models and efficient model updating and / or adaptation to different image domains using an adaptive learning framework.

Background Art

[0002] Digital pathology involves scanning slides (e.g., histopathology or cytopathology glass slides) into digital images interpretable on a computer screen. Tissues and / or cells within the digital images can then be examined by digital pathology image analysis and / or interpreted by a pathologist for various reasons including disease diagnosis, assessment of response to treatment, and drug development to fight diseases. To examine tissues and / or cells (substantially transparent) within the digital image, the pathology slides can be prepared using various staining assays (e.g., immunohistochemistry) that selectively bind to tissue and / or cell components. Immunofluorescence (IF) is a technique for analyzing assays that bind fluorescent dyes to antigens. Multiple assays responsive to different wavelengths can be utilized on the same slide. These multiplexed IF slides enable understanding of the complexity and heterogeneity of the immune landscape of the tumor microenvironment and the potential impact on the tumor's response to immunotherapy. In some assays, the target antigen for a stain in the tissue may be called a biomarker. Digital pathology image analysis can then be performed on the digital image of the stained tissue and / or cells to identify and quantify staining for antigens (e.g., biomarkers indicating various cells such as tumor cells) in the biological tissue.

[0003] Artificial intelligence and machine learning-based methods and / or techniques show great promise in digital pathology image analysis such as cell detection, counting, localization, classification, and patient prognosis. Many computing systems equipped with machine learning techniques including convolutional neural networks (CNNs) have been proposed for image classification and digital pathology image analysis such as cell detection and classification. For example, a CNN can have a series of convolutional layers as hidden layers, and this network structure enables the extraction of representative features for object / image classification and digital pathology image analysis. In addition to object / image classification, machine learning techniques are also implemented for image segmentation. Image segmentation is the process of dividing a digital image into multiple segments (sets of pixels also known as image objects). The purpose of segmentation is to simplify and / or transform the representation of the image into something more meaningful and easier to analyze. For example, image segmentation is typically used to find objects such as cells and boundaries (lines, curves, etc.) within an image. To perform image segmentation for large data (e.g., whole slide pathology images), the image is first divided into many small patches. A computing system equipped with machine learning techniques is trained to classify each pixel of these patches, all pixels of the same class are combined into one segmented area of each patch, and then all segmented patches are combined into one segmented image (e.g., a segmented whole slide pathology image). Thereafter, machine learning techniques are further implemented to predict or further classify the segmented areas (e.g., positive cells for a given biomarker, negative cells for a given biomarker, or cells without staining expression) based on the representative features associated with the segmented areas. Summary of the Invention

[0004] Artificial intelligence and machine learning-based approaches have achieved excellent performance in digital pathology. However, developing such models is very time-consuming and resource-intensive. Hundreds of thousands of annotations are required just to build a highly reliable model from the very beginning in the initial development stage. Moreover, once developed, the model has limited generalizability to unseen data, which leads to an inevitable continuous investment in developing new models even for related tasks. This specification discloses a framework for reducing resource requirements throughout the development process of AI-based digital pathology algorithms, including initial model development and subsequent model updates, improvements, and adaptation to different datasets. Specifically, what is disclosed herein is a model preconditioning stage using existing annotated datasets related to but not necessarily similar to the target dataset for building the model, whereby in the initial model development, only a small number of annotations are required to generate the model with reasonable accuracy. In subsequent model update and adaptation stages, an adaptive learning workflow is used for multiple digital pathology scenarios and strategies to select the best learning method for efficient model updates without having to train all the data from the beginning.

[0005] In various embodiments, a computer-implemented method is provided that, in a data processing system, obtains a first annotated training set of images for training a machine learning algorithm to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within an image, where the first annotated training set of images is within a first image domain; divides the first annotated training set of images into mini-sets of images by the data processing system, where each mini-set represents a distinct modeling subtask and includes a limited number of examples; trains the machine learning algorithm in a first stage using the mini-sets of images by the data processing system to generate a conditional machine learning model configured to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within a new image; labels a limited number of images from a target data set by the data processing system to generate a second annotated training set of images for training a machine learning algorithm to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within an image, where the second annotated training set of images is within a second image domain; and trains the conditional machine learning model in a second stage using the second annotated training set of images by the data processing system to generate a target machine learning model configured to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within a new image, where the number of classes targeted in the first stage is part of or matches the number of classes targeted in the second stage.

[0006] In some embodiments, the first annotated training set of images is a digital pathology image that includes one or more types of cells.

[0007] In some embodiments, partitioning involves, when only one mini-set of images is available, selecting a subset of classes that is a portion of or matches the number of classes targeted in the second stage for each separate modeling subtask, selecting a limited number of examples based on the selected subset of classes, and, when multiple mini-sets of images are available, for each separate modeling subtask, either (i) mixing examples from multiple mini-sets of images, selecting a subset of classes that is a portion of or matches the number of classes targeted in the second stage, and selecting a limited number of examples from the mixed examples based on the selected subset of classes, or (ii) selecting one mini-set of images from the multiple mini-sets of images, selecting a subset of classes that is a portion of or matches the number of classes targeted in the second stage, and selecting a limited number of examples from the selected mini-set of images based on the selected subset of classes.

[0008] In some embodiments, the second stage further includes applying a machine learning model preconditioned to generate a feature vector representation for each example in a second annotated training set of images, combining the feature vector representations from examples of the same class to generate one representation for each target class and using one representation for each target class as a prototype in the target class, generating a feature vector representation for the remaining portion of the unlabeled images or image regions from the target dataset, and comparing each feature vector representation from the unlabeled images to the prototype in the target class based on the distance between the feature vector representation from the unlabeled image and the prototype for the target class.

[0009] In some embodiments, the first stage is an internal learning loop where a machine learning algorithm updates the model weights or parameters on one sub-task for a predetermined or flexible number of epochs to generate a loss in a validation set of the image after model update, which is shown as L-subtask-i for the i-th sub-task, for adaptation to a target dataset, and an external learning loop that searches for a set of model initializations that generate a preconditioned machine learning model when used to update all sub-tasks, each with only a limited number of examples, by finding a model initialization that minimizes the sum of all losses, shown as the sum of L-subtask-i, where i ranges from 1 to the number of sub-tasks calculated from the validation set of the sub-task images for model initialization.

[0010] In some embodiments, training the first stage involves performing an iterative operation to learn a set of parameters to detect, characterize, classify, or a combination thereof, some or all regions or objects within a mini-set of images that maximize or minimize a cost function, where each iteration involves finding a set of parameters for the machine learning algorithm such that the value of the cost function using the set of parameters is greater or less than the value of the cost function using another set of parameters in the previous iteration, and the cost function is constructed to measure the difference between predictions made for some or all regions or objects using the machine learning algorithm and the ground truth labels given to the mini-set of images.

[0011] In some embodiments, training in the second stage involves performing an iterative operation to learn a set of parameters to detect, characterize, classify, or a combination thereof, of some or all regions or objects within a second annotated training set of images that maximizes or minimizes a cost function, where each iteration involves finding a set of parameters for a conditional machine learning model such that the value of the cost function using the set of parameters is greater than or less than the value of the cost function using another set of parameters in the previous iteration, and the cost function is constructed to measure the difference between predictions made for some or all regions or objects using the conditional machine learning model and the ground truth labels given in the second annotated training set of images.

[0012] In some embodiments, the computer-implemented method further includes identifying a digital pathology scenario, selecting an adaptive continuous learning method for updating a target machine learning model based on the digital pathology scenario, and updating the target machine learning model based on the adaptive continuous learning method to generate an updated machine learning model.

[0013] In some embodiments, the digital pathology scenario is a data incremental scenario, a domain incremental scenario, a class incremental scenario, or a task incremental scenario.

[0014] In some embodiments, the adaptive continuous learning method is selected from the group including Elastic Weight Consolidation (EWC), Learning Without Forgetting (LWF), Incremental Classifier and Representation Learning (iCaRL), Continuous Prototype Evaluation (CoPE), A-GEM, and parameter separation methods.

[0015] In some embodiments, the computer-implemented method further includes providing the target machine learning model and / or the updated machine learning model.

[0016] In some embodiments, providing includes deploying a target machine learning model and / or an updated machine learning model to a digital pathology system.

[0017] In some embodiments, a computer-implemented method further includes receiving, by a data processing system, a new image; inputting the new image into a target machine learning model or an updated machine learning model; detecting, characterizing, classifying, or a combination thereof, some or all regions or objects within the new image by the target machine learning model or the updated machine learning model; and outputting, by the target machine learning model or the updated machine learning model, an inference based on the detection, characterization, classification, or a combination thereof.

[0018] In some embodiments, the computer-implemented method further includes determining, by a user, a diagnosis of a subject related to the new image, the diagnosis being determined based on an inference output by the target machine learning model or the updated machine learning model.

[0019] In some embodiments, the computer-implemented method further includes treating, by a user, the subject based on (i) an inference output by the target machine learning model or the updated machine learning model, and / or (ii) a diagnosis of the subject.

[0020] In some embodiments, training a machine learning algorithm includes implementing meta-learning principles to enable a first stage to generate a preconditioned machine learning model using a limited number of examples.

[0021] In some embodiments, training a preconditioned machine learning model includes implementing meta-learning principles to enable a second stage to generate a target machine learning model using a limited number of images.

[0022] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more data processors, cause the one or more data processors to perform some or all of the one or more methods disclosed herein.

[0023] In some embodiments, a computer program product is provided that is tangibly embodied on a non-transitory machine-readable storage medium and contains instructions configured to cause one or more data processors to perform some or all of the one or more methods disclosed herein.

[0024] The terms and expressions used are used as terms for explanation and not for limitation, and there is no intention to exclude any equivalents of the features shown and described, and it is recognized that various changes are possible within the scope of the claimed invention. Therefore, although the invention described in the claims is specifically disclosed by the embodiments and any features, changes and modifications of the concepts disclosed herein may be reclassified by those skilled in the art, and such changes and modifications are to be considered within the scope of the invention as defined by the appended claims.

Brief Description of the Drawings

[0025] Aspects and features of various embodiments will become more apparent from the description of the embodiments in connection with the accompanying drawings.

[0026]

Figure 1

[0027]

Figure 2

[0028]

Figure 3

[0029]

Figure 4

[0030]

Figure 5

[0031]

Figure 6

[0032]

Figure 7

[0033]

Figure 8

[0034]

Figure 9

[0035]

Figure 10

[0036]

Figure 11

[0037]

Figure 12

[0038]

Figure 13

[0039]

Figure 14

[0040]

Figure 15

[0041]

Figure 16

[0042]

Figure 17

[0043]

Figure 18

[0044]

Figure 19

[0045]

Figure 20

[0046]

Figure 21

[0047]

Figure 22

[0048]

Figure 23

[0049]

Figure 24

[0050]

Figure 25

DETAILED DESCRIPTION OF THE INVENTION

[0051] Specific embodiments are described, but these embodiments are presented by way of example only and are not intended to limit the scope of protection. The apparatuses, methods, and systems described herein can be embodied in various other forms. Further, various omissions, substitutions, and changes in the form of the exemplary methods and systems described herein may be made without departing from the scope of protection.

[0052] I. Overview Artificial intelligence and machine learning-based approaches have achieved unprecedented performance in solving complex problems in digital pathology-based analysis, thereby automatically suggesting potential lesions and thus improving the reliability in diagnosis and reducing the workload of pathologists by reducing subjectivity. A common strategy for ensuring the performance of highly reliable digital pathology algorithms is to initially collect and train models with a large number of annotations for each target image domain in digital pathology both (a) during the initial model training phase and (b) during the model update phase for performance improvement or adaptation to different image domains. However, basic tasks for histopathological image analysis such as image classification, semantic segmentation, and object detection require manual creation and curation of annotations by pathologists. Such time-consuming and resource-intensive model development has brought about two major challenges in digital pathology.

[0053] First, there is the problem of efficiently and effectively establishing an artificial intelligence model for a target digital pathology dataset. Due to the large amount of data required and the limited generalizability to invisible data distributions, most deep learning models take months to develop when a large set of annotations specific to the target dataset needs to be generated and verified by experts. However, in such a time-consuming and resource-intensive development process, there is a challenge in meeting the ever-increasing demand for such models. Considering new multiplexing technologies in the market, in the future, a variety of new assays will be developed along with deep learning-based digital pathology analysis. Furthermore, the ever-evolving patient population and unpredictable disease pandemics require efficient model development strategies to address changing needs in diagnosis.

[0054] Second, there is the problem of efficiently adapting existing model systems to related but different datasets. Image digitization conditions are constantly evolving, especially with the progress of technology. The use of different staining chromogens, changes in staining, digital scanners and vendor platforms result in changes in the appearance of digitized images. The popularity and availability of digital pathology also bring about a nearly continuous stream of data with biological variation due to heterogeneous disease samples to be analyzed. The ever-increasing data volume and complexity require the development of artificial intelligence algorithms that can adapt and maintain performance under various conditions. Conventional techniques of retraining with manually task-specific annotations on newer data batches rapidly increase the requirements for data storage, training time and computing power, subsequently increasing costs, delaying product releases, and ultimately reaching prohibitively high costs for the organization, as shown in Figure 1.

[0055] Active learning-based approaches have been designed since before the deep learning era to reduce the number of annotations required for model development. These approaches utilize heuristic scoring strategies to query a small subset of the most beneficial unlabeled examples within a dataset. By repeatedly adding only such a small set of selected examples for model retraining, active learning aims to gradually improve model performance and thus avoid the need to annotate many potentially less informative examples. However, models developed using active learning are applicable only to test images with the same or an identical distribution as the training images. For example, a model trained by active learning in one immunohistochemistry (IHC) assay (domain 1) cannot be generalized to another domain, e.g., another IHC assay (domain 2). As another example, a model trained on one tissue type cannot be easily transferred to other tissue types. Thus, such approaches cannot fully address the issues of model inflexibility and lack of model adaptability, unlike the adaptive learning framework disclosed herein. Additionally, active learning requires iterative model retraining using all existing examples and newly annotated examples in each development iteration and thus cannot help reduce the computational power and memory space requirements.

[0056] Alternatively, transfer learning and domain adaptation have been used to improve model generalization to some extent. In transfer learning, all or part of the model weights are targeted with a large annotated dataset applied to a new image domain, while domain adaptation aims to use only a small amount of annotations from the target domain for model development or no annotations at all. However, both frameworks can suffer from catastrophic forgetting, where the performance of a model targeted at one image domain (source domain) degrades significantly when utilized to train in another image domain (target domain), thus "forgetting" the knowledge learned from the previous training process on source domain data. Domain adaptation algorithms aim to generate models with good performance for both the source and target domains even without annotations from the target domain (unsupervised domain adaptation), but current practice still relies on a validation set from the target domain for model selection, which inevitably tends to overfit to the validation set.

[0057] To address these and other issues, various embodiments disclosed herein relate to methods, systems, and computer-readable storage media for (1) reducing the resource requirements necessary to develop an artificial intelligence model for the unseen distribution of digital pathology data in an initial model development stage, and (2) reducing the resource requirements for subsequent iterations of model development aimed at improving or adapting an initial artificial intelligence model constructed according to (1) to a related but not identical dataset.

[0058] To reduce resource requirements in the initial model development stage, techniques are implemented that precondition an artificial intelligence system to learn useful features from existing digital pathology data. Such designs utilize existing annotated digital pathology data sets that are related to but not necessarily similar to the target data, allowing the artificial intelligence system to distill its learning skills through a pre-training stage using these related data sets. In this specification, "learning skills" include one or more of the following, namely, the best set of model initializations, the best set of model weights that can generalize to unseen data, the best set of model architectures, etc. With these learning skills, the preconditioned model requires only a small number of annotation sets to achieve reasonable performance. For example, 75% accuracy is achieved with <50 annotated images for classifying tissue types from tumor types for which the model has not been trained, and >2000 images for training a model specific to this tumor type using a conventional artificial intelligence model.

[0059] To reduce resource requirements in subsequent model development or model adaptation to different data sets, continuous learning techniques and algorithms are implemented to enable model updates using sequentially acquired data without training the model from scratch on all existing and new data sets. Continuous learning algorithms provide solutions for learning from a sequence of data streams, but the challenge is to prevent old knowledge from being forgotten (catastrophic forgetting). To select the most effective continuous learning algorithms for diverse digital pathology data, a set of strategies is designed that target the various model update requirements commonly encountered in digital pathology applications, and the corresponding algorithms are implemented to continuously learn without training from scratch and without a decrease in model performance for previously encountered data (e.g., catastrophic forgetting by a machine learning model).

[0060] These various techniques and algorithms are implemented in an adaptive learning framework that includes the following features and advantages.

[0061] (1)Effectively reduce resource requirements throughout the entire development process of artificial intelligence-based digital pathology algorithms. The adaptive learning framework addresses the challenges of the process that requires resources for both initial model development and subsequent model updating / adaptation, as shown in Figure 2. The adaptive learning framework enables efficient initial model development with a small number of annotations (e.g., less than 50) and faster model improvement (adaptive learning dot) to reach convergence to good model performance compared to conventional artificial intelligence strategies (training from scratch). The arrows indicate the model update process, and various methods can be applied, including self-supervised learning, continuous adaptive learning, conditional continuous adaptive learning (with a small set of annotations), and conventional training (with sufficient annotations).

[0062] (2) The adaptive learning framework can be applied to model development targeting images from different image domains in digital pathology. In this specification, an image domain refers to a set of images having a specific sample distribution. Examples of different image domains include different image modalities (IHC images vs. H&E images), IHC assays (IHC targeting Ki67 vs. CK7 vs. PDL1-CK7), tumor types and subtypes (breast cancer vs. lung cancer), scanners from different vendors, etc.

[0063] (3) The adaptive learning framework utilizes existing digital pathology datasets (from one or more image domains) during the preconditioning stage for initial model development, and the corresponding preconditioned model can be applied to the same or related but different image domains as the image domain in the preconditioning stage.

[0064] (4) The designed algorithm for model update strategy selection enables effective expansion of the existing model to adapt to different image domains after initial model development.

[0065] (5) Users such as developers can flexibly apply only the initial model development strategy, only the model update strategy, or both, according to the usability of the initial model and the need to update the model to a new image dataset or a new domain. Alternatively, self-managed pre-training can be combined with preconditions and model update strategies in one or all of the multiple model development iterations.

[0066] (6) The adaptive learning framework is domain-independent and can be applied to other imaging modalities as well as other computational studies such as multimodal analysis, gene sequencing signal analysis, and survival modeling.

[0067] Figure 3 shows a comparison of the resource requirements for training from scratch, transfer learning, and the adaptive learning framework described herein. The resource (y-axis) refers to the total number of images that need to be computed for model training of various subsequent model versions (x-axis). If one image is processed N times, it is counted as N images. Training from scratch consumes most resources and the demand increases rapidly. Transfer learning only requires resources to compute a new batch of data. Adaptive learning requires similar computational resources as transfer learning and its increase is negligible.

[0068] In one exemplary embodiment, a computer-implemented process is provided that includes obtaining, in a data processing system, a first annotated training set of images for training a machine learning algorithm to detect, characterize, classify, or combinations thereof, some or all regions or objects within an image, where the first annotated training set of images is within a first image domain; splitting, by the data processing system, the first annotated training set of images into mini-sets of images, where each mini-set represents a distinct modeling sub-task and includes a limited number of examples; training, in a first stage, the machine learning algorithm using the mini-sets of images to generate a pre-conditioned machine learning model configured to detect, characterize, classify, or combinations thereof, some or all regions or objects within a new image; labeling, by the data processing system, a limited number of images from a target data set to generate a second annotated training set of images for training the machine learning algorithm to detect, characterize, classify, or combinations thereof, some or all regions or objects within an image, where the second annotated training set of images is within a second image domain; and training, in a second stage, the pre-conditioned machine learning model using the second annotated training set of images to generate a target machine learning model configured to detect, characterize, classify, or combinations thereof, some or all regions or objects within a new image, where the number of classes targeted in the first stage is a subset of or matches the number of classes targeted in the second stage.

[0069] In some embodiments, the computer-implemented process further includes identifying a digital pathology scenario, selecting an adaptive continuous learning method for updating a target machine learning model based on the digital pathology scenario, and updating the target machine learning model based on the adaptive continuous learning method to generate an updated machine learning model.

[0070] Preferably, the various techniques described herein can improve the robustness of the machine learning model (e.g., improve the accuracy of cell classification).

[0071] II. Definitions As used herein, when an act is “based on” something, this means that the act is at least partially based on at least a portion of that something.

[0072] As used herein, the terms “substantially,” “approximately,” and “about” are defined as being mostly what is specified, but not necessarily completely what is specified (and including what is completely specified), as would be understood by one of ordinary skill in the art. In any of the disclosed embodiments, the term “substantially,” “approximately,” or “about” may be replaced with “within [percentage]” of what is specified, where the percentage includes 0.1, 1, 5, and 10%.

[0073] As used herein, the terms "sample", "biological sample", "tissue", or "tissue sample" refer to any sample containing biomolecules (e.g., proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof) obtained from any organism, including viruses. Other examples of organisms include mammals (e.g., veterinary animals such as humans, cats, dogs, horses, cows, and pigs, as well as laboratory animals such as mice, rats, and primates), insects, annelids, spiders, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections and needle biopsies of tissue), cell samples (such as cytological smear specimens like Pap smears or blood smears, or samples of cells obtained by microdissection), or cell fractions, fragments, or organelles (such as those obtained by lysing cells and separating their components by centrifugation, etc.). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucus, tears, sweat, pus, biopsy tissue (e.g., obtained by surgical biopsy or needle biopsy), nipple aspirate, earwax, milk, vaginal fluid, saliva, swabs (such as oral swabs), or any material containing biomolecules derived from an initial biological sample. In certain embodiments, the term "biological sample" as used herein refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or a part thereof obtained from a subject.

[0074] As used herein, the terms "biological material", "biological structure", or "cellular structure" refer to natural materials or structures that include all or part of a biological structure (e.g., cell nucleus, cell membrane, cytoplasm, chromosomes, DNA, cells, cell aggregates, etc.).

[0075] As used herein, the term "digital pathology image" refers to a digital image of a stained sample.

[0076] As used herein, the term "cell detection" refers to the detection of the location and characteristics of cell or cellular structure (e.g., cell nucleus, cell membrane, cytoplasm, chromosomes, DNA, cells, cell aggregates, etc.) pixels.

[0077] As used herein, the term "target region" refers to a region of an image that includes the image data intended to be evaluated in an image analysis process. The target region includes any region, such as a tissue region of an image, that is intended to be analyzed in an image analysis process (e.g., tumor cells or staining expression).

[0078] As used herein, the term "tile" or "tile image" refers to a single image corresponding to a part of the whole image or the whole slide. In some embodiments, a "tile" or "tile image" refers to a region of a whole slide scan or an area of interest having (x,y) pixel dimensions (e.g., 1000 pixels × 1000 pixels). For example, considering an entire image divided into M columns of tiles and N rows of tiles, each tile within the M×N mosaic includes a part of the whole image; that is, the tile at position M1,N1 includes the first part of the image, the tile at position M1,N2 includes the second part of the image, and the first part is different from the second part. In some embodiments, the tiles can each have the same dimensions (pixel size × pixel size). In some examples, the tiles can partially overlap and can represent overlapping regions of the whole slide scan or the area of interest.

[0079] As used herein, the term "patch", "image patch", or "mask patch" refers to a container of pixels corresponding to a part of the whole image, whole slide, or whole mask. In some embodiments, a "patch", "image patch", or "mask patch" refers to a region of an image or mask or an area of interest having (x,y) pixel dimensions (e.g., 256 pixels × 256 pixels). For example, a 1000 pixel × 1000 pixel image divided into 100 pixel × 100 pixel patches will contain 10 patches (each patch containing 1000 pixels). In other embodiments, a patch has (x,y) pixel dimensions and each "patch", "image patch", or "mask patch" overlaps with one or more pixels shared with another "patch", "image patch", or "mask patch".

[0080] III. Generation of Digital Pathology Images Digital pathology involves the interpretation of digitized images in order to accurately diagnose a subject and guide treatment decisions. In digital pathology decision-making, an image analysis workflow can be established to automatically detect or classify biological objects of interest, such as positive and negative tumor cells. An exemplary digital pathology decision-making workflow includes obtaining a tissue slide and scanning a preselected area or the entire tissue slide with a digital image scanner (e.g., a whole slide image (WSI) scanner) to obtain a digital image, performing image analysis on the digital image using one or more image analysis algorithms, and potentially detecting and quantifying each object of interest based on the image analysis (e.g., quantitative or semi-quantitative scoring such as positive, negative, moderate, weak, etc.) (e.g., counting or identifying the object-specific area or cumulative area of each object of interest).

[0081] Figure 4 shows an exemplary network 400 for generating digital pathology images. A fixation / embedding system 405 uses a fixative (e.g., a liquid fixative such as a formaldehyde solution) and / or an embedding material (e.g., a histological wax such as paraffin wax, and / or one or more resins such as styrene or polyethylene) to fix and / or embed a tissue sample (e.g., a sample containing at least a portion of at least one tumor). Each sample can be fixed by exposing the sample to the fixative for a predetermined period (e.g., at least 3 hours) and then dehydrating the sample (e.g., via exposure to an ethanol solution and / or a clearing intermediate agent). The embedding material can infiltrate when the sample is in a liquid state (e.g., during heating).

[0082] The fixation and / or embedding of samples is used to preserve the samples and delay the decomposition of the samples. In histology, fixation generally refers to an irreversible process that uses chemicals to retain the chemical composition, preserve the natural sample structure, and maintain the cell structure from decomposition. Fixation may also harden the cells or tissues for sectioning. Fixatives may enhance the preservation of samples and cells by using cross-linked proteins. Fixatives may bind to and cross-link some proteins and denature other proteins by dehydration, which can inactivate enzymes that would otherwise decompose the tissue and may harden the tissue. Fixatives may also kill bacteria.

[0083] Fixatives can be administered, for example, by perfusion and immersion of the prepared samples. A variety of fixatives can be used, including methanol, Bouin's fixative, and / or formaldehyde fixatives, such as neutral buffered formalin (NBF) or paraffin-formaldehyde (paraformaldehyde-PFA). If the sample is a liquid sample (e.g., a blood sample), the sample may be smeared onto a slide and dried before fixation. The fixation process can help preserve the structure of the samples and cells for the purposes of histological testing, but fixation can mask tissue antigens, thereby reducing antigen detection. Therefore, fixation is generally considered a limiting factor in immunohistochemistry because formaldehyde can cross-link antigens and mask epitopes. In some cases, additional processes are performed to reverse the effects of cross-linking, including treating the fixed sample with anhydrous citraconic acid (a reversible protein cross-linking agent) and heating.

[0084] Embedding can involve infiltrating a sample (e.g., a fixed tissue sample) with a suitable histological wax such as paraffin wax. The histological wax can be insoluble in water or alcohol, but may be soluble in a paraffin solvent such as xylene. Thus, it may be necessary to replace the water in the tissue with xylene. To do so, the sample may first be dehydrated by gradually replacing the water in the sample with alcohol, which can be achieved by passing the tissue through ethyl alcohol of increasing concentration (e.g., 0 to about 100%). After replacing the water with alcohol, the alcohol can be replaced with xylene, which is miscible with alcohol. Since the histological wax can be soluble in xylene, the melted wax can fill the space that was previously filled with water and is filled with xylene. A block formed by cooling and hardening the sample filled with wax can be clamped to a microtome, vibratome, or compressotome to cut sections. In some cases, deviating from the above exemplary procedure may result in infiltration of paraffin wax and may inhibit the penetration of antibodies, chemicals, or other fixatives.

[0085] Next, a tissue slicer 410 can be used to slice the fixed and / or embedded tissue sample (e.g., a tumor sample). Slicing is the process of cutting thin slices of the sample (e.g., 4-5 μm thick) from the tissue block and mounting them on a microscope slide for examination. Slicing may be performed using a microtome, vibratome, or compressotome. In some cases, the tissue can be rapidly frozen in dry ice or isopentane and then cut with a cold knife in a refrigerated cabinet (e.g., a cryostat). Other types of coolants such as liquid nitrogen can be used to freeze the tissue. Sections for use with brightfield and fluorescence microscopes are generally about 4-10 pm thick. In some cases, the sections can be embedded in epoxy resin or acrylic resin, which may enable thinner sections (e.g., <2 μm) to be cut. These sections may then be mounted on one or more glass slides. A coverslip may be placed on top to protect the sample sections.

[0086] Since tissue sections and the cells within them are substantially transparent, slide preparation typically further includes staining the tissue sections (e.g., automated staining) to make the relevant structures more visible. In some cases, the staining is performed manually. In some cases, the staining is performed semi-automatically or automatically using a staining system 415. The staining process includes exposing the tissue sample or sections of the fixed liquid sample to one or more different stains (e.g., sequentially or simultaneously) to manifest different characteristics of the tissue.

[0087] For example, staining can be used to mark specific types of cells and / or to flag specific types of nucleic acids and / or proteins to assist microscopy. The staining process generally involves adding a dye or stain to a sample to confirm or quantify the presence of a specific compound, structure, molecule, or feature (e.g., an intracellular feature). For example, staining can help identify or highlight specific biomarkers from tissue sections. In other examples, staining can be used to identify or highlight cell organelles within biological tissues (e.g., muscle fibers or connective tissue), cell populations (e.g., different blood cells), or individual cells.

[0088] One exemplary type of tissue staining is histochemical staining, which uses one or more chemical dyes (e.g., acidic dyes, basic dyes, chromogens) to stain tissue structures. Histochemical staining may be used to show general aspects of tissue morphology and / or cell microanatomy (e.g., distinguish cell nuclei from cytoplasm, show lipid droplets, etc.). An example of histochemical staining is H&E. Other examples of histochemical staining solutions include trichrome staining solutions (e.g., Masson's trichrome), periodic acid Schiff (PAS), silver staining solutions, and iron staining solutions. The molecular weight of histochemical staining reagents (e.g., dyes) is generally 500 kilodaltons (kD) or less, although some histochemical staining reagents (e.g., alcian blue, phosphomolybdic acid (PMA)) may have a molecular weight of up to 2000 or 3000 kD. An example of a high molecular weight histochemical staining reagent is α-amylase (about 55 kD), which may be used to show glycogen.

[0089] Another type of tissue staining is IHC, also called "immunostaining", which uses a primary antibody that specifically binds to a target antigen of interest (also called a biomarker). IHC can be either direct or indirect. In direct IHC, the primary antibody is directly conjugated to a label (e.g., a chromophore or a fluorophore). In indirect IHC, first the primary antibody is bound to the target antigen, and then a secondary antibody conjugated to a label (e.g., a chromophore or a fluorophore) is bound to the primary antibody. Since antibodies have a molecular weight of about 150 kD or more, the molecular weight of IHC reagents is much larger than that of histochemical staining reagents.

[0090] For performing staining, various types of staining protocols may be used. For example, an exemplary IHC staining protocol involves using a hydrophobic barrier line around the sample (e.g., tissue section) to prevent leakage of reagents from the slide during incubation, treating the tissue section with reagents to block endogenous sources of non-specific staining (e.g., enzymes, free aldehyde groups, immunoglobulins, other unrelated molecules that can mimic specific staining), incubating the sample with a permeabilization buffer to facilitate penetration of antibodies and other staining reagents into the tissue, incubating the tissue section with a primary antibody at a specific temperature (e.g., room temperature, 6 - 8 °C) for a certain period (e.g., 1 - 24 hours), rinsing the sample using a wash buffer, then incubating the sample (tissue section) with a secondary antibody at another specific temperature (e.g., room temperature) for another period, rinsing the sample again with a water buffer, incubating the rinsed sample with a chromogen (e.g., DAB: 3,3’-diaminobenzidine), and washing away the chromogen to stop the reaction. In some examples, subsequently, counterstaining is used to identify the overall “landscape” of the sample and serves as a reference for the main color used for detecting the tissue target. Counterstaining agents can include, for example, hematoxylin (a blue to purple stain), methylene blue (a blue stain), toluidine blue (a stain that stains nuclei dark blue and polysaccharides pink to red), nuclear fast red (also called Kernechtrot dye, a red stain), methyl green (a green stain), non-nuclear chromogenic stains, such as eosin (a pink stain), and the like. As will be understood by those skilled in the art, other immunohistochemical staining techniques can be implemented to perform the staining.

[0091] In another example, an H&E staining protocol can be implemented for tissue section staining. The H&E staining protocol involves applying a hematoxylin stain mixed with a metal salt or mordant to the sample. The sample can then be rinsed with a weak acid solution to remove excess staining (differentiation), and subsequently blued in a weak alkaline water. After the application of hematoxylin, the sample can be counterstained with eosin. It will be understood that other H&E staining techniques can be implemented.

[0092] In some embodiments, depending on which features of interest are targeted, various types of staining agents can be used to perform staining. For example, DAB can be used for various tissue sections for IHC staining, and DAB provides a brown color that represents the features of interest in the stained image. In another example, since the DAB color may be masked by melanin pigment, alkaline phosphatase (AP) can be used for skin tissue sections for IHC staining. Regarding primary staining techniques, applicable staining agents can include, for example, basophilic and eosinophilic staining, hematoxylin and hematein, silver nitrate, trichrome staining agents, etc. Acidic dyes can react with cationic or basic components in tissues or cells, such as proteins and other components in the cytoplasm. Basic dyes can react with anionic or acidic components in tissues or cells, such as nucleic acids. As described above, an example of a staining system is H&E. Eosin may be a negatively charged pink acidic dye, and hematoxylin may be a purple or blue basic dye containing hematein and aluminum ions. Other examples of staining can include periodic acid-Schiff reaction (PAS) staining, Masson's trichrome, alcian blue, Fontana-Masson, reticulin staining, etc. In some embodiments, different types of staining agents can be used in combination.

[0093] Next, the sections may be attached to corresponding slides, and then the imaging system 420 can scan or image to generate raw digital pathology images 425a - n. To magnify the stained sample, a microscope (e.g., an electron microscope or an optical microscope) can be used. For example, an optical microscope can have a resolution of less than 1 μm, such as about several hundred nanometers. An electron microscope can be used to observe finer details in the nanometer or sub - nanometer range. The imaging device (combined with or separated from the microscope) images the magnified biological sample to acquire image data such as a multi - channel image (e.g., multi - channel fluorescence) having several (e.g., 10 - 16, etc.) channels. The imaging device can include, but is not limited to, a camera (e.g., an analog camera, a digital camera, etc.), optical elements (e.g., one or more lenses, a sensor focus lens group, a microscope objective lens, etc.), an imaging sensor (e.g., a charge - coupled device (CCD), a complementary metal - oxide - semiconductor (CMOS) image sensor, etc.), photographic film, etc. In a digital embodiment, the imaging device can include a plurality of lenses that cooperate to demonstrate on - the - fly focusing. An image sensor, such as a CCD sensor, can image a digital image of a biological sample. In some embodiments, the imaging device is a bright - field imaging system, a multi - spectral imaging (MSI) system, or a fluorescence microscope system. The imaging device can utilize invisible electromagnetic radiation (e.g., UV light) or other imaging techniques to capture an image. For example, the imaging device may comprise a microscope and a camera configured to capture an image magnified by the microscope. The image data received by the analysis system may be the same as the raw image data captured by the imaging device and / or may be derived from the raw image data.

[0094] Next, an image of the stained section can be stored in a storage device 425, such as a server. The image can be stored locally, remotely, and / or in a cloud server. Each image can be stored in association with an identifier of the subject and a date (e.g., the date the sample was collected and / or the date the image was captured). The image can further be transmitted to another system (e.g., a system associated with a pathologist, an automated or semi-automated image analysis system, or a machine learning training and deployment system, as described in more detail herein).

[0095] It is understood that modifications to the processes described with respect to network 400 are contemplated. For example, if the sample is a liquid sample, embedding and / or sectioning may be omitted from the process.

[0096] IV. Adaptive Learning Framework FIG. 5 shows that the adaptive learning framework includes two components: (505) development of an initial model, and (510) model update and / or adaptation to different image domains. The two components 505; 510 can be applied separately or all together, such as for initial model development and subsequent model update and / or adaptation to different image domains.

[0097] Development of the Initial Model For efficient development of the initial model, the two-stage development strategy (shown in FIG. 5) is as follows. Stage 1 Model preconditioning, where an artificial intelligence system (e.g., the artificial intelligence system described in detail with respect to FIG. 10) utilizes existing annotated datasets and improves learning skills through training using these datasets. Stage 2 Target model training, where the artificial intelligence system utilizes the learning skills learned in Stage 1 to expand itself to a different image domain (the target domain) that requires fewer annotations in the target domain than conventional learning methods.

[0098] The aforementioned "learning skills" include, among others, one or more of the following, namely, the best set of model initializations, the best set of model weights that can generalize to unseen data, the best set of model architectures, etc. The best set is determined using one or more metrics for measuring model performance, such as accuracy or area under the curve (AUC). The learning skills are then applied to the target image domain, such that fewer annotations are required for the target domain than with conventional machine learning.

[0099] To achieve model preconditioning, a meta-learning strategy may be adopted, where an existing dataset (e.g., a training dataset) is split into mini-sets, each mini-set representing a distinct modeling subtask, containing only a few examples, and thus forming a large number of subtasks. The splitting can be performed during training time, and different data splits can be executed in each model training iteration. By training an artificial intelligence system using these subtasks, the artificial intelligence system can search for excellent solutions in the model optimization landscape, which can generalize to any related small subtasks without overfitting to a specific subtask, and thus precondition the artificial intelligence system for stage 2 for unseen target domains. Generally, the number of classes in stage 1 training is set to match a part or the number of classes in stage 2 training. In the stage 2 class incremental scenario, the number of classes in stage 1 can be increased, and in the stage 2 domain and data incremental scenario, the number of classes from stage 1 needs to be made to match. In a specific example, the number of classes in stage 2 can be more than that in stage 1 (explained in detail in relation to Figure 9). In classification tasks (e.g., binary and multi-class classification), the class labels are at the image level, for example, the class (or classes) is set for the entire image. In prediction tasks such as image segmentation and object detection (e.g., dense prediction), the class labels are at a finer level. For example, in the case of object detection, each distinct object within the image has a label for its class and location. The training datasets used for model preconditioning are related to each other but can have various degrees of similarity to those of the target domain.

[0100] To sample an existing annotated dataset, the following criteria may be implemented. (1) If only one annotated dataset is available, for each subtask, a subset of classes may be randomly selected to be a part of or match the number of classes in stage 2, and several examples may be randomly selected from the selected classes. In this scenario, there are multiple subtasks with different or partially different classes. (2) If multiple annotated datasets are available, for each subtask, examples can be mixed from multiple datasets, and then the mixed dataset can be implemented in the same way as in (1), or a specific dataset can be randomly selected first, and then a subset of classes within that dataset can be randomly selected to be a part of or match the number of classes in stage 2. In both scenarios, (a) sampling strategies other than complete randomness can also be adopted, for example, some classes or some datasets can be sampled more frequently than others, and / or (b) the examples within each subtask may be divided into a training subset and a validation subset.

[0101] More specifically, in the case of image-level prediction, for example in an image classification task, an "example" refers to an image having that class label within the dataset, i.e., each subtask consists of a set of images all belonging to the selected class. Within each subtask, the image class labels are redefined for the stage 1 training process. For example, 5 images per class, e.g., 15 images in total for 3 classes, may be selected, and regardless of the original class labels of each image, for stage 1 training, a random ordering of the classes is generated, whereby any of the 3 classes can be set as class No.0, another class can be set as class No.1, and the last class can be set as class No.2. In the next subtask, another 15 images can be selected from a different set of classes, and then their class labels are reset in a random order to again be 0, 1, and 2. In this way, the model becomes class-independent in the sense that the model is not focused on learning information from each specific class, but rather learns a way to improve the learning skills for all possible subtasks it encounters.

[0102] In the case of a prediction task (e.g., dense prediction), an example includes all annotated entities (regions or objects) of the same class within the image along with their labels, and when sampling several examples from a selected subset of classes, several images containing at least one labeled entity from the selected class are first selected, and then the labeled entities that are not from the selected class are set as the background class. For example, in an image segmentation task, the class labels may include a foreground class and a background class (i.e., regions of no interest for modeling purposes), and if the selected classes are tumor foci and blood vessels, all regions within the image that do not belong to these two classes are relabeled as the background class, all regions belonging to the tumor foci are randomly relabeled as class No.0 or No.1, and the blood vessel regions are relabeled as whatever class index remains after relabeling the tumor regions.

[0103] The learning method of the subtask can be considered as follows. A metric-based method that learns the representation of existing datasets and distills the skill of comparing the similarity between examples of any class from an unseen target domain from these datasets. In these methods, the representation learned in stage 1 can be applied to stage 2 in the following way. That is, (a) select a small number of images in the target domain, (b) annotate these images and apply a conditional model for each example to generate a feature vector representation (e.g., the output of the last layer in a convolutional network before the classification layer), (c) combine the representations from examples of the same class to generate one representation for each target class, (d) use these processed representations as prototypes of the target classes, (e) generate the feature vector representations of the remaining images or image regions of the unlabeled target domain (or other entities in the case of a high-density prediction task), (f) compare each feature vector representation from the unlabeled target domain with the prototypes by calculating the distance between these vectors, e.g., cosine distance, and then assign the class label in the unlabeled image or image region as the prototype class with the minimum distance (e.g., the most similar). Other techniques for learning the subtask may be used in combination with the metric-based method. For example, an adversarial generation model may be used, and the adversarial generation model synthesizes images based on the distribution of existing images to increase the number of examples for each class in the target image domain.

[0104] Alternatively, optimization-based methods for learning the best model weights to initialize the model can efficiently adapt to an unseen target dataset with only a small number of examples. In these methods, the training in stage 1 involves two model optimization loops, namely, (a) one inner training loop where the artificial intelligence model updates its model weights on one subtask for a predetermined or flexible number of epochs to generate the loss in its validation set as shown as L-subtask-i for the i-th subtask, and (b) one outer training loop where the goal is to search for a set of model initializations that generate the best model when used to update all subtasks, with each subtask accompanied by only a small number of annotated examples, which is achieved by finding a model initialization that minimizes the sum of all losses (summing L-subtask-i where i ranges from 1 to the number of subtasks) calculated from the validation sets of the subtasks on their validation sets for those model initializations. Other techniques for learning subtasks may be used in combination with the optimization-based methods. For example, not only search for the best model initialization but also search for the best model architecture (e.g., perform neural architecture search).

[0105] The workflows described herein for the design of initial model development were applied to cell detection in brightfield IHC assays and tissue type classification in H&E assays as follows. However, similar workflows can also be applied to other staining methods, such as special staining in brightfield assays (e.g., Trichrome Masson assay that stains muscle, collagen fibers, red blood cells, and cell nuclei simultaneously) and fluorescence IHC assays.

[0106] (i) Model preconditioning for cell detection in IHC assays: The aim is to identify the staining phenotype, cell type, and cell location in each image. For example, in the case of a DAB-Ki67 IHC assay, tumor cells stained positively with Ki67 (Ki67+ tumors), tumor cells stained negatively with Ki67 (Ki67- tumors), and all other cell types / staining types can be identified together with the position of the center of each cell nucleus in a single pixel, or together with the bounding box of each cell nucleus (e.g., the rectangular pixel positions circumscribing each cell nucleus), or together with the pixels of each cell nucleus (e.g., a nuclear segmentation mask). Thus, a cell detection model can be designed.

[0107] Generally, images from brightfield IHC assays are related and have a certain level of similarity in their appearance. That is, hematoxylin staining is used in most of these assays to stain cell nuclei and serves as a pointer as to where cells are within the whole slide image. However, one or more biomarkers are targeted by the IHC staining protocol, and correspondingly, the cells expressing these biomarkers show color upon application of the chromogen. A "staining pattern" refers to the appearance of image regions positive for a target biomarker with respect to (a) intracellular and / or subcellular localization of the biomarker, (b) stained cell type, (c) staining intensity, (d) incidence frequency, and (e) spatial distribution of the positive regions. For example, Ki67 has a nuclear staining pattern. For instance, the positive staining signal in an IHC image from a Ki67 assay is observed to be located in cell nuclei, mainly in either scattered tumor cells or clustered tumor nests, and the positive staining signal ranges from low staining intensity to very high intensity.

[0108] Workflows for preconditioning may be designed using images from various IHC assays having various chromogens and biomarker staining patterns (see Table 1 for an example assay). Exemplary workflows may include the following. (1) When one or more chromogens used in an existing annotated dataset are the same as those of one or more target IHC assays: (1.1) If annotations from only one assay are available, in each model training iteration, split this dataset into subtasks, sample one subtask for training, which has a small number of images with cell annotations from one or more classes, and ensure that cells of all classes selected for this development stage are present in at least one of the images. For example, apply to the preconditions of the DAB-Ki67 assay and the DAB-PDL1 assay. Figure 6 shows an example of a DAB-Ki67 IHC image (A), where the brown signal is the image area where the chromogen DAB produced color, the Ki67 expressed in the cell nuclei indicates these cells detected by this IHC assay, and the grayish blue signal is the hematoxylin-stained cell nuclei. An exemplary ground truth (B) is also shown for cell detection in the image of (A), and the differently colored dots overlaid on the cell nuclei centers indicate the class labels of each cell. (1.2) If annotations from multiple assays are available, in each model training iteration, an assay may first be selected from all assays to train an artificial intelligence system preconditioned to be as assay-independent as possible, and then several examples may be sampled from this assay. For example, the artificial intelligence system may be preconditioned for the DAB-Ki67 and DAB-CK7 assays and applied to the DAB-PDL1 assay.

[0109] (2) When the chromogens used in the existing annotated dataset are different from those of the target IHC assay: A workflow similar to (1.1) or (1.2) can be executed. Furthermore, to match the chromogens in both stage 1 and stage 2, dye unmixing can be performed to decompose the IHC image into each dye component, and these components can be remixed with the color vectors extracted from the target IHC assay. For example, an artificial intelligence system may be a prerequisite for the Tamra-PDL1 / Dabsyl-CK7 duplex assay and may also be applied to the DAB-Ki67 assay. The training system or user can choose to decompose the Tamra-PDL1 / Dabsyl-CK7 duplex image into PDL1, CK7, and hematoxylin grayscale intensity images and remix them with the color vectors of DAB to obtain two sets of images for the preconditioning stage, namely, the synthetic DAB-PDL1 single-chain image and the synthetic DAB-CK7 single-chain image. Figure 6 shows an exemplary Tamra-PDL1-Dabsyl-CK7 IHC image (C). An exemplary ground truth (D) is also shown for the DAB-PR IHC image where dots of different colors overlaid at the center of the cell nucleus indicate the class label of each cell.

[0110] The assay used in the preconditioning stage does not necessarily have to have a high similarity to the target domain assay, but if there is some level of similarity, it can be beneficial for the accumulation of both the learned skills and the learned knowledge in stage 1.

Table 1

[0111] (ii) Model preconditioning for tissue type classification in the H&E assay: In this case, the goal is to generate a model for identifying the tissue type of each image tile from all slide images. For example, a model can be constructed that classifies each image tile into tumor, stroma, normal tissue, and other types. Model preconditioning can be performed using the same workflow described herein with existing datasets from other disease types (e.g., different tumor types), different disease stages, etc., outside the target domain. Each subtask can be sampled from the same dataset or, if available, a mixture of different datasets.

[0112] Model Update and Adaptation To efficiently perform model update and adaptation to new datasets after initial model development, techniques have been developed to identify scenarios commonly encountered in digital pathology settings and select the adaptation learning method optimal for the scenario. The tissue type classification of H&E images is used as an example to illustrate these techniques. However, it should be understood that the techniques described for model update and adaptation can be applied to various other scenarios commonly encountered in digital pathology.

[0113] When the model is continuously trained, the data stream can vary in different ways. The situations commonly encountered in digital pathology are classified into the following scenarios. 1. Data incremental scenario: Creating an annotated dataset curated by a pathologist is a time-consuming process, and it is preferable to train the model in batches when data arrives and when it arrives. Incoming data considered a new stream for training the model has a minimal difference from the previous data and is usually from the same underlying distribution and has all classes as seen previously. 2. Domain Incremental: When the continuous stream of data is from different domains or distributions, it is considered a domain incremental scenario. This scenario is similar to the data incremental scenario in that all data streams have all the classes they represent. 3. Class Incremental: A class incremental scenario is when the model is extended with respect to the number of classes. The clinical need for this scenario arises when a batch of annotated images contains different subsets of classes or when the model definition changes to classify data into more tissue subtypes. It should be understood that each data stream may include images from new unseen classes as well as images of the seen classes introduced in previous data streams. 4. Task Incremental: When each data stream is defined as a new task, any of the above scenarios can be considered task incremental. This scenario uses different architecture designs and the network has shared layers between tasks and task-specific layers.

[0114] Each incremental data stream is called an experience. Each experience is then split into training, validation, and test streams. The model is targeted in the training stream, validated in the validation stream at the end of all epochs, and evaluated in the test stream at the end of each experience. Model performance is evaluated in the test stream from all experiences at the end of training for all experiences to study forward and backward movement.

[0115] (i) Data Incremental Digital Pathology Model Update: (1.1) Digital Pathology Scenario: After the initial development of an artificial intelligence-based digital pathology algorithm, it may be necessary to incorporate more annotated data (e.g., with slight differences and mostly similar; from the same underlying data distribution) of the same domain into the model. In a digital pathology setting, it is common to expect annotated data to arrive in batches because expert-annotated and / or expert-verified annotations can be time-consuming and may not be available at the time of initial model development. Additionally, due to the large number of parameters in artificial intelligence models, especially deep neural networks, it is common to prevent overfitting of the model by learning from additional data. For example, an initial model may be trained to classify tumor tissue versus normal tissue from colon cancer images (shown in Figure 7 - the plot in the lower left corner shows the class labels and the number of examples for each model version), and in subsequent model development / updating, similar examples of the same class may be integrated into the model, and such examples may arrive in batches over time.

[0116] (1.2) Workflow for selecting the optimal learning method: (1.2a) To efficiently and effectively update the model without training from scratch using existing data and newly arrived data, a benchmark dataset is set from a given digital pathology domain, the dataset is randomly split into batches, and the case of sequentially arriving data is simulated. (1.2b) To evaluate how effective the candidate learning methods are in the presence of small random variations in the individual examples within each batch, an extended dataset may be created to simulate such small variations, and the image order may be randomly shuffled so that each data batch contains images extended in different ways. The various types and degrees of variations to the original data may be carefully balanced so that the distribution of the entire dataset is not distorted. The data augmentation can be performed dynamically during training or pre-executed for each iteration of training. Exemplary augmentation techniques include (i) changing the color space by first performing stain decomposition (e.g., unmixing), and then remixing each stain with a predetermined color vector to change the hue, saturation, and intensity of each stain, (ii) increasing the image resolution by resizing the image and then returning it to the original size, (iii) the same, or (iv) any combination thereof.

[0117] (1.3) Artificial intelligence-based method for adaptive learning: The batches of benchmark data generated from (1.2b) can be executed against candidate learning methods as follows.

[0118] (1.3a) Normalization-based methods: These methods mainly focus on privacy prioritization and memory reduction. Privacy is maintained by avoiding saving raw inputs. Among different normalization-based methods, two common methods are the Elastic Weight Consolidation (EWC) method and the Learning Without Forgetting (LWF) method. EWC is a previously noted method where model parameters are used a priori when learning from new data, and the method estimates the distribution over the model parameters. On the other hand, LWF is a data-focused method. The main design of the data-focused method is based on knowledge distillation from the previous model to the current model targeted with new data. This concept is also introduced in LWF which addresses forgetting with a knowledge transfer mechanism. These methods prevent the model from forgetting the knowledge it has learned by applying constraints to the current model weights, and as a result, when updating the current model, the model weights do not deviate significantly from their current version.

[0119] (1.3b) Replay-based methods: These methods retain the most beneficial examples or their feature representations (e.g., feature vectors extracted from the hidden layers of a neural network) from existing data batches and revisit them during the training of newly arriving data. To overcome forgetting, in the process of learning a new task, previous samples of the task are replays. Among the various methods belonging to this category are Incremental Classifier and Representation Learning (iCaRL), Continuous Prototype Evaluation (CoPE), and A-GEM. The iCaRL method stores a subset of the best exemplars for each class selected according to the approximate class means in the learned feature space. The upper limit of this method is determined by the co-training of past and current tasks. CoPE is an online data incremental learner that has prototypes that permanently represent the most prominent features of a class population. The rapidly evolving prototypes enable learning and prediction at any given time. CoPE is robust against class imbalance using the replay and balanced memory population approach. GEM is designed based on the task incremental setting. This method focuses only on new tasks by restricting their updates and thus does not interfere with previous tasks. This is achieved by projecting the gradient direction calculated onto the feasible region outlined by the previous task gradient by means of a first-order Taylor series approximation. The A-GEM method is an improved version of the GEM method. A-GEM helps alleviate the problem of projecting in one direction estimated by samples randomly selected from the previous task data buffer. A-GEM provides similar performance accuracy to GEM with similar computational and memory efficiency as the regularization method.

[0120] (1,3c) Methods combining regularization and replay

[0121] (1.3d) Methods that utilize meta-learning principles to enable model updates with a few examples in each batch (see the "CoPE" method implemented in the experimental section)

[0122] (1.3e) Parameter Isolation Methods: These methods assign different model parameters to each task to address forgetting. Since there are no constraints on the size of the architecture, these methods do not have a fixed architecture. Therefore, new branches may grow for new tasks. This can be achieved by freezing the previous task parameters or making model copies dedicated to each task. Architectures of these types are called dynamic architectures.

[0123] (1.3f) Other Techniques Applied with the Aforementioned Methods. For example, before applying the aforementioned method, perform self - supervised learning on new unlabeled data. As another example, run a adversarial generative model to generate examples similar to these from previous data batches.

[0124] (ii) Update of the Domain - Incremental Digital Pathology Model: (2.1) Digital Pathology Scenario: After the initial development of an artificial - intelligence - based digital pathology algorithm, it may be necessary to incorporate annotated data from different domains (e.g., those from similar but different underlying data distributions) into the model. In a digital pathology setting, it is a practical necessity to be able to generate a model that is highly robust to (a) changes in experimental settings, such as changes or differences in (i) staining reagents (ii) staining protocols or equipment, (iii) scanner suppliers, (iv) sample sources (e.g., different clinical sites or tissue banks), etc., and (b) changes in other aspects such as tissue type, organ type, patient population, disease stage, disease subtype (e.g., tumor type and subtype).

[0125] (2.2) Workflow for selecting the optimal adaptation learning method: (2.2a) To efficiently and effectively update the model without training from scratch using both existing data and newly arrived data, set a benchmark dataset from the DP domain and generate synthetic images simulating realistic changes as data from different sources / domains. (2.2b) To evaluate which candidate adaptation learning method is most effective for this scenario, an extended dataset may be created to simulate multiple types of domain shifts, and how well the candidate method functions with such changes may be evaluated. Data augmentation can be performed dynamically during training or pre-executed for each iteration of training. Exemplary augmentation techniques include (i) changing the color space by first performing color decomposition (e.g., unmixing), then remixing each color with a predetermined color vector to change the hue, saturation, and intensity of each color, (ii) increasing the image resolution by resizing the image and then returning it to the original size, (iii), etc., or (iv) any combination thereof. One or more types of changes can be applied to each subset to simulate a series of consecutive domain shifts. For example, for H&E images having the same tissue type for a classification task, the dataset may be randomly split and the following augmented subsets may be generated. That is, (i) due to a change in protocol, one of the stains becomes stronger, (ii) due to the aging of the prepared slides, both stains fade (i.e., a decrease in staining intensity), (iii) due to a change in the stain, one stain is more saturated than the other in the HSV color space, (iv) due to a change in the scanner, both stains change hue in the HSV color space. Figure 8 shows the changes described in (i) and (ii) above (the plot in the lower left corner shows the class labels and the number of examples in each model version).

[0126] (2.3) Artificial intelligence-based methods for adaptive learning: The batch of benchmark data generated from (2.2b) may be executed against candidate learning methods as follows. (2.3a) Normalization-based methods: These methods prevent the model from forgetting the knowledge learned by applying constraints to the current model weights, so that when updating the current model, the model weights do not deviate significantly from their current version. (2.3b) Replay-based methods: These methods retain the most beneficial examples or their feature representations (e.g., feature vectors extracted from the hidden layers of a neural network) from existing data batches and revisit them during the training of newly arrived data. (2.3c) Methods that combine regularization and replay. (2.3d) Methods that utilize meta-learning principles to enable model updates with a few examples of each batch (see the "CoPE" method implemented in the experimental section). (2.3e) Parameter isolation methods. (2.3f) Other techniques applied together with the aforementioned methods. For example, perform self-supervised learning on newly unlabeled data before applying the aforementioned methods. As another example, execute a adversarial generative model to generate examples similar to these from previous data batches.

[0127] (iii) Class-incremental digital pathology model updates: (3.1) Digital pathology scenario: After the initial development of an artificial intelligence-based digital pathology algorithm, it may be necessary to expand the model with respect to the number of classes. In a digital pathology setting, this scenario is commonly encountered for reasons such as (i) changes in end-user needs, (ii) annotated data is provided in batches, and each batch has a set of classes that are partially or completely different from the classes of the existing data, and / or (iii) it has been found that for model design options, e.g., a model initially designed to classify tumor vs. normal, it is necessary to identify necrosis and lymphocyte clusters to ensure good model performance.

[0128] (3.2) Workflow for selecting the optimal learning method: (3.2a) To efficiently and effectively update the model without training from scratch both the existing data and the new data, a benchmark dataset may be set from the digital pathology domain, and the classes may be split into several different subsets, and thus the corresponding examples are split according to their class labels. For example, Figure 9 shows such class splitting for each model version (the lower left plot shows the class labels and the number of examples in each model version). (3.2b) Next, evaluate the candidate adaptation learning methods to determine which adaptation learning method is most effective for learning the classes in an incremental manner, and the performance of learning different sets of class orders for each data batch may be compared. (3.3) AI-based methods for adaptation learning: Batches of the benchmark data generated from (3.2b) may be run against candidate learning methods as follows. (3.3a) Normalization-based methods: These methods prevent the model from forgetting the knowledge it has learned by applying constraints to the current model weights, so that when updating the current model, the model weights do not deviate significantly from their current version. (3.3b) Replay-based methods: These methods preserve the most beneficial examples or their feature representations (e.g., feature vectors extracted from the hidden layers of a neural network) from the existing data batches and revisit them during the training of the newly arrived data. (3.3c) Methods that combine regularization and replay. (3.3d) Methods that utilize meta-learning principles to enable model updates with some examples of each batch (see the "CoPE" method implemented in the experimental section). (3.3e) Parameter isolation methods. (3.3e) Other techniques applied together with the aforementioned methods. For example, perform self-supervised learning on the newly unlabeled data before applying the aforementioned methods. As another example, run a adversarial generative model to generate examples similar to these from previous data batches.

[0129] (iv) Update of the task-incremental digital pathology model: (4.1) In each of the foregoing scenarios, a decision may be made as to whether to formulate each different data batch as a new task, and thus whether to apply a task-incremental approach. Such an approach may specify a separate model component for each task (e.g., a separate set or layer of neurons in a neural network model), and in each model iteration, only the model-specific component is trained and the remainder of the previous trained components remains unchanged. Thus, the model can adapt to changes in data with separate parts and avoid forgetting the knowledge learned from previous tasks.

[0130] Application of the Adaptive Learning Framework to Other Modalities The adaptive learning framework may be applied to other imaging modalities and other research fields. The adaptive learning framework is domain-independent in the following aspects: (1) The preconditioning strategy for initial model training can be applied to other types of data that require a large number of annotations to generate an initial model. (2) The three adaptive learning scenarios for model update / adaptation have commonalities with the scenarios encountered in other computational biomedicine research, and thus the adaptive learning method selection strategy can be utilized by these studies.

[0131] The adaptive learning framework may be applied to federated learning. Federated learning aims to update a global model without sharing data from individual data sources and without explicitly sharing local models. The adaptive learning framework can be utilized by federated learning in the following ways: That is, (1) to precondition local models and / or global models to enable more effective and efficient model updates with a smaller number of annotations within the target image domain, and (2) to continuously update the model without retraining previous data and to select the best learning method by applying the model selection workflow described herein, local models and / or global models are updated via one or more of the adaptive learning methods.

[0132] The adaptive learning framework may be applied to multi-modal learning. Multi-modal learning aims to integrate knowledge learned from different modalities of data. The adaptive learning framework can be utilized by multi-modal learning in the following ways. That is, (1) preconditioning a model from one or more data modalities to enable more effective and efficient model updates with fewer annotations within the target image domain, (2) updating from one or more data modalities via one or more adaptive learning methods to continuously update the model without retraining previous data and selecting the best learning method by applying the model selection workflow described herein, and (3) generating a model for adaptively integrating representations from data of multiple modalities by applying the adaptive learning framework for the first iteration of learning and / or continuous updates / adaptations.

[0133] V. Exemplary Systems for Adaptive Learning FIG. 10 shows a block diagram illustrating a computing environment 1000 for processing digital pathology images using an artificial intelligence system (e.g., one or more machine learning models). As further described herein, processing digital pathology images can include training a machine learning algorithm using digital pathology images and / or converting some or all of the digital pathology images into one or more results using a trained (or partially trained) version of the machine learning algorithm (i.e., a machine learning model).

[0134] As shown in FIG. 10, the computing environment 1000 includes several stages, namely an image storage stage 1005, a preprocessing stage 1010, a labeling stage 1015, a data augmentation stage 1017, a training stage 1020, and a result generation stage 1025.

[0135] The image storage stage 1005 includes one or more image data stores 1030 (e.g., the memory device 430 described in connection with FIG. 4) that are accessed (e.g., by the preprocessing stage 1010) to provide a set of digital images 1035 of a preselected area from a biological sample slide (e.g., a tissue slide) or of the entire biological sample slide. Each digital image 1035 stored in each image data store 1030 and accessed at the image storage stage 1010 may include a digital pathology image generated according to some or all of the processes described with respect to the network 400 shown in FIG. 4. In some embodiments, each digital image 1035 includes image data from one or more scanned slides. Each of the digital images 1035 may correspond to image data from a single specimen and / or image data from a single day from which the underlying image data corresponding to the image was collected.

[0136] Image data can include an image, along with any information regarding color channels or color wavelength channels, and details regarding the imaging platform on which the image was generated. For example, a tissue section may need to be stained by the application of a staining assay that includes one or more different biomarkers associated with a chromogenic stain for brightfield imaging or a phosphor for fluorescence imaging. The staining assay can use a chromogenic stain for brightfield imaging, an organic phosphor for fluorescence imaging, quantum dots, or a combination of an organic phosphor and quantum dots, or any other combination of stains, biomarkers, and observation or imaging devices. Examples of biomarkers include biomarkers such as estrogen receptor (ER), human epidermal growth factor receptor 2 (HER2), human Ki-67 protein, progesterone receptor (PR), programmed cell death protein 1 (PD1), where the tissue section is detectably labeled with respective binding agents (e.g., antibodies) for ER, HER2, Ki-67, PR, PD1, etc. In some embodiments, digital image and data analysis operations such as classification, scoring, Cox modeling, and risk stratification are dependent on the type of biomarker used as well as field of view (FOV) selection and annotation. Further, a typical tissue section is processed on an automated staining / assay platform that applies the staining assay to the tissue section, thereby obtaining a stained sample. There are various commercially available products suitable for use as a staining / assay platform, one example being the VENTANA® SYMPHONY® product of the assignee, Ventana Medical Systems, Inc. The stained tissue section may be supplied, for example, to an imaging system of a microscope or a whole slide scanner having a microscope and / or imaging components, one example being the VENTANA® iScan Coreo® / VENTANA® DP200 product of the assignee, Ventana Medical Systems, Inc. Multiple tissue slides may be scanned with an equivalent multiple slide scanner system.Additional information provided by the imaging system may include any information regarding the staining platform, including the concentration of the chemicals used in the staining, the reaction time of the chemicals applied to the tissue in the staining, and / or pre - analysis conditions of the tissue such as the age of the tissue, the fixation method, the period, the embedding method of the section, the cutting method, etc.

[0137] In the pre - processing stage 1010, one, a plurality, or all of the set of digital images 1035 are each pre - processed using one or more techniques to generate corresponding pre - processed images 1040. The pre - processing may include trimming the images. In some examples, the pre - processing may further include standardization or rescaling (e.g., normalization) to make all features the same scale (e.g., the same size scale or the same color scale or saturation scale). In certain cases, the images are resized such that the minimum size (width or height) is a predetermined number of pixels (e.g., 2500 pixels) or the maximum size (width or height) is a predetermined number of pixels (e.g., 3000 pixels), and optionally maintained at the original aspect ratio. The pre - processing may further include removing noise. For example, the images may be smoothed, such as by applying a Gaussian function or Gaussian blur, to remove unwanted noise.

[0138] The pre - processed images 1040 may include one or more training images, validation input images, and unlabeled images. It should be understood that the pre - processed images 1040 corresponding to the training group, validation group, and unlabeled group do not need to be accessed simultaneously. For example, an initial set of training and validation pre - processed images 1040 may first be accessed and used to train the machine learning algorithm 1055, and subsequently, unlabeled input images may be accessed or received (e.g., once or multiple times later) and used by the trained machine learning model 1060 to provide a desired output (e.g., cell classification).

[0139] In some examples, the machine learning algorithm 1055 is trained using supervised training, and some or all of the pre-processed image 1040 is manually, semi-automatically, or automatically partially or fully labeled at the labeling stage 1015 with labels 1045 that identify the “correct” interpretation (i.e., “ground truth”) of the various biological substances and structures within the pre-processed image 1040. For example, the label 1045 can identify features of interest (e.g.), cell classification, a binary indication as to whether a given cell is a particular type of cell, a binary indication as to whether the pre-processed image 1040 (or a particular region having the pre-processed image 1040) includes a particular type of indication (e.g., necrosis or artifact), a categorical feature of a slide-level or region-specific indication (e.g., identifying a particular type of cell), a number (e.g., identifying the amount of a particular type of cell within a region, the amount of artifact displayed, or the amount of necrosis region), the presence or absence of one or more biomarkers, etc. In some cases, the label 545 includes a position. For example, the label 1045 can identify the point location of the nucleus of a particular type of cell, or the point location of a particular type of cell (e.g., a live dot label). As another example, the label 1045 may include a border or boundary of a drawn tumor, blood vessel, necrosis region, etc. As another example, the label 1045 may include one or more biomarkers identified based on a biomarker pattern observed using one or more stains. For example, a tissue slide stained for a biomarker, such as programmed cell death protein 1 (“PD1”), may be observed and / or processed to label cells as either positive or negative cells considering the expression level and pattern of PD1 in the tissue. Depending on the feature of interest, a given labeled pre-processed image 1040 can be associated with a single label 1045 or multiple labels 1045. In the latter case, each label 1045 can be associated with an indication (e.g.) as to which position or portion within the pre-processed image 1045 the label corresponds to.

[0140] The label 1045 assigned at the labeling stage 1015 can be identified based on inputs from a human user (e.g., a pathologist or an image scientist) and / or an algorithm (e.g., an annotation tool) configured to define the label 1045. In some examples, the labeling stage 1015 can include sending and / or presenting some or all of the one or more pre-processed images 1040 to a computing device operated by a user. In some examples, the labeling stage 1015 includes using (e.g., using an API) an interface presented by a labeling controller 1050 in a computing device operated by a user, the interface including input components for receiving inputs to identify the label 1045 for features of interest. For example, a user interface enabling selection of an image or region of an image (e.g., FOV) for labeling can be presented by the labeling controller 1050. A user operating the terminal can select an image or FOV using the user interface. Some image or FOV selection mechanisms can be provided, such as specifying a known or irregular shape, or defining an anatomical region of interest (e.g., a tumor region). In one example, the image or FOV is the entire tumor region selected on an IHC slide stained with a combination of H&E stains. The selection of the image or FOV can be performed by a user or by an automated image analysis algorithm such as tumor region segmentation on an H&E tissue slide. For example, the user may select that the image or FOV can be automatically designated as the entire slide or tumor, or the entire slide or tumor region as the image or FOV using a segmentation algorithm. Thereafter, the user operating the terminal may select one or more labels 1045 to be applied to the selected image or FOV, such as point positions on cells, positive markers for biomarkers expressed by cells, negative biomarkers for biomarkers not expressed by cells, boundaries around cells, etc.

[0141] In some examples, the interface may identify which particular label 1045 is required and / or to what extent it is required, which may be communicated to the user via (for example) text instructions and / or visualization. For example, a particular color, size, and / or symbol may indicate that label 1045 is required for a particular display (e.g., a particular cell or region or staining pattern) within an image relative to other displays. If multiple labels 1045 corresponding to multiple displays are required, the interface may identify each of the displays simultaneously, or sequentially (such that identification of the next display for labeling is triggered upon providing a label to one identified display). In some examples, each image is presented until the user identifies a particular number (e.g., a particular type) of labels 1045. For example, a given whole slide image or a given patch of a whole slide image may be presented until the user identifies the presence or absence of three different biomarkers, at which point the interface may present different whole slide images or images of different patches (e.g., until a threshold number of images or patches are labeled). Thus, in some examples, the interface is configured to request and / or accept labels 1045 for an incomplete subset of features of interest, and the user may determine which of potentially many displays are to be labeled.

[0142] In some examples, the labeling stage 1015 includes a labeling controller 1050 that implements an annotation algorithm to semi-automatically or automatically label various features of an image or region of interest within the image. The labeling controller 1050 annotates the image or FOV on the first slide and maps the annotation across the remainder of the slide according to user input or the annotation algorithm. Depending on the defined FOV, several methods for annotation and alignment are possible. For example, the entire tumor region annotated on an H&E slide from among a plurality of consecutive slides can be selected automatically or by the user on an interface such as VIRTUOSO / VERSO (trademark). Since other tissue slides correspond to consecutive sections from the same tissue block, the labeling controller 1050 performs a marker-to-marker alignment operation to map the entire tumor annotation from the H&E slide and transfer it to each of the remaining series of IHC slides. Exemplary methods for marker-to-marker alignment are described in more detail in International Publication No. WO 2014 / 140070, "Whole slide image registration and cross-image annotation devices, systems and methods," filed Mar. 12, 2014, by the same applicant, which is hereby incorporated by reference in its entirety for all purposes. In some embodiments, any other method for image alignment and generation of whole tumor annotation may be used. For example, a qualified reader such as a pathologist can annotate the entire tumor region on any other IHC slide and run the labeling controller 1050 to map the whole tumor annotation on other digitized slides. For example, a pathologist (or an automatic detection algorithm) can annotate the entire tumor region on an H&E slide to trigger the analysis of all adjacent consecutively sectioned IHC slides and determine a whole slide tumor score for the annotated regions on all slides.

[0143] In the augmentation stage 1017, the training set of labeled or unlabeled images (original images) from the pre-processed image 1040 is augmented with synthetic images 1052 generated using augmentation control 1054 that executes one or more augmentation algorithms. Augmentation techniques are used to artificially increase the amount and / or type of training data by adding slightly modified synthetic copies of existing training data, or newly created synthetic data from existing training data. As described herein, differences between scanners and between examination rooms can cause variations in intensity and color within digital images. Further, insufficient scanning can result in gradient changes and blurring effects, and in assay staining, staining artifacts such as background washing can occur, and differences in tissue / patient samples can result in variations in cell size. These variations and disruptions can potentially have an adverse impact on the quality and reliability of deep learning and artificial intelligence systems. The augmentation techniques implemented in the augmentation stage 1017 function as regularizers for these variations and disruptions and can help reduce overfitting when training machine learning models. Examples of augmentation techniques include (i) varying within the staining space, where first staining decomposition (e.g., unmixing) is performed and then each stain is remixed with a predetermined color vector to change the hue, saturation, and intensity of each stain, (ii) increasing the image resolution by resizing the image and then returning it to its original size, (iii) the same, or (iv) any combination thereof.

[0144] In training stage 1020, label 1045 and corresponding preprocessed image 1040 can be used by training controller 1065 to train machine learning algorithm 1055 according to various workflows described herein. For example, to train algorithm 1055, preprocessed image 1040 may be divided into a subset 1040a of images for training (e.g., 90%) and a subset 1040b of images for validation (e.g., 10%). The division may be performed randomly (e.g., 90 / 10% or 70 / 30%), or according to more complex validation techniques such as K-fold cross-validation, leave-one-out cross-validation, leave-group-out cross-validation, nested cross-validation, etc. to minimize sampling bias and overfitting. The division may also be performed based on including augmented or synthetic images 1052 in preprocessed image 1040. For example, it may be beneficial to limit the number or ratio of synthetic images 1052 included within the subset of training images 1040a. In some examples, the ratio of the original image 1035 to the synthetic image 1052 is maintained at 1:1, 1:2, 2:1, 1:3, 3:1, 1:4, or 4:1.

[0145] In some examples, the machine learning algorithm 1055 includes a convolutional neural network (CNN), a modified CNN with an encoding layer replaced by a residual neural network (“Resnet”), or a modified CNN with an encoding layer and a decoding layer replaced by Resnet. In other examples, the machine learning algorithm 1055 is any suitable machine learning algorithm configured to localize, classify, and / or analyze the pre-processed image 1040, such as a two-dimensional CNN (“2DCNN”), Mask R-CNN, U-Net, feature pyramid network (FPN), dynamic time warping (“DTW”) technique, hidden Markov model (“HMM”), a pure attention-based model, or a combination of one or more of such techniques, such as a vision transformer, CNN-HMM, or MCNN (multi-scale convolutional neural network). The computing environment 1000 may employ the same type of machine learning algorithm, or different types of machine learning algorithms trained to detect and classify different cells. For example, the computing environment 1000 can include a first machine learning algorithm (e.g., U-Net) for detecting and classifying PD1. The computing environment 500 can also include a second machine learning algorithm (e.g., 2DCNN) for detecting and classifying the differentiation cluster 68 (“CD68”). The computing environment 1000 can also include a third machine learning algorithm (e.g., U-Net) for detecting and classifying the combination of PD1 and CD68. The computing environment 1000 can also include a fourth machine learning algorithm (e.g., HMM) for diagnosing a disease for treating or prognosticating a subject such as a patient. In other examples according to the present disclosure, still other types of machine learning algorithms can be implemented.

[0146] The training process of the machine learning algorithm 1055 includes selecting hyperparameters of the machine learning algorithm 1055 from the parameter data store 1063, inputting a subset 1040a of images (e.g., labels 1045 and corresponding preprocessed images 1040) into the machine learning algorithm 1055, and performing an iterative operation to learn a set of parameters (e.g., one or more coefficients and / or weights) of the machine learning algorithm 1055. Hyperparameters are settings that can be adjusted or optimized to control the behavior of the machine learning algorithm 1055. Most algorithms explicitly define hyperparameters that control different aspects of the algorithm, such as memory or execution cost. However, additional hyperparameters can be defined to adapt the algorithm to a specific scenario. For example, hyperparameters can include the number of hidden units of the algorithm, the learning rate of the algorithm (e.g., 1e-4), the convolutional kernel width, or the number of kernels of the algorithm. In some examples, the number of model parameters decreases for each convolutional layer and deconvolutional layer, and / or the number of kernels is reduced by half for each convolutional layer and deconvolutional layer compared to a typical CNN.

[0147] The subset 1040a of images can be input into the machine learning algorithm 1055 as a batch of a predetermined size. The batch size limits the number of images presented to the machine learning algorithm 1055 before parameter updates can be performed. Alternatively, the subset 1040a of images can be input into the machine learning algorithm 1055 as a time series or sequentially. In either case, if the augmented image or synthetic image 1052 is included within the preprocessed image 1040a, the number of original images 1035 per number of synthetic images 1052 included within each batch, or the way the original images 1035 and the phenotypic images 1052 are supplied to the algorithm (e.g., every other batch or image is the original batch or image of images) can be defined as a hyperparameter.

[0148] Each parameter is a variable that can be adjusted such that the value for the parameter is adjusted during training. For example, a cost function or objective function can be configured to optimize the accurate classification of the presented representations, optimize the characterization of a given type of feature (e.g., characterization of shape, size, uniformity, etc.), optimize the detection of a given type of feature, and / or optimize the accurate localization of a given type of feature. Each iteration can include learning a set of parameters of the machine learning algorithm 1055 that minimizes or maximizes the cost function of the machine learning algorithm 1055, such that the value of the cost function using the set of parameters is made smaller or larger than the value of the cost function using another set of parameters in the previous iteration. The cost function can be configured to measure the difference between the output predicted using the machine learning algorithm 1055 and the label 1045 included in the training data. For example, in the case of a supervised learning-based model, the goal of training is to learn a function “h()” (sometimes called a hypothesis function) that maps the training input space X to the target value space Y, h:X→Y, such that h(x) is a good predictor of the corresponding value of y. Various different techniques may be used to learn this hypothesis function. In some techniques, as part of deriving the hypothesis function, a cost function or loss function may be defined that measures the difference between the ground truth value for an input and the predicted value for that input. As part of training, techniques such as backpropagation, random feedback, direct feedback alignment (DFA), indirect feedback alignment (IFA), Hebbian learning, etc. are used to minimize this cost or loss function.

[0149] The training iterations continue until the stop condition is met. The training completion conditions can be configured to be met when, for example, a predetermined number of training iterations are completed, when a statistical value generated based on testing or validation exceeds a predetermined threshold value (e.g., a classification accuracy threshold), when a statistical value generated based on a confidence measurement criterion (e.g., the average or median of a confidence metric exceeding a specific value or a percentage of the confidence metric) exceeds a predetermined confidence threshold, and / or when the user device involved in the training review closes the training application executed by the training controller 1065. When a set of model parameters is identified through training, the machine learning algorithm 1055 is trained, and the training controller 1065 executes an additional process of testing or validation using a subset 1040b of the images (test or validation dataset). The validation process may include iterative operations of inputting images from the subset 1040b of the images to the machine learning algorithm 1055 using validation techniques such as k-fold cross-validation, leave-one-out cross-validation, leave-group-out cross-validation, nested cross-validation, etc., to adjust hyperparameters and ultimately find an optimal set of hyperparameters. When an optimal set of hyperparameters is obtained, a reserved test set of images from the subset 1040b of the images is input to the machine learning algorithm 1055 to obtain an output, and the output is evaluated against the ground truth by calculating performance metrics such as error, accuracy, precision, recall, receiver operating characteristic curve (ROC), etc., using correlation techniques such as the Bland-Altman method and Spearman's rank correlation coefficient. Optionally, a new training iteration may be started in response to receiving a corresponding request or trigger condition from the user device (e.g., initial model development, model update / adaptation, continuous learning, drift determined within the trained machine learning model 1060, etc.).

[0150] As will be appreciated, other training / validation mechanisms are contemplated to be implemented within computing environment 1000. For example, with respect to images from image subset 1040a, machine learning algorithm 1055 may be trained and hyperparameters may be adjusted, and images from image subset 1040b may be used only to test and evaluate the performance of machine learning algorithm 1055. Further, the training mechanisms described herein focus on the training of new machine learning algorithm 1055. These training mechanisms can also be utilized for the development of an initial model, the updating / adapting of a model, and continuous learning of an existing machine learning model 1060 trained from other datasets, as will be described in detail herein. For example, in some cases, machine learning model 1060 may be conditioned on images of other objects or biological structures or on sections from other subjects or studies (e.g., human trials or mouse experiments). In those cases, machine learning model 1060 can be used for the development of an initial model, the updating / adapting of a model, and continuous learning using preprocessed images 1040.

[0151] Next, (in result generation stage 1025) using the trained machine learning model 1060, the new preprocessed image 1040 can be processed to predict the cell center and / or position probability, classify the cell type, generate a cell mask (e.g., a segmentation mask for each pixel of the image), predict the diagnosis or prognosis of a disease of a subject such as a patient, or a combination thereof, to generate predictions or inferences. In some examples, the mask identifies the displayed cell positions associated with one or more biomarkers. For example, based on tissue stained for a single biomarker, the trained machine learning model 1060 can be configured to (i) infer the cell center and / or position, (ii) classify the cells based on the characteristics of the staining pattern associated with the biomarker, and (iii) output a cell detection mask for positive cells and a cell detection mask for negative cells. As another example, based on tissue stained for two biomarkers, the trained machine learning model 1060 can be configured to (i) infer the cell center and / or position, (ii) classify the cells based on the characteristics of the staining patterns associated with the two biomarkers, and (iii) output a cell detection mask for cells positive for the first biomarker, a cell detection mask for cells negative for the first biomarker, a cell detection mask for cells positive for the second biomarker, and a cell detection mask for cells negative for the second biomarker. As another example, based on tissue stained for a single biomarker, the trained machine learning model 1060 can be configured to (i) infer the cell center and / or position, (ii) classify the cells based on the cell characteristics and the staining pattern associated with the biomarker, and (iii) output a cell detection mask for positive cells, a cell detection mask for negative cell codes, and a masked cell classified as a tissue cell.

[0152] In some examples, the analysis controller 1080 generates analysis results 1085 for the entity that requested the processing of the underlying image. The analysis results 1085 may include a mask output from the trained machine learning model 1060 overlaid on the new preprocessed image 1040. Additionally or alternatively, the analysis results 1085 may include information calculated or determined from the output of the trained machine learning model, such as an overall slide tumor score. In an exemplary embodiment, the automated analysis of tissue slides uses VENTANA's FDA-cleared 510(k) approved algorithm, which is the assignee. Alternatively or additionally, any other automated algorithm may be used to analyze selected regions of the image (e.g., the masked image) to generate a score. In some embodiments, the analysis controller 1080 may further respond to instructions received from a computing device, such as a pathologist, physician, investigator (e.g., associated with a clinical trial), subject, medical professional, etc. In some examples, the communication from the computing device includes identifiers for each of a specific set of subjects and corresponds to a request to perform an iteration of analysis for each subject represented by that set. The computing device can further perform an analysis based on the output of the machine learning model and / or the analysis controller 1080, and / or provide recommended diagnoses / treatments to the subject.

[0153] The computing environment 1000 is exemplary, and it will be understood that computing environments 1000 having different stages and / or using different components are contemplated. For example, in some examples, the network may omit the preprocessing stage 1010, whereby the images used to train the algorithm and / or the images processed by the model are raw images (e.g., from an image data store). As another example, it will be understood that each of the preprocessing stage 1010 and the training stage 1020 can include a controller for performing one or more operations described herein. Similarly, the labeling stage 1015 is shown in relation to the labeling controller 1050, and the result generation stage 1025 is shown in relation to the analysis controller 1080, but the controller associated with each stage can further or alternatively facilitate other operations described herein other than the generation of labels and / or the generation of analysis results. As yet another example, the display of the computing environment 1000 shown in FIG. 10 lacks the displayed representation of devices associated with a programmer (e.g., who selected the architecture of the machine learning algorithm 1055 that defines how various interfaces function), a device associated with a user who provides an initial label or label review (e.g., in the labeling stage 1015), and a device associated with a user who requests model processing of a given image (which may be the same user or a different user than the user who provided the initial label or label review). Despite the absence of the display of these devices, the computing environment 1000 may include the use of one, multiple, or all of the devices, and in fact, may include the use of multiple devices associated with corresponding multiple users who provide an initial label or label review and / or multiple devices associated with corresponding multiple users who request model processing of various images.

[0154] VI. Techniques for Training a Machine Learning Algorithm Using an Adaptive Learning Framework FIG. 11 is a flowchart showing a process 1100 for using a training set of images to train a machine learning algorithm according to various embodiments. The process 1100 shown in FIG. 11 may be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of each system, hardware, or combinations thereof. The software may be stored in a non-transitory storage medium (e.g., a memory device). The process 1100 presented in FIG. 11 and described below is intended to be exemplary and non-limiting. Although FIG. 11 shows various processing steps performed in a particular sequence or order, it is not limited thereto. In certain alternative embodiments, the steps may be performed in some different order or some steps may be performed in parallel. In certain embodiments, such as the embodiments shown in FIGS. 4 and 10, the processing shown in FIG. 11 is part of a training phase (e.g., algorithm training 1020) for training a machine learning algorithm using a training set of images to generate a machine learning model configured to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within an image.

[0155] Process 1100 starts at block 1105, where a first annotated training set of images is obtained to train a machine learning algorithm to perform detection, characterization, classification, or combinations thereof of some or all regions or objects within the images. The first annotated training set of images is in a first image domain (e.g., the pre-processed image 1040 of the computing environment 1000 described with respect to FIG. 10). In some cases, the first annotated training set of images is a digital pathology image containing one or more types of cells. The first annotated training set of images may depict cells having staining patterns related to biomarkers. In some cases, the first annotated training set of images shows cells having multiple staining patterns related to multiple biomarkers. The first annotated training set of images may be annotated with training labels (e.g., supervised, semi-supervised, or weakly supervised).

[0156] At block 1110, the first annotated training set of images is divided into mini-sets of images, each mini-set representing a distinct modeling sub-task and containing a limited number of examples. In some cases, the division includes, for each distinct modeling sub-task, selecting a subset of classes to be a portion of or match the number of classes targeted in the second stage when only one mini-set of images is available, and selecting a limited number of examples based on the selected subset of classes, and, when multiple mini-sets of images are available, for each distinct modeling sub-task, either (i) mixing examples from multiple mini-sets of images, selecting a subset of classes to be a portion of or match the number of classes targeted in the second stage, and selecting a limited number of examples from the mixed examples based on the selected subset of classes, or (ii) selecting one mini-set of images from the multiple mini-sets of images, selecting a subset of classes to be a portion of or match the number of classes targeted in the second stage, and selecting a limited number of examples from the selected mini-set of images based on the selected subset of classes.

[0157] In block 1115, the machine learning algorithm is trained in a first stage using the mini-set of images to generate a preconditioned machine learning model configured to detect, characterize, classify, or a combination thereof, some or all regions or objects in new images. In some cases, the first stage further includes an inner learning loop, in which the machine learning algorithm generates a loss on a validation set of images after model updating, denoted as L-subtask-i for the i-th subtask, updates model weights or parameters on one subtask for a predetermined or flexible number of epochs to initialize the preconditioned machine learning model for adaptation to the target dataset, and an outer learning loop, in which the goal is to search for a set of model initializations that, when used to update all subtasks with only a limited number of examples each, generates a preconditioned machine learning model by finding a model initialization that minimizes the sum of all losses, denoted as L-subtask-i summed together, where i ranges from 1 to the number of subtasks calculated from the validation set of images for the subtask for model initialization.

[0158] In some cases, training the first stage involves performing an iterative operation to learn a set of parameters to detect, characterize, classify, or a combination thereof, some or all of the regions or objects in the mini-set of images that maximizes or minimizes a cost function, where each iteration involves finding a set of parameters for the machine learning algorithm such that the value of the cost function using the set of parameters is greater or smaller than the value of the cost function using another set of parameters in a previous iteration, where the cost function is constructed to measure the difference between predictions made for some or all of the regions or objects using the machine learning algorithm and ground truth labels assigned to the mini-set of images.

[0159] In block 1120, a limited number of images from the target dataset are labeled to generate a second annotated training set of images for training a machine learning algorithm to perform detection, characterization, classification, or combinations thereof, of some or all regions or objects within the images. The second annotated training set of images is in a second image domain (different from the first image domain). In a particular example, the limited number of images is less than 50 images, 30 images, or 20 images.

[0160] In block 1125, a preconditioned machine learning model is trained in a second stage using the second annotated training set of images to generate a target machine learning model configured to perform detection, characterization, classification, or combinations thereof, of some or all regions or objects within new images. The number of classes targeted in the first stage is part of or matches the number of classes targeted in the second stage. In some cases, the second stage further includes applying the preconditioned machine learning model to generate a feature vector representation for each example within the second annotated training set of images, combining the feature vector representations from examples of the same class to generate one representation per target class for use as a prototype in the target class, generating a feature vector representation for the remaining portion of the images or image regions of the images not labeled from the target dataset, and comparing each feature vector representation from the unlabeled images to the prototype in the target class based on the distance between the feature vector representation from the unlabeled images and the prototype for the target class.

[0161] In some embodiments, training in the second stage involves performing detection, characterization, classification, or combinations thereof of some or all regions or objects within a second annotated training set of images that maximize or minimize a cost function, and performing an iterative operation to learn a set of parameters so that, for each iteration, a value of the cost function using the set of parameters is greater or less than a value of the cost function using another set of parameters in the previous iteration, which involves finding a set of parameters for a conditional machine learning model, and the cost function is constructed to measure the difference between predictions made for some or all regions or objects using the conditional machine learning model and the ground truth labels given in the second annotated training set of images.

[0162] In optional block 1130, a target machine learning model is provided. For example, the target machine learning model may be deployed for execution in an image analysis environment as described with respect to FIG. 10.

[0163] In block 1135, a digital pathology scenario is identified. The digital pathology scenario can be a data incremental scenario, a domain incremental scenario, a class incremental scenario, or a task incremental scenario.

[0164] In block 1140, an adaptive continuous learning method is selected to update the target machine learning model based on the digital pathology scenario. In some examples, the adaptive continuous learning method is selected from the group including elastic weight consolidation (EWC), learning without forgetting (LWF), incremental classifier and representation learning (iCaRL), continuous prototype evaluation (CoPE), A-GEM, and parameter separation methods. In other examples, the adaptive continuous learning method includes EWC, LWF, iCaRL, CoPE, A-GEM, parameter separation methods, similar continuous learning methods, or any combination thereof.

[0165] In block 1145, the target machine learning model is updated based on an adaptive continuous learning method to generate an updated machine learning model.

[0166] In optional block 1150, an updated machine learning model is provided. For example, the updated machine learning model may be deployed for execution in an image analysis environment as described with respect to FIG. 10.

[0167] In block 1155, a new image is received. The new image may be divided into image patches of a predetermined size. For example, a full slide image usually has a random size, and machine learning algorithms such as a modified CNN learn better (e.g., parallel computation for batches of images of the same size, memory constraints) with a normalized image size, and thus the image may be divided into image patches having a specific size to optimize the analysis. In some embodiments, the image is divided into image patches having a predetermined size of 64 pixels × 64 pixels, 128 pixels × 128 pixels, 256 pixels × 256 pixels, or 512 pixels × 512 pixels.

[0168] In block 1160, the new image or image patch is input into the target machine learning model or the updated machine learning model. In block 1165, the target machine learning model or the updated machine learning model performs detection, characterization, classification, or a combination thereof of some or all regions or objects within the new image or image patch, and outputs an inference based on the detection, characterization, classification, or a combination thereof.

[0169] In optional block 1170, a diagnosis of the object associated with the image or image patch is determined based on the inference output by the revised machine learning model.

[0170] In an optional block 1175, a treatment is applied to a subject associated with an image or an image patch. In some examples, the treatment is applied based on (i) an inference output by a machine learning model or a revised machine learning model, and / or (ii) a diagnosis of the subject determined at block 1170.

[0171] VII. Examples The systems and methods implemented in various embodiments may be better understood by reference to the following examples.

[0172] Data CRC: In the following experiments, for training, 100,000 non - overlapping patches were used from H&E - stained histological images of human colorectal cancer (CRC) composed of 9 tissue classes including adipose tissue (ADI), background (BACK), debris (DEB), lymphocytes (LYM), mucus (MUC), smooth muscle (MUS), normal colon mucosa (NORM), cancer - associated stroma (STR), and colorectal adenocarcinoma epithelium (TUM). Some exemplary images are shown in FIG. 12. The test set included 7,180 image patches that did not overlap with the training data. All images were color - normalized using Macenko's method.

[0173] The CRC dataset was augmented by varying the staining intensity, color, and saturation individually and in combination to simulate data collected from different stainers, scanners, and chromogens. The images were not mixed using non - negative matrix factorization. Four settings of color, saturation, and intensity were applied to individual stains from non - overlapping subsets of the original dataset. Along with the original dataset, each synthetic setting was used to create different adaptive learning scenarios. 1. Each extended setting was recognized as a domain shift from the original dataset. Thus, the dataset can be split into five data streams, each representing a different extension or domain setting that results in five experience domain incremental scenarios (although five were used in the experiments disclosed herein, it should be understood that the dataset can be split into any number "n" of data streams, each representing a different extension or domain setting that results in "n" experience domain incremental scenarios). The model learns to classify images in the new domain setting across all the experiences the model is targeted at. 2. The different extensions were also evenly mixed across the entire classes. Using this mixed dataset, five data streams or experiences of equal size with representations from all classes were created (although five were used in the experiments disclosed herein, it should be understood that the dataset can be split into any number "n" of data streams or experiences of equal size with representations from all classes). This constitutes a data incremental scenario where each subsequent experience is added to the model's training data. 3. The evenly mixed dataset was also split into subsets or experiences, each containing different classes, forming a class incremental scenario. Depending on how the data is split, the model is exposed to different classes in each experience and targeted at all classes in the training process.

[0174] Continuous learning scenarios: Each synthetic setting can be used to create different continuous learning scenarios together with the original dataset. Each augmentation setting is a domain shift from the original dataset. They were used separately as experiences in individual data streams or domain incremental settings. Images from different augmentation settings were uniformly mixed across classes and split into experiences with equal class representation to form a data incremental scenario. The uniformly mixed dataset was also split into experiences each containing different classes to form a class incremental scenario.

[0175] PatchCam: The PatchCam benchmark dataset consists of 327,680 patches extracted from 400 H&E stained whole slide images of lymph node sections from breast tissue, sized 96×96 pixels at a magnification of 10x. The 75 / 12.5 / 12.5% training / validation / test split was selected using a hard negative mining regime. The dataset has two classes (normal and tumor) to indicate the presence of metastatic tissue. For consistency and easy comparison with the CRC dataset, this dataset was also normalized using Macenko's method. Normalized examples from both classes are shown in Figure 13. The upper row contains samples from the normal class and the lower row contains examples from the tumor class.

[0176] Continuous learning scenario: A dramatic domain shift in the data stream was evaluated by training the model using the original CRC images (stain normalized) in the first experience and the normalized PatchCam dataset in the second experience.

[0177] Method The following sequential learning methods were experimented with, dealing with three scenarios: EWC and online EWC, LwF, iCaRL, CoPE and A-GEM. All methods were compared to two baselines: 1) training from scratch (upper bound), where an 18-layer ResNet with the same network architecture was targeted on all data available from all experiences seen so far, and 2) transfer learning or fine-tuning (i.e., lower bound), where the model was trained with the same design as sequential learning where the model was exposed only to the data available during a particular experience, but instead of using a strategy to mitigate forgetting, the model was only fine-tuned to adapt to new classes. For all experiments, the training epochs were set to 15 and the batch size to 16. The same ResNet architecture was used in a multi-head setting where each head was used for different tasks when testing with A-GEM according to their findings. The stochastic gradient descent optimizer was used starting with a learning rate of 0.1, a momentum of 0.9, and a weight decay of 0.00001 applied after epochs 10 and 13.

[0178] Example 1. - Data incremental setting The first experiment included a data incremental scenario where more data was sequentially fed to the model and then the model was updated based only on the latest data without accessing any of the older datasets it was previously trained on. The newer data stream has the same classes as the older stream but may have a shift in distribution. The mixed dataset used in this experiment should have a uniform distribution across experiences.

[0179] To train the model with this setting, a method called Continuous Prototype Evaluation (CoPE) was used. CoPE is an online data incremental algorithm that uses prototypes to represent the most important features from the data. The prototypes evolve continuously as the model learns to keep up with the changes in the data and can make accurate predictions. CoPE also incorporates balanced replay to ensure that all classes are well represented in the replay population. The data was supplied in an online manner as mini-batches or mini-experiences, i.e., the model saw each data sample only once and was thus trained in a single epoch. To reduce forgetting, a mini-experience size of 128 samples (i.e., each mini-experience had only 128 samples) was used for training with a batch size of 10 and a momentum of 0.99. Thus, for the data incremental scenario created for the extended dataset, each of the five experiences had 99 mini-experiences.

[0180] At the end of training, the test stream with samples from different experiences had an average accuracy of 76%. Figure 14 shows the accuracy of the test stream at the end of each major experience under data incremental settings using the extended CRC dataset with CoPE. Experiences were divided into user-defined sizes of 128 images that resulted in 99 mini-experiences within each experience, which were fed to the model in an online manner. The numbers shown in the plot were at the end of the 99th mini-experience or at the end of all major experiences. The classification accuracy increased gradually, indicating that the model benefited from more data. The accuracy of individual test streams also increased for models learned from newer training streams, indicating that the model did not forget what it had learned previously. Equal representation from all synthetic domains in each experience led to transfer learning and test streams that benefited from all the data. It is also worth noting that the model performed similarly on test streams from subsequent experiences. For example, at the end of primary experience 1, the model performed similarly on test stream 1 as well as test streams 2, 3, and 4, showing some transfer learning. The mixed extended dataset used here has equal representation from all synthetic domains, resulting in similar distributions across experiences. The accuracy also improved by less than 1% between experiences 3 and 4, indicating that the model performed similarly with 80% of the data.

[0181] Example 2. - Domain Incremental Setting The second experiment involved a domain incremental scenario with five experiences that varied the hue, saturation, and intensity values of the two stains, eosin and hematoxylin, to different degrees to mimic images obtained using different stains, scanners, and reagents. Examples from the five experiences are shown in FIG. 15. Each row corresponds to one of the nine tissue types, and three exemplary images are plotted from each augmentation setting (i.e., domain). Domain 0 (columns 1 - 3): Staining normalization CRC dataset. Domain 1 (columns 4 - 6): Simulates scenarios of increasing eosin staining intensity, increasing eosin solution concentration, or extending the staining time. Domain 2 (columns 7 - 9): Simulates scenarios of decreasing eosin intensity and aged slides with faded stain. Domain 3 (columns 10 - 12): Difference in hue change between hematoxylin and eosin. Domain 4 (columns 13 - 15): Hue change and increase in saturation level for both eosin and hematoxylin. Each of the three - column sections can be considered a separate domain in the domain incremental scenario. The domains can be mixed and split into experiences that each have representations from all classes for the data incremental scenario, or each can be split into mixed - domain experiences that each have a non - overlapping subset of classes for the class incremental scenario.

[0182] A method called Learning without Forgetting (LwF) was adopted for this scenario. LwF is a combination of fine - tuning and distillation. LwF uses only the most recent data corresponding to the current task to learn the task - specific parameters of the new / current task without degrading the performance of the old tasks. Different from conventional regularization that penalizes parameter changes based on their importance, LwF penalizes changes in the mapping from input to output. The loss function consists of two terms: the cross - entropy loss for the current task and the distillation loss to prevent the previously acquired knowledge from being forgotten.

[0183] LwF functioned well with an accuracy exceeding 86% in three out of five predetermined domains. The evaluation accuracies for Domains 1 and 2 during the experience were 88% and 93% respectively. Although the model was targeted in the corresponding domains, the retention of knowledge was insufficient. However, when specific domain data became no longer available, especially in Domains 1 and 2, approximately 28% of the acquired domain-specific knowledge was forgotten. The results are shown in Figure 16, which presents the accuracy (left) and forgetting (right) of the test stream at the end of each major experience under the domain-incremental setting using the extended CRC dataset with LwF. At the end of training, the model was executed with an accuracy of 86 - 94% in three out of five predetermined domains. The model functioned well in Domain 1 (88%) and Domain 2 (93%), but the acquired domain-specific knowledge was forgotten during the experience after the specific domain data became no longer available. The test stream forgetting metric was approximately 28% at the end of training, indicating the accuracy loss from Domains 1 and 2 over the course of training tires. This experiment demonstrated that a model presented with data from different domains within a continuous data stream can be trained to reasonably adapt well to new domains while still maintaining performance in previous domains that it can no longer access.

[0184] Example 3. - Class-incremental setting In the third experiment, the model was trained in the class incremental setting described herein using three experiences and six classes such that the model could access data from only two classes during each experience and new classes were gradually added in each experience. Incremental Class & Representation Learning (iCaRL) was the adaptation learning strategy used here. iCaRL dynamically selects exemplars from the data stream and each class has its own set of exemplars. iCaRL updates both the parameters and the exemplars when new data is seen. iCaRL performs classification by nearest neighbor averaging of exemplars. iCaRL includes representation learning by distillation and prototype rehearsal, and the extended dataset includes data from the current task, and the stored exemplars and model parameters were updated based on the cross-entropy loss for the newer classes and the distillation loss for the previously learned classes.

[0185] There are two baselines compared to the iCaRL algorithm. One of them is - 1) training from scratch where the same neural net is targeted with all the data available up to a particular experience (upper bound). That is, the iCaRL algorithm is training from scratch targeted at two classes during the first experience, four classes during the second experience, and all six classes during the third experience, and the other baseline is - 2) transfer learning or fine-tuning (lower bound) where the model was trained with the same design as adaptation learning that exposes only two classes between each of the three experiences, but instead of using a strategy to mitigate forgetting, the model was only fine-tuned to adapt to the newer classes. The results are shown in Figure 17. As shown, it was observed that iCaRL functions equivalently to the upper bound of training from scratch using only a portion of the data and thus also provides a significant computational advantage. Transfer learning worked well at experience 0 but forgot the previously acquired knowledge with the exposure and learning of newer classes. Transfer learning performance was insufficient with catastrophic forgetting, though computationally comparable to iCaRL.

[0186] Example 4. - Comparison of Methods Continuous learning using an extended CRC dataset. For a fair comparison, the domain and data incremental experiments had five experiences, and the class incremental experiments had four experiences where the first three experiences each had two classes and the last experience had the remaining three classes. A-GEM was treated as a task incremental method having each experience of introducing a new set of classes to the model with a task ID since it was found to provide the best results with the task descriptor. The hyperparameters of each method were determined by grid search. CoPE and A-GEM were treated as online, few-shot methods and trained for only 1 epoch. iCaRL was experimented with in three settings. The first setting had four experiences where the first three each had two classes and the last experience had the remaining three classes. The second setting also had four experiences, but had three classes in the first experience and two classes each in the remaining experiences. The final setting had three experiences with three classes each. The class order was the same across settings and in ascending order. LwF was experimented with for continuous learning in the domain incremental setting using the original CRC dataset and the normalized PatchCam dataset, where one tumor type was considered one domain.

[0187] The evaluation accuracy at the end of training for three designed scenarios (data, domain, and class incremental) for continuous learning methods is shown in Fig. 18. In the data and domain incremental scenarios, LwF and iCaRL were comparable to the upper bound baseline. The class incremental scenario was a more difficult task to learn overall. The accuracy of iCaRL was the highest at 83%. A-GEM was the only method tested and evaluated in the task incremental scenario.

[0188] Data incremental scenario: LwF had an overall accuracy of 93% at the end of training and forgot <1% of the previously acquired knowledge. This was 4% better than the lower bound and within 0.5% of the upper bound. The accuracy per experience is shown in Figure 19. Specifically, Figure 19 shows the data incremental experiences tested with different CL methods, as listed in the legend including two baselines. Each experience is composed of its own test stream that includes examples from a smaller batch of data that exclusively belongs to that experience. Each subplot shows how the model performed on the test stream evaluated at the end of training for all experiences. The gray areas indicate experiences that the model has not yet been targeted at, and the model is expected to not function well for these experiences. However, once targeted at a specific experience, ideally it should not forget what it has learned and should retain knowledge throughout the rest of the training process, i.e., the accuracy should remain high in the non-greyed out areas for all test streams. LwF had the highest overall accuracy at the end of training.

[0189] The classification accuracy increased gradually, indicating that the model benefited from more data. The accuracy of individual test streams also increased for models that learned from newer training streams, indicating that the model did not forget what it had learned previously. Another observation was the performance of the iCaRL method. Although iCaRL was designed as a class incremental method, the concept of storing exemplars representing classes within each domain should theoretically have worked better than EWC and LwF. The maximum memory size tested was not sufficient to store only "n" classes as in class incremental, but it was possible to store "n" classes × "d" domains.

[0190] Domain Incremental Scenario: As shown in Figure 20, iCaRL functioned better than EWC and LwF with the best hyperparameters selected from grid search. As shown in Figure 20, different CL methods were tested with a 5-domain incremental experience, as listed in the legend including two baselines. Each experience was composed of its own test stream containing examples from the domain exclusively belonging to that experience. Each subplot shows how the model was executed against the test stream evaluated at the end of training for all experiences. The gray area indicates the experiences not yet targeted by the model, and the model is expected not to function well against these experiences. However, once targeted in a specific experience, ideally, it should not forget what it has learned and should retain knowledge throughout the rest of the training process, i.e., the accuracy should remain high in the non-gray-out areas for all test streams. It can be seen that iCaRL is comparable to the upper bound baseline, while the performance of EWC and LwF was close to the lower bound baseline. Note that in the first experience, since the number of training examples was small (20% of all examples), the upper bound was low (accuracy 0.57).

[0191] Interestingly, 1) the model retained knowledge from some domains more than others. Specifically, it was more difficult to retain knowledge learned on datasets where eosin intensity increased (domain 1; columns 4 - 6 in FIG. 15) or decreased (domain 2; columns 7 - 9 in FIG. 15), but there was some transfer learning between domains for changes in the hue and / or saturation of the stain (domains 3 and 4, not shown in FIG. 15). Even after training only on domain 0, the model functioned well on the test stream corresponding to domain 4. The performance on these two domains remained high even at the end of the training process. Since the fourth domain was generated using hue changes to the stain, these results suggest that the model can handle the range of hue changes when continuously learning and retaining knowledge from sequential data streams.

[0192] Class incremental scenario: iCaRL functioned significantly better than other methods. None of the tested methods, including the lower baseline, could retain the knowledge about the classes learned in previous experiences. The overall accuracy of iCaRL was 88%, which was approximately 6% lower than the joint training upper bound. Considering the low data storage and resource load, it was still beneficial. Figure 21 shows the experience of 4-class incremental tested with different CL methods, as listed in the legend including two baselines. Each experience was composed of its own test stream that included examples from the classes exclusively belonging to that experience. Each subplot shows how the model was executed against the test stream evaluated at the end of training for all experiences. The grayed-out areas indicate the experiences on which the model has not yet been trained, and the model is expected to function poorly against these experiences. However, when the model is targeted at a specific experience, ideally, the model should not forget what it has learned and should retain the knowledge throughout the rest of the training process. That is, the accuracy should remain high in the non-grayed-out areas for all test streams. It can be seen that iCaRL is the only CL method that functions well in the current experience and does not completely forget the previous knowledge. Experience 0, the first experience the model was targeted at, was the most forgotten experience with 64% accuracy at the end of the training process.

[0193] Few-shot Online Continual Learning: The largest gains by both CoPE and A-GEM are associated with the amount of training data having each experience. A-GEM was tested using a class-incremental setting with task IDs, and the model updates were based only on 128 randomly selected examples memorized in memory in one epoch while producing results comparable to iCaRL trained over 15 epochs instead of an online method. The overall accuracy at the end of training was 79%, which was more than 50% better than the lower-bound baseline. The detailed results are shown in Figure 22 - The dataset was initially designed as a class-incremental scenario, with each experience assigned a separate task ID and targeted with a multi-head architecture in an online fashion. The overall accuracy was approximately 79%, and the baseline lower bound was 27%. A-GEM also deteriorates significantly when used without task IDs. A-GEM was unable to retain any of the previously acquired knowledge.

[0194] CoPE was tested in a domain incremental scenario, where the dataset was split into mini-experiences each having the same number of examples as the mini-batch size to simulate online training. This had an overall accuracy that was 67% - 11% better than the lower bound baseline. It is interesting to compare the results from CoPE with those of LwF and EWC from the domain incremental scenario. In both of the latter methods, the model did not function well on the test streams from domains 1 and 2. CoPE did not help in retaining knowledge from domain 1 (less than 25% accuracy), but the model had 60% accuracy from domain 2. The online setting might have helped in retaining more information. CoPE was also found to be sensitive to the softmax temperature. In contrast to using a temperature > 1 as in other distillation methods, lower temperatures were tested, as recommended in the literature, for a harder softmax distribution. A finer sweep of the hyperparameters could yield better results. Another point to note is how each experience is split into mini-batches or mini-experiences in CoPE. Each experience in the tested setting had 128 samples or examples. Not all classes were equally represented in each mini-experience, which could also affect the overall accuracy.

[0195] The comparison between CoPE and the baseline is shown in Figure 23. Experiences were split with a user-defined size of 128 images, and within each experience, 99 mini-experiences were obtained and fed to the model in an online fashion. The numbers shown in the left plot are at the end of the 99th mini-experience or at the end of all major experiences. The right side shows the results from the fine-tuning / naive baseline. Equal presentation from all synthetic domains in each experience led to transfer learning and test streams that benefited from all the data for both CoPE and the baseline, but the overall accuracy was better for CoPE.

[0196] Impact of Class Grouping on Sequential Learning The results of the third experiment using different class grouping settings are shown in Figure 24. Three settings were tested: A. Four experiences of two classes out of the first three classes and four classes of the last three classes; B. Four experiences of three classes in the first experience and two classes in the remaining experiences; C. Three experiences with three classes each. Comparing the first two subplots, both having four experiences, the overall accuracy drops by about 6.5%. The difference between the two settings is the number of classes the model was targeted at in the first experience. The hypothesis here is based on curriculum learning that knowledge is better captured when a more difficult task follows an easier task. Here, starting with three classes may make it more difficult for the model to learn and retain the knowledge reflected in the accuracy drop despite the remaining experiences where it was targeted at only two classes each in both settings. The same is true for the third setting. Despite having fewer experiences, the model had to learn more - three classes per experience for the two classes in setting 1, and the accuracy dropped by nearly 8%.

[0197] Sequential Learning from Multiple Tumor Types Both EWC and LwF were evaluated in this experiment, and LwF yielded slightly better results, shown in Figure 25. CRC was designed as the first domain and PatchCam as the second domain. The model started well with an accuracy of over 90% in the CRC test stream but, by the end of training, forgot some of the knowledge obtained in the first experience. At the end of the training process, the model produced an accuracy of about 70% in the CRC test stream and about 76% in the PatchCam test stream.

[0198] Conclusion This systematic study characterized the performance of various continuous learning methods for different scenarios using extended digital pathology images and evaluated the models when different tumor types were presented. The dataset was evaluated with regularization and replay methods. EWC and LwF performed relatively well in data and domain incremental scenarios, while rehearsal methods such as iCaRL and A-GEM were necessary to prevent catastrophic forgetting in more difficult class incremental scenarios. The few online methods tested required additional fine-tuning of hyperparameters and experimental settings to fully understand their effectiveness. Furthermore, it is interesting to investigate how changes in images from a clinical perspective due to shifted patient populations, disease progression, and / or disease (sub)types affect the performance of these CL methods, which provides insights into the feasibility of applying these methods in the clinical setting. In these experiments, it was found that while it is difficult to retain knowledge regarding staining intensity, the model appears to be less sensitive to hue changes within the tested range. Some results indicate the difficulty of learning tumor classification from DP images, but this study demonstrates the potential for continuous learning when adapting to changes in clinical histopathological image acquisition factors.

[0199] VIII. Further Considerations Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed by one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein and / or some or all of one or more processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein and / or some or all of one or more processes.

[0200] The terms and expressions used are for illustrative purposes only and not for limitation, and there is no intention to exclude any equivalents of the features shown and described, and it is recognized that various changes are possible within the scope of the claimed invention. Accordingly, although the invention described in the claims is specifically disclosed by embodiments and any features, modifications and variations of the concepts disclosed herein may be reclassified by those skilled in the art, and such modifications and variations are to be considered within the scope of the invention as defined by the appended claims.

[0201] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments provides those skilled in the art with a possible explanation for implementing various embodiments. It is understood that the functions and arrangements of the elements may be varied in various ways without departing from the spirit and scope as described in the appended claims.

[0202] In the following description, specific details are provided to provide a comprehensive understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as not to obscure the embodiments.

Claims

1. In a data processing system, obtaining a first annotated training set of the images for training a machine learning algorithm to perform detection, characterization, classification, or a combination thereof on some or all regions or objects within the images, wherein the first annotated training set of the images is within a first image domain; obtaining the first annotated training set of the images; dividing, by the data processing system, the first annotated training set of the images into mini-sets of the images, wherein each mini-set represents a distinct modeling sub-task and includes a limited number of examples; dividing the first annotated training set of the images into mini-sets of the images; training, by the data processing system, the machine learning algorithm in a first stage using the mini-sets of the images to generate a pre-adjusted machine learning model configured to perform detection, characterization, classification, or a combination thereof on some or all regions or objects within new images; labeling, by the data processing system, a limited number of images from a target data set to generate a second annotated training set of the images for training a machine learning algorithm to perform detection, characterization, classification, or a combination thereof on some or all regions or objects within the images, wherein the second annotated training set of the images is within a second image domain; labeling a limited number of images from the target data set; training, by the data processing system, the pre-adjusted machine learning model in a second stage using the second annotated training set of the images to generate a target machine learning model configured to perform detection, characterization, classification, or a combination thereof on some or all regions or objects within the new images, wherein some of the classes targeted in the first stage are part of or match some of the classes targeted in the second stage; training the pre-adjusted machine learning model; A computer-implemented method comprising the steps above.

2. The computer-implemented method according to claim 1, wherein the first annotated training set of the images is a digital pathology image including one or more types of cells.

3. The dividing step is When only one mini - set of images is available, for each separate modeling sub - task, select a subset of classes such that the subset is part of or matches some of the classes targeted in the second stage, select the limited number of examples based on the selected subset of classes, and When multiple mini - sets of images are available, for each separate modeling sub - task, either (i) mix examples from the multiple mini - sets of images, select a subset of classes such that the subset is part of or matches some of the classes targeted in the second stage, and select the limited number of examples from the mixed examples based on the selected subset of classes, or (ii) select one mini - set of images from the multiple mini - sets of images, select a subset of classes such that the subset is part of or matches some of the classes targeted in the second stage, and select the limited number of examples from the selected mini - set of images based on the selected subset of classes, The computer - implemented method according to claim 1 or 2, comprising.

4. The second stage is Applying the pre - adjusted machine - learning model to generate a feature - vector representation for each example in the second annotated training set of images; Generating one representation for each target class and combining the feature - vector representations from examples of the same class to use the one representation for each target class as a prototype for the target class; Generating a feature - vector representation for the remaining part of the unlabeled images or image regions from the target data set; Comparing each feature - vector representation from the unlabeled images with the prototype for the target class based on the distance between the feature - vector representation from the unlabeled images and the prototype for the target class; The computer - implemented method according to any one of claims 1 to 3, further comprising.

5. The first stage is An inner learning loop, in which the machine learning algorithm generates a loss in a validation set of images after model update, denoted as L-subtask-i for the i-th subtask, and updates model weights or parameters on one subtask for a predetermined or flexible number of epochs for adaptation to the target dataset, initializing the pre-adjusted machine learning model. An outer learning loop, which aims to search for a set of model initializations that generate the pre-adjusted machine learning model when used to update all subtasks, each with only a limited number of examples, by finding a model initialization that minimizes the sum of all losses, denoted as the sum of L-subtask-i, where i ranges from 1 to the number of subtasks calculated from the validation set of images of the subtasks for model initialization. The computer-implemented method according to any one of claims 1 to 3, further comprising.

6. The training in the first stage includes performing an iterative operation to learn a set of parameters to detect, characterize, classify, or a combination thereof, some or all regions or objects within the mini-set of images that maximize or minimize a cost function, where each iteration involves finding a set of parameters for the machine learning algorithm such that the value of the cost function using the set of parameters is greater than or less than the value of the cost function using another set of parameters in the previous iteration, and the cost function is constructed to measure the difference between predictions made for some or all of the regions or objects using the machine learning algorithm and the ground truth labels given to the mini-set of images.

7. Said training in the second stage involves performing detection, characterization, classification, or a combination thereof of some or all regions or objects within said second annotated training set of images that maximize or minimize a cost function, by executing an iterative operation to learn a set of parameters, each iteration involving finding a set of parameters for said pre - adjusted machine learning model such that the value of the cost function using said set of parameters is greater than or less than the value of the cost function using another set of parameters in the previous iteration, said cost function being constructed to measure the difference between predictions made for some or all of said regions or said objects using said pre - adjusted machine learning model and the ground - truth labels given in said second annotated training set of images. The computer - implemented method according to claim 1 or 2.

8. Identifying a digital pathology scenario; Selecting an adaptive continuous learning method for updating a target machine learning model based on said digital pathology scenario; Updating said target machine learning model based on said adaptive continuous learning method to generate an updated machine learning model; The computer - implemented method according to any one of claims 1 to 7, further comprising.

9. Said digital pathology scenario is a data - incremental scenario, a domain - incremental scenario, a class - incremental scenario, or a task - incremental scenario. The computer - implemented method according to claim 8.

10. Said adaptive continuous learning method is selected from the group comprising Elastic Weight Consolidation (EWC), Learning Without Forgetting (LWF), incremental classifier and representation learning (iCaRL), continuous prototype evaluation (CoPE), A - GEM, and parameter separation methods. The computer - implemented method according to claim 8.

11. The computer - implemented method according to claim 1 or 8, further comprising providing a target machine learning model and / or an updated machine learning model.

12. Said providing includes deploying said target machine learning model and / or said updated machine learning model to a digital pathology system. The computer - implemented method according to claim 11.

13. Receiving, by said data processing system, a new image; Inputting the new image into a target machine learning model or an updated machine learning model; Detecting, characterizing, classifying, or a combination thereof of some or all regions or objects within the new image by the target machine learning model or the updated machine learning model; Outputting an inference based on the detection, characterization, classification, or a combination thereof by the target machine learning model or the updated machine learning model; The computer-implemented method according to any one of claims 1 to 12, further comprising.

14. The computer-implemented method according to claim 13, further comprising determining, by a user, a diagnosis of a target related to the new image, wherein the diagnosis is determined based on the inference output by the target machine learning model or the updated machine learning model.

15. The computer-implemented method according to claim 14, further comprising treating the target by the user based on (i) an inference output by the target machine learning model or the updated machine learning model, and / or (ii) the diagnosis of the target.

16. Training the machine learning algorithm includes implementing a meta-learning principle to enable the first stage to generate the pre-adjusted machine learning model using the limited number of examples. The computer-implemented method according to claim 1.

17. Training the pre-adjusted machine learning model includes implementing a meta-learning principle to enable the second stage to generate a target machine learning model using the limited number of images. The computer-implemented method according to claim 1.

18. One or more data processors; A non-transitory computer-readable storage medium containing instructions, which, when executed on the one or more data processors, cause the one or more data processors to execute any of the steps of the method according to any one of claims 1 to 17. A non-transitory computer-readable storage medium; A system comprising.

19. A computer program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to execute any of the steps of the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Distributed and self-validating computer vision for dense object detection in digital images

    US20200242357A1