Federated learning system for training machine learning algorithms while maintaining patient privacy
The federated learning system addresses the challenge of data privacy in digital pathology by allowing local training of models on client devices, aggregating updates to improve the global model, thus enhancing image analysis accuracy without sharing patient data.
Patent Information
- Application Number
- JP2022547853
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-11
- Filing Date
- 2021-02-10
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-02-10
AI Technical Summary
The challenge in digital pathology is the difficulty in obtaining large amounts of training data due to privacy concerns, which hinders the effective training of machine learning classifiers for image analysis tasks such as tumor region identification and metastasis detection.
A federated learning system is employed, where a global model is distributed to client devices for local training on available data, allowing for model updates without sharing actual patient data. The updated models are then aggregated to generate an improved global model.
This approach enables the improvement of machine learning models for digital pathology without compromising patient data privacy, facilitating more accurate image analysis and classification tasks.
Smart Images

Figure 0007695256000001 
Figure 0007695256000002 
Figure 0007695256000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to digital pathology, and more particularly to machine learning techniques for federated learning.
Background Art
[0002] Digital pathology involves scanning a pathology slide (e.g., a histopathology or cytopathology glass slide) having tissue and / or cells into a digital image for evaluation. The tissue and / or cells within the digital image are then examined using digital pathology image analysis and / or can be interpreted by a pathologist for various reasons including disease diagnosis, evaluation of response to treatment, and development of drugs to fight disease. To examine the tissue and / or cells within the digital image (which is substantially transparent), the pathology slide can be prepared using a staining dye (e.g., an immunostain) that selectively binds to tissue and / or cell components. Immunohistochemistry (IHC) is a common use of immunostaining and involves a process of selectively identifying antigens (proteins) in cells of a tissue section by utilizing the principle of antibodies and other compounds (or substances) that specifically bind to antigens in a biological tissue. In some assays, the target antigen for a stain in a specimen may be referred to as a biomarker. Digital pathology image analysis can then be performed on the digital image of the stained tissue and / or cells to identify and quantify the staining for antigens (e.g., biomarkers indicative of tumor cells) in the biological tissue.
[0003] Machine learning techniques have shown great promise in digital pathology image analysis, such as tumor region identification, metastasis detection, and patient prognosis. For image classification and digital pathology image analysis, such as tumor region and metastasis detection, many computing systems equipped with machine learning techniques, including convolutional neural networks (CNNs), have been proposed. For example, a CNN can have a series of convolutional layers as hidden layers, and this network structure enables the extraction of representative features for object / image classification and digital pathology image analysis. In addition to object / image classification, machine learning techniques have also been implemented for image segmentation. Image segmentation is the process of dividing a digital image into multiple segments (sets of pixels, also known as image objects). A typical goal of segmentation is to simplify and / or transform the representation of an image to be more meaningful and easier to analyze. For example, image segmentation is often used to find objects such as tumors (or other tissue types) and boundaries (lines, curves, etc.) within an image. To perform image segmentation for large data (e.g., whole slide pathology images), the image is first divided into many small patches. A computing system equipped with machine learning techniques is trained to classify these patches, and all patches within the same class are combined into one segmented region. Subsequently, machine learning techniques can be further implemented to predict or classify the segmented regions (e.g., negative tumor cells or tumor cells without staining expression) based on the representative features associated with the segmented regions.
[0004] Various machine learning techniques require training data to establish ground truth for performing classification. In the medical field, due to privacy concerns and legal requirements, patient data is often difficult to obtain. Therefore, properly training a classifier can be a challenging problem. Federated learning is a decentralized machine learning technique that involves providing a base classifier to one or more client devices. Each device can then operate using the base classifier. When the classifier is utilized at each device, the user provides input regarding the output provided by the classifier. The user can provide input to each respective classifier based on the output, and each respective classifier can be updated according to the user input. The updated classifier can then be provided to update the base classifier. The updated classifier can then be distributed to the client devices. Thus, a federated learning system can be updated without the need to pass data between entities.
Summary of the Invention
[0005] In various embodiments, a computer-implemented method is provided.
[0006] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that, when executed on the one or more data processors, includes instructions to cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0007] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0008] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes.
[0009] Some embodiments of the present disclosure include a computer-implemented method for using an ensemble learning classifier. The method includes delivering, by a central server, a global model configured to classify pathological images to a plurality of client devices; receiving, by the central server, an updated model from at least one of the plurality of client devices, the updated model having been further trained at at least one of the plurality of client devices using a plurality of slide images and a plurality of corresponding annotations; aggregating, by the central server, the updated model with the global model to generate an updated global model; and delivering the updated global model to at least one of the plurality of client devices.
[0010] Some embodiments of the present disclosure include a computer-implemented method, wherein aggregating the updated model with the global model to generate an updated global model includes performing an averaging of at least one weight of the global model by at least one weight of the updated model.
[0011] Some embodiments of the present disclosure include a computer-implemented method that performing averaging comprises performing a weighted average of at least one weight of an updated model and at least one weight of a global model according to the number of a plurality of slide images used to further train the updated model and the total number of images used to train the global model.
[0012] Some embodiments of the present disclosure include a computer-implemented method that an annotation is provided by a user observing an output of a global model on a slide image, and the annotation includes a change to the output generated by the global model.
[0013] Some embodiments of the present disclosure further include receiving, by a centralized server, metadata related to a plurality of slide images, and aggregating further includes normalizing a model further trained according to the metadata, including a computer-implemented method.
[0014] Some embodiments of the present disclosure further include validating, by a centralized server, an improved performance of an updated global model against the global model using a validation dataset, including a computer-implemented method.
[0015] Some embodiments of the present disclosure include a computer-implemented method for using an federated learning classifier by a client device. The method includes receiving a global model configured to classify pathological images from a central server, receiving a stained tissue image, where the stained tissue image is divided into image patches, using the global model to perform image analysis on the image patches, training the global model using the image patches and at least one corresponding user annotation to generate an updated model, where the at least one corresponding user annotation includes a correction to a classification generated by the global model, sending the updated model to the central server, receiving the updated global model, and validating an improvement in performance of the updated global model using a client-specific validation dataset.
[0016] Some embodiments of the present disclosure include a computer-implemented method where a correction to a classification generated by the global model is a reclassification of at least one of a cell type, a tissue type, or a tissue boundary.
[0017] Some embodiments of the present disclosure include a computer-implemented method where the updated model does not include individual patient information.
[0018] Some embodiments of the present disclosure further include generating metadata related to a plurality of images and providing the metadata to a central server, and include a computer-implemented method.
[0019] Some embodiments of the present disclosure include a computer-implemented method where the metadata includes at least one of a region of a slide or tissue to which the image corresponds, a type of staining performed, a concentration of the staining, and a device used for the staining or scanning.
[0020] Some embodiments of the present disclosure include a computer-implemented method in which transmitting an updated model is performed after a threshold, number of iterations, length of time, or after the model has changed by a threshold amount.
[0021] Some embodiments of the present disclosure include a computer-implemented method for using an associative learning classifier in digital pathology. The method includes distributing a global model to a plurality of client devices by a central server, training the global model using a plurality of images from the plurality of client devices by the client devices to generate at least one further trained model, wherein one or more of the plurality of images includes at least one annotation, providing the at least one further trained model to the central server by the client devices, aggregating the at least one further trained model by the central server with the global model to generate an updated global model, and distributing the updated global model to the plurality of client devices.
[0022] Some embodiments of the present disclosure further include a computer-implemented method in which the client device generates metadata related to the plurality of images and provides the metadata to the central server, and the central server aggregates the at least one further trained model with the global model to generate an updated global model, further including normalizing the at least one further trained model according to the metadata.
[0023] Some embodiments of the present disclosure include a computer-implemented method in which the metadata includes at least one of the slide or tissue region to which the image corresponds, the type of staining performed, the concentration of the staining, and the device used for the staining or scanning.
[0024] Some embodiments of the present disclosure include a computer-implemented method further configured by a central server to verify the performance of an updated global model using a validation dataset.
[0025] Some embodiments of the present disclosure include a computer-implemented method further configured to roll back an update to the global model if the performance of the updated global model is lower than the global model.
[0026] Some embodiments of the present disclosure include a computer-implemented method including performing an averaging of at least one weight of the global model by at least one weight of the updated model, where aggregating the updated model with the global model generates an updated global model.
[0027] Some embodiments of the present disclosure include a computer-implemented method including performing a weighted average of at least one weight of the updated model and at least one weight of the global model according to the number of a plurality of slide images used to further train the updated model and the total number of images used to train the global model, where performing the averaging is used to further train the updated model.
[0028] Some embodiments of the present disclosure include a computer-implemented method where transmitting the updated model is performed after a threshold, number of iterations, length of time, or after the model has changed by a threshold amount.
[0029] The terms and expressions used are used as terms of explanation and not of limitation, and in the use of such terms and expressions, there is no intention to exclude equivalents or portions thereof of any feature shown and described, but it is recognized that various modifications are possible within the scope of the invention as set forth in the claims. Accordingly, while the invention as set forth in the claims is specifically disclosed by embodiments and any features, modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and it should be understood that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
Brief Description of the Drawings
[0030] This disclosure is described in conjunction with the following accompanying drawings:
[0031]
Figure 1
[0032]
Figure 2
[0033]
Figure 3
[0034]
Figure 4
[0035]
Figure 5
[0036]
Figure 6
[0037]
Figure 7
[0038]
Figure 8
[0039]
Figure 9
[0040] In the accompanying drawings, similar components and / or features can have the same reference labels. Further, various components of the same type can be distinguished by following the reference label with a dash and a second label that distinguishes the similar components. When only the first reference label is used in this specification, the description is applicable to any of the similar components having the same first reference label regardless of the second reference label.
Mode for Carrying Out the Invention
[0041] The present disclosure describes techniques for a digital pathology (DP) federated learning (FL) system. FL is a distributed machine learning approach in which multiple client devices collaboratively train a deep learning model (global model) for performing image analysis without sharing training data. The server is configured to distribute the global model to one or more clients. The server is configured to maintain, update, and redistribute the global model as part of an iterative process. In each iteration (or round), each client can receive the global model and perform DP image analysis on local data (e.g., patient data including pathology slides). The client can further train the global model using locally available data (e.g., patient data and user input). The updated model can be periodically sent from one or more clients to the server. The updated model may be incorporated into the global model to generate an updated global model. The updated global model can then be distributed to the clients. The iteration continues indefinitely or, for example, until training converges. In some examples, the received updated model may not be incorporated into the global model.
[0042] Immunohistochemical (IHC) slide staining is widely used in the study of different types of cells, such as cancerous cells and immune cells in biological tissues, because it can be used to identify proteins within cells of tissue sections. It is possible to evaluate IHC-stained cells of tissue sections at high magnification under a microscope and / or automatically analyze digital images of biological specimens using digital pathology algorithms. In many cases, in whole slide analysis, the evaluation of a stained biological specimen requires segmentation of the region of the stained biological specimen, including identification of target regions (e.g., positive and negative tumor cells) and exclusion of non-target regions (e.g., normal tissue or blank slide regions). In some examples, non-target regions to be excluded include biological substances or structures that may be very difficult to distinguish from other biological substances or structures in the target region, and thus are excluded from the evaluation of the biological specimen. As a result, in such cases, pathologists usually provide manual tumor annotation while excluding non-target regions. However, manual tumor annotation is error-prone, pathologist-biased, and time-consuming due to the large size of whole slide images at high magnification and the large amount of data to be processed.
[0043] Automated segmentation and classification of tumors and tumor cells can be difficult for various reasons. For example, tumors and tumor cells can vary widely among patients with respect to size, shape, and localization. This prohibits the use of strong prior information regarding shape and location identification, which is commonly used for robust image analysis in many other applications such as face recognition or navigation. As a result, conventional image analysis algorithms usually provide undesirable detection results (e.g., over-detection or misclassification) for these difficult regions.
[0044] To address these limitations and problems, a wide variety and large amount of training data are required. Considering privacy concerns regarding medical data, it has been found that it is difficult to obtain a large amount of training data. The technology of the FL DP system of the present embodiment includes the use of a machine learning architecture that enables the use of data at the client location for training without the need to send data to a centralized location. Therefore, the patient's personal information does not leave its original location, and privacy concerns are reduced. One exemplary embodiment of the present disclosure relates to a computer-implemented method for automatically performing image analysis on a pathology slide, including performing preprocessing, image analysis, and postprocessing. For example, the FL DP system can include one or more deep learning architectures that utilize FL to improve performance while not transferring underlying training data between entities. For example, the FL DP system can include a deep learning preprocessing system (for performing image segmentation, e.g., to remove or mask a specific region), a deep learning system for image processing (for identifying regions of an image having desired features), and / or a deep learning system for performing postprocessing (for performing further analysis using the identified regions of the image). Therefore, the FL DP system can include multiple models in each client device, and each model can utilize FL.
[0045] In some embodiments, the computer-implemented method can include the use of one or more models. The model can have, for example, a convolutional neural network (CNN) architecture or model that utilizes a two-dimensional segmentation model (e.g., a modified U-Net or other suitable architecture) to automatically detect and exclude biological structures or non-tumor cells before executing a standard image analysis algorithm for learning and recognizing a target region. Thereafter, post-analysis can be performed to provide or assist in a diagnosis or a further series of operations. The convolutional neural network architecture or model can be trained using pre-labeled images. As a result, a model (e.g., the architecture or model of a trained convolutional neural network) can be used to segment non-target regions, which can then be masked from whole slide analysis before, during, or after inputting the image to an image analysis algorithm. The image analysis model (e.g., a CNN) performs a classification task and outputs a tumor call for the target region. A post-processing model performs further classification based on the tumor call. Advantageously, this proposed architecture and technique can improve the accuracy of tumor cell classification by improving the models used at all stages of image analysis.
[0046] As used herein, when an operation is “based on” something, this means that the operation is at least partially based on at least a part of something.
[0047] As used herein, the terms "substantially", "approximately", and "about" are defined as being largely as specified, but not necessarily completely as specified (including being completely as specified), as would be understood by one of ordinary skill in the art. In any of the disclosed embodiments, the terms "substantially", "approximately", or "about" can be replaced with "within [percentage]" of the specified amount, where the percentages include 0.1, 1, 5, and 10%.
[0048] As used herein, the terms "sample", "biological sample", or "tissue sample" refer to any sample that contains biomolecules (such as proteins, peptides, nucleic acids, lipids, carbohydrates, or combinations thereof) obtained from any organism, including viruses. Other examples of organisms include mammals (such as veterinary animals like humans, cats, dogs, horses, cows, and pigs, as well as laboratory animals like mice, rats, and primates), insects, annelids, arachnids, marsupials, reptiles, amphibians, bacteria, and fungi. Biological samples include tissue samples (such as tissue sections or tissue needle biopsies), cell samples (such as cytological smears like Pap smears or blood smears, or samples of cells obtained by microdissection), or cell fractions, fragments, or organelles (obtained by lysing cells and separating their components by centrifugation or the like). Other examples of biological samples include blood, serum, urine, semen, feces, cerebrospinal fluid, interstitial fluid, mucosa, tears, sweat, pus, biopsy tissue (such as obtained by surgical biopsy or needle biopsy), nipple aspirate, earwax, milk, vaginal fluid, saliva, rinse fluid (such as cheek swabs), or any material containing biomolecules derived from an initial biological sample. In certain embodiments, the term "biological sample" as used herein refers to a sample (such as a homogenized or liquefied sample) prepared from a tumor or a part thereof obtained from a subject.
[0049] As used herein, the term "biological material or structure" refers to a natural material or structure that includes all or a part of a biological structure (e.g., cell nucleus, cell membrane, cytoplasm, chromosome, DNA, cell, cell mass, etc.).
[0050] As used herein, the term "non-target region" refers to a region of an image that has image data not intended to be evaluated in an image analysis process. The non-target region can include, for example, an unorganized region of an image corresponding to a substrate such as glass without a sample when only white light from the imaging source is present. The non-target region can additionally or alternatively include a tissue region of an image corresponding to one or more biological materials or structures that are not intended to be analyzed in the image analysis process or are difficult to distinguish from biological materials or structures within the target region (e.g., lymphoid aggregates).
[0051] As used herein, the term "target region" refers to a region of an image that includes image data intended to be evaluated in an image analysis process. The target region includes any region such as a tissue region of an image intended to be analyzed in the image analysis process.
[0052] As used herein, the term "tile" or "tile image" refers to a single image corresponding to a part of the entire image or the entire slide. In some embodiments, a "tile" or "tile image" refers to a region of a whole-slide scan or a region of interest having (x,y) pixel dimensions (e.g., 1000 pixels × 1000 pixels). For example, considering an entire image divided into M columns of tiles and N rows of tiles, each tile within the M×N mosaic includes a part of the entire image, i.e., the tile at position M1, N1 includes a first part of the image, the tile at position M3, N4 includes a second part of the image, and the first part and the second part are different. In some embodiments, the tiles can each have the same dimensions (pixel size × pixel size).
[0053] As used herein, the term "patch" or "image patch" refers to a container of pixels corresponding to a part of a tiled image, the entire image, or the entire slide. In some embodiments, a "patch" or "image patch" refers to a region of interest of a tiled image having (x,y) pixel dimensions (e.g., 256 pixels × 256 pixels). For example, a 1000 pixel × 1000 pixel tiled image divided into 100 pixel × 100 pixel patches will contain 100 patches (each patch containing 1000 pixels). In other examples, the patches may overlap.
[0054] In some embodiments, a federated learning (FL) system for digital pathology (DP) is utilized to generate and distribute a global model (e.g., an aggregated global model) without exchanging confidential or identifying data (e.g., patient data) between clients and / or a central system (e.g., a server). The server is configured to maintain and distribute the global model in an iterative process when updated models are received from clients. FIG. 1 depicts an exemplary FL DP system 100 that includes one or more servers 110 configured to maintain and distribute one or more global models 112, 114. The server 110 communicates with one or more client systems 120, 130, 140, each of which can include various DP devices such as workstations 122, 132, 142, microscopes 124, 134, 144, digital slide scanners 126, 136, 146, and any other necessary equipment as understood by those skilled in the art. Each of the client systems can utilize one or more local models 128, 138, 148, 150 based on the global models 112, 114. The client systems 120, 130, 140 can be further utilized to train the local models 128, 138, 148, 150. For example, the client systems 120, 130, 140 can receive patient data, classify the patient data using the local models 128, 138, 148, 150, receive user input regarding the classified patient data (e.g., from a pathologist or other medical professional utilizing a graphical user interface that displays the classified data), and update the local models 128, 138, 148, 150 based on the user input (e.g., each client retrains the global model using a local training dataset). In various embodiments, the client devices are configured to periodically provide their local models 128, 138, 148, 150 to the central server 110.Next, the central server 110 can update the global models 112, 114 using the local models 128, 138, 148, 150 (e.g., by updating the weights in the global models), and distribute the updated global models 112, 114 to the client systems 120, 130, 140.
[0055] In some embodiments, after each iteration, the performance of each of the updated local models 128, 138, 148, 150 can be verified using a validation dataset. If it is determined that the local models 128, 138, 148, 150 provide improved performance with respect to the validation dataset, the local models can be incorporated into the global models 112, 114. The performance of the updated global models 112, 114 may also be verified using the validation dataset. If the global models 112, 114 are improved, the updated global models 112, 114 may be distributed to all or some of the client devices 120, 130, 140. In some embodiments, the clients can choose not to share their updated local models 128, 138, 148, 150 but still receive the updated global models 112, 114. In other embodiments, the clients can choose to share their local models 128, 138, 148, 150 but not receive the updated global models 112, 114. In other embodiments, the clients can choose not to share their updated local models 128, 138, 148, 150 and not receive the updated global models 112, 114. Thus, the models generated at the client sites are not controlled by the central server 110 and are shared with the central server 110 based on the discretion of the clients. Each client can have an independent validation dataset and can use the validation dataset to inspect the performance of the model based on their quality criteria. Based on this verification, the client can decide whether to deploy the global models 112, 114.
[0056] In some embodiments, after each iteration, the performance of each of the updated local models 128, 138, 148, 150 can be confirmed using a validation dataset. If it is determined that the local models 128, 138, 148, 150 provide improved performance with respect to the validation dataset, the local models can be incorporated into the global models 112, 114. The performance of the updated global models 112, 114 may also be verified using the validation dataset. If the global models 112, 114 are improved, the updated global models 112, 114 may be distributed to all or some of the client devices 120, 130, 140. In some embodiments, the clients can choose not to share their updated local models 128, 138, 148, 150 but still receive the updated global models 112, 114. In other embodiments, the clients can choose to share their local models 128, 138, 148, 150 but not receive the updated global models 112, 114. In other embodiments, the clients can choose not to share their updated local models 128, 138, 148, 150 and not receive the updated global models 112, 114. Thus, the models generated at the client site are not controlled by the central server 110 and are shared with the central server 110 based on the discretion of the client. Each client can have an independent validation dataset and can use the validation dataset to inspect the performance of the model based on their quality criteria. Based on this verification, the client can determine whether to deploy the updated global models 112, 114.
[0057] Figure 2 shows a block diagram showing a computing environment 200 for non-tumor segmentation and image analysis using deep convolutional neural networks according to various embodiments. The computing environment 200 can include an analysis system 205 for training and executing a prediction model, such as a two-dimensional CNN model. More specifically, the analysis system 205 can include training subsystems 210a-n (where "a" and "n" represent any natural numbers) for constructing and training respective prediction models 215a-n (which can be referred to individually as prediction model 215 or collectively as prediction model 215) used by other components of the computing environment 200. The prediction model 215 can be a machine learning ("ML") or deep learning ("DL") model, such as a deep convolutional neural network (CNN) like a U-Net neural network, a first-movement neural network, a residual neural network ("Resnet"), or a recurrent neural network, such as a long short-term memory ("LSTM") model or a gated recurrent unit ("GRU") model. The prediction model 215 can also be any other suitable ML model trained to segment non-target regions (e.g., lymphoid aggregate regions), segment target regions, or provide image analysis of target regions, such as a two-dimensional CNN ("2DCNN"), dynamic time warping ("DTW") technology, a hidden Markov model ("HMM"), etc., or a combination of one or more of such technologies, such as CNN-HMM or MCNN (multi-scale convolutional neural network). The computing environment 200 can use the same type or different types of prediction models trained to segment non-target regions, segment target regions, or provide image analysis of target regions. For example, the computing environment 200 can include a first prediction model (e.g., U-Net) for segmenting non-target regions (e.g., lymphoid aggregation regions, necrosis regions, or any other suitable regions).The computing environment 200 can also include a second prediction model (e.g., 2DCNN) for segmenting a target region (e.g., a region of tumor cells). The computing environment 200 can also include a third model (e.g., CNN) for image analysis of the target region. The computing environment 200 can also include a fourth model (e.g., HMM) for diagnosing a disease for treating or prognosticating a subject such as a patient. In other examples according to the present disclosure, even other types of prediction models may be implemented. Furthermore, multiple models can be used to classify different cell types and regions.
[0058] In various embodiments, each prediction model 215a - n corresponding to classifier subsystems 210a - n can be based on global models 112, 114 provided by server 110. In various embodiments, each prediction model 215a - n corresponding to classifier subsystems 210a - n is further separately trained based on one or more sets of input image elements 220a - n. In some embodiments, each of the input image elements 220a - n includes image data from one or more scanned slides. Each of the input image elements 220a - n can correspond to a single specimen from which underlying image data corresponding to the image was collected and / or image data from one day. The image data can include the image as well as any information regarding the imaging platform on which the image was generated. For example, tissue sections may need to be stained by application of a staining assay that includes one or more different biomarkers associated with chromogenic staining for brightfield imaging or fluorophores for fluorescence imaging. The staining assay can use chromogenic staining for brightfield imaging, organic fluorophores, quantum dots, or organic fluorophores together with quantum dots for fluorescence imaging, or any other combination of stains, biomarkers, and observation or imaging devices. Further, a typical tissue section is processed on an automated staining / assay platform that applies the staining assay to the tissue section to obtain a stained sample. There are various commercially available products suitable for use as staining / assay platforms, and one example is the VENTANA SYMPHONY product of assignee Ventana Medical Systems, Inc. The stained tissue section can be supplied, for example, to an imaging system on a microscope or a whole slide scanner having a microscope and / or imaging components, and one example is the VENTANA iScan Coreo product of assignee Ventana Medical Systems, Inc. Multiplex tissue slides can be scanned with an equivalent multiplex slide scanner system.Additional information provided by the imaging system can include any information regarding the staining platform, including the concentration of the chemicals used in the staining, the reaction time of the chemicals applied to the tissue in the staining, and / or the pre - analysis conditions of the tissue, such as the age of the tissue, the fixation method, the period, how the sections were embedded, cut, etc.
[0059] The input image elements 220a - n can include one or more training input image elements 220a - d, validation input image elements 220e - g, and unlabeled input image elements 220h - n. It should be understood that the input image elements 220a - n corresponding to the training group, validation group, and unlabeled group do not need to be accessed simultaneously. For example, a set of training and validation input image elements 220a - n may first be accessed and used to further train the prediction model 215, and the unlabeled input image elements may subsequently be accessed or received (e.g., at one or more subsequent times) and used by the further - trained prediction model 215 to provide a desired output (e.g., segmentation of non - target regions). In some examples, the prediction models 215a - n are trained using supervised training, and each of the training input image elements 220a - d and optional validation input image elements 220e - g is associated with one or more labels 225 that identify non - target regions, the "correct" interpretation of target regions, and the identification of various biological substances and structures within the training input image elements 220a - d and validation input image elements 220e - g. The labels may alternatively or additionally be used to classify the corresponding training input image elements 220a - d and validation input image elements 220e - g or the pixels therein with respect to the presence and / or interpretation of staining related to normal or abnormal biological structures (e.g., tumor cells). In a particular example, alternatively or additionally, labels are used to classify the corresponding training input image elements 220a - d and validation input image elements 220e - g at the time corresponding to when the underlying image was taken or at a subsequent time point (e.g., this is a predetermined period following the time when the image was taken).
[0060] In some embodiments, classifier subsystems 210a - n include a feature extractor 230, a parameter data store 235, a classifier 240, and a trainer 245, which are used collectively to train a prediction model 215 based on training data (e.g., training input image elements 220a - d) and to optimize the parameters of the prediction model 215 during supervised or unsupervised training. In some examples, the training process includes iterative operations to find a set of parameters of the prediction model 215 that minimize a loss function of the prediction model 215. Each iteration can include finding a set of parameters of the prediction model 215 such that the value of the loss function using the set of parameters is less than the value of the loss function using another set of parameters in the previous iteration. The loss function can be constructed to measure the difference between the output predicted using the prediction model 215 and the labels 225 included in the training data. Once the set of parameters is identified, the prediction model 215 is trained and can be utilized for segmentation and / or prediction as designed.
[0061] In some embodiments, classifier subsystems 210a-n access training data from training input image elements 220a-d in an input layer. Feature extractor 230 can preprocess the training data to extract relevant features (e.g., edges, colors, textures, or any other suitable relevant features) detected in specific portions of training input image elements 220a-d. Classifier 240 receives the extracted features and converts the features into one or more output metrics that segment non-target or target regions according to weights associated with a set of hidden layers within one or more prediction models 215, provides image analysis, provides a diagnosis of a disease for treatment or prognosis of a subject such as a patient, or provides a combination thereof. Trainer 245 can train feature extractor 230 and / or classifier 240 by facilitating learning of one or more parameters using training data corresponding to training input image elements 220a-d. For example, trainer 245 can use backpropagation techniques to facilitate learning of weights associated with a set of hidden layers of prediction model 215 used by classifier 240. Backpropagation can cumulatively update the parameters of the hidden layers, for example, using a Stochastic Gradient Descent (SGD) algorithm. The learned parameters can include, for example, weights, biases, and / or other hidden layer-related parameters, which can be stored in parameter data store 235.
[0062] Individual or ensembles of trained prediction models can be deployed to process unlabeled input image elements 220h-n to segment non-target or target regions, provide image analysis, provide a diagnosis of a disease for the treatment or prognosis of a subject such as a patient, or provide combinations thereof. More specifically, a trained version of the feature extractor 230 can generate a feature representation of unlabeled input image elements that can then be processed by a trained version of the classifier 240. In some embodiments, image features can be extracted from unlabeled input image elements 220h-n based on one or more convolutional blocks, convolutional layers, residual blocks, or pyramid layers that leverage the expansion of the prediction model 215 within the classifier subsystems 210a-n. The features can be compiled into a feature representation such as a feature vector of the image. The prediction model 215 can be trained to learn the feature type based on the classification and subsequent adjustment of the parameters in the hidden layer including the fully connected layer of the prediction model 215.
[0063] In some embodiments, the image features extracted by a convolutional block, convolutional layer, residual block, or pyramid layer include a feature map that is a matrix of values representing one or more portions of a specimen slide on which one or more image processing operations (e.g., edge detection, sharpening of image resolution) have been performed. These feature maps can be flattened for processing by a fully connected layer of a prediction model 215 that outputs one or more metrics corresponding to a non-target region mask, target region mask, or current or future prediction regarding the specimen slide. For example, input image elements can be supplied to an input layer of the prediction model 215. The input layer can include nodes corresponding to specific pixels. The first hidden layer can include a set of hidden nodes, each of which is connected to a plurality of input layer nodes. Nodes within subsequent hidden layers can similarly be configured to receive information corresponding to a plurality of pixels. Thus, the hidden layers can be configured to learn to detect features that extend across a plurality of pixels. Each of the one or more hidden layers can include a convolutional block, convolutional layer, residual block, or pyramid layer. The prediction model 215 can further include one or more fully connected layers (e.g., a softmax layer).
[0064] At least a portion of the training input image elements 220a - d, the validation input image elements 220e - g, and / or the unlabeled input image elements 220h - n can be elements of the analysis system 205, but can include data obtained directly or indirectly from a source that is not necessarily so, or may be derived from them. In some embodiments, the computing environment 200 includes an imaging device 250 that images a sample to obtain image data such as a multi - channel image (e.g., a multi - channel fluorescence or bright - field image) having several (e.g., between 10 and 16) channels. The imaging device 250 includes, without limitation, a camera (e.g., an analog camera, a digital camera, etc.), an optical system (e.g., one or more lenses, a sensor focus lens group, a microscope objective lens, etc.), an imaging sensor (e.g., a charge - coupled device (CCD), a complementary metal - oxide - semiconductor (CMOS) imaging sensor, etc.), photographic film, etc. In digital embodiments, the image - capturing device can include a plurality of lenses that cooperate to demonstrate on - the - fly focusing. An imaging sensor, e.g., a CCD sensor, can capture a digital image of a specimen. In some embodiments, the imaging device 250 is a bright - field imaging system, a multispectral imaging (MSI) system, or a fluorescence microscopy system. The imaging device 250 can capture images using invisible electromagnetic radiation (e.g., UV light) or other imaging techniques. For example, the imaging device 250 can include a microscope and a camera configured to capture an image magnified by the microscope. The image data received by the image analysis system 205 may be the same as the raw image data captured by the imaging device 250 and / or may be derived from the raw image data.
[0065] In some examples, the labels 225 associated with the training input image elements 220a - d and / or the validation input image elements 220e - g may be received, or may be derived from data received from one or more provider systems 255 that are each associated with a particular subject (e.g., a doctor, nurse, hospital, pharmacist, etc.). The received data can include, for example, one or more medical records corresponding to a particular subject. The medical records can indicate, for example, whether the subject had a tumor and / or the stage of progression of the subject's tumor (e.g., along a standard scale and / or by identifying a metric such as total metabolic tumor volume (TMTV)) with respect to a period corresponding to the time when one or more input image elements related to the subject were collected or a defined period thereafter. The received data can further include pixels of the location of a tumor or tumor cells within one or more input image elements related to the subject. Thus, the medical records can include or be used to identify one or more labels for each training / validation input image element 220a - g. The medical records can further indicate each of one or more treatments (e.g., medications) the subject has received and the period during which the subject has received the treatment. In some examples, the images or scans input to one or more classifier subsystems are received from the provider system 255. For example, the provider system 255 can receive an image from the imaging device 250 and then transmit the image or scan to the analysis system 205 (e.g., along with a subject identifier and one or more labels).
[0066] In some embodiments, data received or collected at one or more of the imaging devices 250 may be aggregated with data received or collected at one or more of the provider systems 255. For example, the analysis system 205 can identify corresponding or identical identifiers of the subject and / or period to associate the image data received from the imaging device 250 with the label data received from the provider system 255. The analysis system 205 can further process the data using metadata or automated image analysis to determine which classifier subsystem to supply a particular data component to. For example, the image data received from the imaging device 250 can correspond to whole slides or multiple regions of a slide or tissue. Metadata, automated alignment, and / or image processing can indicate, for each image, which region of the slide or tissue the image corresponds to, the type of stain performed, the concentration of the stain used, the laboratory that performed the stain, the timestamp, the type of scanner used, or any other suitable data as understood by one of ordinary skill in the art. Automated alignment and / or image processing can include detecting whether the image has image characteristics corresponding to a slide substrate or has biological structures and / or shapes associated with specific cells such as white blood cells. The label-related data received from the provider system 255 can be slide-specific, region-specific, or subject-specific. If the label-related data is slide-specific or region-specific, metadata or automated analysis (e.g., using natural language processing or text analysis) can be used to identify which region the particular label-related data corresponds to. If the label-related data is subject-specific, the same label data (for a given subject) can be supplied to each classifier subsystem 210a-n during training.
[0067] In some embodiments, the computing environment 200 can further include a user device 260 that can be associated with a user who requests and / or coordinates the execution of one or more iterations of the analysis system 205 (e.g., each iteration corresponds to one execution of the model and / or one generation of the output of the model). The user can correspond to a physician, a research practitioner (e.g., associated with a clinical trial), a patient, a medical professional, and the like. Thus, in some examples, it will be understood that the provider system 255 can include and / or function as the user device 260. Each iteration can be associated with a particular subject (e.g., a person) who may be different from the user (although this is not required). The iteration request can include and / or be accompanied by information about the particular subject (e.g., the name or other identifier of the subject such as an un-identified patient identifier). The iteration request can include an identifier of one or more other systems that collect data such as input image data corresponding to the subject. In some examples, the communication from the user device 260 includes an identifier for each subject represented in a particular set of subjects in response to a request to perform an iteration for each subject in the set.
[0068] Upon receiving a request, the analysis system 205 can send the request for unlabeled input image elements (e.g., including a subject identifier) to one or more corresponding imaging systems 250 and / or provider systems 255. The trained prediction model 215 can then process the unlabeled input image elements to segment non-target or target regions, provide image analysis, provide a diagnosis of a disease for treatment or prognosis of a subject such as a patient, or provide a combination thereof. The results for each identified subject can include or be based on the segmentation and / or one or more output metrics from the trained prediction model 215 deployed by the classifier subsystems 110a-n. For example, the segmentation and / or one or more output metrics can include or be based on the output generated by the fully connected layers of one or more CNNs. In some examples, such output may be further processed using (e.g.) a softmax function. Further, the output and / or the further processed output can then be aggregated using an aggregation technique (e.g., random forest aggregation) to generate one or more subject-specific metrics. One or more results (e.g., including plane-specific output and / or one or more subject-specific outputs and / or processed versions thereof) can be sent to and / or utilized by the user device 260. In some examples, some or all of the communication between the analysis system 205 and the user device 260 is performed via a website. It will be appreciated that the CNN system 205 can gate access to results, data, and / or processing resources based on authentication analysis.
[0069] Although not explicitly shown, it will be understood that the computing environment 200 can further include a developer device associated with a developer. Communications from the developer device can indicate to each prediction model 215 within the analysis system 205 which type of input image elements should be used, the number of neural networks that should be used, the configuration of each neural network including the number of hidden layers and hyperparameters, and how data requests should be formatted and / or which training data should be used (e.g., and how to access the training data).
[0070] FIG. 3 shows an exemplary schematic diagram 300 representing a model architecture for non-target region segmentation (e.g., part of the analysis system 205 described with respect to FIG. 2) according to various embodiments. The model architecture can include an image acquisition module 310 for generating or obtaining an input image including single image data (e.g., images each having a single stain) and / or multi-image data (e.g., images having multiple stains), an optional image annotation module 315 for electronically annotating a portion of the input image for further analysis, such as a portion indicating a tumor region or an immune cell region, and an optional unmixing module 320 for generating image channel images corresponding to one or more stain channels present in the multi-image, and can include a preprocessing stage 305. The model architecture can further include a processing stage 325 including an image analysis module 330 for detecting and / or classifying biological substances or structures including cells or nuclei (e.g., tumor cells, stromal cells, lymphocytes, etc.) based on features within the input image (e.g., within a hematoxylin and eosin stained image, a biomarker image, or an unmixed image channel image).
[0071] The model architecture can further comprise a post-processing stage 335 comprising any scoring module 340 for deriving expression predictions and / or scores for each biomarker in each of the identified regions or biological structures, and any metric generation module 345 for deriving a metric that describes the variability between the derived expression predictions and / or scores in different regions or biological structures and optionally providing a diagnosis of a disease for treatment or prognosis of a subject such as a patient. The model architecture can further comprise a segmentation and masking module 350 for segmenting regions or biological structures such as lymphocyte aggregates or clusters of tumor cells in the input image and generating a mask based on the segmented regions or biological structures, and any alignment module 355 for mapping the identified regions or biological structures (such as tumor cells or immune cells) from a first image or first set of images in the input image to at least one additional image or a plurality of additional images. The segmentation and masking module 350 and any alignment module 355 may be implemented within the pre-processing stage 305, the processing stage 325, the post-processing stage 335, or any combination thereof.
[0072] In some embodiments, the image acquisition module 310 generates or acquires an image or image data of a biological sample having one or more stains (e.g., the image may be a single image or a multiple image). In some embodiments, the generated or acquired image is an RGB image or a multispectral image. In some embodiments, the generated or acquired image is stored in a memory device. The image or image data (used interchangeably herein) can be generated or acquired, for example, in real time using an imaging device (e.g., the imaging device 250 described with respect to FIG. 2). In some embodiments, the image is generated or acquired from a microscope or other device capable of imaging image data of a microscope slide holding a specimen, as described herein. In some embodiments, the image is generated or acquired using a 2D scanner such as one capable of scanning image tiles. Alternatively, the image may be an image previously generated (e.g., scanned), stored in a memory device (or, in that case, acquired from a server via a communication network).
[0073] In some embodiments, the image acquisition module 310 is used to select portions of a biological sample from which one or more images or image data are to be acquired. For example, the image acquisition module 310 can receive an identified region of interest or field of view (FOV). In some embodiments, the region of interest is identified by a user of the system of the present disclosure or another system communicatively coupled to the system of the present disclosure. Alternatively, in other embodiments, the image acquisition module 305 obtains a region or location or identification of interest from a storage / memory device. In some embodiments, the image acquisition module 310 automatically generates a field of view or region of interest (ROI) via, for example, the method described in PCT / EP2015 / 062015, the content of which is hereby incorporated by reference in its entirety for any purpose. In some embodiments, the ROI is automatically determined by the image acquisition module 305 based on some predetermined criteria or characteristics within the image or of the image (e.g., for a biological sample stained with three or more stains, identifying a region of the image that contains only two of the stains). In some examples, the image acquisition module 310 outputs the ROI.
[0074] In some embodiments, the image acquisition module 310 generates or acquires at least two images as input. In some embodiments, the images generated or acquired as input are obtained from consecutive tissue sections, e.g., consecutive sections from the same tissue sample. Generally, at least two images received as input each contain signals corresponding to stains (including chromogens, fluorophores, quantum dots, etc.). In some embodiments, one of the images is stained with at least one primary stain (hematoxylin or eosin (H&E)), and another one of the images is stained with at least one of an IHC assay or an in-situ hybridization (ISH) assay for identifying a specific biomarker. In some embodiments, one of the images is stained with both hematoxylin and eosin, and another one of the images is stained with at least one of an IHC assay or an ISH assay for identifying a specific biomarker. In some embodiments, the input images are multiplex images, e.g., stained for multiple different markers with a multiplex assay according to methods known to those skilled in the art.
[0075] In some embodiments, the generated or acquired images are optionally annotated by a user (e.g., a medical professional such as a pathologist) for image analysis using the image annotation module 315. In some embodiments, the user identifies portions (e.g., sub-regions) of the image that are suitable for further analysis. The targeted or non-targeted regions (e.g., tumor regions or immune regions) that are annotated to generate a slide score may be either the entire tissue region or a specified set of regions on the digital slide. For example, in some embodiments, the identified portion represents an overexpressive tumor region of a particular biomarker, e.g., a particular IHC marker. In other embodiments, the user, medical professional, or pathologist can annotate lymphocyte aggregate regions within the digital slide. In some embodiments, the annotated representative fields can be selected by a pathologist to reflect biomarker expression used by the pathologist for overall slide interpretation. The annotation may be drawn using an annotation tool provided in a viewer application (e.g., VENTANA VIRTUOSO software), and the annotation may be drawn at any magnification or resolution. Alternatively or additionally, image analysis operations are used to automatically detect targeted and non-targeted regions or other regions using automated image analysis operations such as segmentation, thresholding, edge detection, etc., and fields of view (FOVs - portions of the image having a predetermined size and / or shape) automatically generated based on the detected regions. In some embodiments, user annotations can be utilized to further train one or more of the models.
[0076] In some embodiments, the generated or acquired image may be a multiplexed image, i.e., the received image is an image of a biological sample stained with two or more dyes. In these embodiments, prior to further processing, each multiplexed image is first unmixed into its constituent channels using, for example, an unmixing module 320, and each unmixed channel corresponds to a particular stain or signal. In some embodiments, the unmixed images (often referred to as "channel images" or "image channel images") can be used as inputs to each of the modules described herein. For example, a model architecture can be implemented to evaluate marker-to-marker heterogeneity (an indicator of the amount of protein expression heterogeneity of biomarkers in a sample) determined using a first H&E image, a second multiplexed image stained for a plurality of differentiation marker clusters (such as CD3, CD8, etc.), and a plurality of single images each stained for a particular biomarker (such as ER, PR, Ki67, etc.). In this example, the multiplexed image is first unmixed into its constituent channel images, and those channel images can be used along with the H&E image and the plurality of single images to determine the heterogeneity between markers.
[0077] Subsequent to image acquisition and / or unmixing, the input image or the unmixed image channel images are processed by an image analysis algorithm provided by the image analysis module 330 to identify and classify cells and / or nuclei. The procedures and algorithms described herein can be adapted to identify and classify various types of cells or cell nuclei based on features within the input image, including the identification and classification of tumor cells, non-tumor cells, stromal cells, lymphocytes, non-target staining, etc. Those skilled in the art should understand that the nuclei, cytoplasm, and membranes of cells have different characteristics, and different stained tissue samples can exhibit different biological characteristics. Specifically, those skilled in the art should understand that certain cell surface receptors can have staining patterns that are localized to the membrane or to the cytoplasm. Thus, a "membrane" staining pattern is analytically distinct from a "cytoplasm" staining pattern. Similarly, a "cytoplasm" staining pattern and a "nucleus" staining pattern are analytically distinct. Each of these distinct staining patterns can be used as a feature for identifying cells and / or nuclei. For example, stromal cells can be strongly stained by FAP, whereas tumor epithelial cells can be strongly stained by EpCAM, while cytokeratin can be stained by panCK. Thus, by utilizing different stainings, different cell types can be identified and distinguished during image analysis to provide a classification solution.
[0078] Methods for identifying, classifying, and / or scoring nuclei, cell membranes, and cell cytoplasm in images of biological samples having one or more stains are described in U.S. Patent No. 7,760,927, the “’927 Patent,” the content of which is incorporated herein by reference in its entirety for all purposes. For example, the ’927 Patent contemplates considering a first color plane of a plurality of pixels in the foreground of an input image for simultaneous identification of cytoplasm and cell membrane pixels, the consideration being processed to remove background portions of the input image and to remove reversely stained components of the input image, determining a threshold level between cytoplasm and cell membrane pixels in the foreground of the digital image, and using the determined threshold level to simultaneously determine whether a selected pixel is a cytoplasm pixel, a cell membrane pixel, or a transition pixel in the digital image, along with the selected pixel and its eight adjacent pixels from the foreground. The ’927 Patent describes an automated method for simultaneously identifying a plurality of pixels in an input image of a biological tissue stained with a biomarker. In some embodiments, tumor nuclei are automatically identified by first identifying candidate nuclei and then automatically distinguishing tumor nuclei from non-tumor nuclei. Many methods for identifying candidate nuclei in images of tissue are known in the art. For example, automated candidate nucleus detection can be performed by applying Parvin's method based on radial symmetry, such as in a hematoxylin image channel or biomarker image channel after deconvolution (see Parvin, Bahram et al., “Iterative voting for inference of structural saliency and characterization of subcellular events,” Image Processing, IEEE Transactions on 16.3 (2007): 615-623, the disclosure of which is incorporated herein by reference in its entirety).
[0079] For example, in some embodiments, the image acquired as input is processed to detect the center (seed) of the nucleus and / or to segment the nucleus, etc. For example, instructions for detecting the nucleus center based on radial symmetry voting using the technology of Parvin (above) can be provided and executed. In some embodiments, the nucleus is detected using radial symmetry to detect the center of the nucleus, and then the nucleus is classified based on the intensity of the staining around the cell center. In some embodiments, as described in International Publication No. WO 2014 / 140085 pamphlet, a co-pending patent application by the same applicant, the entire content of which is incorporated herein by reference in its entirety for all purposes, a nucleus detection operation based on radial symmetry is used. For example, the size of the image can be calculated within the image, and one or more votes at each pixel are accumulated by adding up the total size within the selected region. Mean shift clustering can be used to find the local center within the region, and the local center represents the location of the actual nucleus. Nucleus detection based on radial symmetry voting is performed on color image intensity data, explicitly utilizing the prior domain knowledge that the nucleus is an elliptical blob with various sizes and eccentricities. To achieve this, along with the color intensity of the input image, image gradient information is used in the radial symmetry voting and combined with an adaptive segmentation process to accurately detect and locate the cell nucleus. As used herein, "gradient" is, for example, the pixel intensity gradient calculated for the particular pixel by taking into account the intensity value gradients of the set of pixels surrounding the particular pixel. Each gradient can have a specific "direction" with respect to a coordinate system whose x-axis and y-axis are defined by two orthogonal edges of the digital image. For example, nucleus seed detection includes defining a seed as a point that is assumed to be inside the cell nucleus and functions as a starting point for identifying the location of the cell nucleus. The first step is to detect the seed points associated with each cell nucleus using a very robust approach based on radial symmetry and detect an elliptical blob that is a structure similar to the cell nucleus. The radial symmetry approach operates on the gradient image using a kernel-based voting procedure.The voting response matrix is created by processing each pixel that accumulates votes through a voting kernel. The kernel is based on the gradient direction calculated at that particular pixel, the expected range of the minimum and maximum kernel sizes, and the voting kernel angle (usually in the range of [p / 4, p / 8]). In the resulting voting space, the maximum positions with voting values higher than a predefined threshold are saved as seed points. Unrelated seeds can be discarded later during subsequent segmentation or classification processes. Other methods are described in U.S. Patent Application Publication No. 2017 / 0140246, the disclosure of which is incorporated herein by reference.
[0080] After the candidate kernels are identified, the candidate kernels are further analyzed to distinguish tumor kernels from other candidate kernels. The other candidate kernels can be further classified (e.g., by identifying lymphocyte nuclei and stromal nuclei). In some embodiments, as further described herein, a learned supervised classifier is applied to identify tumor kernels. For example, a learned supervised classifier is trained on nuclear features to identify tumor kernels and is then applied to classify nuclear candidates in the test image as either tumor or non-tumor kernels. Optionally, the learned supervised classifier may be further trained to distinguish different classes of non-tumor kernels such as lymphocyte nuclei and stromal nuclei. In some embodiments, the learned supervised classifier used to identify tumor kernels is a random forest classifier. For example, a random forest classifier can be trained by: (i) creating a training set of tumor and non-tumor kernels, (ii) extracting the features of each kernel, and (iii) training the random forest classifier to distinguish between tumor and non-tumor kernels based on the extracted features. The trained random forest classifier can then be applied to classify the nuclei in the test image as tumor and non-tumor kernels. Optionally, the random forest classifier may be further trained to distinguish different classes of non-tumor kernels such as lymphocyte nuclei and stromal nuclei.
[0081] The nuclei can be identified using other techniques known to those skilled in the art. For example, the size of an image can be calculated from one particular image channel of the FI&E or IHC image, and each pixel around the specified size can be assigned a vote count based on the sum of the sizes within the region around the pixel. Alternatively, an average shift clustering operation can be performed to find local centers within the vote image that represent the actual positions of the nuclei. In other embodiments, nucleus segmentation can be used to segment the entire nucleus based on the centers of currently known nuclei via morphological operations and local thresholding. In yet other embodiments, model-based segmentation can be utilized to detect nuclei (i.e., learn a shape model of the nuclei from a training dataset and use it as prior knowledge for segmenting the nuclei in the test image).
[0082] Next, in some embodiments, the nuclei are then segmented using a threshold calculated individually for each nucleus. For example, Otsu's method can be used for the segmentation of the region around the identified nuclei since the pixel intensity of the nuclear region is considered to vary. As understood by those skilled in the art, Otsu's method is used to determine an optimal threshold by minimizing the within-class variance and is known to those skilled in the art. More specifically, Otsu's method is used to perform clustering-based image thresholding or to automatically reduce a grayscale image to a binary image. This algorithm assumes that the image contains two classes of pixels (foreground pixels and background pixels) following a bimodal histogram. Next, an optimal threshold is calculated to separate the two classes such that their combined spread (within-class variance) is minimized or equalized (since the sum of the pairwise squared distances is constant) so that the between-class variance is maximized.
[0083] In some embodiments, the system and method further include automatically analyzing spectral and / or shape characteristics of the identified nuclei in the image to identify nuclei of non-tumor cells. For example, blobs can be identified in the first digital image of the first step. As used herein, a "blob" can be, for example, a region of a digital image where some characteristics, such as intensity or gray value, are constant or vary within a specified range of values. All pixels within a blob can be considered to be similar to each other in a sense. For example, blobs can be identified using differentiation methods based on derivatives of functions of positions on the digital image and methods based on local extrema. A nuclear blob is a blob whose pixels and / or its contour shape indicate that the blob was likely generated by a nucleus stained by the first staining. For example, it is possible to evaluate the radial symmetry of a blob to determine whether to identify the blob as a nuclear blob or as any other structure, such as a staining artifact. For example, if a blob has an elongated shape and is not radially symmetric, the blob can be identified as a staining artifact rather than as a nuclear blob. Depending on the embodiment, the blob identified as a "nuclear blob" can represent a set of pixels that are identified as candidate nuclei and can be further analyzed to determine whether the nuclear blob represents a nucleus. In some embodiments, any type of nuclear blob is directly used as an "identified nucleus". In some embodiments, the filtering operation is applied to the identified nuclei or nuclear blobs to identify nuclei that do not belong to biomarker-positive tumor cells, and to remove the identified non-tumor nuclei from the list of already identified nuclei, or not to add the nuclei to the list of nuclei identified initially. For example, additional spectral and / or shape characteristics of the identified nuclear blobs can be analyzed to determine whether the nuclei or nuclear ear blobs are nuclei of tumor cells. For example, the nuclei of lymphocytes are larger than the nuclei of other tissue cells, such as lung cells.When the tumor cells are derived from lung tissue, the lymphocyte nuclei are identified by discriminating all nuclear lobes of a minimum size or diameter that is significantly larger than the average size or diameter of normal lung cell nuclei. The identified nuclear lobes associated with lymphocyte nuclei can be removed (i.e., "filtered out") from the already identified set of nuclei. By removing the nuclei of non-tumor cells, the accuracy of the method can be increased. Depending on the biomarker, non-tumor cells may also express the biomarker to some extent and thus may generate an intensity signal in the first digital image not derived from tumor cells. By identifying and filtering out nuclei that do not belong to tumor cells from the entire set of already identified nuclei, the accuracy of identifying biomarker-positive tumor cells can be increased. These and other methods are described in U.S. Patent Application Publication No. 2017 / 0103521, the content of which is hereby incorporated by reference in its entirety for all purposes. In some embodiments, when a seed is detected, a locally adaptable thresholding method can be used and lobes are created around the detected center. In some embodiments, other methods can also be incorporated, such as using a marker-based watershed algorithm to identify nuclear lobes around the detected nuclear center. These and other methods are described in PCT / EP2016 / 051906, published as WO 2016 / 120442, the content of which is hereby incorporated by reference in its entirety for all purposes.
[0084] In some embodiments, various marker expression scores are calculated for each stain or biomarker within each cell cluster in each image (a single image from a multiplexed image or a deconvolved image channel image) using the scoring module 340. The scoring module 340 utilizes, in some embodiments, data obtained during the detection and classification of cells by the image analysis module 330. For example, the image analysis module 330 can include a series of image analysis algorithms and can be used to determine the presence of one or more of nuclei, cell walls, tumor cells, or other structures within the identified cell clusters, as described herein. In some embodiments, the derived staining intensity values and counts of specific nuclei for each field of view can be used by the scoring module 340 to determine various marker expression scores such as positive percent or H-score. The scoring methods are described in more detail in International Publication No. WO 2014 / 102130, "Image analysis for breast cancer prognosis," a co-pending application by the same applicant filed on December 19, 2013, and International Publication No. WO 2014 / 140085, "Tissue object-based machine learning system for automated scoring of digital whole slides," filed on March 12, 2014, the contents of each of which are incorporated herein by reference in their entirety. For example, the automated image analysis algorithms within the image analysis module 330 can be used to interpret each of a series of IFIC slides to detect tumor nuclei that are stained positive and negative for specific biomarkers such as Ki67, ER, PR, HER2. Based on the detected positive and negative tumor nuclei, various slide-level scores such as marker percent positive, H-score, etc. can be calculated using the scoring module 340.
[0085] In some embodiments, the expression score is the H-score used to assess the percentage of tumor cells with cell membrane staining graded as "weak", "medium", or "strong". The grades are summed to give an overall maximum score of 300 and a cutoff point of 100 to distinguish between "positive" and "negative". For example, the membrane staining intensity (0, 1+, 2+, or 3+) is determined for each cell within a fixed field of view (or here, each cell within a tumor or cell cluster). The H-score can be based simply on the dominant staining intensity, or more complexly, can include the sum of the individual H-scores for each observed intensity level. In other embodiments, the expression score is the Allred score. The Allred score is a scoring system that examines the percentage of cells tested positive for a hormone receptor and how well the receptor appears after staining (this is referred to as "intensity"). In other embodiments, the expression score is the positive percent. In the context of scoring breast cancer samples stained for the PR and Ki-67 biomarkers on PR and Ki-67 slides, the positive percent on a single slide is calculated as follows (e.g., the total number of nuclei of cells stained positive (e.g., malignant cells) in each field of view of a digital image of the slide is added and divided by the total number of nuclei stained positive and negative from each field of view of the digital image): positive percent = number of cells stained positive / (number of cells stained positive + number of cells stained negative). In other embodiments, the expression score is an IHC combination score, which is a prognostic score based on the number of IHC markers, where the number of markers is greater than 1. IHC4 is one such score based on four measured IHC markers in breast cancer samples, namely ER, HER2, Ki-67, and PR (e.g., Cuzick et al., J. Clin. Oncol. 29:4273-8, 2011, and Barton et al., Br. J. Cancer 1-6, April 24, 2012, both incorporated herein by reference).
[0086] Following the image analysis of each marker within each identified or mapped cluster and determination of the expression score, metrics can be derived from the various identified clusters and biological structures using a metric generation module 345. In some examples, morphological metrics can be calculated by applying various image analysis algorithms to the pixels contained within or surrounding a nuclear blob or seed. In some embodiments, morphological metrics include area, length of the short and long axes, perimeter length, radius, solidity, and the like. At the cellular level, such metrics can be used to classify nuclei as belonging to healthy or diseased cells. At the tissue level, statistics of these features across the tissue are utilized for classification of whether the tissue is diseased. In some examples, appearance metrics can be calculated for a particular nucleus by comparing the pixel intensity values of the pixels contained within or surrounding the nuclear blob or seed used to identify the nucleus, where the pixel intensities being compared are derived from different image channels (e.g., background channel, channel for staining of a biomarker, etc.). In some embodiments, metrics derived from appearance features are calculated from percentile values (e.g., 10th percentile value, 50th percentile value, and 95th percentile value) of pixel intensities and magnitude of gradients calculated from different image channels. For example, first, the number P(X = 10, 50, 95) of X percentile values of the pixel values of each of the image channels (e.g., three channels: HTX, DAB, luminance) of a plurality of ICs within a nuclear blob representing the nucleus of interest is identified. Calculating appearance feature metrics can be advantageous as the derived metrics can describe the characteristics of the nuclear region and the membrane region around the nucleus.
[0087] In some examples, background metrics can be calculated that indicate the appearance and / or presence of staining in the cytoplasm and cell membrane features of cells, including nuclei, extracted from an image. The background features and corresponding metrics are calculated, for example, by identifying a nuclear blob or seed representing the nucleus and analyzing a pixel region (e.g., a 20-pixel ribbon approximately 9 microns thick around the nuclear blob boundary) directly adjacent to the identified set of cells, and thus imaging the appearance and presence of staining in the cytoplasm and membrane of the cells using this nucleus along with the region directly adjacent to the cell, and can be calculated for the nuclei and corresponding cells depicted in the digital image. In some examples, the color metric may be derived from a color including a color ratio R / (R+G+B) or a color principal component. In other embodiments, the color metric derived from the color includes local statistics (mean / median / variance / standard deviation) of each color and / or color intensity correlations within a local image window. In some examples, the intensity metric may be derived from a group of adjacent cells having a specific characteristic value set between the dark and light color tones of the gray cells represented in the image. The correlation relationship of the color features can define an instance of a size class, and thus, in this way, determine cells whose intensity is affected by the surrounding cluster of dark cells among these colored cells.
[0088] In some examples, other features may be considered and used as a basis for calculating metrics such as texture features or spatial features. As another example, expression scoring can be utilized as a predictive measure or to guide treatment. For example, in the context of breast cancer and its association with ER and PR biomarkers, a sample testing positive can lead to the decision to provide hormone therapy during the treatment process. One of ordinary skill in the art will also understand that not all clusters within a biological sample may have the same score for any given marker. By being able to determine a heterogeneity score or metric that describes the variability between clusters, additional guidance can be provided for making information-based treatment decisions. In some embodiments, the heterogeneity is determined to measure how different clusters are compared to each other. The heterogeneity can be measured, for example, by a variability metric that describes how protein expression levels between various identified and mapped clusters differ from each other, as described in International Publication No. WO 2019 / 110567, the entire contents of which are incorporated herein by reference for all purposes. In some embodiments, the heterogeneity is measured across all identified clusters. In other embodiments, the heterogeneity is measured only between a subset of the identified clusters (e.g., clusters that meet certain predetermined criteria).
[0089] In some embodiments, the image received as input can be segmented and masked by a segmentation and masking module 350. For example, an architecture or model of a trained convolutional neural network can be used to segment non-target regions and / or target regions, and then the image can be masked for analysis before, during, or after inputting the image into an image analysis algorithm. In some embodiments, the input image is masked such that only tissue regions are present within the image. In some embodiments, a tissue region mask is generated to mask non-tissue regions from the tissue regions. In some embodiments, a tissue region mask can be created by identifying tissue regions and excluding background regions (e.g., regions of a whole slide image corresponding to glass without a sample, such as when only white light from the imaging source is present).
[0090] In some embodiments, a segmentation technique is used to generate a tissue region mask image by masking the tissue regions from the non-tissue regions within the input image. In some embodiments, image segmentation techniques are utilized to distinguish digitized tissue data from the slides, tissue corresponding to the foreground, and slides corresponding to the background within the image. In some embodiments, the segmentation and masking module 350 calculates an area of interest (AOI) to detect all tissue regions within the AOI of the whole slide image while limiting the amount of non-tissue regions of the background to be analyzed. A wide range of image segmentation techniques (e.g., HSV color-based image segmentation, Lab image segmentation, mean shift color image segmentation, region growing, level set method, fast marching method, etc.) can be used to determine, for example, the boundary between tissue data and non-tissue or background data. Based at least in part on segmentation, the segmentation and masking module 350 can generate a tissue foreground mask that can be used to identify portions of the digitized slide data corresponding to the tissue data. Alternatively, the component can generate a background mask that can be used to identify portions of the digitized slide data that do not correspond to the tissue data.
[0091] This identification can be enabled by image analysis operations such as edge detection. Using the tissue region mask, non-tissue background noise in the image, for example, non-tissue regions, can be removed. In some embodiments, the generation of the tissue region mask comprises one or more of the following operations (however, it is not limited to the following operations): calculating the luminance of the low-resolution input image, generating the luminance image, applying a standard deviation filter to the luminance image, generating the filtered luminance image, and applying a threshold to the filtered luminance image such that pixels having a luminance exceeding a given threshold are set to 1 and pixels below the threshold are set to zero, generating the tissue region mask. Additional information and examples regarding the generation of the tissue region mask are disclosed in PCT / EP / 2015 / 062015 entitled "An Image Processing Method and System for Analyzing a Multi-Channel Image Obtained from a Biological Tissue Sample Being Stained by Multiple Stains", the content of which is incorporated herein by reference in its entirety for all purposes.
[0092] In addition to masking the non-tissue region from the tissue region, the segmentation and masking module 350 can also mask other regions of interest, such as non-target regions or a portion of the tissue identified as belonging to a specific tissue type (e.g., lymphoid aggregate region), or a portion of the tissue identified as belonging to the target region or a specific tissue type (e.g., suspicious tumor region), as needed. In various embodiments, non-target region segmentation, such as lymphoid aggregate region segmentation, is performed by a CNN model (e.g., the CNN model associated with the classifier subsystem 210a described with respect to FIG. 2). In some embodiments, the CNN model is a two-dimensional segmentation model. For example, the CNN model may be a U-Net having residual blocks, dilations, and depthwise convolutions. Pre-processed or processed image data (e.g., two-dimensional regions or whole slide images) can be used as input to the U-Net. The U-Net includes a contracting path complemented by an expanding path, and the pooling operations of successive layers within the expanding path are replaced by upsampling operators. Thus, these successive layers increase the resolution of the output. Based at least in part on the segmentation, the U-Net can generate a non-target region foreground mask that can be used to identify the portion of the digitized slide data corresponding to the non-target region data. Alternatively, the component can generate a background mask that can be used to identify the portion of the digitized slide data that does not correspond to the non-target region data. The output of the U-Net may be a foreground non-target region mask representing the location of the non-target region present in the underlying image, or a background non-target region mask representing the portion of the digitized slide data that does not correspond to the non-target region data (e.g., the target region).
[0093] In some embodiments, the alignment module 355 and alignment process are used to map biological materials or structures, such as tumor cells or cell clusters identified in one or more images, to one or more additional images. Alignment is the process of transforming different datasets, here images, or cell clusters within an image, to a single coordinate system. More specifically, alignment is the process of aligning two or more images, and generally involves designating one image as a reference (also called a reference image or fixed image), and applying geometric transformations to the other images so that they are aligned with the reference. Geometric transformations map the location of one image to a new location in another image. The step of determining the correct geometric transformation parameters is the key to the image alignment process. In some embodiments, image alignment is performed using the method described in International Publication No. WO 2015 / 049233, entitled "Line-Based Image Registration and Cross-Image Annotation Devices, Systems and Methods," filed on September 30, 2014, the contents of which are hereby incorporated by reference in their entirety for all purposes. International Publication No. WO 2015 / 049233 describes an alignment process that includes a coarse alignment process used alone or in combination with a refined alignment process. In some embodiments, the coarse alignment process can include selecting digital images for alignment, generating a foreground image mask from each of the selected digital images, and matching the resulting tissue structures between the foreground images. In further embodiments, generating a foreground image mask can include generating a soft-weighted foreground image from a whole-slide image of a stained tissue section, and applying an OTSU threshold to the soft-weighted foreground image to generate a binary soft-weighted image mask.In yet further embodiments, generating a foreground image mask includes generating a binary soft-weighted image mask from a whole-slide image of a stained tissue section, separately generating a gradient amplitude image mask from the same whole-slide image, applying an OTSU threshold to the gradient image mask to generate a binary gradient amplitude image mask, and combining the binary soft-weighted image and the binary gradient amplitude image mask using a binary OR operation to generate the foreground image mask. As used herein, a "gradient" is, for example, a pixel intensity gradient calculated for a particular pixel by considering the intensity value gradients of a set of pixels surrounding the particular pixel. Each gradient can have a particular "direction" with respect to a coordinate system whose x-axis and y-axis are defined by two orthogonal edges of a digital image. A "gradient orientation feature" can be a data value indicating the orientation of the gradient within the coordinate system.
[0094] In some embodiments, matching tissue structures includes calculating line-based features from the respective boundaries of the resulting foreground image masks, calculating global transformation parameters between a first set of line features on a first foreground image mask and a second set of line features on a second foreground image mask, and globally aligning the first and second images based on the transformation parameters. In yet another embodiment, the coarse alignment process includes mapping a selected digital image to a common grid based on the global transformation parameters, where the grid can encompass the selected digital image. In some embodiments, the fine alignment process can include identifying a first sub-region of a first digital image within a set of aligned digital images, identifying a second sub-region of a second digital image within the set of aligned digital images, where the second sub-region is larger than the first sub-region and the first sub-region is substantially located within the second sub-region on the common grid, and calculating an optimized position for the first sub-region within the second sub-region.
[0095] Figure 4 depicts an example of staining variation across different H&E slide images 410, 420, 430, 440. In various examples, the H&E slides can have different colors and brightness. For example, different pathology laboratories and / or pathologists can choose to stain tissue samples based on individual preferences, different staining processes, and / or different staining / scan devices. Additionally, the H&E slide images can be of different types of tissue (e.g., tumor, stroma, and necrosis) and / or different organs (e.g., liver, prostate, breast, etc.). Thus, the global models 112, 114 should be appropriately trained to be general enough such that the models still operate accurately despite variations in color, tissue, and organ, or multiple models can be utilized.
[0096] Figure 5 shows a process for training a prediction model according to various embodiments.
[0097] The process for training begins at block 500, where multiple tile images of specimens are accessed. One or more of the multiple tile images include an annotation of the one or more tile images (e.g., to identify regions with tumor cells, to segment non-target and target regions, or any other suitable annotation). At block 510, one or more tile images can be divided into image patches (e.g., of size 256 pixels × 256 pixels). At block 520, a prediction model such as a two-dimensional segmentation model is trained using the one or more tile images or image patches. In some examples, the two-dimensional segmentation model is a modified U-Net model that includes a downsampling path and an upsampling path, each of the downsampling path and the upsampling path having a maximum of 256 channels, and one or more layers of the downsampling path performing spatial dropout. Training can include performing iterative operations to find a set of parameters of the prediction model that minimizes the loss function of the prediction model. Each iteration can include finding a set of parameters of the prediction model such that the value of the loss function using the set of parameters is less than the value of the loss function using another set of parameters in the previous iteration. The loss function is configured to measure the difference between the output predicted using the prediction model and the annotation included in the one or more tile images or image patches. In some examples, training further includes adjusting the learning rate of the modified U-Net by reducing the learning rate according to a predetermined schedule. The predetermined schedule can be a step decay schedule that reduces the learning rate by a predetermined factor every predetermined number of epochs to optimize the loss function. In a particular example, the loss function is a binary cross-entropy loss function. At block 530, the further trained prediction model can be provided to the aggregation server after a number of iterations, a length of time, or after the model has been modified by a threshold amount. For example, the further trained prediction model can be deployed to run in the FL image analysis environment as described with respect to FIGS. 2 and 3.
[0098] FIG. 6 shows a process for rounds of FL training of a prediction model according to various embodiments.
[0099] The FL process for a round of training starts at block 600, where each of the client devices is provided with one or more global models for use in classification. Each of the client devices can access local data that can be used for further training of the provided global models. One or more tile images from the local data include annotations of the one or more tile images (e.g., to identify regions with tumor cells, to segment non-target and target regions, or any other suitable annotation). As described above, the one or more tile images can be split into image patches. At block 610, the prediction model (e.g., the global model) is further trained on the one or more tile images or image patches. At block 620, after the local training data has been exhausted, the further trained prediction model is provided to the aggregation server. At block 630, the server can receive the one or more further trained models and aggregate the weights from those models into the global model. The weights can be aggregated by performing an average, a weighted average, or any other suitable method for combining weights as understood by those skilled in the art. For example, in some embodiments, the weights may be incorporated into the global model based on a weighted average based on the number of training rounds (e.g., slides analyzed) performed.
[0100] FIG. 7 shows the results generated after multiple rounds of FL training of a prediction model according to various embodiments.
[0101] The improved accuracy provided by multiple training rounds can be visualized. For example, H&E image 700 can be used to validate the training of the FL system. Ground truth 710 can be provided for comparison with the output of the model. In this example, the image is colored blue to indicate tumors and purple for all other tissues. Exemplary results 720 using a model trained with centralized data are also provided. In this example, 6 rounds of classification and training are performed and the resulting classifications 730 generated by each round are shown. After each round of FL, the global model is further trained in one or more client systems and the results converge to the ground truth 710.
[0102] Figure 8 shows a process for rounds of FL training of a prediction model according to various embodiments.
[0103] In various embodiments, the FL process for a round of training begins at block 800, where each of the client devices is provided with one or more global models for use in classification. As described above, each of the client devices can access local data that can be used for further training of the provided global models, and one or more tile images from the local data include annotations (e.g., to identify regions with tumor cells, to segment non-target and target regions, or any other appropriate annotation). Additionally, the local data may include metadata that further describes that local data. For example, the metadata can include information about how the sample was prepared (e.g., the staining applied, the staining concentration, and / or any other relevant information related to sample preparation), the equipment used (e.g., staining equipment, scanning equipment, etc.), and further patient information. At block 810, the metadata may be raised to determine whether data compensation or normalization needs to be managed. For example, certain scanning devices may introduce artifacts that require compensation. In another example, some staining concentrations may result in overly bright or dark coloring that can be compensated for. Thus, at block 820, the system can compensate for data imbalance using the metadata or other information. At block 830, the model is further trained on one or more tile images or image patches, and the updated model is provided to the central server, updating the global model. At block 840, the updated global model is tested using a validation dataset to confirm improvement of the model. When the global model is improved, the changes can be saved. At block 850, the server can distribute the updated model to each client device.
[0104] FIG. 9 shows a process for receiving an updated model from a client according to various embodiments.
[0105] In various embodiments, the centralized server receives updated models and metadata from client devices. As described above, at block 910, the system can evaluate the metadata associated with the local training data. In various embodiments, the system may be configured to have a plurality of global classifiers selected according to various metadata. For example, separate classifiers can be used for locations that utilize specific devices or staining techniques. Thus, at block 920, the system can be configured to determine whether the updated classifier should be used to update one of the plurality of global models or whether a new global model should be added. At block 930, the received updated model is normalized and used to update one of the global models. At block 940, the newly updated model is verified using a validation dataset. At block 950, it is determined that a new global model needs to be added. Thereby, the received updated model is verified. Next, at block 960, the verified model is added to the plurality of global models. At block 970, the updated model is distributed to appropriate client devices.
[0106] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system is a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods and / or some or all of one or more of the processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more of the methods and / or some or all of one or more of the processes disclosed herein.
[0107] The terms and expressions used are used as terms of explanation and not of limitation, and in the use of such terms and expressions, there is no intention to exclude equivalents or portions thereof of any features shown and described, but it is recognized that various modifications are possible within the scope of the invention described in the claims. Accordingly, while the invention described in the claims is specifically disclosed by the embodiments and any features, modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and it should be understood that such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.
[0108] The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the preferred exemplary embodiments provides those skilled in the art with a possible explanation for implementing various embodiments. It is understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope described in the appended claims.
[0109] Specific details are provided in the following description to provide a complete understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as not to obscure the embodiments.
Claims
1. A computer-implemented method for using a federated learning classifier, comprising: distributing a global model by a central server to a plurality of client devices, wherein the global model takes as input slide images containing stained tissues and / or cells, and outputs classification results of objects contained in the slide images; receiving, by the central server, a local model from at least one of the plurality of client devices, wherein the local model is a model further trained on the global model using a plurality of slide images containing stained tissues and / or cells and a plurality of corresponding annotations identifying objects contained in the slide images in at least one of the plurality of client devices; receiving, by the central server, metadata from at least one of the plurality of client devices, wherein the metadata includes at least one of the type of staining, the concentration of staining, and the type of device used for staining performed on the plurality of slide images used in training in at least one of the plurality of client devices; normalizing, by the central server, the local model using the metadata; updating, by the central server, the global model using the normalized local model; and distributing the updated global model to at least one of the plurality of client devices.
2. The computer-implemented method according to claim 1, wherein updating the global model includes performing a weighted average of the normalized local model and the global model.
3. The weighted average is executed with weights according to the number of the plurality of slide images used in training in at least one of the plurality of client devices and the total number of images used to train the global model before the update. The computer-implemented method according to claim 2.
4. A computer-implemented method for using an association learning classifier, distributing a global model by a central server to a plurality of client devices, wherein the global model is a model that takes as input a slide image including tissues and / or cells and outputs a classification result of an object included in the slide image; distributing the global model; receiving, by the central server, a local model from at least one of the plurality of client devices, wherein the local model is a model further trained on the global model using a plurality of slide images including tissues and / or cells and a plurality of corresponding annotations identifying objects included in the slide images in at least one of the plurality of client devices; receiving the local model; updating, by the central server, the global model using the local model; distributing the updated global model to at least one of the plurality of client devices; comprising updating the global model includes performing a weighted average of the local model and the global model, and the weighted average is executed with weights according to the number of the plurality of slide images used in training in at least one of the plurality of client devices and the total number of images used to train the global model before the update. A computer-implemented method.
5. The computer-implemented method according to any one of claims 1 to 4, wherein the annotation is provided by a user observing the output of the global model for the slide image in at least one of the plurality of client devices, and the annotation includes a change to the output generated by the global model.
6. The computer-implemented method according to claim 5, wherein the change to the output generated by the global model includes at least one change of cell type, tissue type, or tissue boundary.
7. The computer-implemented method according to any one of claims 1 to 6, wherein the local model does not include personal information of a patient.
8. The computer-implemented method according to any one of claims 1 to 7, further comprising determining, by the centralized server, whether the local model provides improved performance over the global model using a set of labeled validation input images.
9. The computer-implemented method according to claim 8, wherein when it is determined that the local model provides improved performance over the global model, the global model is updated, and when it is determined that the performance of the local model is inferior to that of the global model, the global model is not updated.
Citation Information
Patent Citations
Distributed model learning
JP2017519282A
Biospecimen classification methods and systems, including analysis optimization and correlation exploitation
JP2018502275A
Distributed clinical workflow training of deep learning neural networks
WO2018098039A1
Using machine learning and / or neural networks to validate stem cells and their derivatives for use in cell therapy, drug discovery, and diagnostics
WO2019178561A2
Application development platform and software development kits that provide comprehensive machine learning services
WO2019216938A1