Sample processing independent image representation learning for digital pathology
By receiving image data processed by different slide preparation machines, using student-teacher framework and adversarial learning methods to train the encoder, the problem of convolutional neural networks being difficult to distinguish features in digital pathology is solved, and effective image analysis under different scanners and processing technologies are achieved and classification accuracy is improved.
Patent Information
- Application Number
- CN202380081893.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-12-01
- Publication Date
- 2025-07-08
AI Technical Summary
In digital pathology, existing convolutional neural networks are difficult to effectively distinguish the features in digital pathology images without special pre-training, and the differences caused by different scanners and processing technologies affect downstream classification tasks, resulting in a degradation of model performance.
By receiving image data including processed by different slide preparation machines, generating enhanced views, and training encoders using student-teacher frameworks and adversarial learning methods, reducing the impact of device features such as scanners and dyers, developing machine learning models independent of sample processing.
Effective analysis of digital pathological images under different scanners and processing technologies is achieved, and the accuracy and consistency of downstream classification tasks are improved.
Smart Images

Figure CN120283260A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 385,931, filed on December 2, 2022, entitled "SAMPLE PROCESSING AGNOSTIC IMAGE REPRESENTATION LEARNING FOR DIGITAL PATHOLOGY", the entire disclosure of which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to digital pathology, and more particularly, to tools for analyzing and classifying digital pathology images. Background Art
[0004] With the advancement of digital scanning technology, many laboratories and other institutions that store and preserve tissue specimens on glass slides have been scanning these slides to generate digital images of tissue samples. Pathologists or other trained experts typically evaluate whole - slide images (WSIs) for evidence of abnormalities in the depicted tissue. Digital WSIs are typically very large; for example, each of several color channels has 100,000 pixels by 100,000 pixels, making it difficult to effectively analyze WSIs at the global level without applying advanced computer image - analysis techniques.
[0005] Traditional image - analysis techniques typically begin by obtaining a convolutional neural network (CNN) pre - trained on a general natural - image set (e.g., ImageNet) and then applying transfer - learning techniques. However, since general natural images typically do not depict objects that appear in WSIs, the value of such CNNs may be limited, and considering the unique and specific aspects of WSIs, such CNNs may not help in distinguishing target features. For example, most structures in hematoxylin and eosin (H&E) - stained slides are in the color gradations of blue, purple, and pink.
[0006] For digital pathology, it is desirable to obtain a CNN that has been specifically pre - trained to extract features (e.g., H&E slide images) from WSIs. However, unfortunately, most large WSI datasets are unlabeled, making it difficult to utilize supervised - learning techniques.
[0007] In addition, in digital pathology, different techniques can be used to process tissue samples. These processing techniques can impart differences in the generated digital pathology images. For example, digital pathology slides obtained from different scanners may have visible (and in some cases invisible) differences due to settings, features, components, or other aspects. These different scanners can use different light sources to capture images, have different lenses, use different software or different software versions, have lenses formed from different materials, or have other differences or combinations of the above differences. As another example, the differences can be attributed to different slide preparation machines (e.g., stainers) used to prepare digital pathology slides, different batches of colorants, different stains, or in some cases, two different scanning machines of the same model may just be calibrated differently. As yet another example, the differences can be attributed to different magnifications used by the same or different scanners, different sample thicknesses being prepared, etc.
[0008] When training a feature encoder, using images obtained via different tissue sample processing techniques (e.g., from different scanners) can preserve the specific features of the tissue sample processing technique (e.g., scanner) used to obtain the corresponding images, which may hinder downstream classification tasks and analysis. For example, the quality of the features derived from the images may be degraded because the representation of the images has a relatively small relationship with the biological attributes depicted therein.
[0009] Accordingly, there is a need to develop and train machine learning models for performing digital pathology analysis that are independent of the tissue sample processing techniques used to prepare and capture the corresponding images. Summary of the Invention
[0010] Systems, methods, and programming are described herein for training and implementing machine learning models for performing digital pathology analysis that are independent of the tissue sample processing techniques used to prepare and capture the corresponding digital pathology images.
[0011] Some aspects include receiving image data including a first set of images and a second set of images. The first set of images and the second set of images can include digitized images of a plurality of digital pathology slides processed using a first slide preparation machine and a second slide preparation machine, respectively. The first slide preparation machine and the second slide preparation machine can each have a set of attributes, where the value of at least one of the attributes can be different between the first slide preparation machine and the second slide preparation machine. A first set of enhanced views and a second set of enhanced views can be generated using the image data. The first enhanced view and the second enhanced view can be generated based on one or more enhancements applied to each image in the first set of images and the second set of images. For each of the digital pathology slides, a first vision transformer can be trained. Training the first vision transformer can include: generating, using the first vision transformer, a first representation of an enhanced view in the first set of enhanced views, and enhancing the similarity between the first representation and a second representation of an enhanced view in the second set of enhanced views. The second representation can be generated via a second vision transformer. Both the first representation and the second representation can correspond to the same digital pathology slide.
[0012] Some additional aspects include receiving training data including images of a plurality of biological samples processed using a first slide preparation machine or a second slide preparation machine. The first slide preparation machine and the second slide preparation machine can each have a set of attributes. The value of at least one of the attributes can be different between the first slide preparation machine and the second slide preparation machine. For each of the biological samples, a first encoder can be used to generate a first representation of each image of the biological sample processed using the first slide preparation machine. The first representation can be provided to a discriminator to produce a prediction as to whether the biological sample corresponding to one of the images was processed using the first slide preparation machine or the second slide preparation machine. One or more parameters of the discriminator can be updated based on a first loss, which is calculated based on the generated prediction and metadata associated with one of the images. The metadata can indicate whether the biological sample was processed using the first slide preparation machine or the second slide preparation machine. For example, the metadata can indicate that the biological sample was processed using the first slide preparation machine. One or more parameters of the first encoder can be updated based on the first loss. The updated first encoder can be trained to generate an updated first representation of one of the images and to enhance the similarity between the updated first representation and a second representation of the same biological sample in an image generated using a second encoder.
[0013] The embodiments disclosed above are merely examples, and the scope of the present disclosure is not limited thereto. A particular embodiment can include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed above. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] This patent or application document contains at least one color drawing. After a request is made and the necessary fees are paid, the Patent Office will provide a copy of the published patent or patent application with one or more color drawings.
[0015] Figure 1A Schematic diagram showing an exemplary system for training and using a model independent of sample processing according to various embodiments.
[0016] Figure 1B Schematic diagram of an exemplary process for obtaining an image depicting a biological sample according to various embodiments.
[0017] Figures 2A to 2B Schematic diagram of exemplary image data and exemplary processing technique tags of image data according to various embodiments.
[0018] Figure 3 Schematic diagram of exemplary training data for training a model independent of sample processing according to various embodiments.
[0019] Figure 4 Schematic diagram showing an exemplary framework of a model serving as the basis for a model independent of sample processing according to various embodiments.
[0020] Figure 5 Schematic diagram showing an exemplary framework of another model serving as the basis for a model independent of sample processing according to various embodiments.
[0021] Figure 6A Schematic diagram of an exemplary image patch according to various embodiments, the exemplary image patch being derived from images captured using different tissue processing techniques and being included in an image set for training a model.
[0022] Figure 6B Schematic diagram showing different slide preparation machines and different settings associated with these slide preparation machines for generating images representing the same biological sample according to various embodiments.
[0023] Figure 7 Flowchart of an exemplary process for training a model independent of sample processing using average similarity training techniques according to various embodiments.
[0024] Figure 8 For training based on the framework of Figure 4 Schematic diagram of an exemplary adversarial model for training a model independent of sample processing built on the framework of
[0025] Figures 9 to 10 For showing the training of Figure 8Schematic diagram of an exemplary training process of an adversarial model.
[0026] Figure 11 Another schematic diagram of an exemplary adversarial model for training a model independent of sample processing built on the framework of Figure 5 according to various embodiments.
[0027] Figures 12A to 12B Schematic diagram showing an exemplary training process of an adversarial model for training Figure 11 according to various embodiments.
[0028] Figure 13 Flowchart showing an exemplary process for training a model independent of sample processing using adversarial training techniques according to various embodiments.
[0029] Figure 14A Schematic diagram of an exemplary downstream classification subsystem for training a downstream classifier according to various embodiments.
[0030] Figure 14B Flowchart showing an exemplary process for training a downstream classifier according to various embodiments.
[0031] Figure 15A Schematic diagram of exemplary training and validation data for training a model according to various embodiments.
[0032] Figure 15B Schematic diagram of exemplary downstream classifier training and validation data for training a downstream classifier according to various embodiments.
[0033] Figure 16 Flowchart showing an exemplary process for analyzing images of biological samples according to various embodiments.
[0034] Figure 17 Images of different biological samples captured using different tissue processing techniques according to various embodiments are shown.
[0035] Figure 18A Graph showing the embeddings generated from multi-scanner data according to various embodiments.
[0036] Figure 18B Graphs showing various representations created by an encoder for biological samples included in the validation data of an encoder trained using the average similarity method according to various embodiments.
[0037] Figure 19 Graph showing the representation created by an encoder for biological samples included in the validation data of an encoder trained using the adversarial method according to various embodiments.
[0038] Figures 20A to 20B Graphs showing the standard deviation in training of embeddings of the mean embedding similarity method and the adversarial method according to various embodiments are presented.
[0039] Figure 21 A data table showing an image of a biological sample and sample data according to various embodiments is presented.
[0040] Figures 22A to 22B Images obtained by two different scanners of the same biological sample according to various embodiments are presented.
[0041] Figures 23A to 26B Various exemplary embedding graphs according to various embodiments are presented.
[0042] Figure 27 An exemplary computing system that can implement one or more embodiments described herein is presented. DETAILED DESCRIPTION
[0043] In digital pathology, biological samples (such as tissue samples) can be processed using a slide preparation machine. The slide preparation machine may include a slide scanner (also interchangeably referred to herein as a “scanner”), a slide stainer, or other machines and systems. Digital images of biological samples can be captured using a slide scanner. The images can be analyzed to detect possible tissue abnormalities or other anomalies. Traditionally, these images were analyzed by humans (such as doctors), but the emergence of machine learning and artificial intelligence has enabled faster and more powerful digital pathology analysis. In addition, machine learning models and artificial intelligence can detect associations and / or anomalies that cannot be detected by traditional manual review.
[0044] Machine learning models can be trained to perform tasks. For example, an image classifier can be trained to receive an image, generate a computer - understandable representation of the image, and output a classification result based on the generated representation. The classification result can indicate one or more classes of objects that the image is determined to depict, the likelihood (e.g., probability) that the image depicts an object belonging to one or more specific classes, or other information.
[0045] To train a machine learning model to perform classification, training data including images of biological samples (e.g., whole slide images) can be provided to the model. In some embodiments, images of each biological sample can be captured. Additionally, in some embodiments, each biological sample's image can be captured using each slide scanner from a set of slide scanners. Slide scanners can include different attributes such as magnification, lens material, lens thickness, software, etc. The purpose of training the model using images captured with different slide scanners is to ensure that the model does not learn features associated with the specific slide scanner used and / or to minimize the impact of the features of a specific scanner on the classification results. Additionally, different processing techniques can be used to process biological samples prior to imaging. For example, different slide stainers can be used to prepare biological samples for digitization. Different slide stainers can use different stains, different batches of stains, or other differences or combinations thereof. Images of biological samples processed using different slide preparation machines can be used to train the model to ensure that the model does not learn features associated with the specific slide preparation machine used and / or to minimize the impact of these features on downstream classification results. Tissue cutting can also introduce differences. For example, one biological sample can be prepared with a first thickness while a second biological sample can be prepared with a second thickness.
[0046] In some embodiments, the backbone of the trained model includes a pre-trained feature encoder. The feature encoder is trained to generate meaningful image representations. The representation describes aspects of the underlying biological sample and can be analyzed by a computer to perform one or more tasks (e.g., tissue classification). As defined herein, an "embedding" refers to a vector representation of the features that describe an image. An embedding is meaningful if it has features suitable for downstream classification or regression tasks. The function of the encoder is to "encode" the image data representing the image into an embedding in the embedding space. Thus, the encoder can be trained to focus on features that describe the morphological properties of what the image depicts while also not focusing on features that are irrelevant or of minimal value to the downstream classification task. Regression can also be performed. For example, the progression of a disease can be predicted, resulting in a continuous output that can be analyzed via regression techniques.
[0047] In some cases, the data used to train a feature encoder (e.g., a classifier) includes images of biological samples. For example, images of tissue samples depicting tumors or other anatomical / biological abnormalities can be used to train a classifier to detect their presence in images of other biological samples. Tissue abnormalities that can be detected from WSI include, by way of example and not limitation, inflammation, pigmentation, degeneration, anisocytosis, hypertrophy, increased mitosis, monocytic infiltration, inflammatory cell infiltration, inflammatory cell foci, glycogenopenia, glycogen accumulation (diffuse or focal), extramedullary myelopoiesis, extramedullary hematopoiesis, extramedullary erythropoiesis, single cell necrosis, diffuse necrosis, marked necrosis, coagulative necrosis, apoptosis, nuclear gigantism, pericholecystic, increased cellularity, glycogen deposition, lipid deposition, microgranuloma, hyperemia, Kupffer cell pigmentation, increased hemosiderin, histiocytosis, hyperplasia, or vacuolization, etc.
[0048] In some embodiments, a given biological sample is disposed on a slide prepared using one or more slide preparation mechanisms. Slide preparation can include, for example, one or more slide stainers (also interchangeably referred to herein as "stainers"), one or more slide scanners (e.g., scanners), or other machines. Each slide preparation machine can have a set of attributes. On different slide preparation machines, the values of each set of attributes can be the same or similar. Some embodiments include that the value of at least one of the attributes of two separate slide preparation machines is the same (or substantially the same). Some embodiments include that the value of at least one of the attributes of the same slide preparation machine used to prepare and digitize an image of a biological sample is the same (or substantially the same). For example, a first batch of images can be captured using a scanner for biological samples using one stain, and a second batch of images can be captured using a scanner for biological samples using another stain. The set of attributes can indicate the stain (such as hematoxylin and eosin) used to stain the slide, the tissue thickness for the biological sample, or other attributes.
[0049] One or more slide scanners (e.g., scanners) can be used to image the slide. The attributes of the slide scanner can include magnification, lens material, lens thickness, scanning software, or other attributes. In some embodiments, the value of at least one attribute of a first scanner can be different from the value of at least one attribute of a second scanner. Generally, the values of the attributes are independent of the morphological characteristics of the underlying biological sample. The scanner can capture an image of the biological sample, and in some embodiments, such an image can be referred to as a whole slide image (WSI).
[0050] As described herein, a WSI is a digital image of an extremely large format (e.g., 100,000 x 100,000 pixels), which can be generated by digitizing a physical glass slide of a biological sample (e.g., tissue) into a high-resolution image file, or can be directly output by a medical scanning device. Due to the nature of the images captured, and to avoid misdiagnosis of the tissue depicted in the WSI due to artifacts typically caused by image compression and manipulation, WSIs are typically saved in the highest resolution format possible. WSIs typically include several orders of magnitude more pixels than typical digital images and can include a resolution of 100,000 pixels x 100,000 pixels (e.g., 10,000 megapixels) or higher.
[0051] Analysis of WSIs is a labor-intensive process that requires highly specialized personnel with knowledge and flexibility to review WSIs, identify and recognize abnormalities, classify the abnormalities, label the WSIs, and potentially render a diagnosis of the tissue depicted by the WSIs. Additionally, because WSIs are used for a wide range of tissue types, personnel with knowledge and skills in recognizing abnormalities must be further specialized in order to provide accurate analysis and diagnosis.
[0052] Accordingly, due to the labor- and knowledge-intensive nature of the work, WSIs are considered candidates for automating certain functions. However, the large size of WSIs renders typical image analysis techniques ineffective, slow, and expensive. Performing standard image recognition and deep learning techniques is impractical as these techniques require multiple rounds of analysis of many samples of the WSIs to increase accuracy. The techniques described herein are aimed at solving the problem of automating feature recognition in WSIs and enabling the development of novel data analysis and rendering techniques that have not been previously optimized for these images (e.g., WSIs) that have different characteristics compared to natural images.
[0053] Intrinsic differences can occur when classifiers are trained using images captured from different sources, or when classifiers are obtained using different methods and / or protocols. For example, an image of a biological sample captured using one scanner may be different from an image of the same biological sample captured using a different scanner. As an example, reference Figure 22A and 22B, Images 2200 and 2250 illustrate examples of the same biological sample captured by two different slide scanners. Although the magnifications in Images 2200 and 2250 are different, there are other differences between the features detected by either scanner when imaging the biological sample, which may affect the resulting embeddings and lead to errors in downstream classification. As another example, an image of a biological sample prepared using a first slide preparation technique may be different from an image of the same biological sample captured using a second slide preparation technique. These differences may affect the embeddings generated by the encoder, and thus may affect downstream classification. Accordingly, techniques for training an encoder for use in a downstream biological sample image classification task that is independent of the slide preparation machine (e.g., scanner, stainer) used to capture the underlying image of the biological sample, and / or alternatively independent of the characteristics of other specific sources of the image, would be beneficial.
[0054] Certain embodiments may repeat one or more steps of the exemplary process as appropriate. Although the present disclosure describes and illustrates the specific steps of the exemplary process as occurring in a particular order, the present disclosure contemplates any suitable steps of the exemplary process occurring in any suitable order. Additionally, although the present disclosure describes and illustrates the exemplary process, the present disclosure encompasses any suitable process that includes any suitable steps, which may, as appropriate, include all or some of the steps of the exemplary process, or exclude steps. Additionally, although the present disclosure describes and illustrates specific components, devices, or systems for performing the specific steps of the exemplary process, the present disclosure contemplates any suitable combination of any suitable components, devices, or systems for performing any suitable steps of the exemplary process.
[0055] In all of the exemplary embodiments described herein, appropriate options, features, and system components may be provided to enable information collection, storage, transmission, information security measures (e.g., encryption, authentication / authorization mechanisms), anonymization, pseudonymization, isolation, and aggregation in compliance with applicable laws, regulations, and rules. In all of the example embodiments described herein, appropriate options, features, and system components may be provided to enable the protection of the privacy of a particular individual, including (by way of example and not limitation) generating reports regarding what personal information is being or has been collected and how that information is being used, such that any personal information collected can be deleted or purged and / or such that the purpose for which any personal information collected is used can be controlled.
[0056] Figure 1ASchematic diagram showing an exemplary system 100 for training and using a model independent of sample processing according to various embodiments. System 100 may include a computing system 102, a slide preparation machine 120 (e.g., slide preparation machines 120-1 - 120-M), client devices 130 (e.g., client devices 130-1 - 130-N), a database 140 (e.g., image database 142, training data database 144, model database 146, validation data database 148), or other components. In some embodiments, the components of system 100 may communicate with each other using a network 150 such as the Internet.
[0057] In some embodiments, the slide preparation machine 120 may use one or more slide preparation techniques to prepare and digitize slides of biological samples. These slide preparation techniques may impart different characteristics to the resulting representations of the biological samples, which may affect downstream classification tasks. As an example, the slide preparation machines 120 may be of the same type / model, but some of the slide preparation machines 120 may be different. Specifically, each slide preparation machine 120 may have a set of attributes, and the values of these attributes may be the same or different between different slide preparation machines or even the same slide preparation machine. The slide preparation machine may be used to prepare slides of biological samples for digitization, scan the slides to obtain images of the biological samples, or perform other tasks or combinations thereof.
[0058] For example, the slide preparation machine 120 may include a slide staining machine (e.g., a stainer) and a slide scanner (e.g., a scanner) made by the same or different manufacturers. The slide preparation machine 120 may have different attributes, such as settings, features, components, or other differences. For example, different scanners may use different light sources to capture images, have different lenses, use different software or different software versions, have lenses formed of different materials, have different lens thicknesses, or have other differences or combinations thereof. In some cases, two different slide scanners of the same model may have been calibrated differently. As another example, different stainers may use different stains, different batches of the same stain, sample different tissue thicknesses, or have other differences or combinations thereof. For example, a batch of coloring agent may be used for biological samples captured by one or more slide preparation machines 120. In some embodiments, one or more steps may be performed on the biological sample before being processed by the slide preparation machine 120. For example, the biological sample may be cut into different thicknesses. Typically, the cutting process is performed by a person, however some embodiments include using an automated sample cutter. Whether the tissue is cut via a human or a machine, each cut and each tissue sample may be different. Differences in tissue thickness can also affect digital pathology analysis. Thus, the techniques described herein may also be able to minimize the impact that variations in tissue sample thickness may have on downstream classification tasks.
[0059] The client device 130 may include client devices 130-1 through 130-N. Each client device 130 may be capable of communicating with one or more components of the system 100 via the network 150 and / or via a direct connection. The client device 130 may refer to a computing device capable of docking with various components of the system 100 to control one or more tasks, cause one or more actions to be performed, or implement other operations. For example, the client device 130 may be configured to receive and display images of scanned biological samples. Exemplary computing devices that the client device 130 may correspond to include, but are not limited to (this does not mean that other lists are restrictive), desktop computers, servers, mobile computers, smart devices, wearable devices, cloud computing platforms, or other client devices. In some embodiments, each client device 130 may include one or more processors, memory, communication components, display components, audio capture / output devices, image capture components, or other components or combinations thereof. Each client device 130 may include any type of wearable device, mobile terminal, fixed terminal, or other device.
[0060] Note that although one or more operations are described herein as being performed by specific components of computing system 102, in some embodiments, these operations may be performed by other components of computing system 102 or other components of system 100. As an example, although one or more operations are described herein as being performed by components of computing system 102, in some embodiments, these operations may be performed by components of slide preparation machine 120 and / or client device 130. Note that although some embodiments are described herein with respect to machine learning models, in other embodiments, other predictive models (e.g., statistical models or other analytical models) may be used in place of or in addition to machine learning models (e.g., in one or more embodiments, a statistical model that replaces a machine learning model and a non-statistical model that replaces a non-machine learning model).
[0061] Computing system 102 may include one or more subsystems, such as, for example, training data generation subsystem 110, model training subsystem 112, downstream classification subsystem 114, or other subsystems.
[0062] Training data generation subsystem 110 may be configured to generate training data for training a model, generate validation data for validating the model after training, update the training / validation data, or perform other operations related to preparing data for training a model. In some embodiments, training data generation subsystem 110 may be configured to obtain images depicting biological samples. One of the slide preparation machines 120 (e.g., scanner 124) may be used to capture each image. For example, a first slide preparation machine may be used to capture image 1, while a second slide preparation machine may be used to capture image 2, and both image 1 and image 2 are images of the same biological sample (e.g., tissue sample). Slide preparation machine 120 may include a set of attributes, and in some embodiments, the value of at least one of the attributes may be the same. For example, a first scanner A of a first slide preparation machine may be used to capture an image of a biological sample, and a second scanner B of a second slide preparation machine may be used to capture an image of the same biological sample. In some embodiments, slide preparation machine 120 may employ different slide preparation processes. For example, a first stainer of a first slide preparation machine may use a first stain to stain a slide of a biological sample, and a second stainer of a second slide preparation machine may use a second stain to stain a slide of the biological sample.
[0063] In some embodiments, some or all of the captured images may include metadata indicating the slide preparation machine used for those images. Associated metadata may indicate scanner source, stainer source, slide batch information, or other information to classify attributes of the slide preparation process used to prepare the slide and / or its captured images. An example of image database 142 may be at Figure 2Afound in
[0064] Figure 1B is a schematic diagram of an exemplary process 180 for obtaining an image depicting a biological sample according to various embodiments. As described above, the slide preparation machine 120 may include a stainer 122, a scanner 124, or other devices for preparing and capturing images for digital pathology analysis. Although Figure 1B the slide preparation machine 120 including a stainer 122 and a scanner 124 is depicted, each of them may be separate components (e.g., a stainer coupled to the slide preparation machine 120 that includes a scanner; a scanner coupled to the slide preparation machine 120 that includes a stainer). In some embodiments, the process 180 may include a sample 126 provided to the slide preparation machine 120. In particular, the sample 126 may be provided to the stainer 122, which may perform one or more staining processes to prepare the slide for digitization. The prepared slide may then be scanned via the scanner 124 to obtain an image 128.
[0065] In some embodiments, the stainer 122 may apply a stain (e.g., H&E) to the slide of the sample 126. Different stains may be used by different stainers 122. In addition, even for the same stain, the concentration, amount, or other aspects of the application of the stain may vary from sample to sample and from stainer to stainer.
[0066] Before staining, the biological sample can be cut into different thicknesses. The cutting can be performed by a machine, but a person can also cut the biological sample. Whether performed by a machine or a person, the cutting itself may introduce variations that affect downstream analysis. For example, one biological sample can be cut to a thickness of 2 mm, while another biological sample can be cut to a thickness of 4 mm. The captured digital pathology images of these biological samples, even from biopsies of the same source, may contain different characteristics. However, the basic morphological properties of the samples should remain the same because these samples are from the same source.
[0067] In addition to the different tissue processing techniques that can be used to process a sample with a stainer 122, the slide preparation machine 120 can also impart different tissue processing characteristics. Similar to the tissue processing techniques described above, the characteristics specific to the slide preparation machine 120 can also be encoded into the image 128, which can affect downstream classification. For example, different scanners 124 can capture an image 128 of the sample 126 using different lenses, different magnification settings, different lens thicknesses, different software, or different digitization techniques. Some embodiments include an image 128 having multiple magnifications (e.g., 5x, 10x, 20x, 40x, etc.). In this case, the image 128 can form a pyramid-like format. In some embodiments, one of the magnifications can be selected as the target magnification (e.g., 40x). Tiles can be obtained from the image at this fixed magnification. Some embodiments include using images from multiple magnifications together. Additionally, the slide preparation machine 120 can have different brands, models, versions, or other differences that can be imparted to the image 128.
[0068] As described herein, a scanner refers to a computing system and imaging system capable of scanning, digitizing, compressing, storing, retrieving, and / or viewing a slide of a biological sample. A scanner can include a portion into which the sample can be loaded into the machine, an image capture component, a processor, a memory, a network interface, a display (or other input / output device), or other components or combinations thereof. Some scanners are capable of loading different numbers of slides. For example, a scanner can load 100 or more slides, 200 or more slides, 500 or more slides, or other numbers. Depending on the scanner, software, lens, or other considerations, the image capture speed of these slides and the magnification used to capture these images can vary.
[0069] As an example, refer to Figure 2A, the image database 142 can be configured to store image data 200. The image data 200 can include N images 202 (e.g., whole slide images) depicting P biological samples. For each biological sample, one or more images can be captured. For example, for each of the P biological samples, one or more images (e.g., ten or more images, one hundred or more images, one thousand or more images, etc.) can be captured. Each image can be processed and digitized using one of a set of predefined slide preparation machines (e.g., slide preparation machine 120). For example, the slide preparation machine 120 can include M slide preparation machines. Each slide preparation machine can include a slide stainer and a slide scanner. Thus, there can be M slide stainers and M slide scanners. In some embodiments, some or all of the M slide stainers can be used to prepare each biological sample. For example, a first stainer can use a first staining technique to prepare a slide of a first biological sample, while a second stainer can use a second staining technique to prepare a slide of the first biological sample. In some embodiments, some or all of the M scanners can be used to scan each biological sample. For example, for a first biological sample, scanner 1 can be used to capture image 1 of biological sample 1, while scanner 2 can be used to capture image 2 of biological sample 1.
[0070] In some embodiments, the image data 200 can be organized in a data structure based on the corresponding biological samples represented by the images. For example, all the captured images of biological sample 1 can be stored in association with each other, while all the captured images of biological sample N can be stored in association with each other. Those of ordinary skill in the art will recognize that alternative organization schemes can be used, and the above scheme is only an example.
[0071] In some embodiments, each image 202 can have metadata associated therewith. For example, each image 202 can include processing technique tags (e.g., scanner tags, staining tags), sample identifiers, timestamps indicating when a particular image was captured, slide tags indicating the slide used to capture the image, tissue thickness, geographical location, operator identifiers, or other information or combinations thereof. Alternatively, only some of the images 202 can include processing technique tags, sample identifiers, or other metadata.
[0072] In some embodiments, the processing technique tags can indicate values of attributes associated with the processing of a particular image. For example, the processing technique tag 204-1 can indicate one or more processing techniques used to create the image 202-1. Although only a single processing technique tag is shown in Figure 2A only, those of ordinary skill in the art will recognize that multiple processing technique tags can be associated with a given image. By way of example, referring to Figure 2B, the processing technology tag 204 may include a scanner tag 252, a stainer tag 254, a sample thickness tag 256, or other metadata indicating an attribute value associated with a slide preparation machine used to prepare the slide and the corresponding image. For example, the scanner tag 252 may indicate the scanner used to capture an image of a given biological sample; the stainer tag 254 may indicate the stain and / or process used to stain the slide of the biological sample; the sample thickness tag 256 may indicate the thickness of the biological sample, and so on. In some embodiments, a magnification tag indicating a fixed magnification used for training is also included.
[0073] The scanner tag 252 may include metadata associated with the image, which may indicate the scanner used to capture the corresponding image. For example, the processing technology tag 204-1 associated with the image 202-1 may include a first scanner tag (e.g., the scanner tag 252) indicating that the first scanner is used to capture the image 202-1. Similarly, the processing technology tag 204-2 associated with the image 202-2 may include a second scanner tag indicating that the second scanner is used to capture the image 202-2, and the processing technology tag 204-M associated with the image 202-N may include an Mth scanner tag indicating that the Mth scanner is used to capture the image 202-N.
[0074] In some embodiments, two or more images 202 may be prepared using the same slide preparation machine or the same components of the slide preparation machine. For example, two or more images 202 may be captured by the same scanner. In this case, the corresponding scanner tags may include the same information. For example, if both the images 202-1 and 202-2 are captured using Scanner 1, the corresponding scanner tags may be equivalent or otherwise specify the same scanner (e.g., device identifier, port address, etc.).
[0075] The sample ID 206 indicates the biological sample depicted by the corresponding image. For example, the image 202-1 may include a sample ID 206-1 indicating the biological sample depicted by the image 202-1. Similarly, the image 202-2 may include a sample ID 206-2 indicating the biological sample depicted by the image 202-2, and the image 202-P may include a sample ID 206-P indicating the biological sample depicted by the image 202-P. In some embodiments, two or more of the images 202 may depict the same biological sample. In this case, the corresponding sample IDs may include the same information. For example, if both the images 202-1 and 202-2 are images of a first biological sample, the sample ID 206-1 and the sample ID 206-2 may be equivalent or otherwise specify the same biological sample (e.g., slide number, clinical trial information, etc.).
[0076] In some embodiments, the training data generation subsystem 110 may be configured to create training data and / or validation data for training a machine learning model based on the image data 200. For example, the training data generation subsystem 110 may organize the image data 200 such that images related to the same biological sample are grouped together, images captured by the same scanner are grouped together, biological samples related to the same clinical trial or treatment group are grouped together, etc. As an example, referring to Figure 3 , the training data database 144 may include training data 300 generated by the training data generation subsystem 110 based on the image data 200. In some embodiments, the training data generation subsystem 110 may be configured to identify the biological sample associated with each of the images 202. For example, a sample identifier indicating the biological sample associated with a given image may be extracted from the metadata of each image. The training data generation subsystem 110 may select images having similar sample IDs and may group these images together. For example, as Figure 3 shown, the images 202-1, 202-2, 202-X, and 202-Y may each be grouped together based on each of these images having the same sample ID 206-1. The sample ID 206-1 may refer to biological sample 1, and thus each of the images associated with biological sample 1 may be grouped together.
[0077] In some embodiments, the training data generation subsystem 110 may be configured to further organize the images based on the slide preparation machine used to prepare and / or digitize the biological sample. As an example, the training data generation subsystem 110 may organize the images based on the scanner source. The training data generation subsystem 110 may detect the scanner tag associated with the image 202 and group together the images captured by the same scanner. As an example, referring to Figure 3, images 202-1 and 202-2 can be grouped into a first image set 310a, and images 202-X and 202-Y can be grouped into a second image set 310b. Each of images 202-1, 202-2, 202-X, and 202-Y is related to the same biological sample indicated by sample ID 206-1. However, in this image collection, the training data generation subsystem 110 can group the images based on the corresponding processing technology label (e.g., scanner label) of each image. For example, the first image set 310a can include images 202-1 and 202-2 based on each of images 202-1 and 202-2, and images 202-1 and 202-2 have the same processing technology label, i.e., processing technology label 204-1. This can indicate that images 202-1 and 202-2 are prepared and processed by the same slide preparation mechanism. For example, in addition to being images depicting the same biological sample (e.g., biological sample 1), images 202-1 and 202-2 can be captured using the same scanner. As another example, the second image set 310b can include images 202-X and 202-Y based on each of images 202-X and 202-Y, and images 202-X and 202-Y have the same processing technology label, i.e., processing technology label 204-2. This can indicate that images 202-X and 202-Y are prepared and processed using the same slide preparation mechanism. For example, in addition to being images depicting the same biological sample (e.g., biological sample 1), images 202-X and 202-Y can be captured by the same scanner.
[0078] In some embodiments, the training data 300 can include images related to N biological samples. For each biological sample, one or more image sets can be generated by the training data generation subsystem 110. In addition, each slide preparation machine can prepare and process the same biological sample. For example, each scanner can be used to capture one or more images of the same biological sample (however, only some scanners can be used to capture some biological samples).
[0079] In some embodiments, different tissue schemes can be used by the training data generation subsystem 110. For example, in some embodiments, an image set including images captured by multiple scanners can be generated - similar to the combined first image set 310a and second image set 310b. In addition, an image set including images depicting multiple biological samples can be generated. For example, an image set including images of two or more biological samples can be generated.
[0080] Each image set can include one or more, one hundred or more, one thousand or more, one million or more, or other numbers of images. Thus, although image sets 310a and 310b are shown as including two images, each corresponding image set can include more images. Additionally, the image sets can include different numbers of images. The number of biological samples scanned can be any number of samples and can depend on the clinical trial. For example, N biological samples can include one or more biological samples, ten or more biological samples, one hundred or more biological samples, one thousand or more biological samples, or other numbers.
[0081] It should be understood that, as described herein, the processing techniques associated with different slide preparation machines can be different. For example, different scanners can have different settings, perform different tasks, be made by different entities, or in different ways, and different stainers can use different staining techniques, different stains, different tissue thicknesses, or have other differences. Thus, one of ordinary skill in the art will recognize that other differences can be used to develop a sample-independent model. For example, a sample-independent model can be independent of the stain, staining technique, tissue thickness, etc. In these cases, the image can include associated metadata that indicates, for example, the stain used for the biological sample depicted by the image.
[0082] In some embodiments, the training data generation subsystem 110 can be configured to generate validation data for validating (e.g., testing) a model. The validation data can include images of biological samples captured using a known slide preparation machine. The validation data can be used to test the accuracy of the model to correctly guess the corresponding slide preparation machine (e.g., scanner source) of the image, or to be unaffected by the slide preparation machine of the image and / or the process performed thereby when used for downstream classification.
[0083] As described above, the inherent differences between data sets captured using different methods or protocols can affect the ability of the model to perform downstream classification. A known problem with traditional digital pathology analysis is that the features learned by the model can be scanner-independent. For example, reference Figure 18A, Figure 1800 shows an exemplary t-Distributed Stochastic Neighbor Embedding (TSNE) map embedded from multi-scanner data. Figure 1800 can be created by generating an embedding from image data, including images of biological samples captured using two scanners (however, other properties of the slide preparation process may also vary). Each scanner is represented by one of the colors in Figure 1800. For example, the embedding of one scanner is represented by the "yellow" or light-colored clusters, and the embedding of the other scanner is represented by the "purple" or dark-colored clusters. As shown in Figure 1800, these two clusters are easily distinguishable, indicating that although the same biological sample is imaged by the same two scanners, the scanner affects the embedding generation process, and the affected embedding affects the downstream classification task. Therefore, there are technical problems in generating a model that is configured to perform downstream digital pathology classification tasks and is independent of the scanner used to capture the images and / or other non-morphological characteristics of the images (such as stain, magnification, tissue sample thickness, etc.).
[0084] Described herein are technical solutions for overcoming the above technical problems. In some embodiments, a sample-processing-independent model for performing digital pathology classification tasks is developed. The model can adopt a student-teacher framework as the backbone. An example of such a backbone is the "Unsupervised Label Propagation" or DINO framework. Another example of the backbone is the "Self-Guided Latent Representation" or BYOL framework. Both frameworks and enhancements to the frameworks are described herein to develop a sample-processing-independent model.
[0085] In some embodiments, the encoder selected as the model architecture is a convolutional neural network. For example, the ResNet-18 architecture can be used as the base framework for the encoder. ResNet-18 includes 18 layers, which are organized into four residual blocks. A residual block is a block that applies an identity mapping: the input of one layer is also directly passed to another layer. In some embodiments, each residual block is connected to the next layer in the network as well as a layer skipped further down. The connection between the residual block and the next network layer is called a shortcut or skip connection, which can bypass one or more layers. Mathematically, if the input x is the input of a layer and the output is F(x), the output of the residual block can be expressed as Y = F(x) + x. In particular, ResNet-18 is a trained convolutional neural network CNN that is trained on images from the ImageNet database to classify images into one of 1000 categories. The input image size of ResNet-18 is 256x256.
[0086] DINO Base Framework
[0087] In some embodiments, the underlying framework for building a model independent of sample processing is the DINO framework. DINO is a self-supervised learning method based on knowledge distillation that enables a network to mimic the output of another network, resulting in information propagation from a small set of annotations to a large unlabeled database. The knowledge distillation process can even be extended to use cases where images lack labels. For a complete description of the DINO framework, reference is made to "Emerging Properties in Self-Supervised Vision Transformers" by Caron et al. in 2021, the content of which is incorporated herein by reference in its entirety.
[0088] As an example, referring to Figure 4 , the DINO framework 400 includes two networks, a student network formed by a student encoder 410 and a softmax layer 412, and a teacher network including a teacher encoder 420, a softmax layer 422, and an intermediate layer 424. An image 402, such as an image of a biological sample, can be obtained (e.g., from the image database 142). In some embodiments, the image 402 may include metadata indicating the scanner used to scan the biological sample and generate the image 402. Additionally or alternatively, the image 402 may include metadata indicating the stainer used to apply the stain to the biological sample, the thickness of the biological sample, the magnification used, or other slide preparation information. In some embodiments, one or more augmentations may be performed on the image 402 to obtain augmented images 404a and 404b. Various augmentations that can be performed on the image 402 include, but are not limited to, flipping, blurring, cropping, color jittering, or other image augmentations or combinations thereof. In some embodiments, the training data generation subsystem 110 may be configured to perform augmentations on the image 402 to obtain augmented images (which may also be interchangeably referred to as "augmented views") 404a, 404b. In some embodiments, the augmented images 404a and 404b may have different augmentations performed on them. For example, the augmented image 404a may include a small crop (which may be interchangeably referred to herein as a "local crop"), while the augmented image 404b may include a large crop (which may be interchangeably referred to herein as a "global crop"). By providing the local crop to the student encoder 410 while providing the global crop (and in some embodiments, the local crop) to the teacher encoder 420, a local-to-global mapping can be obtained.
[0089] The student encoder 410 and the teacher encoder 420 may have the same architecture. For example, the student encoder 410 may include a first set of parameters θ s , and the teacher encoder 420 may include a second set of parameters θ t。The student encoder 410 is trained to match the output of the teacher encoder 420, and its hyperparameters are updated based on the average value (e.g., exponential moving average) of the student encoder 410. In other words, the parameter θ can be learned based on the parameter θ s to learn the parameter θ t 。In some embodiments, the parameter θ can be learned by minimizing the loss function (Equation 1) using stochastic gradient descent s :
[0090]
[0091] The model training subsystem 112 can be configured to train the student encoder 410 and the teacher encoder 420. In some embodiments, each of the student encoder 410 and the teacher encoder 420 can output an embedding representing the input image. The model training subsystem 112 can be configured to pass the embedding from the student encoder 410 to the softmax layer 412, which outputs p1. The output p1 can include elements that indicate the likelihood that the image depicts content related to K categories. This likelihood can be calculated using Equation 2:
[0092]
[0093] In Equation 2, τ represents the temperature parameter that controls the sharpness of the output distribution, and g represents the student or teacher network.
[0094] The model training subsystem 112 can be configured to pass the embedding output from the teacher encoder 420 to the centering layer 424. The centering layer 424 can be used to center the output embedding from the teacher encoder 420 with the average value calculated over a batch of augmented views.
[0095] The model training subsystem 112 can be configured to maximize the similarity of the softmax temperatures of the student network and the teacher network, as shown by the outputs p1 and p2. Using the notation of Equation 2, the outputs p1 and p2 can be represented as probabilities P s (x1) and P t (x2). The similarity can be measured as the cross-entropy loss, represented by Equation 3:
[0096]
[0097] where H(a,b) = -a log(b) Equation 3.
[0098] In some embodiments, the technical problem of Equation 3 can be applicable to more views than the local-global cropping strategy of the DINO framework. For example, different views of an image (e.g., enhanced image 404a, image 404b) can be created by applying one or more image augmentations (such as, for example, blur, flip, rotation, color distortion, and / or cropping). The enhanced images can also be interchangeably referred to as "enhanced views" herein. The augmentation can be performed by the training data generation subsystem 110 and / or the model training subsystem 112.
[0099] For one sample, the result of augmentation is a set V of distorted views consisting of two global crops (denoted as x g1 and x g2 ) and several local crops. In some embodiments, all crops can be provided to the student encoder 410, while only the global crops can be passed to the teacher encoder 420. This enables learning of local-to-global correspondences.
[0100] The model training subsystem 112 can be further configured to solve the optimization problem of DINO, which can be represented by Equation 4:
[0101]
[0102] In Equation 4, x and x' refer to the enhanced views of the images provided to the teacher encoder 420 and the student encoder 410, respectively.
[0103] In some embodiments, the model training subsystem 112 can be configured to propagate gradients through the student encoder 410. In some cases, the gradients can be propagated only through the student encoder 410. The model training subsystem 112 can apply gradient stopping to the teacher network (e.g., teacher encoder 420, centering layer 424, softmax layer 422) to prevent backpropagation. In some embodiments, the model training subsystem 112 can use a momentum encoder to dynamically construct the teacher network (e.g., teacher encoder 420). As an example, the parameter θ t can be updated with the exponential moving average (ema) of the weights (parameters θ s ) of the student encoder 410.
[0104] BYOL framework
[0105] In some embodiments, the basic framework for constructing a model independent of sample processing is the BYOL framework. The BYOL model is a self-supervised image representation learning process. The BYOL architecture includes two neural networks: an "online" neural network and a "target" neural network. The online and target neural networks interact and learn from each other. For example, for a given image, an augmented version of the image can be created, and the first augmented version of the image is used to train the online neural network to predict the target neural network representation of the second augmented version of the image. The principle behind this process is that the representation of one augmented view of an image should predict the representation of a different augmented view of the same image. Thus, the BYOL process includes training a model to generate enhanced representations by predicting target representations using the target model.
[0106] As an example, referring to Figure 5 , the BYOL framework 500 is presented. In some embodiments, the framework 500 is configured to learn a representation y θ (e.g., for classifying images). The framework 500 may have a first encoder 510 and a second encoder 530. The first encoder 510 and the second encoder 530 may be interchangeably referred to herein as the "online encoder 510" and the "target encoder 530". In some embodiments, the first encoder 510 and / or the second encoder 530 may be a convolutional neural network.
[0107] The first encoder 510 may be defined by a set of weights θ and may include a first stage: encoding f θ , a second stage: projection g θ and a third stage: prediction q θ . The second encoder 530 may be defined by another set of weights ξ and may include similar first and second stages: encoding f ξ and projection g ξ . In some embodiments, the model training subsystem 112 may configure the second encoder 530 to train the first encoder 510, where the weight ξ is the exponential moving average of the weight θ. Thus, after each round of training, the model training subsystem 112 may update the weight ξ according to Equation 5:
[0108] ξ←τξ+(1 - τ)θξ←τξ+(1 - τ)θ Equation 5.
[0109] In some embodiments, the model training subsystem 112 can be configured to use a framework 500 to take an image 502 and apply two (or more) image augmentations to the image 502 to generate two augmented versions of the image 502 (augmented view 504, v and augmented view 506, v'). The model training subsystem 112 can randomly select the image 502 from a set of training images (e.g., a set of whole slide images depicting biological samples). Some embodiments include the image 502 as a sample image to be analyzed by the framework 500, or a downstream classifier trained based on the framework 500. For example, the first encoder 510 can be used to classify images after being trained. The augmentations t and t' applied to the image 502 produce a first augmented view 504, v, and a second augmented view 506, v', of the input image 502. In some embodiments, the model training subsystem 112 can be configured to select the augmentations t, t' from a set of predefined augmentations. The augmentation t can be selected from a first set of predefined augmentations, while the augmentation t' can be selected from a second set of predefined augmentations. In some cases, the first and second sets of predefined augmentations can share one or more common types of augmentations (e.g., both sets can include rotation augmentation), however, the sets of these predefined augmentations may not share any common types of augmentations.
[0110] The model training subsystem 112 can be configured to provide the first augmented view v and the second augmented view v' to the encoder 510 and the encoder 530, respectively. Then, the first augmented view v and the second augmented view v' can be encoded using the encoder 510 and the encoder 530 to produce a first representation y θ (such as an embedding 512) and a second representation y' ξ (such as an embedding 532). For example, the encodings of the first encoder 510 and the second encoder 530 are defined by equations 6a, 6b, respectively:
[0111]
[0112]
[0113] The model training subsystem 112 can provide the first representation y θ (e.g., an embedding 512) and the second representation y' ξ (e.g., an embedding 532) to the first (online) projector 514 and the second (target) projector 534. The embeddings 512 and 532 can then be used by the first projector 514 and the second projector 534 to generate a first projection 516 and a second projection 536, respectively, where the first projection 516 is denoted as z θ and the second projection 536 is denoted as z' ξFor example, the projections of the first projector 514 and the second projector 534 are defined by equations 7a and 7b, respectively:
[0114]
[0115] In some embodiments, the online predictor 518 may be configured to generate a prediction 520 represented by q θ (z θ ). The prediction 520 is a prediction of the second projection 536. After generating the prediction 520, the model training subsystem 112 may be configured to normalize both the prediction 520 and the projection 536 (e.g., l2 normalization) to obtain:
[0116]
[0117] The loss function (also interchangeably referred to as the BYOL loss) is defined as:
[0118]
[0119] After each training step, the model training subsystem 112 may update the weights according to equations 5 and 10:
[0120]
[0121] In equation 11, ηη is the learning rate, which may be predefined, and where refers to the symmetric loss function obtained by feeding the first augmented view 504 into the first encoder 510 and the second view 506 into the second encoder 530.
[0122] After training is complete, the model training subsystem 112 may store the encoding f of the first encoder 510 θ in the model database 146. The encoding f of the first encoder 510 θ can be used to train a classifier to classify images. In particular, the downstream classification subsystem 114 may be configured to train a classifier to classify images of the biological sample encoding f of the first encoder 510 stored in the model database 146 θ . In some embodiments, the encoding f of the first encoder 510 θ may be stored in the model database 146 together with metadata. For example, the metadata may indicate the time when the first encoder 510 was trained and / or validated.
[0123] In some embodiments, the augmentations that can be applied to the input image 502 can include, but are not limited to, this does not mean that other lists are random cropping, limiting, flipping around one or more axes, color distortion, adjustment of at least one of image brightness, contrast, saturation, or hue, conversion to grayscale, blurring, or other augmentations or combinations thereof.
[0124] In some embodiments, the encoders 510 and 530 can be convolutional neural networks (CNNs) with multiple layers. For example, the encoders 510 and 530 can include 3 or more layers, 10 or more layers, 50 or more layers, 100 or more layers, 200 or more layers, etc. An exemplary deep learning model that can be used for the encoders 510 and 530 is ResNet (e.g., ResNet-18).
[0125] Two independent techniques for developing sample - processing - agnostic models are described herein. The first technique, namely the mean - embedding similarity method, can be implemented without modifying the DINO or BYOL frameworks. The second technique, namely the adversarial method, produces more accurate results when used for downstream classification but requires modification of the DINO or BYOL frameworks.
[0126] Mean - embedding similarity method.
[0127] In some embodiments, the available data (such as the image data 200) can include images of biological (e.g., tissue) samples that have been processed using two (or more) slide preparation machines (e.g., scanned with different scanners, stained with different stainers, etc.). For example, as Figure 3 shown, the training data 300 can be organized such that each biological sample includes one or more images captured by each available slide preparation machine (e.g., slide preparation machine 120). In some embodiments, each slide preparation machine can be configured to stain biological samples and / or capture images of slides including stained biological samples. The training data generation subsystem 110 can be configured to organize the images 202 - 1 and 202 - 2 into a first image set 310a because both the images 202 - 1 and 202 - 2 have the same processing technique label 204 - 1 (e.g., both the images 202 - 1 and 202 - 2 can be captured by a first scanner using a first stain, a first magnification, a first tissue thickness, etc.). Similarly, the images 202 - X and 202 - Y of biological sample 1 can also have the same processing label 204 - 2 (e.g., both the images 202 - X and 202 - Y are captured by a second scanner using a second stain, a second magnification, a second tissue thickness, etc.).
[0128] In some embodiments, the training data generation subsystem 110 can be configured to divide each image (e.g., image 202) into a plurality of tiles. Each tile can have the same size and / or shape and, in some cases, can overlap with each other. For example, each tile can represent a 512x512 pixel block. In some embodiments, the image from which the tiles are obtained can have a larger size (e.g., 100,000x100,000 pixels). For example, each tile can be formed by dividing a WSI image into a set of tiles, each having the same size (which may or may not overlap). In some embodiments, the tiles can be randomly selected from a larger image.
[0129] In some embodiments, random augmentation can be performed on the tiles. For example, an image (e.g., a whole slide image) can be divided into tiles, and some or all of the tiles can be augmented using one or more predefined image augmentations. Alternatively, the image can be augmented and then tiles can be obtained therefrom. For this, in some embodiments, the images included in the training data 300 can correspond to image patches (e.g., obtained by dividing an image into tiles), however, in some cases, the training data 300 can include images (e.g., whole slide images) and / or image patches. Additionally, the training data 300 can further include augmented views of the images and / or image patches. For example, image 202-1 can correspond to an image patch to which one or more augmentations have been applied.
[0130] To develop a slide preparation machine-independent model, the model needs to be trained to correctly identify the outputs of two tiles from the same tissue but prepared using different slide preparation machines (e.g., scanned by different scanners). A technical solution to this technical problem is to modify the loss function associated with the model framework (e.g., framework 400 or framework 500) such that embeddings that are too far from other embeddings corresponding to the same tissue sample are penalized.
[0131] A technique that can be used to perform the penalty is a tile-level technique. A tissue sample scanned by a scanner can be considered an image augmentation, which can be used during the creation of a set of augmented views at the beginning of the framework. Similarly, a tissue sample prepared by one slide preparation machine with a set of attributes can be considered a form of image augmentation. In some embodiments, the set V of crops (e.g., views) can include tiles corresponding to two tiles from images prepared using different slide preparation machines (e.g., different scanners). As an example, refer to Figure 6A, process 600 includes a biological sample 602 (e.g., a tissue sample) provided to a first slide preparation machine 610 and a second slide preparation machine 620. As described above, the first slide preparation machine 610 may include a first slide staining machine (e.g., a stainer) and a first slide scanning machine (e.g., a scanner), and the second slide preparation machine 620 may include a second slide staining machine and a second slide scanning machine. The attributes of each slide preparation machine may be different. For example, the scanner associated with the first slide preparation machine 610 may be different from the scanner associated with the second slide preparation machine 620. As another example, the stainer associated with the first slide preparation machine 610 may be different from the stainer associated with the second slide preparation machine 620. In some embodiments, the biological sample 602 may be prepared on a slide and stained with a slide stained with hematoxylin and eosin (H&E). The slide preparation machines 610 and 620 (or the scanners associated therewith) may be configured to output a first image 612 and a second image 622, respectively, each of the images depicting the biological sample 602.
[0132] In some embodiments, the image 612 may be divided into a plurality of image patches 614a to 614d and 624a to 624d, respectively (e.g., via the training data generation subsystem 110). Although only four image patches are presented in Figure 6A , those of ordinary skill in the art will recognize that other numbers of image patches may be generated for a given image, and the image patches may overlap. In some embodiments, corresponding image patches from different images captured using different slide preparation machines but from the same biological sample may be identified and included in the image set 630. For example, the image patch 614d from the image 612 and the image patch 624d from the image 622 may each be included in the image set 630. In some embodiments, each image patch in the image set 630 may include an identifier indicating the slide preparation machine used to capture and / or prepare the image and an identifier indicating the biological sample represented by the corresponding image of the patch.
[0133] In some embodiments, one or more image enhancements may be used to enhance some or all of the image patches included in the image set 630. For example, the image patches 614d and / or 624d may be rotated, flipped, blurred, have their colors jittered, etc. The enhanced tiles may also include several global and local crops of the image. However, this technique may not be feasible because the two tiles being analyzed should contain the same part of the biological sample (e.g., the image patches 614d and 624d should depict the same part of the biological sample 602). Registering the similarity in tissue samples depicted by different image patches may be impractical for tracking a large number of images and image patches, so slide-level techniques may be used instead.
[0134] In slide-level techniques, some embodiments do not penalize the distance between embeddings at the tile level (e.g., measure the similarity difference as a distance in the latent space), but rather penalize the distance at the slide level. For example, a model in a tile-level technique is configured to generate embeddings based on image tiles (e.g., tiles 614a to 614d of image 612). Then, the parameters of the model can be adjusted such that the similarity of the embeddings of tiles representing a common part of the biological sample (e.g., tiles 614d and 624d) is maximized. In slide-level techniques, the similarity between the embeddings of each image generated by each scanner can be maximized. For example, the similarity between the embedding generated based on image 612 and the embedding generated based on image 622, each of the embeddings representing biological sample 602, can be maximized. In some embodiments, the maximized similarity is the average tile embedding of two corresponding whole-slide images. For example, the embeddings of each of image tiles 614a to 614d are calculated and then averaged together to obtain a first representation of image 612. A similar process can be performed on image tiles 624a to 624d to obtain a second representation of image 622.
[0135] In some embodiments, a loss can be calculated based on the first and second representations. In some embodiments, a tile batch including image tiles can be created (and for calculating the loss, new shards of the tiles in the mini-batch can be performed in each training epoch. The batch can be formed by tiles corresponding to the same biological sample but from different source images, rather than randomly creating a batch of tiles. For example, a batch of image tiles can be created that represents a part of biological sample 602 but is captured by a first scanner (associated with first slide preparation machine 610) and a second scanner (associated with second slide preparation machine 620). For example, the batch of image tiles can include image tile 614d and image tile 624d. In some embodiments, if multiple available image tiles are to be selected from two or more images, where each tile corresponds to a tile from another image, the image tiles included in the batch can be randomly selected. As an example, image tiles 614a to 614d and image tiles 624a to 624d can be created, each image tile corresponding to images 612 and 622 captured and / or prepared by slide preparation machines 610 and 620 of biological sample 602, respectively. Thus, to create a batch, the training data generation subsystem 110 can randomly select one or more pairs of image tiles, such as tiles 614a and 624a, 614b and 624b, 614c and 624c, and 614d and 624d. Each batch can include images - represented by {x 1 ,,x 2n}, where each element x i has at least one corresponding image x j(e.g., captured by the i-th and j-th scanners). On a given batch, the loss function can be minimized. An example of the loss function is Equation 12:
[0136]
[0137] where
[0138] In Equation 12, g s represents the student encoder, θ s represents the parameters of the student encoder, and λ represents the proportion of the added term. The first term, L DINO is given by Equation 4.
[0139] Figure 6B FIG. is a schematic diagram showing different slide preparation machines according to various embodiments and different settings associated with these slide preparation machines for generating images representing the same biological sample. In system 650, a biological sample can be provided to one or more of slide preparation machines 120. For example, biological sample 652 can be prepared on a slide and provided to a first slide preparation machine 120-1, a second slide preparation machine 120-2, and an M-th slide preparation machine 120-M. Each of slide preparation machines 120-1, 120-2, 120-M can include a slide stainer and a slide scanner. For example, slide preparation machine 120-1 can include stainer 122-1 and scanner 124-1, slide preparation machine 120-2 can include stainer 122-2 and scanner 124-2, and slide preparation machine 120-M can include stainer 122-M and scanner 124-M.
[0140] Each of the stainers 122-1, 122-2, 122-M can have different settings 660-1, 660-2, 660-M, respectively. Additionally, each of the scanners 124-1, 124-2, 124-M can have different settings 664-1, 664-2, 664-M. The settings 660-1, 660-2, 660-M can indicate different stains, tissue thickness, staining levels, or other properties of the corresponding stainer when preparing slides of a subsequently imaged biological sample (e.g., biological sample 652). When capturing an image of a biological sample slide, the settings 664-1, 664-2, 664-M can indicate different scanner types, models, manufacturers, software, software versions, light sources, lenses, lens materials, lens thicknesses, magnifications, or other properties of the corresponding scanner. In some embodiments, the slide preparation machines 120-1, 120-2, and 120-M can capture one or more images, such as images 654-1, 654-2, and 654-M, respectively. For example, a slide prepared by stainer 122-1 based on setting 660-1 can be provided to scanner 124-1, which can capture image 654-1 based on setting 664-1. Images 654-1, 654-2, and 654-M each depict biological sample 652; however, the images may include differences that affect downstream classification tasks if not compensated for by making the classification model independent of the slide preparation machine and the corresponding slide preparation and processing techniques used.
[0141] As an example, Tables 1 and 2 below show datasets that can be used to train a machine learning model. As can be seen from Table 1, different datasets may stain biological samples with different stainers, have tissue samples of different thicknesses, and have different staining levels and other differences.
[0142]
[0143] Table 1.
[0144]
[0145] Table 2.
[0146] Figure 7A flowchart showing an exemplary process 700 for training a model independent of sample processing using average similarity training techniques according to various embodiments. Process 700 may begin at operation 710. At operation 710, image data may be received. In some embodiments, the image data may include a first image set 712 and a second image set 716. The first image set 712 may include slides 714 that represent images (e.g., whole slide images) captured and prepared using a first slide preparation machine (e.g., slide preparation machine 120-1). For example, slide 714 may include a slide prepared using a first stainer and / or a slide captured using a first scanner. Each image in the first image set 712 may represent an image of one of a plurality of biological samples (such as tissue samples). For example, if ten biological samples are to be imaged, slide 714 may include ten slides, each depicting one of the ten biological samples. The second image set 716 may include slides 718 that represent images (e.g., whole slide images) captured using a second slide preparation machine (e.g., slide preparation machine 120-M). For example, slide 718 may include a slide prepared using a second stainer and / or a slide captured using a second scanner. Each image in the second image set 716 may represent an image of one of a plurality of biological samples (such as tissue samples). Continuing with the previous example, if ten biological samples are to be imaged, slide 718 may include ten slides, each depicting one of the ten biological samples. In some embodiments, each biological sample may have images captured and / or prepared using each available slide preparation machine. For example, for one biological sample, the first image set 712 may include one image depicting the biological sample, and the second image set 716 may include one image depicting the biological sample. The similarities and differences between these two images may be utilized to train a model independent of scanner-related and / or stainer-related embedding drivers to improve downstream biological sample classification. In some embodiments, operation 710 may be performed by a subsystem that is the same as or similar to the training data generation subsystem 110.
[0147] In operation 720, a first enhanced view set and a second enhanced view set can be generated. In some embodiments, the first enhanced view set and the second enhanced view set can correspond to the first image set 712 and the second image set 716, respectively. In some embodiments, the first enhanced view set and the second enhanced view set can include the images included in the first image set 712 and the second image set 716 (e.g., slides 714 and 718, respectively), with one or more image enhancements performed thereon. For example, the image enhancements that can be performed include rotation, flipping, blurring, color jittering, cropping, or others. In some embodiments, slides 714 and 718 can be divided into image patches. The image patches can overlap. In some embodiments, the image patches can have enhancements performed thereon. However, alternatively (or additionally), the enhancements can be performed before dividing the slides into image patches. In some embodiments, each enhanced view set can include images depicting each biological sample. For example, the first enhanced view set can include images of a first biological sample (including the image patches of the image), and the second enhanced view set can include images of the first biological sample (including the image patches of the image). In some embodiments, operation 720 can be performed by a subsystem that is the same as or similar to the training data generation subsystem 110.
[0148] In operation 730, a biological sample can be selected. The selected biological sample can be one of the biological samples whose images are captured by the first scanner and the second scanner. For example, biological sample 1 can be selected. Based on this selection, one or more enhanced views from the first enhanced view set corresponding to biological sample 1 and one or more enhanced views from the second enhanced view set corresponding to biological sample 1 can be selected. In some embodiments, operation 730 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0149] In operation 740, a first encoder can be trained. The first encoder can be implemented by a convolutional neural network, such as ResNet, a vision transformer, or other machine learning models. The first encoder can be part of the DINO framework 400, the BYOL framework 500, or other frameworks. In some embodiments, operation 740 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0150] Operation 740 may include one or more sub-steps. For example, operation 740 may include sub-step 742. In sub-step 742, a first representation of an enhanced view in a first enhanced view set may be generated using a first encoder. In some embodiments, the first representation may be an embedding. An embedding is a mapping of variables to vectors (arrays of numbers). An embedding is a vector representation that describes image features. An embedding is meaningful if it has features suitable for downstream classification or regression tasks. As an example, the embedding may be a 2048-dimensional vector, while the corresponding medical image (e.g., a whole slide image depicting a biological sample) includes data corresponding to a large number of pixels (e.g., tens of thousands of pixels, hundreds of thousands of pixels, millions of pixels, etc.). In some embodiments, the enhanced views in the first enhanced view set may include enhanced views of image patches of the corresponding image. For example, a first whole slide image from the first image set 712 may be divided into image patches, and then the image patches may have enhancements applied to them. Each of the image patches of the first whole slide image may be provided to the first encoder, and an embedding representing each image patch may be generated. In some embodiments, the embeddings representing the image patches may be combined to generate an average embedding representation of the corresponding image. As another example, a first whole slide image from the first image set 712 may have one or more enhancements applied to it and then subsequently be divided into image patches. The image patches may each be provided to the first encoder to generate embeddings, which may then be combined to generate an average embedding representation of the corresponding image. In an embodiment, if multiple images of the same biological sample are captured by a single scanner, the same process may be repeated, and an overall average embedding representation may be generated based on the corresponding average embedding representations of each image.
[0151] In sub-step 744, a second representation of the enhanced views in the second enhanced view set can be generated using a second encoder. In some embodiments, the second representation can be an embedding. In some embodiments, the enhanced views in the second enhanced view set can include enhanced views of image patches of the corresponding image. For example, a second whole slide image from the second image set 716 can be divided into image patches, and then the image patches can have enhancements applied to them. Each of the image patches of the second whole slide image can be provided to the second encoder, and an embedding representing each image patch can be generated. In some embodiments, the embeddings representing the image patches can be combined to generate an average embedding representation of the corresponding image. As another example, a second whole slide image from the second image set 712 can have one or more enhancements applied to it and then subsequently be divided into image patches. The image patches can each be provided to the second encoder to generate embeddings, which can then be combined to generate an average embedding representation of the corresponding image. In an embodiment, if multiple images of the same biological sample are captured by a single scanner, the same process can be repeated, and an overall average embedding representation can be generated based on the corresponding average embedding representations of each image.
[0152] In sub-step 746, the similarity between the first representation and the second representation can be calculated. In some embodiments, the similarity can be calculated based on a loss function. For example, the loss function used can be the DINO loss function or the BYOL loss function, depending on the framework used for training. For example, the loss function of the DINO framework is represented by Equation 12, while the loss function of the BYOL framework is represented by Equation 10 ( where refers to the symmetric loss function obtained by feeding the first view 504 into the first encoder 510 and the second view 506 into the second encoder 530).
[0153] In sub-step 748, the similarity between the first representation and the second representation can be maximized. Maximizing the similarity can include updating the hyperparameters of the first encoder and the second encoder. In some embodiments, backpropagation can be used to perform the update. In some embodiments, an optimizer (e.g., see Equation 11) can be used to optimize the loss function. For example, the Adam optimizer can be used.
[0154] In some embodiments, after performing operation 740 on the corresponding images (e.g., enhanced views of image patches) of a biological sample, the parameters of the second encoder can be updated based on the update of the first encoder parameters. For example, the exponential moving average (ema) of the weights and biases of the first encoder can be used to update the weights and biases and other parameters of the second encoder.
[0155] In operation 750, it can be determined whether any additional biological samples are to be analyzed. For example, if there are N biological samples, the first enhanced view set and the second enhanced view set can each include N whole slide images (e.g., enhanced versions of each of whole slide image 714 and whole slide image 718). Thus, for a selected biological sample, a round of training of the first encoder can be performed using the enhanced views (e.g., tiles obtained by partitioning the enhanced views) of the images of the selected biological sample from each of the first enhanced view set and the second enhanced view set. In some embodiments, the determination can be made after the parameters of the first encoder (and in some cases, the second encoder) have been updated. If it is determined in operation 750 that there are additional biological samples to be analyzed, process 700 can return to operation 730, where another biological sample can be selected and the first encoder can be trained using the enhanced views from each of the first enhanced view set and the second enhanced view set representing the biological sample (e.g., by repeating sub-steps 742 to 748 using the representations of the first enhanced view and the second enhanced view). However, if in operation 750 it is determined that there are no additional biological samples to be analyzed, process 700 can proceed to operation 760. In operation 760, process 700 can end and the trained model can be stored (e.g., in model database 146). In some embodiments, operations 750 and 760 can be performed by a subsystem that is the same as or similar to model training subsystem 112.
[0156] In some embodiments, if it is determined in operation 750 that there are no additional samples to be analyzed, test data (which can also be interchangeably referred to as validation data herein) can be used to test / validate the first encoder. For example, the first encoder can be tested using the validation data to determine the accuracy of the trained model for a new data classification task. In some embodiments, the accuracy of the trained encoder relative to the validation can be compared to a threshold accuracy level. If the accuracy is determined to be greater than or equal to the threshold accuracy level, the model can be stored in model database 146. For example, the trained model can be used for further downstream biological sample classification tasks.
[0157] Adversarial method
[0158] In some embodiments, the DINO framework and / or the BYOL framework can be modified to facilitate adversarial learning. Traditionally, the DINO and BYOL frameworks are designed for unlabeled data. However, for digital pathology tasks where the dataset includes images stained with multiple stainers, images captured using multiple scanners, etc., underlying labels may still exist. For example, metadata indicating the scanner label can be included in each image (and / or image patch) indicating the scanner used to capture the image. The scanner label can be used to ensure that the embeddings encoded by the corresponding encoder of the DINO / BYOL framework only retain morphological characteristics and do not retain or minimize the impact of the scanner source.
[0159] The DINO / BYOL framework can be modified to form an adversarial learning process where two networks compete with each other. These two networks include a "generator" and a "discriminator". The goal of adversarial learning is for the generator to generate results that can deceive the discriminator, while the discriminator aims to distinguish the output of the generator.
[0160] In the examples described herein, the student encoder of the DINO framework (e.g., student encoder 410) and the online encoder of the BYOL framework (e.g., online encoder 510) can be used as the generators of the adversarial DINO framework and the adversarial BYOL framework, respectively. In some embodiments, the encoder of the DINO / BYOL framework can be trained asynchronously with respect to the discriminator added for adversarial learning. For example, the parameters of the discriminator can be updated, and then the parameters of the encoder from the DINO / BYOL framework can be updated.
[0161] Adversarial DINO method
[0162] Figure 8 An exemplary adversarial framework 800 is shown according to various embodiments, which is used to train a model independent of scanner type when performing digital pathology classification. In some embodiments, the adversarial framework 800 can be built on the DINO framework depicted in the Figure 5 framework 400.
[0163] For example, the adversarial framework 800 can include two paths: one path includes a student encoder 410 and a softmax layer 412 (e.g., "student network"), and the other path includes a teacher encoder 420, a centering layer 424, and a softmax layer 422 (e.g., "teacher network"). An image (such as image 402) can have one or more image augmentations performed on it to obtain enhanced images 404a and 404b. In some embodiments, the enhanced images 404a, 404b can represent image patches of the image 402 on which image augmentation has been performed. As described above, reference Figure 5, the framework 400 is a self-supervised learning method based on knowledge extraction, which enables the network to imitate the output of another network, resulting in the spread of information from a small set of annotations to a large unlabeled database. The knowledge distillation process can even be extended to use cases where the images lack labels.
[0164] In some embodiments, the adversarial framework 800 adjusts the framework 400 such that the student encoder 410 acts as the generator 820 of the adversarial learning process. The generator 820 can be configured to encode the enhanced image 404a (which can be an image patch) to obtain the embedding 830. In some embodiments, the embedding 830 is the same as or similar to the embedding generated by the student encoder 410 using the framework 400 (for the same image). The adversarial framework 800 can provide the embedding 830 to the discriminator 860, which can be configured to produce a prediction 840. The prediction 840 includes the discriminator 860's prediction regarding which slide preparation machine was used to generate the image 402. For example, the prediction 840 can include a prediction of whether the image 402 was captured via a first scanner or a second scanner. As another example, the prediction 840 can include a prediction of whether the image 402 was stained with a first stain or a second stain, a first magnification or a second magnification, a first tissue thickness or a second tissue thickness, or other attributes of the slide preparation machine used to prepare and capture the image 402. In some embodiments, the discriminator 860 can be constructed as a linear layer with an input size equal to the dimension of the output embedding from the generator 820 (e.g., the embedding 830). For example, if the ResNet-18 architecture is used to implement the generator 820 (the student encoder 410), the dimension of the embedding 830 will be 512 dimensions, and the discriminator 860 can have an input size of 512 dimensions.
[0165] In some embodiments, the model training subsystem 112 can be configured to train the discriminator 860 to determine the embedding 830 of the slide preparation machine type output by the generator 820. Additionally, the model training subsystem 112 can be configured to train the generator 820 to produce a meaningful representation (e.g., an embedding) of the biological sample that can deceive the discriminator 860. The generator 820 utilizes the framework 400 (e.g., the DINO framework), and the presence of the adversary (e.g., the discriminator 860) can improve the created embedding (e.g., the embedding 830) such that it has greater value. However, in some embodiments, in addition to the first enhanced view 404a, the image 402 can also be provided to the student encoder 410. Additionally, as described in more detail below, in some embodiments, the adversarial framework 800 can include a gradient stop on the teacher network (indicated by the label "sg").
[0166] Compared to the average embedding similarity method, where the loss function is the only modification, the adversarial framework 800 requires a new approach to calculate the loss. The reason behind this is that another network, namely the discriminator 860, needs to be trained. Additionally, the training of the discriminator 860 and the student encoder 410 (and the teacher encoder 420) is not done simultaneously. Therefore, at each step (each mini - batch) of the training process, two different loss functions can be created and two backpropagations can be performed. Thus, some embodiments include a model training subsystem 112 that performs two training steps, as detailed below and with reference to Figure 9 and Figure 10 .
[0167] During the first step of training, the discriminator 860 can be updated. In some embodiments, the model training subsystem 112 can be configured to freeze the student encoder 410 (and the teacher encoder 420) so that only the parameters θ of the discriminator 860 are updated D . As Figure 9 shown in the first training step 900 of, the model training subsystem 112 can freeze the student encoder 410 and the teacher encoder 420, as indicated by the dashed lines, while allowing the reverse gradient propagation to be performed with respect to the discriminator 860. To freeze the student encoder 410 and the teacher encoder 420, the model training subsystem 112 can be configured to apply gradient stop "sg". Training the discriminator 860 is a fully - supervised classification task. The discriminator 860 is trained using the model training subsystem 112 to predict the type of the slide preparation machine (e.g., scanner) used to capture the image 402 based on the embedding 830. For example, the image 402 can be enhanced and tiled to obtain the enhanced image 404a, which can then be provided to the student encoder 410 to be encoded as the embedding 830.
[0168] In some embodiments, the biological sample depicted in the image can include a slide preparation label (e.g., slide preparation label 204 - 1). For example, the image 402 as described above can correspond to an image patch to which one or more image enhancements have been applied and can include a scanner label (e.g., scanner label 252) indicating whether the first scanner or the second scanner was used to capture the image depicting the biological sample. Mathematically, for a given sample (e.g., image 402) x that can include a scanner label y (or another label), the discriminator 860 can be configured to output an n - dimensional vector D(g s(x)). In this example, n refers to the number (i.e., category) of scanners used to capture images of biological samples. For example, if there are only two scanners, n = 2, the output of the discriminator 860 can be an indicator of whether the discriminator 860 predicts, based on the embedding 830, that the input image x was captured using the first scanner or the second scanner. In some cases, n can refer to the number of samples used to train the discriminator 860. In some embodiments, the model training subsystem 112 can be configured to minimize a loss function for training the discriminator 860. The loss function to be minimized can be a cross-entropy loss, defined by Equation 13:
[0169]
[0170] During the second step of the training process, the generator (e.g., generator 820) can be updated. In some embodiments, the model training subsystem 112 can be configured to perform a second training step 1000, as Figure 10 shown. In particular, during the second training step 1000, the student encoder 410 (and the teacher encoder 420) can be updated. During the second training step 1000, the model training subsystem 112 can freeze the discriminator 860 such that only the student encoder 410 and the teacher encoder 420 are updated, as indicated by the dashed lines for the discriminator 860. In some cases, the student encoder 410 can be updated and subsequently the teacher encoder 420 can be updated. For example, the parameters of the teacher encoder 420 can be updated with an exponential moving average of the parameters of the student encoder 410. The generator 820 (e.g., the student encoder 410) can be trained to produce a meaningful representation of the input image (e.g., input image patch) while also being able to deceive the discriminator 860 into incorrectly predicting the scanner used to capture the image.
[0171] In some embodiments, the model training subsystem 112 can be configured to combine the loss function previously used for the discriminator 860 with a conventional DINO loss function. The combined loss function is described by Equation 14:
[0172]
[0173] The combined loss function described in Equation 14 can be used by the model training subsystem 112 to minimize the DINO loss (e.g., the loss function of the framework 400) while maximizing the loss of the discriminator 860. In Equation 14, μ represents a parameter used to scale the additional loss (e.g., discriminator loss) compared to the DINO loss.
[0174] Adversarial BYOL method
[0175] Figure 11Another exemplary adversarial framework 1100 is shown in accordance with various embodiments for training a slide preparation machine - agnostic model when performing digital pathology classification. In some embodiments, the adversarial framework 1100 can be built on the BYOL framework depicted in the framework 500 of Figure 5 the framework 500.
[0176] For example, the adversarial framework 1100 can include two paths: one path includes an online encoder 510, an online (first) projector 514, and an online predictor 518 (e.g., "online network"), and the other path includes a target encoder 530 and a target (second) projector 534 (e.g., "target network"). An image (such as image 502) can have one or more image enhancements performed on it to obtain enhanced views 504 and 506. In some embodiments, the enhanced views 504, 506 can represent image patches of the image 502 on which image enhancement has been performed.
[0177] In some embodiments, the adversarial framework 1100 adjusts the framework 500 such that the online encoder 510 acts as the generator 1120 of the adversarial learning process. The generator 1120 can be configured to encode the enhanced view 504 (which can be an image patch) to obtain an embedding 512. The adversarial frame 1100 can provide the embedding 512 to a discriminator 1150, which can be configured to produce a prediction 1154. The prediction 1154 can include the discriminator 1150's prediction regarding which slide preparation machine was used to produce the image 502. For example, the prediction 1154 can include a prediction of whether the image 502 was captured via a first scanner or a second scanner. As another example, the prediction 1154 can include a prediction of whether the image 502 was prepared via a first stainer or a second stainer (or via a first staining agent or a second staining agent, which can be applied by the same or different slide staining machines). In some embodiments, the discriminator 1150 can be built as a linear layer of an input size equal to the dimension of the output embedding from the generator 1120 (e.g., embedding 512). For example, if a vision transformer architecture is used to implement the generator 1120 (student encoder 410), the dimension of the embedding 512 will be 512 - dimensional, and the discriminator 1150 can have an input size of 512 dimensions.
[0178] In some embodiments, the model training subsystem 112 can be configured to train the discriminator 1150 to determine the embedding 512 of the slide preparation machine type output by the generator 1120. Additionally, the model training subsystem 112 can be configured to train the generator 1120 to produce a representation (e.g., embedding) of a biological sample that is meaningful and can deceive the discriminator 1150. The generator 1120 utilizes the framework 500 (e.g., BYOL framework), and the presence of the adversary (e.g., discriminator 1150) can improve the created embedding (e.g., embedding 512) such that it has greater value.
[0179] Similar to the adversarial DINO method, the adversarial BYOL method also requires calculating losses. Additionally, the training of the discriminator 1150 and the online encoder 510 (and the target encoder 530) cannot be completed simultaneously. Therefore, at each step (each mini - batch) of the training process, two different loss functions can be created, and two backpropagations can be performed. Thus, some embodiments include the model training subsystem 112 that performs two training steps, as detailed below and with reference to Figure 12A and 12B .
[0180] During the first step of training, the discriminator 1150 can be updated. In some embodiments, the model training subsystem 112 can be configured to freeze the online encoder 510 (and the target encoder 530) such that only the parameters θ of the discriminator 1150 are updated. D . As Figure 12A shown in the first training step 1200 of, the model training subsystem 112 can freeze the online encoder 510 and the target encoder 530, as indicated by the dashed lines, while allowing the execution of backpropagation of gradients with respect to the discriminator 1150. To freeze the online encoder 510 and the target encoder 530, the model training subsystem 112 can be configured to apply the gradient stop "sg". In some embodiments, the discriminator 1150 is trained for a fully - supervised classification task. The discriminator 1150 is trained using the model training subsystem 112 to predict the scanner type used to capture the image 502 based on the embedding 512. For example, the image 502 can be augmented and tiled to obtain the augmented image 504, which can then be provided to the online encoder 510 to be encoded as the embedding 512.
[0181] In some embodiments, the biological sample depicted by the image may include a slide preparation label (e.g., slide preparation label 204-1). For example, the image 502 as described above may correspond to an image patch to which one or more image enhancements have been applied, and may include a scanner label (e.g., scanner label 252) indicating whether the first scanner or the second scanner was used to capture the image depicting the biological sample. Mathematically, for a given sample (e.g., image 502) v that may include label 1152, the discriminator 1150 may be configured to output an n-dimensional vector D(g s (x)), such as prediction 1154. In this example, n is the number (i.e., the class) of scanners used to capture the image of the biological sample. For example, if there are only two scanners, n = 2, the output of the discriminator 1150 may be an indicator of whether the discriminator 1150 predicts the input image v as being captured using the first scanner or the second scanner based on the embedding 512. In some embodiments, the model training subsystem 112 may be configured to minimize the loss function used to train the discriminator 1150. The loss function to be minimized may be the cross-entropy loss, defined by Equation 13.
[0182] During the second step of the training process, the generator (e.g., generator 1120) may be updated. In some embodiments, the model training subsystem 112 may be configured to perform a second training step 1250, as Figure 12B shown. In particular, during the second training step 1250, the online encoder 510 (and the target encoder 530) may be updated. During the second training step 1250, the model training subsystem 112 may freeze the discriminator 1150 such that only the online encoder 510 and the target encoder 530 are updated, as indicated by the dashed lines for the discriminator 1150. In some cases, the online encoder 510 may be updated and subsequently the target encoder 530 may be updated. For example, the parameters of the target encoder 530 may be updated with an exponential moving average of the parameters of the online encoder 510. The generator 1120 (e.g., the online encoder 510) may be trained to produce a meaningful representation of the input image (e.g., input image patch) while also being able to deceive the discriminator 1150 into incorrectly predicting the slide preparation machine used to prepare and capture the image.
[0183] In some embodiments, the model training subsystem 112 may be configured to combine the loss function previously used for the discriminator 1150 with a conventional BYOL loss function. The combined loss function may be described by Equation 14, except that the DINO loss term is replaced with the BYOL loss term. In some embodiments, the model training subsystem 112 may be used to update the BYOL framework (e.g., framework 500) through backpropagation of the combined loss of the BYOL network and the discriminator.
[0184] Figure 13 A flowchart showing an exemplary process 1300 for training a model independent of sample processing using adversarial training techniques according to various embodiments. Process 1300 may begin at operation 1310. At operation 1310, training data 1302 may be received. In some embodiments, training data 1302 may include a first image set 1304 and a second image set 1306. The first image set 1304 may include images (e.g., whole slide images) captured using a first slide preparation machine (e.g., slide preparation machine 120-1). For example, the images included in the first image set 1304 may include images captured using a first scanner and / or prepared using a first stainer. Each image in the first image set 1304 may represent one of a plurality of biological samples (such as tissue samples). For example, if ten biological samples are to be imaged, the image sets 1304 and 1306 may each include ten slides, each slide depicting one of the ten biological samples. The second image set 1306 may include images (e.g., whole slide images) captured using a second slide preparation machine (e.g., slide preparation machine 120-M). For example, the images included in the first image set 1306 may include images captured using a second scanner and / or prepared using a second stainer. Each image in the second image set 1306 may represent one of a plurality of biological samples (such as tissue samples).
[0185] In some embodiments, each biological sample may have images captured using each available scanner. For example, for one biological sample, the first image set 1304 may include one image depicting the biological sample, and the second image set 1306 may include one image depicting the biological sample. The similarities and differences between these two images can be utilized to train a model independent of scanner-related and / or stainer-related embedding drivers to improve downstream biological sample classification. In some embodiments, operation 1310 may be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0186] In some embodiments, some or all of the enhanced views of the images are included in the first image set 1304 and the second image set 1306. Each image in the first image set 1304 may include a first slide preparation label that indicates that the first slide preparation machine was used to prepare and capture the image, and each image in the second image set 1306 may include a second slide preparation label that indicates that the second slide preparation machine was used to prepare and capture the image. In some embodiments, the images included in each of the image sets 1304 and 1306 may correspond to enhanced views of the respective images. For example, the training data 1302 may include a first enhanced view set and a second enhanced view set. In some embodiments, the first enhanced view set and the second enhanced view set may correspond to the first image set 1304 and the second image set 1306, respectively. In some embodiments, the first and second enhanced view sets may include the images contained in the first image set 1304 and the second image set 1306 on which one or more image enhancements have been performed. For example, the image enhancements that may be performed include rotation, flipping, blurring, color jittering, cropping, or others. In some embodiments, the images in the first image set 1304 and the second image set 1306 may be divided into image patches. The image patches may overlap. In some embodiments, the image patches may have enhancements performed on them; however, alternatively (or additionally), the enhancements may be performed before the slide is divided into image patches. In some embodiments, each enhanced view set may include images depicting each biological sample. For example, the first enhanced view set may include images of a first biological sample (including the image patches of the image), and the second enhanced view set may include images of a first biological sample (including the image patches of the image). In some embodiments, the training data 1302 may include image patches (including the applied image enhancements), each image patch having a label associated with it that indicates one or more attributes associated with the slide preparation machine used to prepare and / or capture the corresponding image of the image patch. For example, each image may have a label indicating the scanner used to capture the corresponding image.
[0187] In some embodiments, a biological sample may be selected. The selected biological sample may be one of the biological samples prepared and / or captured via the first slide preparation machine or the second slide preparation machine. In one example, the selected biological sample may be one of the biological samples captured by the first scanner and the second scanner. For example, biological sample 1 may be selected. Based on this selection, one or more enhanced views (e.g., whole slide images, image patches) from the first enhanced view set corresponding to biological sample 1 and one or more enhanced views from the second enhanced view set corresponding to biological sample 1 may be selected.
[0188] In operation 1312, an image representation can be generated using a first encoder. Some embodiments include generating an image representation from a subset of images from the first image set 1304 and / or the second image set 1306. For example, the first encoder can refer to the online encoder 510. The images for generating the image representation can correspond to whole slide images or image patches, which can also include image enhancements applied thereto. In some embodiments, the images can depict a selected biological sample prepared and / or captured using a first slide preparation mechanism, where the first slide preparation machine is associated with a first label. Additionally or alternatively, the images can depict a selected biological sample prepared and / or captured using a second slide preparation mechanism, where the second slide preparation machine is associated with a second label. In some embodiments, the image representation can correspond to a generated embedding to represent an image depicting a selected biological sample (e.g., an image of a biological sample including the first label). As an example, the embedding 512 can be produced by the online encoder 510. In some embodiments, the embedding can be a 2048-dimensional vector. In some embodiments, operation 1312 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0189] In operation 1314, the image representation can be provided to a discriminator that is configured to predict the slide preparation machine or its attributes used to prepare and / or capture the image of the selected biological sample. As an example, the discriminator 1150 can receive the embedding 512 and can attempt to determine whether the image 502 (represented by the embedding 512) was imaged by a first scanner or a second scanner. In some embodiments, the discriminator 1150 can be constructed as a linear layer of the input size that is equal to the dimension of the output embedding from the generator 1120 (e.g., the embedding 512). For example, if the generator 1120 (online encoder 510) is implemented using a ResNet-18 architecture, the dimension of the embedding 512 will be 512 dimensions, and the discriminator 1150 can have an input size of 512 dimensions. In some embodiments, operation 1314 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0190] In operation 1316, the loss of the discriminator can be calculated. The discriminator loss can be calculated based on the prediction 1154 obtained from the discriminator 1150. As previously mentioned, the prediction 1154 is the prediction of the discriminator 1150 as to which of n possible slide preparation machines and / or attributes associated with these slide preparation machines was used to prepare and / or capture the image of the biological sample from which the input embedding was derived. For example, the prediction 1154 can indicate whether the image 502 was prepared using a first slide stainer or a second slide stainer, captured using a first slide scanner or a second slide scanner, etc. In some embodiments, operation 1316 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0191] In operation 1318, the discriminator can be updated based on the loss computed in the previous step. Updating the discriminator can include updating the parameters of the discriminator based on the computed discriminator loss. For example, the parameter θ can be updated based on the computed loss. D . In some embodiments, backpropagation can be used to update the parameters. As described above, during the update of the parameters θ D of the discriminator 1150, the parameters of the online encoder 510 and the target encoder 530 can remain frozen. Updating the parameters of the discriminator represents the training operation of the discriminator. In other words, the discriminator is trained by updating the parameters of the discriminator based on the discriminator loss computed in operation 1316. In some embodiments, operation 1318 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0192] In operation 1320, an updated image representation of the image can be generated using the first encoder. For example, the same image that was passed to the encoder during operation 1312 can be passed to the first encoder again. The image can be passed to the encoder again because during the previous step (e.g., updating the discriminator), the BYOL / DINO framework was detached, i.e., not updated. Thus, now that the discriminator has been updated, the image can be encoded by the first encoder in order to train the components of the previously frozen BYOL / DINO framework. In some embodiments, operation 1320 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0193] In operation 1322, the updated image representation can be provided to the updated discriminator. For example, the image representation of the image that was input to the first encoder after the discriminator update (in operation 1318) can be obtained and passed to the updated discriminator. In operation 1324, an updated loss of the updated discriminator can be computed. Operation 1324 can be substantially similar to operation 1316, except that in operation 1324, the parameters of the discriminator have been updated and the updated image representation can be used. In some embodiments, operations 1322 and 1324 can be performed by a subsystem that is the same as or similar to the model training subsystem 112.
[0194] In operation 1326, a main loss can be computed. The main loss function used can be the DINO loss function or the BYOL loss function, depending on the framework used for training. For example, the loss function of the DINO framework is represented by Equation 12, while the loss function of the BYOL framework is represented by Equation 10 ( where (The symmetric loss function obtained by feeding the first view 504 to the first encoder 510 and the second view 506 to the second encoder 530). In some embodiments, operation 1326 may be performed by a subsystem that is the same as or similar to the model training subsystem 112.)
[0195] In operation 1328, the adversarial loss may be determined based on the updated discriminator loss and the main loss. In some embodiments, a weight applied to the updated discriminator loss may be used to calculate the adversarial loss. To calculate the adversarial loss, the updated (weighted) discriminator loss may be subtracted from the main loss. For example, as shown in Equation 14, a weight μ may be applied to the discriminator loss term, which may be subtracted from the adversarial loss term. In some embodiments, operation 1328 may be performed by a subsystem that is the same as or similar to the model training subsystem 112.)
[0196] In operation 1330, the first encoder may be updated based on the determined adversarial loss. For example, the parameters θ of the student encoder 410 may be updated S (e.g., using backpropagation). As another example, for the adversarial framework 1100 built on the BYOL framework, the BYOL loss may be calculated, and the parameters θ of the online encoder 510 may be updated 在线 (e.g., using backpropagation). In some embodiments, operation 1330 may be performed by a subsystem that is the same as or similar to the model training subsystem 112.)
[0197] In operation 1332, the second encoder may be updated based on the update made to the first encoder. In some embodiments, the parameters of the second encoder may be updated based on the updated parameters of the first encoder. For example, based on the parameters θ of the student encoder 410 S and the parameters θ of the online encoder 510 在线 the parameters θ of the teacher encoder 420 may be updated T and the parameters θ of the target encoder 530 may be updated 目标 . For example, the exponential moving average of the parameters θ of the student encoder 410 S (where the parameters θ of the student encoder 410 have been updated based on the discriminator loss) S may be used to update the parameters θ of the teacher encoder 420 T . In some embodiments, operation 1332 may be performed by a subsystem that is the same as or similar to the model training subsystem 112.)
[0198] In some embodiments, process 1300 can be repeated for additional sample images. For example, after updating the second encoder in operation 1332, additional images of the biological sample included in the training data 1302 can be retrieved. In some embodiments, it can be determined whether any additional biological samples are to be analyzed. For example, if there are N biological samples, the first enhanced view set and the second enhanced view set can each include N whole slide images (e.g., enhanced versions / tiles of each of the whole slide images from the first image set 1304 and / or enhanced versions / tiles of the whole slide images from the second image set 1306). Thus, for the selected biological sample (the image of which is used at operation 1312 and then again at operation 1320), a round of training of the first encoder can be performed using the enhanced views (e.g., tiles obtained by partitioning the enhanced views) of the image of the selected biological sample from each of the first enhanced view set and the second enhanced view set. In some embodiments, the determination can be made after the parameters of the first encoder (and in some cases, the second encoder) have been updated. If it is determined that there are additional biological samples to be analyzed, process 1300 can return to operation 1312, where another biological sample can be selected and process 1300 can be repeated.
[0199] In some embodiments, if it is determined that there are no additional samples to be analyzed, test data (which can also be interchangeably referred to herein as validation data) can be used to test / validate the first encoder. For example, the first encoder can be tested using the validation data to determine the accuracy of the trained model for the new data classification task. In some embodiments, the accuracy of the trained encoder can be compared to a threshold accuracy level with respect to the validation. If the accuracy is determined to be greater than or equal to the threshold accuracy level, the model can be stored in the model database 146. For example, the trained model can be used for further downstream biological sample classification tasks. In some embodiments, the validation step can be performed on the second encoder instead of or in addition to the first encoder.
[0200] The above training techniques (e.g., mean similarity embedding method, adversarial method) can be used to "pre-train" the encoder to perform downstream classification tasks. After pre-training, each framework can be tested. In some cases, the only hyperparameter that changes is the scale term in the global loss function (λ and μ for the mean similarity embedding method and the adversarial method, respectively). For example, a higher value of the scale term can increase the model's ability to be independent of the scanner origin. In some embodiments, other hyperparameters can remain unchanged. In some cases, the optimizer used is the Adam optimizer, as described, for example, in "Adam: A method for stochastic optimization" by Kingma et al., the disclosure of which is incorporated herein by reference in its entirety. As an example, for stochastic gradient descent, the learning rate is fixed, while the Adam optimizer calculates the exponential moving averages of the gradients and the squared gradients, and the decay rates of the moving averages are controlled by two parameters.
[0201] downstream task
[0202] In some embodiments, the downstream task can be set up to evaluate the quality of the pre-trained encoder. The pre-trained encoder can alternatively or additionally be used for many independent tasks, such as survival prediction, cell detection and segmentation, mitosis detection, or other tasks. The downstream task is a fully supervised classification task for detecting whether a biological sample depicted in an image of the biological sample prepared and captured by a given slide preparation mechanism depicts a predefined characteristic associated with a disorder. For example, the downstream task can be to determine whether an image of a tissue sample includes a depiction of a tumor. In some embodiments, the downstream task is calculated by obtaining a trained encoder from the mean similarity embedding method and a trained encoder from the adversarial method and connecting each encoder to a fully connected classifier. As an example, referring to Figure 14A , the downstream classification subsystem 114 can be configured to perform downstream classification of images depicting biological samples to evaluate the quality of each pre-trained encoder (for the mean similarity embedding method, DINO framework, and BYOL framework, and for the adversarial method, adversarial DINO framework, and adversarial BYOL framework). For example, the trained encoder 1410 can refer to framework 400 (e.g., DINO framework), framework 500 (e.g., BYOL framework), framework 800 (e.g., adversarial DINO framework), or framework 1100 (e.g., adversarial BYOL framework).
[0203] The training encoder 1410 can be provided with an image 1402 depicting a biological sample. For example, the image 1402 can depict a tissue sample. The tissue sample can be a tissue sample including a tumor or other abnormality, or the tissue sample can depict "normal" tissue (e.g., without abnormality). The training encoder 1410 can be configured to generate an embedding 1412 representing the image 1402 in the latent space. For example, the embedding 1412 can be an n-dimensional vector. In some embodiments, the embedding 1412 can be input into a fully connected classifier 1420, which is configured to output a result 1422. The result 1422 can indicate whether the biological sample depicted by the image 1402 represents a specific abnormality or depicts another predefined characteristic as determined by the classifier 1420. In some embodiments, the result 1422 can be a binary flag indicating whether the tissue sample depicted by the image 1402 is a tumor.
[0204] Figure 14B FIG. 1450 is an illustrative flow chart of an exemplary process for training a classifier according to various embodiments. In some embodiments, the process 1450 can start at operation 1452. In operation 1452, a model can be selected to be used as a pre-trained encoder to support a downstream classifier. The model can be selected from a model database 146. For example, a pre-trained encoder with one of the frameworks, such as framework 400, 500, 800, or 1100, can be selected. The selected framework can be connected to a fully connected classifier for a downstream classification task, e.g., determining whether a whole slide image of a tissue sample depicts abnormal tissue, such as a tumor. In some embodiments, operation 1452 can be performed by a subsystem that is the same as or similar to the downstream classification subsystem 114.
[0205] In operation 1454, training data for training the classifier can be retrieved. For example, the training data for training the classifier can be retrieved from a training data database 144. In some embodiments, the training data can include images (e.g., image patches) depicting biological samples prepared and / or captured using one slide preparation mechanism. This is to determine / ensure that the model is independent of different slide preparation machines, as detailed below. The images used for training the classifier may be different from the images used for training the encoder. For example, each of the images used for training the classifier may not have a corresponding image included in the images used for training the encoder. In some embodiments, operation 1454 can be performed by a subsystem that is the same as or similar to the training data generation subsystem 110, the model training subsystem 112, and / or the downstream classification subsystem 114.
[0206] In operation, the retrieved training data and the pre-trained encoder can be used to train the classifier. For example, the retrieved training data and the pre-trained encoder 1410 can be used to train the classifier 1420. In some embodiments, training the classifier can include one or more sub-steps, such as sub-steps 1470 to 1474. In sub-step 1470, the training data can be provided to the pre-trained encoder to generate embeddings representing each image included in the training data. For example, the training data can include images / image patches depicting biological samples captured using a scanner and / or a stainers. In some embodiments, the training data can include images captured using only one scanner, only one stainer, or only one slide preparation machine (which can include a scanner and a stainer). In sub-step 1472, the embeddings generated by the pre-trained encoder can be provided to the classifier. The classifier can be configured to generate classification results. The classification results can include a label indicating whether the given biological sample depicted by the given image / image patch depicts an abnormal biological sample or a normal biological sample. For example, the label can indicate whether the image patch depicts a tissue sample including a tumor or normal tissue. The classification result can be a vector or an array indicating the likelihood that the image depicts one of a set of predefined classes. For example, if there are two classes (e.g., a normal tissue class or a tumor tissue class), each value in the vector or array can represent the probability that the given image depicts the two different classes. In sub-step 1474, the classifier can be updated based on the classification results. For example, the weights and biases of the classifier can be updated based on the classification results. In some embodiments, each of the sub-steps 1470 to 1474 can be repeated for each image included in the training data.
[0207] In operation 1458, the accuracy of the trained classifier can be calculated. In some embodiments, the accuracy can be calculated based on validation data. The validation data can include images / image patches depicting biological samples with known classifications. The accuracy can be calculated by determining the number or percentage of images correctly classified by the classifier. In some embodiments, operation 1458 can be calculated by a subsystem that is the same as or similar to the model training subsystem 112 and / or the downstream classification subsystem 114.
[0208] In operation 1460, the accuracy of the classifier can be compared with a predefined threshold accuracy score to determine whether the accuracy of the classifier is greater than the threshold accuracy score. For example, the threshold accuracy score can be 75% or higher, 85% or higher, 90% or higher, 95% or higher, or other values. If the accuracy score is determined to be less than (or equal to) the threshold accuracy score, process 1450 can return to operation 1456, where the classifier can be retrained or further trained. In some embodiments, process 1450 can return to operation 1454, where new or updated training data is selected and / or retrieved, and the classifier is further trained based on the new or updated training data. Some embodiments can include resetting the parameters of the classifier and / or the pre-trained encoder, or selecting a different pre-trained encoder if the accuracy score is less than the predefined threshold accuracy score. For example, if the accuracy score is less than the threshold accuracy score for more than N iterations, the parameters of the classifier can be reset. In some embodiments, operation 1460 can be performed by a subsystem that is the same as or similar to model training subsystem 112 and / or downstream classification subsystem 114.
[0209] In operation 1462, the trained classifier can be stored. For example, the trained classifier can be stored in model database 146. In some embodiments, operation 1460 can be performed by a subsystem that is the same as or similar to model training subsystem 112 and / or downstream classification subsystem 114.
[0210] In some embodiments, downstream classification subsystem 114 can be designed to serve two functions: (1) verifying the basic quality of the trained encoder, and (2) testing the ability of the trained encoder to be independent of the slide preparation machine. To this end, model training subsystem 112 can be configured to train classifier 1420 with data from one slide preparation machine (e.g., one scanner), while testing the classifier on data from multiple (or all) slide preparation machines (e.g., multiple scanners). This is different from the pre-training of the encoder, which uses data from each slide preparation machine to perform both training and validation. As an example, referring to Figure 15A , during the pre-training of the encoder, training dataset 1500 can include a first training set 1502 and a second training set 1504. The first training set 1502 includes images of biological samples captured using a first scanner (indicated by the dashed line), and the second training set 1504 includes images of biological samples captured using a second scanner (indicated by the solid line). As another example, the first training set 1502 can include images of biological samples captured using a first slide preparation machine, and the second training set 1504 can include images of biological samples captured using a second slide preparation machine.
[0211] In some embodiments, each image included in the first training set 1502 may include the same slide preparation label (e.g., a scanner label corresponding to scanner 124-1), while each image included in the second training set 1504 may include the same slide preparation label (e.g., a scanner label corresponding to scanner 124-2). The tiled and augmented versions of the images may be provided to the encoder to be trained. For example, images 402, 502 may each be an image included in the first training set 1502 or the second training set 1504. The validation data 1510 may (similar to the training data set 1500) include a first test set 1512 and a second test set 1514, where the first test set 1512 includes images depicting biological samples prepared and captured by a first slide preparation mechanism, and the second test set 1514 includes images depicting biological samples prepared and captured by a second slide preparation mechanism. In some embodiments, each image included in the first test set 1512 may include the same slide preparation label (e.g., a scanner label corresponding to scanner 124-1), while each image included in the second test set 1514 may include the same slide preparation label (e.g., a scanner label corresponding to scanner 124-2). In some embodiments, the images included in the validation data 1510 may have their labels masked. The tiled and augmented versions of the images may be provided to the trained encoder to test the accuracy of the trained encoder. For example, images 402 and 502 may each be an image included in the first test set 1512 or the second test set 1514.
[0212] As Figure 15B shown, the downstream training of the classifier 1420 may be trained and validated using different data. For example, the training data for training the classifier may include a training data set 1522, which may include images depicting biological samples captured using a first scanner. Alternatively, the training data set 1522 may include images depicting biological samples but using a second scanner.
[0213] In some embodiments, the downstream classifier (e.g., classifier 1420) may be tested using a first test set 1524 (including images prepared and captured by a first slide preparation mechanism), and may also be tested using a second test set 1526 (including images prepared and captured by a second slide preparation mechanism). If more slide preparation machines are used to prepare and / or capture images of biological samples, test sets including images depicting these slide preparation machines may also be used during testing of the trained classifier.
[0214] The data for creating the training data and validation data for pre-training the encoder and training the downstream classifier is described in detail below.
[0215] Encoder Pretraining Data
[0216] During the pre-training phase, which is self-supervised training, the images stored in the image database 142 can be sourced from multiple slide preparation machines. For example, images captured by two scanners can be stored in the image database 142, as shown in Table 3.
[0217] Scanner Slide Tile Scanner 1 271 2288954 Scanner 2 271 1750349
[0218] Table 3.
[0219] In some embodiments, each image can depict a biological sample, such as a tissue sample. Some or all of the biological samples can represent tumor samples from various clinical trials (e.g., datasets from various cancer studies). In some embodiments, the scanners of each slide preparation machine can be set to the same magnification setting. For example, the slides captured by Scanner 1 and Scanner 2 in Table 3 can be captured at 20x magnification. In some embodiments, for each slide, tiles can be extracted. For example, the training data generation subsystem 110 can be configured to generate training data and validation data for training and validating the frameworks 400, 500, 800, or 1100. The tiles can overlap. Subsequently, the tiles can be divided into training data and validation data and stored in the training data database 144 and the validation data database 148.
[0220] Downstream Training
[0221] In some embodiments, some or all of the images stored in the image database 142 can include labels indicating whether the biological sample depicted by a given image is "normal" or "abnormal". For example, the image can include a label indicating whether the tissue sample depicted by the image is a normal tissue sample or a tumor sample, and for downstream tasks, the training data generation subsystem 110 and / or the downstream classification subsystem 114 can be configured to select images including the "normal" label and images including the "abnormal" label from the available images. By selecting images containing these labels, the number of samples used to train, validate, and test a scanner-independent model can be significantly reduced.
[0222] Table 4 below includes example sizes of the training data, validation data, and test data after filtering out images lacking the "normal" or "abnormal" label.
[0223] Slide Total number of tiles Abnormal tiles Normal tiles Training Scanner 1 100 397615 368023 29592 Scanner 2 101 352294 340455 11839 Validation Scanner 1 21 83329 77466 5863 Scanner 2 21 77637 73776 3861 Testing Scanner 1 22 77444 70787 6657 Scanner 2 22 71336 68631 2705
[0224] Table 4.
[0225] The training data generation subsystem 110 can be configured to split the tile sets into subsets based on the slide preparation machine (e.g., scanner 1 or scanner 2) and the role during the training process (e.g., training, validation, testing). Thus, for a dual-scanner system (e.g., the slide preparation machine 120 includes the slide preparation machine 120-1 and the slide preparation machine 120-2), six subsets of samples are created. Additionally, the training data generation subsystem 110 can be configured to split the various tiles among different roles during the training process such that the corresponding slides belong to the same corresponding set. Thus, during the training portion of the training process, the tiles in the test subset are not included and the corresponding tiles are not used.
[0226] Figure 16 FIG. 1600 is a flow diagram illustrating an exemplary process 1600 for analyzing an image of a biological sample according to various embodiments. In some embodiments, process 1600 may begin at operation 1602. In operation 1602, an image depicting a biological sample may be received. The biological sample may be prepared on a slide using a slide preparation machine, and the image may be captured using a slide scanner (the slide stainer and the slide scanner may be part of the same slide preparation machine). In some embodiments, the type, settings, or other characteristics of the slide preparation machine may be unknown. For example, during the classification of biological samples, the scanner may be one of a group of scanners used to train a model. In some embodiments, the received image may be one of a set of images captured by a scanner depicting a particular biological sample. In some embodiments, operation 1602 may be performed by a subsystem that is the same as or similar to the downstream classification subsystem 114.
[0227] In operation 1604, the image may be provided to a trained classifier. The trained classifier may be trained using a backbone framework (such as framework 800 (e.g., the adversarial DINO framework) or framework 1100 (e.g., the adversarial BYOL framework)). For example, the image received at operation 1602 may be provided to the classifier 1420. In some embodiments, operation 1604 may be performed by a subsystem that is the same as or similar to the downstream classification subsystem 114.
[0228] In operation 1606, one or more classifications of the image may be determined by the trained classifier. For example, the trained classifier may use a trained encoder (e.g., encoder 1410) to generate an embedding representing the image, and the embedding generated by the trained encoder may be provided to the trained classifier to obtain a classification result. The classification result may include one or more classifications of the biological sample. For example, the classification result may include a classification indicating that the biological sample represents an abnormal biological sample (e.g., a tumor). In some embodiments, operation 1606 may be performed by a subsystem that is the same as or similar to the downstream classification subsystem 114.
[0229] Figure 17 A schematic diagram showing an exemplary cross-scanner reference dataset is presented. Two rows of image patches within dataset 1700 can represent images of biological samples captured via one of two different scanners. For example, the top row can represent image patches of a biological sample captured using a first scanner, and the bottom row can represent image patches of a biological sample captured using a second scanner. Image pairs 1702, 1704, and 1706 can represent various pairs of image patches. For example, image pair 1702 can include two image patches representing the same slide (e.g., the same slide of a biological sample) captured using two different scanners. Image pair 1706 can include two image patches representing the same slide captured using the same scanner. Image pair 1704 can include two image tiles representing the same tile identified by two different scanners. Image pair 1704 can be a registration pair. A registration pair can include images that should have the same content, and the only difference can be the color appearance (e.g., if the images are from different scanners). Image pair 1702 may not be a registration pair because the content of the two images can be different. When generating embeddings for each image pair, the distances in the latent space between the given pair of embeddings can be different. For example, the average L2 distance of the embeddings generated for the image patches of image pair 1702 can be greater than the average L2 distance of the embeddings generated for the image patches of image pair 1704. However, the average L2 distance of the embeddings generated for the image patches of image pair 1704 can be less than the average L2 distance of the embeddings generated for the image patches of image pair 1706.
[0230] In some embodiments, a cross-scanner ratio can be calculated based on the average distance between registration image pairs of the same blocks from the same scanner (e.g., image pair 1704) and the average distance between unregistered image pairs of the same slides from different scanners (e.g., image pair 1702). For example, the cross-scanner ratio can be:
[0231] Cross-scanner ratio = (average distance between registration pairs) / (average distance between unregistered pairs, same slide, different scanners)
[0232] Reaction 15.
[0233] The same scanner ratio can be calculated based on the average distance between registration image pairs of the same tiles from the same scanner (e.g., image pair 1704) and the average distance between unregistered pairs of the same slides of the same scanner. The lower these ratios are, the better the trained model is independent of various slide preparation techniques (e.g., scanners). For example, the same scanner ratio can be: Same scanner ratio = (average distance between registration pairs) / (average distance between unregistered pairs, same slide, same scanner)
[0234] Reaction 16
[0235] Exemplary results
[0236] Results are provided below with reference to Tables 5 and 6, including comparisons of different methods using the DINO framework (e.g., frameworks 400 and 800). Additionally, some results describe the results of the ResNet-18 model pre-trained on the ImageNet dataset. In some embodiments, one or more hyperparameters are tuned. For example, the weight decay can be adjusted to avoid overfitting, and the learning rate can also be adjusted. In some cases, the Adam optimizer is used.
[0237] DINO Cross-scanner ratio Same-scanner ratio Base model 0,937 1.772 Adversarial model 0,832 1.246
[0238] Table 5
[0239] BYOL Cross-scanner ratio Same-scanner ratio Base model 0,768 0.883 Adv model 0,543 0.548
[0240] Table 6
[0241] Figure 18B Various FIGS. 1810 to 1840 of representations created by an encoder for biological samples included in validation data trained using the average similarity method are shown according to various embodiments. Each of FIGS. 1810 to 1840 can be depicted as a t-SNE plot and can be used to highlight whether the encoder trained using the average embedding similarity method is independent of the slide preparation machine source. FIGS. 1810 to 1840 show the features learned by various encoders at different ratios. For example, the ratio can refer to λ in Equation 12. FIGS. 1810 to 1840 show the learned features of different versions of the average similarity method for different ratios. For example, FIG. 1810 includes λ = 1.0, FIG. 1820 includes λ = 10.0, FIG. 1830 includes λ = 100.0, and FIG. 1840 includes λ = 1000.0. Some embodiments include (for each of FIGS. 1810 to 1840) data points each corresponding to an embedding representing an image patch, and each color corresponding to a specific slide preparation machine type. For example, yellow (or light-colored) data points can represent the embeddings of image patches prepared and captured using a first slide preparation machine, while purple (or dark-colored) data points can represent the embeddings of image patches captured using a second slide preparation machine.
[0242] Figure 19Figures 1900 to 1950 show representations created by an encoder for biological samples included in validation data of an encoder trained using an adversarial method. Each of Figures 1900 to 1950 can be depicted as a t-SNE plot and can be used to highlight whether an encoder trained using an adversarial method is independent of the slide preparation machine source. Figures 1900 to 1950 show the features learned by various encoders. For example, Figures 1900 to 1950 show the learned features for different versions of the adversarial method at different ratios. Here, the ratio refers to the weight applied to the discriminator loss (e.g., μ in Equation 14). For example, Figure 1900 includes μ = 1.0, and Figure 1950 includes μ = 10.0. Some embodiments include (for each of Figures 1900 to 1950) data points each corresponding to an embedding representing an image patch, and each color corresponding to a specific slide preparation machine type. For example, yellow (or light-colored) data points can represent the embeddings of image patches captured using a first slide preparation machine, while purple (or dark-colored) data points can represent the embeddings of image patches captured using a second slide preparation machine.
[0243] Figure 18B and Figure 19 show that the higher the ratio, the more "independent of the slide preparation machine" (or "independent of slide handling") the model may be. For example, for a high ratio, the t-SNE plot is observed to be more mixed. The reason behind this is that, compared to traditional losses (e.g., BYOL loss, DINO loss), this ratio sets the importance of the penalty for being independent of the slide preparation machine. In particular, for the adversarial framework, the t-SNE plot for a higher ratio value (e.g., Figure 1950) shows fragmentation of the clusters that leads to the clusters starting to merge significantly. For example, for the average embedding similarity method, the mixing is less obvious.
[0244] Figure 20A Figure 2000 shows the standard deviation of the embeddings of the average embedding similarity method during training according to various embodiments. Self-supervised techniques can pose challenges to ensure that the model does not collapse. A collapsed model can refer to a model that produces a constant output regardless of the input. The DINO framework (e.g., Framework 400) and the BYOL framework (e.g., Framework 500) are designed to avoid model collapse. For example, the embeddings produced by the teacher encoder (in the DINO framework) go through a centering layer and a sharpening step, which helps avoid model collapse. However, to ensure that the model is independent of the slide preparation machine type, penalty modifications are made to the DINO framework, and the risk of model collapse is greater. Model collapse can be detected based on the distribution of the learned features. For example, the computing system 102 can be configured to monitor the standard deviation of the embeddings produced during training.
[0245] Figure 20BFIG. 2050 shows the standard deviation of the embeddings of the average embedding similarity method and the adversarial method during training according to various embodiments. The different traces in FIG. 2050 may correspond to different training methods used. For example, one trace shows the standard deviation of the embeddings generated via the mean similarity method with a ratio of λ = 1.0. Other traces show the standard deviation of the embeddings generated via the adversarial method with different ratios, such as μ = 1.0 and μ = 10.0. Another trace shows the standard deviation of the embeddings generated via the DINO framework.
[0246] In some embodiments, several metrics can be used to evaluate the training process. For the example datasets described above with respect to Tables 3 and 4, the data is imbalanced, so the various frameworks can be ranked based on their AUC-PR scores. Tables 7 and 8 below show the results of various metrics for the test datasets. For example, the downstream classification subsystem 114 can be configured to use the test data (as included in Table 7) to evaluate the trained model. In Table 7, the test data used to validate the trained model and the training data used to train the model can include images captured using the same scanner. In Table 8, the test data used to validate the trained model and the training data used to train the model can include images captured using different scanners.
[0247]
[0248] Table 7.
[0249]
[0250] Table 8.
[0251] Figure 20A It shows that the encoder trained with the maximum scaling factor seems to collapse. While this may happen in some extreme cases, other versions do not seem to collapse. For example, as Figure 20B shown, the standard deviation of the embeddings seems to be lower than that from DINO and Adversarial DINO. This indicates that when applying the penalty, the features may lose track of useful information. However, the standard deviation of the adversarial method does not seem to indicate that a collapse will occur.
[0252] As shown in Tables 7 and 8, for downstream classification tasks, the overall performance of the encoder trained using the mean embedding similarity method is not better than that of the basic DINO. The modification of the framework does not seem to improve the performance but rather imposes too large a penalty on training. On the other hand, the results seem to indicate that the modification of the framework for the adversarial method (e.g., Framework 800, Adversarial DINO framework) enhances the scanner-type-independent ability, and / or alternatively, seems to produce more meaningful improved embeddings. The encoder trained using the adversarial method further shows a level similar to that of the ResNet-18 model trained on ImageNet and has improved results compared to the DINO framework 400. In addition, the adversarial method, especially the Adversarial DINO framework (e.g., Framework 800), shows that the features added from the framework improve the model's scanner-type-independent ability.
[0253] Figure 21 An image 2100 of a biological sample according to various embodiments is shown. The data for developing the training data for training the classifier (and the encoder) can include 516K images and / or image patches of tissue samples depicting tumors and 42K images and / or image patches of normal tissue samples captured by a first slide preparation machine (e.g., Scanner 1), as well as 483K images and / or image patches of tissue samples depicting tumors and 18K images and / or image patches of normal tissue samples captured by a second slide preparation machine (e.g., Scanner 2). A similar process as described above can be used for the training and testing of the classifier. As shown in the image 2100, the downstream classifier can determine that one or more regions of the tissue sample represent a tumor based on the image 2100. In some embodiments, detecting a tumor or other abnormality in the biological sample can cause the computing system 102 to output the classification result to the client device 130. For example, the classification result can include an image of the tissue sample, the region where the tissue abnormality is detected, and an indication of how the image depicts the abnormality.
[0254] Figures 23A to 26B Various exemplary embedding diagrams according to various embodiments are shown. Figures 23A to 23B TSNE diagrams 2300 and 2350 of the embeddings obtained from the validation set analyzed using the Adversarial DINO technique are shown. In the TSNE diagrams 2300 and 2350, each color represents a unique tissue processing condition, which can be interchangeably referred to herein as a slide preparation machine or a set of attributes of the slide preparation machine (e.g., stain, magnification, scanner, etc.). In addition, the embeddings can be generated at the tile level. The TSNE diagram 2300 shows the distribution of various embeddings before performing the adversarial training. In other words, without filtering the sample processing-related characteristics, the TSNE diagram 2300 describes how various slide preparation machine attributes affect the resulting embeddings. However, the TSNE diagram 2350 can show the distribution of the embeddings after performing the adversarial training.
[0255] Figures 24A to 24B Show the TSNE plots 2400 and 2450 of the embeddings obtained from different datasets of a specific clinical trial analyzed using the adversarial DINO technique. In the TSNE plots 2400 and 2450, each color represents a unique dataset. Additionally, the embeddings can be generated at the tile level. The TSNE plot 2400 shows the distribution of the embeddings before performing adversarial training. In other words, without filtering the sample processing-related characteristics, the TSNE plot 2400 describes how the dataset affects the generated embeddings. However, the TSNE plot 2450 can show the distribution of the embeddings after performing adversarial training.
[0256] Figures 25A to 25B Show the TSNE plots 2500 and 2550 of the embeddings obtained from the validation set using the adversarial BYOL technique. In the TSNE plots 2500 and 2550, each color represents a unique tissue processing condition, which can be interchangeably referred to in this article as a slide preparation machine or a set of attributes of the slide preparation machine (e.g., stain, magnification, scanner, etc.). Additionally, the embeddings can be generated at the tile level. The TSNE plot 2500 shows the distribution of the embeddings before performing adversarial training. In other words, without filtering the sample processing-related characteristics, the TSNE plot 2500 describes how the dataset affects the generated embeddings. However, the TSNE plot 2550 can show the distribution of the embeddings after performing adversarial training.
[0257] Figures 26A to 26B Show the TSNE plots 2600 and 2650 of the embeddings obtained from different datasets of a specific clinical trial analyzed using the adversarial BYOL technique. In the TSNE plots 2600 and 2650, each color represents a unique dataset. Additionally, the embeddings can be generated at the tile level. The TSNE plot 2600 shows the distribution of the embeddings before performing adversarial training. In other words, without filtering the sample processing-related characteristics, the TSNE plot 2600 describes how the dataset affects the generated embeddings. However, the TSNE plot 2650 can show the distribution of the embeddings after performing adversarial training.
[0258] Figure 27An exemplary computer system 2700 is shown. In certain embodiments, one or more computing systems 2700 perform one or more steps of one or more of the methods described or shown herein. In certain embodiments, one or more computing systems 2700 provide the functionality described or shown herein. In certain embodiments, software running on one or more computing systems 2700 performs one or more steps of one or more of the methods described or shown herein, or provides the functionality described or shown herein. Certain embodiments include one or more portions of one or more computing systems 2700. Herein, references to a computing system may, where appropriate, include a computing device, and vice versa. Further, references to a computer system may, where appropriate, include one or more computer systems.
[0259] The present disclosure contemplates any suitable number of computing systems 2700. The present disclosure contemplates computing systems 2700 in any suitable physical form. By way of example and not limitation, the computing system 2700 may be an embedded computer system, a system on a chip (SOC), a single board computer system (SBC) (such as, for example, a computer on module (COM) or system on module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system grid, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more thereof. Where appropriate, the computing system 2700 may include one or more computing systems 2700; may be integrated or distributed; may span multiple locations; may span multiple machines; may span multiple data centers; or may reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computing systems 2700 may perform one or more steps of one or more of the methods described or shown herein without substantial spatial or temporal limitation. By way of example and not limitation, one or more computing systems 2700 may perform one or more steps of one or more of the methods described or shown herein in real time or in batch mode. Where appropriate, one or more computing systems 2700 may perform one or more steps of one or more of the methods described or shown herein at different times or in different locations.
[0260] In certain embodiments, the computing system 2700 includes a processor 2702, a memory 2704, a storage device 2706, an input / output (I / O) interface 2708, a communication interface 2710, and a bus 2712. Although the present disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, the present disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0261] In certain embodiments, the processor 2702 includes hardware for executing instructions, such as those that make up a computer program. By way of example and not limitation, to execute instructions, the processor 2702 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 2704, or the storage device 2706; decode and execute the instructions; and then write one or more results to an internal register, an internal cache, the memory 2704, or the storage device 2706. In certain embodiments, the processor 2702 may include one or more internal caches for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 2702 that includes any suitable number of any suitable internal caches. By way of example and not limitation, the processor 2702 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction cache may be copies of the instructions in the memory 2704 or the storage device 2706, and the instruction cache may accelerate the retrieval of those instructions by the processor 2702. The data in the data cache may be: a copy of the data in the memory 2704 or the storage device 2706 for operations by instructions being executed at the processor 2702; the result of a previous instruction executed at the processor 2702 for access or writing to the memory 2704 or the storage device 2706 by a subsequent instruction executed at the processor 2702; or other suitable data. The data cache may accelerate the read or write operations of the processor 2702. The TLB may accelerate the virtual address translation of the processor 2702. In certain embodiments, the processor 2702 may include one or more internal registers for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 2702 that includes any suitable number of any suitable internal registers. In appropriate instances, the processor 2702 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 2702. Although the present disclosure describes and shows particular processors, the present disclosure contemplates any suitable processor.
[0262] In certain embodiments, the memory 2704 includes a main memory that stores instructions for execution by the processor 2702 or data for the processor 2702 to operate on. By way of example and not limitation, the computing system 2700 can load instructions from the storage device 2706 or another source (such as, for example, another computing system 2700) into the memory 2704. The processor 2702 can then load the instructions from the memory 2704 into internal registers or an internal cache. To execute the instructions, the processor 2702 can retrieve the instructions from the internal registers or internal cache and decode the instructions. During or after instruction execution, the processor 2702 can write one or more results (which can be intermediate or final results) to the internal registers or internal cache. The processor 2702 can then write one or more of those results to the memory 2704. In certain embodiments, the processor 2702 only executes instructions in one or more internal registers or internal cache or in the memory 2704 (and not in the storage device 2706 or elsewhere) and only operates on data in one or more internal registers or internal cache or in the memory 2704 (and not in the storage device 2706 or elsewhere). One or more memory buses (which can each include an address bus and a data bus) can couple the processor 2702 to the memory 2704. The bus 2712 can include one or more memory buses, as described below. In certain embodiments, one or more memory management units (MMUs) reside between the processor 2702 and the memory 2704 and facilitate access to the memory 2704 requested by the processor 2702. In certain embodiments, the memory 2704 includes random access memory (RAM). This RAM can be volatile memory. In suitable cases, this RAM can be dynamic RAM (DRAM) or static RAM (SRAM). Additionally, in suitable cases, the RAM can be single-port or multi-port RAM. The present disclosure contemplates any suitable RAM. In suitable cases, the memory 2704 can include one or more memories. Although the present disclosure describes and illustrates particular memories, the present disclosure contemplates any suitable memory.
[0263] In certain embodiments, storage device 2706 includes a mass storage device for data or instructions. By way of example and not limitation, storage device 2706 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a tape, or a universal serial bus (USB) drive or a combination of two or more thereof. In appropriate cases, storage device 2706 can include removable or non-removable (or fixed) media. In appropriate cases, storage device 2706 can be internal or external to computing system 2700. In certain embodiments, storage device 2706 is a non-volatile solid-state memory. In certain embodiments, storage device 2706 includes read-only memory (ROM). In appropriate cases, the ROM can be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically re-writable ROM (EAROM), or flash memory or a combination of two or more thereof. The present disclosure contemplates a mass storage device 2706 in any suitable physical form. In appropriate cases, storage device 2706 can include one or more memory control units that facilitate communication between processor 2702 and storage device 2706. In appropriate cases, storage device 2706 can include one or more storage devices 2706. Although the present disclosure describes and illustrates particular storage devices, the present disclosure contemplates any suitable storage device.
[0264] In certain embodiments, I / O interface 2708 includes hardware, software, or both that provide one or more interfaces for communication between computing system 2700 and one or more I / O devices. In appropriate cases, computing system 2700 can include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computing system 2700. By way of example and not limitation, the I / O devices can include a keyboard, a keypad, a microphone, a monitor, a mouse, a printer, a scanner, a speaker, a still camera, a stylus, a tablet computer, a touch screen, a trackball, a video camera, another suitable I / O device, or a combination of two or more thereof. The I / O devices can include one or more sensors. The present disclosure contemplates any suitable I / O devices and any suitable I / O interface 2708 therefor. In appropriate cases, I / O interface 2708 can include one or more device or software drivers that enable processor 2702 to drive one or more of these I / O devices. In appropriate cases, I / O interface 2708 can include one or more I / O interfaces 2708. Although the present disclosure describes and illustrates particular I / O interfaces, the present disclosure encompasses any suitable I / O interface.
[0265] In certain embodiments, communication interface 2710 includes hardware, software, or both that provide one or more interfaces for communication (such as, for example, packet-based communication) between computing system 2700 and one or more other computing systems 2700 or one or more networks. By way of example and not limitation, communication interface 2710 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (such as a WI-FI network). The present disclosure contemplates any suitable network and any suitable communication interface 2710 therefor. By way of example and not limitation, computing system 2700 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, computing system 2700 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN or an ultra-wideband WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless networks, or a combination of two or more of these. In appropriate instances, computing system 2700 may include any suitable communication interface 2710 for any of these networks. In appropriate instances, communication interface 2710 may include one or more communication interfaces 2710. Although the present disclosure describes and shows particular communication interfaces, the present disclosure contemplates any suitable communication interface.
[0266] In certain embodiments, bus 2712 includes hardware, software, or both that couple the components of computing system 2700 to each other. By way of example and not limitation, bus 2712 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus, or a combination of two or more of these. In appropriate instances, bus 2712 may include one or more buses 2712. Although the present disclosure describes and shows particular buses, the present disclosure contemplates any suitable bus.
[0267] In this document, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (such as, for example, a field-programmable gate array (FPGA) or an application-specific IC (ASIC)), a hard disk drive (HDD), a hybrid hard drive (HHD), an optical disc, an optical disc drive (ODD), a magneto-optical disc, a magneto-optical disc drive, a floppy disk, a floppy disk drive (FDD), a magnetic tape, a solid state drive (SSD), a RAM drive, a SECURE DIGITAL card or drive, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more thereof. Where appropriate, the computer-readable non-transitory storage medium may be a volatile storage medium, a non-volatile storage medium, or a combination of a volatile storage medium and a non-volatile storage medium.
[0268] In this document, "or" is inclusive and not exclusive, unless expressly stated otherwise or the context otherwise indicates. Thus, in this document, "A or B" means "A, B, or both", unless expressly stated otherwise or the context otherwise indicates. Additionally, in this document, "and" is both conjunctive and disjunctive, unless expressly stated otherwise or the context otherwise indicates. Thus, in this document, "A and B" means "A and B, jointly or severally", unless expressly stated otherwise or the context otherwise indicates.
[0269] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications that would be understood by a person of ordinary skill in the art to the exemplary embodiments described or illustrated herein. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Additionally, although the present disclosure describes and illustrates the corresponding embodiments herein as including particular components, elements, features, functions, operations, or steps, any one of these embodiments may include any combination or arrangement of any components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by a person of ordinary skill in the art. Further, any reference herein to a device or system or a component of a device or system being adapted to, arranged to, capable of, configured to, enabled to, operable to, or operating to perform a particular function encompasses that device, system, or component whether or not that particular function is activated, turned on, or unlocked, so long as the device, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operating. Additionally, although the present disclosure describes or illustrates a particular embodiment as providing a particular advantage, a particular embodiment may not provide that advantage, some of the advantages, or all of the advantages.
[0270] Exemplary embodiment
[0271] The embodiments disclosed herein may include:
[0272] 1. A computer-implemented method, comprising: receiving image data comprising a first image set and a second image set, wherein the first image set and the second image set comprise digital images of a plurality of digital pathology slides processed using a first slide preparation machine and a second slide preparation machine, respectively, wherein the first slide preparation machine and the second slide preparation machine each have a set of attributes, and wherein the value of at least one of the attributes is different between the first slide preparation machine and the second slide preparation machine; based on the image data, generating a first enhanced view set and a second enhanced view set based on one or more enhancements applied to each image in the first image set and the second image set; and for each of the digital pathology slides: training a first vision transformer to: generate a first representation of an enhanced view in the first enhanced view set using the first vision transformer; and reinforce the similarity between the first representation and a second representation of an enhanced view in the second enhanced view set, the second representation being generated via a second vision transformer, and both the first representation and the second representation corresponding to the same digital pathology slide.
[0273] 2. The method according to embodiment 1, wherein both the first slide preparation machine and the second slide preparation machine are slide scanners.
[0274] 3. The method according to any one of embodiments 1 to 2, wherein both the first slide preparation machine and the second slide preparation machine are slide stainers.
[0275] 4. The method according to embodiment 3, wherein the first slide preparation machine uses a first staining technique and the second slide preparation machine uses a second staining technique different from the first staining technique.
[0276] 5. The method according to embodiment 1, wherein the first slide preparation machine and the second slide preparation machine are the same machine, and wherein the value of at least one of the attributes changes over time.
[0277] 6. The method according to any one of embodiments 1 to 5, wherein the digital pathology slides comprise whole slide images of multiple types of tissue.
[0278] 7. The method according to any one of embodiments 1 to 6, wherein the digital pathology slides are images of tissue stained with hematoxylin and eosin.
[0279] 8. The method according to any one of embodiments 1 to 7, wherein the first vision transformer and the second vision transformer have the same architecture.
[0280] 9. The method according to any one of embodiments 1 to 8, wherein the one or more augmentations include at least one of the following: blurring the image, flipping the image, rotating the image, distorting one or more colors of the image, or cropping the image.
[0281] 10. The method according to any one of embodiments 1 to 9, further comprising: dividing each image of the first augmented view set into a first plurality of tiles; and dividing each image of the second augmented view set into a second plurality of tiles, wherein the first representation is generated based on the first plurality of tiles, and the second representation is generated based on the second plurality of tiles.
[0282] 11. The method according to embodiment 10, further comprising: generating a first plurality of embeddings, each embedding corresponding to one of the first plurality of tiles, wherein the first representation is generated based on the first plurality of embeddings.
[0283] 12. The method according to embodiment 11, further comprising: calculating a first average embedding based on the first plurality of embeddings, the first representation including the calculated first average embedding.
[0284] 13. The method according to any one of embodiments 11 to 12, further comprising: generating a second plurality of embeddings, each embedding corresponding to one of the second plurality of tiles, wherein the second representation is generated based on the second plurality of embeddings.
[0285] 14. The method according to embodiment 13, further comprising: calculating a second average embedding based on the second plurality of embeddings, the second representation including the calculated second average embedding.
[0286] 15. The method according to any one of embodiments 10 to 14, wherein the first augmented view set and the second augmented view set include tiles randomly selected for the first plurality of tiles and the second plurality of tiles.
[0287] 16. The method according to any one of embodiments 1 to 15, wherein maximizing the similarity between the first representation and the second representation includes minimizing a loss function.
[0288] 17. The method according to embodiment 16, wherein the loss function is:
[0289]
[0290] 18. The method according to any one of embodiments 1 to 17, further comprising: training a classifier based on the first vision transformer for image classification of biological sample slides.
[0291] 19. The method according to embodiment 18, further comprising: receiving an image depicting a biological sample to be classified into at least one of a plurality of tissue categories; and providing the image to the trained classifier to determine one or more of the tissue categories for classifying the biological sample. The image is provided to the trained classifier to determine one or more of the tissue categories for classifying the biological sample.
[0292] 20. The method according to any one of embodiments 1 to 19, wherein the first vision transformer has a first set of hyperparameters and the second vision transformer has a second set of hyperparameters.
[0293] 21. The method according to embodiment 20, wherein the second set of hyperparameters is adjusted based on the first set of hyperparameters.
[0294] 22. The method according to embodiment 21, wherein an exponential moving average is applied to the values of the first set of hyperparameters to obtain the values of the second set of hyperparameters.
[0295] 23. The method according to any one of embodiments 20 to 22, wherein the values of the first set of hyperparameters are tuned using backpropagation based on the similarity between each corresponding first representation and second representation.
[0296] 24. The method according to any one of embodiments 1 to 23, wherein the first vision transformer includes an encoder and a softmax layer.
[0297] 25. The method according to embodiment 24, wherein training the first vision transformer includes: tuning the hyperparameters of the encoder.
[0298] 26. The method according to embodiment 25, wherein the second vision transformer includes an encoder, a centering layer, and a softmax layer, wherein the hyperparameters of the encoder of the second vision transformer are trained based on the tuned hyperparameters of the encoder of the first vision transformer.
[0299] 27. The method according to any one of embodiments 25 to 26, wherein the encoder of the first vision transformer and the encoder of the second vision transformer are each implemented using a residual neural network.
[0300] 28. The method according to any one of embodiments 26 to 27, wherein training the first vision transformer includes: tuning the hyperparameters of the encoder of the first vision transformer to generate the same representations of the same images processed by the first slide preparation machine and the second slide preparation machine.
[0301] 29. The method according to any one of embodiments 1 to 28, further comprising: evaluating the quality of the trained first vision transformer for a downstream classification task.
[0302] 30. The method according to embodiment 29, further comprising: connecting a fully connected classifier to the trained first vision transformer; and training the fully connected classifier using training data to classify biological sample slides.
[0303] 31. The method according to embodiment 30, wherein the training data includes at least some of the first image set.
[0304] 32. The method according to any one of embodiments 30 to 31, wherein the training data does not include any second image set.
[0305] 33. The method according to any one of embodiments 30 to 32, further comprising: testing the trained classifier using test data, wherein the test data includes at least some of the first image set and at least some of the second image set.
[0306] 34. The method according to embodiment 33, wherein the test data is derived from the image data, and the test data includes images indicating the classification of the biological sample slides and associated metadata.
[0307] 35. A method comprising: classifying digital pathology images using a trained classifier, wherein the trained classifier is trained using the method according to any one of embodiments 1 to 34.
[0308] 36. One or more computer-readable non-transitory storage media comprising instructions that, when executed by one or more processors, are configured to cause the one or more processors of the system to perform the method according to any one of embodiments 1 to 35.
[0309] 37. A system comprising: one or more processors and one or more computer-readable non-transitory storage media, the computer-readable non-transitory storage media being coupled to one or more of the processors and comprising instructions that, when executed by one or more of the processors, are operable to cause the system to perform the method according to any one of embodiments 1 to 35.
[0310] 38. A computer-implemented method, comprising: receiving training data including images of a plurality of biological samples processed using a first slide preparation machine or a second slide preparation machine, wherein the first slide preparation machine and the second slide preparation machine each have a set of attributes, and wherein the value of at least one of the attributes is different between the first slide preparation machine and the second slide preparation machine; for each of the biological samples: using a first encoder to generate a first representation of one of the images of the biological sample processed using the first slide preparation machine; providing the first representation to a discriminator to produce a prediction as to whether the biological sample corresponding to the one of the images was processed using the first slide preparation machine or the second slide preparation machine; updating one or more parameters of the discriminator based on a first loss, the first loss being calculated based on the generated prediction and metadata associated with the one of the images, the metadata indicating that the biological sample was processed using the first slide preparation machine; updating one or more parameters of the first encoder based on the first loss; training the updated first encoder to: generate an updated first representation of the one of the images, and enhance the similarity between the updated first representation and a second representation of the same biological sample in another of the images generated using a second encoder.
[0311] 39. The method according to embodiment 38, wherein both the first slide preparation machine and the second slide preparation machine are slide scanners.
[0312] 40. The method according to any one of embodiments 38 to 39, wherein both the first slide preparation machine and the second slide preparation machine are slide stainers.
[0313] 41. The method according to embodiment 40, wherein the first slide preparation machine uses a first staining technique and the second slide preparation machine uses a second staining technique different from the first staining technique.
[0314] 42. The method according to embodiment 38, wherein the first slide preparation machine and the second slide preparation machine are the same machine, and wherein the value of at least one of the attributes changes over time.
[0315] 43. The method according to any one of embodiments 38 to 42, further comprising: for each image: dividing the image into a plurality of tiles; selecting a set of tiles from the tiles; for each of the selected set: performing one or more enhancements on the tile to obtain a first enhanced view of the tile, wherein the first representation of the one of the images is generated based on the first enhanced view.
[0316] 44. The method according to embodiment 43, wherein the one or more augmentations include at least one of the following: blurring the image, flipping the image, rotating the image, distorting one or more colors of the image, or cropping the image.
[0317] 45. The method according to any one of embodiments 43 to 44, wherein the first representation is further generated based on corresponding tiles of the first augmented view and the selected tile group.
[0318] 46. The method according to any one of embodiments 43 to 45, wherein the one or more augmentations of the tiles further obtain a second augmented view of the tiles, and the second representation is generated based on the second augmented view.
[0319] 47. The method according to any one of embodiments 38 to 46, wherein the first encoder acts as a generator for adversarial learning with the discriminator.
[0320] 48. The method according to any one of embodiments 38 to 47, wherein the discriminator includes a linear layer having an input size equal to the number of dimensions of the embedding generated via the first encoder.
[0321] 49. The method according to embodiment 48, wherein the input size is 512.
[0322] 50. The method according to any one of embodiments 38 to 49, wherein the discriminator includes a linear layer having an output size equal to the number of slide preparation machines for processing the biological sample and obtaining the digital image.
[0323] 51. The method according to embodiment 50, wherein the output size is 2.
[0324] 52. The method according to any one of embodiments 38 to 51, wherein one or more parameters of the second encoder are updated based on the one or more updated parameters of the first encoder.
[0325] 53. The method according to embodiment 52, wherein one or more parameters of the second encoder are tuned based on an exponential moving average of the one or more parameters of the first encoder.
[0326] 54. The method according to any one of embodiments 38 to 53, wherein the first encoder and the second encoder are trained asynchronously with the discriminator.
[0327] 55. The method according to any one of embodiments 38 to 54, wherein one or more parameters of the discriminator are updated using backpropagation.
[0328] 56. The method according to embodiment 55, wherein the first loss is calculated based on a first loss function of the discriminator:
[0329]
[0330] 57. The method according to embodiment 56, wherein a loss function is used to calculate the similarity between the updated first representation and the second representation:
[0331]
[0332] 58. The method according to embodiment 57, wherein μ is a parameter for scaling the increased loss.
[0333] 59. The method according to any one of embodiments 38 to 57, wherein the first encoder has a first set of hyperparameters, and the second encoder has a second set of hyperparameters, the first set of hyperparameters including the one or more parameters of the first encoder.
[0334] 60. The method according to embodiment 59, wherein the second set of hyperparameters is adjusted based on the first set of hyperparameters.
[0335] 61. The method according to any one of embodiments 59 to 60, wherein the values of the first set of hyperparameters are tuned using backpropagation based on the similarity between each corresponding first representation and second representation.
[0336] 62. The method according to any one of embodiments 38 to 61, wherein a first vision transformer is used to implement the first encoder, and a second vision transformer is used to implement the second encoder.
[0337] 63. The method according to embodiment 62, wherein the first vision transformer further includes a softmax layer.
[0338] 64. The method according to any one of embodiments 62 to 63, wherein the second vision transformer further includes a centering layer and a softmax layer.
[0339] 65. The method according to any one of embodiments 62 to 64, wherein the first encoder and the second encoder are constructed using a residual neural network architecture.
[0340] 66. The method according to any one of embodiments 56 to 59, wherein training the first encoder comprises: adjusting the one or more hyperparameters of the first encoder to generate the same representation of the same biological sample processed using the first slide preparation machine or the second slide preparation machine.
[0341] 67. The method according to any one of embodiments 38 to 66, further comprising: evaluating the quality of the first encoder for a downstream classification task.
[0342] 68. The method according to embodiment 67, further comprising: connecting a fully-connected classifier to the trained first encoder; and training the fully-connected classifier using downstream training data to classify biological sample slides.
[0343] 69. The method according to embodiment 68, wherein the downstream training data includes at least some of the images, the images including metadata indicating that a corresponding biological sample has been processed using the first slide preparation machine.
[0344] 70. The method according to embodiment 69, wherein the downstream training data does not include any images having metadata indicating that a corresponding biological sample has been processed using the second slide preparation machine.
[0345] 71. The method according to any one of embodiments 68 to 70, further comprising: testing the trained classifier using test data, wherein the test data includes at least some of the images, the images including the metadata indicating that a corresponding biological sample has been processed using the first slide preparation machine, and at least some of the images, the images including the metadata indicating that a corresponding biological sample has been processed using the second slide preparation machine.
[0346] 72. The method according to embodiment 71, wherein the test data is derived from the image data, the test data including images including additional metadata indicating the classification of the biological sample.
[0347] 73. A non-transitory computer-readable medium storing computer program instructions, the computer program instructions when executed implementing the method according to any one of embodiments 38 to 72.
[0348] 74. A system, comprising: a memory storing computer program instructions; and one or more processors configured to execute the computer program instructions to perform the method according to any one of embodiments 38 to 72.
[0349] 75. The system according to embodiment 74, further comprising: a first slide scanner; and a second slide scanner.
[0350] 76. The system according to embodiment 75, wherein the first slide preparation machine includes the first slide scanner, and the second slide preparation machine includes the second slide scanner.
[0351] 77. The system according to embodiment 76, further comprising: a first slide staining machine; and a second slide staining machine.
[0352] 78. The system according to embodiment 77, wherein the first slide preparation machine includes the first slide staining machine, and the second slide preparation machine includes the second slide staining machine.
[0353] 79. The system according to embodiment 78, wherein the first slide staining machine uses a first slide staining technique, and the second slide staining machine uses a second slide staining technique different from the first slide staining technique.
Claims
1. A computer-implemented method, comprising: Receiving image data comprising a first image set and a second image set, wherein the first image set and the second image set respectively comprise digital images of a plurality of digital pathology slides processed using a first slide preparation machine and a second slide preparation machine, wherein the first slide preparation machine and the second slide preparation machine each have a set of attributes, and wherein the value of at least one of the attributes is different between the first slide preparation machine and the second slide preparation machine; Based on the image data, generating a first enhanced view set and a second enhanced view set based on one or more enhancements applied to each image in the first image set and the second image set; And For each of the digital pathology slides: Training a first vision transformer to: Generate a first representation of an enhanced view in the first enhanced view set using the first vision transformer; And Reinforce the similarity between the first representation and a second representation of an enhanced view in the second enhanced view set, the second representation being generated via a second vision transformer, and both the first representation and the second representation corresponding to the same digital pathology slide.
2. The method according to claim 1, wherein both the first slide preparation machine and the second slide preparation machine are slide scanners.
3. The method according to claim 1, wherein both the first slide preparation machine and the second slide preparation machine are slide stainers.
4. The method according to claim 3, wherein the first slide preparation machine uses a first staining technique and the second slide preparation machine uses a second staining technique different from the first staining technique.
5. The method according to claim 1, wherein the first slide preparation machine and the second slide preparation machine are the same machine, and wherein the value of at least one of the attributes changes over time.
6. The method according to any one of claims 1 to 5, wherein the digital pathology slides comprise whole slide images of multiple types of tissue.
7. The method according to any one of claims 1 to 6, wherein the digital pathology slides are images of tissue stained with hematoxylin and eosin.
8. The method according to any one of claims 1 to 7, wherein the first vision transformer and the second vision transformer have the same architecture.
9. The method according to any one of claims 1 to 8, wherein the one or more enhancements comprise at least one of the following: blurring an image, flipping an image, rotating an image, distorting one or more colors of an image, or cropping an image.
10. The method according to any one of claims 1 to 9, further comprising: Dividing each image in the first enhanced view set into a first plurality of tiles; And Dividing each image in the second enhanced view set into a second plurality of tiles, wherein the first representation is generated based on the first plurality of tiles and the second representation is generated based on the second plurality of tiles.
11. The method according to claim 10, further comprising: Generate a first plurality of embeddings, each embedding corresponding to one of the first plurality of tiles, wherein the first representation is generated based on the first plurality of embeddings.
12. The method according to claim 11, further comprising: Calculating a first average embedding based on the first plurality of embeddings, the first representation including the calculated first average embedding.
13. The method according to any one of claims 11 to 12, further comprising: Generate a second plurality of embeddings, each embedding corresponding to one of the second plurality of tiles, wherein the second representation is generated based on the second plurality of embeddings.
14. The method according to claim 13, further comprising: Calculating a second average embedding based on the second plurality of embeddings, the second representation including the calculated second average embedding.
15. The method according to any one of claims 10 to 14, wherein the first enhanced view set and the second enhanced view set include randomly selected tiles for the first plurality of tiles and the second plurality of tiles.
16. The method according to any one of claims 1 to 15, wherein maximizing the similarity between the first representation and the second representation includes minimizing a loss function.
17. The method according to any one of claims 1 to 16, further comprising: Training a classifier based on the first vision transformer to perform image classification of slides of biological samples.
18. The method according to claim 17, further comprising: Receiving an image depicting a biological sample to be classified as at least one of a plurality of tissue categories; and Providing the image to the trained classifier to determine one or more of the tissue categories for classifying the biological sample.
19. The method according to any one of claims 1 to 18, wherein the first vision transformer has a first set of hyperparameters and the second vision transformer has a second set of hyperparameters.
20. A method, comprising: Classifying digital pathology images using the trained classifier, wherein the trained classifier is trained using the method according to any one of claims 1 to 19.
21. A computer-implemented method, comprising: Receiving training data, the training data including images of a plurality of biological samples processed using a first slide preparation machine or a second slide preparation machine, wherein the first slide preparation machine and the second slide preparation machine each have a set of attributes, and wherein the value of at least one of the attributes is different between the first slide preparation machine and the second slide preparation machine; For each of the biological samples: Generating a first representation of one of the images of the biological sample processed using the first slide preparation machine using a first encoder; Providing the first representation to a discriminator to produce a prediction as to whether the biological sample corresponding to the one of the images was processed using the first slide preparation machine or the second slide preparation machine. Update one or more parameters of the discriminator based on a first loss, the first loss being calculated based on the generated prediction and metadata associated with the one in the image, the metadata indicating that the biological sample is processed using the first slide preparation machine; Update one or more parameters of the first encoder based on the first loss; Train the updated first encoder to: Generate an updated first representation of the one in the image, and Reinforce the similarity between the updated first representation and a second representation of another one in the image of the same biological sample generated using a second encoder.
22. A non-transitory computer-readable medium storing computer program instructions that, when executed, implement the method according to any one of claims 1 to 20 or 21.
23. A system comprising: A memory storing computer program instructions; And One or more processors configured to execute the computer program instructions to perform the method according to any one of claims 1 to 20 or 21.
24. The system according to claim 23, further comprising: A first slide scanner; And A second slide scanner.
25. The system according to claim 24, wherein the first slide preparation machine includes the first slide scanner, and the second slide preparation machine includes the second slide scanner.
26. The system according to claim 25, further comprising: A first slide staining machine; And A second slide staining machine.
27. The system according to claim 26, wherein the first slide preparation machine includes the first slide staining machine, and the second slide preparation machine includes the second slide staining machine.
28. The system according to claim 27, wherein the first slide staining machine uses a first slide staining technique, and the second slide staining machine uses a second slide staining technique different from the first slide staining technique.