System and method for detecting three-stage lymphatic structures

By blocking digital histological images and classifying neural networks, the time-consuming and labor-intensive TLS detection is solved, and efficient and accurate TLS automated detection is achieved to support personalized treatment decisions for cancer patients.

CN120303695APending Publication Date: 2025-07-11OWKIN INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082361.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-30
Filing Date
2023-11-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the pathological evaluation of tertiary lymphoid structures (TLS) is time-consuming and laborious, and it is difficult to detect efficiently and accurately in histopathological images, especially in cancer patients, affecting the evaluation and prediction of immunotherapy effects.

Method used

Using a computer-implemented method, automated detection and segmentation of TLS states are achieved by receiving digital histological images, processing images in chunks, extracting feature vectors, and classifying using neural network models, including a combination of the first neural network and the second neural network.

Benefits of technology

It improves the efficiency and accuracy of TLS detection, reduces the demand for computing resources, and can quickly identify the presence or absence of TLS without relying on local annotation, supporting personalized treatment decisions for cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303695A_ABST
    Figure CN120303695A_ABST
Patent Text Reader

Abstract

Computer-implemented methods, machine learning models, and machine-readable media are provided for detecting three-level lymphatic structures (TLSs) and / or the likelihood of the presence or absence of TLSs in histological images and / or subjects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to machine learning and computer vision, and more particularly to image preprocessing and classification. Background Art

[0002] Histopathological image analysis (HIA) is a key element in the diagnosis of many medical fields including oncology. Tertiary lymphoid structures (TLS) are ectopic lymphoid structures that develop at sites of chronic inflammation including tumors. TLS exist in tumors in different states of maturity, ultimately forming germinal centers. In patients with solid tumors, the presence of mature TLS, characterized by mature follicles containing germinal centers, in the tumor area is associated with an increased survival rate after receiving cancer immunotherapy. Thus, TLS can be used as a predictor for patients more likely to benefit from immune checkpoint inhibitors. However, the pathological assessment of TLS status is still time-consuming and typically requires additional analysis, including immunohistochemical staining. Therefore, there is a need for methods and systems for accurately and efficiently detecting TLS in histological images and subjects. Summary of the Invention

[0003] Describe a method and apparatus for classifying an image.

[0004] In one aspect, a computer-implemented method is disclosed for detecting the presence or absence of tertiary lymphoid structures (TLS) in a subject, comprising:

[0005] Receiving a digitized histological image of a sample obtained from a subject;

[0006] Tiling the histological image into a set of tiles;

[0007] Extracting a plurality of feature vectors from each of the tiles, wherein each feature in the one or more feature vectors represents a local descriptor of the tile; and

[0008] Classifying the histological image for TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with a training set of histological images with known TLS annotations, wherein the TLS status indicates the presence or absence of TLS in the subject.

[0009] In some embodiments, the classifying step comprises:

[0010] Applying a first neural network to the one or more of the plurality of feature vectors, wherein the first neural network assigns a tile score to each tile in the set of tiles based on the one or more of the plurality of feature vectors, wherein the tile score represents the likelihood that the tile includes TLS; and

[0011] Apply a second neural network to each of the tiles, where the second neural network aggregates a subset of the tile scores of the set of tiles and determines the TLS status in the histological image, where:

[0012] Train the first neural network using a training set of histological images that includes known local annotations of the presence or absence of TLS at the tile level; and

[0013] Train the first neural network and the second neural network using a training set of histological images that includes known global annotations of the presence or absence of TLS at the histological image level.

[0014] In some embodiments, the first neural network includes a 1D convolutional layer. In some embodiments, the first neural network uses MoCo features as input and outputs the likelihood that the tile includes a TLS. In some embodiments, the second neural network includes a multi-layer perceptron model.

[0015] In some embodiments, the computer-implemented method provided herein further includes:

[0016] Detect one or more locations in the histological image where a TLS is present.

[0017] In some embodiments of the computer-implemented method provided herein, each tile includes a plurality of pixels. In some embodiments, the method further includes:

[0018] Receive a digitized histological image of a sample obtained from a subject;

[0019] Partition the histological image into a set of tiles;

[0020] Extract a segmentation mask from each tile;

[0021] Apply a third neural network to each tile, where the third neural network uses the extracted segmentation mask to detect pixel scores; and

[0022] Detect a pixel-based TLS segmentation within the tile or image, where the third neural network is trained using a training set of tiles that includes known pixel-based TLS segmentation masks within the tiles.

[0023] In some embodiments, the third neural network assigns a pixel score to each pixel of each tile in the set of tiles and determines the pixel-based segmentation of the TLS within the tile, where the pixel score represents the likelihood that the pixel includes a TLS. In some embodiments, the third neural network is a U-NET semantic segmentation neural network.

[0024] In some embodiments, the extraction step is performed by a ResNet50 neural network. In some embodiments, the extraction of the plurality of feature vectors includes a Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2) algorithm.

[0025] In some embodiments of the computer-implemented method provided herein, the sample is cancer. In some embodiments, the cancer is selected from the group consisting of lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma, kidney cancer, head and neck cancer, liver cancer, breast cancer, gastrointestinal stromal tumor, cervical cancer, endometrial cancer, gastric cancer, thyroid cancer, cholangiocarcinoma, prostate cancer, anal cancer, vulvar cancer, skin cancer, parotid gland cancer, digestive tract cancer, penile cancer, esophageal cancer, and cancer of unknown primary site.

[0026] In some embodiments of the computer-implemented method provided herein, the training set of patches and / or the training set of histological images are digital images of histological sections of heterologous cancers (e.g., lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma).

[0027] In some embodiments, the histological image is a digital whole slide image (WSI). In some embodiments, the digital histological image is a digital image of a histological section stained with a dye. In some embodiments, the dye is hematoxylin and eosin (H&E).

[0028] In some embodiments, the classification step further includes:

[0029] Sorting the set of patches as follows:

[0030] Selecting the patch containing the highest TLS patch score, and

[0031] Selecting the patch containing the lowest TLS patch score.

[0032] In some embodiments, the computer-implemented method provided herein includes:

[0033] For each histological image among a plurality of histological images, repeating all of the following steps:

[0034] Receiving a digital histological image of a sample obtained from a subject;

[0035] Partitioning the histological image into a set of patches;

[0036] Extracting a plurality of feature vectors from each of the patches, wherein each feature of the one or more feature vectors represents a local descriptor of the patch; and

[0037] Classify histological images for TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with an imaging training set having known TLS annotations, and

[0038] Process the TLS status of the plurality of histological images, thereby detecting the presence or absence of TLS in a subject.

[0039] In some embodiments, the histological images lack local annotations of histopathological features. In some embodiments, each of the set of tiles comprises approximately 224×224 pixels.

[0040] In one aspect, a machine-readable medium is disclosed herein having executable instructions to cause one or more processing units to perform a method for detecting the presence or absence of tertiary lymphoid structures (TLS) in a subject, the method comprising:

[0041] Receiving a digitized histological image of a sample obtained from a subject;

[0042] Partitioning the histological image into a set of tiles;

[0043] Extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; and

[0044] Classify the histological image for TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with a training set of histological images having known TLS annotations, wherein the TLS status indicates the presence or absence of TLS in the subject.

[0045] In some embodiments, the classifying step comprises:

[0046] Applying a first neural network to one or more of the plurality of feature vectors, wherein the first neural network assigns a tile score to each tile in the set of tiles based on one or more of the plurality of feature vectors, wherein the tile score represents the likelihood that the tile comprises TLS; and

[0047] Applying a second neural network to each tile, wherein the second neural network aggregates a subset of the tile scores of the set of tiles and determines the TLS status in the histological image, wherein:

[0048] The first neural network is trained using a training set of histological images having known local annotations of the presence or absence of TLS at the tile level; and

[0049] Train the first neural network and the second neural network using a training set of histological images with known global annotations of the presence or absence of TLS at the histological image level.

[0050] In some embodiments, the first neural network includes a 1D convolutional layer. In some embodiments, the first neural network uses MoCo features as input and outputs the likelihood that the tile includes a TLS. In some embodiments, the second neural network includes a multi-layer perceptron model.

[0051] In some embodiments, the method performed by the machine-readable medium provided herein further includes:

[0052] Detect one or more locations where TLS is present in the histological image.

[0053] In some embodiments of the method performed by the machine-readable medium provided herein, each tile includes a plurality of pixels. In some embodiments, the method further includes:

[0054] Receiving a digitized histological image of a sample obtained from a subject;

[0055] Partitioning the histological image into a set of tiles;

[0056] Extracting a segmentation mask from each tile;

[0057] Applying a third neural network to each tile, where the third neural network uses the extracted segmentation mask to detect a pixel score; and

[0058] Detecting a pixel-based TLS segmentation within the tile or image, where the third neural network is trained using a training set of tiles that includes known pixel-based TLS segmentation masks within the tiles.

[0059] In some embodiments, the third neural network assigns a pixel score to each pixel of each tile in the set of tiles and determines the pixel-based segmentation of the TLS within the tile, where the pixel score represents the likelihood that the pixel includes a TLS. In some embodiments, the third neural network is a U-NET semantic segmentation neural network.

[0060] In some embodiments, the extraction step is performed by a ResNet50 neural network. In some embodiments, the extraction of the plurality of feature vectors includes the Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2) algorithm.

[0061] In some embodiments, the extraction of the plurality of feature vectors includes the Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2) algorithm.

[0062] In some embodiments of the method executed by the machine-readable medium provided herein, the sample is cancer. In some embodiments, the cancer is selected from the group consisting of lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma, kidney cancer, head and neck cancer, liver cancer, breast cancer, gastrointestinal stromal tumor, cervical cancer, endometrial cancer, gastric cancer, thyroid cancer, cholangiocarcinoma, prostate cancer, anal cancer, vulvar cancer, skin cancer, parotid gland cancer, digestive tract cancer, penile cancer, esophageal cancer, and cancer of unknown primary origin.

[0063] In some embodiments of the method executed by the machine-readable medium provided herein, the training set of tiles and / or the training set of histological images is a digital image of a histological section of a heterologous cancer (e.g., lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma).

[0064] In some embodiments, the histological image is a digital whole slide image (WSI). In some embodiments, the digital histological image is a digital image of a histological section stained with a dye. In some embodiments, the dye is hematoxylin and eosin (H&E).

[0065] In some embodiments, the classification step further includes:

[0066] Sorting the set of tiles as follows:

[0067] Selecting the tile that contains the highest TLS tile score, and

[0068] Selecting the tile that contains the lowest TLS tile score.

[0069] In some embodiments, the method executed by the machine-readable medium provided herein includes:

[0070] For each of the plurality of histological images, repeating all of the following steps:

[0071] Receiving a digital histological image of a sample obtained from a subject;

[0072] Partitioning the histological image into a set of tiles;

[0073] Extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; and

[0074] Classifying the histological image for TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with an imaging training set having known TLS annotations, and

[0075] Processing the TLS status of the plurality of histological images, thereby detecting the presence or absence of TLS in the subject.

[0076] In some embodiments, the histological image lacks local annotations of histopathological features. In some embodiments, each of the set of tiles includes approximately 224×224 pixels. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] The present invention is illustrated by way of example and not limitation in the figures, in which like reference numerals indicate similar elements.

[0078] Figure 1 An example flowchart showing a process for detecting the TLS status of an image and / or the presence or absence of TLS in a subject using a machine learning model according to an embodiment of the present disclosure.

[0079] Figure 2 An example flowchart showing a process for classifying the TLS status of an image and optionally predicting the location of TLS within the image and / or optionally predicting a pixel-based TLS segmentation according to an embodiment of the present disclosure.

[0080] Figure 3 An example flowchart showing a training and validation process of a machine learning model for detecting the TLS status of an image and / or the presence or absence of TLS in a subject according to an embodiment of the present disclosure.

[0081] Figure 4 A receiver operating characteristic (ROC) curve depicting the prediction of the presence or absence of TLS in a subject during cross-validation of a machine learning model according to an embodiment of the present disclosure.

[0082] Figure 5 A ROC curve depicting the prediction of the presence or absence of TLS in a validation cohort of a subject using a machine learning model according to an embodiment of the present disclosure.

[0083] Figure 6A and Figure 6B Depicting the location of TLS in a histological image according to an embodiment of the present disclosure. Figure 6A A H&E stained histological image depicting the location of TLS manually annotated by a pathologist and marked by a blue line. Figure 6B Depicting the location of TLS in a histological image analyzed by a machine learning model, where brighter regions are assigned higher tile scores to indicate a higher likelihood of TLS.

[0084] Figure 7 An example flowchart showing a process 700 for detecting pixel-based TLS segmentation in an image using a machine learning model according to an embodiment of the present disclosure.

[0085] Figure 8An example flowchart depicting the training and validation process of a machine learning model for detecting pixel-based TLS segmentation in an image according to an embodiment of the present disclosure.

[0086] Figure 9A and Figure 9B Depicting pixel-level TLS segmentation in a histological image. Figure 9A Depicting the extraction of tiles from an H&E-stained histological image in which a pathologist has manually annotated the TLS locations marked by blue lines and their segmentation masks. Figure 9B Depicting the process of training a machine learning model to predict TLS segmentation within a tile according to an embodiment of the present disclosure.

[0087] Figure 10 Showing an example of a computer system that can be used in conjunction with the embodiments described herein. Detailed Description

[0088] Describes computer-implemented methods, related systems, devices, and computer-readable media for detecting the presence or absence of tertiary lymphoid structures (TLS). In certain aspects, diagnostic tools are provided herein that apply machine learning to digital images of tissue sections, such as whole slide images (WSI), to identify subjects with or without TLS associated with their tumors and / or to assist in making treatment decisions.

[0089] In the following description, numerous specific details are set forth to provide a thorough explanation of embodiments of the invention. However, it will be clear to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other instances, well-known components, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0090] Histology is the field of study related to the microscopic features of biological samples. Histopathology refers to the microscopic examination of samples (e.g., tissues) obtained or otherwise derived from a subject (e.g., a patient) for the purpose of assessing a disease state. Histopathology samples are typically derived from processing the sample (e.g., tissue) in such a way as to attach the sample (e.g., tissue) or a portion thereof to a microscope slide. For example, a thin section of a tissue sample can be obtained using a microtome or other suitable device, and the thin section can be fixed to the slide. Optionally, to aid in the visualization of the sample, the sample can be further processed, for example, by applying a stain. A variety of stains have been developed for visualizing cells and tissues. These include, but are not limited to, hematoxylin and eosin (H&E), methylene blue, Masson's trichrome stain, congo red, oil red O, and safranin. Pathologists typically use H&E to aid in visualizing cells in tissue samples. Hematoxylin stains the cell nucleus blue, and eosin stains the cytoplasm and extracellular matrix pink. A pathologist visually inspecting an H&E-stained slide can use this information to assess the morphological features of the tissue. However, H&E-stained slides typically contain insufficient information to assess the presence or absence of a specific biomarker by visual inspection. Visualization of specific biomarkers (e.g., protein or RNA biomarkers) can be achieved by other staining techniques that rely on the use of labeled detection reagents that specifically bind to the biomarker of interest, e.g., immunofluorescence, immunohistochemistry, in situ hybridization, etc. These techniques can be used to determine the expression of individual genes or proteins, but are not suitable for assessing complex expression patterns involving a large number of biomarkers. Global expression analysis can be achieved by genomics and proteomics methods that use independent samples derived from the same tissue source as the samples used for histopathological analysis. Nevertheless, this approach is both expensive and time-consuming, requires the use of specialized equipment and reagents, and does not provide any information relating biomarker expression to specific regions within the tissue sample (e.g., specific regions within an H&E-stained image).

[0091] As used herein, the "tertiary lymphoid structure" (TLS) refers to ectopic lymphoid structures that develop in non-lymphoid tissues at sites of chronic inflammation, including autoimmune diseases, transplant rejection, and cancer. TLS is similar to lymph nodes both structurally and developmentally, and the organization and integrity of TLS are supported by stromal cells. TLS includes a heterogeneous cell population, including B cells (e.g., germinal center B cells), T cells (e.g., type 1 T helper (Th1) cells, T follicular helper (Tfh) cells, regulatory T (Treg) cells), dendritic cells (e.g., follicular dendritic cells (FDC), mature dendritic cells), plasma cells, neutrophils, and macrophages. The features of TLS also include the presence of high endothelial venules (HEV), which are blood vessels adapted for lymphocyte trafficking. Well-developed TLS contains B cell follicles with actively replicating B cell germinal centers surrounded by T cell areas. As used herein, a "germinal center" refers to a transiently formed microscopic structure within the B cell area (follicle) of secondary lymphoid organs, in which mature B cells are activated, proliferate, differentiate, and mutate their antibody genes. Scattered throughout the TLS are high endothelial venules and dendritic cell lysosome-associated membrane protein (DC-LAMP) and dendritic cells. TLS is not encapsulated and occurs in various non-lymphoid tissues such as epithelial tissues and stroma. Compared with secondary lymphoid organs (SLO), which are well-defined and refer to specific structures such as lymph nodes, TLS refers to structures with different organizations. TLS can be simple lymphocyte aggregates or more organized structures that exist within non-lymphocyte structures (Munoz-Erazo, L. et al. 2020 Cell. Mol. Immunol. 17:570-575).

[0092] The use of TLS as a prognostic indicator for cancer has been proposed (Colbeck, E. J. et al. 2017 Front. Immunol. 8, 1830; Trajkovski, G. et al. 2018 Open Access Maced. J. Med. Sci. 6, 1824 - 1828). The presence or induction of TLS after cancer treatment can predict treatment response and is generally associated with favorable treatment response and / or clinical outcome. The presence or induction of TLS can also be a prognostic factor, typically indicating a good prognosis for various cancers. In some embodiments, there is no correlation between the presence or induction of TLS in hepatocellular carcinoma (HCC) and favorable treatment response, clinical outcome, or prognosis. For example, in patients with solid tumors, the presence of mature TLS characterized by mature follicles containing germinal centers in the tumor area is associated with increased survival after receiving cancer immunotherapy. Thus, TLS can be used to identify subjects who are more likely to benefit from immune checkpoint inhibitors. The number, density, and location of TLS can affect the prognosis of a subject or the response to treatment (Munoz-Erazo, L. et al. 2020 Cell. Mol. Immunol. 17:570 - 575).

[0093] In addition, the exogenous induction of TLS can provide therapeutic benefits. The formation of TLS by various pharmacological methods can promote lymphocyte infiltration, tumor antigen activation, and differentiation to enhance the anti-tumor immune response and / or increase the sensitivity of immune-cold tumors to immunotherapy when combined with immune checkpoint blockade, vaccines, viruses, local intratumoral drugs, or intervention therapies. In immunocompetent tumors with a disrupted tumor microenvironment and strong chronic inflammation, angiogenesis, and fibrotic stroma, the use of anti-angiogenic and anti-immune inhibitors may help normalize the immune environment, thus facilitating TLS formation and the treatment response to immune checkpoint blockade. Several methods of using chemokines, cytokines, antibodies, antigen-presenting cells, or synthetic scaffolds to induce TLS formation are being developed. Strategies aimed at inducing de novo TLS in hypo-immunogenic and hyper-immunogenic tumors (in this case, in combination with therapeutic agents that suppress the inflammatory environment and / or with immune checkpoint inhibitors) represent promising avenues for cancer treatment.

[0094] The detection of TLSs, including their location and quantity, has been performed based on the morphological assessment of TLSs, thus relying on local annotation of important regions in the images by expert pathologists. For example, H&E staining enables the detection of TLSs in formalin-fixed paraffin-embedded tumor sections. Mature TLSs correspond to lymphoid follicles, including dense cell aggregates similar to germinal centers seen in secondary lymphoid structures (SLOs). Less differentiated structures, such as lymphoid aggregates and lymphoid follicles without germinal centers, can also be detected by pathological examination of histological slides stained with H&E. Currently, this pathological detection and evaluation of TLSs is still time-consuming and labor-intensive.

[0095] Other analyses, including immunohistochemistry, can be used to assist in TLS detection. Immunohistochemistry (IHC) on serial tumor sections or dual or multiplex labeling techniques using markers present in the cells or tissues containing TLSs can also be used to detect TLSs, and then, for example, the TLS density, size, and cell content on scanned images can be evaluated using quantitative digital pathology software. The common cell types present in TLSs and exemplary markers that can be used to detect TLSs are set forth in Table 1 below.

[0096] Table 1 Cell types present in TLSs and their markers

[0097]

[0098] Various gene signatures of TLSs can also be used to aid in the pathological detection of TLSs in samples (e.g., histological samples of cancer). Exemplary gene signatures of TLSs are provided in Table 2 below. The properties, functions, roles, and implications of TLSs have been discussed in the following documents, the entire contents of which are incorporated herein by reference: Sautès-Fridman et al. 2019 Nat. Rev. Cancer 19(6); Sautès-Fridman et al. 2016 Front. Immunol. 7:407; Vanherseche et al. 2021 Nat. Cancer 2(8):794 - 802.

[0099] Table 2 Gene signatures for TLS detection

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] Even when using histological analysis with IHC or the gene signature analysis provided herein, the detection of TLSs remains slow, laborious, and expensive, and thus is not well-suited for high-throughput applications. To overcome this problem, the present disclosure provides an image processing pipeline for analyzing histopathology images without using local annotations. The pipeline initially involves segmenting a large image (e.g., a WSI) into smaller images, e.g., images of 224×224 pixels, and detecting regions of interest in the images, on which classification is performed using Otsu's method. Thus, such classification is applicable to small images with a much lower computational cost than a single large image. These smaller images can be fed into a ResNet-type convolutional neural network to extract feature vectors from each small image, where the feature vectors include local descriptors of the small image. A score is calculated for each small image from the extracted feature vectors as a local tile-level (instance) descriptor. The top and bottom instances are used as inputs to a multi-layer perceptron (MLP) to perform classification on them.

[0106] In one embodiment, the device uses one or more neural network models to classify a histology image (e.g., a cancer histology image) to determine a label for the image. In this embodiment, the histology image can be a large image, where it is computationally impractical to process the image as a whole using only a neural network model. In particular, the device reduces the amount of computational resources (e.g., time and / or memory requirements) needed to perform the image classification task on these large images. This reduction in resources further improves the performance of the device when performing the image classification task. Additionally, even when a whole-slide image (WSI) is too large to fit into the memory of a graphics processing unit typically used for training machine learning models, the device can classify this type of image. In another embodiment, the device reduces the dimensionality of the data, thereby giving a better generalization error and being more effective in terms of model accuracy.

[0107] The mention of “one embodiment” or “some embodiments” in the specification means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase “in one embodiment” appearing in various places in the specification does not necessarily refer to the same embodiment. The term “exemplary” used herein means “an example” rather than “ideal”. It can be understood from the present disclosure that the present invention is not limited to the examples described herein.

[0108] For any method described herein, unless the context otherwise requires or dictates, the order of steps presented, whether in the text or in the accompanying flowcharts, should not be construed as meaning that the steps must be performed in the order presented. Instead, the order of steps presented provides one example of how the method may be performed, and generally the steps may alternatively be performed in a different order or simultaneously. The processes described in the following figures may be executed by processing logic that includes hardware (e.g., circuits, dedicated logic, etc.), software (such as that running on a general-purpose computer system or a dedicated machine), or a combination of both. Although the processes are described below in terms of some sequential operations, it should be understood that some of the operations described may be performed in a different order. Additionally, some operations may be performed in parallel rather than sequentially.

[0109] Computational methods for implementing the methods provided herein may include, for example, machine learning, artificial intelligence (AI), deep learning (DL), neural networks, classification and / or clustering algorithms, and regression algorithms.

[0110] The terms “server,” “client,” and “device” generally refer to data processing systems and not specifically to a particular form factor of a server, client, and / or device.

[0111] As used in the specification, “local annotation” means metadata (e.g., text, markers, numbers, and / or another type of metadata) that applies to a portion of an image rather than the entire image. For example, in one embodiment, a local annotation may be a marker for a region of interest in an image such as a histological image. Exemplary local annotations include markers that outline or otherwise identify a portion of the image, e.g., a tumor region of the image, a stromal region of the image, identification of a cell type within the image, identification of a biological structure (e.g., TLS) composed of multiple cells within the image. In contrast, “global annotation” as used in the specification refers to metadata that applies to the entire image. Exemplary global annotations include labels that identify the entire image, data on how the image was acquired, labels that identify characteristics of the subject from which the image was obtained (e.g., labels indicating the age, gender, diagnosis, etc. of the subject from which the image was obtained), and / or any other data that applies to the entire image. In some embodiments, the global annotation may indicate the presence, quantity, or location of TLSs known or understood to be present in the subject from which the image was obtained. In other embodiments, the global annotation may indicate known characteristics of the subject from which the image was obtained, such as survival time (e.g., survival time after obtaining the sample represented in the image) or response to a given treatment. In some embodiments described herein, an image containing global annotations may be used in the absence of local annotations.

[0112] "Patient" refers to a subject who exhibits symptoms and / or complications of a disease or disorder (e.g., a malignancy, cancer), is being treated by a clinician (e.g., an oncologist), has been diagnosed with a disease or disorder, and / or is at risk of developing a disease or disorder. The term "patient" includes human and veterinary subjects. Unless the context clearly indicates otherwise, any reference to a subject in this disclosure should be understood to include the possibility that the subject is a "patient".

[0113] As used herein, a "subject" is an animal that benefits from the methods according to this disclosure, such as a mammal, including primates (such as humans, monkeys, and chimpanzees) or non-primates (such as cows, pigs, and horses). In some aspects of the invention, the subject is a human, such as a human diagnosed with cancer. The subject can be a woman. The subject can be a man. In some aspects, the subject is an adult subject.

[0114] As used herein, in the context of this disclosure, "predict" or "forecast" refers to determining the likelihood of the presence or absence of a certain condition (e.g., TLS) in the past, present, or future. In some embodiments, a model (e.g., a 1D convolutional layer with MoCo features, followed by a multi-layer perceptron) can predict the likelihood of the TLS state through one or more of the following test accuracy metrics:

[0115] The odds ratio is greater than 1, preferably about 2 or greater or about 0.5 or less, about 3 or greater or about 0.33 or less, about 4 or greater or about 0.25 or less, about 5 or greater or about 0.2 or less, or about 10 or greater or about 0.1 or less;

[0116] The specificity is greater than 0.5, preferably at least about 0.6, at least about 0.7, at least about 0.8, at least about 0.9, or at least about 0.95, and the corresponding sensitivity is greater than 0.2, preferably at least about 0.3, at least about 0.4, at least about 0.5, at least about 0.6, at least about 0.7, at least about 0.8, at least about 0.9, or at least about 0.95;

[0117] The sensitivity is at least 0.5, preferably at least about 0.6, at least about 0.7, at least about 0.8, at least about 0.9, or at least about 0.95, and the corresponding sensitivity is at least 0.2, preferably at least about 0.3, at least about 0.4, at least about 0.5, at least about 0.6, at least about 0.7, at least about 0.8, at least about 0.9, or at least about 0.95;

[0118] At least about 75% sensitivity, combined with at least about 75% specificity;

[0119] The positive likelihood ratio [calculated as sensitivity / (1 - specificity)] is greater than 1, preferably at least about 2, at least about 3, at least about 4, at least about 5, at least about 10; or

[0120] The negative likelihood ratio [calculated as (1 - sensitivity) / specificity] is less than 1, preferably about 0.5 or less, about 0.33 or less, or about 0.25 or less, or about 0.1 or less.

[0121] As used herein, "tumor" refers to an abnormal growth of cells or tissue and / or a mass resulting therefrom. In some embodiments, the tumor tissue or tumor cells are malignant (e.g., cancerous). As used herein, "cancer" or "malignant tumor" refers to multiple cells in which abnormal cells divide without control and have the ability to invade or metastasize to adjacent or distant tissues or organs. Exemplary malignant tumors or cancers include lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, and melanoma.

[0122] I. Methods for Detecting TLS in a Subject

[0123] In some aspects, provided herein are computer-implemented methods for detecting the presence or absence of TLS in a subject. In one aspect, the computer-implemented method includes:

[0124] Receiving a digitized histological image of a sample obtained from the subject;

[0125] Partitioning the histological image into a set of tiles;

[0126] Extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; and

[0127] Classifying the histological image for TLS status using at least the plurality of feature vectors and a classification model trained with an imaging training set having known TLS annotations, wherein the TLS status indicates the presence or absence of TLS in the subject. In some embodiments, the histological image lacks local annotations of histopathological features and / or global annotations of TLS status.

[0128] As used herein, "tile" refers to a sub-part of an image. As used herein, "partitioning" refers to dividing an image or region of interest into tiles.

[0129] As used herein, the term "digital image" or "digitized image" refers to an electronic image represented by a collection of pixels that can be viewed, processed, and / or analyzed by a computer. In some aspects of the present disclosure, digital images of histological slides (e.g., H&E stained slides), as a supplement or alternative to visual inspection by a pathologist, permit computational evaluation of tissue samples. In some embodiments, digital images can be acquired by a digital camera or other optical device capable of capturing digital images from a slide or a portion thereof. In other embodiments, digital images can be acquired by scanning a non-electronic image of a slide or a portion thereof. In some embodiments, the digital images used in the applications provided herein are whole slide images. As used herein, the term "whole slide image (WSI)" refers to an image that includes all or substantially all of a tissue section (e.g., a tissue section present on a histological slide). In some embodiments, the WSI includes an image of the entire slide. In other embodiments, the digital images used in the applications provided herein are a selected portion of a tissue section (e.g., a tissue section present on a histological slide). In some embodiments, digital images are acquired after treating a tissue section with a stain (e.g., H&E).

[0130] As used herein, the "TLS" status refers to the presence, location, or quantity of TLS in a subject, histological image, tile within a histological image, or pixel within a tile. The TLS status can be expressed in binary form (e.g., present or absent), categorically, as a continuous range (e.g., a quantity represented by a number), as a description (e.g., a narrative description of the location or abundance of TLS), or any combination thereof. The TLS status can be based on a likelihood score assigned to a subject, histological image, tile, or pixel or its constituent elements. For example, at the subject level, the TLS status can refer to the presence or absence of TLS in a subject (e.g., at a cancer site) and can be determined based on a likelihood score (e.g., an image score) assigned to a histological image of a sample obtained from the subject (e.g., a tumor sample). At the histological image level, the TLS status can refer to the presence or absence, quantity, or location of TLS in the image (e.g., segmentation within the image) and can be determined based on an image score assigned to the image or a tile score assigned to a tile within the image. At the tile level, the TLS status can refer to the presence or absence, quantity, or location of TLS in the tile (e.g., segmentation within the tile) and can be determined based on a tile score assigned to the tile or a pixel score assigned to a pixel within the tile. At the pixel level, the TLS status can refer to the presence or absence or quantity of TLS in the pixel and can be determined based on a pixel score assigned to the pixel.

[0131] As used herein, "score", "likelihood score", or "risk score" refers to the likelihood of a situation, such as the presence of TLS, or the presence of treatment relevance (e.g., present at least in a certain amount of clinical relevance, present at least at a location with clinical relevance). In some embodiments, the score is expressed as a classification. In other embodiments, the score is expressed as a continuous range. In one embodiment, the score represents the likelihood that TLS is present in a pixel, tile, image, or subject.

[0132] According to one embodiment, the device classifies at least one histological input image by applying a first convolutional neural network to segment the image between at least one region of interest that contains information that can be used for classification and at least one background region that contains little or no information that can be used for classification. The device also tiles the region of interest of the image into a set of tiles. Additionally, the device extracts a feature vector for each tile by applying a second convolutional neural network, where the feature is a local descriptor of the tile. Additionally, to classify the image, the device processes the feature vectors of the extracted tiles. In one embodiment, by segmenting the input image, the device processes a reduced number of tiles and avoids processing the entire image.

[0133] In one embodiment, the first convolutional network is a semantic segmentation neural network that classifies the pixels of the input image into one of two categories: (a) region of interest; and (b) background region. Additionally, step (b) of tiling can be performed by applying a fixed tiling grid to the image such that the tiles have a predetermined size. Additionally, at least one level of scaling can be applied to the tiles. For example, multi-level scaling can be applied to the tiles and tiles of different scaling levels can be combined. Additionally, the device can optionally randomly sample the tiles and / or fill a set of tiles with blank tiles such that the set of tiles includes a given number of tiles.

[0134] In another embodiment, the second convolutional neural network can be a residual neural network, such as a ResNet50 residual neural network or a ResNet101 residual neural network, where the last layer is removed using the previous layer as the output, or a VGG neural network. The second convolutional neural network can be a pre-trained neural network, allowing the use of state-of-the-art neural networks without the need to have a large-scale image database and computational resources to train the neural network.

[0135] In one embodiment, the device may calculate at least one score for a tile from the extracted feature vectors, where the tile score represents the contribution of the tile to the image classification. Using the tile scores, the device may sort a set of tile scores and select a subset of the tile scores based on their values and / or their rank in the sorted set; and to classify the image, apply a classifier to the saved tile scores. The device may also apply the classification to multiple input images, where the device may aggregate a population of corresponding tiles from different input images.

[0136] In an alternative embodiment, the device may also aggregate clusters of adjacent tiles. In this embodiment, aggregating clusters of tiles may include tiles of connected clusters, selecting a single tile from the cluster according to a given criterion, using the cluster as a multi-dimensional object, or aggregating values, for example, by mean or max pooling operations. Additionally, the device may apply an autoencoder to the extracted feature vectors to reduce the dimensionality of the features. In one embodiment, the image may be a histopathology slide, the region of interest is the tissue region, and the classification of the image is a diagnostic classification.

[0137] In an alternative embodiment, when local annotations are available, such as the presence of a tumor in a slide region, a hybrid technique may be used to take these annotations into account. To this end, the device may train a machine learning model for two concurrent tasks: (1) locally predicting the presence or absence of a tumor and / or other macroscopic properties on each tile, and predicting a set of global labels. The device (or devices) may use a complex architecture that, on the one hand, involves the classification system described above to process a set of 2048 features. On the other hand, the device applies a convolutional neural network to transform the features of N tiles into N*2048 feature vectors. Based on this vector, the device trains the convolutional neural network to predict the presence or absence of a tumor (or some other macroscopic property) for each tile. The device may take both the predicted output and the N*2048 feature vectors, and apply an operation of weighted pooling to the concatenation of these two vectors to obtain 2048 feature vectors of the input image. The device concatenates the output of the classification model and the obtained 2048 features, and attempts to predict a set of global labels for the image (e.g., survival, tumor size, necrosis, and / or other types of predictions) based on this vector. The loss of the model involves both global and local predictions. In this embodiment, by adding the information derived from the local annotations to the computational flow, the performance of the entire model can be improved.

[0138] Figure 1FIG. 0 shows an example flowchart of process 100 according to an embodiment of the present disclosure, which uses a machine learning model to detect the TLS status of an image and / or the presence or absence of TLS in a subject. At block 102, process 100 receives an image, one or more machine learning (ML) models, and optional other inputs. In some embodiments, the image is a digitized histological image of a sample obtained from a subject, such as a digitized whole slide image (WSI).

[0139] At block 104, process 100 detects a region of interest (ROI) within the image. As used herein, a "region of interest" (ROI) of an image can be any region that is semantically relevant to the task to be performed, particularly a region corresponding to tissue, organ, bone, cell, body fluid, etc. in the context of histopathology. In another embodiment, process 100 segments the image into regions of interest and background regions. In this embodiment, by extracting the regions of interest from the input image, the computational amount required to classify the input image can be reduced. For example, in one embodiment, a histopathology slide (or other type of image) may include blank regions in the image with little or no tissue, so it is useful to introduce so-called "tissue detection" or "substance detection" methods to evaluate whether a certain region of the slide contains any tissue. More generally, when the goal is to classify a large image, it makes sense to identify the regions of interest in the image and distinguish them from the background regions. These regions of interest are the regions in the image that contain valuable information for the classification process. In addition, the background region is the region in the image that contains little or no valuable information, where the background region can be regarded as noise for the current task. To achieve this task, various different types of image segmentation schemes can be used. For example, in one embodiment, the Otsu method can be used to segment the image, where the Otsu method is a simple threshold method based on the intensity histogram of the image. In this embodiment, when the image contains two classes of pixels (e.g., foreground pixels and background pixels, or more specifically tissue and non-tissue) that follow a bimodal distribution, using the Otsu method to segment the image shows quite good results. However, when it cannot be assumed that the histogram of the intensity levels has a bimodal distribution, this method is known to perform poorly on complex images. To improve the overall efficiency of the method, more robust techniques are needed.

[0140] In another embodiment, to improve the robustness of image segmentation and be able to process complex images (such as histopathology images), a semantic segmentation neural network for segmenting images can be used, such as a U-NET semantic segmentation neural network, SegNet, DeepLab, or another type of semantic segmentation neural network. In this embodiment, a semantic segmentation neural network that does not depend on a specific distribution in the intensity histogram can be used. Additionally, using such a neural network allows image segmentation to consider multi-channel images, such as red-green-blue (RGB) images. Thus, the segmentation not only depends on the histogram of pixel intensities but can also utilize the semantics of the image. In one embodiment, the semantic segmentation neural network is trained to separate the tissue of a histopathology image from the background of the image, to distinguish between stained or unstained tissue and the background.

[0141] In another embodiment, another advantage of using a U-NET segmentation neural network is that this network type was developed for biomedical image segmentation and thus conforms to the usual constraints of biomedical data, namely, small data sets with very high dimensions. In fact, the U-NET segmentation neural network is a model with very few training parameters, enabling the network to be trained with fewer training examples. Additionally, in another embodiment, using data augmentation techniques on the training data can produce very good results because this architecture allows for obtaining more training examples from the same training set.

[0142] In another embodiment, to reduce the computational cost of the image segmentation step, the original image can be downsampled. As described further below, in one embodiment, some of the image analysis is performed at the tile level (sub-parts of the image), and using semantic segmentation on the downsampled version of the image does not degrade the quality of the segmentation. This allows for using the downsampled image without degrading the segmentation quality. In one embodiment, to obtain the segmentation mask of the original full-resolution image, process 100 only needs to upsample the segmentation mask generated by the neural network.

[0143] At block 106, process 100 chunks the ROI into a set of tiles. Chunking the image can include dividing the original image into smaller, more manageable images, called tiles. In one embodiment, the chunking operation is performed by applying a fixed grid to the whole-slide image, using the segmentation mask generated by the segmentation method, and selecting the tiles that contain tissue or any other region of interest. To further reduce the number of tiles to be processed, in one embodiment, additional or alternative selection methods, such as random subsampling, can be used to retain only a given number of slides.

[0144] Tiling can be performed to improve the ability to preprocess images. For example, in one embodiment, due to the large size of whole slide images, using a tiling method helps in histopathological analysis. More generally, when processing professional images such as histopathological slides, satellite images, or other types of large images, the resolution of the image sensors used in these fields can grow as rapidly as the capacity of the random access memory associated with the sensor. With this increase in image size, it becomes difficult to store batches of images, and sometimes even a single image, in the random access memory of a computer. This difficulty is exacerbated if one attempts to store these large images in the dedicated memory of a graphics processing unit (GPU). This situation makes it computationally difficult to process an entire image slide or any other image of a similar size.

[0145] In one embodiment, tiling (or image minus background) of an image addresses this challenge by dividing the original image (or image minus background) into smaller, more manageable images (i.e., tiles). In one embodiment, the tiling operation is performed by applying a fixed grid to the whole slide image, using a segmentation mask generated by a segmentation method, and selecting tiles that contain tissue or any other type of region of interest for subsequent classification processes. To further reduce the number of tiles to be processed, additional or alternative selection methods such as random subsampling can be used to retain only a given number of slides. For example, in one embodiment, the image (or image minus background) is divided into tiles of a fixed size (e.g., each tile is 224×224 pixels in size). Within a slide, the tiles can be of any uniform size. The tiles can be square. For example, the tiles can have a width and / or depth of approximately 10 - 500 μm. Each side of a tile can have approximately 20 - 1000 pixels. The number of tiles generated depends on the size of the substance being detected and can vary from a few hundred tiles to 50,000 or more tiles. In one embodiment, the number of tiles is limited to a fixed number (e.g., 10,000 tiles) that can be set at least based on computational time and memory requirements. In a particular embodiment, the tiles have a size of approximately 224×224 pixels and approximately 112 μm×112 μm. In these embodiments, each image of the digitized whole slide image can have approximately 10,000 tiles.

[0146] In one embodiment, augmentation can be applied to each of multiple sets of tiles. In some embodiments, a first set of features is extracted from a first batch of augmented tiles. A second set of features is extracted from a second batch of augmented tiles. In some embodiments, the augmented tiles include magnified or rotated views, or views with color augmentation. For example, since orientation is not important in histological slides, the slides can be rotated to varying degrees. The slides can also be enlarged or magnified. To make the matching pairs of tiles closer and the different pairs of tiles farther apart, the contrastive loss between the first and second sets of extracted features can be used. To focus on the positive pairs rather than the negative pairs extracted from the first and second sets of features, a contrastive loss can be applied.

[0147] At block 108, process 100 extracts feature vectors from the respective tiles. In one embodiment, the feature extraction is performed by a ResNet50 neural network. In one embodiment, each of the features is extracted by applying a trained feature extractor that is trained using a machine learning algorithm with a training set of images. For example, in one embodiment, the machine learning algorithm is Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2). In one embodiment, the trained machine learning model is the machine learning model trained as described herein Figure 3 in the machine learning model trained herein.

[0148] In some embodiments, a machine learning algorithm, such as MoCo or MoCo v2, extracts multiple feature vectors from the respective tiles. The extraction of the multiple feature vectors can be performed using a convolutional neural network, such as a ResNet50 neural network. The multiple feature vectors can be any number, such as approximately 1000, approximately 1500, approximately 2000, approximately 3000, approximately 4000, or more.

[0149] In some embodiments, machine learning algorithms, such as MoCo or MoCo v2, extract multiple feature vectors from digital images and use a first convolutional neural network to extract multiple feature vectors. In one embodiment, process 100 can use any feature extraction neural network to extract features, such as ResNet-based architectures (ResNet-50, ResNet-101, ResNetX, etc.), Visual Geometry Group (VGG) neural networks, Inception neural networks, or custom neural networks designed specifically for the task. In some embodiments, non-neural network feature extractors, such as SIFT or CellProfiler, can be used to extract features. Additionally, the feature extraction neural network used can be a pre-trained feature extraction neural network, as these neural networks are trained on very large-scale datasets and thus have optimal generalization accuracy. In one embodiment, the first neural network includes a 1D convolutional layer. In one embodiment, the 1D convolutional layer uses MoCo features as input and outputs the likelihood of the tile including TLS.

[0150] In one embodiment, process 100 uses a ResNet-50 neural network because this neural network can provide features that are well-suited for image analysis without requiring too much computational resources. For example, in one embodiment, ResNet-50 can be used for histopathology image analysis. In this example, the ResNet-50 neural network relies on residual blocks that allow the neural network to be deeper and still improve its accuracy because simple convolutional neural network architectures may result in the worst accuracy when the number of layers grows too large. In one embodiment, the weights of the ResNet-50 neural network can be the weights used for feature extraction, which are from pre-training on the ImageNet dataset because this dataset is a truly general image dataset. In one embodiment, using a neural network pre-trained on a large independent image dataset provides good features that are independent of the image type, even when the input images are specialized, such as histopathology images (or other types of images). In one embodiment, process 100 uses a ResNet-50 convolutional neural network to extract 2048 features for each tile. If process 100 extracts, for example, 10,000 tiles, then process 200 generates a matrix of 2048×10,000. Additionally, if process 200 is executed with several images as input, then process 100 generates a tensor with dimensions of number of images×number of features / number of tiles×number of tiles.

[0151] In one embodiment, to extract features of a given slide, process 100 processes each of the selected tiles to output a feature vector for the tile through a ResNet-50 neural network. In this embodiment, the feature vector can be a 2048-dimensional vector or a vector of other sizes. Additionally, process 100 can apply an autoencoder to the feature vector to further provide dimensionality reduction (e.g., reducing the dimension of the feature vector to 256 or another dimension). In one embodiment, an autoencoder can be used when a machine learning model may be vulnerable to overfitting. For example, in one embodiment, process 100 can reduce the length of 2048 feature vectors to 512-length feature vectors. In this example, process 100 uses an autoencoder including a single hidden layer architecture (512 neurons). This can prevent the model from overfitting by finding several singular features in the training dataset, and also reduce the computation time and required memory. In one embodiment, the classification model is trained on a small subset of the image tiles, e.g., on 200 tiles randomly selected from each slide (out of a total of 411,400 tiles).

[0152] To derive the minimum number of features, process 100 can optionally perform a zero-padding operation on the feature vectors. In this embodiment, if the number of feature vectors is below the minimum number of feature vectors, process 100 can perform zero-padding to add feature vectors to a set of feature vectors of the image. In this embodiment, each zero-padded feature vector has a null value.

[0153] At block 110, process 100 uses the extracted feature vectors and a machine learning model to perform TLS status classification on the image. The computer-implemented method can process all tiles within the ROI to extract feature vectors and / or perform classification. In some embodiments, the computer-implemented method further includes selecting a subset of the tiles to be applied to the machine learning model. In some embodiments, the subset of tiles is selected by random sampling. In this embodiment, process 100 can output that the digital image is TLS positive or TLS negative, where this designation is a global classification of the digital image. Alternatively, process 100 can determine which tiles are TLS positive based on the tile scores of the tiles. In one embodiment, TLS image classification is further discussed in Figure 2 below.

[0154] In some embodiments, the classification step, such as block 110 in process 100, includes:

[0155] applying a first neural network to one or more of the plurality of feature vectors, wherein the first neural network assigns a tile score to each tile in the set of tiles based on one or more of the plurality of feature vectors, where the tile score represents the likelihood that the tile includes a TLS; and

[0156] Apply a second neural network to each of the tiles, where the second neural network aggregates a subset of the tile scores of the set of tiles and determines the TLS status in the histological image, where:

[0157] Train the first neural network using a training set of histological images with known local annotations of the presence or absence of TLS at the tile level; and

[0158] Train the first neural network and the second neural network using a training set of histological images with known global annotations of the presence or absence of TLS at the histological image level.

[0159] Figure 2 FIG. 11 shows an example flowchart of a process 200 according to an embodiment of the present disclosure, the process 200 classifying the TLS status of an image and optionally predicting the TLS location within the image. In one embodiment, the process 100 performs Figure 2 to classify the TLS status of an image.

[0160] At block 202, the process 200 receives features extracted from the tiles, the machine learning model, and optional other inputs. In one embodiment, the extracted features are the extracted features determined in block 108 above. At block 204, the process 200 determines the tile score for each tile based on the extracted feature vectors by using a first neural network trained using a training set of images with local annotations. In some embodiments, the first neural network is a 1D convolutional layer. In some embodiments, the tile score represents the likelihood that the tile includes a TLS. In one embodiment, the tile score is represented as a continuous range. For example, in one embodiment, the process 200 determines the tile score for each tile as any number between 0 and 1, where a score of 0 represents the lowest likelihood that the tile includes a TLS and a score of 1 represents the highest likelihood that the tile includes a TLS. In one embodiment, the process 200 uses a connected neural network to reduce each of the feature vectors to one or more scores. In one embodiment, the process 200 can use a fully connected neural network to reduce the feature vector to a single score, or use a fully connected neural network that outputs various scores or multiple fully connected neural networks that output different scores respectively to reduce the feature vector to multiple scores representing various characteristics of the tile. These scores associated with a tile are sorted, and a subset of the tiles is selected for image classification.

[0161] For example, in one embodiment, process 200 may use a 1D convolutional layer to create scores for each tile. In this example, as described above, with a feature vector of length 2048, this convolutional layer performs a weighted sum among the 2048 features of the tile to obtain the score, where the weights of the sum are learned by the model. Additionally, since the 1D convolutional layer is unbiased, the score of a zero-padded tile is zero and thus serves as a reference for a completely uninformative tile. Process 200 selects the top N scores and the bottom M scores and uses them as inputs for the classification described below. This architecture ensures which tiles are used for prediction.

[0162] In one embodiment, process 200 uses the tile score vector as the input to a dense multi-layer neural network to provide a desired classification (e.g., the TLS status of an image). The classification can be any task of associating a label with the data given as the input to the classifier. In one embodiment, a trained classifier is used for the digital image input because the input data is derived from the entire pipeline, so the classifier is able to label the histopathology image given as the input without the need to process the full image, which may be computationally infeasible. For example, in one embodiment, the label can be any type of label, such as: a binary value representing a given pathological prognosis; a numerical label representing a score, probability, or prediction of a physical quantity (such as a survival prediction or a response to a treatment prediction); and / or a scalar label as described above or a vector, matrix, or tensor of such labels representing structured information. For example, in one embodiment, process 200 may output a TLS positive or TLS negative status, thereby indicating that the digital image is predicted to include or not include TLS. In one embodiment, process 200 uses a multi-layer perceptron (MLP) with two fully connected layers having 200 and 100 neurons, respectively, with sigmoid activation. In this embodiment, the MLP serves as the core of the prediction algorithm that converts the tile scores into labels. Although in one embodiment, process 200 predicts a single label for the image (e.g., a probability score), in an alternative embodiment, process 200 may predict which of the tiles may include TLS. In this embodiment, process 200 may determine a certain score as the threshold for determining the presence of TLS in a tile.

[0163] Histological images can be classified based on at least one set of tile scores derived from the image tile feature vectors generated by a neural network. At block 204, process 200 calculates the tile scores for each tile using the relevant feature vectors of the tile. For example, in one embodiment, process 200 can use a 1D convolutional layer to create scores for each tile. In the example where the length of the above-mentioned feature vector is 2048, this convolutional layer performs a weighted sum among all 2048 features of the tile to obtain the score, where the weights of the sum are learned by the model. Additionally, since the 1D convolutional layer is unbiased, the score of a zero-padded tile is zero and thus serves as a reference for a completely uninformative tile.

[0164] At block 206, process 200 sorts the tiles by tile score. In one embodiment, process 200 sorts a set of tiles to determine the top N and / or bottom M scores at block 206.

[0165] At block 208, process 200 selects the highest tile score (N) and the lowest tile score (M). In one embodiment, process 200 selects a subset of tiles for a later classification step. In one embodiment, this subset of tiles can be tiles with the top N highest scores and the bottom M lowest scores, the top N highest scores, the bottom M highest scores, and / or any weighted combination of scores. In one embodiment, the values of N and / or M can have the same or different ranges. Additionally, the N and / or M ranges can be a static number range (e.g., 10, 20, 100, or some other number), a range, percentage, label (e.g., small, large, or some other label), and / or some other value suitable for setting through a user interface component (slide, user input, and / or another type of user interface component). In one embodiment, process 200 additionally concatenates these scores into an image score vector that can be taken as an input for image classification.

[0166] In some embodiments, the classification step further includes:

[0167] sorting the set of tiles in the following manner:

[0168] selecting the tile that contains the highest TLS tile score, and

[0169] selecting the tile that contains the lowest TLS tile score.

[0170] In one embodiment, when studying histological whole slide images (or slides), a single patient can be associated with multiple slides taken at different locations in the same sample, from multiple organs, or at different time points with various staining methods. In this embodiment, slides from a single patient can be aggregated in a variety of ways. In one embodiment, to form a larger slide processed (segmented, tiled, feature extracted, and classified) in the same or a similar manner as a normal slide, process 200 can concatenate the slides.

[0171] In another embodiment, a slide may not contain enough useful tissue to extract tiles for applying the feature extraction step on them and thus feed features to a classifier. In such a case, the input to the classifier is zero-padded, which means that for each missing tile, features consisting of zeros are added to the true features computed by the feature extractor.

[0172] At block 210, process 200 applies a second neural network to the tiles using tile scores. In one embodiment, the second neural network is a multi-layer perceptron model. In one embodiment, process 200 uses a subset of the tiles, namely tile scores N and M. At block 212, process 200 determines the TLS global score of the histological image using the result of the second neural network applied to the tile scores. In one embodiment, the TLS global score can be a positive presence of TLS or a negative presence of TLS. In this embodiment, the number of tiles in the image can be on the order of 10,000 tiles. In another embodiment, the number of tiles in the image can be more or less. In one embodiment, to reduce computational complexity, the classification system samples the tiles to reduce the number of tiles used in the neural network computation. In one embodiment, the classification system samples the tiles randomly or using some other type of sampling mechanism. For example, in one embodiment, the classification system samples the tiles randomly to reduce the number of tiles from an order of 10,000 tiles to an order of a few thousand tiles (e.g., 3,000 tiles).

[0173] In some embodiments, the computer-implemented method provided herein further includes detecting one or more locations in the histological image where TLS is present.

[0174] For example, process 200 can optionally proceed to block 214, where process 200 outputs the predicted TLS locations within the histological image. In some embodiments, at block 214, process 200 superimposes the tile scores assigned to the respective tiles onto the histological image, visualizing the regions assigned high and low tile scores. For example, as Figure 6BAs shown, process 200 can detect the TLS locations in histological images. Regions are color-coded based on the assigned tile scores, such that brighter regions correspond to higher tile scores, indicating a higher likelihood of TLS. Detecting the TLS locations corresponds to the manual annotation of the TLS locations by a pathologist in Figure 6A where the regions marked by the blue lines indicate the TLS locations.

[0175] In some embodiments, the computer-implemented methods provided herein include:

[0176] For each of a plurality of histological images, repeating all of the following steps:

[0177] Receiving a digitized histological image of a sample obtained from a subject;

[0178] Partitioning the histological image into a set of tiles;

[0179] Extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; and

[0180] Classifying the histological image for TLS status using at least the plurality of feature vectors and a classification model of an imaging training set with known TLS annotations, and

[0181] Processing the TLS status of the plurality of histological images to detect the presence or absence of TLS in the subject.

[0182] The machine learning models of the computer-implemented methods provided herein can be trained and validated. In a particular embodiment, a first neural network is trained using a training set of histological images including known local annotations of the presence or absence of TLS at the tile level; and a first neural network and a second neural network are trained using a training set of histological images including known global annotations of the presence or absence of TLS at the histological image level.

[0183] In one embodiment, at Figure 1 and Figure 2The model used in [description] can be trained and validated using a training set of images. In one embodiment, process 300 trains a one-dimensional convolutional neural network that generates scores and an MLP classification model using the input labels of a training set of images. In one embodiment, the process iteratively trains the model by calculating a set of scores for the training images, predicting labels, determining the difference between the predicted labels and the input labels, and optimizing the model based on the difference (e.g., calculating new weights for the model) until the difference is within a threshold. In one embodiment, process 300 trains the model to predict a single label for an image (e.g., tile score). In alternative embodiments, the process can be trained to predict multiple global labels for an image. In one embodiment, the process can be trained to perform multi-task learning to predict multiple global labels. For example, in one embodiment, a machine learning model (e.g., MLP and / or the model described elsewhere) can be trained to predict multiple labels at once in a multi-task learning environment using the resulting feature vectors.

[0184] Figure 3 A flowchart illustrating one embodiment of process 300 for training and validating first and second neural networks to detect the TLS state in histological images or a subject. In one embodiment, the neural network includes one or more separate models for the classification processes described in Figure 1 and Figure 2 In Figure 3 process 300 begins by receiving a training set of histological images, local and global annotations, a machine learning model, and optionally other inputs at block 302. In one embodiment, process 300 receives images, local and global annotations, the first and second neural networks used in the above Figure 1 and Figure 2 and optionally other inputs described in Figure 1 In one embodiment, the first neural network is a 1D convolutional layer. In one embodiment, the 1D convolutional layer uses MoCo features as input and outputs tiles including the likelihood of TLS. In one embodiment, the second neural network is a multi-layer perceptron model. In one embodiment, a training set of histological images including known local annotations of the presence or absence of TLS at the tile level is used to train the first neural network; and a training set of histological images including known global annotations of the presence or absence of TLS at the histological image level is used to train the first neural network and the second neural network. Process 300 performs a processing loop (blocks 304 - 314) on each training image to generate a set of feature vectors from the tiles and predict the TLS state of the image. At block 306, process 300 detects a region of interest (ROI) in the training image. In one embodiment, as described above with respect to Figure 1As described in block 104, process 300 detects the ROI. At block 308, process 300 chunks the ROI into a set of tiles. In one embodiment, as described above with respect to Figure 1 block 106, process 300 chunks the ROI. At block 310, process 300 extracts feature vectors from the tiles. In one embodiment, as described above with respect to Figure 1 block 108, process 300 extracts feature vectors. For example, in one embodiment, as described above with respect to Figure 1 process 300 uses a ResNet-50 convolutional neural network to determine the feature vectors of the tiles in the segmented image of the chunked ROI. In one embodiment, process 300 generates a set of feature vectors for the training images. Additionally, process 300 can perform data augmentation during the training of the method to improve the generalization error. This data augmentation can be accomplished by applying various transformations to the tiles, such as rotation, translation, flipping, cropping, blurring, adding noise to the image, modifying the intensity of a particular color, or changing the contrast. In one embodiment, as described above with respect to Figure 2 block 204, process 300 uses a first neural network to determine the tile scores. In one embodiment, as described above with respect to Figure 2 block 206, process 300 continues to sort the blocks by the block scores. In one embodiment, as described above with respect to Figure 2 block 208, process 300 continues to select N (the highest tile scores) and M (the lowest tile scores). In one embodiment, as described above with respect to Figure 2 block 210, process 300 continues to apply a second neural network to N, M, and the tiles. At block 312, process 300 uses the extracted feature vectors and a machine learning model to predict the TLS status (e.g., the presence, quantity, or location of TLS) in each histological image. In one embodiment, as described above with respect to Figure 1 block 110, process 300 uses the extracted feature vectors and a machine learning model to predict the TLS status (e.g., the presence, quantity, or location of TLS) in each histological image, as described above with respect to Figure 2 block 212, outputs the TLS global score, and / or, as described above with respect to Figure 2 block 214, outputs the TLS location in the histological image. The process loop ends at 314.

[0185] To determine the sufficiency of training, at block 316, process 300 evaluates whether the predictions of the machine learning model for the TLS status and the global and local annotations of the TLS status in the training dataset have converged. If they have converged, at block 318, process 300 validates the machine learning model. If they have not converged, at block 320, process 300 adjusts the machine learning model and re-executes the processing loop (blocks 304 - 314).

[0186] In Figure 3 , process 300 trains a machine learning model for classifying images. The accuracy of the machine learning model can be evaluated by using an image training set as input to validate the machine learning model and calculate one or more labels. The validation process can be performed by receiving a validation image set and the trained machine learning model and processing the validation image set. In one embodiment, the validation image set is the same as the training set. In another embodiment, the validation set can be different from the training image set. For example, in one embodiment, an image set labeled for a particular type of image (e.g., histological image) can have some images selected for training the model, while other images from the set are used to validate the trained model. In one embodiment, the model is a machine learning model, such as the MLP model described above.

[0187] In some embodiments of the computer-implemented method provided herein, each tile includes a plurality of pixels. In some embodiments, the method further includes detecting a pixel-based segmentation of a TLS within the tile. This determination step can include applying a machine learning model to each tile in the set of tiles, where the model is trained using a tile training set that includes known pixel-based TLS segmentation masks within the tiles. As used herein, "TLS segmentation" or "segmentation of TLS" refers to the localization of a TLS in a given region. The TLS segmentation in a given region (e.g., tile, histological image) can be expressed by the presence or absence of a TLS in a unit region (e.g., pixel) that partitions the given region. In some embodiments, the model assigns a pixel score to each pixel of each tile in the set of tiles and determines the pixel-based segmentation of the TLS within the tile, where the pixel score represents the likelihood that the pixel includes a TLS.

[0188] For example, as Figure 9A shown, the training set of images can be based on manual annotations by a pathologist of the TLS positions marked by blue lines, with local manual annotations of tiles from H&E-stained histological images and their segmentation masks. As Figure 9B shown, the machine learning model can be trained to predict TLS segmentation by processing the training set of images with known pixel-based segmentation masks and comparing the output with the annotations.

[0189] Figure 7 FIG. 700 is an example flowchart showing a process for detecting pixel-based TLS segmentation in an image using a machine learning model according to an embodiment of the present disclosure. At block 702, process 700 receives an image, one or more machine learning (ML) models, and optional other inputs. In some embodiments, the image is a digitized histological image of a sample obtained from a subject, such as a digitized WSI. At block 704, as described with respect to Figure 1 block 104 of FIG. Figure 1 FIG., process 700 detects a region of interest (ROI) within the image. At block 706, as described with respect to Figure 8 block 106 of FIG.

[0190] Figure 8 FIG., process 700 tiles the ROI into a set of tiles. At block 708, process 700 extracts a pixel-based segmentation mask from each tile. At block 710, process 700 determines the pixel scores of each pixel within the tile. At block 712, process 700 uses a third neural network to detect pixel-based segmentation in the tile or the image. In some embodiments, as described herein and for example Figure 7 as shown in FIG. Figure 7 FIG., the determination of pixel scores and / or the detection of pixel-based segmentation in the tile or the image is performed using a third neural network that is trained using an image training set with local annotations of pixel-based segmentation masks. In one embodiment, the third neural network is a semantic segmentation neural network, such as a U-NET semantic segmentation neural network or another type of semantic segmentation neural network. In one embodiment, a semantic segmentation neural network that does not rely on a specific distribution in the intensity histogram can be used. Additionally, using such a neural network allows image segmentation to consider multi-channel images, such as red-green-blue (RGB) images. Thus, the segmentation not only depends on the histogram of pixel intensities but can also utilize the semantics of the image. At block 714, process 700 outputs pixel-based TLS segmentation in the histological image. Figure 8 FIG. Figure 7 FIG. shows a flowchart of one embodiment of a process 800 for training and validating a machine learning model to detect TLS segmentation in a histological image or a subject. In one embodiment, the machine learning model to be trained includes one or more individual models for the segmentation detection process described in Figure 7Other inputs described in. In one embodiment, the third neural network is a semantic segmentation neural network, such as a U-NET semantic segmentation neural network. In one embodiment, as Figure 9A shown, a training set of histological images with known local manual annotations including the TLS positions (i.e., TLS segmentation) in the images is used to train the third neural network. Process 800 performs a processing loop (blocks 804-816) on each training image to generate a set of segmentation masks from the tiles and predict the pixel-based TLS segmentation of the image. At block 806, process 800 detects a region of interest (ROI) in the training image. In one embodiment, as described above with respect to Figure 7 block 704, process 800 detects the ROI. At block 808, process 800 chunks the ROI into a set of tiles. In one embodiment, as described above with respect to Figure 7 block 706, process 800 chunks the ROI. At block 810, process 800 extracts segmentation masks from the tiles. In one embodiment, as described above with respect to Figure 7 block 708, process 800 extracts the segmentation masks. In one embodiment, process 800 generates a set of segmentation masks for the training image. At block 812, process 800 uses the extracted segmentation masks and the third neural network to detect pixel scores. In one embodiment, as described above with respect to Figure 7 block 712, process 800 uses the extracted segmentation masks and the third neural network to detect pixel scores. At block 814, process 800 predicts the pixel-based TLS segmentation in the image. In one embodiment, as described with respect to Figure 7 block 714, process 800 predicts the pixel-based TLS segmentation in the image. The process loop ends at 816.

[0191] To determine the sufficiency of training, at block 818, process 800 evaluates whether the prediction of the TLS segmentation by the machine learning model and the local manual annotation of the TLS status in the training dataset have converged. If they have converged, then at block 820, process 800 validates the machine learning model. If they have not converged, then at block 822, process 800 adjusts the machine learning model and re-executes the processing loop (blocks 804-816). As described above, the validation process can be performed by receiving a validation image set and the trained machine learning model and processing the validation image set.

[0192] In some embodiments, a trained feature extractor can be achieved after training a certain number of rounds. In some embodiments, training is performed until the accuracy is 1 or close to 1 (or 100%), until the AUC is 1 or close to 1 (or 100%), or until the loss is close to zero. In some embodiments, during the training of a feature extractor with contrastive loss, a large number of useful metrics may not be obtainable. Thus, one of the available metrics of the downstream task, such as AUC, can be monitored to understand how the feature extractor performs. In one example, to evaluate performance, the feature extractor trained at a certain round can be used to train a downstream weakly-supervised task. If additional training can improve the downstream performance, then such additional training may be necessary.

[0193] In some embodiments of the computer-implemented methods provided herein, the sample is cancer. The methods provided herein can be used for cancer samples of heterogeneous primary sites, stages, pathological types, subject profiles, or clinical / therapeutic courses. The methods provided herein can also be used for cancer samples of specific primary sites, statuses, pathological types, subject profiles, or clinical / therapeutic courses.

[0194] For example, the methods provided herein can be used for any cancer originating from any organ or tissue in the body. In some embodiments, the cancer is selected from the group consisting of lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, and melanoma. A cancer sample can be obtained from a subject at any time related to cancer diagnosis. For example, a sample can be obtained from a subject before or after the pathological diagnosis of cancer. A sample can be obtained from a subject before or after cancer treatment (e.g., immunotherapy, chemotherapy, surgery, radiation).

[0195] In some embodiments of the computer-implemented methods provided herein, the training set of tiles and / or the training set of histological images are digitized images of histological sections of heterogeneous cancers (e.g., lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma).

[0196] In some embodiments, the histological image is a digitized whole slide image (WSI). In some embodiments, the digitized histological image is an image of a histological section of a sample that has been stained with a dye to visualize the underlying tissue structure. The dye can be hematoxylin and eosin (H&E). Other common stains that can be used to visualize tissue structure in the input image include, for example, Masson's trichome stain, Periodic Acid Schiff stain, Prussian Blue stain, Gomori trichome stain, Alcian Blue stain, or Ziehl-Neelsen stain.

[0197] In some embodiments, the machine learning model is a deep multi-instance learning model. In some embodiments, the machine learning model is a Weldon model. In some embodiments, the machine learning model is applied to an entire group of tiles. In some embodiments, the machine learning model is applied to a subset of tiles. The training images can include digital images of tissue sections obtained from several control subjects. In some cases, the training images lack local annotations. The training images can include images associated with one or more global labels that indicate one or more TLS features of the patient from whom the sample was derived.

[0198] The machine learning model described herein can identify the TLS locations in histological images. The TLS locations can be identified, for example, by selecting queues of tiles having the highest M and lowest N scores. The highest and lowest tile queues identified by the model as having the best correlation with the presence or absence of TLS can be analyzed by a pathologist to determine the TLS features within the tiles. For example, in some embodiments, the TLS features can be determined by analyzing queues of tiles having M scores in the top 20% (e.g., top 15%, top 10%, top 5%, top 2%, top 1%, etc.) and / or N scores in the bottom 20% (e.g., bottom 15%, bottom 10%, bottom 5%, bottom 2%, bottom 1%, etc.) among all the tiles evaluated by the model.

[0199] TLS features include one or more of the following features and combinations thereof:

[0200] a. Structural features found in secondary lymphoid organs (SLOs) (e.g., lymph nodes, tonsils, spleen, Peyer's patches, or mucosa-associated lymphoid tissue), such as lymphoid follicles including dense cell aggregates similar to germinal centers in SLOs;

[0201] b. Non-encapsulated, non-lymphoid tissues, such as Peyer's patches and pre-existing lymphoid follicles;

[0202] c. Lymphocyte aggregates;

[0203] d. B cell follicles with actively replicating B cell germinal centers surrounded by T cell areas;

[0204] e. Mature dendritic cells in the T cell area and / or follicular dendritic cells in the B cell area (i.e., follicles); and

[0205] f. Heterogeneous cell populations and structures, including discrete B cell areas, T cell areas, marginal zones with activated macrophages and dendritic cells, a reticular fibroblast cell (RFC) network (or an RFC-like stromal network), a vasculature allowing extravasation of immune cells (e.g., high endothelial venules that are blood vessels expressing peripheral lymph node addressin (PNAd) and are specialized for extravasation of circulating immune cells), dendritic cell lysosome-associated membrane protein (DC-LAMP), and at least one of dendritic cells;

[0206] In some embodiments, histological images that measure the presence of TLSs contain one or more, two or more, three or more, four or more, five or more, or six of the above TLS features. In some embodiments, the histological image is a whole slide image. In other embodiments, the histological image is a part of a whole slide image, e.g., a tile derived from a whole slide image.

[0207] The presence or absence of the TLS features described herein can be determined in an image obtained from a subject's tissue. The image can be, for example, a whole slide image (WSI) or a part thereof, e.g., a tile derived from a WSI. In an exemplary embodiment, the tissue is derived from a biopsy obtained from a subject, e.g., a cancer biopsy. Suitable tissue sources for biopsies are known in the art and include, but are not limited to, tissue samples obtained from needle biopsies, endoscopic biopsies, or surgical biopsies. In an exemplary embodiment, the image is derived from a thoracentesis biopsy, a thoracoscopic biopsy, a thoracotomy biopsy, a needle biopsy, a laparoscopic biopsy, or a laparotomy biopsy.

[0208] Any suitable method and stain for histopathological analysis can be used to perform image analysis processing on tissue sections. For example, tissue sections can be stained with hematoxylin and eosin, alkaline phosphatase, methylene blue, Hoechst stain, and / or 4′,6-diamidino-2-phenylindole (DAPI).

[0209] A classification algorithm can calculate a TLS score for a subject, which TLS score indicates the TLS status. The TLS score can be a classification, e.g., the presence or absence of TLS. The TLS score can also be a continuous likelihood score. The continuous TLS score of a subject (e.g., a cancer subject) can be plotted relative to scores obtained from multiple subjects with known TLS status to determine the TLS status of the test subject.

[0210] II. Computer System and Machine - Readable Medium

[0211] As Figure 10 shown, computer system 1000 in the form of a data processing system includes a bus 1003 that couples to a microprocessor 1005, a ROM (Read - Only Memory) 1007, a volatile RAM 1009, and a non - volatile memory 1013. The microprocessor 1005 can include one or more CPUs, GPUs, specialized processors, and / or combinations thereof. The microprocessor 1005 can communicate with a cache 1004 and can retrieve instructions from memories 1007, 1009, 1013 and execute the instructions to perform the actions described above. The bus 1003 interconnects these different components together and also interconnects these components 1005, 1007, 1009, and 1013 to a display controller and display device 1015 and to peripheral devices, such as input / output (I / O) devices 1011 that can be a mouse, keyboard, modem, network interface, printer, and other devices well - known in the art. Typically, the input / output devices 1011 are coupled to the system through an input / output controller 1017. The volatile RAM (Random Access Memory) 1009 is typically implemented as a dynamic RAM (DRAM) that requires continuous power to refresh or maintain the data in the memory.

[0212] The non - volatile memory 1013 can be, for example, a magnetic hard disk drive or a magnetic optical disk drive or an optical disk drive or a DVD RAM or a flash memory or other type of memory system that retains data (e.g., large amounts of data) even after the system power is turned off. Typically, the non - volatile memory 1013 will also be a random access memory, although this is not required. While Figure 10 illustrates that the non - volatile memory 1013 is a local device directly coupled to the remaining components in the data processing system, it can be understood that the present invention can utilize non - volatile storage that is remote from the system, such as a network storage device coupled to the data processor system through a network interface such as a modem, an Ethernet interface, or a wireless network. The bus 1003 can include one or more buses interconnected by various bridges, controllers, and / or adapters well - known in the art.

[0213] The above-described portions may be implemented using a logic circuit, such as a dedicated logic circuit, or a microcontroller or other form of processing core that executes program code instructions. Thus, the processes taught by the above discussion may be executed by program code, such as machine-executable instructions, that cause a machine that executes the instructions to perform certain functions. In this case, the "machine" may be a machine that converts intermediate form (or "abstract") instructions into processor-specific instructions (e.g., an abstract execution environment, such as a "virtual machine" (e.g., Java virtual machine), interpreter, common language runtime, high-level language virtual machine, etc.) and / or an electronic circuit disposed on a semiconductor chip (e.g., a "logic circuit" implemented with transistors) designed to execute instructions, such as a general-purpose processor and / or a dedicated processor. The processes taught by the above discussion may also be executed by (in place of or in combination with the machine) an electronic circuit designed to execute the process (or a portion thereof) without executing program code.

[0214] The invention also relates to an apparatus for performing the operations described herein. The apparatus may be specially constructed for the required purpose or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium (such as, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory ("ROM"), random access memory ("RAM"), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions and coupled respectively to a computer system bus).

[0215] Machine-readable media include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, machine-readable media include: read-only memory ("ROM"); random access memory ("RAM"); magnetic disk storage media; optical storage media; flash memory devices; and the like.

[0216] An article of manufacture may be used to store program code. The article of manufacture storing the program code may be embodied as, but not limited to, one or more memories (e.g., one or more flash memories, random access memories (static, dynamic, or otherwise)), optical disks, CD-ROMs, DVD-ROMs, EPROMs, EEPROMs, magnetic or optical cards, or other types of machine-readable media suitable for storing electronic instructions. The program code may also be downloaded from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a data signal embodied in a propagation medium (e.g., via a communication link (e.g., a network connection)).

[0217] Algorithms and symbolic representations of operations on data bits within a computer memory are presented in the foregoing detailed description. These algorithmic descriptions and representations are tools used by those skilled in the data processing art to most effectively convey the substance of their work to other such skilled persons in the art. Here, an algorithm is generally considered to be a self-consistent sequence of operations leading to an expected result. An operation is an operation that requires a physical manipulation of physical quantities. These quantities, although not necessarily but typically, take the form of electrical or magnetic signals capable of being stored, transmitted, combined, compared, and otherwise manipulated. It has proven convenient, mainly for common reasons, to sometimes refer to these signals as bits, values, elements, symbols, characters, terms, or numbers, etc.

[0218] However, it should be borne in mind that all such and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise specifically stated as apparent from the foregoing discussion, it should be understood that throughout the specification, discussions using terms such as "partitioning", "chunking", "receiving", "computing", "extracting", "processing", "applying", "enhancing", "normalizing", "pre-training", "sorting", "selecting", "aggregating", or "classifying" refer to the actions and processes of a computer system or similar electronic computing device, which actions and processes will manifest as the manipulation and transformation of data represented as physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented as physical quantities in the memories or registers of the computer system or other such information storage, transmission, or display devices.

[0219] The processes and displays presented herein have no inherent relationship to any particular computer or other device. A variety of general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the described actions. From the description herein, the structure required for various such systems is apparent. In addition, the present invention is not described with reference to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the present invention described herein.

[0220] III. Product

[0221] In some aspects, the present disclosure provides products capable of detecting the TLS status (e.g., the presence, quantity, or location of TLS) in a subject, a histological image, a tile in a histological image, or a pixel in a tile. In some aspects, the product is connected to a scanner. In some aspects, the scanner is capable of scanning a pathology slide, such as an H&E slide. In some aspects, the product is particularly applicable to medical institutions, clinics, or providers, including those without cancer pathology expertise (including diagnosing or prognosticating cancer (e.g., lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, and melanoma) or diagnosing TLS). In some aspects, the product can be used to identify personalized drugs or targeted treatment options for cancer in a subject.

[0222] IV. General Considerations

[0223] As used herein, the articles "a" and "an" refer to one or more than one (i.e., at least one) of the grammatical object of the article. For example, "an element" means one element or more than one element, e.g., a plurality of elements.

[0224] The term "comprising" is used herein to mean the phrase "including but not limited to" and can be used interchangeably therewith. The term "containing" does not necessarily mean that there must be other elements in addition to the stated elements.

[0225] When referring to a number or a numerical range, the term "about" or "approximately" means that the indicated number or numerical region is an approximation within experimental variability (or within statistical experimental error), and thus, the numerical interval can vary, for example, between 1% and 20% of the indicated number or numerical range. In some aspects, "about" represents a value within 20% of the stated value. In preferred aspects, "about" represents a value within 10% of the stated value. In more preferred aspects, "about" represents a value within 1% of the stated value.

[0226] Unless otherwise specified, all numbers expressing component amounts, properties such as molecular weight, and reaction conditions used in the specification and claims should be understood to be modified in all instances by the term "about". Accordingly, unless otherwise indicated, the numerical properties set forth in the following specification and claims are approximations that can vary depending on the desired properties sought to be obtained in aspects of the present invention. Although numerical ranges and parameters setting forth the broad scope of the present invention are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. However, any numerical value inherently contains certain errors necessarily resulting from the errors found in their respective measurements.

[0227] The term "at least" before a number or series of numbers is understood to include the number adjacent to the term "at least" and all subsequent numbers or integers that can be clearly understood as logically included from the context. When "at least" precedes a series of numbers or a range, it is understood that "at least" can modify each of the numbers in the series or range.

[0228] As used herein, "not greater than" or "less than" is understood, based on the context logic, to be the value adjacent to the phrase and the logically lower numerical value or integer, down to zero (if a negative value is not possible). When "not greater than" precedes a series of numbers or a range, it is understood that "not greater than" can modify each of the numbers in the series or range.

[0229] As used herein, in the context of non - negative integers, "up to 10" in "up to 10" is understood to be up to and including 10, i.e., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10.

[0230] Where a range of values is provided, it is understood that each intermediate value between the upper and lower limits of the range (e.g., tenths of a unit of the lower limit, unless the context clearly dictates otherwise) and any other stated or intermediate value within the said range is included in the present invention. The upper and lower limits of these smaller ranges can be independently included in the smaller ranges and are also included in the present invention, subject to any specifically excluded limitations within the said range. Where the range includes one or both of the limits, the present invention also includes ranges excluding one or both of these included limits.

[0231] Examples

[0232] Example 1: Detection of TLS in a cohort of cancer patients

[0233] Using a pan-cancer cohort of 289 H&E-stained whole-slide images (WSIs) obtained from 289 patients, one slide per patient from Institut Bergonié, cross-validation of a machine learning model for predicting TLS status was performed. On these WSIs, the TLS status was manually examined and annotated by expert pathologists. The cohort consisted of WSIs from the following patients: 113 patients with non-small cell lung cancer (NSCLC) (39.1%), 45 patients with sarcoma (15.6%), 30 patients with bladder cancer (10.4%), 26 patients with colorectal cancer (9.0%), 10 patients with kidney cancer (3.7%), 10 patients with head and neck cancer (3.7%), 9 patients with ovarian cancer (3.1%), 7 patients with liver cancer, 5 patients with breast cancer, 5 patients with gastrointestinal stromal tumor, 4 patients with cervical cancer, 4 patients with endometrial cancer, 4 patients with gastric cancer, 3 patients with thyroid cancer, 2 patients with cholangiocarcinoma, 2 patients with prostate cancer, 2 patients with anal cancer, 2 patients with vulvar cancer, 1 patient with skin cancer, 1 patient with parotid gland cancer, 1 patient with gastrointestinal tract cancer, 1 patient with penile cancer, 1 patient with cancer of unknown primary site, and 1 patient with esophageal cancer. A deep learning (DL) model was trained on the WSIs to predict the TLS status (presence or absence of TLS) at the patient level.

[0234] The model was evaluated using five-fold cross-validation. The best-performing DL model provided two main components—a prediction score for the presence of TLS in a small region of the WSI with dimensions 112 μm × 112 μm (i.e., one tile) (tile score), and then an aggregation at the patient level—the ROC AUC score was 0.917 (standard deviation 0.036) ( Figure 4 ). The trained deep learning (DL) model provided the TLS status of the subjects, with sensitivities and specificities of 90% sensitivity and 68% specificity, 85% sensitivity and 85% specificity, and 80% sensitivity and 87% specificity, respectively.

[0235] The transferability of the DL model was evaluated using a validation cohort (PEMBROSARC) of 236 sarcoma WSI (subjects), which included 47 WSI (subjects) with a positive TLS status (i.e., presence of TLS) (19.9% of the entire cohort). The PEMBROSARC study was the first clinical trial to implement TLS status as an inclusion criterion (Italiano A. et al. 2022 Nat. Med. 28:1199-1206). The DL model detected the TLS status of the subjects, with an ROC AUC score of 0.89, and its sensitivity and specificity were: sensitivity 90%, specificity 64%; sensitivity 85%, specificity 88%; sensitivity 80%, specificity 86%. In summary, this study demonstrated the predictive ability of the DL model to detect the TLS status of subjects based on images of H&E-stained histological slides. The DL model provided in this article can be implemented as an efficient pre-screening tool for the TLS status of subjects in pathology laboratories and medical institutions.

[0236] The foregoing discussion describes only some exemplary embodiments of the present invention. Those skilled in the art will readily recognize from such discussion, the accompanying drawings, and the claims that various modifications can be made without departing from the spirit and scope of the invention.

Claims

1. A computer-implemented method for detecting the presence or absence of tertiary lymphoid structures (TLS) in a subject, comprising: Receiving a digitized histological image of a sample obtained from the subject; Partitioning the histological image into a set of tiles; Extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; And Classifying the histological image for the TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with a training set of histological images with known TLS annotations, wherein the TLS status indicates the presence or absence of TLS in the subject.

2. The computer-implemented method according to claim 1, wherein the classification comprises: Applying a first neural network to the one or more of the plurality of feature vectors, wherein the first neural network assigns a tile score to each tile in the set of tiles based on the one or more of the plurality of feature vectors, wherein the tile score represents the likelihood that the tile includes TLS; And Applying a second neural network to each tile, wherein the second neural network aggregates a subset of the tile scores of the set of tiles and determines the TLS status in the histological image, wherein: The first neural network is trained using a training set of histological images with known local annotations of the presence or absence of TLS at the tile level; and The first neural network and the second neural network are trained using a training set of histological images with known global annotations of the presence or absence of TLS at the histological image level.

3. The computer-implemented method according to claim 2, wherein the first neural network comprises a 1D convolutional layer.

4. The computer-implemented method according to claim 2 or 3, wherein the second neural network comprises a multi-layer perceptron model.

5. The computer-implemented method according to any one of claims 1 to 4, further comprising Detecting one or more locations where TLS is present in the histological image.

6. The computer-implemented method according to any one of claims 1 to 5, wherein each tile comprises a plurality of pixels, and wherein the method further comprises: Receiving a digitized histological image of a sample obtained from the subject; Partitioning the histological image into a set of tiles; Extracting a segmentation mask from each tile; Applying a third neural network to each tile, wherein the third neural network detects a pixel score using the extracted segmentation mask; And Detecting a pixel-based TLS segmentation within the tile or image, wherein the third neural network is trained using a training set of tiles comprising known pixel-based TLS segmentation masks within the tile.

7. The computer-implemented method according to claim 6, wherein the third neural network assigns a pixel score to each pixel of each tile in the set of tiles and determines the pixel-based segmentation of TLS within the tile, wherein the pixel score represents the likelihood that the pixel includes TLS.

8. The computer-implemented method according to claim 6 or 7, wherein the third neural network is a U-NET semantic segmentation neural network.

9. The method according to any one of claims 1 to 8, wherein the extraction step is performed by a ResNet50 neural network and / or a Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2) algorithm.

10. The computer-implemented method according to any one of claims 1 to 9, wherein the sample is cancer.

11. The computer-implemented method according to claim 10, wherein the cancer is selected from the group consisting of lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma, kidney cancer, head and neck cancer, liver cancer, breast cancer, gastrointestinal stromal tumor, cervical cancer, endometrial cancer, gastric cancer, thyroid cancer, cholangiocarcinoma, prostate cancer, anal cancer, vulvar cancer, skin cancer, parotid gland cancer, digestive tract cancer, penile cancer, esophageal cancer, and cancer of unknown primary site.

12. The computer-implemented method according to any one of claims 1 to 11, wherein the training set of the tiles and / or the training set of the histological images is a digital image of histological sections of heterologous cancers.

13. The computer-implemented method according to any one of claims 1 to 12, wherein the histological image is a digital whole slide image (WSI).

14. The computer-implemented method according to any one of claims 1 to 13, wherein the digital histological image is a digital image of a histological section stained with a dye.

15. The computer-implemented method according to claim 14, wherein the dye is hematoxylin and eosin (H&E).

16. The computer-implemented method according to any one of claims 2 to 15, wherein the classification step further comprises: sorting the set of tiles in the following manner: selecting the tile that contains the highest TLS tile score, and selecting the tile that contains the lowest TLS tile score.

17. The computer-implemented method according to any one of claims 1 to 16, comprising: repeating all the steps of claim 1 in a plurality of histological images, and processing the TLS status of the plurality of histological images, thereby detecting the presence or absence of TLS in a subject.

18. The computer-implemented method according to any one of claims 1 to 17, wherein the histological image lacks local annotation of histopathological features.

19. The computer-implemented method according to any one of claims 1 to 18, wherein each of the set of tiles comprises approximately 224×224 pixels.

20. A machine-readable medium having executable instructions to cause one or more processing units to perform a method for detecting the presence or absence of tertiary lymphoid structures (TLS) in a subject, the method comprising: receiving a digital histological image of a sample obtained from a subject; tiling the histological image into a set of tiles; extracting a plurality of feature vectors from each of the tiles, wherein each feature of the one or more feature vectors represents a local descriptor of the tile; and Classify the histological image for the TLS status using at least the plurality of feature vectors and a classification model, the classification model being trained with an imaging training set having known TLS annotations, wherein the TLS status indicates the presence or absence of TLS in a subject.

21. The machine-readable medium according to claim 20, wherein the classification comprises: Applying a first neural network to one or more of the plurality of feature vectors, wherein the first neural network assigns a tile score to each tile in the set of tiles based on one or more of the plurality of feature vectors, wherein the tile score represents the likelihood that the tile includes TLS; and Applying a second neural network to each tile, wherein the second neural network aggregates a subset of the tile scores of the set of tiles and determines the TLS status in the histological image, wherein: The first neural network is trained using a training set of histological images having known local annotations of the presence or absence of TLS at the tile level; and The first neural network and the second neural network are trained using a training set of histological images having known global annotations of the presence or absence of TLS at the histological image level.

22. The machine-readable medium according to claim 21, wherein the first neural network comprises a 1D convolutional layer.

23. The machine-readable medium according to claim 21 or 22, wherein the second neural network comprises a multi-layer perceptron model.

24. The machine-readable medium according to any one of claims 20 to 23, further comprising Detecting one or more locations where TLS is present in the histological image.

25. The machine-readable medium according to any one of claims 20 to 24, wherein each tile comprises a plurality of pixels, and wherein the method further comprises: Receiving a digitized histological image of a sample obtained from a subject; Partitioning the histological image into a set of tiles; Extracting a segmentation mask from each tile; Applying a third neural network to each tile, wherein the third neural network detects pixel scores using the extracted segmentation mask; and Detecting a pixel-based TLS segmentation within the tile or image, wherein the third neural network is trained using a training set of tiles having known pixel-based TLS segmentation masks within the tile.

26. The machine-readable medium according to claim 25, wherein the third neural network assigns a pixel score to each pixel of each tile in the set of tiles and determines the pixel-based segmentation of TLS within the tile, wherein the pixel score represents the likelihood that the pixel includes TLS.

27. The machine-readable medium according to claim 25 or 26, wherein the third neural network is a U-NET semantic segmentation neural network.

28. The machine-readable medium according to any one of claims 20 to 27, wherein the extraction step is performed by a ResNet50 neural network and / or a Momentum Contrast (MoCo) or Momentum Contrast v2 (MoCo v2) algorithm.

29. The machine-readable medium according to any one of claims 20 to 28, wherein the sample is cancer.

30. The machine-readable medium according to claim 29, wherein the cancer is selected from the group consisting of lung cancer, sarcoma, bladder cancer, colorectal cancer, ovarian cancer, pancreatic cancer, melanoma, kidney cancer, head and neck cancer, liver cancer, breast cancer, gastrointestinal stromal tumor, cervical cancer, endometrial cancer, gastric cancer, thyroid cancer, cholangiocarcinoma, prostate cancer, anal cancer, vulvar cancer, skin cancer, parotid gland cancer, digestive tract cancer, penile cancer, esophageal cancer, and cancer of unknown primary origin.

31. The machine-readable medium according to any one of claims 20 to 30, wherein the training set of the patches and / or the training set of the histological images are digital images of histological sections of heterologous cancers.

32. The machine-readable medium according to any one of claims 20 to 31, wherein the histological image is a digital whole slide image (WSI).

33. The machine-readable medium according to any one of claims 1 to 32, wherein the digital histological image is a digital image of a histological section stained with a dye.

34. The machine-readable medium according to claim 33, wherein the dye is hematoxylin and eosin (H&E).

35. The machine-readable medium according to any one of claims 21 to 34, wherein the classification step further comprises: sorting the set of patches in the following manner: selecting the patch that contains the highest TLS patch score, and selecting the patch that contains the lowest TLS patch score.

36. The machine-readable medium according to any one of claims 19 to 35, comprising: repeating all the steps according to claim 17 in a plurality of histological images, and processing the TLS status of the plurality of histological images, thereby detecting the presence or absence of TLS in a subject.

37. The machine-readable medium according to any one of claims 19 to 36, wherein the histological image lacks local annotation of histopathological features.

38. The method according to any one of claims 19 to 37, wherein each of the set of patches comprises approximately 224×224 pixels.