Analysis of histopathology samples
A computer-implemented method using deep learning and machine learning algorithms addresses the limitations of manual CD8+ cell analysis in histopathology, providing accurate immunophenotyping and predictive metrics for cancer therapy outcomes across multiple cancer types.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-09
AI Technical Summary
Existing automated approaches fail to accurately capture the manual immunophenotyping of CD8+ cells in histopathology images, which is subjective and labor-intensive, limiting their predictive capability for cancer therapy outcomes.
A computer-implemented method using deep learning and machine learning algorithms to identify CD8+ cells and classify tumor and stromal tissues, determining immunophenotypes as 'desert', 'excluded', or 'inflamed' based on CD8+ cell densities and distributions, with optional features like heatmaps and diagnostic plots.
Provides reproducible immunophenotype scores and diagnostic metrics, addressing subjectivity and labor issues, and is applicable across various cancer types, enhancing predictive capabilities for treatment responses.
Smart Images

Figure EP2025078232_09042026_PF_FP_ABST
Abstract
Description
[0001] Analysis of histopathology samples
[0002] Field of the Disclosure
[0003] The present disclosure relates to the analysis of histopathology samples. In particular, the present disclosure relates to methods of analysing histopathology images by identifying an immunophenotype and / or identifying CD8+ cells in an image of a tumour sample.
[0004] Background
[0005] Analysis of histopathology images is a cornerstone of modern pathology, particularly but not exclusively in the context of cancer. Particularly in the context of cancer immunotherapy, relationships between CD8+ T cells in the tumour environment and therapeutic outcome have been repeatedly demonstrated. Therefore, analysis of the presence, density and distribution of CD8+ T cells in histopathology images of tumours is particularly important to understand, characterise and predict patient response.
[0006] In 2023, Li et al. proposed a characterisation of tumours as belonging to one of three immune phenotypes based on the density and spatial distribution of CD8+ cells in and around a tumour: desert, excluded, and inflamed. They showed that manual immunophenotyping of tumours is predictive for clinical benefit by checkpoint inhibition-targeted therapy in two independent cohorts of two distinct indications.
[0007] However, manual immunophenotyping is associated with problematic levels of subjectivity and is extremely labour intensive.
[0008] Summary
[0009] The present inventors recognised that none of the existing automated approaches to capture aspects of CD8 cell density appropriately captures the manual approach, and that as a result, the manual approach remains the standard despite the severe limitations associated with the subjective and labour intensive nature of the process. Therefore, the present inventors set out to develop a new method for tumour immunophenotyping that better captures the manual approach, without suffering from the same limitations. The goal of the work was to compute reproducible immunophenotype scores identifying samples as Inflamed, Excluded and Desert, as well as providing diagnostic report metrics in both tumour nests and stroma compartments including e.g. densities, areas, ratios, CD8+ infiltration heatmaps, and other diagnostic plots. Further, the inventors also provided a new segmentation algorithm that delineates tumour nests and stroma areas as well as a CD8 cell identification method that is able to robustly address the variability of tissue appearance and is applicable various indications i.e., breast, colorectal, melanoma, malignant, non-small cell lung cancer, ovarian, pancreas, prostate, and renal cell carcinoma. Accordingly, a first aspect of the invention provides a computer-implemented method of identifying an immunophenotype for a tumour, the method comprising: receiving one or more whole slide images of a sample of the tumour that has been stained using at least a CD8 label; identifying, using a deep learning model, areas of the one or more whole slide images that show tumour nest tissue and areas that show stromal tissue; identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label; classifying each area of a plurality of tiles of the one or more whole slide images that show tumour nest tissue and each area of the plurality of tiles that show stromal tissue between at least a first class and a second class, wherein the tiles have a predetermined tile size and the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, and the second class comprises areas with a number of CD8 cells below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size; and classifying the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype based on the proportion of the area of the tiles that show stroma tissue classified in the first class and the proportion of the area of the tiles that show tumour nest tissue classified in the first class.
[0010] The methods of the first aspect may have any one or more of the following optional features.
[0011] Classifying the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype based on the proportion of the area of the tiles that show stroma tissue classified in the first class and the proportion of the area of the tiles that show tumour nest tissue classified in the first class may comprise: (i) classifying the one or more whole slide images as inflamed when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is at or above a second predetermined threshold; (ii) classifying the one or more whole slide images as desert when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is below the second predetermined threshold and the proportion of the area of the tiles that show stromal tissue classified in the first class is below a third predetermined threshold; and (iii) classifying the one or more whole slide images as excluded when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is below the second predetermined threshold and the proportion of the area of the tiles that show stromal tissue classified in the first class is at or above the third predetermined threshold. The second and third predetermined thresholds may be the same. The second and third predetermined thresholds may be 20%. In embodiments, classifying each area of a plurality of tiles of the one or more whole slide images between at least a first class and a second class comprises classifying the areas of the plurality of tiles between a first class, a second class and a third class, wherein the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, the second class comprises areas with a number of CD8 cells within a first predetermined range below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size, and the third class comprises areas with a number of CD8 cells within a second predetermined range below the first predetermined range scaled by the relative size of the area and the predetermined tile size.
[0012] The predetermined tile size, first predetermined threshold, and optional first and second predetermined ranges may have been identified by selecting a plurality of sets of candidate values of the predetermined tile size, first predetermined threshold, and optional first and second predetermined ranges, and: (a) for each candidate set of candidate values, displaying a heatmap of tile area classifications for one or more samples to a user; and selecting a set of candidate values that is identified by the user as accurately capturing areas of immune infiltration in the one or more samples; and / or (b) for each candidate set of candidate values, classifying a plurality of samples between whole slide images as one of a desert, excluded or inflamed immunophenotype using the method of any preceding claim; and selecting a set of candidate values that maximises the number of the plurality of samples that have the same classification as corresponding ground truth classification, optionally wherein ground truth classifications are obtained from annotations by multiple pathologists.
[0013] In embodiments, the predetermined tile size is between 100 pm x 100 pm and 200 pm x 200 pm, optionally 119.04 pm x 119.04 pm. In embodiments, the first predetermined threshold is 8, and / or the first predetermined range is 4 to 7 and the second predetermined range is 0 to 3. The first prredetermined threshold may 7 corresp rond to a CD8 density 7 of —# CP8 cMs« - — — — or tile area (mm2) (0.119042)
[0014] . . # CDS cells 8 . # CDS cells 8 between - ~ - and - ~ . tile area (mm2) (0.1002) tile area (mm2) (0.2002)
[0015] The tumour may be a solid tumour, optionally a squamous cell cancer and / or an anal carcinoma, cervical carcinoma, esophagus carcinoma, head and neck carcinoma, Thymoma or thymic carcinoma, melanoma, or non small cell lung cancer. The sample may be a formaldehyde-fixed paraffin-embedded (FFPE) tissue section. The whole slide images may be immunohistochemistry images.
[0016] In embodiments, the one or more whole slide images are images of a sample of the tumour that has been stained using a CD8 label, a proliferative cell label, optionally Ki67, and a nuclear label, optionally hematoxylin. In embodiments, the one or more whole slide images are images of a sample of the tumour that has been stained using a plurality of stains associated with respective colours, and the method comprises performing colour deconvolution prior to identifying CD8 cells and areas of the images that show tumour nest tissue or stromal tissue, thereby obtaining single channel images each corresponding to a respective stain, optionally wherein the plurality of stains comprise hematoxylin, and identifying, using a deep learning model, areas of the one or more whole slide images that show tumour nest tissue and areas that show stromal tissue is performed using single channel images corresponding to the hematoxylin stain.
[0017] In embodiments, classifying each area of a plurality of tiles of the one or more whole slide images between at least a first class and a second class, comprises: (i) obtaining an adjusted first predetermined threshold for a tile area based on the proportion of the area of the tile that shows tumour tissue, and classifying the tile area showing tumour tissue as a tile area that shows tumour tissue in the first class when the number of CD8 cells in the tumour tissue area in the tile is at or above the adjusted first predetermined threshold, wherein the adjusted first predetermined threshold corresponds to the first predetermined threshold scaled by the relative size of the tile area and the predetermined tile size, and (ii) obtaining an adjusted first predetermined threshold for the tile area based on the proportion of the area of the tile that shows stromal tissue, and classifying the tile area showing stromal tissue as a tile area that shows stromal tissue in the first class when the number of CD8 cells in the stromal tissue area in the tile is at or above the adjusted first predetermined threshold, wherein the adjusted first predetermined threshold corresponds to the first predetermined threshold scaled by the relative size of the tile area and the predetermined tile size.
[0018] In embodiments, the deep learning model is a deep learning model that has been trained to take as input a whole slide image and produce as output a pixel classification for each of a plurality of tiles of the whole slide image between a first class associated with tumour nest tissue, a second class associated with stromal tissue, and a third class associated with signal other than tumour nest tissue and stromal tissue. In embodiments, the method further comprises training the deep learning model using a plurality of training whole slide images and corresponding ground truth annotations indicating areas associated with tumour nest tissue and stromal tissue. The deep learning model may have been trained using a plurality of training whole slide images and corresponding ground truth annotations indicating areas associated with tumour nest tissue and stromal tissue. The training whole slide images may have any of the features described herein in relation to the one or more whole slides images that are being analysed. The training whole slide image may comprise whole slide images that have not been stained using a CD8 label. The training whole slide image may comprise H&E stained images. In embodiments, the deep learning model is a deep neural network, optionally a convolutional neural network. In embodiments, the method further comprises producing a heatmap that represents, for each area of at least one of the one or more whole slide images corresponding to one of the plurality of tiles, the classification associated with the tile areas, individually for areas that show tumour nest and / or for areas that show stromal tissue. In embodiments, the method further comprises producing a histogram of the number of tiles showing tumour nest tissue classified in each of the first, second and optionally third class and / or producing a histogram of the number of tiles showing stromal tissue classified in each of the first, second and optionally third class,. In embodiments, the method further comprises producing a histogram of the number of tiles showing tumour tissue comprising a number of CD8 cells within each of a plurality of bins. In embodiments, the method further comprises producing a histogram of the number of tiles showing stromal tissue comprising a number of CD8 cells within each of a plurality of bins. In embodiments, the method further comprises identifying a prognosis or treatment response for the tumour based on the classification of the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype.
[0019] In embodiments, identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label comprises: (i) applying a nucleus detection algorithm, (ii) quantifying, for each detected nucleus, a set of features comprising a plurality of nucleus morphology features, a plurality of nucleus appearance features, a plurality of nucleus background features, and a plurality of nucleus context features, and (iii) classifying each nucleus between a plurality of classes comprising at least a CD8- class and a CD8+ class using a machine learning model trained to classify a nucleus between the plurality of classes using as input the set of features for the nucleus.
[0020] In embodiments, the machine learning model comprises a support vector machine, a logistic regression model or a random forest model. In embodiments, the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model has been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label.
[0021] In embodiments, the one or more whole slide images are images of a sample of the tumour that has been stained using a plurality of stains associated with respective colours and comprising a nuclear stain, and identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label comprises performing colour deconvolution prior to applying the nucleus detection algorithm thereby obtaining single channel images each corresponding to a respective stain, wherein the nucleus detection algorithm is applied to the single channel image corresponding to the nuclear stain.
[0022] In embodiments, the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model has been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label using a two-step process comprising: a first step of classifying each nucleus between a plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label; and a second step of classifying each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label. The machine learning model may have been trained using training data comprising a plurality of training images associated with ground truth labels indicating whether a plurality of cells in the training images are CD8+ cells or CD8- cells. The method may further comprise training the machine learning model using training data comprising a plurality of training images associated with ground truth labels indicating whether a plurality of cells in the training images are CD8+ cells or CD8- cells. The ground truth labels may further indicate whether the plurality of cells are positive or negative for a proliferative cell label.
[0023] The machine learning model may have been trained to classify nuclei between 5 classes: a CD8+ class that is positive for the proliferative cell label (Ki67+CD8+), a CD8+ class that is negative for the proliferative cell label (Ki67-CD8+), a CD8- class that is positive for the proliferative cell label (Ki67+CD8-), a CD8- class that is negative for the proliferative cell label (Ki67-CD8-; non-target), and a class for artifacts. The machine learning model may comprise a plurality of machine learning models each configured (i.e. trained) to classify nuclei between a plurality of classes comprising a subset of the 5 classes.
[0024] In embodiments, the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67. The machine learning model may comprise: (i) a first machine learning model configured to classify each nucleus between a first plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label, and (ii) a second machine learning model configured to classify each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a second plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative forthe proliferative cell label. The first machine learning model may be configured to classify nuclei between: a class that is positive forthe proliferative cell label and negative for CD8, a class that is negative for the proliferative cell label and negative for CD8, and a class for all other nuclei (e.g. artifacts and any CD8+ cells). The first machine learning model may be configured to classify nuclei between: a CD8- class that is positive for the proliferative cell label, a CD8- class that is negative for the proliferative cell label, and a class for all other nuclei (e.g. any nucleus that is not CD8- or that is an artifact). The second machine learning model may be configured to classify nuclei between: a CD8+ class that is positive for the proliferative cell label, a CD8+ class that is negative for the proliferative cell label, and a class for all other nuclei (e.g. artifacts). In embodiments, the machine learning model comprises a first machine learning model trained to classify nuclei between a first plurality of classes, and a second machine learning model trained to classify nuclei between a second plurality of classes, wherein the method comprises classifying each nucleus using the first machine learning model, selecting nuclei classified in a subset of the first plurality of classes, and classifying the selected nuclei between the second plurality of classes using the second machine learning model.
[0025] In embodiments, for a nucleus of interest, the plurality of nucleus morphology features comprise features selected from: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus. In embodiments, for a nucleus of interest, the plurality of nucleus appearance features comprise features selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution. In embodiments, for a nucleus of interest, the plurality of nucleus background features comprise features computed in a ribbon of a predetermined number of pixels around the nucleus boundary, optionally wherein the features are selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, and the gradient magnitude in one or more channels after colour deconvolution. In embodiments, for a nucleus of interest, the plurality of nucleus context features comprise the number or proportion of nuclei within a predetermined distance from the nucleus of interest that are assigned to each of a plurality of clusters identified by applying a clustering algorithm to the plurality of nucleus morphology features, the plurality of nucleus appearance features and the plurality of nucleus background features for each of a plurality of nuclei in one or more training whole slide images. The plurality of nucleus morphology features may comprise features including or consisting of all of: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus. The plurality of nucleus morphology features may be computed using a segmented nucleus area based on an identified nucleus centre.
[0026] The plurality of nucleus appearance features may comprise features including or consisting of all of: the 10th, 50th and 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution. The plurality of nucleus appearance features may be computed using a segmented nucleus area based on an identified nucleus centre. The one or more channels may include one or more of: a nuclear marker channel (e.g. a hematoxylin channel), a luminance channel, and a marker of proliferative cells (e.g. Ki67 channel). The plurality of nucleus background features may comprise or consist of features computed in a ribbon of a predetermined number of pixels around the nucleus boundary, wherein the features include all of: the 10th, 50th and 95th percentile values of pixel intensities in one or more channels after colour deconvolution, and the gradient magnitude in one or more channels after colour deconvolution. The ribbon of a predetermined thickness may be a ribbon of approximately 75 pixels, or a number of pixels corresponding to approximately 30, 31 , 32, 33, 34, 35, 40, 45 or 50 pm from the nucleus. The plurality of nucleus background features may be computed using a segmented nucleus area based on an identified nucleus centre. The one or more channels may include one or more of: a nuclear marker channel (e.g. a hematoxylin channel), a luminance channel, and a marker of proliferative cells (e.g. Ki67 channel). Nuclei within a predetermined distance from the nucleus of interest may be nuclei within 20 pixels or a number of pixels corresponding to approximately 8, 9, 10, 11 , 12, 13, 14, 15 or 20 pm from the nucleus of interest. The clustering algorithm may be a k-means algorithm. The predetermined distance may be measured from detected nuclei centres.
[0027] In embodiments, the nucleus detection algorithm uses a radial-symmetry-based method using a series of kernels and iterative voting to identify nuclei centres. The nucleus detection algorithm may use Otsu’s method in a region centred around each identified nucleus centre to obtain a segmented nucleus area for the respective nucleus.
[0028] According to a second aspect, there is provided a method of identifying CD8 cells in one or more whole slide images of a tumour, the method comprising: (i) receiving the one or more whole slide images, wherein the one or more whole slide images are images of a sample of the tumour that has been stained using at least a CD8 label; (ii) applying a nucleus detection algorithm, (iii) quantifying, for each detected nucleus, a set of features comprising a plurality of nucleus morphology features, a plurality of nucleus appearance features, a plurality of nucleus background features, and a plurality of nucleus context features, and (iv) classifying each nucleus between a plurality of classes comprising at least a CD8- class and a CD8+ class using a machine learning model trained to classify a nucleus between the plurality of classes using as input the set of features for the nucleus.
[0029] Methods according to the present aspect may have any of the features described in relation to the first aspect. In particular, methods of the present aspect may have any one or more of the following optional features.
[0030] The machine learning model may comprise a support vector machine, a logistic regression model or a random forest model. The machine learning model may be a random forest model. The machine learning model may comprise a first machine learning model trained to classify nuclei between a first plurality of classes, and a second machine learning model trained to classify nuclei between a second plurality of classes, wherein the method comprises classifying each nucleus using the first machine learning model, selecting nuclei classified in a subset of the first plurality of classes, and classifying the selected nuclei between the second plurality of classes using the second machine learning model. The first and second machine learning models may be random forest models. The images may be images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model may have been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label. The images may be images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67. The machine learning model may have been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label using a two-step process comprising: a first step of classifying each nucleus between a plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label; and a second step of classifying each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label.
[0031] The machine learning model may have been trained using training data comprising a plurality of training images associated with ground truth labels indicating whether a plurality of cells in the training images are CD8+ cells or CD8- cells. The method may further comprise training the machine learning model using training data comprising a plurality of training images associated with ground truth labels indicating whether a plurality of cells in the training images are CD8+ cells or CD8- cells. The ground truth labels may further indicate whether the plurality of cells are positive or negative for a proliferative cell label. The machine learning model may have been trained to classify each nucleus between 5 classes: a CD8+ class that is positive for the proliferative cell label (Ki67+CD8+), a CD8+ class that is negative for the proliferative cell label (Ki67-CD8+), a CD8- class that is positive for the proliferative cell label (Ki67+CD8-), a CD8- class that is negative for the proliferative cell label (Ki67-CD8-; non-target), and a class for artifacts. The machine learning model may comprise a plurality of machine learning models each configured (i.e. trained) to classify nuclei between a plurality of classes comprising a subset of the 5 classes. The machine learning model may comprise: (i) a first machine learning model configured to classify each nucleus between a first plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label, and (ii) a second machine learning model configured to classify each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a second plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label. The first machine learning model may be configured to classify nuclei between: a class that is positive for the proliferative cell label and negative for CD8, a class that is negative for the proliferative cell label and negative for CD8, and a class for all other nuclei (e.g. artifacts and any CD8+ cells). The first machine learning model may be configured to classify nuclei between: a CD8- class that is positive for the proliferative cell label, a CD8- class that is negative for the proliferative cell label, and a class for all other nuclei (e.g. any nucleus that is not CD8- or that is an artifact). The second machine learning model may be configured to classify nuclei between: a CD8+ class that is positive for the proliferative cell label, a CD8+ class that is negative for the proliferative cell label, and a class for all other nuclei (e.g. artifacts).
[0032] The one or more whole slide images may be images of a sample of the tumour that has been stained using a plurality of stains associated with respective colours and comprising a nuclear stain, and the method may comprise performing colour deconvolution prior to applying the nucleus detection algorithm thereby obtaining single channel images each corresponding to a respective stain, wherein the nucleus detection algorithm is applied to the single channel image corresponding to the nuclear stain.
[0033] The plurality of nucleus morphology features may comprise features selected from: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus. The plurality of nucleus appearance features may comprise features selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution. The plurality of nucleus background features may comprise features computed in a ribbon of a predetermined number of pixels around the nucleus boundary. The features may be selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, and the gradient magnitude in one ormore channels after colour deconvolution. The plurality of nucleus context features may comprise the number or proportion of nuclei within a predetermined distance from the nucleus of interest that are assigned to each of a plurality of clusters identified by applying a clustering algorithm to the plurality of nucleus morphology features, the plurality of nucleus appearance features and the plurality of nucleus background features for each of a plurality of nuclei in one or more training whole slide images. The plurality of nucleus morphology features may comprise features including or consisting of all of: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus. The plurality of nucleus morphology features may be computed using a segmented nucleus area based on an identified nucleus centre. The plurality of nucleus appearance features may comprise features including or consisting of all of: the 10th, 50th and 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution. The plurality of nucleus appearance features may be computed using a segmented nucleus area based on an identified nucleus centre. The one or more channels may include one or more of: a nuclear marker channel (e.g. a hematoxylin channel), a luminance channel, and a marker of proliferative cells (e.g. Ki67 channel). The plurality of nucleus background features may comprise or consist of features computed in a ribbon of a predetermined number of pixels around the nucleus boundary, wherein the features include all of: the 10th, 50th and 95th percentile values of pixel intensities in one or more channels after colour deconvolution, and the gradient magnitude in one or more channels after colour deconvolution. The ribbon of a predetermined thickness may be a ribbon of approximately 75 pixels, or a number of pixels corresponding to approximately 30, 31 , 32, 33, 34, 35, 40, 45 or 50 pm from the nucleus. The plurality of nucleus background features may be computed using a segmented nucleus area based on an identified nucleus centre. The one or more channels may include one or more of: a nuclear marker channel (e.g. a hematoxylin channel), a luminance channel, and a marker of proliferative cells (e.g. Ki67 channel). Nuclei within a predetermined distance from the nucleus of interest may be nuclei within 20 pixels or a number of pixels corresponding to approximately 8, 9, 10, 11 , 12, 13, 14, 15 or 20 pm from the nucleus of interest. The clustering algorithm may be a k-means algorithm. The predetermined distance may be measured from detected nuclei centres.
[0034] The nucleus detection algorithm may use a radial-symmetry-based method using a series of kernels and iterative voting to identify nuclei centres. The nucleus detection algorithm may use Otsu’s method in a region centred around each identified nucleus centre to obtain a segmented nucleus area for the respective nucleus. The sample may be a formaldehyde-fixed paraffin- embedded (FFPE) tissue section. The whole slide images may be immunohistochemistry images. The one or more whole slide images may be images of a sample of the tumour that has been stained using a CD8 label, a proliferative cell label, optionally Ki67, and a nuclear label, optionally hematoxylin
[0035] According to a third aspect, there is provided a computer implemented method of determining a prognosis or treatment prediction response for a subject, the method comprising: identifying an immunophenotype for a tumour of the subject using one or more whole slide images of a sample of the tumour and the method of any embodiment of the first or second aspects; and identifying a prognosis or treatment response based on the identified immunophenotype.
[0036] Any of the methods described herein may be computer-implemented.
[0037] According to a fourth aspect, there is provided a system including: at least one processor; and at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to implement any of the methods described herein. For example, the system may be configured to implement the methods of any embodiment of any of the first, second, and / or third aspects.
[0038] According to a fifth aspect, there is provided a non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform any of the methods described herein. For example, the instructions may cause the at least one processor to implement the methods of any embodiment of any of the first, second, and / or third aspects. According to a sixth aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform any of the methods described herein. For example, the code may cause the computer to implement the methods of any embodiment of any of the first, second, and / or third aspects.
[0039] Brief Description of the Figures
[0040] Figure 1 is a flowchart illustrating schematically a method identifying an immunophenotype for a tumour, according to the disclosure.
[0041] Figure 2 is a flowchart illustrating schematically a method of identifying CD8 cells in a sample, according to the disclosure.
[0042] Figure 3 shows an embodiment of a system for analysing histopathology samples.
[0043] Figure 4 shows data on inter-observer correlation agreements among three observers for evaluating Ki67+CD8-, Ki67+CD8+, and Ki67-CD8+.
[0044] Fig. 4A shows a scatterplot of agreement between 3 observers for identification of Ki67+CD8- cells.
[0045] Fig. 4B shows a scatterplot of agreement between 3 observers for identification of Ki67+CD8+ cells.
[0046] Fig. 4C shows a scatterplot of agreement between 3 observers for identification of Ki67- CD8+ cells.
[0047] Fig. 4D shows summary statistics for the agreement between 3 observers in the identification of Ki67+CD8-, Ki67+CD8+, and Ki67-CD8+.
[0048] Figure 5 shows data on agreement between cell types identified by an algorithm as described herein and an average observer.
[0049] Fig. 5A shows a scatterplot of agreement between cell types identified by an algorithm as described herein and an average observer for Ki67+CD8- cells.
[0050] Fig. 5B shows a scatterplot of agreement between cell types identified by an algorithm as described herein and an average observer for Ki67+CD8+ cells.
[0051] Fig. 5C shows a scatterplot of agreement between cell types identified by an algorithm as described herein and an average observer for Ki67-CD8+ cells.
[0052] Fig. 5D shows summary statistics for the identification of Ki67+CD8-, Ki67+CD8+, and Ki67-CD8+ using an algorithm as described herein, compared to corresponding labels from three observers.
[0053] Figures 6A, B, C show examples of cell identification performed by an algorithm as described herein (left) and three observers for the same sample. Fig. 7 shows a flowchart illustrating a general approach for tumour immunophenotyping.
[0054] Fig. 8A illustrates the general principles of a deep learning colour-invariant segmentation approach for bright field images (immunohistochemistry and H&E).
[0055] Fig. 8B illustrates the principle of colour deconvolution and colour convolution.
[0056] Fig. 9 illustrates the general principles of an immunophenotyping algorithm described herein.
[0057] Fig. 10 is a flowchart illustrating the classification of samples using an immunophenotype score.
[0058] Fig. 11 shows an example of a tile comprising both tumour and stroma, illustrating the application of the algorithm of Fig. 9 to such a tile.
[0059] Fig. 12 shows heatmaps of CD8 cell density for the same whole slide image, with densities (and therefore heatmap colours) obtained using increasing tile sizes from left to right.
[0060] Fig. 13A shows an example of a stroma CD8 infiltration heatmap created by the immunophenotyping algorithm described herein.
[0061] Fig. 13B shows an example of a tumour nest CD8 infiltration heatmap created by the immunophenotyping algorithm described herein.
[0062] Fig. 14 shows examples heatmaps, histograms and statistical metrics for a sample classified as “Desert”.
[0063] Fig. 14A shows a histogram of tile counts by CD8+ cell count bins, in stromal regions.
[0064] Fig. 14B shows the stroma CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 14A.
[0065] Fig. 14C shows a histogram of tile counts by CD8+ cell count bins, in tumour nest regions.
[0066] Fig. 14D shows the tumour nest CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 14C.
[0067] Fig. 14E shows on the left the stroma CD8 infiltration heatmap and on the right the immunophenotype score (Desert = Immunophenotype 0) and statistical metrics computed for the tumour nest and stroma regions.
[0068] Fig. 15 shows examples heatmaps, histograms and statistical metrics for a sample classified as “Excluded”.
[0069] Fig. 15A shows a histogram of tile counts by CD8+ cell count bins, in stromal regions.
[0070] Fig. 15B shows the stroma CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 15A.
[0071] Fig. 15C shows a histogram of tile counts by CD8+ cell count bins, in tumour nest regions. Fig. 15D shows the tumour nest CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 15C.
[0072] Fig. 15E shows on the left the stroma CD8 infiltration heatmap and on the right the immunophenotype score (Excluded = Immunophenotype 1) and statistical metrics computed for the tumour nest and stroma regions.
[0073] Fig. 16 shows examples heatmaps, histograms and statistical metrics for a sample classified as “Inflamed”.
[0074] Fig. 16A shows a histogram of tile counts by CD8+ cell count bins, in stromal regions.
[0075] Fig. 16B shows the stroma CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 16A.
[0076] Fig. 16C shows a histogram of tile counts by CD8+ cell count bins, in tumour nest regions.
[0077] Fig. 16D shows the tumour nest CD8 infiltration heatmap of the sample with colours corresponding to the CD8+ density bins of Fig. 16C.
[0078] Fig. 16E shows on the left the stroma CD8 infiltration heatmap and on the right the immunophenotype score (Inflamed = Immunophenotype 2) and statistical metrics computed for the tumour nest and stroma regions.
[0079] Fig. 17 illustrates a process for nuclei detection using an iterative voting method as described in Parvin et al. 2007.
[0080] Fig. 18 shows agreement scores among raters for an immunphenotyping method, using data from Study 1. Cohen’s Kappa scores are computed to measure the agreement between 2 raters and Fleiss Kappa scores for the agreement among 3 or more raters. All refers to CellCarta, Algorithm (according to an embodiment of the disclosure), Pathologist 1 and Pathologist 2.
[0081] Fig. 19 shows agreement scores among raters for an immunophenotyping method, using data from Study 2. Cohen’s Kappa scores are computed to measure the agreement between 2 raters and Fleiss Kappa scores for the agreement among 3 or more raters. All refers to Algorithm (according to an embodiment of the disclosure), Pathologist 1 and Pathologist 2.
[0082] Detailed Description
[0083] Certain aspects and embodiments of the invention will now be illustrated by way of example and with reference to the figures described above.
[0084] In describing the present invention, the following terms will be employed, and are intended to be defined as indicated below.
[0085] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0086] A “sample” as used herein may be a cell or tissue sample (e.g. a biopsy), from which histopathology images can be obtained, typically after fixation and / or labelling. The sample may be a cell or tissue sample that has been obtained from a subject. In particular, the sample may be a tumour sample. The sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to making a determination (e.g. frozen, fixed or subjected to one or more purification, enrichment or labelling steps). Alternatively, the sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells. The sample may be a eukaryotic cell sample or a prokaryotic cell sample. A eukaryotic cell sample may be a mammalian sample, such as a human or mouse sample. Further, the sample may be transported and / or stored, and collection may take place at a location remote from the sample processing location (e.g. labelling, fixation etc.), and / or the image data acquisition location. Further, the computer-implemented method steps may take place at a location remote from the sample collection and / or processing location(s) and / or remote from the image data acquisition location. For example, the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider.
[0087] A sample may be a histopathology sample, specifically a histopathology sample (e.g. a tissue section, tissue sample, tissue microarray or biopsy). A histopathology sample may be a tumour section, tumour tissue microarray or tumour biopsy. A tumour section may be a whole-tumour section. A whole-tumour section is typically a section cut from a surgically resected tumour, thus representing the characteristics of the whole tumour. Thus, a whole-tumour section may be a surgically resected section. A histopathology sample may be a biopsy obtained from a tumour. A histopathology sample may be a sample forming part of a tissue microarray (also referred to as a tissue section in a microarray block). A tissue microarray is a slide comprising multiple tissue sections from multiple original samples or locations within a sample.
[0088] The histopathology sample may be a stained sample. Staining facilitates morphological analysis of tumour sections by colouring cells, subcellular structures and organelles. Any type of staining that facilitates morphological analysis may be used. For example, the sample may have been stained using Haematoxylin and eosin staining, Papanicolaou (PAP) staining (a combination of haematoxylin, Orange G, eosin Y, Light Green SF yellowish, and sometimes Bismarck Brown Y), Masson's trichrome stain, Romanowsky stain, Silver staining, or any combination of one or more known histological dyes such as acridine orange, Bismarck brown, carmine, Coomassie blue, cresyl violet, DAPI, eosin, ethidium bromide, Hoechst stains, iodine, malachite green, methyl green, methylene blue, neutral red, nile blue, nile red, osmium tetraoxide, propium iodide, rhodamine, and safranine. The histopathology sample may be stained with hematoxylin and eosin (H&E). H&E stain is the most commonly used stain in histopathology for medical diagnosis, particularly for the analysis of biopsy sections of suspected cancers by pathologists. Thus H&E- stained histopathology samples are usually readily available as part of large datasets collated for the study of cancer. The histopathology sample may instead or in addition be a sample that has been stained using one or more labels that are associated with specific cellular markers of interest. A common way to associate a label with a cellular marker is through the use of labelled affinity reagents, such as labelled antibodies. Thus, the cellular markers of interest may also be referred to as “antigens”. The labels may be detectable through a chromogenic reaction, where the affinity reagent is conjugated to an enzyme (e.g. a peroxidase) that can catalyse a colour-producing reaction. Alternatively, the labels may be detected through fluorescence, where the affinity reagent is associated with a fluorophore (such as e.g. fluorescein or rhodamine). The latter may be referred to as immunofluorescence. The former may be referred to as chromogenic immunohistochemistry, or simply immunohistochemistry. The term “immunohistochemistry” is commonly used as an umbrella term to encompass detection via chromogenic immunohistochemistry and via immunofluorescence.
[0089] The methods described herein apply to images of histopathology samples. A “whole slide image” (WSI) refers to an image that shows s substantial portion of a histopathology sample. This is by contrast to a “tile” or “patch” which refers to a subsection of a whole slide image. Thus, a whole slide image refers to an image of a histopathology sample from which a plurality of tiles can be obtained. Tiles are typically obtained as fixed size subsections of an image by applying a grid or stride of a predetermined size, or equivalently by dividing an image of a predetermined size in a predetermined number of areas of predetermined (optionally equal) size. For example, tiles can be obtained by applying a grid of 32 x 32 tiles. Tiles can be non-overlapping or partially overlapping. Partially overlapping tiles can overlap by a predetermined percentage of their surface area with one or more adjacent tiles. A whole slide image as used herein for the purpose of analysing a sample or training a model does not necessarily comprise all of the tiles of an original whole slide image. For example, a set of tiles obtained from an original whole slide image may be filtered, e.g. to remove any tiles where the percentage area of the tile that shows a tissue is below a predetermined threshold. As another example, instead or in addition to this, a section (also referred to as a region) of an original whole slide image may be selected and analysed individually. Such a section may also be referred to as a “whole slide image”.
[0090] A plurality of machine learning models are referred to herein, including a deep learning model used to segment stroma / tumour nest areas, and a machine learning model used to classify nuclei.
[0091] A deep learning model used to identify areas of the one or more whole slide images that show tumour tissue (also referred to herein as “tumour nests”) and areas that show stromal tissue (also referred to herein as a “stoma / tumour segmentation model) is a machine learning model that has been trained in a supervised manner to classify pixels or image tiles between a plurality of classes comprising at least a tumour nest class and a stroma class. Any machine learning model that can be used for image segmentation can be used for this purpose. This model is typically a deep neural network, such as e.g. a convolutional neural network (CNN) such as e.g. ResNet (e.g. ResNet50, or derivatives thereof such as DeepLabv3, FCN-Resnet50, etc), or a transformer-based model such as a ViT. The tumour / stroma segmentation model is configured to take as input an image or image tile, and produce as output a classification associated with the image or image tiles (specifically each pixel in the image or tile).
[0092] A machine learning model used to classify nuclei is a classification model. Any feature-based classifier known in the art may be used, including e.g. a random forest classifier, a support vector machine, a logistic regression model, a decision tree, a neural network, a k-nearest neighbour classifier, a naive Bayes classifier, etc. The nuclei classification model is configured to take as input a feature vector associated with a segmented nucleus, and produce as output a classification associated with the segmented nucleus.
[0093] As used herein, the terms “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above described embodiments. For example, a computer system may comprise a central processing unit (CPU), one or more graphics processing units (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. Preferably the computer system has a display or comprises a computing device that has a display to provide a visual output display (for example in the design of the business process). The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.
[0094] As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.
[0095] As the skilled person understands, the complexity of the operations described herein (due at least to the amount of data that is analysed and the complexity of the machine learning models used) are such that they are beyond the reach of a mental activity. Thus, unless context indicates otherwise (e.g. where sample preparation or acquisition steps are described), all steps of the methods described herein are computer implemented. Previous automated histopathology image analysis methods that used CD8+ cells identification have been proposed. For example, WO 2022 / 017666 describes a method that uses CD8+ cell densities in tumour nests vs CD8+ cell densities in stroma (e.g. number of stained CD8+ cells per unit area in the segmentation set corresponding to tumour nest tissue and number of stained CD8+ cells per unit area in the segmentation set corresponding to stroma tissue - overall, not per tile), before and after treatment. However, this method does not consider local densities or do an immunophenotyping classification. As such, it is unable to replicate the gold standard manual pathologist approach, which takes into account the proportion of tissue with high density of CD8+ cells (with both the proportion of tissue and the “high density” being qualitatively estimated). ). Li et al. 2023 describes a manual approach for immunophenotyping between desert, inflamed and excluded tumours. Other groups have used the same 3 categories since then, but there is no clear and established consensus on how the 3 phenotypes should be scored. The authors in Li et al. 2023 also proposed a series of automated schemes as alternative to the manual one, none of which capture the features that are relevant in the manual scheme, as further described below. Dejardin et al. 2024 describes a composite decision rule based on CD8+ T-cell density in paired biopsies taken before and on-treatment, as an indicator of treatment response (in particular correlating with progression free survival). This combines a threshold on the fold change on a natural log scale in CD8+ T cells from baseline, and a threshold on geometric mean of the on treatment density of T cells. This does not perform area-specific CD8+ densities counts or immunophenotyping.
[0096] Analysing pathology images
[0097] The present disclosure provides method for analysing histopathology samples, using image data acquired from the sample, including in particular a method of identifying an immunophenotype for a tumour sample, and a method of identifying CD8+ cells in an image of a histopathology sample. An illustrative method of identifying an immunophenotype for a tumour will be described by reference to Figure 1. Figure 1 describes steps performed in analysing a histopathology sample, and Figure 2 describes steps performed in a method of identifying CD8+ cells in an image of a histopathology sample. The methods of Figures 1 and 2 may each be performed independently, or may both be performed. For example, the steps of Figure 2 may optionally be performed in a method of analysing a histopathology sample according to Figure 1.
[0098] At step 10, an image of a histopathology sample is obtained. The sample may be a stained sample, for example an IHC stained sample as described elsewhere herein. The image may be referred to as a digital pathology image. The image may have been acquired using any imaging technology suitable for obtaining images of histopathology samples. The image may be received from a user interface, a data store or an image acquisition means. The image may be a multicolour image such as an RGB image. The sample is a sample of a tumour to be analysed, which has been stained using at least a CD8 label. A CD8 label is any detectable label that selectively labels cells that express the CD8 protein. The tumour sample may be associated with one or more whole slide images (WSI), and the method may be repeated for each of these images individually, or combined. The following will refer to a single WSI for simplicity, but the term encompasses performing the method fora plurality of WSIs. The WSI may be an image of a sample of the tumour that has been stained using a plurality of stains associated with respective colours. For example, the plurality of stains may include a counterstain such as H&E, and a CD8 stain. The plurality of stain may further include a proliferative cell label, such as Ki67. At optional step 12, the colour deconvolution is performed, thereby obtaining single channel images each corresponding to a respective stain. This may be performed prior to identifying CD8 cells and / or prior to identifying areas of the images that show tumour nest tissue or stromal tissue. Indeed, either or both of these steps may use colour deconvoluted images.
[0099] At step 14, areas of the one or more whole slide images that show tumour nest tissue and areas that show stromal tissue are identified, using a deep learning model. At step 16, CD8 cells are identified in the one or more whole slide images using a signal associated with the CD8 label. This can be performed using an algorithm as described by reference to Fig. 2, or using any other algorithm or expert annotations. At step 18, a plurality of tiles of a predetermined size are obtained.
[0100] Step 20 comprises classifying each area of a plurality of tiles of the one or more whole slide images that show tumour nest tissue and each area of the plurality of tiles that show stromal tissue between at least a first class (hot) and a second class (non-hot), wherein the tiles have a predetermined tile size and the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, and the second class comprises areas with a number of CD8 cells below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size. In embodiments, classifying each area of a plurality of tiles of the one or more whole slide images between at least a first class and a second class comprises classifying the areas of the plurality of tiles between a first class (hot), a second class (medium) and a third class (cold), wherein the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, the second class comprises areas with a number of CD8 cells within a first predetermined range below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size, and the third class comprises areas with a number of CD8 cells within a second predetermined range below the first predetermined range scaled by the relative size of the area and the predetermined tile size.
[0101] At optional step 22, the total proportion of stroma area classified in the first class (i.e. % of the stroma area that is classified in the first class) and the total proportion of tumour nest area classified in the first class (i.e. % of the tumour nest area that is classified in the first class) are calculated. At step 24, the WSI is classified as one of a desert, excluded or inflamed immunophenotype based on the proportion of the area of the tiles that show stroma tissue classified in the first class and the proportion of the area of the tiles that show tumour nest tissue classified in the first class. At optional step 26, one or more heatmaps may be produced. A first heatmap may be produced that represents, for each area of at least one of the one or more whole slide images corresponding to one of the plurality of tiles, the classification associated with the tile areas, individually for areas that show tumour nest. A second heatmap may be produced that represents, for each area of at least one of the one or more whole slide images corresponding to one of the plurality of tiles, the classification associated with the tile areas, individually for areas that show stromal tissue. At steo 26, one or more histograms may be produced, or summary statistics associated with such histograms. This may include producing a histogram of the number of tiles showing tumour nest tissue classified in each of the first, second and optionally third class and / or producing a histogram of the number of tiles showing stromal tissue classified in each of the first, second and optionally third class and / or producing a histogram of the number of tiles showing tumour tissue comprising a number of CD8 cells within each of a plurality of bins and / or producing a histogram of the number of tiles showing stromal tissue comprising a number of CD8 cells within each of a plurality of bins.
[0102] At optional step 28, a prognosis or treatment response for the tumour may be identified based on the classification of the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype. For example, a tumour identified as inflamed prior to immunotherapy may be more likely to respond to immunotherapy than a tumour identified as desert or excluded and / or has a better prognosis than a tumour identified as desert or excluded. The method may further comprise selecting a subject from whom the sample has been obtained for treatment using said identified treatment response. For example, the subject may be selected for immunotherapy when the tumour is classified as inflamed. The subject may be selected for treatment other than immunotherapy when the subject is identified as desert. The immunotherapy may be an immune checkpoint inhibitor therapy, such as e.g. a PD-L1 inhibitor. At step 30, one or more results of the analysis may be provided to a user. This may include any of the information obtained at step 20, or any information derived therefrom, such as e.g. the information mentioned at steps 22, 24, 26 and / or 28.
[0103] The method used at step 16 may be a method as described in relation to Figure 2, or any other method for identifying cells positive for a label using images of a sample stained with the label, preferably methods specifically developed to identify CD8+ cells in tumours. At step 110, one or more histopathology images are obtained. These may be the images being analysed in the method of Fig. 1 , or may be images being analysed in a different context. These images may have any of the features described herein in relation to histopathology images, and in particular any of the features described in relation to the image used in Figure 1 . The images are images of a sample of a tumour that has been stained using a CD8 label and a proliferative cell label, such as e.g. Ki67. At step 112, colour deconvolution is optionally performed if not already performed at step 12 in the context of a method of Fig. 1. At step 114, a nucleus detection algorithm is applied. The nucleus detection algorithm may use a radial-symmetry-based method using a series of kernels and iterative voting to identify nuclei centres. The nucleus detection algorithm may use Otsu’s method in a region centred around each identified nucleus centre to obtain a segmented nucleus area for the respective nucleus. Any nucleus detection (also referred to as nucleus segmentation) algorithm known in the art may be used instead. The nucleus detection algorithm may be applied to the single channel image corresponding to a nuclear stain, when colour deconvolution is performed at step 112.
[0104] At step 116, a set of features is quantified for each detect nucleus. This includes a plurality of nucleus morphology features, a plurality of nucleus appearance features, a plurality of nucleus background features, and a plurality of nucleus context features. The plurality of nucleus morphology features may comprise features selected from: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus. The plurality of nucleus appearance features may comprise features selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution. The plurality of nucleus background features may comprise features computed in a ribbon of a predetermined number of pixels around the nucleus boundary, optionally wherein the features are selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, and the gradient magnitude in one or more channels after colour deconvolution. The plurality of nucleus context features may comprise the number or proportion of nuclei within a predetermined distance from the nucleus of interest that are assigned to each of a plurality of clusters identified by applying a clustering algorithm to the plurality of nucleus morphology features, the plurality of nucleus appearance features and the plurality of nucleus background features for each of a plurality of nuclei in one or more training whole slide images.
[0105] At step 118, each nucleus is classified between a plurality of classes comprising at least a CD8- class and a CD8+ class using a machine learning model trained to classify a nucleus between the plurality of classes using as input the set of features for the nucleus. Step 118 may comprise using a first machine learning model to classify each nucleus between a plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label, and using a second machine learning model configured to of classify each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label. At step 120, the results of the classification (or parts thereof) are provided to a user or computer processor or memory, for example for use as part of a method as described in relation to Fig. 1. This may include the locations of nuclei identified as CD8+ cells.
[0106] The methods of the present invention are performed on samples (such as e.g. histopathology samples) and in particular on images of such samples. The methods of the present invention are therefore ex vivo or in vitro methods, that is, the methods of the present invention are not practiced on the human body.
[0107] Applications
[0108] The above methods find applications in a variety of clinical contexts. For example determining the immunophenotype or CD8 cell presence I density in a tumour sample may provide an indication of the immune status of the tumour. This has been shown to have prognostic value for a variety of cancers. Thus, also described herein are methods of providing a prognosis for a subject that has been diagnosed as having a cancer, the method comprising identifying an immunophenotype for a tumour from the subject, and / or identifying and analysing the presence of CD8+ cells in a sample of a tumour from the subject. Tumours classified as desert may be associated with poorer prognosis than tumours classified as inflamed. Poorer prognosis may be associated with reduced relapse free survival and / or reduced overall survival.
[0109] Further, the presence of CD8 cells in the tumour environment, and the immunophenotype of a tumour being inflamed have been shown to be associated with an increased likelihood of response to immunotherapy, including in particular checkpoint inhibitors (CPI), but also any other forms of immunotherapy that leverage an existing immune function such as T-cell transfer therapy (tumourinfiltrating lymphocytes (or TIL) therapy and CAR T-cell therapy), therapeutic antibodies, cancer treatment vaccines and immune system modulators (e.g. interferons and interleukins). Thus, also described herein are methods of determining whether a subject that has been diagnosed as having a cancer is likely to benefit from treatment with an immunotherapy, the method comprising determining the immunophenotype or identifying and analysing CD8+ cells in a tumour sample from the subject. The immunotherapy may be any therapy that modulates the function of the immune system to treat cancer. Indeed, any such therapy is believed to be likely to be influenced by the presence or absence of immune cells within the tumour. In embodiments, the immunotherapy is a CPI therapy. CPI therapy includes for example treatment with an anti-CTL4 or anti-PDL1 drug. The method may further comprise classifying the subject between a group that is likely to respond to the immunotherapy therapy, and a group that is not likely to respond to the immunotherapy. For example, subjects with tumours classified as inflamed or tumours having a predetermined pattern of CD8+ cells (e.g. increase in CD8+ cells upon treatment above a threshold and / or CD8+ cell density above a threshold, as described in Dejardin et al 2024) may be more likely to benefit from immunotherapy than subjects with tumours that do not satisfy these criteria. A subject may then be classified in the group that is not likely to respond to the immunotherapy (e.g. CPI therapy) if one or more samples from the subject satisfy these criteria.
[0110] Further, also described herein are methods of treating a subject that has been diagnosed as having a cancer with a particular treatment, or identifying a subject that has been diagnosed as having a cancer as likely to benefit from a particular treatment. The method may comprise analysing a histopathology sample from the subject as described herein to obtain a predicted treatment response or prognosis associated with the sample. The method may further comprise treating the subject with a particular treatment if the histopathology sample is classified in a class that is likely to benefit from the particular treatment. For example, subjects identified as likely to respond to immunotherapy may be treated with the immunotherapy. Conversely, a subject who is identified as unlikely to respond to immunotherapy may be selected for treatment or treated with a treatment that is not an immunotherapy.
[0111] The cancer may be solid cancer. The cancer may be a squamous cancer. The cancer may be cancer type or subtype selected from, breast cancer (including e.g. Breast invasive carcinoma), central nervous system cancer (including glioblastoma multiforme and brain lower grade glioma), endocrine cancer (including adrenocortical carcinoma, thyroid carcinoma, paraganglioma & pheochromocytoma), gastrointestinal cancer (including Cholangiocarcinoma, Colon Adenocarcinoma, Esophageal carcinoma, Rectum adenocarcinoma, Stomach adenocarcinoma), gynaecologic I reproductive system cancer (including Cervical squamous cell carcinoma and endocervical adenocarcinoma, Ovarian Serous Cystadenocarcinoma, Uterine Corpus Endometrial Carcinoma, Testicular Germ Cell Tumors), Liver cancer (such as Liver Hepatocellular Carcinoma), pancreatic cancer (including Pancreatic Ductal Adenocarcinoma), head and neck cancer (including Head and Neck Squamous Cell Carcinoma, Uveal Melanoma), skin cancer (including Skin Cutaneous Melanoma), soft tissue cancer (including Sarcoma), thoracic cancer (including Lung Adenocarcinoma (LUAD), Lung Squamous Cell Carcinoma (LUSC), Mesothelioma) and urologic cancer (including Chromophobe Renal Cell Carcinoma, Clear Cell Kidney Carcinoma, Kidney renal papillary cell carcinoma, Prostate Adenocarcinoma, Urothelial Bladder Carcinoma).
[0112] Systems
[0113] Figure 3 shows an embodiment of a system according to the present disclosure. The system comprises a computing device 1 , which comprises a processor 11 and computer readable memory 12. In the embodiment shown, the computing device 1 also comprises a user interface 13, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network, to image acquisition means 3, such as a microscope or scanner, and / or to one or more data storing means 2 storing image data. The computing device may be a smartphone, tablet, personal computer, cloud computer, or other computing device. The computing device is configured to implement a method for analysing images, as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method of analysing images, as described herein. In such cases, the remote computing device may also be configured to send the result of the method of analysing images to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network such as e.g. over the public internet.
[0114] The image data acquisition means may be in wired connection with the computing device 1 , or may be able to communicate through a wireless connection as illustrated, such as e.g. through WiFi. The connection between the computing device 1 and the image data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer). The image data acquisition means 3 are configured to acquire image data from samples, for example histopathology samples or other samples of cells and / or tissues. In some embodiments, the sample may have been subject to one or more preprocessing steps such as fixation, labelling, etc. The image acquisition means 3 acquire signal associated with cells or cellular structures in the sample. The samples may be labelled with one or more histological stains or dyes that are detectable by the image data acquisition means 3. In some embodiments, it may not be necessary to actively label the target. Signal detection may be done by any method for detecting electromagnetic radiation (e.g., light) such as a method selected from far-field optical microscopy, near-field scanning optical microscopy, epi-fluorescence microscopy, confocal microscopy, two-photon microscopy, optical microscopy, and total internal reflection microscopy, where the cell or cellular structure is labelled with an electromagnetic radiation emitter. Thus, the image acquisition means 3 may comprise a far-field optical microscope, a near-field scanning optical microscope, an epi-fluorescence microscope, a confocal microscope, a two-photon microscope, an optical microscope, or a total internal reflection microscope. Conveniently, the image acquisition means may comprises an optical or confocal microscope. The image acquisition means 3 may be configured to acquire optical images. The methods of the present disclosure find particular use in the context of analysis of histopathology images obtained by optical imaging of stained samples, such as e.g. hematoxylin and eosin stained samples. Thus, any image acquisition mean that is suitable to acquire images of samples stained with a tissue stain can be used.
[0115] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims. Examples
[0116] Introduction
[0117] The present examples describe a new method of identifying CD8+ cells in a histopathology image of a tumour sample, as well as a new tumour immunophenotyping method.
[0118] Example 1 - Identification of CD8+ cells
[0119] The present inventors designed a new algorithm to identify Ki67 positive tumor cells as well as CD8 single-positive (Ki67-CD8+) and double-positive (Ki67+CD8+) cells in generic tissue types. The algorithm is able to achieve this even in samples with dense tumor nuclei, e.g., small-cell- lung-cancer (SCLC). The algorithm identifies nuclei and classifies them in one of 5 classes: Ki67+CD8+, Ki67-CD8+, Ki67+CD8-, Ki67-CD8- (non-target), and artifacts. The algorithm results are exported as Ki67+CD8- tumor cells, Ki67+CD8+ (double-positive CD8) and Ki67-CD8+ (single positive CD8) in 3 classes.
[0120] The algorithm uses a machine learning method based on Random Forest classifiers to detect and classify tumour and T-cells (quantifying tumor cells expressing Ki67 and T-cells expressing CD8 and optionally Ki67). The approach uses robust nucleus detection and supervised cell-by-cell classification algorithms with a combination of nucleus and contextual features.
[0121] The algorithm was trained based on ground truth obtained from experts, to automatically identify tumor cells and T-cells using a supervised-generation rule. The algorithm reported tumor cells expressing Ki67, T-cells expressing CD8, and double positive Ki67 / CD8 T cells in whole slide images (although cells are classified in 5 classes, as explained above, these are the informative classes reported). An example output of the algorithm is shown on Figure 3. This shows an example of a Ki67 / CD8 immunohistochemistry whole slide image (left) overlaid with the algorithmimage based analysis results (cells identified as Ki67+CD8-, Ki67+CD8+ and Ki67-CD8+ in respective colours), and the corresponding H&E (right) whole slide image, for a melanoma sample.
[0122] Methods
[0123] Tissue Staining and Digital Images. Using 2.5-pm formaldehyde-fixed paraffin-embedded (FFPE) tissue sections, A total of 66 human tissue slides were included in this study, with 48 slides used for algorithm training and 18 slides used for algorithm verification.
[0124] A total of 48 slides of Ki67CD8 stained in different tissue types was included in the algorithm training illustrated in Table 1.
[0125] Table 1. Detailed information of Samples.
[0126] A total of 48 slides containing 44 fields of view (FOVs) for Ki67 tumor and CD8+ / - were used for training. The training process consists of two main steps. The inventors used a dedicated graphical user interface (GUI) software with an image overlay indicating the nuclear positions (x, y). A trained expert manually annotated each nucleus as either tumor cells (Ki67+ / CD8-) or not tumor cells (Ki67- / CD8-). The second step involves first removing the Ki67+ / CD8- cells from the candidate nuclei identified in the initial classification step. Next, the inventors annotated the remaining cells to determine whether they are double-positive tumor cells (Ki67+ / CD8+) or single-positive CD8+ cells (Ki67- / CD8+). The different GUI was used, with label classes as follows: Ki67+ / CD8+, Ki67- / CD8+, non-target, and artifact.
[0127] Preprocessing. Unmixing of three chromogens of Purple (Ki67), Yellow (CD8), and hematoxylin (HTX, nuclei) in multicolor IHC-stained slides was performed using the color-deconvolution method (as described in Ruifrok 2001) to decompose an RGB image into its individual-constituent dyes for each biomarker (i.e. , Ki67, CD8, nucleus counterstain).
[0128] Automated cell analysis. This step focuses on detecting and isolating tumor cells with Ki67+ / CD8- from a mixture of biological cells, including lymphocytes, stromal cells, and others. The automated quantification of tumor cells employs nucleus detection and supervised classification algorithms. The classification was performed in two steps: 1 . The inventors classified each nucleus as either tumor cells (Ki67+ / CD8-) or not tumor cells (Ki67- / CD8-). 2. The second step involves first removing the Ki67+ / CD8- cells from the candidate nuclei identified in the initial classification step. Next, the remaining cells are classified to determine whether they are double-positive cells (Ki67+ / CD8+, i.e. proliferating T cells) or single-positive CD8+ cells (Ki67- / CD8+).
[0129] Automated nucleus detection. Robust automated detection of cell nuclei was performed using an algorithm detector adapted from the radial-symmetry-based method (Loy & Zelinsky, 2003) on an unmixed HTX-image channel, as described in Parvin et al. 2007. The method is illustrated on Fig. 17, and provides a method for detecting the centres of mass of objects with radial symmetry in an image. The method comprises iterative detection of radial symmetry where gradient magnitude (here the gradient magnitude of the hematoxylin channel) is projected along the radial direction according to a kernel function. The kernel function topography becomes more focused and dense at each iteration. For an initial image l(x,y), we can associated each point (x,y) with a voting direction a(x, y)=(cos(9(x,y), sin(0(x,y))) with an angle 6(x,y) that varies with the image location. An angular range A and a radial range {rmin, rmax} are set as parameters. V(x,y; rmin, rmax, A) is a voting image given these parameters. A(x,y; rmin, rmax, A) is a local voting area, defined at each image point (x,y) as A(x,y; rmin, rmax, A) :={(x±r cos<t>, y±r cos<t>)|rmin<r<rmax and 0(x,y) - A<(D< 9(x,y )+ A }. K(x,y;o, a, A) is a 2D Gaussian kernel with variance o2). The voting direction a(x, y) is calculated based on the image gradient Vl(x,y) and its magnitude ||VI(x,y) || using the equation on Fig. 17. The votes at iteration n V(x,y; rmin, rmax, An) are then computed using the equation on Fig. 17. These are updated at each iteration to refine the angular range. Finally, the centers of mass of the objects are defined by thresholding the vote image. Using a series of kernels and iterative voting, this approach was able to solve these difficulties in the detection of nuclei: 1) an inherent diversity of the appearance of epithelial cancerous nuclei, which varied from almost perfectly round to highly irregularly shaped and enlarged nuclei with coarse and marginalized chromatin and prominent nucleoli (small round structures inside the nuclei), 2) different nucleus types, such as elongated fibroblasts and lymphocyte nuclei that often appeared together with epithelial nuclei, and 3) nuclei that overlapped, clustered, or tightly clumped, which made them difficult to separate.
[0130] Feature Extraction. A set of the combination of nucleus features and contextual features based on a context-Bag-of-Words (context-BoW) method was extracted using signal from each of the 3 channels (purple (Ki67), yellow (CD8), and HTX (nuclei)). This context-BoW method uses the observation that the contextual information (Cl) of a nucleus of interest (Nol) can be described via the appearance of its neighbours. In these examples, neighbours were defined as nuclei within a distance d = 75 pixels. To capture the appearance of neighbours as contextual information, the BoW model is used to quantize their appearance, following the approach described in Nguyen et al. 2015. As explained above, nucleus centres are detected based on radial-symmetry voting. Then, nuclei are segmented using Otsu’s method in a region around the nucleus. Then, nuclear features are computed for each nucleus. The nucleus features consisted of:
[0131] 1) morphology features: area, minor and major axis lengths, radius, perimeter and solidity,
[0132] 2) appearance features: the 10th, 50th and 95th percentile values of pixel intensities (to estimate the min, mean and max values in a way that is robust against noise) and gradient magnitudes computed from hematoxylin (HTX) image channel, luminance and DAB channel (from Ki67 stained images - referred to herein as Ki67 channel) in the segmented nuclei area, and
[0133] 3) background features: appearance features computed in a ribbon of 20 pixels (~9pm) thickness around each nucleus boundary.
[0134] The context-BoW method is then implemented as follows:
[0135] 1. Extract nuclear features of all nuclei in training images. 2. Perform a clustering procedure using the K-means algorithm on all the training nuclear features to obtain C clusters.
[0136] 3. Given a nucleus of interest, assign its neighbours to the closest cluster centres using the Euclidian distance of each neighbour’s nuclear feature vector to the C centres, and compute a histogram of the cluster assignment. This histogram can be considered as contextual features of the nucleus of interest.
[0137] 4. Combine the nuclear features and contextual features into a single feature vector for training and classification. The cluster assignment of the nuclei in the neighbourhood is a mean to describe the environment around the nucleus of interest.
[0138] Note that although the context-BoW approach was used to identify contextual features in the present examples, any of the other approaches described in Nguyen et al. 2015 may be used instead, such as the context-texture method and the context-CRF method. Use of the context-BoW approach was found to result in particularly good classification accuracy in the present context of tumour / CD8+ cell classification.
[0139] Cell classification. A supervised machine learning classification based on the combination of nucleus and contextual features was performed to identify tumor cells from all detected cells. A training set was created by manually annotating regions on tissue slides. The ground-truth set consisted of cancerous cells stained with Ki67 (Ki67+CD8-), T-cells stained with CD8 (Ki67+CD8+ or Ki67-CD8+), non-tumor cells, e.g., necrosis, non-target background cells, and image / tissue artifacts. A total of 240 field-of-views (FOVs) from 48 slides was used as a training set. Five FOVs were selected per slide. Each FOV was chosen by a trained expert to ensure that these FOVs accurately represent the entire slide for training purposes. Using this training set, different classifiers, were assessed based on 5-fold cross-validation accuracy (80% training, 20% validation, randomly selected from the 240 FOVs). For our 5-class classification, a random forest classifier was chosen to train the data, and this has an accuracy above 85%. Other classifiers may be used, such as e.g. SVM, L1 -logistic regression, etc. Classification was performed in two steps: 1. A first step classifies each nucleus as either tumour cells (Ki67+ / CD8-) or not tumour cells (Ki67- ZCD8-). 2. A second step involves first removing the Ki67+ / CD8- cells from the candidate nuclei identified in the initial classification step. Next, the remaining cells are classified to determine whether they are double-positive CD8+ cells (Ki67+ / CD8+) or single-positive CD8+ cells (Ki67- / CD8+). The random forest classifier was implemented as two separate random forest classifiers, individually trained using the same training data to perform the classification in steps 1 and 2 above, respectively. In particular, a first random forest classifier was trained to classify nuclei between 3 categories: tumour cell (Ki67+ / CD8-), not tumour cell (non-target, Ki67- / CD8-, cells that are neither tumour cells nor CD8+ cells), and other (which includes artifacts and all CD8+ cells). A second random forest classifier was trained to classify nuclei between 3 categories: double positive CD8+ cells (Ki67+CD8+), single positive CD8+ cells (Ki67-CD8+), and other (artifacts). As a result of both classifications, nuclei are classified in 5 classes. Nuclei that are identified as tumour cells (Ki67+CD8-, from step 1), Ki67+CD8+ (double-positive CD8, from step 2) and Ki67-CD8+ (single positive CD8, from step 2) are identified for further analysis. This 2-step classification was found to result in particularly high accuracy for identification of these 3 disease-relevant classes of cells.
[0140] Validation. A total of 44 FOV images from 18 slides were included to verify the accuracy of the cell detection and classification of Ki67-positive (Ki67+CD8-), Ki67 / CD8-double-positive (Ki67+CD8+), and CD8-positive (Ki67-CD8+) cells. The inventors selected at least 2 FOVs per slide (more than 2 FOVs were selected from some slides). Each FOV was chosen by a trained expert to ensure that these FOVs accurately represent the entire slide for verification purposes. Using a dedicated GUI, a total of three qualified readers (two pathologists and one trained expert) provided a cell-by-cell correction on the algorithm results in each FOV image. As a result, the expert’s corrected results became a ground-truth cell count for each FOV image (i.e. the corrected verification data was included as additional ground truth to retrain the model).
[0141] Results
[0142] The inter-observer agreement among three observers for evaluating Ki67+CD8-, Ki67+CD8+, and Ki67-CD8+, respectively was quantified in order to establish a benchmark for the algorithm. As shown in Figure 4, the correlation of the agreement was R2= 0.86, 0.97, and 0.80, concordance correlation coefficient (CCC) = 0.86, 0.86, and 0.86 for Ki67+CD8-, Ki67+CD8+, and Ki67-CD8+, respectively. Figure 4A-C presents the scatter plots illustrating the inter-observer agreements among the three readers. Figure 4D presents summary statistics for these.
[0143] Next, a comparison between the results of the algorithm and the Pathologist’s Ground-Truth Counts was evaluated. A total of 44 FOV images by extracting 1-2 FOV per slide were included to verify the accuracy of the cell detection and classification for the developed image analysis algorithm. Three observers provided a cell-by-cell correction on the algorithm results in each FOV image. The results of this are shown on Figure 5, and summarized below:
[0144] For Ki67+ Tumor: R2= 0.99, concordance correlation coefficient (CCC) = 0.99, Slope = 0.998, offset = 6.369, 16,684 cells Algorithm for 16,683 cells Corrected;
[0145] - For Ki67+ / CD8+: R2= 0.99, CCC = 0.95, Slope = 0.842, offset = 2.396, 1 ,526 cells Algorithm for 1 ,936.7 cells Corrected;
[0146] - For Ki67- / CD8+: R2= 0.99, CCC = 0.98, Slope = 0.901 , offset = 0.761 , 2,272 cells algorithm for 2,558.7 cells Corrected.
[0147] The R2 and CCC are based on the comparison of the cells identified by the algorithm and the corrected cells (results of 3 pathologists averaged and compared to the algorithm results). The numbers of cells provided represent the total cell count (e.g., Ki67+ / CD8+) as determined by the algorithm (e.g., 1 ,526 cells) and the average cell count corrected by three pathologists (e.g., 1 ,936.7) in the verification data. The higher corrected number indicates that some desired cells were missed by the algorithm in the verification data. The data show that the algorithm performed extremely well at identifying these 3 types of cells (Ki67+ / CD8-, Ki67+ / CD8+, and Ki67- / CD8+).
[0148] Figure 6 shows examples of cell identification performed by the algorithm and the readers, respectively. These examples illustrate areas with a high density of tumour cells, as well as regions containing both stroma and tumour tissue.
[0149] Example 2 - development of a tumour-stroma segmentation method
[0150] For tumour nest / stroma segmentation, the inventors developed a deep learning colour-invariant segmentation approach for bright field images (immunohistochemistry and H&E). It is colourinvariant as the model uses only the hematoxylin channel (found in any IHC as the counter stain and H&E) to learn the features. A flowchart is illustrated on Fig. 8A.
[0151] IHC staining. Staining of immune cell infiltrate was perform on 2.5-3 pm FFPET sections for CD8 / Ki-67 with the anti-human CD8 antibody (Spring Biosciences, SP239, 1 :12.5) and the antibody against Ki-67 (RRID:AB_2625821 , Ventana Medical Systems, 30-9, RTU). The sections were stained as a double chromogenic assay using Ventana Discovery Ultra, Discovery XT, or Benchmark XT automated stainers (Ventana Medical Systems) with NEXES version 10.6 software. Signal detection was performed with Discovery purple for Ki67 and with / Discovery yellow (both VMS) for CD8. All sections with Hematoxylin as the nuclear counterstaining. The slides were digitally scanned at 20* magnification and image resolution of 0.465 pm / pixel with the high- throughput iScan HT (Ventana Medical Systems). Further details can be found in Zwing et al. 2020.
[0152] Dataset and annotations. A collection of 111 WSI of squamous cell carcinomas with manual immunophenotype scores were selected for annotations. The manual immunophenotype scores were obtained from CellCarta certified pathologists. Images were selected intentionally to comprise different diagnoses and immunophenotypes. A more detailed description of the dataset can be seen in Table 2.
[0153] Table 2. Characteristics of samples used for training of tumour / stroma segmentation model.
[0154] For all images, automatic tissue detection is performed. Tumour areas are manually annotated by board-certified pathologists making it the only regions of interest in the biopsy to be analysed. Necrosis and artefact areas are also annotated by pathologists to exclude these areas for any analysis. 5 fields of view (FOVs) of 500x500 pm (~ 1076x1076 pixels at 20x) were drawn per image for annotations. Six pathologists have reached a consensus on the annotation classes and follow an annotation protocol to draw 'Tumor nest', 'Stroma', 'Necrosis', 'Exclude', 'Tissue fold', 'Artifact', 'Hemorrhage'. The total number of slides were distributed among 5 pathologists and some FOVs were discarded due to missing valid tissue coverage (valid FOVs should contain only tumour tissue and have at least 50% non artefact tissue). A total of 538 FOVs were annotated with the 7 classes mentioned above. These set of FOVs will be referred to herein as “Dataset 1 ”. Additionally, a single pathologist annotated 105 WSIs from the same WSI collection (Table 2) with 14 classes ('Tumor nest', 'Stroma', 'Necrosis', 'Artifact', 'Blur', 'Calcification', 'Crush', 'Foreign object', 'Hemorrhage', 'In situ', 'Keratin', 'Normal', 'Tissue fold', 'Exclude') making a total of 518 FOVs. This is from herein referred to as “Dataset 2”. Both sets of FOVs (with annotations by multiple pathologists and with annotations by a single pathologist) were used to train two separate CNNs for comparison purposes. Further details are explained in the following section.
[0155] As explained above, the dataset used for training comprise Hematoxylin only channel (i.e. after colour deconvolution as explained above). Segmentation using greyscale versions of the RGB H&E images was also tested instead of deconvoluted H channel. This was found to be not as accurate as the use of the H channel only. This is believed to be at least in part due to high levels of variability in greyscale intensities in H&E images.
[0156] Method. As mentioned above, the inventors develop a semantic segmentation model to separate tumour nests from stroma areas in a tumour sample using only the hematoxylin channel. Deeplabv3 with ResNet50 as backbone (see aihub.qualcomm.com / models / deeplabv3_resnet50) was chosen as the CNN due to its proven feature use efficiency. Deeplabv3 is a deep convolutional neural network with atrous convolutions in an atrous spatial pyramid module described in Chen et al. 2017. The model was modified to accept as input a single channel instead of 3 and was trained to output 3 classes: Tumour nest, Stroma and Exclusion areas. From the annotated FOVs, tiles of 512x512 pixels at 10x magnification (~ 476.16x476.16 pm) were extracted to build the training datasets. For each tile, color deconvolution is applied to extract the hematoxylin channel using the following stain matrix:
[0157] Colour deconvolution was performed as described in Ruifrok & Johnston (2001 ). This method uses colour vectors of specific stains to be deconvoluted (e.g. measured on single stain specimens), which represent the optical density of the stain in each of the 3 RGB channels to obtain a normalised optical density matrix, the inverse of which is a colour deconvolution matrix from which RGB values at each pixel can be transformed to corresponding stain channel values by matrix multiplication. In the example illustrated on Fig. 8B the IHC slide images are deconvoluted between a hematoxylin channel (blue), a Ki67 channel (purple), and a CD8 channel (yellow).
[0158] In order to generate the training semantic masks for each tile with the 3 classes required for the algorithm (1) Tumour nest, (2) Stroma and (3) Exclusion; the following classes were grouped for the Stroma class: 'Stroma', 'Necrosis', 'Artifact', 'Blur', 'Calcification', 'Crush', 'Foreign object', 'Hemorrhage', 'In situ', 'Keratin', 'Normal', 'Tissue fold'.
[0159] Both datasets (Dataset 1 - 538 tiles with annotations by multiple pathologists and Dataset 2 - 518 FOVs with annotations by a single pathologist) were split into 70-10-20% for train-validation-test sets stratified by immunoscores and constrained to patientjd. The split sets for each dataset contain the same patients for comparison purposes while keeping all tiles from the same patient in the same split. Details of the train, validation and test sets are provided in Table 3A and 3B for Dataset 1 and Dataset 2 respectively.
[0160] Table 3A. Details of the sets used for training, validation and testing of the tumour / stroma segmentation algorithm. Dataset 1 - FOVs annotated by multiple pathologists.
[0161] Table 3B. Details of the sets used for training, validation and testing of the tumour / stroma segmentation algorithm. Dataset 2 - FOVs annotated by a single pathologist.
[0162] A different architecture was also tested using FCN-resnet50 (see pytorch.org / vision / main / models / generated / torchvision.models.segmentation.fcn_resnet50.html). FCN-resnet50 is a fully convolutional neural network described in Long et al. 2014. All results shown in these examples were obtained using the Deeplabv3 model.
[0163] Two Deeplabv3 models were trained with the 2 different datasets. The CNNs weights were initialized randomly, setting up 250 epochs for training. Early stopping was employed wherein training was halted if after 50 epochs no improvement was noted on the validation set. Adam optimizer was employed with a learning rate of 1 *10 -4 and weight decay of 1 *10 -4, with a batch size of 8. The following data augmentations were used: random horizontal and vertical flip, as well as contrast and brightness adjustments. Implementation of the CNNs was done in the PyTorch framework. The results for each of these are shown in Tables 4A and 4B, respectively.
[0164] Table 4A. Results for Deeplabv3-resnet50 on Dataset 1 . Table 4B. Resu ts for Deeplabv3-resnet50 on Dataset 2.
[0165] Based on the results above, all classes were scored higher in Dataset 2. Qualitative evaluation was also performed by a pathologist comparing the segmentation results from both models on randomly selected FOVs from the test set. The pathologists observed better segmentation of the tumour nests when the model was trained with annotations made a by a single pathologist. The evaluator also observed that some areas annotated as tumour nests by other pathologists were not consistent as well as the level of precision was very variable; while some were very precise, others performed more coarse annotations. Such findings confirmed the importance of having consistent ground truth to train a robust model.
[0166] Therefore, the model Deeplabv3 with ResNet50 as backbone trained with Dataset 2 was selected as the model to separate tumour nests from stroma areas.
[0167] Inference on a WSI using the tumour-stroma segmentation model. In order to segment tumour nest from stroma areas in a Ki67-CD8 WSI, the tissue area, the tumour area annotation made by a pathologist and if available the necrosis and exclusion annotations are taken into account to generate tiles of 512x512 pixels at 10x magnification. An overlapping area of 25% among tiles is taken into account to avoid harsh delineations. For each tile, the positional coordinates (x,y) from the reference WSI are saved. Necrosis and exclusion areas are not analyzed by any algorithm.
[0168] The segmentation model takes the RGB tiles, performs color deconvolution of the hematoxylin channel and takes them as input to generate semantic masks with 3 classes. A 10x magnification mask is constructed based on the outputs and coordinates of each tile. The final mask is colored for output visualization purposes with a color dictionary defined by the user. In addition, the tumour nest mask is converted into annotation coordinates for further evaluation.
[0169] The WSI segmentation mask is displayed to the user together with the tumour nest annotations extracted from the segmentation mask. If the pathologist considers that some tumour nests are not properly identified, he / she will modify the corresponding annotation area and will flag the slide. In this way the immunophenotyping algorithm will know which input to take into account (segmentation mask or annotations).
[0170] Example 3 - Immunophenotyping
[0171] In 2023, Li et al. proposed a characterisation of tumours as belonging to one of three immune phenotypes based on the density and spatial distribution of CD8+ cells in and around a tumour: desert, excluded, and inflamed. They showed that manual immunophenotyping of tumors is predictive for clinical benefit by checkpoint inhibition-targeted therapy in two independent cohorts of two distinct indications. This approach is illustrated on Fig. 1 of Li et al. 2023, and comprised manually grading the density of CD8+ cells within tumour area (defined as viable tumour cells with associated stroma, excluding areas of necrosis or intraluminal aggregates of CD8+ cells) as:
[0172] “0” (single dispersed cells with a density too sparse to allow identification of a distinct pattern) - classified as “Desert”;
[0173] 1+, 2+ or 3+ (an arbitrary scale representing increasing CD8+ densities) with distribution limited to the CK- stroma - classified as “Excluded”; and
[0174] 1+, 2+ or 3+ (an arbitrary scale representing increasing CD8+ densities) co-localised with CK+ tumor cells either as a diffuse infiltrate with or without involvement of CK- stroma, or with a predominantly stromal distribution of CD8+ cells with “spill-over” into CK+ tumor cell aggregates - classified as “Inflamed”.
[0175] To address intratumoral heterogeneity of density and pattern of the infiltrate, the authors further implemented a 20% cut-off; patterns occupying less than 20% of the tumor area were not considered for categorization. Cases showing an inflamed phenotype in >20% of the tumor area were labeled “Inflamed”, independent of the pattern(s) observed in the remaining areas. Implementation of the cutoff was based on a subjective estimate by the inspecting pathologist, as was implementation of the judgement of distribution and density. This is summarised in Table 1 below.
[0176] Table 1. Manual immunophenotyping scheme from Li et al. 2023.
[0177] Li et al. 2023 also proposed alternative automated approaches to attempt to capture these phenotypes, namely the use of (i) the ratio of the areas segmented as CD8+ pixels relative to the areas of CK+ pixels, with the lowest 20% of samples in a cohort being classified as “desert”, the next 40% as “excluded” and the highest 40% as “inflamed” (density based method); (ii) a support vector machine classifier trained to classify between these 3 classes based on a 20-element vector for each slide containing the number of tiles in each of 10 categories defined by respective ranges of the density of CD8+ pixels in the CK+ area and the CK- area of the tile, respectively (binned CD8 density); (iii) a random forest classifier trained to classify between these 3 classes based on 50 spatial features that capture co-localisation of CK and CD8 in each slide of distribution of CK or CD8 in each slide, based on the ratios of the number of CK+ to CD8+ pixels in each tile of the slide (random forest classifier with spatial statistics); and (iv) a random forest classifier trained to classify between these 3 classes based on the concatenation of the feature vector in (ii) and the feature vector in (iii) (random forest classifier with spatial statistics and binned CD8 T-cell density). All tilebased approaches used a fixed tile size of 280x280 pixels at 0.46 pm per pixel. The first and fourth of these could differentiate between the 3 types of immunophenotypes. However, the binned CD8 density features and the spatial statistics both individually or together failed to cluster into the expected 3 phenotype categories. A supervised classification approach was necessary to be able to use these features. Of these, the random forest classifier with spatial statistics and binned CD8 T-cell density was shown to provide a classification that is associated with response to checkpoint inhibition. However, none of these approaches appropriately captures the manual approach (and the classification features used did not in themselves enable the authors to see distinctions between the 3 phenotypes - leading to low explainability of the best performing machine learning model). As a result, the manual approach remains the standard despite the severe limitations associated with the subjective and labour intensive nature of the process.
[0178] Therefore, the present inventors set out to develop a new method for tumour immunophenotyping that better captures the manual approach, without suffering from the same limitations. The goal of the work was to compute reproducible immunophenotype scores identifying samples as Inflamed, Excluded and Desert, as well as providing diagnostic report metrics for the tumour nest and stroma compartments including e.g. densities, areas, ratios, CD8+ infiltration heatmaps, and other diagnostic plots. The general approach is illustrated on Fig. 7 and comprises: parallel application of (a) a tumour nest / stroma segmentation algorithm, and (b) a CD8+ detection algorithm. Nevertheless, the immunophenotyping algorithm could score a tumour sample if (1) the cell centroids of the CD8+ cells are provided, eitherfrom an algorithm output or by manual annotations; and (2) the areas of tumour nest are defined either by masks or annotations coming from an algorithm or delineated by expert pathologists.
[0179] Data All images analysed were RGB images of IHC slides stained with Ki67 and CD8 antibodies, and Hematoxylin counterstain. Two early phase clinical studies with a total of 58 WSI were selected for algorithm development and a set of 213 WSI were selected for validation. The samples are stained with Ki67-CD8 duplex as described in section ‘IHC staining’ in Example 2. The selection of samples comprises different indications and immunophenotype distribution as shown in Table 5.
[0180] Table 5. Description of the samples used for the immunophenotyping algorithm.
[0181] Every biopsy tumour sample was annotated by board certified pathologists with the Tumour area (area of interest for analysis) as well as Necrosis and Exclusion areas. Exclusion areas comprises either tissue artefacts like tissue folds, out of focus areas or any other area that the pathologist considers not suitable for evaluation. Necrosis and Exclusion areas are excluded from any analysis.
[0182] Study 1 and the validation set samples have immunophenotype scores performed manually by CellCarta pathologists.
[0183] Tumour / stroma segmentation
[0184] For Study 1 and 2, Roche board-certified pathologists have annotated the tumour nest areas in all samples. Annotations were performed in order to have a valid ground truth for algorithm development. On the contrary, for the validation set, the areas of tumour nest were separated from the stroma using the algorithm described in Example 2. A pathologist evaluated the outputs of the segmentation algorithm and approved the results.
[0185] CD8 detection
[0186] CD8 detection is performed using the algorithm described in Example 1. All CD8+ cells (i.e. both the proliferating and non-proliferating class) were considered together (i.e. no distinction was made between the two).
[0187] Immunophenotyping algorithm
[0188] Fig. 9 illustrates the general principles of the immunophenotyping algorithm. This analyses the level of immune CD8 cell infiltration in the segmented regions (either provided by an algorithm output or by manual annotations) identified as stroma and those identified as tumour nest separately to categorize the image sample as Inflamed, Excluded or Desert. In a first step, the WSI is subdivided into tiles and for each tile, the number of CD8+ cells (both 'MKI67+ CD8A+', and 'MKI67- CD8A+') in the tumour nest area and the number of CD8+ cells in the stroma area are evaluated. The tile area from both regions (compartments) is then classified as cold, medium or hot based on predetermined ranges of number of CD8+ cells. The predetermined ranges were identified fora specific tile size (i.e. coupling tile size and identification of a corresponding threshold for cold, medium or hot) and assuming the tile contains 100% a single region (tumour nest or stroma) so as to reproduce the qualitative assignment of cold / medium / hot regions that are performed by pathologists when manually classifying immunophenotypes. In the illustrated example, all tiles with 0 to 3 CD8+ cells were classified as “cold”, all tiles with 4 to 7 CD8+ cells were classified as “medium”, and all tiles with more than 7 CD8+ cells were classified as hot. These thresholds provided the best replication of pathologist classifications when in combination with tile sizes of 119 x 119 pm in squamous cancers. Finally, a score is calculated based on the percentage of hot areas per compartment (tumour nest and stroma, respectively), as illustrated on Fig. 10. Specifically, samples with at least 20% of “hot” tumour nest area were classified as “Inflamed”. Samples with less than 20% of “hot” tumour nest area and less than 20% of “hot” stroma area were classified as “Desert”. Samples with less than 20% of “hot” tumour nest area and at least 20% of “hot” stroma area were classified as “Excluded”. If a tile contains areas from both compartments (tumor nest and stroma), as shown on Fig. 11 , a normalization is considered between the area and CD8 threshold to determine if such area will be hot or not. In other words, a tile which stroma area is 50% of the tile will be classified as “hot” for stroma if such area comprises at least 4 CD8+ T cells. Similarly a tile which tumour nest area is 50% will be classified as “hot” for tumour nest if such area comprises at least 4 CD8+ T cells. Evaluation could also be translated into cell densities with a tile threshold for “hot” being at least 564.55 cells / mm2(see Formula 2).
[0189] # CDS cells 8 hot density = — ; - - - 7 « . — — (2) tile area (mm2) (0.119042)
[0190] Note that only the hot vs not hot (i.e. cold or medium) classification is used to determine the immunophenotype. However, statistics about the areas classified as hot, medium and cold may be provided for a more granular view of the sample.
[0191] The tile size and CD8 thresholds mentioned previously and shown in Figure 9 were chosen based on the evaluation of different tile sizes and number of CD8 cells. Two pathologists assessed 3 different tile sizes (0.05952 mm, 0.11904 mm and 0.23808 mm) on tumour biopsy samples from Study 1 to define areas with high immune CD8 infiltration (hot areas). A tile size of 0.11904 mm x 0.11904 mm (~ 119.04 pm x 119.04 pm) was chosen as the standard size for evaluation due to the variability of small tissue present in some samples. The number of 8 CD8+ cells was selected from the evaluation of tumour nest areas in Inflamed cases to define high infiltration. At last, 20% of highly infiltrated areas was defined as the final cut-off for decision making between phenotypes as used in several studies (e.g. Li et al. 2024).
[0192] Study 1 and 2 were scored automatically using as input: (1) the manual annotations performed by the pathologists for the tumour nests creating segmentation masks of tumour nest and stroma areas, and (2) the CD8+ cell position coordinates (MKI67+ CD8A+', and 'MKI67- CD8A+'). The segmentation masks take into account only the tissue within the “Tumour” annotation area and exclude all areas annotated as “Necrosis” and / or “Exclude”. From Study 1 (18 WSI), the immunophenotype scores between the manual (CellCarta scores) and algorithm approaches were different in more than 50% of the samples. These 10 samples were scored manually by 2 Roche pathologists to compare the results (see Table 6). In 7 / 10 samples, both of the pathologists agreed with the Algorithm scores, while in 1 / 10 at least one of the pathologists agreed with the algorithm. In 2 / 10 cases, both pathologists agreed with the manual CellCarta scores and these were further evaluated. Pathologists looked at the biopsies from patient 6 together with the algorithm statistics. The amount of high infiltrated area in the tumour nest was near but less than the 20% cut-off, which was the reason for the algorithm to score the samples as Desert instead of Inflamed. Calculating visually the total area percentage is subjective and will always lead to different scores.
[0193] Table 6. Immunophenotype score evaluation from Study 1.
[0194] Cohen’s Kappa scores (Cohen 1960) were computed among 2 raters as well as Fleiss Kappa scores (Fleiss et al. 1981) to measure the agreement among 3 or more raters (Figure 18). The highest agreements are reached among both Pathologists and the Algorithm (moderate to substantial agreement) while the worse are those among the pathologist and CellCarta scores (less than chance, poor and slight agreement).
[0195] Immunophenotype visual scoring for the 20 patients with paired biopsies in Study 2 was performed by 2 pathologists to compare the outputs between theirs and the Algorithm scores. Cohen and Fleiss kappa scores were computed to measure the reliability of agreement among the immunophenotypes (see Figure 19). Agreement ~ 0.7 was reached among all of them. The agreement scores in both Study 1 and 2, between the Algorithm and Pathologist 1 and 2, range in the Moderate category allowing the inventors to conclude that the proposed algorithm generates confident and reproducible results.
[0196] The approach was additionally evaluated in a collection of 213 images (validation set in Table 5) stained with Ki67CD8 yellow / purple. The immunophenotype algorithm was run and the results were compared against the visual (manual) scores performed by CellCarta. The algorithm scored the same inmmunophenotypes for 122 / 213 (57.28%) while 91 images (42.72%) had different computed scores. Six samples, 2 for each algorithm’s immunophenotype category, were randomly chosen from the entire validation set and evaluated by two separate pathologists (Table 7). These results showed high discordance between the two pathologists, with one agreeing with the algorithm in 5 out of 6 cases, and the other one disagreeing with the algorithm in 5 out of 6 cases. This indicates that there is no “ground truth” classification against which the algorithm should be compared to, and a pathologist would not always be certain of the score to report. Particularly, this is the case when the distribution of CD8+ cells in both compartments (tumor nest and stroma) is heterogeneous creating high uncertainty and difficulty for visual evaluation which is problematic in clinical trials. Hence, the inventor’s proposed algorithm tackles the need for a rigorous quantitative and reproducible analysis.
[0197] Table 7. Immunophenotype score evaluation of a subset of 6 samples from the validation set.
[0198] Heatmaps of CD8 cells in tumour nest and stroma
[0199] In addition to the immunophenotype score, the algorithm creates two separated heatmaps to visualize the amount of immune CD8 infiltration in the tumour nest and stroma areas (compartments) within the tumour tissue (Tumour area annotated by pathologists). This provides an alternative and complementary way to analyze a whole slide image describing the local behaviour of CD8 cells in the different compartments.
[0200] Additionally, the heatmaps produced were used as tools to identify the size of tile that is necessary to capture the presence of CD8+ enriched regions from which the immunophenotypes are derived. In general, heatmaps show how a phenomenon is clustered or varies over space within its own distribution. For example, in the case of cells in a tissue, it shows the local cell distribution in a certain area with colors ranging from cold (blue - min number of cells) to hot (red - max number of cells). Heatmaps are usually normalized from 0 to 1 where 1 represents the maximum number of cells in a region. If there are clusters of cells and very few cells in most of the tissue, the heatmap should show red in the clustered areas whereas in the rest little or no signal will be seen. However, if the tile size (local area) over which the number of cells is calculated is too big, then small local clusters can fail to be detected. This is illustrated with an example on Fig. 12 which shows heatmaps of CD8 cell density for the same whole slide image, with densities (and therefore heatmap colours) obtained using increasing tile sizes from left to right. The final tile size that was used for this cancer type (squamous cell cancer) for the immunophenotyping algorithm is the one illustrated in the middle image. Using the CD8 threshold (8+ cells) defined by the expert pathologists to consider an area highly infiltrated (described in subsection Immunophenotyping algorithm), heatmaps per compartment are created to allow visual comparison across patients. The heatmaps are built analyzing local areas defined by the tile size (0.119 mm x 0.119 mm = 0.0141705 mm2) and calculating the number of CD8+ cells in such an area. If the total number of CD8+ cells exceeds the CD8 threshold to consider a hot area (8+ cells), such an area adopts the threshold, otherwise, the area keeps the numberof cells found. A mask with the same size as the WSI is built with such information applying a Gaussian (smoothing) filter and normalizing the outputs between 0 and 1. The total area is calculated in mm2and reported in ranges from 0.1 to 1 .0.
[0201] Figure 13A and 13B show an example WSI scored as Inflamed with the CD8 infiltration heatmaps created by the immunophenotyping algorithm.
[0202] Statistical metrics and histograms in tumour nest and stroma
[0203] Additional diagnostic metrics and plots provided from outputs of the immunophenotyping algorithm include, per compartment, the total area reported in mm2and percentages, cold and hot area percentages, total amount of CD8+ cells, CD8 densities (cells / mm2), number of hotspots and histograms of the number of tiles containing CD8+ cells. The histograms are divided into 10 bins (# CD8 cells) and show as well the arbitrary thresholds defined by the pathologists to quantify tiles as “hot”, “medium” or “cold”. The colour of each bin of the histogram matches the colour used in the heatmaps. Examples are shown on Fig. 14, Fig. 15 and Fig. 16 for samples classified as “Desert”, “Excluded” and “Inflamed” respectively according to the immunophenotyping algorithm..
[0204] Conclusion
[0205] The present examples describe and demonstrate the use of a new method for identification of CD8 cells and tumour cells in a histopathology image, a colour-invariant segmentation algorithm for separating the tumour nests from stroma, as well as a method for immunophenotyping of tumours from histopathology images that is based on a manual immunophenotyping approach that has become standard in the art but alleviates problems of subjectivity and processivity associated with the manual approach.
[0206] References
[0207] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.
[0208] Xiao Li, et al. Automated tumor immunophenotyping predicts clinical benefit from anti-PD-L1 immunotherapy, bioarxiv. April 07, 2023. doi.org / 10.1101 / 2023.04.03.535467 Li, Xiao, Jeffrey Eastham, Jennifer M. Giltnane, Wei Zou, Andries Zijlstra, Evgeniy Tabatsky, Romain Banchereau et al. "Automated tumor immunophenotyping predicts clinical benefit from anti-PD-L1 immunotherapy." The Journal of Pathology 263, no. 2 (2024): 190-202.
[0209] Zwing, Natalie, et al. "Analysis of spatial organization of suppressive myeloid cells and effector T cells in colorectal cancer — a potential tool for discovering prognostic biomarkers in clinical research." Frontiers in Immunology 11 (2020): 550250.
[0210] Liang-Chieh Chen George Papandreou Florian Schroff Hartwig Adam. Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv:1706.05587v3 [cs.CV] 5 Dec 2017
[0211] Ruifrok AC, Johnston DA. Quantification of histological staining by color deconvolution. Anal Quant Cytol Histol 23: 291-299, 2001.
[0212] Jonathan Long, Evan Shelhamer, Trevor Darrell. Fully Convolutional Networks for Semantic Segmentation. arXiv:1411.4038v2 [cs.CV] 8 Mar 2015
[0213] David Dejardin, et al. A Composite Decision Rule of CD8|D T-cell Density in Tumor Biopsies Predicts Efficacy in Early-stage, Immunotherapy Trials. Clin Cancer Res 2024;30:877-82.
[0214] B. Parvin, Q. Yang, J. Han, H. Chang, B. Rydberg and M. H. Barcellos-Hoff, "Iterative Voting for Inference of Structural Saliency and Characterization of Subcellular Events," in IEEE Transactions on Image Processing, vol. 16, no. 3, pp. 615-623, March 2007.
[0215] Nguyen, K., Bredno, J.., Knowles, D., “Using Contextual Information to Classify Nuclei in Histology Images,” Proc. Int. Symp. Biomed. Imaging, 995-998 (2015).
[0216] Loy G, Zelinsky A (2003) Fast radial symmetry for detecting points of interest. IEEE Trans Pattern Anal Machine Intell 25: 959-973.
[0217] Cohen, Jacob. "A coefficient of agreement for nominal scales." Educational and psychological measurement 20, no. 1 (1960): 37-46.
[0218] Fleiss, Joseph L., Bruce Levin, and Myunghee Cho Paik. "The measurement of interrater agreement." Statistical methods for rates and proportions 2, no. 212-236 (1981): 22-23.
[0219] The specific embodiments described herein are offered by way of example, not by way of limitation. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Any sub-titles herein are included for convenience only, and are not to be construed as limiting the disclosure in any way. The methods of any embodiments described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described above. Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described. Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments of the invention may be readily combined, without departing from the scope or spirit of the invention. It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. Other aspects and embodiments of the invention provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of’ or ’’consisting essentially of’, unless the context dictates otherwise. The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.
Claims
44Claims:1 . A computer-implemented method of identifying an immunophenotype for a tumour, the method comprising: receiving one or more whole slide images of a sample of the tumour that has been stained using at least a CD8 label; identifying, using a deep learning model, areas of the one or more whole slide images that show tumour nest tissue and areas that show stromal tissue; identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label; classifying each area of a plurality of tiles of the one or more whole slide images that show tumour nest tissue and each area of the plurality of tiles that show stromal tissue between at least a first class and a second class, wherein the tiles have a predetermined tile size and the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, and the second class comprises areas with a number of CD8 cells below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size; and classifying the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype based on the proportion of the area of the tiles that show stroma tissue classified in the first class and the proportion of the area of the tiles that show tumour nest tissue classified in the first class.
2. The method of claim 1 , wherein classifying the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype based on the proportion of the area of the tiles that show stroma tissue classified in the first class and the proportion of the area of the tiles that show tumour nest tissue classified in the first class comprises: classifying the one or more whole slide images as inflamed when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is at or above a second predetermined threshold; classifying the one or more whole slide images as desert when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is below the second predetermined threshold and the proportion of the area of the tiles that show stromal tissue classified in the first class is below a third predetermined threshold; and classifying the one or more whole slide images as excluded when the proportion of the area of the tiles that show tumour nest tissue classified in the first class is below the second predetermined threshold and the proportion of the area of the tiles that show stromal tissue classified in the first class is at or above the third predetermined threshold; optionally wherein the second and third45 predetermined thresholds are the same, and / or wherein the second and third predetermined thresholds are 20%.
3. The method of any preceding claim, wherein classifying each area of a plurality of tiles of the one or more whole slide images between at least a first class and a second class comprises classifying the areas of the plurality of tiles between a first class, a second class and a third class, wherein the first class comprises areas with a number of CD8 cells at or above a first predetermined threshold associated with the predetermined tile size scaled by the relative size of the area and the predetermined tile size, the second class comprises areas with a number of CD8 cells within a first predetermined range below the first predetermined threshold scaled by the relative size of the area and the predetermined tile size, and the third class comprises areas with a number of CD8 cells within a second predetermined range below the first predetermined range scaled by the relative size of the area and the predetermined tile size.
4. The method of any preceding claim, wherein the predetermined tile size, first predetermined threshold, and optional first and second predetermined ranges have been identified by selecting a plurality of sets of candidate values of the predetermined tile size, first predetermined threshold, and optional first and second predetermined ranges, and:(a) for each candidate set of candidate values, displaying a heatmap of tile area classifications for one or more samples to a user; and selecting a set of candidate values that is identified by the user as accurately capturing areas of immune infiltration in the one or more samples; and / or(b) for each candidate set of candidate values, classifying a plurality of samples between whole slide images as one of a desert, excluded or inflamed immunophenotype using the method of any preceding claim; and selecting a set of candidate values that maximises the number of the plurality of samples that have the same classification as corresponding ground truth classification, optionally wherein ground truth classifications are obtained from annotations by multiple pathologists.
5. The method of any preceding claim, wherein the predetermined tile size is between 100 pm x 100 pm and 200 pm x 200 pm, optionally 119.04 pm x 119.04 pm and / or the first predetermined threshold is 8, and / or the first predetermined range is 4 to 7 and the second predetermined range is 0 to 3.
6. The method of any preceding claim, wherein the tumour is a solid tumour, optionally a squamous cell cancer and / or an anal carcinoma, cervical carcinoma, esophagus carcinoma, head and neck carcinoma, Thymoma or thymic carcinoma, melanoma, or non small cell lung cancer, and / or wherein the sample is a formaldehyde-fixed paraffin-embedded (FFPE) tissue section, and / or wherein the whole slide images are immunohistochemistry images.
467. The method of any preceding claim, wherein: (i) the one or more whole slide images are images of a sample of the tumour that has been stained using a CD8 label, a proliferative cell label, optionally Ki67, and a nuclear label, optionally hematoxylin; and / or (ii) the one or more whole slide images are images of a sample of the tumour that has been stained using a plurality of stains associated with respective colours, and the method comprises performing colour deconvolution prior to identifying CD8 cells and areas of the images that show tumour nest tissue or stromal tissue, thereby obtaining single channel images each corresponding to a respective stain, optionally wherein the plurality of stains comprise hematoxylin, and identifying, using a deep learning model, areas of the one or more whole slide images that show tumour nest tissue and areas that show stromal tissue is performed using single channel images corresponding to the hematoxylin stain.
8. The method of any preceding claim, wherein classifying each area of a plurality of tiles of the one or more whole slide images between at least a first class and a second class, comprises: obtaining an adjusted first predetermined threshold for a tile area based on the proportion of the area of the tile that shows tumour tissue, and classifying the tile area showing tumour tissue as a tile area that shows tumour tissue in the first class when the number of CD8 cells in the tumour tissue area in the tile is at or above the adjusted first predetermined threshold, wherein the adjusted first predetermined threshold corresponds to the first predetermined threshold scaled by the relative size of the tile area and the predetermined tile size, and obtaining an adjusted first predetermined threshold for the tile area based on the proportion of the area of the tile that shows stromal tissue, and classifying the tile area showing stromal tissue as a tile area that shows stromal tissue in the first class when the number of CD8 cells in the stromal tissue area in the tile is at or above the adjusted first predetermined threshold, wherein the adjusted first predetermined threshold corresponds to the first predetermined threshold scaled by the relative size of the tile area and the predetermined tile size.
9. The method of any preceding claim, wherein the deep learning model is a deep learning model that has been trained to take as input a whole slide image and produce as output a pixel classification for each of a plurality of tiles of the whole slide image between a first class associated with tumour nest tissue, a second class associated with stromal tissue, and a third class associated with signal other than tumour nest tissue and stromal tissue, optionally wherein the method further comprises training the deep learning model using a plurality of training whole slide images and corresponding ground truth annotations indicating areas associated with tumour nest tissue and stromal tissue.
10. The method of any preceding claim, further comprising:(i) producing a heatmap that represents, for each area of at least one of the one or more whole slide images corresponding to one of the plurality of tiles, the classification associated with the tile areas, individually for areas that show tumour nest and / or for areas that show stromal tissue, and / or(ii) producing a histogram of the number of tiles showing tumour nest tissue classified in each of the first, second and optionally third class and / or producing a histogram of the number of tiles showing stromal tissue classified in each of the first, second and optionally third class, and / or(iii) producing a histogram of the number of tiles showing tumour tissue comprising a number of CD8 cells within each of a plurality of bins, and / or(iv) producing a histogram of the number of tiles showing stromal tissue comprising a number of CD8 cells within each of a plurality of bins, and / or(v) identifying a prognosis or treatment response for the tumour based on the classification of the one or more whole slide images as one of a desert, excluded or inflamed immunophenotype.
11. The method of any preceding claim, wherein identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label comprises: applying a nucleus detection algorithm, quantifying, for each detected nucleus, a set of features comprising a plurality of nucleus morphology features, a plurality of nucleus appearance features, a plurality of nucleus background features, and a plurality of nucleus context features, and classifying each nucleus between a plurality of classes comprising at least a CD8- class and a CD8+ class using a machine learning model trained to classify a nucleus between the plurality of classes using as input the set of features for the nucleus.
12. The method of claim 11 , wherein the machine learning model comprises a support vector machine, a logistic regression model or a random forest model, and / or wherein the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model has been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label; and / or wherein the one or more whole slide images are images of a sample of the tumour that has been stained using a plurality of stains associated with respective colours and comprising a nuclear stain, and identifying CD8 cells in the one or more whole slide images using a signal associated with the CD8 label comprises performing colour deconvolution prior to applying the nucleus detection algorithm thereby obtaining single channel images each corresponding to a respective stain, wherein the nucleus detection algorithm is applied to the single channel image corresponding to the nuclear stain.
13. The method of claim 11 or claim 12, wherein the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model has been trained to classify each nucleus between a plurality of classes comprising at least a CD8- class, a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label using a two-step process comprising: a first step of classifying each nucleus between a plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label; and a second step of classifying each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label; and / or wherein the images are images of a sample of the tumour that has been stained using a CD8 label and a proliferative cell label, optionally Ki67, and the machine learning model, and the machine learning model comprises: a first machine learning model configured to classify each nucleus between a plurality of classes comprising a CD8- class that is positive for the proliferative cell label, and a CD8- class that is negative for the proliferative cell label, and a second machine learning model configured to of classify each nucleus that is not in the CD8- class that is positive for the proliferative cell label between a plurality of classes comprising a CD8+ class that is positive for the proliferative cell label, and a CD8+ class that is negative for the proliferative cell label; and / or wherein the machine learning model comprises a first machine learning model trained to classify nuclei between a first plurality of classes, and a second machine learning model trained to classify nuclei between a second plurality of classes, wherein the method comprises classifying each nucleus using the first machine learning model, selecting nuclei classified in a subset of the first plurality of classes, and classifying the selected nuclei between the second plurality of classes using the second machine learning model.
14. The method of any of claims 11 to 13, wherein, for a nucleus of interest:(i) the plurality of nucleus morphology features comprise features selected from: the area of the nucleus, the length of the major axis of the nucleus, the length of the minor axis of the nucleus, the radius of the nucleus, the perimeter of the nucleus, and the solidity of the nucleus, and / or(ii) the plurality of nucleus appearance features comprise features selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or more channels after colour deconvolution, the gradient magnitude in one or more channels after colour deconvolution, and / or(iii) the plurality of nucleus background features comprise features computed in a ribbon of a predetermined number of pixels around the nucleus boundary, optionally wherein the features are selected from: the 10th, 50th and / or 95th percentile values of pixel intensities in one or morechannels after colour deconvolution, and the gradient magnitude in one or more channels after colour deconvolution; and / or(iv) the plurality of nucleus context features comprise the number or proportion of nuclei within a predetermined distance from the nucleus of interest that are assigned to each of a plurality of clusters identified by applying a clustering algorithm to the plurality of nucleus morphology features, the plurality of nucleus appearance features and the plurality of nucleus background features for each of a plurality of nuclei in one or more training whole slide images.
15. A system comprising: at least one processor; and at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to implement the method of any preceding claim.
Citation Information
Patent Citations
Method of generating a metric to quantitatively represent an effect of a treatment
WO2022017666A1
Tumor immunophenotyping based on spatial distribution analysis
US20240104948A1
Cell localization signature and immunotherapy
WO2022047412A1