Ultrafast evolving loci as dynamic predictive biomarkers
By measuring microsatellite diversity at coding homopolymer loci using SDI and JSD, the method addresses the challenge of predicting ICB response and CRC prognosis in MMRd patients, offering precise mutation rate estimation and tumor classification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UCL BUSINESS LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-30
AI Technical Summary
Current methods fail to accurately predict the response of mismatch repair deficient (MMRd) colorectal cancer patients to immune checkpoint blockade (ICB) therapy and provide insights into genomic mutation dynamics during tumor progression, lacking a companion test to account for spatial and temporal variation in subclonal mutation rates.
Utilizing nucleic acid sequencing data to measure microsatellite diversity at specific coding homopolymer loci, employing metrics like Shannon Diversity Index (SDI) and Jensen-Shannon Distance (JSD), to estimate mutation rates and predict ICB response, CRC prognosis, and classify MSH6 proficiency/deficiency.
Provides a sensitive proxy for genomic mutation rates and intra-tumor heterogeneity, enabling accurate prediction of ICB response and CRC prognosis, distinguishing between MSH6 proficient and deficient samples, and offering insights into tumor evolution.
Smart Images

Figure IMGF000010_0001 
Figure IMGF000011_0001 
Figure IMGF000011_0002
Abstract
Description
[0001] 008613259
[0002] Ultrafast Evolving Loci As Dynamic Predictive Biomarkers
[0003] This application claims priority from GB 2501140.4 filed 27 January 2025 the contents and elements of which are herein incorporated by reference for all purposes.
[0004] Technical Field
[0005] The present invention relates to methods for cancer prognosis, drug response prediction and particularly, although not exclusively, to methods in which diversity at particular loci is used to classify colorectal cancer patients.
[0006] Background
[0007] Loss of the DNA mismatch repair (MMR) machinery occurs in -15% of colorectal cancers (CRCs). Subsequent inability to repair replication - associated errors in these tumours drives accumulation of single - nucleotide mismatches and frameshift variants primarily occurring within short repetitive genomic sequences known as microsatellites, leading to microsatellite instability (MSI). Although mismatch repair deficient (MMRd) CRCs are immunogenic, only 50% of patients respond to immune checkpoint blockade (ICB), with no method or clinical marker (s) to identify those who benefit up front. This complicates translation to the neoadjuvant space where patients might opt to prioritise resection over systemic treatment.
[0008] In human cells, DNA MMR is performed by protein complexes consisting of MutL homolog 1 (MLH1) and PMS1 homolog 2 (PMS2), known as MutLa, and MutS homolog 2 (MSH2 ) and MutS homolog 6 (MSH6), known as MutSa (Kunkel et al., 2005). Alternatively, MSH2 can pair with MSH3 in a complex called MutSp. MutSa and MutSp each function as DNA mismatch detection modules with008613259
[0009] partially overlapping specificities, whereas MutL (MLH1 / PMS2) executes mismatch repair.
[0010] Previous research has identified a mechanism, termed adaptive mutagenesis, whereby hypermutable homopolymer tracts in the minor MMR genes MSH3 and MSH6 dynamically move in and out of reading frame to control DNA mismatch repair gene expression in response to immune selection (Kayhanian et al., 2024).
[0011] Secondary loss of these genes in MMR lesions with prior MLH1 / PMS2 truncal loss increases mutation rate and burden, as well as shifts in mutation bias with a decrease in T> C transitions in favour of C> T transitions and C> A transversions.
[0012] Even in the absence of purifying selection, exons incur fewer mutations than expected in several tumour types due to the MMR machinery prioritizing mutation repair of coding over noncoding regions. Enrichment for histone mark H3 Lysine 36 tri methylation (H3K36me3 ) drives enhanced MMR activity in exonic regions. Multi - region whole exome sequencing (WES) confirms that coding microsatellites in MLH1 - def icient / MSH6 - def icient clones undergo faster erosion compared to the same sites in MLH1 - def icient / MSH6 -prof icient ancestor lineages.
[0013] W02024 / 105220 relates to microsatellite markers and their uses, in particular for determining the microsatellite status of a tumour. WO2023 / 052795 relates to methods for evaluating levels of microsatellite instability in a sample and evaluating the biological significance of sequence variations identified in a sample during sequencing. Kondelin et al., " Comprehensive Evaluation of Protein Coding Mononucleotide Microsatellites in Microsatellite-Unstable Colorectal Cancer." Cancer research vol. 77, 15 (2017): 4078 - 4088.
[0014] doi: 10. 1158 / 0008 - 5472. CAN- 17 - 0682 relates to a statistical008613259
[0015] model for the somatic background indel mutation rate of microsatellites to assess mutation significance. Maley et al., " Genetic clonal diversity predicts progression to esophageal adenocarcinoma." Nat Genet 38, 468-473 (2006). https: / / doi. org / 10.1038 / ngl768 show that clonal diversity measures adapted from ecology and evolution can predict progression to adenocarcinoma in the premalignant condition known as Barrett' s esophagus, even when controlling for established genetic risk factors.
[0016] Tumour mutation burden (TMB) is currently used in the NHS as a biomarker in predicting ICB response. However, this only gives insight into historic mutation accumulation, rather than providing estimates of evolutionary dynamics within tumour regions at present. However, mutation rates vary spatially and temporally during tumour progression. Therefore, despite these advances, there is a need for a companion test to aid with MMRd CRC diagnosis and ICB response prediction, accounting for variation in subclonal mutation rates.
[0017] The present invention has been devised in light of the above considerations.
[0018] Summary of the Invention
[0019] The present inventors have identified close to 300 coding homopolymers (6 - 11 bp in length) that show ultrarapid evolution in the context of MSH3 / 6 status. As described in detail herein, it was found that population diversity metrics (e.g. Shannon diversity index (SDI) and Jensen - Shannon distance (JSD) ) across these loci serve as a sensitive proxy for genomic mutation rate and thereby provide a convenient measure of intra - tumour heterogeneity (ITH) and sub - clonal complexity. Moreover, diversity measures of subsets of these loci exhibit strong separation between MSH6 -prof icient and008613259
[0020] MSH6 - def icient CRC MMRd samples even using as few as 10 of the loci (see Fig. 4N). A measure of diversity, such as SDI and JSD, at 10 or more of the coding homopolymer loci set forth in Table 1 provides a prediction of ICB response, CRC prognosis, progression of Lynch Syndrome to CRC and other clinically actionable statuses of subj ects, particularly human patients having MMRd CRC.
[0021] Accordingly, in a first aspect the present invention provides a method for estimating the mutation rate, or mutation rate-influenced property, of a tumour of a subj ect, comprising: providing nucleic acid sequencing data obtained by sequencing DNA or RNA derived from one or more cells of the tumour ("sample sequencing data" );
[0022] measuring or assessing microsatellite diversity of a plurality of, e.g. at least 10, erosion sensitive loci in the sample sequencing data to provide one or more diversity metrics;
[0023] correlating the one or more diversity metrics (per loci or grouped) to an estimated mutation rate or mutation rate-influenced property of the tumour.
[0024] In some embodiments, the mutation rate - inf luenced property may be selected from: response to an anti - cancer immunotherapy, such as ICB therapy, prognosis of the cancer, such as the overall survival time or progression - free survival time, likelihood of progressing to a more advanced stage of the cancer or from a pre - cancerous state to the onset of malignancy, MSH3 - and / or MSH6 - def iciency, and prediction of increasing intra - tumoral heterogeneity (ITH).
[0025] In some embodiments, the plurality of, e.g. at least 10, erosion sensitive loci are selected from the loci set forth in Table 1. In some embodiments, the at least 10 erosion008613259
[0026] sensitive loci comprise at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, at least 270 or substantially all of the loci set forth in Table 1. As shown in Figure 4N, the inventors have found that methods of the present disclosure are discriminative using even as f ew as 10 of their erosion sensitive loci. The inventors have demonstrated that using j ust 10 of their loci could discriminate between MSH6 / MSH3 prof icient and MSH6 / MSH3 def icient samples, f inding a surprisingly high - log10 (p - value) of approximately 5 ( see Examples ). They therefore contemplate the use of at least 10 of their identif ied loci (Table 1 ) in methods of the present disclosure. As shown in Figure 4N, discriminatory power initially increased with the number of Table 1 loci, with the use of >250 of the Table 1 loci still providing high discriminatory power. The inventors therefore also contemplate the use of their methods wherein the plurality of erosion sensitive loci comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270 or substantially all of the loci set forth in Table 1. The inventors found the use of 10 - 200 erosion sensitive loci to be highly discriminative, particularly 20 -140 loci. The plurality of erosion sensitive loci may pref erentially comprise 10 - 200, 20 - 150, 30 - 100, 40 - 80, 50 - 80, 50 - 70, 60 - 80 and / or 60 - 70 erosion sensitive loci set out in Table 1.
[0027] As shown in Figure 4N, the inventors have found that methods of the present disclosure are discriminative using even as f ew as the top 10 of their erosion sensitive loci. The inventors have demonstrated that using j ust these 10 of their loci could008613259
[0028] discriminate between MSH6 / MSH3 prof icient and MSH6 / MSH3 def icient samples, f inding a surprisingly high - log10 (p - value) of approximately 5 ( see Examples ). They therefore contemplate the use of these identif ied loci (Table 1 ) in methods of the present disclosure. As shown in Figure 4N, discriminatory power increased with use of the top 20, top 30, etc. loci until approximately the top 70 loci ranked in Table 1, with the use of the top - ranked >250 of the Table 1 loci still providing high discriminatory power. Accordingly, in some embodiments, the at least 10 erosion sensitive loci comprise the top 10, top 20, top 30, top 40, top 50, top 60, top 70, top 80, top 90, top 100, top 110, top 120, top 130, top 140, top 150, top 160, top 170, top 180, top 190, top 200, top 210, top 220, top 230, top 240, top 250, top 260, top 270, or top 280 loci by ranking as set forth in Table 1. In some embodiments, the at least 10 erosion sensitive loci comprise the top 80, top 90 or top 100 loci by ranking as set forth in Table 1. As shown in Fig. 4N, numbers in the range of 50 - 100 may be particularly advantageous for separation of prof icient vs. def icient tumours. In some cases, the at least 10 erosion sensitive loci comprise the top 50, top 60, top 70, top 80, top 90, or top 100 loci according to the ranking of Table 1.
[0029] In some cases, the method employs not more than 500, not more than 400, not more than 300, not more than 200 or not more than 150 erosion sensitive loci. In some cases, the method employs not more than 100, not more than 90, not more than 80, not more than 70, not more than 60, not more than 50, not more than 40, not more than 30, not more than 20, or not more than 10 erosion sensitive loci.
[0030] In some embodiments, the at least 10 erosion sensitive loci comprise all of the loci shown as having length 7 in Table 1 and / or all of the loci shown as having length 8 in Table 1. In008613259
[0031] some embodiments, the at least 10 erosion sensitive loci comprise at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, or substantially all of the loci shown as having length 7 in Table 1. In some embodiments, the at least 10 erosion sensitive loci comprise at least 5, at least 10, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or substantially all of the loci shown as having length 8 in Table 1. In some embodiments, the at least 10 erosion sensitive loci comprise at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180 or substantially all of the loci shown as having length 7 or length 8 in Table 1.
[0032] In some embodiments, the at least 10 erosion sensitive loci are located on at least 2, at least 5, at least 10, at least 15 or at least 20 chromosomes. For example, the at least 10 erosion sensitive loci may comprise 5 loci on a f irst chromosome and 5 loci on a second chromosome, or one locus on each of 10 dif f erent chromosomes or any other loci - chromosome relationship such that not all the loci are found on the same chromosome. The at least 2, at least 5, at least 10, at least 15 or at least 20 chromosomes may be human chromosomes. The chromosomes may be distinct types of chromosomes, such as each of chromosomes 1 - 22, X and Y when the tumour is a human tumour. For example, where the at least 10 erosion sensitive loci are located on chromosomes 1 and 2 where the tumour is a human tumour, the at least 10 erosion sensitive loci may be regarded as being located on 2 chromosomes. Conversely, where the at least 10 erosion sensitive loci are located on two copies of the same type of chromosome ( e. g. the maternal and paternal copy of chromosome 1 ), the at least 10 erosion008613259
[0033] sensitive loci may be regarded as being located on one chromosome only. Having loci located on multiple chromosomes may advantageously give a broader view of the genome -wide mutation rate or mutation rate - inf luenced property.
[0034] In some embodiments, the at least 10 erosion sensitive loci comprise loci located in genes. In some embodiments, the at least 10 erosion sensitive loci are located in genes. In some embodiments, the at least 10 erosion sensitive loci comprise loci located in exons. In some embodiments, the at least 10 erosion sensitive loci are located in exons. Erosion sensitive loci located in exons may be referred to as exonic loci or coding loci. In some embodiments, the at least 10 erosion sensitive loci are located in at least 2, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 40 or at least 50 genes. In some embodiments, the at least 10 erosion sensitive loci are located in at least 2, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 40 or at least 50 exons. Erosion sensitive loci located in genes, particularly those located in exons, may advantageously exhibit higher responsiveness to deficiencies in mismatch repair machinery, such as MSH- 6 and / or MSH- 3 deficiency. Accordingly, the use of erosion sensitive loci located in genes may provide a more sensitive proxy for the mutation rate or mutation rate - inf luenced property. Moreover, embodiments in which the erosion sensitive loci are located in, that is spread across, 2 or more genes may provide for a broader view of mutation rate affecting more than a single gene. For example, the at least 10 erosion sensitive loci may comprise 5 loci located in a first gene or exon and 5 loci located in a second gene or exon, or may comprise one locus on each of 10 different genes or 10 different exons (of one or more genes) or any other loci -gene or loci - exon relationship such that not all the loci are008613259
[0035] located in the same gene or not all the loci are located in the same exon.
[0036] In some embodiments, the at least 10 erosion sensitive loci comprise loci with approximately 5 - 11 bp repeat length, e.g. 5 - 11 bp, 5 - 10 bp, 6 - 11 bp, or 6 - 10 bp. In some embodiments, the at least 10 erosion sensitive loci comprise loci with 6 - 10 bp repeat length. The repeat length may be the length of an individual repeat unit in base pairs.
[0037] In some embodiments, measuring microsatellite diversity comprises measuring the diversity of microsatellite repeat length. For example, measuring may comprise generating a read count distribution of the microsatellite repeat lengths at said loci.
[0038] In some embodiments the greater the one or more diversity metrics, the higher the estimated mutation rate of the tumour. For example, where the measured diversity at the loci is higher than a reference value (such as the median value measured for a plurality of reference tumours known to be MSH3 -prof icient / MSH6 -prof icient or the median value measured for a plurality of reference tumours known to be MSH3 -deficient / MSH6 - def icient), the mutation rate of the sample tumour may be inferred to be high. Where the measured diversity at the loci is lower than a reference value (such as the median value measured for a plurality of reference tumours known to be MSH3 -prof icient / MSH6 -prof icient or the median value measured for a plurality of reference tumours known to be MSH3 - def icient / MSH6 - def icient), the mutation rate of the sample tumour may be inferred to be low. Samples or tumours with measured diversity at the loci lower than a reference value (such as the median value measured for a plurality of reference tumours known to be MSH3 -prof icient / MSH6 -008613259
[0039] proficient or the median value measured for a plurality of reference tumours known to be MSH3 - def icient / MSH6 - def icient) may be inferred to have a lower mutation rate than samples or tumours with measured diversity at the loci lower than a reference value.
[0040] In some embodiments, the one or more diversity metrics comprises Shannon Diversity Index (SDI). In particular, the SDI may be determined according to the formula:
[0041] R
[0042] Shannon diversity =
[0043]
[0044] i=l
[0045] wherein pi = the proportion of total sequence reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite.
[0046] In some embodiments, the one or more diversity metrics comprises a measure of the divergence of: (i) the probability distribution of microsatellite lengths at each of said loci for the tumour; and (ii) the probability distribution of microsatellite lengths of the same loci for sequence data obtained by sequencing DNA or RNA derived from one or more normal (non- cancer) cells of the subj ect.
[0047] In some embodiments, the one or more cells may be a population of cells. For example, the one or more cells may be a population of tumour cells, a population of normal (non- cancer or non- tumour) cells, or a population comprising a mixture of tumour cells and normal cells. In such embodiments, the microsatellite diversity may be measured or assessed for the alleles of at each of the at least 10 erosion sensitive loci present in the population of cells. The microsatellite008613259
[0048] diversity may account for both the number of different alleles and their relative frequencies in the population of cells.
[0049] In some embodiments, the one or more diversity metrics comprises the Jensen - Shannon Distance (JSD) according to the formula:
[0050] JSD(P II Q) = |r>(P II M) + |r>(Q II M)
[0051]
[0052] wherein P and Q are the probability distributions of microsatellite lengths in tumour and normal samples,
[0053]
[0054] respectively, M is the average distribution defined as M = — — and D(X II K) is the Kullback - Leibler divergence between X and Y.
[0055] In particular, the Kullback - Leibler divergence (DKL) may be given by the formula:
[0056]
[0057] In the context of the present disclosure, x represents microsatellite length, and X and Y represent probability distributions of microsatellite lengths. Hence, the Kullback-
[0058]
[0059] To / over all values of x (i. e. all microsatellite lengths observed for X and Y). X and Y may, in some cases, be P and Q, respectively. X and Y may, in some cases, be P and M, respectively. X and Y may, in some cases, be Q and M, respectively.
[0060] In some embodiments, the one or more diversity metrics are determined at each of said loci and combined to provide a global measure of diversity across said loci, optionally wherein the global measure is the median or the mean. Median was chosen as a preferred metric instead of mean in order to008613259
[0061] avoid disproportionately reflecting outliers in the population. However, other measures of central tendency, such as mean are contemplated herein.
[0062] In some embodiments, the global measure of diversity across said loci (e.g. the median SDI) is compared to a reference value. The tumour may subsequently be classified into one of a plurality of classes comprising a first class and a second class. The tumour may be classified into the first class when the global measure of diversity is above the reference value. The tumour may be classified into the second class when the global measure of diversity is below the reference value.
[0063] Tumours classified into the first class may have a higher mutation rate than tumours classified into the second class. In some embodiments, the global measure of diversity across said loci (e.g. the median SDI) is compared to a reference value, wherein the global measure of diversity being above the reference value indicates that the mutation rate of the tumour is high, and wherein the global measure of diversity being at or below the reference value indicates that the mutation rate of the tumour is low. Typically, the reference value will be selected or derived from analysis of tumour samples of known tumour MSH3 / MSH6 proficiency or deficiency. For example, the median or an upper quartile or a lower quartile may be determined for each of a proficient group and a deficient group of tumour standards. The reference value will generally be chosen so as to maximise the reliability of classification of an unknown sample.
[0064] In some embodiments, correlating the one or more diversity metrics to the estimated mutation rate or mutation rate-influenced property of the tumour comprises: application of a binary classification model, application of a clustering algorithm or percentile -based stratification. In some008613259
[0065] embodiments, correlating the one or more diversity metrics to the estimated mutation rate or mutation rate - inf luenced property of the tumour comprises: application of a classification model (e.g. binary classification model or multiclass classification model) or a regression model.
[0066] In some embodiments correlating the one or more diversity metrics to the estimated mutation rate or mutation rate-influenced property of the tumour comprises:
[0067] inputting sample data comprising the one or more diversity metrics (e.g. SDI) determined at each of said loci into a machine learning classifier, wherein the machine learning classifier has been trained on a training dataset comprising the same one or more diversity metrics (e.g. SDI) determined at each of said loci for a plurality of samples of cancers known to be of low mutation rate (e.g. MSH3 / MSH6 proficient) and a plurality of samples of cancers known to be of high mutation rate (e.g. MSH3 / MSH6 deficient); and causing the machine learning classifier to classify the tumour as high or low mutation rate based on the inputted sample data.
[0068] In some embodiments correlating the one or more diversity metrics to the estimated mutation rate or mutation rate-influenced property of the tumour comprises:
[0069] inputting the one or more diversity metrics into a machine learning model, wherein the machine learning model has been trained to take as input the one or more diversity metrics and output the estimated mutation rate or mutation rate - inf luenced property of the tumour.
[0070] In some embodiments, the sample data input into the machine learning model comprises the one or more diversity metrics (e.g. SDI) determined at each of the at least 10 erosion008613259
[0071] sensitive loci. In some embodiments, the sample data input into the machine learning model comprises a global measure of the one or more diversity metrics measured at the at least 10 erosion sensitive loci (e.g. a mean or median of the one or more diversity metrics determined at each of the at least 10 erosion sensitive loci). In some embodiments, the machine learning model is trained to output the estimated mutation rate. The estimated mutation rate may be used to determine a mutation rate - inf luenced property of the tumour. In some embodiments, the machine learning model is trained to output the mutation rate - inf luenced property of the tumour. The mutation rate - inf luenced property may be response to cancer immunotherapy, a prognosis of the tumour, progression of a subj ect having hereditary nonpolyposis colorectal cancer (Lynch Syndrome) to malignancy onset (i. e. malignant CRC), classifying the tumour as MSH3 -prof icient or MSH3 -def icient, and / or classifying the tumour as MSH6 -prof icient or MSH6 -deficient.
[0072] In some embodiments, the machine learning model is a classifier. The machine learning model may be trained to classify a tumour as one of a plurality of classes. The plurality of classes may comprise a first class and a second class. The first class may comprise tumours with a higher mutation rate than tumours in the second class. The first class may comprise tumours with a high mutation rate and the second class may comprise tumours with a low mutation rate. The first class may comprise tumours predicted to respond to a cancer treatment and the second class may comprise tumours predicted not to respond to a cancer treatment, e.g. immunotherapy. The first class may comprise tumours predicted to have a better prognosis than tumours in the second class, for example tumours in the first class may be predicted to have longer (e.g. above a predetermined threshold) overall008613259
[0073] survival, disease free survival and / or time to recurrence than tumours in the second class. The first class may comprise tumours (e.g. polyps) predicted to progress to malignancy (e.g. CRC) and the second class may comprise tumours (e.g. polyps) predicted not to progress to malignancy (e.g. CRC). The first class may comprise MSH6 proficient and / or MSH3 proficient tumours and the second class may comprise MSH6 deficient and / or MSH3 deficient tumours.
[0074] In some embodiments, the machine learning model is a regression model. The machine learning model may be trained to output a mutation rate, probability of response to a cancer treatment (e.g. immunotherapy), overall survival, disease free survival, time to recurrence, probability of progression to malignancy, percentage tumour bed regression in response to cancer treatment (optionally immunotherapy, further optionally ICB therapy), percentage MSH6 function relative to wild- type and / or percentage MSH3 function relative to wild- type. The machine learning model may be a linear regression model. The machine learning model may be a Random Forest or a Gradient Boosting model. In some embodiments, the machine learning model is a Cox proportional hazards regression model. The one or more diversity metrics may be continuous or categorical variables. The machine learning model may be multivariate. The machine learning model may adjust for at least one, two, three, four, five, six, seven or eight variables selected from: age, gender, planned treatment (e.g. Capox vs Folfox), trial arm duration, pathological T stage (e.g. with T4 as reference), pathological N stage (e.g. with NO as reference), tumour differentiation, and / or tumour sidedness. The machine learning model may take as input polyp size, e.g. diameter. In embodiments, survival time may be measured from randomisation to death or last follow-up.008613259
[0075] In some embodiments of any aspect of the present invention the method may be computer implemented.
[0076] In some embodiments, the method comprises a preceding step of sequencing DNA or RNA derived f rom one or more cells of the tumour. In such cases the method may be implemented using laboratory tools ( e. g. a sequencing device) and one or more computers for carrying out the data analysis steps.
[0077] In some embodiments, the step of sequencing DNA or RNA derived f rom one or more cells of the tumour comprises use of a targeted sequencing panel comprising not more than 10000, not more than 1000 or not more than 500 microsatellite loci.
[0078] In some embodiments, the targeted sequencing panel comprises at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, at least 270 or substantially all of the loci set forth in Table 1.
[0079] In some embodiments, the targeted sequencing panel comprises the at least 10 erosion sensitive loci. The targeted sequencing panel may comprise any of the at least 10 erosion sensitive loci disclosed herein.
[0080] In some embodiments, the targeted sequencing panel comprises the at least 10 erosion sensitive loci comprise the top 10, top 20, top 30, top 40, top 50, top 60, top 70, top 80, top 90, top 100, top 110, top 120, top 130, top 140, top 150, top 160, top 170, top 180, top 190, top 200, top 210, top 220, top 230, top 240, top 250, top 260, top 270, or top 280 loci by ranking as set forth in Table 1.
[0081] In some embodiments, the targeted sequencing panel comprises not more than 100, not more than 90, not more than 80, not008613259
[0082] more than 70, not more than 60, not more than 50, not more than 40, not more than 30, not more than 20, or not more than 10 erosion sensitive loci.
[0083] In some embodiments, the targeted sequencing panel comprises at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or substantially all of the loci shown as having length 7 in Table 1.
[0084] In some embodiments, the targeted sequencing panel comprises at least 5, at least 10, at least 30, at least 40, at least 50, at least 60, or substantially all of the loci shown as having length 8 in Table 1.
[0085] In some embodiments, the targeted sequencing panel comprises at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180 or substantially all of the loci shown as having length 7 or length 8 in Table 1.
[0086] In some embodiments, the targeted sequencing panel comprises loci located on at least 2, at least 5, at least 10, at least 15 or at least 20 chromosomes. The at least 2, at least 5, at least 10, at least 15 or at least 20 chromosomes may be human chromosomes. The chromosomes may be distinct types of chromosomes, such as each of chromosomes 1 - 22, X and Y when the tumour is a human tumour. For example, where the at least 10 erosion sensitive loci are located on chromosomes 1 and 2 where the tumour is a human tumour, the at least 10 erosion sensitive loci may be regarded as being located on 2008613259
[0087] chromosomes. Conversely, where the at least 10 erosion sensitive loci are located on two copies of the same type of chromosome (e.g. the maternal and paternal copy of chromosome 1), the at least 10 erosion sensitive loci may be regarded as being located on one chromosome only. Having loci located on multiple chromosomes may advantageously give a broader view of the genome-wide mutation rate or mutation rate - inf luenced property.
[0088] In some embodiments, the targeted sequencing panel comprises loci located in genes. In some embodiments, the targeted sequencing panel comprises loci located in at least 2, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 40 or at least 50 genes. Erosion sensitive loci located in genes, particularly those located in exons, may advantageously exhibit higher responsiveness to deficiencies in mismatch repair machinery, such as MSH- 6 and / or MSH- 3 deficiency. Accordingly, the use of erosion sensitive loci located in genes may provide a more sensitive proxy for the mutation rate or mutation rate - inf luenced property.
[0089] In some embodiments, the sample sequencing data is derived from a tumour tissue sample or a cell - free nucleic acid sample, optionally wherein the sample comprises circulating tumour DNA (ctDNA). A cell - free nucleic acid sample may be any sample of a biological fluid that has been shown to contain cfDNA, including but not limited to plasma, serum, blood, cerebrospinal fluid, ascites, saliva, and urine.
[0090] In some embodiments, the tumour is a cancer commonly associated with microsatellite instability. In some embodiments, the tumour is selected from: a colon cancer, a gastric cancer, a small bowel cancer, a rectal cancer, an008613259
[0091] endometrial cancer, an ovarian cancer, a biliary tract cancer, a urinary tract cancer, a brain cancer, and a skin cancer. In some embodiments, the tumour is selected from: a gastrointestinal cancer, a lung cancer and a breast cancer.
[0092] In some embodiments, the tumour is a colorectal cancer (CRC).
[0093] In some embodiments, the tumour is DNA mismatch repair deficient (MMRd). In some embodiments, the tumour is known or has been previously assessed to be MMRd.
[0094] In some embodiments, the tumour mutation rate - inf luenced property is cancer immunotherapy response (e.g. response to ICB therapy) and / or cancer prognosis (e.g. overall survival, disease free survival and / or time to recurrence) and / or progression of a subj ect having hereditary nonpolyposis colorectal cancer (Lynch Syndrome) to malignancy onset (i. e. malignant CRC). In some embodiments, the tumour mutation rate-influenced property is probability of response to a cancer treatment (optionally immunotherapy, further optionally ICB therapy), probability of progression of a subj ect having hereditary nonpolyposis colorectal cancer to malignancy, and / or percentage tumour bed regression in response to cancer treatment (optionally immunotherapy, further optionally ICB therapy).
[0095] In some embodiments, the method is for predicting response to cancer immunotherapy, optionally immune checkpoint blockade (ICB) therapy, and wherein a higher value of the one or more diversity metrics indicates that the subj ect will respond poorly to the cancer immunotherapy and a lower value of the one or more diversity metrics indicates that the subj ect will respond well to the cancer immunotherapy.008613259
[0096] In some embodiments, the method is for providing a prognosis of the tumour, and wherein a higher value of the one or more diversity metrics indicates that the subj ect has a poor prognosis and a lower value of the one or more diversity metrics indicates that the subj ect has a good prognosis. Poor prognosis may indicate shorter progression - free survival and / or shorter overall survival and / or shorter time to recurrence, and good prognosis may indicate longer progression - free survival and / or longer overall survival and / or longer time to recurrence.
[0097] In some embodiments, the tumour comprises an MMRd CRC tumour, such as a stage II or stage III MMRd CRC tumour, and poor prognosis indicates shorter progression - free survival and / or shorter overall survival and / or shorter time to recurrence, and wherein good prognosis indicates longer progression - free survival and / or longer overall survival and / or longer time to recurrence.
[0098] In some embodiments, the method is for predicting progression of a subj ect having hereditary nonpolyposis colorectal cancer (Lynch Syndrome) to malignancy onset (i. e. malignant CRC). The one or more cells of the tumour may be one or more cells of a colorectal precursor adenoma. A higher value of the one or more diversity metrics indicates that the subj ect is predicted to progress to the onset of CRC malignancy. A lower value of the one or more diversity metrics indicates that the subj ect is not predicted to progress to onset of CRC malignancy.
[0099] In some embodiments, the method is for classifying a MMRd CRC tumour as MSH6 -prof icient or MSH6 -def icient, wherein a higher value of the one or more diversity metrics indicates that the tumour is MSH6 - def icient and a lower value of the one or more diversity metrics indicates that the tumour is MSH6 -008613259
[0100] proficient. MSH6 - def icient tumours are expected to have higher mutation rate than MSH6 -prof icient tumours. In embodiments, the estimated mutation rate may be a classification of the tumour as MSH6 -prof icient or MSH6 - def icient. In embodiments, the estimated mutation rate may be a percentage MSH6 function relative to wild- type.
[0101] In some embodiments, the method is for classifying a MMRd CRC tumour as MSH3 -prof icient or MSH3 -def icient, wherein a higher value of the one or more diversity metrics indicates that the tumour is MSH3 - def icient and a lower value of the one or more diversity metrics indicates that the tumour is MSH3 -proficient. MSH3 - def icient tumours are expected to have higher mutation rate than MSH3 -prof icient tumours. In embodiments, the estimated mutation rate may be a classification of the tumour as MSH3 -prof icient or MSH3 - def icient. In embodiments, the estimated mutation rate may be a percentage MSH3 function relative to wild- type.
[0102] In some embodiments, the mutation rate - inf luenced property may be sub - clonal complexity of the tumour and / or intra - tumoral heterogeneity (ITH). In some embodiments, the method is for estimating sub - clonal complexity of the tumour and / or intra -tumoral heterogeneity (ITH). A higher value of the one or more diversity metrics indicates that the tumour has greater sub-clonal complexity and / or higher ITH. A lower value of the one or more diversity metrics indicates that the tumour has lower sub - clonal complexity and / or lower ITH.
[0103] In some embodiments, the method is for sub - clonal tracking of the tumour and / or temporal tumour evolutionary reconstruction. The sample sequencing data may comprise data for each of a plurality of different locations of the tumour and / or for each of a plurality of different time points in the history of the tumour.008613259
[0104] In some embodiments the one or more diversity metrics comprise the Jensen - Shannon Distance (JSD). The JSDs of the loci may be used to construct a phylogeny of the tumour.
[0105] In a second aspect, the present invention provides a method (e.g. a computer implemented method) for identifying coding microsatellites that are sensitive to changes in tumour mutation rate, comprising:
[0106] providing whole exome sequencing data for each of a plurality of tumours known to be MSH6 -prof icient and / or MSH3 -proficient ("proficient tumours" ) and a plurality of tumours known to be MSH6 - def icient and / or MSH3 - def icient ("deficient tumours" );
[0107] generating for each of a plurality of candidate microsatellites a read count distribution for the proficient tumours and a read count distribution for the deficient tumours;
[0108] deriving from the read count distributions a Shannon diversity index (SDI) for each candidate microsatellite according to the formula:
[0109] R
[0110] Shannon diversity = — piln(pi)
[0111]
[0112] i=l
[0113] wherein pi = the proportion of total sequence reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite;
[0114] ranking the candidate microsatellites based on the change (A) in SDI between proficient and deficient tumours; and selecting based on their ranking, the candidate microsatellites most sensitive to changes in tumour mutation rate.008613259
[0115] In some embodiments, the proficient and deficient tumours may be MMRd CRC tumours.
[0116] In a third aspect, the present invention provides a method of treatment of a subj ect having a cancer, comprising:
[0117] carrying out the method of the first aspect of the invention;
[0118] thereby determining that the subj ect is predicted to respond to anti - cancer immunotherapy (e.g. ICB therapy); and administering said anti - cancer therapy (e.g. ICB therapy) to the subj ect in need thereof.
[0119] In a fourth aspect, the present invention provides an ICB agent for use in a method of treatment of a cancer in a subj ect, wherein the method of the first aspect of the invention has been carried out in respect of a sample obtained from the subj ect and the subj ect has thereby been predicted to respond to the ICB agent.
[0120] In a fifth aspect, the present invention provides use of an ICB agent in the preparation of a medicament for use in a method of treatment of a cancer in a subj ect, wherein the method of the first aspect of the invention has been carried out in respect of a sample obtained from the subj ect and the subj ect has thereby been predicted to respond to the ICB agent.
[0121] In a sixth aspect, the present invention provides a system for use in the method of the first aspect of the invention, the system comprising a processor and memory, wherein the memory stores computer readable instructions that when processed by the processor cause the system to carry out the method of the first aspect of the invention.008613259
[0122] In a seventh aspect, the present invention provides computer readable media on which are stored computer readable instructions that when processed by a computer processor cause the computer to carry out the method of the first aspect of the invention.
[0123] In an eighth aspect, the present invention provides a targeted sequencing panel for not more than 10000, not more than 1000 or not more than 500 microsatellite loci, wherein the targeted sequencing panel comprises probes for at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, at least 270 or substantially all of the loci set forth in Table 1. The targeted sequencing panel may additionally or alternatively comprise probes for any subset of the Table 1 loci disclosed for the aforementioned embodiments.
[0124] The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided.
[0125] Brief Description of Figures
[0126] Fig. 1 is a flow diagram showing, in schematic form, a method of determining or predicting the prognosis of a subj ect and determining or predicting treatment response of the subj ect according to the disclosure.
[0127] Fig. 2 shows an embodiment of a system for implementing methods of the disclosure.
[0128] Fig. 3A (top) shows the targeted re - sequencing panel design and relative proportion of each of the panel groups included (70kb footprint). Fig. 3A (bottom) shows the identification of008613259
[0129] group D microsatellites, which are most responsive to changes in mutation rate, as determined through WES and ranking by change in Shannon diversity between MSH6 / MSH3 -wild- type and MSH6 / MSH3 -mutant samples. 549 microsatellites ("merged targets" ) were included in the final panel.
[0130] Fig. 3B shows a bioinformatics pipeline used to identify microsatellite positions most responsive to changes in mutation rate (group D loci) using multi - region whole- exome sequencing (WES) data. MSlsensor and SciRoKoCo were used to derive instability estimates for each microsatellite in the exome. Downstream steps included identifying microsatellite loci which exhibited the greatest change in Shannon diversity index between tumour areas of high mutation rate and tumour areas of low mutation rate. Mutation rates were determined using MOBSTER mutation rate modelling and underlying MSH6 and MSH3 mutation status.
[0131] Fig. 3C shows a schematic of the protocol used to isolate samples for the WGS -matched cohort (n=21). MSH6 status was determined using immunohistochemistry (IHC). Two or more biopsy punches were taken from within MSH6 -prof icient and MSH6 - def icient subclones (shown in the IHC images).
[0132] Fig. 3D shows a heatmap showing the cumulative sequencing read depths of all microsatellite loci included in the targeted panel. Panel groups A-D are shown on the x axis. Sample groups (normal reference [n=19], WGS -matched [n=21] and WGS - unmatched [n=56] ) are shown on the y axis. The corresponding histogram shows the frequency distribution of raw read counts.
[0133] Fig. 3E shows a schematic method used to extract information from a MSlsensor read count table output to produce density008613259
[0134] lines for microsatellite length dis tributions ( see Fig. 3F and Fig. 3H).
[0135] Fig. 3F (top) shows an MSH6 immunohistochemistry image of a tumour showing subclonal MSH6 loss. An MSH6 -prof icient clone and MSH6 - def icient clone are shown. Fig. 3F (bottom) shows microsatellite length distribution plots for MSH6 and MSH3 in an MSH6 -prof icient clone and an MSH6 - def icient clone. MSH6 and MSH3 mutation status were determined through indels in the C8 and A8 homopolymer regions, respectively.
[0136] Fig. 3G shows a schematic of a genetic deficiency scoring system using a weighted sums approach, using normalised MSlsensor distribution tables.
[0137] Fig. 3H shows microsatellite length distribution plots comparing tumour vs. normal samples for all microsatellite positions in Group A-D.
[0138] Fig. 4A shows an explanation of the Shannon diversity index and method of calculating the median Shannon diversity index for each microsatellite position in the panel using the MSlsensor read distribution table output.
[0139] Figs. 4B-K show violin plots illustrating the difference in median Shannon diversity indices for different groups of microsatellites between proficient and deficient samples. Samples were grouped as proficient or deficient according to the genetic deficiency scoring method described in Fig. 3G and median Shannon diversity indices were calculated as shown in Fig. 4A. The Wilcoxon signed- rank statistic is reported on top of each panel and median values are represented by black horizontal lines.008613259
[0140] Fig. 4B shows the difference in median Shannon diversity indices for group A positions [n=48] between proficient and deficient samples. Fig. 4C shows the difference in median Shannon diversity indices for group B positions [n=24] between proficient and deficient samples. Fig. 4D shows the difference in median Shannon diversity indices for group C positions [n=50] between proficient and deficient samples. Fig. 4E shows the difference in median Shannon diversity indices for group D positions [n=288] between proficient and deficient samples in the WGS -matched cohort. Fig. 4F shows the difference in median Shannon diversity indices for group D positions [n=274] between proficient and deficient samples in the WGS - unmatched cohort.
[0141] Fig. 4G shows the difference in median Shannon diversity indices for length 6 group D positions between proficient and deficient samples in the WGS -matched cohort. Fig. 4H shows the difference in median Shannon diversity indices for length 7 group D positions between proficient and deficient samples in the WGS -matched cohort. Fig. 41 shows the difference in median Shannon diversity indices for length 8 group D positions between proficient and deficient samples in the WGS -matched cohort. Fig. 4J shows the difference in median Shannon diversity indices for length 9 group D positions between proficient and deficient samples in the WGS -matched cohort. Fig. 4K shows the difference in median Shannon diversity indices for length 10 group D positions between proficient and deficient samples in the WGS -matched cohort. All lengths in Figs. 4G-K refer to the number of repeat units in wild- type state.
[0142] Fig. 4L shows regression analysis to identify the correlation between patients ' genetic deficiency scores (Fig. 3G) and008613259
[0143] median Shannon Diversity Index of group D loci for the WGS -matched cohort. R = Spearman ' s Rho correlation coefficient.
[0144] Fig. 4M shows regression analysis to identify the correlation between patients ' genetic deficiency scores (Fig. 3G) and median Shannon Diversity Index of group D loci for the WGS -unmatched cohort. R = Spearman ' s Rho correlation coefficient.
[0145] Fig. 4N shows a plot illustrating the effect of the number of group D loci included in the median Shannon diversity index comparisons between proficient and deficient samples (Figs. 4B-K) on the - loglO p-values of the Wilcoxon signed- rank statistic. The x-axis shows the number of group D loci that were included. The highest - ranking loci (ranked according to shift in median SDI between Proficient and Deficient groups) were included for each comparison.
[0146] Fig. 40 shows an example heatmap illustrating the Jensen-Shannon distances of each of the top 100 group D loci (ranked by shift in median SDI between Proficient and Deficient groups) between normal reference and tumour samples. One matched normal sample (SI) and three multi - region tumour samples derived from a single patient (S2 - 4) are shown.
[0147] Fig. 4P shows the corresponding phylogenetic tree to Fig. 40, constructed using Jensen - Shannon distances matrices (neighbour - j oining method). The plot illustrates the idea that tumours with greater between - sample variation (i. e. JSDs) have longer branches in the star phylogeny.
[0148] Fig. 4Q shows a violin plot illustrating the distribution of differences in Jensen - Shannon distances (normal vs. tumour) for each of the top 100 group D microsatellite position for Proficient and Deficient sample groups in the WGS -matched008613259
[0149] cohort. Values are plotted on a loglO y-axis for clearer visualisation.
[0150] Fig. 4R shows a violin plot illustrating the distribution of differences in Jensen - Shannon distances (normal vs. tumour) for each of the top 100 group D microsatellite position for Proficient and Deficient sample groups in the WGS - unmatched cohort. Values are plotted on a loglO y-axis for clearer visualisation.
[0151] Detailed Description
[0152] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.
[0153] For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations.
[0154] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subj ect matter described.
[0155] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word "comprise" and "include", and variations such as "comprises", "comprising", and "including" will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.008613259
[0156] It must be noted that, as used in the specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent "about," it will be understood that the particular value forms another embodiment. The term "about" in relation to a numerical value is optional and means for example + / - 10%.
[0157] " And / or" where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example " A and / or B" is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein.
[0158] In describing the present invention, the following terms will be employed and are intended to be defined as indicated below.
[0159] A "sample" as used herein may be a cell or tissue sample, a biological fluid, an extract (e.g. a DNA extract obtained from a cell, tissue or fluid sample), from which genomic material can be obtained for genomic analysis, such as genomic sequencing (e.g. whole genome sequencing, or targeted sequencing such as whole exome sequencing or targeted panel sequencing). The sample may be a cell, tissue or biological fluid sample obtained from a subj ect (e.g. a biopsy). Such samples may be referred to as "subj ect samples". The sample may be any sample comprising genomic DNA or cell free DNA. In particular, the sample may be a tumour sample, a biological fluid sample containing DNA or cells, a blood sample008613259
[0160] (including plasma or serum sample), a urine sample, a cervical smear, an ascites fluid sample, or a sample derived therefrom (e.g. after DNA purification). It has been found that urine, ascites fluid and cervical smears contains cells, and so may provide a suitable sample for use in accordance with the present invention. Other sample types suitable for use in accordance with the present disclosure include fine needle aspirates, lymph nodes samples (e.g. aspirates or biopsies), surgical margins, bone marrow or other tissue from a tumour microenvironment, where traces of tumour DNA may be found or expected to be found. A particularly preferred sample type is a tumour sample obtained as a result of surgical resection of a colorectal cancer. The sample may be one which has been freshly obtained from a subj ect or may be one which has been processed and / or stored prior to genomic / transcriptomic analysis (e.g. frozen, fixed or subj ected to one or more purification, enrichment or extraction steps). The sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells or genomic material derived therefrom, whether from a biological sample obtained from a subj ect, or from a sample obtained from e.g. a cell line. In embodiments, the sample is a sample obtained from a subj ect, such as a human subj ect. The sample is preferably from a mammalian (such as e.g. a mammalian cell sample or a sample from a mammalian subj ect, such as a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig), preferably from a human (such as e.g. a human cell sample or a sample from a human subj ect). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the sequence data acquisition (e.g. sequencing) location, and / or any computer - implemented method steps described herein may take place at a location remote from the sample collection location and / or remote from the sequence data acquisition008613259
[0161] (e.g. sequencing) location (e.g. the computer - implemented method steps may be performed by means of a networked computer, such as by means of a "cloud" provider). A sample may be a tissue biopsy, such as a sample of fresh frozen tissue or a sample of formalin - fixed paraffin embedded (FFPE) tissue, or a sample comprising circulating tumour DNA (ctDNA) such as e.g. a sample of biological fluid (sometimes referred to as liquid biopsy). The samples used in methods of the present disclosure are typically samples comprising tumour cells (e.g. a tumour sample or sample comprising circulating tumour cells) or genetic material derived from tumour cells (such as e.g. cell free DNA, or DNA extracted from a sample comprising tumour cells - e.g. from a cell line or tissue sample - or circulating tumour cells). A sample may be a "mixed" sample comprising cells with different genotypes or genetic material derived therefrom. For example, a sample may be a sample comprising tumour cells and normal cells, or DNA derived therefrom, such as e.g. in the context of cell free DNA (cfDNA) samples comprising circulating tumour DNA (ctDNA).
[0162] A "tumour sample" refers to a sample that contains tumour cells or genetic material derived therefrom. The tumour may be benign, precancerous or cancerous / malignant. The tumour sample may be a cell or tissue sample (e.g. a biopsy) obtained directly from a tumour. A tumour sample may be a sample that comprises tumour cell or genetic material derived therefrom, that has not been obtained directly from a tumour. For example, a tumour sample may be a sample comprising circulating tumour cells or circulating tumour DNA. Thus, a tumour sample may also be a biological fluid (e.g. a liquid biopsy such as a blood, urine, or cerebrospinal fluid biopsy). A sample comprising a mixture of tumour cells and other cells (or material genetic derived therefrom) may be subj ect to one or more processing steps, whether prior to or subsequent to008613259
[0163] the acquisition of sequence data, in order to identify sequence data that is representative of the genetic material from the tumour. For example, a sample comprising cells may be subj ect to one or more cell purification steps which selectively enrich the sample for tumour cells. As another example, a sample of genetic material may be subj ect to one or more capture and / or size selection steps to selectively enrich the sample for tumour - derived genetic material. Protocols for doing this are known in the art. As another example, sequence data may be subj ect to one or more filtering steps (e.g. based on fragment length in the context of cell - free DNA samples) to enrich the data for information that relates to tumour - derived genetic material. Protocols for doing this are known in the art. In embodiments, the sample is a sample comprising tumour cells.
[0164] The term "sequence data" refers to information that is indicative of the presence of genetic material in a sample that has a particular sequence. Such information may be obtained using sequencing technologies, such as e.g. next generation sequencing (NGS), for example whole exome sequencing (WES), whole genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing), or using array technologies, such as e.g. SNP arrays, or other molecular counting assays. In embodiments, the sequence data is obtained by DNA sequencing, and particularly next generation sequencing. In such embodiments, the sequence data comprises sequencing reads, or information derived therefrom such as the count of the number of sequencing reads that have a particular sequence. When non-digital technologies are used such as array technology, the sequence data may comprise a signal (e.g. an intensity value) that is indicative of the number of sequences in the sample that have a particular sequence, for example by comparison to an appropriate control.008613259
[0165] Information such as the count of sequencing reads that have a particular sequence can be derived from sequencing reads by mapping the sequence data to a reference sequence, for example a reference genome, using methods known in the art (such as e.g. Bowtie (Langmead, B., Trapnell, C., Pop, M. et al.
[0166] Ultrafast and memory- efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009 ) ) or BWA (Li H, Durbin R. Fast and accurate short read alignment with Burrows -Wheeler transform. Bioinformatics. 2009 Jul
[0167] 15; 25 (14): 1754 - 60). This may result in aligned sequencing reads, for example in the form of a SAM or BAM file. Thus, sequence data may be associated with a particular genomic location (where the "genomic location" refers to a location in the reference genome to which the sequence data was mapped).
[0168] As used herein "erosion sensitive loci" refers to a microsatellite that exhibits sensitivity, in terms of repeat length or other sequence property, to the mutation rate of the cell or cells making up the sample in question, e.g., a tumour sample. As described in detail herein, the inventors devised a method for identifying and ranking such erosion sensitive loci. In particular, one or more erosion sensitive loci may be identified by the following procedure:
[0169] providing whole exome sequencing data for each of a plurality of tumours known to be MSH6 -prof icient and / or MSH- 3 -proficient ("proficient tumours" ) and a plurality of tumours known to be MSH6 - def icient and / or MSH3 - def icient ("deficient tumours" );
[0170] generating for each of a plurality of candidate microsatellites a read count distribution for the proficient tumours and a read count distribution for the deficient tumours;008613259
[0171] deriving from the read count distributions a Shannon diversity index (SDI) for each candidate microsatellite according to the formula:
[0172] R
[0173] Shannon diversity = — piln(pi)
[0174]
[0175] i=l
[0176] wherein pi = the proportion of total sequence reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite;
[0177] ranking the candidate microsatellites base on the change (A) in SDI between proficient and deficient tumours; and selecting based on their ranking, the candidate microsatellites most sensitive to changes in tumour mutation rate. In particular, the proficient and deficient tumours may be MMRd CRC tumours. Specific examples of erosion sensitive loci are shown in Table 1 herein.
[0178] As used herein "mutation rate - inf luenced property" refers to a property, such as a clinically relevant status or prediction, that is dependent on the course, rate, or extent of mutation, whether spatial or temporal. Examples include the likelihood of response to an anti - cancer immunotherapy, such as ICB therapy, tumour mutation burden (TMB), prognosis of the cancer, such as the overall survival time, progression - free survival time, likelihood of progressing to a more advanced stage of the cancer or from a pre - cancerous state to the onset of malignancy, likelihood of metastasis, likelihood of increasing intra - tumoral heterogeneity (ITH), increased sub-clonal complexity, neoantigen load, and susceptibility to DNA damaging agents or DNA repair inhibition. Without wishing to be bound by any particular theory, the present inventors reason that by measuring diversity of erosion sensitive loci,008613259
[0179] the extent or likelihood of such properties can be estimated and monitored over time.
[0180] A composition as described herein may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable carrier, diluent or excipient. The pharmaceutical composition may optionally comprise one or more further pharmaceutically active polypeptides and / or compounds. Such a formulation may, for example, be in a form suitable for intravenous infusion.
[0181] As used herein "treatment" refers to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment.
[0182] As used herein "cancer immunotherapy" refers to refers to a therapeutic approach comprising administration of an immunogenic composition (e.g. a vaccine), a composition comprising immune cells, or an immunoactive drug, such as e.g. a therapeutic antibody or immune checkpoint blockade (ICB) / immune checkpoint inhibitor (ICI), to a subj ect. The term "immunotherapy" may also refer to the therapeutic compositions themselves. In the context of the present invention, the immunotherapy typically targets a neoantigen. For example, an immunogenic composition or vaccine may comprise a neoantigen, neoantigen presenting cell or material necessary for the expression of the neoantigen. As another example, a composition comprising immune cells may comprise T and / or B cells that recognise a neoantigen. The immune cells may be isolated from tumours or other tissues (including but not limited to lymph node, blood or ascites), expanded ex vivo or in vitro and re - administered to a subj ect (a process referred to as "adoptive cell therapy" ). Instead or in addition to this, T cells can be isolated from a subj ect and008613259
[0183] engineered to target a neoantigen (e.g. by insertion of a chimeric antigen receptor that binds to the neoantigen) and re - administered to the subj ect. As another example, a therapeutic antibody may be an antibody which recognises a neoantigen.
[0184] As used herein "immune checkpoint blockade" or " ICB", also known as "immune checkpoint inhibitor" or " ICI" refers to an agent that blocks inhibitory checkpoint receptors of the immune system. Blockade of negative feedback signaling to immune cells thus results in an enhanced immune response against tumours. Examples of ICB agents (such as monoclonal antibodies) include PD- 1 inhibitors, PD-L1 inhibitors, CTLA- 4 inhibitors, LAG- 3 inhibitors, and combinations thereof. In particular, an ICB agent may be selected from: ipilimumab, tremelimumab, pembrolizumab, nivolumab, cemiplimab, atezolizumab, avelumab, durvalumab and relatlimab.
[0185] " Patient" as used herein in accordance with any aspect of the present invention is intended to be equivalent to "subj ect" and specifically includes both healthy individuals and individuals having a disease or disorder (e.g. a proliferative disorder such as a cancer). Preferably, the patient is a human patient. In some cases, the patient is a human patient who has been diagnosed with, is suspected of having or has been classified as at risk of developing, a cancer. In particular, a colorectal cancer (CRC), such as a CRC that is MMRd.
[0186] As used herein "neoplasm" refers to any abnormal or excessive growth of tissue, e.g. forming a benign or malignant tumour.
[0187] As used herein "cancer" refers to any malignant neoplasm. The cancer may be colorectal cancer (CRC), prostate cancer, ovarian cancer, breast cancer, endometrial cancer (uterus / womb008613259
[0188] cancer), kidney cancer (renal cell), lung cancer (small cell, non- small cell and mesothelioma), brain cancer (glioma, astrocytoma, glioblastoma (GBM) ), melanoma, Merkel cell carcinoma, clear cell renal cell carcinoma (ccRCC), lymphoma, gastrointestinal cancer, small bowel cancer (duodenal and j ejunal), leukaemia, pancreatic cancer, hepatobiliary tumours, liver cancer (e.g. hepatocellular carcinoma), germ cell cancer, head and neck cancer, bladder cancer, thyroid cancer, oesophageal cancer (OAC), melanoma (e.g. uveal melanoma), cutaneous squamous cell carcinoma or sarcoma. In particular, the cancer may be CRC. The cancer may be MMRd, for example, MMRd CRC. The inventors have demonstrated that their methods perform well in colorectal cancer. They contemplate that their methods may also be applied to other types of cancer, in particular cancers commonly associated with microsatellite instability (MSI). Accordingly, methods of this disclosure may also be applied to tumours selected from: a colon cancer, a gastric cancer, a small bowel cancer, rectal cancer, an endometrial cancer, an ovarian cancer, a biliary tract cancer, a urinary tract cancer, a brain cancer, and a skin cancer.
[0189] The inventors ' loci may be particularly suited to characterising or classifying tumours known to be MMRd or known to exhibit microsatellite instability (MSI). The use of ecological diversity metrics such as the Shannon Diversity Index (SDI) and Jensen - Shannon distance (JSD) may advantageously allow for a continuous measurement of tumour mutation rates, which may be indicative of tumour evolvability. This may be particularly the case when an aggregate of ecological diversity metrics is measured across multiple of the inventors ' identified loci (Table 1). The inventors have shown that the level of homopolymer erosion at their identified loci is driven by and therefore indicative of MSH3 - induced and MSH6 - induced increases in genome-wide008613259
[0190] mutation rates (i. e. higher erosion at these loci can be used as a proxy for genome-wide mutation rate and evolvability of a tumour sample) (see Examples 1 - 2). Hence, an estimation of the genome -wide mutation rate may be determined by using diversity metrics, such as SDI, as a read-out of the response of the inventors ' identified loci to the genome-wide mutation rate. Consequently, mutation rate - inf luenced properties may be determined using diversity metrics measured for the inventors ' identified loci (see Examples 3 - 5).
[0191] As used herein, a "coding" or "exonic" erosion - sensitive locus may refer to an erosion - sensitive locus that is located at least partially within an exon. A "coding" or "exonic" erosion - sensitive locus may be an erosion - sensitive locus that is located entirely within an exon. Even in the absence of purifying selection, exons incur fewer mutations than expected in several tumour types due to the MMR machinery prioritizing mutation repair of coding over non - coding regions. Enrichment for histone mark H3 Lysine 36 tri -methylation (H3K36me3 ) drives enhanced MMR activity in exonic regions. Consequently, coding microsatellites in MMR-prof icient samples may show relatively narrow length distributions compared to MMR deficient samples. Likewise, coding microsatellites may provide a sensitive readout to differentiate between MMRd samples, such as MMRd samples showing higher genomic mutation rate versus MMRd samples showing lower genomic mutation rate. Accordingly, measuring diversity at coding loci may advantageously offer insight into the dynamic interplay between mutational processes and selection pressures within tumour subclones, rather than relying on the assumption of neutrality for non - coding homopolymer loci.
[0192] Multi - region whole exome sequencing (WES) has confirmed that coding microsatellites in MLH1 - def icient / MSH6 - def icient clones008613259
[0193] undergo faster erosion compared to the same sites in MLH1 -def icient / MSH6 -prof icient ancestor lineages. The inventors therefore hypothesised that exonic homopolymer loci may offer a sensitive readout for genome-wide mutation rates in MMRd samples, in particular the differentiation of MSH- 6 and / or MSH- 3 deficient samples versus MSH- 6 and / or MSH- 3 proficient samples. Consistent with this, the inventors have shown that their identified coding microsatellites provide a sensitive readout of MSH- 6 and / or MSH- 3 status, particularly coding homopolymer loci with approximately 5 - 11 bp repeat length, e.g. 6 - 10 bp repeat length (see group D versus group A-C comparisons, Examples 1 - 2).
[0194] As used herein, " MSH3 -prof icient" may refer to a sample, cell or tumour with wild- type or near wild- type MSH3 function. As used herein, " MSH6 -prof icient" may refer to a sample, cell or tumour with wild- type or near wild- type MSH6 function. As used herein, " MSH3 - def icient" may refer to a sample, cell or tumour with deficient MSH3 function. For example, " MSH3 - def icient" may refer to a sample, cell or tumour in which the MSH3 gene is non - functional or significantly impaired. " MSH3 - def icient" may include, for example, samples, cells or tumours wherein presence of mutations in the MSH3 gene result in loss of function, such as frameshift, nonsense, or missense mutations. A sample, cell or tumour may be considered MSH3 - def icient where one or more frameshift mutations in coding microsatellites in the MSH- 3 gene leads to MSH- 3 inactivation, such as MSH- 3K383fs. " MSH3 - def icient" may additionally and / or alternatively include samples, cells or tumours with reduced or absent MSH3 protein expression compared to wild- type, for example determined using immunohistochemistry. " MSH3 -deficient" may additionally and / or alternatively include samples, cells or tumours with one or more specific mutations008613259
[0195] in the MSH3 gene known to confer impaired DNA mismatch repair activity, such as MSH- 3K383fs.
[0196] " MSH6 - def icient" may include, for example, samples, cells or tumours wherein presence of mutations in the MSH6 gene result in loss of function, such as frameshift, nonsense, or missense mutations. A sample, cell or tumour may be considered MSH3 -deficient where one or more frameshift mutations in coding microsatellites in the MSH- 6 gene leads to MSH- 6 inactivation, such as MSH - 6F1088fs. " MSH6 - def icient" may additionally and / or alternatively include samples, cells or tumours with reduced or absent MSH6 protein expression compared to wild- type, for example determined using immunohistochemistry. " MSH6 -deficient" may additionally and / or alternatively include samples, cells or tumours with one or more specific mutations in the MSH6 gene known to confer impaired DNA mismatch repair activity, such as MSH - 6F1088fs.
[0197] " MSH3 -prof icient" and " MSH6 -prof icient" may refer to at least 50% gene function of MSH3 and MSH6 respectively (indicative of absence of biallelic inactivation). " MSH3 - def icient" and " MSH6 - def icient" may refer to <50% gene function (indicative of biallelic mutation and / or inactivation) or approximately 0% gene function (indicative of biallelic inactivation). For example, proficiency and deficiency may be determined by comparing the percentage protein expression of MSH3 and / or MSH6 to that of a wild- type reference sample using immunohistochemistry (e.g. by comparing staining intensity). " MSH3 -prof icient" and " MSH6 -prof icient" may refer to 50% or fewer of cells in a sample harbouring loss - of - function or inactivating mutations in the MSH- 3 and MSH- 6 genes, respectively. " MSH3 - def icient" and " MSH6 - def icient" may refer to >50% of cells in a sample harbouring loss - of - function or inactivating mutations in the MSH- 3 and MSH- 6 genes,008613259
[0198] respectively. The percentage of cells harbouring loss -of -function or inactivating mutations in the MSH- 3 and MSH- 6 genes may be determined by immunohistochemistry.
[0199] A sample, cell or tumour may be determined to be MSH3 -proficient or -deficient, or MSH6 -prof icient or deficient, by calculating a deficiency score. The deficiency score may provide a continuous measure of MSH- 3 and / or MSH- 6 deficiency. A weighted sums approach may be used. Each normalised microsatellite repeat length relative to wild type may be assigned a weight using a binary gene expression on / off system. A deficiency score may be calculated for MSH6 and / or MSH3 loci for each sample. Lengths corresponding to out -of -frame frameshifts (i. e. +1, +2, - 1, - 2) may assigned a nonzero weight, e.g. positive weight, e.g. 1. Lengths corresponding to wild- type (0) and within- frame frameshifts (+3, - 3 ) may be assigned a weight of 0 (as they have no or reduced effect on gene expression). For MSH6 and MSH3 in each sample, cell or tumour, normalised read depths for each length may be multiplied by the corresponding weight and summed to yield a genetic deficiency score.
[0200] A separate deficiency score may be calculated for MSH6 and MSH3 loci for each sample. The MSH6 score and MSH3 score may be added together to derive an overall deficiency score for each sample, and this combined score may be used to group samples as either " Proficient" or " Deficient", e.g. " MSH3 / MSH6 proficient" or " MSH3 / MSH6 deficient".
[0201] A deficiency score may be thresholded to determine whether a sample, cell or tumour is proficient or deficient for the gene (s) to which the score relates (e.g. MSH3, MSH6,
[0202] MSH3 / MSH6 ). This may be based on a pre-defined deficiency score threshold, for example defined by taking the median008613259
[0203] value of all genetic deficiency scores across a cohort. The cohort may vary depending on the specific application. For example, when classifying tumours as MMRd vs not MMRd, the cohort may comprise samples or tumours known not to be MMRd and / or samples or tumours known to be MMRd. For example, when classifying tumours known to be MMRd by mutation rate or mutation rate - inf luence property, the cohort may comprise samples or tumours known to be MMRd.
[0204] As used herein, " MSH3 -prof icient / MSH6 -prof icient" may refer to tumours or samples known or determined to be MSH3 -proficient and / or MSH6 -prof icient. " MSH3 -prof icient / MSH6 -proficient" alternatively and interchangeably be referred to as " MSH3 / MSH6 proficient". MSH3 -prof icient / MSH6 -prof icient tumours or samples may be MSH3 -prof icient only, MSH6 -proficient only, or MSH3 -prof icient and MSH6 -prof icient. A tumour or sample may additionally or alternatively be determined to be MSH3 -prof icient / MSH6 -prof icient based on a combined score of MSH3 function or mutational status and MSH6 function or mutational status (see Examples, Genetic deficiency scoring).
[0205] As used herein, " MSH3 - def icient / MSH6 - def icient" may refer to tumours or samples known or determined to be MSH3 - def icient and / or MSH6 - def icient. MSH3 - def icient / MSH6 - def icient" alternatively and interchangeably be referred to as " MSH3 / MSH6 deficient". MSH3 - def icient / MSH6 - def icient tumours or samples may be MSH3 - def icient only, MSH6 - def icient only, or MSH3 -deficient and MSH6 - def icient. A tumour or sample may additionally or alternatively be determined to be MSH3 -deficient / MSH6 - def icient based on a combined score of MSH3 function or mutational status and MSH6 function or mutational status (see Examples, Genetic deficiency scoring).008613259
[0206] As used herein, the terms "computer system" includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the embodiments described herein. For example, a computer system may comprise a central processing unit (CPU) and / or a graphic processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. A computer system may comprise a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other non - transitory computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.
[0207] As used herein, the term "computer readable media" includes, without limitation, any non - transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.
[0208] Embodiments of the present disclosure make use of machine learning models. A machine learning model may be a mathematical model trained (i. e. parameterised) to classify observations between a plurality of classes, i. e. classifiers (also referred to as "classification model" ). A classification model may provide as output a classification label and / or one or more probabilities of an observation belonging to008613259
[0209] respective one or more classes. A binary classification model may provide as output a single probability or score indicating the probability that the observation belongs to a positive class (e.g. a class with a poorer prognosis, rather than a class with a better prognosis). A predetermined threshold may be applied to the one or more probabilities or the score, to assign a class label. The threshold may be determined based on desired characteristics of the classification, such as e.g. a desired level of specificity, sensitivity or any accuracy performance that combines aspects of specificity (precision) and sensitivity (recall), such as accuracy and Fl score or balanced versions thereof that take into account the proportions of observations in the training data in each of the classes. Alternatively, an observation may simply be assigned to the class with highest probability. When using a binary classifier this may be equivalent to assigning the class for which the probability is over 50%. Embodiments of the present disclosure make use of one or more binary classifiers. The one or more binary classifiers may each output a score or probability. Any machine learning classifier known in the art may be used in the context of the present disclosure. The term "machine learning algorithm" or "machine learning method" refers to an algorithm or method that trains and / or deploys a machine learning model. Machine learning models of embodiments of the present disclosure are trained by supervised learning, i. e. using training data comprising observations for each of set of predictive variables for each of a plurality of genes, and a ground truth label in the form of a class or predicted variable. An optimality criterion may be the minimisation of a loss function that quantifies the model prediction error based on the observed (ground truth) and predicted values of the predicted variables. Suitable loss functions for use in training machine learning models are known in the art and include the mean squared error, the mean008613259
[0210] absolute error, and regularised versions thereof. Any of these can be used according to the present disclosure. Regularised loss functions are functions that include a loss function as described above, and one or more terms penalizing model complexity in order to reduce the risk of overfitting.
[0211] Overfitting is a phenomenon that occurs where a machine learning model is trained to very closely reproduce the features of a training data set, resulting in poorer performance on other datasets that do not have the same characteristics (i. e. poor generalizability). Ll regularisation (also known as " Lasso" in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of absolute value of the coefficients of the model. L2 regularisation (also known as " Ridge" in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of squared value of the coefficients of the model. Ll regularisation can be used as a feature selection method as it minimizes the coefficients associated with less informative predictive features. In embodiments, the machine learning model is a regularized model. In embodiments, the machine learning model is a regularized tree-based model. In embodiments, the machine learning model is a regularized gradient boosted decision tree model. Examples of such models are available in the XGBoost software library (xgboos t. ai / ). Such models may be referred to as " XGBoost" models, although any other implementation of regularized gradient boosted models may equally be used. In embodiments, a machine learning model comprises an ensemble of models whose predictions are combined. Alternatively, a machine learning model may comprise a single model. Random forest models and gradient boosted tree models (such as XGBoost) are ensemble models. Ensemble versions of any models can be constructed. A machine learning model may be a mathematical model trained (i. e. parameterised)008613259
[0212] to predict the values of one of more predicted variables, i. e. regression models (also referred to as "regressors" ).
[0213] Regression models aim to predict the value of a continuous variable associated with observations. They are typically suitable when the output to be predicted is a continuous value in the training data. A machine learning model as described herein may be selected from: decision trees and variants thereof including regularized and / or gradient boosted decision trees and random forest models, regularized discriminant analysis, logistic regression models, generalized linear models, artificial neural networks (ANNs) including multilayer perceptrons (with linear or non- linear activation functions) and deep learning models (e.g. long short- term memory networks (LSTMs), Recurrent neural networks (RNNs) ), naive Bayes classifiers, and support vector machines (SVM, using linear or non- linear kernels such as radial basis function). Non- linear models, such as decision trees and variants thereof (including in particular random forests and gradient boosted trees), SVM with a non- linear kernel, and ANNs (e.g. multilayer perceptrons) with non- linear activation functions may be particularly performant.
[0214] As depicted schematically in Fig. 2, a system comprises a computing device 1, which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network, to sequence data acquisition means 3, such as a sequencer and / or to one or more databases 2 storing sequence data. The one or more databases 2 may further store one or more of: parameters of a trained008613259
[0215] machine learning model, patient and / or sample related information, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network 6 such as e.g. over the public internet. The sequence data acquisition means may be in wired connection with the computing device 1, or may be able to communicate through a wireless connection, such as e.g. through WiFi and / or over the public internet, as illustrated. The connection between the computing device 1 and the sequence data acquisition means 3 and / or databases 2 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 3 are configured to acquire sequence data from samples, for example tumour samples. This information may be used in embodiments of the disclosure to measure diversity (e.g. SDI) at a plurality of the loci set forth in Table 1 and to derive from the measured diversity an estimate of the mutation rate, intra - tumoral heterogeneity, sub - clonal complexity, disease prognosis, treatment response prediction and / or to identify a treatment strategy for a subj ect.
[0216] The methods described herein are computer implemented unless context indicates otherwise. Indeed, genomic data analysis, and the process of training machine learning models is of a complexity, and in particular requires the analysis of large008613259
[0217] amounts of data through complex mathematics, that places the methods described herein far beyond the capability of mental investigation. Thus, any method described herein may be implemented in a computer system (in particular in computer hardware or in computer software) in addition to the structural components and user interactions described. The term "computer system" includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above - described embodiments. For example, a computer system may comprise one or more processing units such as central processing units (CPU) and / or graphics processing units (GPU), input means, output means and data storage. Preferably the computer system has a monitor to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that the computer system may consist of or comprise a cloud computer. The methods described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method (s) described herein.
[0218] Examples
[0219] In the work described in the following examples, the inventors implement high- depth targeted re- sequencing of a key set of microsatellite loci across two independent cohorts to profile differences in microsatellite length distribution between multiple tumour regions. Quantification of underlying the MSH3 / 6 status allowed erosion rates of microsatellite loci to be compared between tumour areas of low mutation rate and008613259
[0220] tumour areas of high mutation rate. Estimates made using the panel developed by the inventors were correlated to in vivo mutation rates mathematically modelled using whole-genome sequencing (WGS). Aspects of the correlation are further described in Kayhanian et al., 2024 where MSH3 / MSH6 mutated clones were found to having higher mutation rates compared to MSH6 / MSH3 wild- type (as calculated using the MOBSTER tool). In the present work the correlation between the estimates made using the panel and the in vivo mutation rates led to the derivation of a small panel of homopolymers capable of acting as a sensitive proxy read-out for sub - clonal mutation rate. The inventors believe that their panel may be applied to MMRd cancers in the prediction of patient response to immunotherapy treatment, prediction of treatment - naive MMRd colorectal cancer (CRC) prognosis, risk prediction for familial CRC cases and measurement of temporal tumour evolution.
[0221] Example 1: Panel Design
[0222] In this example, the inventors identified a set of coding microsatellites exhibiting pronounced differences in length distribution between MSH6 mutant tumour regions and MSH6 wildtype tumour regions. Reasoning that these microsatellites may act as a proxy for MMRd status, they designed a targeted re - sequencing panel including their identified loci for further investigation.
[0223] Methods
[0224] Whole exome sequencing (WES) - pre-processing. Multi - region whole exome sequencing was performed on 22 treatment - naive colorectal cancer resection FFPE specimens.
[0225] Immunohistochemically- determined MSH6 -prof icient and MSH6 -deficient clones were isolated using laser capture microdissection (LCM) from sequential 10pm PEN-membrane slides008613259
[0226] with 46 individual clones included in total. DNA was extracted using the Chemagic Prepito automated instrument and sequenced at 50x target coverage. FastQ sequencing files were aligned to the Hgl9 reference genome using BWA-mem. Aligned sequencing files were converted to BAM files followed by sorting and indexing of reads using Samtools. PicardTools was used to mark duplicates and GATK was used for local indel realignment.
[0227] PicardTools, GATK and FastQC were used to produce quality control metrics. Of the 46 clones, 7 were MSH6 / MSH3 -wildtype (proficient), 17 were MSH6 -mutated, 9 were MSH3 -mutated, and 13 were MSH3 / MSH6 -mutated. For each patient, a normal proximal / dis tai resection margin was also sequenced as a germline reference.
[0228] Whole exome sequencing (WES) - downstream steps. Analysis -ready matched normal and tumour BAM files were processed using MSIsensor to quantify microsatellite instability and report somatic status of corresponding positions in the genome.
[0229] Default parameters were used, except the setting of microsatellite minimum and maximum lengths. MSIsensor was used to output a read count distribution table for each tumour sample, which provides information on the number of alleles supporting alternative microsatellite repeat lengths. Clonespecific mutation rates were estimated using MOBSTER (Caravagna et al., 2020). MOBSTER calculates the tail of newly acquired passenger mutations in an expanding population using the 1 / f power- law distribution, incorporating the variant allele frequency (VAF) values for each mutation
[0230] detected during sequencing. SciRoKoCo was used as an additional filtering step to curate a list of microsatellites as a starting point for filtering. SciRoKoCo is a bioinformatics tool that scans a genome and provides a complete list of the corresponding microsatellites. This list of microsatellites was compared to a list of coding008613259
[0231] microsatellites and used to filter the WES MSIsensor read depth distribution tables to focus on microsatellites of interest.
[0232] Shannon diversity analysis. To extract microsatellites most responsive to changes in mutation rate from read count distribution tables output by MSIsensor, Shannon diversity indices (SDIs) were derived using the R package vegan (version 2.6 - 8) and manually cross -validated for each coding homopolymer in the exome using the following formula:
[0233] R
[0234] Shannon diversity = — piln(pi)
[0235]
[0236] i=l
[0237] Where pi = the proportion of total reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite. A higher SDI indicates a higher diversity of species in a particular community (in this context, a higher diversity in microsatellite repeat length for a given microsatellite). An SDI value of 0 indicates a community that only has one species (in this context, a microsatellite exhibiting only a single repeat length).
[0238] Based on WES data, samples were sorted based on their underlying MSH6 / MSH3 mutation status into a MSH6 / MSH3 -wildtype group and a MSH6 / MSH3 -mutant group. Clones were originally microdissected according to MSH6 IHC (MSH3 IHC data was not available), then further stratified by the WES mutation output. The median SDI for each microsatellite across the all samples in the MSH6 / MSH3 -wildtype group and across all samples in the MSH6 / MSH3 -mutated group was taken. All microsatellites were ranked based on their change in Shannon diversity indices between MSH6 / MSH3 -wildtype and MSH6 / MSH3 -mutant samples (i. e. tumour areas of low and high mutation rates, respectively). A008613259
[0239] subset of microsatellites (approx. 300) showed distinctly larger shifts in diversity (tail end of the distribution) between MSH6 / MSH3 -wildtype and MSH6 / MSH3 -mutant samples. These were selected for inclusion in the targeted re - sequencing panel (herein referred to as group D) (Fig. 3A).
[0240] Panel positions. The panel design included four groups, A-D (Fig. 3A, Table 2). Group A included 50 non - coding microsatellite loci located in 5 ' UTRs or 3 ' UTRs. Group B comprised 24 non- coding microsatellite loci relevant to clinical cancer diagnostics identified by Gallon et al.
[0241] (2019 ). Group C comprised 53 recurrently mutated coding microsatellite loci identified by Cortes - Ciriano et al.
[0242] (2017). Group D comprised the 300 coding microsatellites (responsive loci) identified by the inventors. Groups A-C were included for comparison with the identified responsive loci (group D).
[0243] Table 2. Genomic position, homopolymer content and length of microsatellites in groups A-D.
[0244] A B C D
[0245] Genomic Intronic Intronic Exonic Exonic position
[0246] Homopolymer A (54%), C A (96%), C A (45%), C A (36%), C content (2%), G (2%), (4%), G ( 0%), (4%), G ( 0%), (15%), G (transcribed T (42%) T ( 0%) T (51%), (16%), T strand) (32%) Length 9 - 11 bp 7 - 12 bp 7 - 23 bp 6 - 10 bp (maj ority 10 (maj ority >12
[0247] bp) bp)
[0248]
[0249] Results
[0250] MSIsensor was employed on matched tumour and normal samples to quantify microsatellite instability and somatic status of008613259
[0251] corresponding positions in the genome. The inventors used the SciRoKoCo tool to scan the hgl9 reference genome and curate a list of all microsatellites present. This was matched against an exonic bed file to produce a list of all coding microsatellites in the genome. This list was used as an additional filtering step for MSIsensor.
[0252] Samples were grouped based on their underlying MMR (specifically, MSH6 / MSH3 ) mutation status, as confirmed through WES data. Shannon diversity indices were derived for each coding homopolymer in the exome to extract microsatellites most responsive to changes in mutation rate, using MSIsensor output read depth distribution tables.
[0253] Microsatellites were ranked based on their change in Shannon diversity indices between MSH6 / MSH3 -wildtype and MSH6 / MSH3 -mutant samples. A subset of microsatellites showed distinctly larger shifts in diversity between MSH6 / MSH3 -wildtype and MSH6 / MSH3 -mutant samples, with the top -300 loci showing markedly high diversity shifts (Fig. 3A). The inventors subsequently reasoned that erosion rates at these loci may act as a proxy for mismatch repair deficiency. The top 300 ranked microsatellites showing the largest increase in SDI with MSH3 / MSH6 loss were therefore selected for inclusion in a targeted re - sequencing panel to test this hypothesis. These microsatellites are herein referred to as group D loci or "responsive loci".
[0254] A panel for targeted re - sequencing of microsatellite regions was designed based on the GRCh37 genome build (Fig. 3A). The panel included four groups (A-D), covering a total of 70 kb of the GRCh37 reference genome. Groups A-C were included for comparison with the identified responsive loci (group D).
[0255] Group A formed 9% of the panel and comprised 50 non - coding microsatellite loci located in 5 ' UTRs or 3 ' UTRs. Group B008613259
[0256] formed 4% of the panel and comprised 24 non - coding cancer diagnostic microsatellite loci. Group C formed 10% of the panel and comprised 53 coding recurrently-mutated microsatellites. Group D represented the novel set of 300 coding microsatellite loci identified by the inventors, forming 77% of the panel.
[0257] Example 2: Pilot Sequencing Round
[0258] Methods
[0259] Sample preparation and sequencing. All University College London Hospitals MMRd / MSI - H (high microsatellite instability) colorectal cancer cases diagnosed from 2017 - 2021 (n=46) were screened. Tumour resection specimens were sectioned and stained for MSH6 using immunohistochemistry. Those that showed evidence of subclonal loss of MSH6 were further analysed.
[0260] Prior to sectioning, PEN membrane slides (Zeiss) were pretreated with poly- 1 - lysine and irradiated with UV light to overcome the hydrophobic nature of the membrane. A total of 13 consecutive sections were cut per tumour region (i. e. per FFPE block), with partitioning 3pm MSH6 IHC slides used to guide accurate sampling of thicker 10pm LCM slides. All slides were baked at 60°C for 1 hr to increase tissue adherence, deparaf f inised using xylene and stained manually with haematoxylin using the following steps: xylene (15 min), xylene (15 min), 100% ethanol (2 min), 100% ethanol (2 min), 90% ethanol (2 min), dH2O (rinse), Gill ' s haematoxylin (1 min), dH2O (rinse). To bypass laser capture microdissection (LCM) and increase sampling efficiency, the inventors developed a biopsy punch method to isolate multi - region FFPE patches. Haematoxylin staining of LCM slides allowed identification of tissue morphology and tumour context.008613259
[0261] WGS-matched cohort (n=21). Cases with sufficient tissue availability for sectioning and exhibiting obvious (approx. 2 -3mm in diameter) MSH6 -prof icient and MSH6 - def icient subclones on scanned slides were amenable to WGS analysis. For the WGS -matched cohort, specific MSH6 -prof icient and MSH6 - def icient clones were mapped and dissected using IHC slides. Two or more punches were taken from within each clone, totalling 21 multi -region samples across 5 patients and matched whole -genome sequencing (WGS) and mutation rate data was obtained.
[0262] WGS-unmatched cohort (n=56). The inventors aimed to validate their results by randomly sampling tumours without using IHC as a reference. For cases that were not suitable for WGS analysis (i. e. those with smaller focal losses and / or embedded reversions), multi - region sampling was performed blind to the underlying MSH6 status. A total of 56 multi - region samples were collected across an additional 14 patients. For these cases, matched WGS and mutation rate data were not available. These cases were included to understand if the approach could be applied to a more clinically relevant situation, where IHC data may not be available as a reference prior to sequencing and samples may be taken across multiple clones (e.g. taking tumour biopsies in the clinic).
[0263] Normal reference samples (n=19). One normal reference sample per patient in the WGS -matched and WGS -unmatched cohorts was derived using normal colonic crypts from distal and proximal resection margins. DNA extraction was carried out using an optimised version of the HighPure FFPET kit (Roche). Samples were sequenced at lOOOx target coverage.
[0264] All tumour - derived and normal samples were sequenced using the targeted re - sequencing panel developed in the work of Example008613259
[0265] 1. Of the 300 original group D loci, 274 were taken forward for analysis (listed in Table 1), due to sites dropping out during the sequencing process because of lower probe efficiency during in hybridisation and / or bias during PCR amplification.
[0266] Data pre-processing. Adapter trimming of targeted resequencing reads was completed using the FastP package prior to alignment to the Hg38 reference genome using BWA-MEM. Postalignment QC and processing was carried out using Picard and GATK. Raw read depths were plotted using the pheatmap package in R to visualise differences in cumulative read counts across all sample and panel groups (Fig. 3D).
[0267] MSlsensor was employed using the same pipeline and parameters as for the WES dataset in Example 1 (default parameters except microsatellite minimum and maximum length), with output read count distribution tables used for downstream analyses. The panel bed file was converted from the Hgl9 to Hg38 genome reference assembly using the UCSC LiftOver Tool. Data was normalised relative to row sum.
[0268] Density line fitting. Density lines were plotted for each microsatellite to visualise microsatellite length distribution, using the workflow shown in Fig. 3E. The MSlsensor output table contains a row for each microsatellite, with a columns corresponding to microsatellite length. This includes both the number of reads for wild- type length, as well as the number of reads for alternative lengths (i. e. changes from wild- type). Microsatellite read depth distribution vectors were derived from the MSlsensor output read count distribution tables, by converting each row into a vector and normalising relative to row sum.008613259
[0269] Equally- spaced bins were created between the minimum and maximum values of each microsatellite read depth distribution vector (i. e. the minimum and maximum read depth values for a given microsatellite, corresponding to either the wild- type length or an alternative length). Probability density functions (PDFs) for a Gaussian distribution based on the mean and standard deviation of the read depth distribution vector were calculated. Density lines were scaled to match area under the histogram of microsatellite length, with probability density along the y axis.
[0270] Genetic deficiency scoring. To incorporate the WGS - unmatched cases and MSH3 frameshift mutation status into the analysis, an approach to derive genetic deficiency scores was developed, based on MSH6 and MSH3 instability within samples. A weighted sums approach was used, where each normalised microsatellite repeat length relative to wild type was assigned a weight using a binary gene expression on / off system (Fig. 3G). A separate deficiency score was first calculated for MSH6 and MSH3 loci for each sample. Lengths corresponding to out -of -frame frameshifts (i. e. +1, +2, - 1, - 2) were assigned a weight of 1. Lengths corresponding to wild- type (0) and within- frame frameshifts (+3, - 3 ) were assigned a weight of 0 (as they have no or reduced effect on gene expression). For MSH6 and MSH3 in each sample, normalised read depths for each length were multiplied by the corresponding weight and summed to yield a genetic deficiency score.
[0271] The MSH6 score and MSH3 score were then added together to derive an overall deficiency score for each sample, and this combined score was used to group samples as either " Proficient" or " Deficient" based on a pre-defined deficiency score threshold, defined by taking the median value of all genetic deficiency scores across the cohort. The method was008613259
[0272] validated by comparing MSH6 deficiency scores to IHC staining information. A median SDI value per microsatellite position was taken across samples labelled as either Proficient or Deficient.
[0273] Shannon diversity analysis. Normalised read depth distribution tables were used to derive Shannon diversity indices for each homopolymer position in the panel (see Example 1 Methods and Fig. 4A). A median Shannon diversity index was taken for each homopolymer position across samples labelled as either proficient or deficient using the genetic deficiency scoring method. This calculation was done using the vegan (version 2.6 - 8) package in R and manually for cross validation.
[0274] Jensen-Shannon Distance (JSD) and Phylogenetic Reconstruction. For each tumour sample, Jensen - Shannon distances were calculated for each microsatellite between the probability distribution of microsatellite lengths in the tumour sample and the corresponding probability distribution in the normal reference genome using the following formula:
[0275] JSD(P || Q) = |r>(P || M) + |r>(Q || M)
[0276]
[0277] where P and Q are the probability distributions of microsatellite lengths in tumour and normal samples, M is the average distribution defined as M = — — and D X II K) is the
[0278]
[0279] Kullback - Leibler divergence between X and Y.
[0280] Leveraging the multi - region sampling method, JSDs were also used to reconstruct phylogenies within the cohort. MSIsensor read depth distribution tables were normalised, and Jensen-Shannon distances (JSDs) calculated for each tumour sample in the cohort relative to its matched normal reference across the008613259
[0281] top 100 microsatellites in group D. Using these pairwise distances, a distance matrix was generated among each pair of samples as per-patient heatmaps using the pheatmap package in R, enabling wi thin-patient genetic divergence between tumour and normal samples to be assessed (Fig. 40). To reconstruct phylogenetic trees, a multi - region sampling approach was used alongside per-patient distance matrices, applying the classical neighbour - j oining tree method implemented in the ape package in R (Fig. 4P).
[0282] For a direct comparison of JSDs between Proficient and Deficient groups, the median USD (relative to the normal reference) across all top 100 microsatellites in group D was calculated for both groups (Figs. 4Q-R).
[0283] Results
[0284] Samples from 46 resection specimens of colorectal cancer were screened, sectioned and stained for MSH6 using immunohistochemistry. Those showing evidence of subclonal loss were further analysed. For cases suitable for WGS analysis, 21 multi - region samples were taken across 5 patients (WGS -matched cohort) (Fig. 3C). For cases that were not suitable for WGS analysis, 56 multi - region samples were collected across an additional 14 patients (WGS - unmatched cohort). One normal reference sample per patient was also derived (n=19).
[0285] Samples were sequenced using the targeted re - sequencing panel developed in the work of Example 1, and MSIsensor output distribution plots were derived from the targeted panel sequencing. Raw read depths were plotted to visualise differences in cumulative read counts across all sample and panel groups (Fig. 3D). Initial quality checks showed that although there was variation in read depths between008613259
[0286] microsatellite loci, this was consistent between panel and sample groups.
[0287] As IHC information was not available for MSH3, it was important to ensure that secondary inactivation of MSH3 did not confound downstream analyses, as exemplified by the data shown in Fig. 3F. Fig. 3F shows microsatellite length distribution plots for MSH6 and MSH3 in an MSH6 -prof icient and MSH6 - def icient clone. Frame shifts at MSH3 and MSH6 microsatellite loci cause secondary inactivation of these genes. The MSH6 - def icient clone harbours a 1 bp deletion in MSH6 (compared to the MSH6 -prof icient clone which is wildtype, i. e. length change 0), as expected. However, at the MSH3 locus, both the MSH6 -prof icient and MSH6 - def icient clone also contain 1 bp deletions in the A8 homopolymer (leading to secondary inactivation of MSH3 ). This highlighted the importance of developing a method to quantify MSH3 mutation status, as MSH3 IHC information was not available.
[0288] Genetic deficiency scores based on MSH6 and MSH3 instability within samples were derived based on a weighted sums method (see Methods and Fig. 3G). The method was validated by comparing MSH6 deficiency scores to MSH6 IHC staining information. MSH6 -prof icient and MSH6 - def icient clones grouped correctly based on their assigned genetic deficiency score.
[0289] Shannon diversity indices were derived for each homopolymer position in the panel (Fig. 4A). The Shannon Diversity index quantifies the diversity of a community, accounting for both the number of species (species richness) and the evenness of species abundances (species evenness). In this context, it allows the diversity of alleles at a particular locus within the population to be measured, considering both the number of different alleles (microsatellite lengths) and their relative008613259
[0290] frequencies. A median Shannon diversity index was taken for each homopolymer position across samples labelled as proficient and across samples labelled as deficient. These median SDIs were compared between the proficient and deficient groups for different subsets of microsatellite loci (Figs. 4B-K). Out of groups A, B, C and D (see Example 1), only group D loci showed significant differences in median SDI between the proficient and deficient groups (Figs. 4B-F). Microsatellites with length 7 in wild- type showed the most significant differences in median SDI between proficient and deficient groups, with microsatellites of length 8 (in wild- type) ranking next by p-value (Figs. 4G-K).
[0291] To further investigate the behaviour of homopolymer positions in the panel, a second ecologic diversity metric was incorporated. Jensen - Shannon distances (JSD) were derived for the length distribution of each microsatellite from their corresponding normal reference genome. JSD allows quantification of the similarity between two probability distributions (derived from the Kullback - Leibler divergence), as opposed to the Shannon diversity index, which measures the diversity within a single distribution.
[0292] Density lines were plotted for each microsatellite to visualise microsatellite length distribution (Fig. 3E and Fig.
[0293] 3H). As expected, microsatellite loci within normal samples centred around length 0 (i. e. wild- type, unmutated), whereas tumour samples showed shifts in microsatellite length distribution. Preferential contraction was observed.
[0294] To determine the optimal number of microsatellites to be included in downstream analyses (e.g. prediction of ICB response, cancer prognosis, Lynch syndrome progression), the effect of the number of group D loci included in the median008613259
[0295] Shannon diversity indices comparisons between Proficient and Deficient samples on the - loglO p-values of the Wilcoxon signed- rank statistic was determined (Fig. 4N).
[0296] Microsatellites were ranked based on their shift in SDI between Proficient and Deficient groups, then the top 10, top 20, etc. were taken to assess the effect on the p-value for the Proficient vs. Deficient comparison. Based on this analysis, the inventors chose to proceed with the top 100 ranked loci.
[0297] Example 3: Clinical Validation as a Predictive Biomarker
[0298] Methods
[0299] To test the targeted re - sequencing technologies ' potential as a predictive biomarker of immunotherapy response, analyses were expanded to 32 patients enrolled on the NEOPRISM clinical trial (NCT05197322). These included patients with radiological high- risk stage II or stage III colorectal cancer confirmed as MMRd using IHC or MSI -H testing by PCR. Those deemed as TMB-high through FM1 testing went on to receive three cycles of an immune checkpoint inhibitor, pembrolizumab, prior to surgery, with or without adjuvant FOLFOX or CAPOX. DNA was extracted from pre - treatment PAXgene biopsies using an optimised version of the PAXgene Tissue DNA kit (Qiagen) (n = 32). Normalised samples were sequenced at lOOOx target coverage.
[0300] Initial quality checks on paired FastQ sequencing files (R1 and R2) were carried out using the FastQC package. Adapter trimming was completed using the Fastp package prior to alignment to the Hg38 reference genome using BWA-MEM. Aligned FastQ files were converted to BAM files using SamTools and validated using Picard ValidateSamFile. GATK was used to derive a recalibration table for subsequent Base Quality Score008613259
[0301] Recalibration (BQSR). Post - alignment QC and assessment of sequencing metrics was carried out using Picard MarkDuplicates and CollectHsMetrics. MSlsensor -pro (vl.2) was used to quantify microsatellite instability and retrieve read count distribution tables from tumour-only samples.
[0302] For each patient, a median SDI will be derived across all or a subset of (i. e. the top 100 most responsive) group D microsatellites. Linear regression analyses will be undertaken to assess the relationship between median group D SDI and clinical outcomes regarding treatment response (e.g. % tumour bed regression). Additional analyses will be undertaken using an expanded (n = -1500), clinically- annotated genomic database of MMRd CRC samples to further validate the relationship between group D SDl / mutation rate and ICB response metrics.
[0303] Predicted Results
[0304] The inventors reason that it may be possible to predict response to immunotherapy using their framework. A pretreatment tumour biopsy may be collected in the clinic. Either through NGS (DNA or RNA sequencing), or a PCR-based technique, all or a subset of group D homopolymers (e.g. the top 100 most responsive loci) may be profiled, and SDIs may be calculated for each position. A median SDI may be calculated across all or a subset of loci for the sample. To predict ICB response using median SDI, a machine learning model may be trained, such as Random Forest or a Gradient Boosting model, using the SDI values as input features and immunotherapy response (e.g. indicated by percent tumour bed regression or classification as responders vs. non- responders) as the target variable.
[0305] Thresholds for grouping patients as either SDI - low or SDI -high could be gathered from evaluating the performance of the008613259
[0306] classifier using receiver operating characteristic (ROC) analyses, or using a percentile -based approach.
[0307] The inventors reason that higher Shannon diversity indices and JSDs will predict poorer or incomplete outcomes to pembrolizumab / lCB treatment (as determined through correlation with percentage primary tumour bed regression). This is due to higher Shannon diversity indices reflecting greater subclonal intratumour heterogeneity (ITH) (i. e. presence of genetically or phenotypically distinct subpopulations of cells) driven by increased mutation rates and evolvability within these lesions, indicative of MMRd. Despite elevated levels of frameshift mutation and neoantigen accumulation rendering MMRd tumours immunogenic and capable of eliciting a strong local immune response, research has shown that tumours exhibiting increased levels of neoantigen heterogeneity with high subclonal fraction (a high fraction of neoantigens that are present in only a subset of cells within the tumour) compromise efficacy of immune checkpoint blockade (ICB) treatment (McGranahan et al., 2016). This is due to inefficient recognition of tumour neoantigens by T cells because of their diluted frequency, termed "neo - antigen cloaking". Chronic subclonal neoantigen stimulation in this context may also drive T cell exhaustion, leading to decreased effector cytokine production and cytolytic activity (Westcott et al., 2021). In addition, higher mutation rates increase the likelihood of immune escape mutations occurring. These include alterations of the human leukocyte antigen (HLA) class I antigen presentation machinery (APM), either impacting HLA class I light chain or heavy chains directly, or indirectly through interference with antigen processing or formation of HLA class I complexes. Recurrent mutations in MMRd tumours include beta - 2 -microgloublin (B2M), TAPI, TAP2, calreticulin and ERp57 (Kloor et al., 2010).008613259
[0308] Example 4: Use as a Prognostic Biomarker
[0309] Background
[0310] The inventors hypothesize that the ecological diversity metrics, such as Shannon diversity metrics, that may be retrieved using their targeted resequencing approach may serve as a prognostic biomarker in treatment - naive MMRd colorectal cancer (CRC) patients. To test this, samples have been obtained from 227 patients enrolled in the SCOT clinical trial (NCT00749450). This is an international, randomised, phase III, non - inf eriori ty trial involving patients with high risk stage II or stage III colorectal cancer. The trial is aimed at investigating whether 3 months of oxaliplatin - containing chemotherapy is non- inferior to the usual 6 months of treatment.
[0311] Methods
[0312] DNA was extracted from 2x10pm curls from treatment - naive FFPE resection specimens. Samples will be sent for sequencing and the data will be analysed using the data pre-processing approach outlined in Example 2 (see Example 2 Methods).
[0313] Additional multivariate analyses will be conducted to adjust for potential confounders in the patient population.
[0314] A median SDI across all, or a subset of (i. e. the top 100 most responsive) group D microsatellites will be derived. A median group D SDI will be calculated across responsive loci on a per-patient basis. Cox proportional hazards regression analysis will be conducted to evaluate whether median Shannon diversity predicts poorer survival outcomes. The hazard ratio will be used to quantify the relative risk of an event (e.g.008613259
[0315] death or progression) associated with median SDI. Median SDI will be included as an independent variable in the model, and adjustments will be made for potential confounding factors such as age, disease stage, and treatment type.
[0316] Group D SDI will be analysed as both a continuous and categorical variable. In treating group D SDI as a categorical variable, the inventors will use one or more group D SDI thresholds to assign categories. The inventors will test separating patients / samples into categories into tertiles to give an SDI - low, SDI -medium and SDI -high category, with the fertile cut-points calculated using the 33rdand 66thpercentiles. The inventors will also test categorising patients / samples using the median group D SDI value across patients / samples as a cut-point, giving a 50 / 50 split between an SDI - low category and an SDI -high category.
[0317] Overall survival (OS), disease- free survival (DFS) and time-to - recurrence (TTR) will be analysed using Cox proportional hazards regression, with survival time measured from randomisation to death or last follow-up. Kaplan-Meier (KM) survival curves will be generated to visualise survival differences across SDI categorical groups, and log- rank tests will be performed for pairwise comparisons with Bonferroni correction for multiple testing. Multivariate Cox regression models will be fitted adjusting for age at registration, gender, planned treatment (Capox vs Folfox), trial arm duration (12 vs 24 weeks), pathological T stage (with T4 as reference), pathological N stage (with NO as reference), tumour differentiation, and tumour sidedness.
[0318] Predicted Results
[0319] The inventors reason that it may be possible to predict cancer prognosis using their framework. A tumour biopsy may be collected in the clinic. Either through NGS (DNA or RNA008613259
[0320] sequencing), or a PCR-based technique, all or a subset of group D homopolymers (e.g. the top 100 most responsive loci) may be profiled, and SDIs calculated for each position. A median SDI may be calculated across all or a subset of loci for the sample. For individual patient risk stratification, alternative analytical approaches may be employed. A machine learning model may be trained, such as Random Survival Forest (RSF), with SDI values used as input features and survival time (and censoring status) as the target variable. Model performance may be evaluated using concordance index (C-index). For clinical decision-making, thresholds for grouping patients as either SDI - low or SDI -high could be derived from ROC analyses, or using a percentile -based approach.
[0321] The inventors reason that higher Shannon diversity indices will predict poorer patient survival. This metric serves as a proxy measure for intra- tumour heterogeneity (ITH), or subclonal complexity, which is a key driver of tumour evolution. Elevated ITH in various tumour types, including lung, breast, and gastrointestinal cancers, has been linked to poorer survival and an increased risk of metastatic dissemination (Joung et al., 2017). More heterogeneous lesions (i. e. SDI -high), consisting of multiple distinct subclones, provide a substrate for ongoing clonal evolution and facilitate rapid selection of treatment - resis tant subclones. Subclonal neoantigen diversity in this context may also lead to immune exhaustion and dysfunction, compromising immune-mediated tumour control.
[0322] Example 5: Use as a Diagnostic Biomarker
[0323] Background008613259
[0324] Lynch syndrome (LS), previously known as hereditary nonpolyposis colorectal cancer (HNPCC), patients undergo regular colonoscopy surveillance to remove precursor adenomas in an effort to prevent early cancer progression. Despite this, a substantial number of individuals with Lynch syndrome go on to develop CRC, with a reported 10 -year risk of 6 - 11% of developing cancer between colonoscopies (Ahadova et al., 2016). In addition to this, there is significant variation in progression rate between unrelated Lynch families carrying similar germline mutations. Together, this highlights the need for more tailored surveillance strategies based on an individual ' s unique progression risk.
[0325] The inventors hypothesise that the ecological diversity metrics that may be retrieved using their targeted resequencing approach may serve as a diagnostic biomarker for risk stratification in hereditary cancer syndromes. To test this, polyps (specifically, tubular adenomas) resected during routine surveillance colonoscopies from 48 patients with Lynch syndrome have been obtained. For each patient, the St Mark ' s Centre for Familial Intestinal Cancer (SMCFIC) sample database was used to identify Lynch- associated low-grade and high-grade tubular adenomas coupled with clinical information including polyp size and underlying germline MMR mutation. Additional sporadic (i. e. MMR-prof icient / microsatellite stable) tubular adenomas were included to serve as a negative
[0326] control / ref erence group.
[0327] The inventors have shown that in MMRd cancers, tumour cell mutation rate increases with each additional MMR allele that is inactivated (Kayhanian et al., 2024). The inventors aim to investigate whether loss of a single germline MMR gene copy (i. e. 50% gene dosage reduction) may provoke mild microsatellite instability that may predict future cancer008613259
[0328] progression risk in Lynch patients to facilitate tailored surveillance strategies.
[0329] Methods
[0330] DNA will be extracted using an optimised version of the High Pure FFPET Kit (Roche). Samples will be sent for sequencing and data analysed using the data pre-processing approach outlined in Example 2 (see Example 2 Methods).
[0331] For each sample, a median SDI will be calculated across all or a subset (i. e. the top 100) of group D microsatellites to test whether sporadic (MMR-proficient), low-grade (i. e. at least 50% gene dosage, indicative of absence of biallelic inactivation) and high-grade (i. e. MMR- def icient) precursor adenomas can be distinguished. The inventors will also investigate if there are within-group (i. e. between low-grade adenomas, or between high -adenomas) differences in SDI, and run regression analyses to determine whether there is a relationship between median group D SDI and clinical outcomes (i. e. CRC onset or progression) and other clinical features such as polyp size. This could allow patients to be stratified based on their progression risk.
[0332] Additional immunohistochemical staining of the four key MMR genes (MLH1, PMS2, MSH2, MSH6) will be conducted on all samples (sporadic, low-grade and high-grade) to confirm MMR protein expression status. Identification of increased group D SDI and mutation rates in MMR-proficient (i. e. retained MMR protein expression, absence of biallelic inactivation) adenomas may indicate functional penetrance of loss of a single MMR allele (haploinsuf f iciency), and enable risk stratification of early, low-grade precursor lesions.
[0333] Predicted Results008613259
[0334] The inventors hypothesise that their framework may be useful in the prediction of Lynch syndrome precursor lesion progression. During routine surveillance, a polyp biopsy may be collected in the clinic. Either through NGS (DNA or RNA sequencing), or a PCR-based technique, all or a subset of the group D homopolymers (e.g. the top 100 most responsive loci) may be profiled, and SDIs calculated for each position. A median SDI may be calculated across all or a subset of loci for the sample. To predict cancer onset / polyp progression risk, a machine learning model, such as Random Forest or Gradient Boosting, may be trained, using median SDI values as input features and cancer progression (e.g. progression to malignancy or time to adenoma recurrence) as the target variable. Thresholds for grouping polyps and / or patients either SDI - low or SDI -high may be gathered from evaluating the performance of the classifier using ROC analyses, or using a percentile -based approach. The inventors hypothesise that the higher the median SDI, the greater the risk of cancer onset / polyp progression.
[0335] At present, monoallelic germline MMR mutation in Lynch patients is thought to be phenotypically silent, as the remaining wild- type allele safeguards normal DNA repair.
[0336] However, there are important data to suggest that loss of a single MMR allele can provoke context - specif ic DNA repair defects (Sansom et al., 2021). The inventors will compare genomic evolution between sporadic polyps without any MMR loss (100% gene dosage) and Lynch syndrome polyps with low grade dysplasia (50% gene dosage) and Lynch syndrome polyps with high grade dysplasia (complete MMR loss). The rationale behind this is that patients with familial adenomatous polyposis (FAP), another hereditary cancer syndrome caused by autosomal dominant mutations in the APC gene, accumulate hundreds to thousands of polyps in the colon and rectum from childhood008613259
[0337] with a median age of colorectal cancer diagnosis of approximately 45 years. Despite Lynch patients exhibiting a completely different clinical manifestation (characterised by low polyp burden until later life), the median age of CRC is also approximately 45 years. One key reason for this is likely rapid progression of these precursor lesions after normal colonoscopy, with retrospective analysis of the Dutch LS Registry demonstrating that 65% of LS patients who developed CRC surveillance had no adenomas detected at their previous colonoscopy (Argillander et al., 2018). This represents significant acceleration compared to the estimated 10 -year progression timeline in CRC cases.
[0338] As germline monoallelic MMR loss (50% gene dosage) occurs in each of the estimated 10 million crypts in the adult large bowel, even a marginal increase in mutation rates may increase likelihood of oncogenic driver mutations including APC mutation. Indeed, whereas APC mutations in sporadic adenomas generally occur in a hotspot called the mutation cluster region, APC mutations in MMR proficient (50% gene dosage) Lynch adenomas generally occur in a repetitive stretch outside the mutation cluster region (Sekine et al., 2017). Definitive evidence of haploinsuf f iciency (i. e. functional penetrance of 50% gene dosage) would open up new possibilities to forecast future polyp progression risk from lesions removed during surveillance.
[0339] Table 1. Loci of coding homopolymers that show ultrarapid evolution. Rank is based on p- value of Shannon Diversity at the locus as between MSH6 -prof icient and MSH6 - def icient CRC MMRd samples. Genome coordinates are as the GRCh38 human genome assembly ("hg38" ).008613259
[0340] Rank Chromosome Start End Length 1 chr3 123448367 123448373 7 2 chr7 105106537 105106543 7 3 chr2 17517469 17517475 7 4 chrll 62882056 62882063 8 5 chr2 97083987 97083993 7 6 chr6 99934481 99934489 9 7 chrl4 95196611 95196618 8 8 chr8 73595235 73595242 8 9 chrl9 14482827 14482834 8 10 chrl9 32676869 32676876 8 11 chr20 32434638 32434645 8 12 chrlO 96159098 96159106 9 13 chrl7 59169809 59169817 9 14 chrl4 60436846 60436853 8 15 chr2 47803500 47803507 8 16 chrl6 4812227 4812234 8 17 chrl4 73739069 73739075 7 18 chr5 146267756 146267764 9 19 chrl 1180842 1180849 8 20 chr8 112228799 112228805 7 21 chrlO 60064187 60064193 7 22 chrll 1456956 1456963 8 23 chr2 151251573 151251582 10 24 chrl 151346264 151346270 7 25 chrl 93201958 93201966 9 26 chrl 2 121814285 121814292 8 27 chrl 6 10773345 10773353 9 28 chrll 62613611 62613619 9 29 chr4 112657326 112657335 10 30 chrl 8 12122493 12122500 8 31 chrl 4 20198016 20198021 6 32 chrl 6 31759375 31759381 7 33 chr20 13782845 13782851 7 34 chr7 141919402 141919409 8 35 chrlO 124447933 124447939 7
[0341]
[0342] 008613259
[0343] Rank Chromosome Start End Length 36 chrl2 48065112 48065119 8 37 chr3 49013948 49013954 7 38 chr6 42106567 42106573 7 39 chrl7 67903834 67903842 9 40 chrl2 101286378 101286385 8 41 chr22 19524196 19524202 7 42 chrl6 27464650 27464656 7 43 chr4 82864411 82864419 9 44 chrl3 32770749 32770757 9 45 chrlO 102141319 102141326 8 46 chrl9 18777182 18777189 8 47 chrll 63758183 63758190 8 48 chr22 20142998 20143003 6 49 chr5 14476894 14476901 8 50 chrl8 48636671 48636677 7 51 chrlO 88922388 88922395 8 52 chr22 31091812 31091819 8 53 chr2 232523817 232523824 8 54 chrl9 6477239 6477247 9 55 chrl2 88055666 88055673 8 56 chrl 230995820 230995828 9 57 chr7 19729393 19729401 9 58 chrll 105008959 105008968 10 59 chrl 9 464140 464147 8 60 chrl 183546131 183546139 9 61 chr3 43605720 43605728 9 62 chrl 114994979 114994988 10 63 chrll 32614404 32614412 9 64 chrl 2 122867344 122867352 9 65 chr2 210022955 210022962 8 66 chr2 202635176 202635183 8 67 chrl 9 48714848 48714855 8 68 chr8 47954349 47954358 10 69 chrl 9 47694633 47694640 8 70 chr20 49850763 49850772 10
[0344]
[0345] 008613259
[0346] Rank Chromosome Start End Length 71 chr9 133071402 133071409 8 72 chr9 133071369 133071376 8 73 chrl9 49347215 49347222 8 74 chrl6 30030499 30030505 7 75 chrl 27294616 27294623 8 76 chrl7 50356605 50356611 7 77 chrll 107454501 107454508 8 78 chrl 91501799 91501807 9 79 chr20 60012728 60012737 10 80 chrlO 99235745 99235751 7 81 chrl 6 54933218 54933225 8 82 chrl 6 30005331 30005339 9 83 chr2 169140514 169140521 8 84 chrl 7 44678884 44678892 9 85 chrl 5 22851821 22851829 9 86 chrl 9 45768795 45768801 7 87 chrl 9 5587270 5587277 8 88 chr2 27025648 27025655 8 89 chr4 83459705 83459712 8 90 chr22 37011066 37011072 7 91 chrl 6 50711487 50711493 7 92 chrl 8 63917504 63917512 9 93 chrl 3 19467253 19467260 8 94 chr2 70964442 70964448 7 95 chrl 9 11466789 11466797 9 96 chrlO 17321321 17321328 8 97 chr9 93660329 93660336 8 98 chrl 7 66062963 66062969 7 99 chrl 6 71186904 71186912 9 100 chr7 87444965 87444973 9 101 chrl 9 42368602 42368608 7 102 chrll 62831112 62831119 8 103 chrl 3 102729645 102729653 9 104 chr5 24492863 24492870 8 105 chr3 186644754 186644762 9
[0347]
[0348] 008613259
[0349] Rank Chromosome Start End Length 106 chrl5 49016853 49016860 8 107 chrl4 23048304 23048310 7 108 chr20 62064910 62064915 6 109 chrl6 85648683 85648690 8 110 chrll 118899942 118899949 8 111 chrl 64841313 64841320 8 112 chr6 131160135 131160142 8 113 chrl 7 4672598 4672604 7 114 chr20 49241966 49241973 8 115 chr9 129924963 129924970 8 116 chr4 173529042 173529048 7 117 chrl 4 23276100 23276107 8 118 chrl 43384471 43384479 9 119 chrl 1354729 1354735 7 120 chrl 6 67233948 67233955 8 121 chr7 121331824 121331830 7 122 chr2 74460282 74460289 8 123 chrl 5 55620693 55620700 8 124 chrX 108734572 108734579 8 125 chrl 4 35116728 35116736 9 126 chrlO 62399753 62399760 8 127 chr3 77607886 77607893 8 128 chrl 5 93002203 93002211 9 129 chrl 3 26681249 26681256 8 130 chrl 2 8129340 8129347 8 131 chrlO 68422763 68422771 9 132 chrl 5 42450758 42450765 8 133 chr2 110168520 110168527 8 134 chr9 34257624 34257631 8 135 chr3 108510550 108510556 7 136 chrl 7 8320979 8320985 7 137 chr3 183947468 183947476 9 138 chr2 166477463 166477470 8 139 chrl 7 35422473 35422480 8 140 chrll 94496892 94496900 9
[0350]
[0351] 008613259
[0352] Rank Chromosome Start End Length 141 chr7 101163310 101163315 6 142 chr4 80870099 80870106 8 143 chr7 6582217 6582224 8 144 chr2 99162317 99162324 8 145 chrl9 3633461 3633467 7 146 chrl6 70918258 70918265 8 147 chrl2 56165328 56165335 8 148 chrlO 50245433 50245440 8 149 chr7 101544034 101544040 7 150 chrl 118150575 118150581 7 151 chr22 31089935 31089940 6 152 chr2 185797375 185797382 8 153 chr3 42637543 42637550 8 154 chrll 78209809 78209816 8 155 chrlO 89723990 89723998 9 156 chrl 6 30965995 30966001 7 157 chr9 33264655 33264661 7 158 chr5 140669516 140669523 8 159 chr8 66840003 66840009 7 160 chrl 7 79794848 79794855 8 161 chrll 1445649 1445655 7 162 chr3 44244390 44244398 9 163 chrl 8 33333178 33333187 10 164 chr6 107070397 107070403 7 165 chrl 4 88473768 88473775 8 166 chr5 36984980 36984987 8 167 chrl 2 92802642 92802650 9 168 chr8 6445117 6445123 7 169 chrll 66068674 66068681 8 170 chrl 9 51311850 51311857 8 171 chr3 67360635 67360641 7 172 chrl 114510877 114510883 7 173 chrl 8 36625552 36625557 6 174 chr7 107594094 107594100 7 175 chr2 118846931 118846936 6
[0353]
[0354] 008613259
[0355] Rank Chromosome Start End Length 176 chrl6 74942792 74942799 8 177 chr5 144473967 144473975 9 178 chrl7 7294261 7294267 7 179 chrl 17358837 17358844 8 180 chr7 108123138 108123144 7 181 chrl 9 12528356 12528362 7 182 chr4 56354102 56354111 10 183 chrl 4 105469901 105469907 7 184 chrl 4 23359611 23359616 6 185 chrlO 50097794 50097800 7 186 chr4 168396112 168396119 8 187 chrl 5 31483783 31483788 6 188 chrl 150704299 150704305 7 189 chr5 163490419 163490427 9 190 chrll 126267191 126267198 8 191 chr2 43918025 43918032 8 192 chr2 53866207 53866214 8 193 chr3 7579325 7579332 8 194 chr2 164508777 164508785 9 195 chr4 1978831 1978838 8 196 chr8 66576217 66576224 8 197 chrll 75983386 75983395 10 198 chrl 6 77741846 77741853 8 199 chrX 70907767 70907774 8 200 chr3 62874798 62874803 6 201 chrl 3 60008599 60008607 9 202 chrll 4682244 4682251 8 203 chrl 9 48498893 48498900 8 204 chr5 146507159 146507166 8 205 chrl 7 17216394 17216401 8 206 chrl 7 42787851 42787857 7 207 chr2 182757787 182757792 6 208 chrl 2 57028788 57028795 8 209 chrll 5778421 5778430 10 210 chrl 5 101489566 101489571 6
[0356]
[0357] 008613259
[0358] Rank Chromosome Start End Length 211 chrl7 42218217 42218224 8 212 chr9 133071501 133071508 8 213 chr21 30289461 30289468 8 214 chrl6 28902318 28902325 8 215 chrl 35381358 35381366 9 216 chr3 182884885 182884892 8 217 chr7 138918777 138918783 7 218 chr4 2059729 2059735 7 219 chr2 195923649 195923657 9 220 chr5 80675095 80675102 8 221 chr2 162278205 162278212 8 222 chrl 8 22992889 22992897 9 223 chr9 133071600 133071607 8 224 chrl 7 56950654 56950661 8 225 chr22 37373063 37373070 8 226 chrlO 29471186 29471192 7 227 chrl 9 48498956 48498962 7 228 chrl 4 23571226 23571234 9 229 chr2 164694785 164694793 9 230 chr6 161048743 161048750 8 231 chr8 11201065 11201071 7 232 chrll 105007313 105007322 10 233 chrl 6 88624732 88624739 8 234 chr4 3013742 3013750 9 235 chrl 2 54252047 54252054 8 236 chrl 9 48955713 48955720 8 237 chrl 6 28984905 28984912 8 238 chr4 44698597 44698604 8 239 chr2 109544250 109544258 9 240 chrl 4 92130977 92130984 8 241 chr2 178124192 178124200 9 242 chrll 59124903 59124912 10 243 chrlO 76970027 76970035 9 244 chrl 13782253 13782261 9 245 chr7 77794142 77794150 9
[0359]
[0360] 008613259
[0361] Rank Chromosome Start End Length 246 chr4 55470786 55470794 9
[0362] 247 chr4 3430703 3430710 8
[0363] 248 chr7 8158620 8158628 9
[0364] 249 chr2 203057334 203057342 9
[0365] 250 chrl9 14799825 14799834 10
[0366] 251 chrl 1046480 1046487 8
[0367] 252 chrl 160619810 160619818 9
[0368] 253 chr3 157363437 157363445 9
[0369] 254 chrl 7 4972442 4972449 8
[0370] 255 chr4 84635321 84635329 9
[0371] 256 chrl 8 31873454 31873461 8
[0372] 257 chr5 32404055 32404063 9
[0373] 258 chrl 74109528 74109536 9
[0374] 259 chr3 194427120 194427127 8
[0375] 260 chr9 131132605 131132612 8
[0376] 261 chrl 7 18278230 18278236 7
[0377] 262 chr20 31476429 31476436 8
[0378] 263 chrl 28459218 28459227 10
[0379] 264 chr5 16694496 16694503 8
[0380] 265 chr6 158086976 158086983 8
[0381] 266 chr7 48046607 48046614 8
[0382] 267 chrl 7 78050898 78050905 8
[0383] 268 chrl 5 51869364 51869371 8
[0384] 269 chrl 3 52474898 52474905 8
[0385] 270 chrl 22913752 22913760 9
[0386] 271 chr6 80042179 80042187 9
[0387] 272 chrll 118349867 118349875 9
[0388] 273 chr4 15936554 15936562 9
[0389] 274 chr3 136854643 136854651 9
[0390]
[0391] It is noted that an inconsistency in the " Length" column exists between Table 1 of priority application GB 2501140.4 and Table 1 of the present disclosure, whereby all lengths in GB 2501140.4 Table 1 were erroneously shown as 1 bp shorter than the actual loci lengths. As the skilled person understands,008613259
[0392] the length of a genomic locus may be calculated by subtracting the start coordinate from the end coordinate and adding one to account for the first base pair (length = end coordinate -start coordinate + 1). The identity of the loci, ranks, chromosomes, start coordinates and end coordinates remain consistent with GB 2501140.4.008613259
[0393] References
[0394] Ahadova, A., von Knebel Doeberitz, M., Blaker, H. & Kloor, M., 2016. CTNNB1 -mutant colorectal carcinomas with immediate invasive growth: a model of interval cancers in Lynch syndrome. Familial Cancer, 15 ( 1 ), pp. 579 - 586.
[0395] Argillander, T. et al., 2018. Features of incident colorectal cancer in Lynch syndrome. United European Gastroenterology Journal, 6 ( 8 ), p. 1215-1222.
[0396] Caravagna, G. et al. ( 2020 ). Subclonal reconstruction of tumors by using machine learning and population genetics ', Nature Genetics, 52 ( 1 ), pp. 898 - 907.
[0397] Cortes - Ciriano, I., Lee, S., Park, WY. et al. A molecular portrait of microsatellite instability across multiple cancers. Nat Commun 8, 15180 ( 2017 ).
[0398] Gallon, R. et al. ( 2019 ) Sequencing -based microsatellite instability testing using as f ew as six markers for high -throughput clinical diagnostics ', Human Mutation, 41 ( 1 ), pp.
[0399] 332 - 341.
[0400] Joung, J. - G. et al. ( 2017 ) ' Tumor heterogeneity predicts metastatic potential in colorectal cancer ', Clinical Cancer Research, 23 ( 23 ), pp. 7209 - 7216.
[0401] Kayhanian, H., Cross, W., van der Horst, S. E. M. et
[0402] al. Homopolymer switches mediate adaptive mutability in mismatch repair - def icient colorectal cancer. Nat Genet 56, 1420-1433 ( 2024 ).008613259
[0403] Kloor, M., Michel, S., & von Knebel Doeberitz, M. ( 2010 ).
[0404] Immune evasion of microsatellite unstable colorectal cancers. International Journal of Cancer, 127 ( 5 ), 1001 - 1010.
[0405] Kondelin et al., " Comprehensive Evaluation of Protein Coding Mononucleotide Microsatellites in Microsatellite - Unstable Colorectal Cancer. " Cancer research vol. 77, 15 ( 2017 ): 4078 -4088.
[0406] Kunkel, T. A. & Erie, D. A. DNA mismatch repair. Annu Rev Biochem 74, 681–710 ( 2005 ).
[0407] Maley et al., " Genetic clonal diversity predicts progression to esophageal adenocarcinoma. " Nat Genet 38, 468-473 ( 2006 ). https: / / doi. org / 10. 1038 / ngl768.
[0408] McGranahan, N., Furness, A. J. S., Rosenthal, R., Ramskov, S., Lyngaa, R., Saini, S. K., Jamal - Han j ani, M., Wilson, G. A., Birkbak, N. J., Hiley, C. T., Watkins, T. B. K., Shaf i, S., Rizvi, N. A., Snyder, A., Hellmann, M. D., Merghoub, T., Wolchok, J. D., Shukla, S. A., Wu, C. J., Peggs, K. S., Chan, T. A., Hadrup, S. R., Quezada, S. A., & Swanton, C. ( 2016 ). Clonal neoantigens elicit T cell immunoreactivity and sensitivity to immune checkpoint blockade. Science, 351 ( 6280 ), 1463 - 1469.
[0409] Sansom, O. J. et al. ( 2001 ) ' MSH - 2 suppresses in vivo mutation in a gene dose and lesion dependent manner1, Oncogene, 20 ( 27 ), pp. 3580 - 3584.
[0410] Sekine, S. et al. ( 2017 ) ' Mismatch repair def iciency commonly precedes adenoma formation in Lynch syndrome - associated colorectal tumorigenesis1, Modern Pathology, 30 ( 8 ), pp. 1144 -1151.008613259
[0411] Westcott, P. M. K., Muyas, F., Smith, O., Hauck, H., Sacks, N. J., Ely, Z. A., Jaeger, A. M., Rideout Ill, W. M., Bhutkar, A., Zhang, D., Beytagh, M. C., Canner, D. A., Bronson, R. T., Naranj o, S., Jin, A., Patten, J., Cruz, A. M., Cortes - Ciriano, I., & Jacks, T. ( 2021 ). Mismatch repair def iciency is not suf f icient to increase tumor immunogenicity. bioRxiv.
Claims
008613259Claims1. A method for estimating the mutation rate, or mutation rate - inf luenced property, of a tumour of a subj ect, comprising:providing nucleic acid sequencing data obtained by sequencing DNA or RNA derived f rom one or more cells of the tumour ( "sample sequencing data" );measuring microsatellite diversity of at least 10 erosion sensitive loci in the sample sequencing data to provide one or more diversity metrics;correlating the one or more diversity metrics to an estimated mutation rate or mutation rate - inf luenced property of the tumour.
2. The method of claim 1, wherein the at least 10 erosion sensitive loci are selected f rom the loci set forth in Table 1.
3. The method of claim 1 or claim 2, wherein the at least 10 erosion sensitive loci comprise at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, at least 270 or substantially all of the loci set forth in Table 1.
4. The method of claim 3, wherein the at least 10 erosion sensitive loci comprise the top 10, top 20, top 30, top 40, top 50, top 60, top 70, top 80, top 90 or top 100 loci by ranking as set forth in Table 1, optionally the method employs not more than 150 erosion sensitive loci.
5. The method of claim 3 or claim 4, wherein the at least 10 erosion sensitive loci comprise:008613259( i ) loci with 5 - 11 base pair repeat length, optionally 6 - 10 base pair repeat length, and / or( ii ) at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or all of the loci shown as having length 7 in Table 1, and / or ( iii ) at least 5, at least 10, at least 30, at least 40, at least 50, at least 60, or all of the loci shown as having length 8 in Table 1, and / or( iv) at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180 or all of the loci shown as having length 7 or length 8 in Table 1.
6. The method of any one of claims 3 - 5,wherein the at least 10 erosion sensitive loci are located in at least 2, at least 5, at least 10, at least 15 or at least 20 chromosomes, and / orwherein the at least 10 erosion sensitive loci are located in genes, optionally wherein the at least 10 erosion sensitive loci are located in at least 2, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 40 or at least 50 genes, and / orwherein the at least 10 erosion sensitive loci are located in exons optionally wherein the at least 10 erosion sensitive loci are located in at least 2, at least 5, at least 7, at least 10, at least 15, at least 20, at least 30, at least 40 or at least 50 exons.
7. The method of any one of the preceding claims, wherein measuring microsatellite diversity comprises measuring the diversity of microsatellite repeat length, optionally wherein008613259said measuring comprises generating a read count distribution of the microsatellite repeat lengths at said loci.
8. The method of any one of the preceding claims, wherein the greater the one or more diversity metrics, the higher the estimated mutation rate of the tumour.
9. The method of any one of the preceding claims, wherein the one or more diversity metrics comprises Shannon Diversity Index (SDI).
10. The method of claim 9, wherein the SDI is determined according to the formula:RShannon diversity = — piln(pi)i=lwherein pi = the proportion of total sequence reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite.
11. The method of any one of the preceding claims, wherein the one or more diversity metrics comprises a measure of the divergence of: (i) the microsatellite length probability distribution at each of said loci for the tumour; and (ii) the microsatellite length probability distribution of the same loci for sequence data obtained by sequencing DNA or RNA derived from one or more normal (non- cancer) cells of the subj ect.
12. The method of claim 11, wherein the one or more diversity metrics comprises the Jensen - Shannon Distance (JSD) according to the formula:008613259JSD(P || Q) = |r>(P || M) + |r>(Q || M)wherein P and Q are the microsatellite length distributions in tumour and normal samples, respectively, M is the average P+Qdistribution defined as M = — — and D(X || Y) is the Kullback-Leibler divergence between X and Y.
13. The method of claim 12, wherein the Kullback - Leibler divergence (DKL) is given by the formula:
14. The method of any one of the preceding claims, wherein the one or more diversity metrics are determined at each of said loci and combined to provide a global measure of diversity across said loci, optionally wherein the global measure is the median or the mean.
15. The method of claim 14, wherein the global measure of diversity across said loci (e.g. the median SDI) is compared to a reference value, wherein the global measure of diversity being above the reference value indicates that the mutation rate of the tumour is high, and wherein the global measure of diversity being at or below the reference value indicates that the mutation rate of the tumour is low.
16. The method of any one of the preceding claims, wherein correlating the one or more diversity metrics to the estimated mutation rate of the tumour comprises: application of a binary classification model, application of a regression model, application of a clustering algorithm or percentile -based stratification.00861325917. The method of any one of the preceding claims, wherein correlating the one or more diversity metrics to the estimated mutation rate or mutation rate - inf luenced property of the tumour comprises:inputting the one or more diversity metrics into a machine learning model, wherein the machine learning model has been trained to take as input the one or more diversity metrics and output the estimated mutation rate or mutation rate - inf luenced property of the tumour.
18. The method of any one of the preceding claims, wherein correlating the one or more diversity metrics to the estimated mutation rate of the tumour comprises:inputting sample data comprising the one or more diversity metrics (e.g. SDI) determined at each of said loci into a machine learning classifier, wherein the machine learning classifier has been trained on a training dataset comprising the same one or more diversity metrics (e.g. SDI) determined at each of said loci for a plurality of samples of cancers known to be of low mutation rate and a plurality of samples of cancers known to be of high mutation rate; and causing the machine learning classifier to classify the tumour as high or low mutation rate based on the inputted sample data.
19. The method of any one of the preceding claims, wherein the method is computer implemented.
20. The method of any one of the preceding claims, wherein the method comprises a preceding step of sequencing DNA or RNA derived from one or more cells of the tumour.
21. The method of claim 20, wherein the step of sequencing DNA or RNA derived from one or more cells of the tumour008613259comprises use of a targeted sequencing panel comprising not more than 10000, not more than 1000 or not more than 500 microsatellite loci.
22. The method of claim 21, wherein the targeted sequencing panel comprises the at least 10 erosion sensitive loci, optionally comprising at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, at least 270 or substantially all of the loci set forth in Table 1.
23. The method of any one of the preceding claims, wherein the sample sequencing data is derived f rom a tumour tissue sample or a cell - f ree nucleic acid sample, optionally wherein the sample comprises circulating tumour DNA ( ctDNA).
24. The method of any one of the preceding claims, wherein the tumour is selected f rom: a gastrointestinal cancer, a lung cancer and a breast cancer, a colon cancer, a gastric cancer, an endometrial cancer, an ovarian cancer, a biliary tract cancer, a urinary tract cancer, a brain cancer, and a skin cancer.
25. The method of any one of the preceding claims, wherein the tumour is a colorectal cancer (CRC).
26. The method of any one of the preceding claims, wherein the tumour is DNA mismatch repair def icient (MMRd).
27. The method of any one of the preceding claims, wherein the tumour mutation rate - inf luenced property is selected f rom cancer immunotherapy response, cancer prognosis, progression of a subj ect having hereditary nonpolyposis colorectal cancer to malignancy, probability of response to a cancer treatment,008613259probability of progression of a subj ect having hereditary nonpolyposis colorectal cancer to malignancy, and / or percentage tumour bed regression in response to cancer treatment.
28. The method of any one of the preceding claims, wherein the method is for predicting response to cancer immunotherapy, optionally immune checkpoint blockade (ICB) therapy, and wherein a higher value of the one or more diversity metrics indicates that the subj ect will respond poorly to the cancer immunotherapy and a lower value of the one or more diversity metrics indicates that the subj ect will respond well to the cancer immunotherapy.
29. The method of any one of the preceding claims, wherein the method is for providing a prognosis of the tumour, and wherein a higher value of the one or more diversity metrics indicates that the subj ect has a poor prognosis and a lower value of the one or more diversity metrics indicates that the subj ect has a good prognosis.
30. The method of claim 29, wherein the tumour comprises an MMRd CRC tumour, optionally a stage II or stage III MMRd CRC tumour, and wherein poor prognosis indicates shorter progression - free survival and / or shorter overall survival and / or shorter time to recurrence, and wherein good prognosis indicates longer progression - free survival and / or longer overall survival and / or shorter time to recurrence.
31. The method of any one of the preceding claims, wherein the method is for predicting progression of a subj ect having hereditary nonpolyposis colorectal cancer (Lynch Syndrome) to malignancy onset, and wherein said one or more cells of the tumour are one or more cells of a colorectal precursor008613259adenoma, and wherein a higher value of the one or more diversity metrics indicates that the subj ect is predicted to progress to the onset of CRC malignancy and a lower value of the one or more diversity metrics indicates that the subj ect is not predicted to progress to onset of CRC malignancy.
32. The method of any one of the preceding claims, wherein the method is for classifying a MMRd CRC tumour as MSH6 -proficient or MSH6 -def icient, and wherein a higher value of the one or more diversity metrics indicates that the tumour is MSH6 - def icient and a lower value of the one or more diversity metrics indicates that the tumour is MSH6 -prof icient.
33. The method of any one of the preceding claims, wherein the method is for classifying a MMRd CRC tumour as MSH3 -prof icient or MSH3 -def icient, and wherein a higher value of the one or more diversity metrics indicates that the tumour is MSH3 -deficient and a lower value of the one or more diversity metrics indicates that the tumour is MSH3 -prof icient.
34. The method of any one of the preceding claims, wherein the method is for estimating sub- clonal complexity of the tumour and / or intra - tumoral heterogeneity (ITH), and wherein a higher value of the one or more diversity metrics indicates that the tumour has greater sub -clonal complexity and / or higher ITH and a lower value of the one or more diversity metrics indicates that the tumour has lower sub - clonal complexity and / or lower ITH.
35. The method of any one of the preceding claims, wherein the method is for sub - clonal tracking of the tumour and / or temporal tumour evolutionary reconstruction, and wherein the sample sequencing data comprises data for each of a plurality of different locations of the tumour and / or for each of a008613259plurality of different time points in the history of the tumour.
36. The method of claim 35, wherein the one or more diversity metrics comprise the Jensen - Shannon Distance (JSD), optionally wherein the JSDs of the loci are used to construct a phylogeny of the tumour.
37. A computer implemented method for identifying coding microsatellites that are sensitive to changes in tumour mutation rate, comprising:providing whole exome sequencing data for each of a plurality of tumours known to be MSH6 -prof icient and / or MSH3 -proficient ("proficient tumours" ) and a plurality of tumours known to be MSH6 - def icient and / or MSH3 - def icient ("deficient tumours" );generating for each of a plurality of candidate microsatellites a read count distribution for the proficient tumours and a read count distribution for the deficient tumours;deriving from the read count distributions a Shannon diversity index (SDI) for each candidate microsatellite according to the formula:RShannon diversity = — piln(pi)i=lwherein pi = the proportion of total sequence reads represented by the microsatellite length and R = total number of read lengths present at a microsatellite;ranking the candidate microsatellites based on the change (A) in SDI between proficient and deficient tumours; and008613259selecting based on their ranking, the candidate microsatellites most sensitive to changes in tumour mutation rate.
38. The method according to claim 37, wherein the proficient and deficient tumours are MMRd CRC tumours.