Risk stratification for colorectal cancer
By analyzing stool or anal swab samples for colibactin-related mutational signatures, the method identifies individuals at risk for early-onset colorectal cancer, facilitating targeted screening and early intervention.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GENOME RES LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Existing methods fail to identify specific risk factors for early-onset colorectal cancer, particularly those related to environmental and lifestyle carcinogenic exposures, necessitating a new approach to stratify individuals at risk for targeted screening.
Analyze nucleic acid sequence data from stool or anal swab samples to quantify mutations associated with microbial genotoxic compounds like colibactin, using mutational signatures (SBS88 and ID18) to determine an individual's risk of developing colorectal cancer, enabling classification for invasive screening.
Provides a non-invasive method to identify individuals at high risk of colorectal cancer before tumor development, allowing for targeted screening and early intervention.
Smart Images

Figure IMGF000057_0001 
Figure IMGF000064_0001_TABLE 
Figure IMGF000065_0001_TABLE
Abstract
Description
[0001] RISK STRATIFICATION FOR COLORECTAL CANCER
[0002] FIELD OF THE DISCLOSURE
[0003] The present disclosure relates to a method for determining whether a subject is at risk of developing colorectal cancer based on nucleic acid sequence data from a stool sample. It is specifically concerned with a method for identifying whether the sample exhibits mutations associated with exposure to a microbial genotoxic compound, including but not limited to colibactin, and may be used to select subjects for invasive screening tests accordingly.
[0004] BACKGROUND
[0005] Colorectal cancer incidence rates differ markedly by geographic location and have changed substantially in some countries over the last 70 years (Patel et al., 2022). For instance, the ASR for colorectal cancer in North America and in most European countries peaked in the 1980s and 1990s and have been declining since, whereas countries in East Asia such as Japan and South Korea have been steadily increasing over the past seven decadesl . Moreover, In the past 20 years there has been a notable global increase in the incidence of early-onset colorectal cancer (Patel et al., 2022; Torrens et al., 2024), typically defined as colorectal cancer in adults under 50 years of age. Although epidemiological studies have identified multiple risk factors for colorectal cancer, specific risk factors for early-onset colorectal cancer remain largely unidentified, with the exception of family history and hereditary predisposition. The latter is predominantly attributable to Lynch syndrome, which is characterized by DNA mismatch repair deficient cancers of the proximal colon (Spaander et al., 2023; Stigliano et al., 2014) and, therefore, is unlikely to be implicated in the recent increase in early-onset colorectal cancer, which is mainly enriched in sporadic, DNA mismatch repair proficient cancers affecting the distal colon and rectum (Venugopal and Carethers, 2022; You et al., 2012). Despite extensive epidemiological research, the underlying causes for many of these variations remain unclear. However, they are suspected to be due to exogenous environmental or lifestyle carcinogenic exposures, which are, in principle, preventable (Brennan, P. & Davey-Smit, 2022).
[0006] SUMMARY OF THE DISCLOSURE
[0007] The conventional strategy for identifying lifestyle and environmental causes of cancer involves conducting large-scale population-based studies (cohort or case-control) to compare disease incidence between individuals with and without the disease. This method has been crucial in identifying significant cancer causes such as smoking, obesity, alcohol, and infectious agents. However, no major significant cancer causes have been found in the past two decades, highlighting the need for new strategies to systematically identify further cancer causes, particularly in relation to early-onset colorectal cancer. Somatic mutations can arise due to a cell’s exposure to exogenous DNA-damaging agents, or they can be due to internal mutational processes present in all cells of the body. Once generated, mutations are usually permanent, and the somatic mutations found in a normal cell or cancer cell genome have,therefore, accumulated through the full course of an individual’s lifespan. Many known causes of cancer have their own characteristic patterns of mutations, termed mutational signatures. Combining international cancer epidemiology and mutational signatures analysis of cancer genomes provides a complementary approach for revealing the mutational processes that contribute to cancer development, and offers a non-invasive strategy to monitor an individual’s previous exposure to cancer-causing factors and subsequently their future risk of developing cancer.
[0008] Many well-known exogenous carcinogens are also mutagens, which can imprint characteristic patterns of somatic mutations in the genome, known as mutational signatures. Therefore, a complementary approach to conventional epidemiology for investigating unknown causes of cancer is the characterization of mutational signatures in the genomes of cancer and normal cells (Senkin et al. 2024; Moody et al., 2021; Zhang et al. 2021). The Mutographs Cancer Grand Challenge project (Perdomo et al. 2024) has implemented this strategy of “mutational epidemiology” by sequencing cancers from geographic areas of differing incidence rates, using mutational signature analysis to elucidate the mutational processes that have been operative. The present inventors set out to examine colorectal cancer genomes from 11 countries on four continents to investigate whether variation in mutational processes contributes to geographic and age-related differences in incidence rates. In this process, they identified that signatures SBS88 and ID18, caused by the bacteria-produced mutagen colibactin, had higher mutation loads in countries with higher colorectal cancer incidence rates. SBS88 and ID18 were also enriched in early-onset colorectal cancers, being 3.3 times more common in individuals diagnosed before age 40 than in those over 70, and were imprinted early during colorectal cancer development. Colibactin exposure was further linked to APC driver mutations, with ID18 responsible for about 25% of APC driver indels in colibactin-positive cases. This study revealed geographic and age-related variations in colorectal cancer mutational processes, and suggested that early-life mutagenic exposure to colibactin-producing bacteria may contribute to the rising incidence of early-onset colorectal cancer, and be usable as a risk assessment tool for the development of colorectal cancer.
[0009] Although colibactin has been previously postulated to cause mutations that can lead to colorectal cancer and the progression of colorectal cancer, there was previously no indication of when or how such an effect would be detectable, or indeed when in cancer development the effect occurs. The present inventors have been able to show, to the best of their knowledge for the first time, that specific mutational signatures and hallmark mutations and motifs associated with exposure to microbial genotoxic compounds including colibactin are associated specifically with early onset colorectal cancer (as opposed to colorectal cancer in general), and are early events. This provide the first evidence that those would be detectable in non-tumour DNA. i.e. prior to development of the cancer, and could be used as markers to identify individuals who do not yet have colorectal cancer but are at high risk of developing it. Testing of stool samples and / or anal swab samples in subjects that are not known to have colorectal cancer, to identify the presence of these pre-cancerous markers, provides a new and highly beneficial way to stratify patients for cancer screening.Thus, according to a first aspect, there is provided a method of determining whether a subject is at risk of developing colorectal cancer, the method comprising: obtaining nucleic acid sequence data from a stool sample or anal swab sample from said subject; quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound; and determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer.
[0010] The step of determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer may comprise classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer. The step of obtaining genomic sequence data from a stool sample or anal swab sample from said subject may comprise receiving, by a processor, previously acquired genomic sequence data. The steps of quantifying and classifying can be performed at a processor. Thus, the methods of the present aspect may be computer implemented. In other words, also described herein is a computer implemented method of determining whether a subject is at risk of developing colorectal cancer, the method comprising: (i) obtaining genomic sequence data from a stool sample or anal swab sample from said subject; and (ii) quantifying, using said genomic sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound, and determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer, where the determining may comprise classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer.
[0011] The method may have any one or more of the following optional features.
[0012] In embodiments, the microbial genotoxic compound is colibactin. In embodiments, the one or more metrics are selected from: metrics that quantify the presence or activity of a mutational signature associated with exposure to a microbial genotoxic compound, optionally colibactin, and metrics that quantify the presence, number or proportion of mutations associated with exposure to a microbial genotoxic compound, optionally colibactin.
[0013] In embodiments, quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound comprises determining the presence of one or more mutational signatures, the number of mutations associated with one or more mutational signatures, the proportion of mutations associated with one or more mutational signatures, and / or the activity of one or more mutational signatures. In such embodiments, the one or more mutational signature may be mutational signatures associated with exposure to colibactin. A mutational signature may comprise a plurality of proportions for respective categories of mutations, optionally wherein the plurality of categories of mutations are categories of single base substitutions or categories of indels.In embodiments, the one or more mutational signatures are selected from: (i) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS88) is characterised by: (a) a higher proportion of T>C mutations in ATA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in ATA context, T>C mutations in ATT context, T>C mutations in TTT context, and T>G mutations in TTT context than any other categories of single base substitutions; (ii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS89) is characterised by a higher proportion of T>G mutations in GTG context and C>T mutations in ACA context than any other categories of single base substitutions; (iii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS_M) is characterised by: (a) a higher proportion of T>C mutations in CTA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in CTA context, T>C mutations in TTA context, and T>A mutations in CTA context than any other categories of single base substitutions; and (iv) an indel signature comprising proportions for each of a plurality of categories of indels each defined by a unique combination of whether the indel is a deletion or insertion, whether the indel is a 1bp indel or longer, for 1pb indel whether the indel is a C or T and the length of the mononucleotide repeat tract in which they occur, for longer indels whether the indel is an insertion at a repeat region and the deletion length and number of repeat units, whether the indel is a deletion at a repeat region and the insertion length and number of repeat units, and whether the indel is a deletion with microhomology at the deletion boundary and the deletion length and microhomology length, wherein the indel signature (ID18) is characterised by higher proportions of 1bp T deletions than any other categories of indels. In embodiments, the mutational signature is selected from: SBS88, SBS89, SBS_M and ID18, or corresponding signatures obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS88, SBS89 and / or SBS_M, or ID18. In embodiments, the mutational signature is SBS88, or a corresponding signature obtained by: extracting one or more signatures from a cohort of samples using the same plurality of categories, and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS88.
[0014] In embodiments, quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound comprises determining: the percentage or proportion of total W[T>N]W mutations that are within a WAWW[T>N]W motif, where W is A or T, and N is A, C, G or T; and / or the presence of a splicing variant c.835-8A>G in the APC gene; and / or the presence of one or more protein-truncating mutations in the APC gene; and / or the presence of indels in the APC gene associated with 1bp T or A deletions and / or associated with indel signature ID18.In embodiments, the subject is a human subject. In embodiments, the subject is a human subject under the age of 50. In embodiments, the subject is a human subject between the ages of 18 and 50. In embodiments, the subject is a healthy subject. In embodiments, the subject is a subject who has not been diagnosed with Lynch syndrome. In embodiments, the subject is a subject who has not been identified as having one or more germline mutations associated with DNA mismatch repair deficiency. In embodiments, the sample is a stool sample. In embodiments, a subject who is determined to be at risk of developing colorectal cancer (e.g. a subject classified in the first class) is recommended or selected for a further colorectal cancer screening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy. A subject who is determined to be at risk of developing colorectal cancer (e.g. a subject classified in the first class) may be recommended or selected for regular invasive colorectal cancer screening testing, such as e.g. regular colonoscopy. For example, such tests may be recommended to be performed every year, every 2 years, every 3 years, every 4 years or every 5 years. A subject who is determined not to be at risk of developing colorectal cancer (e.g. a subject classified in the second class) may be recommended or selected for standard or care colonoscopy screening associated with the subject’s age. This may not include regular colonoscopy under the age of e.g. 45. In embodiments, the sequence data comprises DNA sequencing reads, or a list of mutations or a mutational profile derived therefrom. In embodiments, the sequence data comprises DNA sequencing reads obtained from bulk sequencing or error-corrected duplex sequencing, or information derived therefrom. In embodiments, the sequence data comprises DNA sequencing read obtained using a sequencing methodology associated with a sequencing error rate below one error per 10 million sequenced base pairs, or information derived therefrom (e.g. a list or table of mutations derived from such data). In embodiments, the sequence data comprises DNA sequencing read obtained using a single molecule duplex sequencing method, or information derived therefrom (e.g. a list or table of mutations derived from such data). A list of mutations and a mutational profile for a subject both comprise information that identifies mutations (e.g. SBS or indels) identified in a subject. The mutations may be somatic mutations. In other words, mutations identified in a sample may be filtered for likely germline mutations. In embodiments, the sequence data comprises DNA sequencing reads from whole genome sequencing, whole exome sequencing or targeted panel sequencing, or information derived therefrom. In embodiments, the sample is a stool sample and the sequence data comprises DNA sequencing reads from a library derived from the sample and subject to a target sequence capture, optionally using an exome capture panel, or information derived therefrom (e.g. a list or table of mutations derived from such data).
[0015] In embodiments, the sequence data comprises sequence data identifying the presence or absence of one or more of: a splicing variant c.835-8A>G in the APC gene, a protein-truncating mutation in the APC gene, and indels in the APC gene associated with 1 bp T or A deletions. Such sequence data may be obtained from e.g. targeted sequencing or any other targeted or untargeted sequence determination technology, such as molecular inversion probes, digital PCR, etc. In embodiments, the colorectal cancer is rectal cancer. In embodiments, the colorectal cancer is early-onset colorectal cancer.In embodiments, the sample is a stool sample that has been previously obtained from a subject, optionally fresh frozen, and from which DNA has been extracted and processed for sequencing. In embodiments, obtaining the sequence data comprises extracting DNA from a stool sample and / or processing extracted DNA from a stool sample for sequencing and / or sequencing the processed extracted DNA from the stool sample. In embodiments, DNA has been or is extracted from the sample using selective lysis of the subject’s cells over microbial cells followed by DNA purification, optionally column-based DNA purification. In embodiments, the method comprises one or more of: receiving a stool sample previously obtained from the subject, optionally in a sterile container; storing the stool sample at -80C; extracting DNA from the sample from the subject, optionally using a stool-specific protocol; extracting DNA from the sample from the subject using a protocol comprising a step of selective lysis of the subject’s cells over microbial cells followed by DNA purification; testing the extracted DNA for DNA integrity and optionally performing an end-repair step when a DNA integrity number for the extracted DNA is above a predetermined threshold; obtaining a DNA sequencing library from DNA extracted from the sample from the subject, optionally including adding unique molecular identifiers to the DNA fragments in the sample and / or performing enrichment of sequences comprising predetermined sequences using target sequence capture; and sequencing a DNA library obtained from the sample from the subject, optionally using standard bulk sequencing or error-corrected duplex sequencing.
[0016] In embodiments, the sequence data comprises DNA sequencing reads and the method comprises: processing said DNA sequencing reads to identify single base substitutions and / or indels present in the sequence data, and optionally obtaining a mutational profile from said identified substitutions and / or indels by classifying identified mutations in a predetermined plurality of categories.
[0017] In embodiments, the sample is an anal swab sample that has been previously obtained from the subject, optionally stored in a cytology preservation solution, and from which DNA has been extracted and processed for sequencing. In embodiments, obtaining the sequence data comprises extracting DNA from an anal sample and / or processing extracted DNA from an anal swab sample for sequencing and / or sequencing the processed extracted DNA from the anal swab sample. In embodiments, the method comprises one or more of: receiving an anal swab sample previously obtained from the subject, optionally in a cytology preservation solution; storing the anal swab sample at 4C; extracting DNA from the sample from the subject; extracting DNA from the sample from the subject, optionally using a standard DNA extraction protocol; testing the extracted DNA for DNA integrity and optionally performing an end-repair step when a DNA integrity number for the extracted DNA is above a predetermined threshold; obtaining a DNA sequencing library from DNA extracted from the sample from the subject, optionally including adding unique molecular identifiers to the DNA fragments in the sample; and sequencing a DNA library obtained from the sample from the subject, optionally using singe molecule duplex sequencing.In embodiments, determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer comprises classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer. In embodiments, classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer comprises comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein values of the metrics or score at or above the respective predetermined thresholds is indicative of a higher risk of developing colorectal cancer than values of the metrics or score below the respective predetermined thresholds. In embodiments, the respective predetermined thresholds are thresholds that have been identified using a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer (e.g. by comparing the values of the one or more metric for the plurality of subjects that have developed early onset colorectal cancer vs. the plurality of subjects that have not developed early onset colorectal cancer). In embodiments, the method further comprises identifying the one or more predetermined thresholds using values of the one or more metrics for a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer. In embodiments, classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer comprises: comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein a subject is: (i) classified in the first class when said score is at or above a predetermined threshold, and classified in the second class otherwise; (ii) classified in the first class when each of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise; or (iii) classified in the first class when a predetermined number or combination of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise. In embodiments, determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer (e.g. by classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer) comprises: inputting said one or more metrics into a machine learning classifier trained to classify subjects between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer using training data comprising the one or more metrics for a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer.
[0018] In embodiments, the method further comprises selecting the subject for a further colorectal cancer screening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy, when the subject is determined to be at risk of developing colorectal cancer (e.g. the subject is classified in the first class). In embodiments, the method further comprises performing a further colorectal cancerscreening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy, on the subject when the subject is determined to be at risk of developing colorectal cancer (e.g. the subject is classified in the first class).
[0019] In embodiments, the method further comprises providing a report comprising an output of the step of determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer, or information derived therefrom selected from: an indication of whether the subject was determined to be at risk of developing colorectal cancer, a classification between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer, a recommendation for a patient monitoring regimen comprising one or more further colorectal cancer screening tests.
[0020] Also described herein according to a second aspect is a method of selecting a subject for invasive screening for early onset colorectal cancer, the method comprising: determining whether the subject is at risk of developing colorectal cancer using the method of any embodiment of the first aspect, and selecting the subject for invasive screening for early onset colorectal cancer when the subject is determined to be at risk of developing colorectal cancer (e.g. the subject is classified in the first class).
[0021] Also described herein according to a third aspect is a method of treating a subject for early onset colorectal cancer, the method comprising: determining whether the subject is at risk of developing colorectal cancer using the method of any embodiment of the first aspect, and performing invasive screening for early onset colorectal cancer when the subject is determined to be at risk of developing colorectal cancer (e.g. the subject is classified in the first class).
[0022] Also described herein according to a fourth aspect is a system comprising: a processor; and a non-transitory computer readable medium comprising instructions that, when executed by the processor, cause the processor to implement the method of any embodiment of the first aspect.
[0023] Also described herein according to a fifth aspect is one or more non-transitory computer readable medium or media comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any embodiment of the first aspect.
[0024] Also described according to a sixth aspect is computer program product comprising instructions that, when executed by a processor, cause the processor to implement the method of any embodiment of the first aspect.
[0025] Also described according to a seventh aspect is a kit comprising reagents specific for measuring a metric according to the first aspect. The kit may comprise reagents specific for measuring a splicing variant c.835-8A>G in the APC gene, a protein-truncating mutation in the APC gene, and indels in the APC gene associated with 1bp T or A deletions.Also described according to an eighth aspect is a system comprising the kit of any embodiment of the seventh aspect, and a computer program product according to the sixth aspect, a computer readable medium according to the fifth aspect, ora system according to the fourth aspect.
[0026] Also described herein according to a ninth aspect is a method of analysing a stool sample or anal swab sample from a subject, the method comprising: obtaining the stool sample or anal swab sample from the subject or receiving a stool sample or anal swab sample previously obtained from the subject, extracting DNAfrom said sample, sequencing said extracted DNA thereby obtaining genomic sequence data associated with the sample, quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound, and determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer (e.g. by classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer). The methods according to the present aspect can have any of the features described in relation to the first aspect. The method may further comprise storing the sample in a container comprising a DNA stabilisation buffer. The method may further comprise storing the sample in a cytology preservation solution (e.g. PreservCyt™). This may be particularly useful for anal swab samples. Said storing may be at temperatures below 0°C, such as e.g. -17°C, -20°C or-80°C. This may be particularly useful for stool samples. Said storing may be at temperatures between 0 and 5°C, e.g.
[0027] 4°C. This may be particularly useful for anal swab samples. The method may further comprise thawing and / or homogenizing the sample prior to DNA extraction. The step of extracting DNA may comprise one or more of: mechanical disruption (e.g. sonication), chemical lysis, and purification of nucleic acids. The step of extracting DNA may comprise a step of selective lysis of the subject’s cells over microbial cells followed by DNA purification. This may be particularly useful for stool samples which frequently have high microbial content. DNA purification may use a column-based DNA purification method. The method may comprise testing the extracted DNA for DNA integrity and optionally performing an endrepair step when a DNA integrity number for the extracted DNA is above a predetermined threshold (e.g. DIN score above 8.5). Sequencing the extracted DNA may comprise preparing a library for sequencing. This may comprise fragmenting the extracted DNA, adding adapters to the fragments, and amplifying the fragments using said adapters. The method may comprise adding unique molecular identifiers (UMIs) to each DNA fragment before amplification. The method may comprise a step of enriching a library for target sequences, such as e.g. exome sequences. Sequencing the extracted DNA may comprise sequencing using a sequencing methodology associated with a sequencing error rate below one error per 10 million sequenced base pairs. Sequencing the extracted DNA may comprise sequencing using a single molecule duplex sequencing methodology, such as e.g. Nanoseqv2 or UDseq. Quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound may comprise identifying somatic mutations present in the genomic sequence data, specifically single base substitutions and / or small insertions and deletions. Quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound may furthercomprise identifying the presence and / or exposure of one or more mutational signatures associated with exposure to a microbial genotoxic compound using said identified somatic mutations.
[0028] BRIEF DESCRIPTION OF THE FIGURES
[0029] Figure 1 illustrates schematically methods of the disclosure.
[0030] Figure 2 illustrates schematically a system for implementing methods of the disclosure.
[0031] Figures 3a-i show the geographic, clinical, and molecular characterization of a colorectal cancer cohort used in Example 1.
[0032] Figure 3a shows the geographic distribution of the 981 patients across four continents and 11 countries, with an indication of the total number of cases as well as the percentage of early-onset (eo) cases below 50 years of age. Countries were colored according to their age-standardized incidence rates (ASR) per 100,000 individuals. The designations employed and the presentation of the material in this publication do not imply the expression of any opinion whatsoever on the part of the authors or their institutions concerning the legal status of any country, territory, city or area or of its authorities, or concerning the delimitation of its frontiers or boundaries.
[0033] Figure 3b shows tumor subsite distribution of the cohort across the coIorectum, with an indication of the total number of cases and the percentage of early-onset cases. Different subsites were colored according to the percentage of early-onset cases. Additional 2 cases had unspecified subsites.
[0034] Figure 3c shows scatter plots indicating the distribution of molecular subgroups across the sequenced tumors according to the total number of single base substitutions (SBS) and small insertions and deletions (indels; ID), as well as the percentage of genome aberrated (PGA). Cases for which tumor purity was insufficient to determine an accurate copy number profile or without large copy number alterations (65 / 981) were excluded from the SBS - PGA panel.
[0035] Figure 3d shows box plots indicating the distribution of SBS and ID across early-onset (under 50 years; purple) and late-onset (50 years or older; green) microsatellite stable (MSS) colorectal tumors. Statistically significant differences were evaluated using multivariable linear regression models adjusted by sex, country, tumor subsite, and tumor purity.
[0036] Figure 3e shows average mutational profiles of early-onset and late-onset MSS colorectal tumors for SBS (SBS-288 mutational context).
[0037] Figure 3f shows average mutational profiles of early-onset and late-onset MSS colorectal tumors for ID (ID-83 mutational context).
[0038] Figure 3g shows average mutational profiles of early-onset and late-onset MSS colorectal tumors for copy number alterations (CN-68 mutational context).
[0039] Figure 3h shows average mutational profiles of early-onset and late-onset MSS colorectal tumors for doublet base substitutions.
[0040] Figure 3i shows average mutational profiles of early-onset and late-onset MSS colorectal tumors for structural variants.Figures 4a-f show geographic variation of mutational signatures in microsatellite stable colorectal cancers analyzed in Example 1.
[0041] Figure 4a shows dot plot indicating the variation of mutational signature prevalence in specific countries compared to all others. Statistically significant enrichments were evaluated using multivariable logistic regression models adjusted by age of diagnosis, sex, tumor subsite, and tumor purity. Firth's bias-reduced logistic regressions were used for regression presenting complete or quasi-complete separation. Data points were colored according to the odds ratio (OR) of the enrichment, with their size representing statistical significance. P-values were adjusted for multiple comparisons based on the total number of mutational signatures considered per variant type and the total number of countries assessed, and reported as q-values. Q-values<0.05 were considered statistically significant and marked in red.
[0042] Figure 4b shows geographic distribution of the ID_J mutational signature. Countries were colored based on the signature prevalence.
[0043] Figure 4c shows geographic distribution of the SBS_F mutational signature. Countries were colored based on the signature prevalence.
[0044] Figure 4d shows volcano plots indicating the association of mutational signature activities with the age-standardized incidence rates. Statistically significant associations were evaluated using multivariable linear regression models adjusted by age of diagnosis, sex, tumor subsite, and tumor purity. P-values were adjusted for multiple comparisons based on the total number of mutational signatures considered per variant type and reported as q-values. Horizontal lines marking statistically significant thresholds were included at 0.05 (dashed orange line) and 0.01 q-values (dashed red line).
[0045] Figure 4e shows scatter plots indicating the association of the mutations attributed to the SBS88 and ID18 mutational signatures with the age-standardized incidence rates (ASR) across countries for colorectal cancer. Data points were colored based on signature prevalence, with their size indicating the total number of cases per country. Statistically significant associations were evaluated using the sample-level multivariable linear regression models used in Fig. 4d.
[0046] Figure 4f shows scatter plots indicating the association of the mutations attributed to the SBS88 and ID18 mutational signatures with the age-standardized incidence rates (ASR) across countries independently for colon and rectal cancers. Data points were colored based on signature prevalence, with their size indicating the total number of cases per country. Statistically significant associations were evaluated using multivariable linear regression models similar to those in Fig. 4d adjusted by age of diagnosis, sex, and tumor purity.
[0047] Figures 5a-e show variation of mutational signatures with age of onset in microsatellite stable colorectal cancers analyzed in Example 1.
[0048] Figure 5a shows volcano plots indicating the enrichment of mutational signature prevalence in early-onset and late-onset cases. Statistically significant enrichments were evaluated using multivariable logistic regression models forage of onset categorized in two subgroups (early-onset, <50 years of age; and late-onset, >50) and adjusted by sex, country, tumor subsite, and tumor purity. Firth's bias-reduced logistic regressions were used for regression presenting complete or quasi-complete separation. P-values were adjusted for multiple comparisons based on the total number of mutational signatures considered per variant type and reported as q-values. Horizontal lines marking statistically significant thresholds were included at 0.05 (dashed orange line) and 0.01 q-values (dashed red line).
[0049] Figure 5b shows line plots indicating mutational signature prevalence trend across ages of onset, using five different age groups. Signatures significantly enriched in early-onset or late-onset cases (as shown in Fig. 5a) were colored in purple and green, respective, whereas signatures not varying significantly with age were colored in grey.
[0050] Figure 5c shows bar plots indicating mutational signature prevalence across age groups, with indication of the total number of cases where signatures were detected. Statistically significant trends were evaluated using multivariable logistic regression models for age categorized in five subgroups (0-39, 40-49, 50-59, 60-69, >70) and adjusted by sex, country, tumor subsite, and tumor purity. Firth's bias-reduced logistic regressions were used for regressions presenting complete or quasi-complete separation.
[0051] Figure 5d shows Box plots indicating the variation in age of onset according to the presence of colibactin mutational signatures (either SBS88, ID18, or both) in all microsatellite stable cases. Statistically significant differences were evaluated using multivariable linear regression models adjusted by sex, country, tumor purity, and tumor subsite.
[0052] Figure 5e shows box plots indicating the variation in age of onset according to the presence of colibactin mutational signatures (either SBS88, ID18, or both) across tumor subsites. Statistically significant differences were evaluated using multivariable linear regression models adjusted by sex, country, and tumor purity.
[0053] Figures 6a-d show data indicating the occurrence of colibactin mutagenesis as an early event in microsatellite stable colorectal cancer evolution.
[0054] Figure 6a shows box plots indicating the fold-change of the relative contribution per sample of each signature between early clonal and late clonal single base substitutions (SBS, left) and small insertions and deletions (ID, right). Signatures that generated early clonal somatic mutations in fewer than 50 samples and also generated late clonal somatic mutations in fewer than 20 samples were excluded from the analysis. Signatures were sorted by their median fold-change.
[0055] Figure 6b shows bar plot indicating the lack of concordance between colibactin exposure status determined by the presence of colibactin-induced mutational signatures SBS88 or ID18, and the microbiome pks status. Statistical significance was evaluated using a multivariable Firth's bias-reduced logistic regression model (due to quasi-complete separation) adjusted by age of diagnosis, sex, country, tumor subsite, and tumor purity.
[0056] Figure 6c shows a distribution of the age of onset based on the detection of colibactin-positive samples using genomic and microbiome status. The genomic status is defined by the presence of SBS88 or ID18 signatures, while the microbiome status (pks) is determined by coverage of at least half of the pks island, and suggests ongoing or active pks+ bacterial infection. Statistical significance was evaluated using a multivariable linear regression model adjusted by sex, country, tumor subsite, and tumor purity.Figure 6d shows a distribution of the cases across age groups based on the detection of colibactin-positive samples using genomic and microbiome status. Genomic status, microbiome status and statistical significance are defined as in Fig. 6c.
[0057] Figures 7a-f show the variation of driver mutations with age of onset and association with colibactin mutagenesis in microsatellite stable colorectal cancers analysed in Example 1.
[0058] Figure 7a shows a bar plot indicating the prevalence of driver mutations affecting the 48 bioinformatically detected driver genes in microsatellite stable colorectal cancers. Genes were colored according to their status as known cancer driver genes for colorectal cancer, known cancer driver genes for other cancer types, or newly detected cancer driver genes.
[0059] Figure 7b shows box plots indicating the distribution of total driver mutations across early-onset (under 50 years of age; purple) and late-onset (50 or over; green) tumors. Statistical significance was evaluated using a multivariable linear regression model adjusted by sex, country, tumor subsite, and tumor purity.
[0060] Figure 7c shows a volcano plot indicating the enrichment of driver mutations in cancer driver genes in early-onset and late-onset cases. Statistically significant enrichments were evaluated using multivariable logistic regression models adjusted by sex, country, tumor subsite, and tumor purity. Firth's bias-reduced logistic regressions were used for regressions presenting complete or quasicomplete separation. P-values were adjusted for multiple comparisons based on the total number of cancer driver genes considered and reported as q-values. Horizontal lines marking statistically significant thresholds were included at 0.05 (dashed orange line) and 0.01 q-values (dashed red line).
[0061] Figure 7d shows a line plot indicating the prevalence of driver mutations in cancer driver genes across ages of onset, using five different age groups. Cancer driver genes significantly enriched in late-onset cases (as shown in Fig. 7c) were colored in green, whereas genes not varying significantly with age of onset were colored in grey.
[0062] Figure 7e shows a bar plot indicating the proportion and number of driver mutations probabilistically assigned to colibactin-induced and other mutational signatures, including single base substitutions. Driver mutations were divided into different groups, including the APC c.835-8A>G splicing-associated driver mutation, other APC driver mutations, TP53 driver mutations, and driver mutations affecting other cancer driver genes.
[0063] Figure 7f shows bar plots indicating the proportion and number of driver mutations probabilistically assigned to colibactin-induced and other mutational signatures, including small insertions and deletions (indels). Driver mutations were divided into different groups, including the APC c.835-8A>G splicing-associated driver mutation, other APC driver mutations, TP53 driver mutations, and driver mutations affecting other cancer driver genes.
[0064] Figure 8A shows a graphical representation of mutational signatures SBS88 and SBS89.
[0065] Figure 8B shows a graphical representation of mutual signature SBS_M.
[0066] Figure 8C shows a graphical representation of mutual signature ID18.Figure 9a-c show the enrichment of colibactin mutagenesis in early-onset colorectal cancers based on motif analysis.
[0067] Fig. 9a shows box plots indicating the percentage of total W[T>N]W mutations with the WAWW[T>N]W motif across different age groups. Statistically significant trend was evaluated using a multivariable linear regression model adjusted by sex, country, tumor subsite, and tumor purity.
[0068] Fig. 9b shows box plots indicating the percentage of total W[T>N]W mutations with the WAWW[T>N]W motif across samples grouped by colibactin exposure status, determined by the presence of signatures SBS88 or ID18. Statistical significance was evaluated using a multivariable linear regression model adjusted by age, sex, country, tumor subsite, and tumor purity.
[0069] Fi. 9c shows bar plots indicating the prevalence of colibactin exposure across age groups, with indication of the total numberof cases where colibactin signatures were detected. Statistically significant trend was evaluated using a multivariable logistic regression model adjusted by sex, country, tumor subsite, and tumor purity.
[0070] Figure 10a-b show enrichment of colibactin mutagenesis as an early clonal event in early-onset and late-onset colorectal cancers.
[0071] Fig. 10a shows box plots indicating the fold-change of the relative contribution per sample of each signature between clonal and subclonal single base substitutions (SBS, left) and small insertions and deletions (ID, right). Signatures that generated clonal somatic mutations in fewer than 10 samples and also generated subclonal somatic mutations in fewer than 10 samples were excluded from the analysis.
[0072] Fig. 10b shows boxplots indicating the fold-change of the relative contribution per sample of each signature between early clonal and late clonal SBS (left) and ID (right) with samples separated by age of diagnosis in early-onset (under 50 years of age; purple) and late-onset (50 or over; green). As in Fig.
[0073] 4a, signatures that generated early clonal somatic mutations in fewer than 50 samples and also generated late clonal somatic mutations in fewer than 20 samples were excluded from the analysis.
[0074] Figure 11a-d show representative microbiome and genomic profiles of colibactin-exposed samples. Fig. 11 a-d shows microbiome and genomic profiles of representative samples corresponding to the four different sample types according to colibactin exposure. The genomic status is defined by the presence of SBS88 or ID18 signatures, while the microbiome status (pks) is determined by the coverage of at least half of the pks island, and suggests ongoing and / or active pks+ bacterial infection. Circos plots display Reads Per Kilobase of transcript per Million (RPKM) values across clb genes within the pks island (left). Bar plots represent the proportion of mutations attributed to SBS88 and ID18 colibactin signatures compared to others (center), and are displayed next to mutational profiles of single base substitutions (SBS-288 mutational context) and small insertions and deletions (ID-83 mutational context) for each sample (right).
[0075] Fig. 11a shows results fora genomic+ and pks+ sample.
[0076] Fig. 11b shows results fora genomic+ and pks- sample.
[0077] Fig. 11c shows results fora genomic- and pks+ sample.
[0078] Fig. 11d shows results fora genomic- and pks- sample.Figure 12 shows driver mutations associated with colibactin mutagenesis in early-onset and late-onset colibactin positive microsatellite stable colorectal cancers. Driver mutations were divided into different groups, including the APC c.835-8A>G splicing-associated driver mutation, other APC driver mutations, TP53 driver mutations, and driver mutations affecting other cancer driver genes.
[0079] Fig. 12a shows bar plots indicating the proportion and number of driver mutations probabilistically assigned to colibactin-induced and SBS mutational signatures with samples separated by age of diagnosis in early-onset (under 50 years of age; left) and late-onset (50 or over; right).
[0080] Fig. 12b shows bar plots indicating the proportion and number of driver mutations probabilistically assigned to colibactin-induced and small insertions and deletions mutational signatures with samples separated by age of diagnosis in early-onset (under 50 years of age; left) and late-onset (50 or over; right).
[0081] DETAILED DESCRIPTION
[0082] The specific embodiments described herein are offered by way of example, not by way of limitation. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Any sub-titles herein are included for convenience only, and are not to be construed as limiting the disclosure in any way. Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the disclosure and apply equally to all aspects and embodiments which are described.
[0083] Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments of the disclosure may be readily combined, without departing from the scope or spirit of the disclosure. The features disclosed in the present description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the disclosure in diverse forms thereof.
[0084] In describing the present disclosure, the following terms will be employed, and are intended to be defined as indicated below.
[0085] “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein. As used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” oneparticular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. Other aspects and embodiments of the disclosure provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of” or ’’consisting essentially of”, unless the context dictates otherwise.
[0086] A “sample” as used herein may be a cell or tissue sample (e.g. a biopsy), a biological fluid, an extract (e.g. a protein or DNA extract obtained from the subject), from which genomic material can be obtained for genomic analysis, such as genomic sequencing (whole genome sequencing, whole exome sequencing, targeted (also referred to as “panel”) sequencing). In particular, the sample may be a blood sample, stool sample, anal swab, normal tissue, ora tumour sample. In embodiments, the sample is a stool sample. In embodiments, the sample is an anal swab. An anal swab sample may be collected, e.g. using a cytobrush, and may therefore also be referred to as an “anal cytobrush” sample. Other types of samples may be used in methods of the disclosure, such as e.g. to obtain germline sequence information or confirm the presence and / or characteristics of a tumour. A sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to making a determination (e.g. frozen, fixed or subjected to one or more purification, enrichment or extractions steps). The sample is preferably from a mammalian (such as e.g. a mammalian cell sample ora sample from a mammalian subject, including in particular a model animal such as mouse, rat, etc.), preferably from a human (such as e.g. a human cell sample ora sample from a human subject). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the genomic sequence data acquisition (e.g. sequencing) location, and / or the computer-implemented method steps may take place at a location remote from the sample collection location and / or remote from the genomic data acquisition (e.g. sequencing) location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider). A “tumour sample” refers to a sample that contains tumour cells or genetic material derived therefrom. The tumour sample may be a cell or tissue sample (e.g. a biopsy) obtained directly from a tumour. Samples as used herein may comprise modified cells or genetic material derived therefrom. Modified cells are cells comprising one or more somatic mutations. Such cells may eventually lead to the development of a tumour, but may not yet have developed into a tumour. For example, samples may be obtained from subjects who have not been diagnosed as having colorectal cancer. Such samples may nonetheless contain modified cells that may lead to the development of colorectal cancer in the future.A “normal sample” (also referred to as “germline sample”) refers to a sample that contains non-tumour or non-modified cells or genetic material derived therefrom. A normal sample may be matched to a particular sample comprising genetic material with one or more somatic mutations (e.g. stool sample, any colon or rectal tissue or cell sample), in the sense that it is obtained from the same biological source (subject or cell line) as the particular sample, but is not expected to contain the same somatic mutations. A normal sample may be a cell or tissue sample obtained from a subject, or a sample of biological fluid. A sample comprising a mixture of normal cells and other cells (or material genetic derived therefrom) may be subject to one or more processing steps, whether prior to or subsequent to the acquisition of sequence data, in order to identify sequence data that is representative of the genetic material from the normal cells (as already described above). For example, a sample comprising modified and nonmodified cells can be subject to one or more purification or selection steps to enrich the sample for nonmodified cells. Similarly, a sample comprising normal and tumour-derived cells can be subject to one or more purification steps which selectively enrich the sample for normal cells.
[0087] The term “sequence data” refers to information that is indicative of the presence and / or amount of genomic material in a sample that has a particular sequence. Such information may be obtained using sequencing technologies, such as e.g. next generation sequencing (NGS, such as e.g. whole exome sequencing (WES), whole genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing)), or using array technologies, such as e.g. SNP arrays, or other molecular counting assays. When NGS technologies are used, the sequence data may comprise a count of the number of sequencing reads that have a particular sequence. When non-digital technologies are used such as array technology, the sequence data may comprise a signal (e.g. an intensity value) that is indicative of the number of sequences in the sample that have a particular sequence, for example by comparison to an appropriate control. Sequence data may be mapped to a reference sequence, for example a reference genome, using methods known in the art (such as e.g. BWA-MEM (Li & Durbin, 2009; Li 2013). Thus, counts of sequencing reads or equivalent non-digital signals may be associated with a particular genomic location. Further, a genomic location may contain a mutation, in which case counts of sequencing reads or equivalent non-digital signals may be associated with each of the possible variants (also referred to as “alleles”) at the particular genomic location. The process of identifying the presence of a mutation at a particular location in a sample is referred to as “variant calling”, and can be performed using methods known in the art (such as e.g. the GATK Mutect2, gatk.broadinstitute.org / hc / en-us / articles / 360037593851-Mutect2; DupCaller see Cheng Y. et al. 2025; etc.). For example, sequence data may comprise a count of the number of reads (or an equivalent non-digital signal) which match a germline (also sometimes referred to as “reference”) allele at a particular genomic location, and a count of the number of reads (or an equivalent non-digital signal) which match a mutated (also sometimes referred to as “alternate”) allele at the genomic location. The term “mutation” refers to a difference in a nucleotide sequence (e.g. DNA or RNA) in a sample compared to a reference. For example, a mutation may be a single nucleotide variant (SNV), multiple nucleotide variants, a deletion mutation, an insertion mutation, a translocation, a missense mutation, atranslocation, a fusion, etc. Mutations may be identified using sequence data. An "indel mutation" (or simply “indel”) refers to an insertion and / or deletion of bases in a nucleotide sequence (e.g. DNA or RNA) of an organism.
[0088] Within the context of the present disclosure, a mutation is typically a somatic mutation, unless the context indicates otherwise. A somatic mutation is a mutation that is acquired during the lifetime of a subject but that is not present in the germline material of the subject. A “somatic mutation” is a mutation that is present in a tumour or modified cell (or genetic material derived therefrom), but not in a corresponding (matched) normal (also referred to as germline) or non-modified cell. Practically speaking, a somatic mutation may be identified by comparing sequence data from a sample comprising modified cells or genetic material derived therefrom to a reference sequence or sequence data from a sample that is assumed not to comprise said modified cells or genetic material derived therefrom. A composition as described herein may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable carrier, diluent or excipient. The pharmaceutical composition may optionally comprise one or more further pharmaceutically active polypeptides and / or compounds. Such a formulation may, for example, be in a form suitable for intravenous infusion.
[0089] As used herein "treatment" refers to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment. The systems and methods described herein may be implemented in a computer system, in addition to the structural components and user interactions described. As used herein, the term “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above-described embodiments. For example, a computer system may comprise a processing unit such as a central processing unit (CPU) and / or graphical processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. Preferably the computer system has a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer. The methods described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described herein. As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media. The methods described herein may be computer implemented unless indicated otherwise. Indeed, the complexity associated with obtaining mutationalprofiles from samples, such as e.g. by analysing sequence data (e.g. thousands to hundreds of thousands of reads per sample) or somatic mutations lists derived therefrom (comprising multiple hundreds or thousands of mutations) and classifying each of these by analysis of the sequence of the mutation and its context, as well as calculating signature exposures using such mutational profiles, is far beyond the capability of the human mind.
[0090] According to the present disclosure, metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound quantified from nucleic acid sequence data, are used to determine the risk of an individual developing colorectal cancer.
[0091] Metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound can be metrics that quantify the presence or activity of a mutational signature associated with exposure to a microbial genotoxic compound. Metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound can be metrics that quantify the presence, number or proportion of mutations associated with exposure to a microbial genotoxic compound. Such metrics may be or may be derived from metrics that associate individual mutations with the exposure to a microbial genotoxic compound (e.g. metrics that assign mutational signatures to individual somatic mutations, which can be used to identify individual somatic mutations that are associated with exposure to a mutational signature associated with exposure to a microbial genotoxic compound). The microbial genotoxic compound may be colibactin. Colibactin is a bacterial toxin that is produced by multiple different bacteria including Escherichia coli and other Enterobacteriaceae ("enteric bacteria"), including multiple species that are commonly found in the human gut microbiome. Colibactin forms DNA interstrand cross-links by alkylation of adenine moieties on opposing DNA strands. Mutational signatures associated with exposure to colibactin, and mutations associated with exposure to colibactin are described herein. Colibactin has chemical formula C37H38N8O7S2 and is available under CHEBI accession number 156303.
[0092] Metrics that quantify the presence, number or proportion of mutations associated with exposure to a microbial genotoxic compound may be selected from: (i) the presence of mutations, percentage of mutations, proportion of mutations or number of mutations associated with one or more mutational signature associated with exposure to a microbial genotoxic compound (as described herein, such as e.g. SBS88, SBS89, ID18 and / or SBS_M); (ii) the percentage or proportion of total W[T>N]W mutations that are within a WAWW[T>N]W motif, where W is A or T, and N is A, C, G or T; (iii) the presence of a splicing variant c.835-8A>G in the APC gene; (iv) the presence of one or more protein-truncating mutations in the APC gene; and (v) the presence of indels in the APC gene associated with 1 bp T or A deletions and / or associated with indel signature ID18, where the reference to indel signature ID18 encompasses a corresponding signature as defined herein. Protein-truncating mutations in the APC gene may be single base substitutions or indels. Protein-truncating mutations in the APC gene may be mutations associated with mutational signatures SBS88, SBS89 or ID18, where the reference to these signatures encompass corresponding signatures as described herein. APC is a tumour suppressorgene that acts as an antagonist of the WNT signalling pathway. It is also known as adenomatous polyposis coli protein, APC regulator of WNT signalling pathway, GS, DP2, DP3, BTPS2, DESMD, DP2.5, PPP1R46. The sequence of human APC is available under Gene ID 324 (NG_008481.4 RefSeqGene).The splicing variant c.835-8A>G in the APC gene is a splice variant located in intron 8 of APC, with sequence available as NM_000038.5: c.835-8A>G (Homo sapiens APC, WNT signaling pathway regulator (APC), transcript variant 3, mRNA) and NM_000038.6(APC):c.835-8A>G, clinVar Variation ID: 418007 (www.ncbi.nlm.nih.gov / clinvar / variation / 418007 / ).Throughout this document, the term “signature” (or “mutational signature”, which encompasses SBS and indel signatures) refers to a set of weights that characterise the imprint of a mutational process on the mutational profile of cancer genomes in terms of the classes and proportions of mutations that are generated by the mutational process. In other words, a mutational signature (SBS or indel signature) comprises, for each of a plurality of categories of indels or SBS, a proportion. The plurality of categories of indels / SBS is the same as the plurality of categories (also referred to herein as “classes”) of indels / SBS present in an indel / SNS profile associated with a sample or subject. An indel / SBS profile may be obtained by classifying a plurality of indels / SBS identified in a genome (e.g. indels / SBS identified in a sample or subject) between a predetermined plurality of classes (i.e. each indel / SBS of the plurality of indels / SBS being assigned a single class of the plurality of classes).
[0093] A mutational signature catalogue can be extracted from a plurality of mutational profiles each associated with respective samples. A mutational signature catalogue can comprise mutational signatures extracted separately from a plurality of respective cohorts of mutational profiles each associated with a respective patient sample. In other words, a mutational signature catalogue can comprise mutational signatures extracted from one or more cohorts of patients, each of which with a sample with a characteristic mutational profile. Mutational signatures can be extracted from cohorts of samples, by identifying characteristic combination of mutation classes (also referred to herein as “categories” or “channels”) that best explain the mutational profiles of the samples in the cohort. This process also results in the quantification of “exposures” to each of the signatures, which quantify the extent of the effect of the respective signatures on the respective mutational profiles. Signatures are defined for a specific type of mutation, e.g. SBS or indels. This is because the classification of mutations is inherently linked to the type of mutation considered. Examples of classifications of single base substitutions are described in Bergstrom et al. 2019 and Islam et al. 2022. For example, SBS mutations are commonly classified based on the mutation content (i.e. which base is mutated to what base) and trinucleotide context (i.e. the identity of the 5’ and 3’ nucleotides surrounding the mutated nucleotide). This typically results in 96 classes, one for each unique combination of content (6 possible mutation contents) and trinucleotide context (16 possible trinucleotides context for each mutation content), and is therefore commonly referred to as SBS-96or96 channels classification of SBS. Such a classification is described in e.g. Alexandrov et al. 2013, and has been widely used since. This is the classification used in the COSMIC database (cancer.sanger.ac.uk / signatures / ; Alexandrov et al. 2020; a widely used database of mutational signatures). A signature using such a classification would consist of a set of weights, one for each of the classes defined using this content + trinucleotide context approach . Other SBS mutationclassifications can take into account only the mutation type, i.e. differentiating only between 6 categories: C:G >A:T, C:G > G:C, C:G > T:A, T:A>A:T, T:A> C:G, and T:A> G:C. This can be referred to as SBS-6 classification. Again, a signature using such a classification would consist of a set of weights, one for each of the classes defined using this content only approach. Other SBS mutation classifications can take into account one or more further features of the mutations, in addition to mutation content and trinucleotide context. For example, transcriptional information can be taken into account, such as e.g. transcriptional strand bias. For example, mutations in transcribed regions of the genomes can be further subclassified depending on whether they occur on the transcribed or nontranscribed stand. Therefore, mutations can be classified as i) transcribed, (ii) un-transcribed, or (iii) non-transcribed; or i) transcribed, (ii) un-transcribed, (iii) bidirectional or (iv) non-transcribed I unknown. Specifically, SBS mutations classified by content and trinucleotide context can be further classified depending on whether the SBS is in non-transcribed / intergenic DNA, on the transcribed strand of a gene, or on the untranscribed strand of the gene, resulting in 288 classes, an expansion of the 96 classes that incorporates transcriptional information. Again, a signature using such a classification would consist of a set of weights, one for each of the classes defined using this content + trinucleotide context + transcriptional information approach. As another example, additional 5’ and 3’ adjacent context can be used. For example, considering two bases 5' and two bases 3' of a mutation results in 256 possible classes for each SBS (16 types of two 5' bases * 16 types of two 3' bases), resulting in a classification with 1536 possible channels (256 different contexts for each of 6 different contents of mutations), also referred to as SBS-1536. Again, a signature using such a classification would consist of a set of weights, one for each of the classes defined using this content + 4 nucleotides (two bases 5' and two bases 3') context approach. Classifications that further distinguishing between mutations with different characteristics in the same class can be referred to as having increased resolution. Classifications (and resulting mutational profiles and signatures) with higher resolution can be trivially mapped to corresponding classifications I profiles I signatures at lower resolution. For example, SBS-1536 can be mapped to SBS-288, SBS-96 and SBS-6, SBS-288 can be mapped to SBS-96 and SBS-6, and SBS-96 can be mapped to SBS-6. By contrast, indels can differ from each other in length, sequence context (where the context that is informative may not be limited to a single base), sequence content, repeatedness and the presence or absence of microhomology. Each of these can be taken into account to generate a plurality of classes of indels, and a signature using such a classification would consist of a set of weights, one for each of the classes defined (where the total number of classes depends on how classes based on unique combinations of these features are combined). Examples of indel classifications are described below and in Bergstrom et al. 2019. The most commonly used indel classification approach comprises 83 channels, is the approach used in the COSMIC database, and is described further below. Note that throughout this disclosure, references to mutations of a T base encompass mutation of the corresponding A base (and vice-versa), and similarly references to mutations of a C base encompass mutation of the corresponding G base (and vice versa). Indeed, as the skilled person understands, and example of a SBS in which a C base mutates to an A base refers to a C:G base pair mutating to an A:T base pair. This can be denoted as C:G>A:T, C>A (pyrimidine base notation) or G>T (purine base notation). By convention, this is typically denoted as C>A as thepyrimidine base notation has been adopted as a community standard. Similarly, when looking at SBS mutations defined by mutation content and trinucleotide context, ACG:TGC > AAG:TTC can be written as ACG > AAG using the pyrimidine base and as CGT > CTT using the purine base (i.e., the reverse complement sequence of the pyrimidine classification), the former having been adopted as community standard.
[0094] A mutational catalogue summarising somatic substitutions associated with a sample or a group of samples may be referred to as a “mutational profile”. The present disclosure relates in particular to indel mutations and single base substitution mutations. As such, a mutational profile may also be referred to as an indel profile or SBS profile in the present disclosure. An indel mutational profile comprises the number of mutations present in a sample within each of a plurality of indel categories. An SBS mutational profile comprises the number of mutations present in a sample within each of a plurality of SBS categories. An indel / SBS mutational profile can be seen as a summary of a list of indel / SBS mutations present in a sample, categorised according to a predetermined set of indel / SBS categories (also referred to as “channels”). A set of mutational signatures may be referred to as a “mutational signatures catalogue”. A set of mutational signatures only contains mutational signatures of the same type that are classified using the same classification scheme, i.e. SBS signatures or indel signatures. Thus, an indel mutational signature catalogue is a set of indel mutational signatures, and a SBS mutational signature catalogue is a set of SBS mutational signatures.
[0095] The methods described herein relate at least in part to determining an indication of whether an indel mutational signature in a catalogue of indel mutational signatures is present in the sample, and / or an indication ofwhethera SBS mutational signature in a catalogue of SBS mutational signatures is present in the sample. The presence of an indel / SBS signature in the sample can be identified by comparing the proportions in the indel / SBS signature (also referred to as “weights”), to the numbers or proportions of indels / SBS in the corresponding classes in the indel / SBS profile of the sample. Proportions may be numbers that sum to 1 over the plurality of categories. Thus, the proportions may indicate the pattern of indel / SBS types that typically arise in a genome when the signature is present. Any individual indel / SBS profile (i.e. any sample or subject) may comprise proportions of mutations in each of a plurality of classes of mutations based on which more indel / SBS signatures were defined, which proportions indicate the presence of one or more indel / SBS signatures. In other words, an indel / SBS profile may represent a combination of indel / SBS signatures. Importantly, the categories in an indel / SBS signature and in an indel / SBS profile to be compared are the same, and are defined using an indel / SBS classification scheme as described herein.
[0096] Determining whether a mutational signature is present in a sample may be based on exposures of the mutational signatures in a mutational signature catalogue (also referred to as “mutational signature metrics”). Methods for determining the exposure to a mutational signature are described in e.g. Alexandrov et al., 2020; Degasperi et al., 2020; Degasperi et al., 2022; Diaz-Gay et al. 2023. In particular, the determination of the exposure to one or more mutational signatures in a set may beperformed by identifying the matrix E that satisfies C~PE where C is a mutational catalogue for one or more samples for which exposure is to be determined, P is a signature matrix comprising the one or more mutational signatures for which exposure is to be determined, and E is an exposure matrix. This may be performed using any NMF implementation with either a Kullback-Leibler divergence or Frobenius norm objective function. Examples include the implementation in the R package NNLM (see cran.nexr.com / web / packages / NNLM / vignettes / Fast-And-Versatile-NMF.html). For example, for each sample a plurality of signature fitting processes (i.e. identifying the matrix E that satisfies C~PE) can be performed using bootstrapped data (perturbing the mutational profile by randomised re-sampling with replacement) in order to obtain a distribution of exposures for the sample, where a point estimate for the exposure is obtained as the median of the distribution. Mutational signatures with exposure of fewer than a predetermined number of mutations (e.g. 50 mutations) may be set to 0 to increase signature detection specificity. Instead or in addition to this, mutational signature exposures with a p-value below a threshold can be to 0 to increase specificity. The p-value may be an empirical p-value estimated from the bootstrapped data. For example, if more than 5% of the bootstrapped exposure values (empirical p-value 0.05) are below the threshold of 5% of total mutations in the sample, then the signature exposure may be set to 0. The determination of the similarity between two mutational profiles, a mutational profile and a mutational signature, or two mutational signatures may be performed by calculating the cosine similarity between the two SBS / indel profiles or SBS / indel mutational signatures. The cosine similarity between two profiles or signatures can be calculated as: sim{S,M') =
[0097]
[0098] S and M are equally-sized vectors with nonnegative components being the respective mutational profiles or mutational signatures (e.g. S being that of a sample and M that of a reference profile such as e.g. a reconstructed profile or signature).
[0099] Thus, a profile (or indel profile, SBS profile, as the case may be) refers to a description of a particular sample I genome in terms of the number of mutations of a particular type (indels or SBS) within each of a plurality of categories (classes, channels), as described herein. The numbers can be normalised to sum to 1 , and therefore represent proportions of mutations of a particular type (indels or SBS) within each of a plurality of categories (classes, channels) as described herein. A signature (or mutational signature or indel mutational signature) refers to a pattern of occurrence of indels / SBS within each of the plurality of categories (channels) that has been previously observed in a plurality of samples I genomes. A mutational profile can result from a combination of one or more mutational signatures. In practice, a profile can be expressed as the number or proportion of all mutations of a particular type (i.e. indels or SBS) observed in a sample I genome in each of a predetermined plurality of categories as described herein. A signature can be expressed as a plurality of weights for each of a predetermined plurality of categories as described herein. The weights can sum to 1 and represent the proportions of the different classes (categories) of mutations that the mutational process associated with the signature is expected to generate, or the probability that the mutational process will cause mutations in the respective categories. A profile can therefore be expressed as a linear combination of one or more signatures. In other words, one or more signatures of a previously identified set of signatures may bedetermined to be present in a sample / genome when the one or more signatures are the signatures of the previously identified set that are such that the linear combination of the one or more signatures best recapitulates the profile of the sample / genome. Additional filters may be applied to determine that a signature is present, such as e.g. filters that apply to the exposure or number of mutations attributed to the signature. The number of mutations in a sample that are attributed to a particular mutational signature may be determined based on the exposure to the mutational signature and the total number of mutations in the mutational profile for the sample, as described further below.
[0100] Embodiments of the disclosure comprise a characterisation of a DNA sample in terms of the indel mutational signatures and / or SBS mutational signatures present in the sample. In embodiments, this is performed by a computer-implemented method or tool that takes as its inputs sequence data from the sample and an indel mutational signatures catalogue comprising one or more indel mutational signatures and / or an SBS mutational signatures catalogue comprising one or more SBS mutational signatures, and produces as output an indication of whether one or more of the indel mutational signatures and / or SBS signatures in the respective catalogues is / are present in the sample. Such an indication, for a particular signature, may referred to as a “mutational signature metric”. A mutational signature metric for a signature may be an exposure to the signature, and / or a number of mutations of the type catalogued in the signature (i.e. indels or SBS) identified in the sample that are attributed to the signature. The exposure can be obtained from the number of mutations attributed to the signature (and vice-versa) by dividing the latter by the total number of mutations of the type catalogued in the signature (i.e. total number of indels or SBS) identified in the sample.
[0101] In embodiments, the computer-implemented method or tool may take as its inputs a list of somatic indel mutations and / or a list of somatic SBS mutations generated from sequence data associated with a stool sample. These somatic mutations can then be analysed to determine the value(s) of the one or more mutational signature metrics. The computer-implemented method or tool may take as its inputs sequence data associated with a stool sample, and may use this data to generate a list of somatic indel mutations using any indel mutation calling process known in the art. The computer-implemented method or tool may take as its inputs sequence data associated with a tumour sample, and may use this data to generate a list of somatic SBS mutations using any SBS mutation calling process known in the art. These somatic indel / SBS mutations can then be analysed to determine the value(s) of the one or more mutational signature metrics. A list of somatic indel / SBS mutations may be obtained by identifying mutations present in sequence data associated with a stool sample, and removing or otherwise excluding indel / SBS mutations that are present or assumed to be present in a corresponding germline genome. Mutations that are present in a corresponding germline genome may be identified by identifying the mutations present in a germline sample obtained from the same subject (also referred to as a “matched germline” or “matched normal” sample). Thus, the computer-implemented method or tool may further take as input sequence data associated with a matched germline sample. Mutations that are assumed to be present in a corresponding germline genome may be identified by identifying germline variants that are present in a reference genome or set of reference genomes. A referencegenome or set of reference genomes may be obtained from one or more reference samples that are not (or not all) matched normal samples. For example, the reference sample(s) may be process matched, or may comprise a plurality of normal (i.e. non-tumour / non-modified) samples not all of which are matched to the sample for which a somatic indel mutational profile is determined (e.g. pooled normal samples may be used as references for a plurality of tumour samples). A reference genome or set of reference genomes may be obtained from one or more databases. For example, a reference genome may be used and all mutations compared to this reference genome may be assumed to be somatic mutations. Alternatively, a set of reference genomes may be obtained from a database as a catalogue of known germline mutations in one or more populations (e.g. a genetic variation database such as dbSNP www.ncbi.nlm.nih.gov / snp / , 1000 genomes www.internationalgenome.org / , etc.). The use of a matched normal sample advantageously provides greatest certainty that the mutations identified in the DNA from the tumour sample are somatic mutations. The use of pooled normal samples comprising a set of unmatched normal samples may provide similar (though less precise information) and may be useful e.g. when sequencing resources are limited. Compared to the use of a matched normal sample, this may risk excluding more somatic mutations as seemingly germline mutations while missing private germline variants. The use of a reference genome or set of reference genome advantageously does not require the acquisition and analysis of a separate normal sample. However, the reference genome or set of reference genome is unlikely to capture all germline mutations present in the subject, and to include mutations that are in fact somatic in the subject. This is particularly true if a single reference genome is used rather than a collection capturing common sequence variation. Thus, this may result in a less accurate identification of somatic mutations.
[0102] The computer-implemented method or tool may take as its inputs an indel and / or an SBS profile. As explained above, an indel profile comprises a count or proportion of indel mutations in each of a plurality of categories, obtained by categorising each somatic indel mutation identified in a sample in a class of a classification as described herein (i.e. a classification comprising the plurality of categories). Similarly, an SBS profile comprises a count or proportion of SBS mutations in each of a plurality of categories, obtained by categorising each somatic SBS mutation identified in a sample in a class of an SBS classification such as the trinucleotide (96 channels) classification described further herein (i.e. a classification comprising the plurality of categories). The computer-implemented method or tool may take as its inputs at list of somatic indel mutations or sequence data associated with a stool sample, and may use this data to generate an indel profile using a classification as described herein. Similarly, the computer-implemented method or tool may take as its inputs a list of somatic SBS mutations or sequence data associated with a stool sample, and may use this data to generate an SBS profile using e.g. the trinucleotide (96 channels) classification described above. The indel / SBS mutations that are included in the indel / SBS profile may be all somatic indel / SBS mutations identified in a sample. The indel / SBS mutations that are included in the indel / SBS profile may be all somatic indel mutations identified in sequence data from a sample. The sequence data may be genome-wide (e.g. from whole genome sequencing or shallow whole genome sequencing), or cover a substantial portion of thegenome, such as e.g. exome wide (e.g. from whole exome sequencing) or from a targeted panel assay targeting a large number of regions (e.g. genes) such as e.g. 100 or more genes.
[0103] Embodiments of the disclosure use mutational signature metrics for one or more SBS signatures and / or one or more indel signatures, that have been previously determined, such as e.g. using a method as described herein. For example, the methods of the present disclosure may start from a list comprising mutational signatures metrics (e.g. exposures and / or number of mutations) associated with a plurality of signatures.
[0104] Embodiments of the methods described herein using indel / SBS profiles are applicable to indel / SBS profiles that have been obtained from any sequencing approach that allows identification of mutations over a substantial part of the genome. This may be achieved through whole genome sequencing (WGS), whole exome sequencing (WES) or any capture sequencing approach (i.e. targeted / bait based sequencing) that captures a portion of the genomes such as e.g. at least 10% of the genome, at least 20% of the genome, at least 30% of the genome, at least 40% of the genome, at least 50% of the genome, at least 60% of the genome, at least 10% of the exome, at least 20% of the exome, at least 30% of the exome, at least 40% of the exome, at least 50% of the exome, at least 60% of the exome, at least 100 genes, at least 200 genes, at least 300 genes, at least 400 genes, at least 500 genes, and / or at least 1000 genes. Further, when incomplete genome sequencing is used, some of the genome may be imputed based e.g. on comparison with corresponding sequences in more complete profiles. In embodiments, the mutational catalogue and / or the mutational signatures catalogue has been determined from whole genome sequencing data or whole exome sequencing data.
[0105] Any set of mutational signatures (indel mutational signatures catalogue or SBS mutational signatures catalogue) may have been extracted from a plurality of indel / SBS profiles comprising counts of mutations for each of the plurality of categories as described herein, using non-negative matrix factorization (NMF). NMF may be used with Kullback-Leibler divergence (KLD) optimization, repeated bootstrapping (such as e.g. at least 500 bootstraps), and removal of local minima. For example, given a matrix of catalogs C, nonnegative matrix factorization (NMF) may be applied to 500 matrices C, bootstrapped from C. The NMF may be solved using an algorithm that optimizes the Kullback-Leibler divergence (KLD). Solving the NMF may produce a matrix of signatures S and a matrix of exposures E for each NMF run, such that C « SE. The NMF may be repeated a number of times (such as e.g. at least 500 times) for each bootstrap matrix, using random initializations. A set of solutions may be selected solutions that have a final KLD within a predetermined percentage (e.g. 0.1%) of the best solution found (the solution with the lowest KLD). Point estimates of exposures may be obtained as the median of the exposures obtained from bootstrapping.
[0106] A metric that quantifies the presence of a mutational signature associated with exposure to a microbial genotoxic compound may be a Boolean metric, i.e. a metric that takes a first value when the mutational signature is present and a second value when the mutational signature is absent. A Boolean metricmay be obtained by quantifying the activity of the signature (also referred to herein as “exposure” of the signature) in the sample and / or a confidence interval around the estimated activity of the signature in the sample and comparing the activity or confidence interval to a respective predetermined threshold. For example, a mutational signature may be determined to be present in a sample when its activity in a sample is above 0, and / or when both boundaries of a confidence interval around the estimated activity in the sample are positive. A metric that quantifies the presence of a mutational signature may be the activity of the signature (also referred to as “exposure”), or a number or proportion of mutations associated with the signature. The number of mutations in a sample that are attributed to a particular mutational signature may be determined by multiplying the activity of the mutational signature and the total number of mutations in the sample. Alternatively, the number of mutations in a sample that are attributed to a particular mutational signature may be determined by attributing signatures to individual somatic mutations in the sample as the signature with the highest probability of being responsible for the mutation. The probability that a mutational signature is responsible for any mutations in a particular category in a sample can be obtained by multiplying the general probability of the signature causing a particular category of mutations (e.g. mutations in a specific mutational context) obtained from the mutational signature profile (i.e. weight for the category), by the activity of the signature in the sample (obtained from the signature activities). This can then be divided by the total number of mutations in the sample, such that the probabilities sum to 1 across all mutational signatures identified in a sample (if the activities themselves are not normalised), or by the total number of mutations in the respective category. Multiplying the probability by the total number of mutations in the category (e.g. total number of mutations corresponding to the specific mutational context), obtained from the mutational profile of the sample (or a reconstructed version thereof which can be obtained e.g. by, for each category, multiplying the total number of mutations in the sample by a weighted sum of the weights of each of the mutational signatures identified to be present in the sample for the category, weighted by their activity) provides the total number of mutations in the category that are attributed to the signature. These can be summed across categories to provide the total number of mutations attributed to a mutational signature.
[0107] The total number of mutations in a sample may be obtained as the total number of mutations that are represented in a mutational profile for the sample. A mutational profile refers to a set of counts of the number of mutations identified in a sample in each of a plurality of categories. The plurality of categories may each be categories of mutations of a predefined type, such as e.g. single base substitutions or indels. Thus, the total number of mutations in a sample may be the total number of single base substitutions in the sample, or the total number of indels in the sample, depending on whether a SBS or an indel signature is considered. The plurality of categories typically corresponding to the plurality of categories for which signature weights (also referred to herein as “proportions”) are provided. A total number of mutations in a sample may also be referred to as “total mutational burden” or simply “mutational burden” of the sample. All mutations considered may be somatic mutations or mutations assumed to be somatic mutations. Somatic mutations may be identified by identifying mutations in a sample assumed and removed mutations assumed to be germline mutations. Mutations assumed to begermline mutations may be mutations identified as germline mutations from a reference sample, e.g. a sample comprising normal DNA from the subject, or mutations from a database or dataset of germline variation (e.g. a SNP database). Alternatively, all mutations compared to a reference genome may be assumed to be somatic mutations. The reference genome may be any reference genome known in the art, such as e.g. GRCh38.
[0108] The activity of a mutational signature may also be referred to as “exposure”. The activity of a set of one or more mutational signatures may be determined by identifying the vector E that satisfies C~PE where C is a mutational profile for the sample, and P is a signature matrix comprising the one or more mutational signatures for which activity is to be determined. The signature matrix may comprise a plurality of previously identified mutational signatures comprising the signature(s) for which activity is to be determined. For example, the signature matrix may comprise a plurality of SBS (single base substitution) signatures in the COSMIC database, or a plurality of indel signatures in the COSMIC database. The signature matrix may comprise all SBS signatures in the COSMIC database or a subset thereof, or all indel signatures in the COSMIC database or a subset thereof. A subset of mutational signatures may be identified using an assignment process as described in Senkin, S. MSA: reproducible mutational signature attribution with confidence based on simulations. BMC Bioinformatics 22, 540 (2021). The subset may be identified for the particular sample, for example by determining the activity of all signatures in a set (e.g. all COSMIC signatures) then calculating the L2 similarity between the mutational profile and a reconstructed mutational profile obtained by linear combination of the signatures weighted by the determined exposures, normalised by its L2 norm, and iteratively removing the lowest activity signature from the set if the removal increases the L2 similarity by less than a predetermined threshold. In embodiments, a plurality of signature fitting processes (i.e. identifying the vector E that satisfies C~PE) can be performed for a sample using bootstrapped data (perturbing the mutational profile by randomised re-sampling with replacement or by sampling from a multinomial distribution Multinomial(M,p1,..., pm), where M is the total mutational burden in the sample, and probabilities pi are normalised mutation counts for each mutation category in the mutational profile) in order to obtain a distribution of exposures for the sample, where a point estimate for the exposure is obtained as the median of the distribution. The total mutational burden in the sample may be quantified as the total number of mutations of the type considered in the signatures, such as e.g. SBS for SBS signatures, and indel for indel signatures. A signature activity and / or confidence interval can be obtained as described in Senkin, S. MSA: reproducible mutational signature attribution with confidence based on simulations. BMC Bioinformatics 22, 540 (2021). Some embodiments of the present disclosure make use of single base substitution (SBS) signatures. Single base substitution signatures characterise the imprint of a mutational process on the mutational profile of cancer genomes in terms of the types and proportions of single base substitutions that are generated by the mutational process. Single base mutational signatures comprise weights for each of a plurality of categories of single base substitution mutations. The plurality of categories for SBS mutations may differ by the nature of the single base substitution (i.e. which base is mutated to which base, or in other words which base is mutated and whether the mutation is a transversion ora transition), and its immediate 3’ and 5’ context (i.e. the basesimmediately 3’ and 5’ of the SBS). The plurality of categories of SBS mutations can comprise 96 categories of single base substitutions defined by mutation content (also referred to herein as “substitution type”) and trinucleotide context (i.e. the base that is mutated and the surrounding 5’ and 3’ nucleotides). A substitution type is selected from C>A, C>G, C>T, T>A, T>C, and T>G. A trinucleotide context is defined by the nucleotides immediately 5’ and 3’ to the substitution. The plurality of categories of single base substitutions may together include all possible combinations of substitution type and trinucleotide context (96 channels signatures). The plurality of categories may be as defined in Table 1. Catalogues of mutational signatures identified in pan-cancer genomes have been described in Alexandrov et al. 2020 and in Degasperi et al., 2022. The signatures described in Alexandrov et al.
[0109] 2020 are available at cancer.sanger.ac.uk / cosmic / signatures / (hereafter “COSMIC” signatures / database). The signatures described in Degasperi et al., 2022 are available at signal.mutationalsignatures.com / explore / cancer (hereafter “Signal” signatures / database). These include signatures that correspond to the vast majority of the COSMIC signatures, and additional signatures. Corresponding signatures in the Signal and COSMIC databases can be assigned the same name (e.g. SBS1 to SBS95 signatures are assigned the same name in the two databases) but do not necessarily have the same name (e.g. in SBS95 to SBS99 in Signal tere are signatures where the same name in signal does not correspond to the equivalent signature in COSMIC). However, methods to identify corresponding signatures based on e.g. cosine similarity are known in the art and the skilled person would be able to identify signatures in one catalogue of mutational signatures that correspond to signatures in another catalogue.
[0110] A SBS mutational signature associated with exposure to a microbial genotoxic compound may be a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature is characterised by: (a) a higher proportion of T>C mutations in ATA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in ATA context, T>C mutations in ATT context, T>C mutations in TTT context, and T>G mutations in TTT context than any other categories of single base substitutions. The mutational signature may be SBS88, or a corresponding signature obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS88. SBS88 is a mutational signature associated with exposure to colibactin. A SBS mutational signature associated with exposure to a microbial genotoxic compound may be a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature is characterised by a higher proportion of T>G mutations in GTG context and C>T mutations in ACA context than any other categories of single base substitutions. The mutational signature may be SBS89, or a corresponding signature obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS89.SBS89 is a signature of unknown aetiology that appears to be most active in the first decade of life. A SBS mutational signature associated with exposure to a microbial genotoxic compound may be a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature is characterised by (a) a higher proportion of T>C mutations in CTA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in CTA context, T>C mutations in TTA context, and T>A mutations in CTA context than any other categories of single base substitutions. The mutational signature may be SBS_M, or a corresponding signature obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS_M. SBS_M is a mutational signature that has been identified as enriched in early-onset colorectal cancer cases. While SBS_M does not have a known aetiology at this point, the present inventors saw a similar enrichment in distal colon and, especially, rectal cases, as the original colibactin-associated SBS88 and ID18 signatures, and therefore postulated that this signature may also be associated with exposure to a microbial genotoxic compound that is also associated with early onset colorectal cancer. SBS88, SBS89 and SBS_M are defined in Table 1, and graphically illustrated on Fig. 8A-B. Corresponding signatures may comprise the same signatures in any version of the COSMIC database (cancer.sanger.ac.uk / signatures / ; Sondka et al. 2024) or the Signal database (signal.mutationalsignatures.com / explore / main / cancer / signatures; Degasperi et al.
[0111] 2020).
[0112] >
[0113] >
[0114] >
[0115] >
[0116] >
[0117] >
[0118] >
[0119] >
[0120] >
[0121] >
[0122] >
[0123] >
[0124] >
[0125] >
[0126] >
[0127] >
[0128] >
[0129]
[0130] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0131]
[0132] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0133]
[0134] >
[0135] >
[0136] >
[0137]
[0138] Table 1. Single base substitution signatures described herein. Definitions from COSMIC v3.4 for SBS_88 and SBS_89- October 2023, based on GRCh37. Proportions (also referred to as “weights”) for each of the plurality of categories are provided.
[0139] Some embodiments of the present disclosure make use of indel signatures. Indel signatures characterise the imprint of a mutational process on the mutational profile of cancer genomes in terms of the types and proportions of indels that are generated by the mutational process. Indel signatures comprise weights for each of a plurality of indel mutation categories.
[0140] A mutational catalogue summarising a list of mutations comprising both somatic insertions and deletions associated with a sample or group of samples may be referred to as an “indel profile”. A classification to obtain indel profiles was described in Alexandrov et al., 2020, and is illustrated in Table 2 (first two columns). The classification is sometimes referred to as the “COSMIC” indel signatures classification as it is described in the COSMIC database at cancer.sanger.ac.uk / cosmic / signatures / ID. According to the COSMIC classification, indels are classified as insertions or deletions, then if they are of a single base, as C or T, then according to the length of the mononucleotide repeat tract in which they occur (e.g. deletion of a T in a repeat tract of 5 Ts - resulting in a sequence of 4 Ts). Longer indels are classified as occurring at repeat regions (also referred to as “repeat mediated” indels) or with overlapping microhomology at deletion boundaries (whether 5’ or 3’), and according to the size of the indel, repeat and microhomology (e.g. deletion of a TC where the immediately following sequence is a TC is a deletion of 2 bp with repeat size 2, deletion of a TC where the immediately following sequence is TA is a deletion of 2bp with a homology size of 1). This classification scheme results in a total of 83 classes (also referred to as “channels”), listed in Table 2 below. Thus, the plurality of categories for indel mutations may differ by: whether they are deletions of 1 bp, insertions of 1 bp, deletions of more than 1bp at repeats, insertions of more than 1bp at repeats, or deletions with microhomology. Such categories may further differ by the base of the indel being a purine or pyrimidine for 1bp indels, the length of the indel for >1bp indels at repeats, and the length of the deletion for deletions with microhomology. Instead or in addition to this, such categories may further differ by the homopolymer length for 1bp indels, the number of repeat units for >1bp indels at repeats, and the microhomology length for deletions with microhomology. The plurality of categories of indels may include: (i) a plurality of categories of 1bp C deletions that differ by the homopolymer length in which the deletion occurred (e.g. this may comprise separate categories for homopolymer lengths of 1, 2, 3, 4, 5 and 6 or more base pairs); (ii) a plurality of categories of 1 bp T deletions that differ by the homopolymer length in which the deletion occurred (e.g. this may comprise separate categories for homopolymer lengths of 1, 2, 3, 4, 5 and 6 or more base pairs); (iii) a plurality of categories of 1bp C insertions that differ by the homopolymer length in which the deletion occurred (e.g. this may comprise separate categories for homopolymer lengths of 0, 1, 2, 3, 4 and 5 or more base pairs); (iv) a plurality of categories of 1bp T insertions that differ by the homopolymer length in which the deletion occurred (e.g. this may compriseseparate categories for homopolymer lengths of 0, 1, 2, 3, 4 and 5 or more base pairs); (v) a plurality of categories of >1bp deletions at repeat regions that differ from each other by the combination of the deletion length and number of repeat units in the sequence in which the deletion occurred (e.g. this may comprise separate categories for deletion lengths of 2, 3, 4 and 5 or more base pairs, in combination with 1, 3, 3, 4, 5 or 6 or more repeat units); (vi) a plurality of categories of >1bp insertions at repeat regions that differ from each other by the combination of the insertion length and number of repeat units in the sequence in which the insertion occurred (e.g. this may comprise separate categories for insertion lengths of 2, 3, 4 and 5 or more base pairs, in combination with 0, 1, 2, 3, 4, 5 or more repeat units); and (v) a plurality of categories of deletions with microhomology at the deletion boundary (either 5’ or 3’) that differ from each other by the combination of deletion length and microhomology length (e.g. this may comprise separate categories for deletions of 2, 3, 4 or 5 or more base pairs in combination with microhomologies of different lengths compatible with the deletion length, e.g. 1 for deletions of 1 bp, 1 or 2 for deletions of 2bp, 1 , 2 or 3 for deletions of 4 bp, and 1 , 2, 3, 4 or 5 or more for deletions of 5 or more bp). The plurality of categories may be as defined in Table 2. Note that as explained above in the context of single base substitutions, references to a deletion of a base (or bases) should be understood to encompass a reference to deletion of the complementary base or reverse complementary bases. For example, reference to 1bp C deletions I insertions encompass 1bp G deletions insertions (as it is in practice a G:C base pair that is deleted / inserted), and similarly reference to 1bpT deletions encompass 1 bp A deletions I insertions (as it is in practice a T:A base pair that is deleted I inserted). By convention in the art the pyrimidine based is used as a shorthand notation to refer to the base pair that is deleted I inserted.
[0141] An indel signature associated with exposure to a microbial genotoxic compound may be an indel signature comprising proportions for each of a plurality of categories of indels each defined by a unique combination of whether the indel is a deletion or insertion, whether the indel is a 1 bp indel or longer, for 1pb indel whether the indel is a C or T and the length of the mononucleotide repeat tract in which they occur, for longer indels whether the indel is an insertion at a repeat region and the deletion length and number of repeat units, whether the indel is a deletion at a repeat region and the insertion length and number of repeat units, and whether the indel is a deletion with microhomology at the deletion boundary and the deletion length and microhomology length, wherein the indel signature is characterised by higher proportions of 1bp T deletions than any other categories of indels. A mutational signature associated with exposure to a bacterial genotoxic compound may be ID18, ora corresponding signature obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained signatures comprising ID18. ID18 is defined in Table 2. Corresponding signatures may comprise the same signature in any version of the COSMIC database (cancer.sanger.ac.uk / signatures / ; Sondka et al.
[0142] 2024) or the signal database (signal.mutationalsignatures.com / explore / main / cancer / signatures; Degasperi et al. 2020).
[0143]
[0144]
[0145]
[0146]
[0147]
[0148] Table 2. Indel signatures described herein. Definitions from COSMIC v3.4 - October 2023, based on GRCh37. Proportions (also referred to as “weights”) for each of the plurality of categories are provided, together with a definition of the category. The COSMIC classification of indels is described in the first two columns. The number of repeats is the total number of repeats including the repeats in the indel and those immediately 3’ of the indel and consecutive with the repeats of the indel. The length of the microhomology is the length of the portion of sequence of the indel that is repeated immediately 5’ or 3’ to the indel, but where the repeat is not immediately adjacent to the corresponding sequence in the indel.
[0149] A subject as described herein may be a human subject. The subject may be a human subject under the age of 50. The subject may be a human subject under the age of 45. The subject may be an adult human subject. The subject may be a human subject between the ages of 18 and 50. The subject may be a human subject between the ages of 18 and 45. The subject may be a healthy subject. The subject may be a subject who has not been diagnosed with Lynch syndrome. Lynch syndrome is a heritable disorder that increases the risk of developing colorectal cancer due to mutations impairing DNA mismatch repair, such as e.g. mutations in MLH1, MSH2, MSH6, or PMS2. The subject may be a subject who has not been identified as having one or more germline mutations associated with DNA mismatch repair deficiency.
[0150] Embodiments of the present disclosure enable determination of whether a subject is at risk of developing colorectal cancer. This may comprise classification of subjects between classes that have different risks of developing colorectal cancer. Thus, determining whether a subject is at risk of developing colorectal cancer may comprise determining whether the subject has a higher than an expected (e.g. age-adjusted) risk of developing colorectal cancer, and / or classifying the subject between a plurality of classes that have different risks of developing colorectal cancer (e.g. a class that shows evidence of the presence of mutations associated with exposure to a microbial genotoxic compound and that has a higher risk of developing colorectal cancer than a second class that does not show or does not show as much evidence of the presence of mutations associated with exposure to the microbial genotoxic compound). The colorectal cancer may be rectal cancer. The colorectal cancer may be early-onset colorectal cancer. The colorectal cancer may be microsatellite stable colorectal cancer. Microsatellite stable colorectal cancer may also be referred to as DNA mismatch repairproficient cancer. Early-onset colorectal cancer may be defined as colorectal cancer that develops before the age of 50. A subject classified as having a high risk of developing colorectal cancer (e.g. a subject classified in a first class associated with a high risk of developing colorectal cancer) may be recommended or selected for regular invasive colorectal cancer screening testing. For example, such tests may be recommended to be performed every year, every 2 years, every 3 years, every 4 years or every 5 years. An invasive colorectal cancer screening test may be any test that requires inserting an instrument into the body, to draw a sample from a subject, such as e.g. a biopsy or liquid biopsy, and / or to acquire imaging data, such as e.g. a colonoscopy. An invasive colorectal cancer screening test may be or may comprise a colonoscopy. An invasive colorectal cancer screening test may be or may comprise a liquid biopsy, such as e.g. a blood test (e.g. the Shield test from Guardant Health; www.fda.gov / medical-devices / recently-approved-devices / shield-p230009). An invasive colorectal cancer screening test may be or may comprise a tissue biopsy. A subject classified as not having a high risk of developing colorectal cancer (e.g. a subject classified in a second class that has a lower risk of developing colorectal cancer) may be recommended or selected for standard of care colorectal cancer screening associated with the subject’s age. This may not include regular colonoscopy under the age of e.g. 45.
[0151] Classification of a subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer using one or more metrics indicative of exposure to a bacterial genotoxic compound may comprise comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein values of the metrics or score at or above the respective predetermined thresholds is indicative of a higher risk of developing colorectal cancerthan values of the metrics or score below the respective predetermined thresholds. The respective predetermined thresholds may have been identified using a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer. Classifying, using said one or more metrics, a subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer may comprise comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein a subject is: (i) classified in the first class when said score is at or above a predetermined threshold, and classified in the second class otherwise; (ii) classified in the first class when each of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise; or (iii) classified in the first class when a predetermined number or combination of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise.
[0152] Classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer may comprise inputting said one or more metrics into a machine learning classifier trained to classify subjects between a first class has a higher risk of developing colorectal cancer, and a second class thathas a lower risk of developing colorectal cancer using training data comprising the one or more metrics for a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer. The machine learning classifier may also be referred to as “machine learning model” or simply “classifier” or “classification model”. Machine learning classifiers are mathematical models trained (i.e. parameterized) to classify observations between a plurality of classes. The machine learning model may be a binary classifier or a multiclass classifier. A binary classifier may be trained to classify subjects between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer. A multiclass classifier may be trained to classify subjects between a plurality of classes comprising a first class that has a higher risk of developing colorectal cancer, a second class that has a lower risk of developing colorectal cancer, and a third class that has a lower risk of developing colorectal cancer than the first class, but higher than the second class. Such classes may for example be referred to as high risk, low risk and medium risk classes, respectively for the first, second and third classes. A classification model may provide as output a classification label and / or one or more probabilities of an observation belonging to respective one or more classes. A binary classification model may provide as output a single probability or score indicating the probability that the observation belongs to a positive class (e.g. a class with a higher risk of developing colorectal cancer, rather than a class with a lower risk of developing colorectal cancer). A predetermined threshold may be applied to the one or more probabilities or the score, to assign a class label. The threshold may be determined based on desired characteristics of the classification, such as e.g. a desired level of specificity, sensitivity or any accuracy performance that combines aspects of specificity (precision) and sensitivity (recall), such as accuracy and F1 score or balanced versions thereof that take into account the proportions of observations in the training data in each of the classes. Alternatively, an observation may simply be assigned to the class with highest probability. When using a binary classifier this may be equivalent to assigning the class for which the probability is over 50%. Any machine learning classifier known in the art may be used in the context of the present disclosure.
[0153] The term “machine learning algorithm” or “machine learning method” refers to an algorithm or method that trains and / or deploys a machine learning model. The machine learning models of the present disclosure are trained by supervised learning, i.e. using training data comprising observations for each of set of predictive variables for each of a plurality of training subjects, and a class (also referred to as subgroup or prognosis group) label (also referred to as ground truth label). The ground truth label indicates the known prognosis group that a training subject belongs to, based on data indicative of whether and / or within what period of time a subject developed colon cancer. A machine classifier may be trained using training data comprising, for a plurality of training subjects, the values of said one or more metrics and a label indicating whether each subject developed colorectal cancer within a predetermined period of time (e.g. by a predetermined age). A binary classifier may be trained using such data for a plurality of subjects comprising a first plurality of subjects that did develop colorectal cancer within the predetermined period of time, and a second plurality of subjects that did not develop colorectal cancer within the predetermined period of time. A multiclass classifier may be trained usingsuch data for a plurality of subjects comprising a first plurality of subjects that did develop colorectal cancer within a first predetermined period of time, a second plurality of subjects that did not develop colorectal cancer within a second, longer predetermined period of time, and a third plurality of subjects that developed colorectal cancer within the second period of time but not within the first predetermined period of time. Supervised learning refers to the parameterisation of a model so as to optimise an optimality criterion based on the comparison of predicted and ground truth labels for observations in the training data. An optimality criterion may be the minimisation of a loss function that quantifies the model prediction error based on the observed (ground truth) and predicted values of the predicted variables. Suitable loss functions for use in training machine learning models are known in the art and include the d error, the mean absolute error, and regularised versions thereof. Any of these can be used according to the present disclosure. Regularised loss functions are functions that include a loss function as described above, and one or more terms penalizing model complexity in order to reduce the risk of overfitting. Overfitting is a phenomenon that occurs where a machine learning model is trained to very closely reproduce the features of a training data set, resulting in poorer performance on other datasets that do not have the same characteristics (i.e. poor generalizability). L1 regularisation (also known as “Lasso” in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of absolute value of the coefficients of the model. L2 regularisation (also known as “Ridge” in the context of regression) add a regularization term to the loss function that penalizes models based on the sum of squared value of the coefficients of the model. L1 regularisation can be used as a feature selection method as it minimizes the coefficients associated with less informative predictive features. In embodiments, the machine learning model is a regularized model. A machine learning model as described herein may be selected from: decision trees and variants thereof including regularised and / or gradient boosted decision trees (e.g. regularized gradient boosted decision tree model, examples of such models are available in the XGBoost software library (xgboost.ai / )) and random forest models, regularised discriminant analysis, logistic regression models, generalised linear models, artificial neural networks (ANNs) including multilayer perceptrons (with linear or non-linear activation functions) and deep learning models (e.g. long short-term memory networks (LSTMs), Recurrent neural networks (RNNs)), naive Bayes classifiers, and support vector machines (SVM, using linear or non-linear kernels such as radial basis function). In embodiments, a machine learning model comprises an ensemble of models whose predictions are combined. Alternatively, a machine learning model may comprise a single model. Random forest models and gradient boosted tree models (such as XGBoost) are ensemble models. Ensemble versions of any models can be constructed. Very simple models such as linear regression models (also referred to herein as “generalised linear models”) can be useful to enable implementation by clinicians. For example, a simple linear model in which values for a plurality of metrics are added up, optionally with a trained coefficient (also referred to as weight) or simply with a coefficient equal to +1 or -1. The coefficient value is chosen as positive or negative depending on whether the value is expected, based on training data, to increase in subjects in a poorer prognosis group compared to subjects in a better prognosis group (positive coefficient) or decrease in subjects in a poorer prognosis group compared to subjects in a better prognosis group (negative coefficient). This results in a simple score that increases with likelihood of poorer prognosis.Figure 1 illustrates schematically embodiments of methods of the disclosure, such as e.g. methods of determining whether a subject is at risk of developing colorectal cancer, methods of analysing a stool sample (or an anal swab sample), methods of selecting a subject for invasive screening for early onset colorectal cancer, and methods of treating a subject for early onset colorectal cancer. At optional step 10, a stool sample and / or an anal swab sample is obtained from a subject. Step 10 may comprise receiving a stool sample previously from the subject. Step 10 may comprise receiving an anal swab sample previously from the subject. The sample may be stored in a sterile container. The sample may be stored in a cytopreservation solution. At optional step 12, the sample is processed to generate nucleic acid sequence data. The sequence data can comprise DNA sequencing reads obtained from bulk sequencing or error-corrected duplex sequencing (e.g. single molecule duplex sequencing, advantageously with ultra low error rates as described herein, e.g. a sequencing error rate below one error per 10 million sequenced base pairs, such as e.g. UDseq or Nanoseqv2), or information derived therefrom. The sequence data can comprise DNA sequencing reads from whole genome sequencing, whole exome sequencing or targeted panel sequencing, or information derived therefrom. In embodiments, the sequence data comprises sequence data identifying the presence or absence of one or more of: a splicing variant c.835-8A>G in the APC gene, a protein-truncating mutation in the APC gene, and indels in the APC gene associated with 1bp T or A deletions. Such sequence data may be obtained from e.g. targeted sequencing or any other targeted or untargeted sequence determination technology, such as molecular inversion probes, digital PCR, etc. Step 12 may comprise optional step 120 of storing a stool sample at temperatures below 0°C, such as e.g. -17°C, -20°C or -80°C, for example at -80C, or above 0°C, such as e.g. between 0°C and 5°C, such as e.g. 4°C. Instead or in addition to storing at temperatures below 0C, storing may be in a container comprising a DNA stabilisation buffer, or a cytopreservation solution. The container may be a sterile container. Step 12 may comprise optional step 122 of extracting DNA from a stool sample from the subject, optionally using a stool-specific protocol. The step of extracting DNA (step 122) may comprise one or more of: mechanical disruption (e.g. sonication), chemical lysis, purification of nucleic acids, and selective lysis of the subject’s cells over microbial cells followed by DNA purification. Step 12 may comprise optional step 124 of obtaining a DNA sequencing library from DNA extracted from a stool sample from the subject. Step 124 may comprise fragmenting the extracted DNA, adding adapters to the fragments, and amplifying the fragments using said adapters. Step 124 may comprise testing the extracted DNA for DNA integrity and optionally performing an end-repair step when a DNA integrity number for the extracted DNA is above a predetermined threshold. Step 124 may optionally include adding unique molecular identifiers to the DNA fragments in the sample. Step 124 may comprise performing enrichment of target sequences. This may be particularly useful in the context of stool samples, where relatively high microbial DNA contamination may subsist. Step 12 may comprise step 126 of sequencing a DNA library obtained from a stool sample from the subject, optionally using standard bulk sequencing or error-corrected duplex sequencing, such as e.g. ultra-low error single molecule duplex sequencing. At step 14, the nucleic acid sequence data is obtained e.g. by a processor. The nucleic acid sequence data can comprise DNA sequencing reads, or a list of mutations or a mutational profile derived therefrom.At step 16, the processor uses the sequence data to quantify one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound. The microbial genotoxic compound may be colibactin. Thus, the one or more metrics may be indicative of the presence of mutations associated with exposure to colibactin. For example, mutational signatures SBS88 and ID18, as well as specific motifs W[T>N]W mutations that are within a WAWW[T>N]W motif, where W is A or T, and N is A, C, G or T; and specific mutations in the APC gene (e.g. a splicing variant c.835-8A>G in the APC gene, multiple protein-truncating mutations in the APC gene, and indels in the APC gene associated with 1 bp T or A deletions) are believed to be frequently associated with exposure to colibactin. The one or more metrics may be indicative of the presence of mutations associated with exposure to unknown enteric bacteria genotoxins. For example, mutational signatures SBS89 and SBS_M share several epidemiological features with colibactin-induced signatures SBS88 and ID18, and are believed to also be likely to be caused by a mutagen originating from the colorectal microbiome. SBS89 mutations show transcriptional strand bias, a common trait of mutations caused by exogenous mutagenic exposures that form bulky covalent DNA adducts. SBS_M is also enriched in early-onset cases shows similar enrichment in distal colon and, especially, rectal cases, as the original colibactin-associated SBS88 and ID18. The one or more metrics may be selected from: metrics that quantify the presence or activity of a mutational signature associated with exposure to a microbial genotoxic compound, optionally colibactin, and metrics that quantify the presence, number or proportion of mutations associated with exposure to a microbial genotoxic compound, optionally colibactin. For example, quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound may comprise determining the presence of a mutational signature, the number of mutations associated with a mutational signature, the proportion of mutations associated with a mutational signature, and / or the activity of a mutational signature. A mutational signature comprises a plurality of proportions for respective categories of mutations, specifically either single base substitutions or indels. The mutational signatures may be selected from: (i) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS88) is characterised by: (a) a higher proportion of T>C mutations in ATA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in ATA context, T>C mutations in ATT context, T>C mutations in TTT context, and T>G mutations in TTT context than any other categories of single base substitutions; (ii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS89) is characterised by a higher proportion of T>G mutations in GTG context and C>T mutations in ACA context than any other categories of single base substitutions; (iii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS_M) is characterised by: (a) a higher proportion of T>C mutations in CTA context than any other categories of single base substitutions, or(b) a higher proportion of T>C mutations in CTA context, T>C mutations in TTA context, and T>A mutations in CTA context than any other categories of single base substitutions; (iv) an indel signature comprising proportions for each of a plurality of categories of indels each defined by a unique combination of whether the indel is a deletion or insertion, whether the indel is a 1 bp indel or longer, for 1pb indel whether the indel is a C or T and the length of the mononucleotide repeat tract in which they occur, for longer indels whether the indel is an insertion at a repeat region and the deletion length and number of repeat units, whether the indel is a deletion at a repeat region and the insertion length and number of repeat units, and whether the indel is a deletion with microhomology at the deletion boundary and the deletion length and microhomology length, wherein the indel signature (ID18) is characterised by higher proportions of 1bp T deletions than any other categories of indels. The mutational signatures may be selected from signatures (i), (ii) and (iv) above. For example, one or more or all of SBS88, SBS89, SBS_M and ID18 (or one or more or all of SBS88, SBS89 and ID18) may be detected in the sample, or corresponding signatures obtained by extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained SBS signatures comprising SBS88, SBS_M and / or SBS89, or extracting one or more signatures from a cohort of samples using the same plurality of categories and mapping the extracted signatures to a catalogue of previously obtained indel signatures comprising ID18. SBS88, SBS89, SBS_M and ID18 are defined in Table 1 (SBS88, SBS89, SBS_M) and Table 2 (ID18). Graphical illustrations of SBS88, SBS89, SBS_M and ID18 are provided in Fig. 8. Corresponding signatures may comprise the same signatures in any version of the COSMIC database (cancer.sanger.ac.uk / signatures / ; Sondka et al. 2024) or the Signal database (signal.mutationalsignatures.com / explore / main / cancer / signatures; Degasperi et al. 2020). Instead or in addition to detecting mutational signatures, a metric may be quantified as the percentage or proportion of total W[T>N]W mutations that are within a WAWW[T>N]W motif, where W is A or T, and N is A, C, G or T. Instead or in addition to these metrics, the presence of specific mutations in the APC gene may be used as metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound. The specific mutations in the APC gene may be: a splicing variant c.835-8A>G; and / or a protein-truncating mutation; and / or indels associated with 1bp T or A deletions; and / or indels associated with indel signature ID18.
[0154] Step 16 may comprise step 162 of identifying mutations present in the sequence data, thereby obtaining a list of mutations (e.g. SBS or indels) identified in the subject sample. For example, the sequence data can comprise DNA sequencing reads and step 162 can comprise: processing said DNA sequencing reads to identify single base substitutions and / or indels present in the sequence data. The mutations may be somatic mutations. In other words, step 162 may comprise identifying mutations compared to a reference genome and filtering the mutations identified to remove likely germline mutations. Step 16 may comprise step 164 of obtaining a mutational profile derived from the sequence data. A list of mutations may comprise information that identifies individual mutations. Thus, step 164 can comprise obtaining a mutational profile from identified mutations (e.g. substitutions and / or indels) by classifying identified mutations in a predetermined plurality of categories. A mutational profile may compriseinformation that identifies the number of mutations in each of a plurality of categories. A mutational profile can also be referred to as a mutational catalogue, or an indel profile or SBS profile (depending on whether SBSs or indels are being analysed). These steps are optional because the sequence data received at step 14 may already be in the form of a list of mutations or a mutational profile. In other words, steps 162 and / or 164 may have been performed prior to the data being provided to the processor for the purpose of quantifying the metric(s) indicative of exposure to a genotoxic microbial compound. Step 16 may comprise step 166 of quantifying exposure to one or more mutational signatures using the mutational profile. Step 16 may comprise step 168 of assigning a mutational signature to each of one or more mutations identified at step 162.
[0155] At step 18, the values of the one or more metrics quantified at step 16 are used, e.g. by the processor, to determine whether the subject is likely to develop colorectal cancer (CRC). This may comprise classifying the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer. Step 18 may comprise step 182 of comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein values of the metrics or score at or above the respective predetermined thresholds is indicative of a higher risk of developing colorectal cancer than values of the metrics or score below the respective predetermined thresholds. The predetermined thresholds may have been identified using a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer. Thus, also described herein are methods comprising performing all of steps 10-16 using stool samples or anal swab samples from such a cohort of subjects, and identifying one or more respective predetermined thresholds such that subjects that have developed early onset colorectal cancer are more likely to be classified in the first class than subjects that have not develop early onset colorectal cancer. Step 182 may comprise comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein a subject is: classified in the first class when said score is at or above a predetermined threshold, and classified in the second class otherwise; classified in the first class when each of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise; or classified in the first class when a predetermined number or combination of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise. Instead or in addition to step 182, step 18 may comprise step 184 of inputting said one or more metrics into a machine learning classifier trained to classify subjects between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer using training data comprising the one or more metrics for a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer. Thus, also described herein are methods comprising performing all of steps 10-16 using stool samples from a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer, and training a machine learning modelto classify subjects between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer.
[0156] At optional step 20, a subject classified in the first class is selected for an invasive colorectal cancer screening test, optionally a colonoscopy, or any other early onset colorectal cancer screening test (e.g. liquid biopsy, tissue biopsy, etc.). At optional step 22, the colonoscopy or other early onset colorectal cancer screening test is performed on the subject selected at step 20. At optional step 24, the results of one or more of steps 14-18 may be provided to a user, for example through a user interface. Step 24 may comprise providing a report comprising an output of the step of determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer, or information derived therefrom selected from: an indication of whether the subject was determined to be at risk of developing colorectal cancer, a classification between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer, a recommendation fora patient monitoring regimen comprising one or more further colorectal cancer screening tests. The report may be provided through a user interface, or may be output to another computing device or memory. Figure 2 shows an embodiment of a system for implementing methods of the disclosure. The system comprises a computing device 1, which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network, to sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 storing sequence data. The one or more databases 2 may further store one or more of: mutational signatures information, training data, parameters (such as e.g. thresholds or parameters of a machine learning model used to classify subjects, etc.), clinical and / or sample related information, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network 6 such as e.g. over the public internet. The sequence data acquisition means may be in wired connection with the computing device 1, or may be able to communicate through a wireless connection, such as e.g. through WiFi and / or over the public internet, as illustrated. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 3 are configured to acquire sequence data from nucleic acid samples, for example genomic DNA samples extracted from stool samples. In some embodiments, the sample may have been subject to one or more preprocessing steps such as DNA extraction / purification, library preparation, target sequence capture (such as e.g. exon capture and / orpanel sequence capture). Any sample preparation process that is suitable for use in the determination of a genomic alteration profile (whether whole genome or sequence specific) may be used within the context of the present invention. The sequence data acquisition means is preferably a next generation sequencer.
[0157] The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.
[0158] EXAMPLES
[0159] Example 1 - Geographic and age-related variations in mutational processes in colorectal cancer
[0160] Introduction
[0161] The age-standardized incidence rates (ASR) for most adult cancers vary across different geographic locations and can change over time (Bray et al., 2024). Despite extensive epidemiological research, the underlying causes for many of these variations remain unclear. However, they are suspected to be due to exogenous environmental or lifestyle carcinogenic exposures, which are, in principle, preventable (Brennan and Davey-Smith, 2022). Many well-known exogenous carcinogens are also mutagens (Ames et al., 1973; Kucab et al., 2019), which can imprint characteristic patterns of somatic mutations in the genome, known as mutational signatures. Therefore, a complementary approach to conventional epidemiology for investigating unknown causes of cancer is the characterization of mutational signatures in the genomes of cancer and normal cells (Moody et al., 2021; Senkin et al., 2024; Zhang et al., 2021).
[0162] In the past 20 years there has been a notable global increase in the incidence of early-onset colorectal cancer (Patel et al., 2022; Torrens et al., 2024), typically defined as colorectal cancer in adults under 50 years of age. This was first reported in the United States (Siegel et al. 2009) and subsequently observed in Australia, Canada, Japan (Sinicrope 2022) and multiple European countries (Vuik et al., 2019). Although epidemiological studies have identified multiple risk factors for colorectal cancer, specific risk factors for early-onset colorectal cancer remain largely unidentified, with the exception of family history and hereditary predisposition. The latter is predominantly attributable to Lynch syndrome, which is characterized by DNA mismatch repair deficient cancers of the proximal colon (Spaander et al., 2023; Stigliano et al., 2014) and, therefore, is unlikely to be implicated in the recent increase in early-onset colorectal cancer, which is mainly enriched in sporadic, DNA mismatch repair proficient cancers affecting the distal colon and rectum (Venugopal and Carethers, 2022; You et al., 2012).
[0163] Previous colorectal cancer whole-genome sequencing studies have incorporated limited numbers of early-onset cases (Alexandrov et al., 2020; Cornish et al., 2024; Degasperi et al., 2022; Nunes et al., 2024; Rosendahl Huber et al., 2024). In the work described in the present example, the inventors examined 981 colorectal cancer genomes (132 early-onset) from 11 countries on four continents toinvestigate whether variation in mutational processes contributes to geographic and age-related differences in incidence rates.
[0164] Results
[0165] Study design
[0166] 981 colorectal cancers (45.7% female) were collected from intermediate-incidence countries with ASRs of 13-20 / 100,000 people (Iran, Thailand, Colombia, Brazil) and high-incidence countries with ASRs >24 (Argentina, Canada, Russia, Serbia, Czech Republic, Poland, Japan), including the highest ASR of 37 in Japan (Fig. 3a Table 3). Of the 981 cases, 320 were from the proximal colon, 333 from the distal colon, 326 from the rectum, and 2 from unspecified subsites (Fig. 3b). There were 132 early-onset cases, which were 1.88-fold enriched in the distal colon and rectum compared to the proximal colon (p=0.006). All cancers and their matched normal samples underwent whole genome sequencing, achieving a median coverage of 53-fold and 27-fold, respectively.
[0167]
[0168] Table 3. Summary of incidence rates, sex, age, tumor subsite, and molecular subgroups across countries included in the Mutographs colorectal cancer cohort. ASR=Age-standardized incidence rate per 100,000 individuals. Chemo=chemotherapy. P. colon=proximal colon.
[0169] Mutation burden and molecular classification
[0170] The 981 colorectal cancers were divided into known molecular subtypes based on their somatic mutation burdens and profiles. Two main subtypes were identified: DNA mismatch repair proficient cancers, also known as microsatellite stable (MSS), and DNA mismatch repair deficient cancers, often referred to as tumors showing microsatellite instability (MSI). MSS samples (n=802, 81.8%; Fig. 3c) were characterized by a lower burden of single base substitutions (SBS; median: 12,054) and small insertions and deletions (ID; median: 1 ,451), and a higher burden of large-scale genomic aberrations (median: 53.5% of genome altered). In contrast, MSI samples (n=153, 15.6%) exhibited higher SBS and ID burdens (median: 95,426 and 125,100, respectively) with limited genomic aberrations (median: 7.0%). As expected, the average mutational profiles of MSS and MSI colorectal tumors were different (Extended Data Fig. 1a-b).
[0171] MSI samples were predominantly found in the proximal colon (OR=12.2, p=3.8x10-27) and were more common in early-onset cases (OR=2.6, p=0.001). Notably, 28 / 153 MSI cases (18.3%), including 13 / 28 MSI early-onset cases (46.4%), carried germline pathogenic variants in DNA mismatch repair genes consistent with Lynch syndrome. After excluding all cases attributed to Lynch syndrome, there was no enrichment of MSI cancers in early-onset cases (p>0.05). Deficiencies of other DNA repair mechanisms were observed in 24 / 981 cancers (2.4%), including ultra-hypermutated cases with mutations in POLE (n=10, 1.0%) and POLD1 polymerases (n=3, 0.3%), homologous recombination deficient (HRD) cases (n=7, 0.7%), and cases with mutations in the base excision repair genes MUTYH (n=1, 0.1%), NTHL1 (n=2, 0.2%), and OGG1 (n=1, 0.1%) (Methods).
[0172] The mutational catalogues of DNA repair deficient cancers are dominated by somatic mutations resulting from the failed repair process, rendering it difficult to characterize mutational processes unrelated to this failure. To enable investigation of mutational processes unrelated to failed DNA repair processes, analyses were focused on DNA repair proficient colorectal cancers. Two cases treated with chemotherapy for prior cancers were also excluded as their mutation profiles were dominated by the mutational signatures of chemotherapy agents. The remaining cohort consisted of 802 treatment-naive DNA repair proficient colorectal cancers, including 97 early-onset cases.
[0173] After adjustment for sex, country, tumor subsite and tumor purity (Methods), early-onset cancers showed reduced burdens of SBS (fold-change [FC]=0.92, p=0.045)and ID (FC=0.90, p=0.018; Fig. 3d) but not of doublet base substitutions (DBS), copy number alterations (CN), or structural variants (SV) when compared to late-onset cases (p>0.05). Nevertheless, the average mutation spectra of early-onset and late-onset cancers were remarkably similar for all types of somatic mutations (cosine similarity>0.97; Fig. 3e-i). Mutation burden also varied substantially for specific countries whencompared to all others, including Canada (lower SBS and ID burdens), Poland (higher SBS and DBS), Japan (lower SBS, ID, DBS), Iran (lower ID), and Brazil (higher ID and CN). However, mutation profiles were generally consistent across all countries.
[0174] Repertoire of mutational signatures
[0175] A total of 16 SBS, 10 ID, 4 DBS, 6 CN, and 6 SV de novo mutational signatures were extracted from the 802 MSS colorectal cancers and subsequently decomposed into a combination of previously reported reference signatures and potential novel signatures (Tables 4-6). The 16 de novo SBS signatures encompassed 15 COSMICv3.4 signatures (Table 6) including those previously associated with clock-like mutational processes (SBS1, SBS5), APOBEC deamination (SBS2, SBS13) and deficient homologous recombination (SBS3) (Nik-Zainal et al., 2012a), in addition to those previously associated with reactive oxygen species (SBS18) (Alexandrov et al., 2013a), exposure to the mutagenic agent colibactin synthesized by Escherichia coli and other microbes carrying a ~40kb polyketide synthase (pks) pathogenicity island (SBS88) (Lee-Six et al., 2019; Pleguezuelos-Manzano et al., 2020), and mutational processes of unknown causes (SBS8, SBS17a / b, SBS34, SBS40a, SBS89, SBS93, SBS94) (Alexandrov et al., 2020; Islam et al., 2022; Lee-Six et al., 2019; Nik-Zainal etal., 2012a; Senkin et al., 2024). Three previously described signatures of unknown origin (SBS_F, SBS_H, SBS_M; Fig.
[0176] 8b; Degasperi et al., 2022) and a novel signature (SBS_O) were also detected. SBS_O corresponds to a refined version of a previously reported signature of unknown etiology (SBS41 ; Methods) (Alexandrov et al., 2020). With respect to ID, DBS, CN, and SV, most de novo extracted mutational signatures were highly similar to, or directly reconstructed by, COSMICv3.4 reference signatures (Table 6) with the exception of a novel ID signature (ID_J), characterized by deletions of isolated Ts and insertions of Ts in long repetitive regions, and three novel signatures from large mutational events (CN_F, SV_B, SV_D), which were extracted due to the extended contexts used in this signature analysis.
[0177] > > > > > > > > > > > > > > >
[0178]
[0179] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0180]
[0181] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0182]
[0183] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0184]
[0185] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0186]
[0187] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0188]
[0189] > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > > >
[0190]
[0191] > > > > > > > > > > > > > > > > > > > > >
[0192]
[0193] Table 4. Mutational profiles of de novo SBS signatures extracted in MSS colorectal cancer cases.
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202] Table 4 (continued). Mutational profiles of de novo SBS signatures extracted in MSS colorectal cancer cases. Rows are the same as above.
[0203]
[0204]
[0205]
[0206] Table 5. Mutational profiles of de novo ID signatures extracted in MSS colorectal cancer cases.
[0207]
[0208]
[0209]
[0210] Table 5 (continued). Mutational profiles of de novo ID signatures extracted in MSS colorectal cancer cases.
[0211] &
[0212] &
[0213] &
[0214] &
[0215] &
[0216] &
[0217] &
[0218] &
[0219] &
[0220] &
[0221] &
[0222]
[0223] & &
[0224] &
[0225] &
[0226] &
[0227] &
[0228] &
[0229] &
[0230] &
[0231] &
[0232] &
[0233] &
[0234] &
[0235] &
[0236] &
[0237] &
[0238] &
[0239] &
[0240] &
[0241] &
[0242] &
[0243] &
[0244] & &
[0245] &
[0246]
[0247] Table 6. Decomposition of de novo MSS colorectal cancer signatures into previously reported signatures.
[0248] Geographic variation in mutational signatures
[0249] Despite the similar mutation profiles across countries, several signatures exhibited varying prevalence when comparing one country to all others (Fig. 4a). Notably, SBS89 (OR=28.0, q=0.001), DBS8 (OR=8.9, q=3.2x10-4), and the novel ID_J (OR=9.6, q=6.2xl0-5) were at elevated frequencies in Argentina when compared to all other countries (Fig.4b). Signatures SBS89, DBS8, and ID_J showed a strong tendency to co-occur (p<1.7x10-11) suggesting they may arise from the same underlying mutational process. In Colombia (Fig. 2c), higher frequencies were observed for SBS94 (OR=19.7, q=3.2x10-5), the novel SBS_F (OR=10.7, q=2.0x10-4), and DBS6 (OR=12.5, q=0.028) when compared to all other countries, with evidence of co-occurrence of SBS94 with SBS_F (p=0.017) and DBS6 (p=1.9x10-4). Enrichments were also found for SBS2 (OR=2.0, q=0.041) and SBS_H (OR=2.3, q=0.001) in Russia and CN_F (OR=3.5, q=3.9x10-4) in Brazil, whereas depletions were identified for DBS2 in Thailand (OR=0.38, q=0.008) and for DBS4 in Colombia (OR=0.06, q=0.034; Fig.4a). Overall, the results indicate international differences in the prevalence of certain mutational processes involved in colorectal cancer development.
[0250] To explore the broader epidemiological implications of international variation in mutational processes, the inventors evaluated the relationships between ASR and mutational signatures (Fig. 4d; Table 7).
[0251] Independent of covariates, colibactin-induced mutational signatures, SBS88 and ID18, as well as clocklike signature SBS1 and novel signature SBS_H, associated with an increasing rate of ASR for colorectal cancer, whereas novel signature CN_F associated with a reduced ASR rate (q<0.05; Fig.
[0252] 4d-e). For SBS88 and ID18, the association was linked with the ASR for rectal cancer (q=0.088 and q=0.008; Fig. 4f; Table 8). In contrast, for SBS1, SBS_H, and CN_F the association was particularly strong for the ASR of colon cancer (q=0.009, q=0.015, and q=0.057).
[0253]
[0254]
[0255] Table 7. Enrichment of mutational signature activities with colorectal cancer incidence in MSS colorectal cancers.
[0256]
[0257]
[0258]
[0259] Table 8. Enrichment of mutational signature activities with colon and rectal cancer incidence in MSS colon and rectal cancers.
[0260] Colibactin induced mutational signatures are enriched in early-onset colorectal cancer
[0261] The substantial number of early-onset colorectal cancer cases enabled evaluation of the association between mutational signatures and age at diagnosis. Although the average mutation profiles of early-onset and late-onset colorectal cancer cases were similar (Fig.3h-l), the prevalence of some mutational signatures was associated with age of diagnosis (Fig. 5a; Table 9). Late-onset cases showed enrichment in signatures known to accumulate in a linear fashion with age in normal colorectal crypts (Alexandrov et al., 2015), including SBS1, SBS5, ID1, and ID2 (Fig. 5a-b). Unknown etiology indel signatures ID4, ID9, and ID10 also showed associations with late-onset cases (Fig. 5a-b).
[0262]
[0263]
[0264] Table 9. Enrichment of mutational signature prevalence in late-onset compared to early-onset MSS colorectal cancers.
[0265] By contrast, enrichment in early-onset cancers was observed for colibactin-induced signatures, SBS88 (OR=2.5, q=0.006) and ID18 (OR=4.0, q=3.7x10-7; Fig. 5a-b). The primary associations of early-onset cases with SBS88 and ID18 were further supported by the successive decline in the prevalence of these signatures with increasing age of diagnosis (p-trend=1.3x10-4 and p-trend=2.0x10-7, respectively; Fig. 5c; Table 10). Signatures SBS88 and ID18 were 3.3 and 3.6 times more common, respectively, in colorectal cancer diagnosed under the age of 40 than those over the age of 70 (Fig.
[0266] 5c). A similar effect was observed using a complementary motif enrichment analysis for detecting SBS88 (p-trend=1.0x10-7; Fig. 9a-b). Based on the strong co-occurrence of SBS88 and ID18 (p=7.4x10-63), as well as previous functional (Pleguezuelos-Manzano et al., 2020) and population studies (Cornish et al., 2024; Lee-Six et al., 2019; Rosendahl Huber et al., 2024), exposure to colibactin was defined by the presence of either SBS88 or ID18. Colibactin exposure was found in 21.1% of all colorectal cancers (169 / 802) and was associated with earlier age of onset (median age: 62 vs. 67, p=1.6x10-8; Fig. 5d), an effect more evident in the distal colon (median age: 57 vs. 66, q=5.2x10-7) and rectum (median age: 63 vs. 66, q=0.025; Fig. 5e). Overall, colibactin exposure had a strong inverse correlation with age, being 3.3 times more common in colorectal cancers diagnosed in individuals younger than 40 compared to those over 70 (p-trend=2.7x10-7 Fig. 9c).
[0267] Signatures of unknown etiology SBS_M and ID14 (Fig. 5a-c) were also enriched in early-onset cases, and SBS89 similarly exhibited a higher prevalence in younger individuals (5.4 times more patients younger than 40 compared to older than 70, p-trend=0.047), albeit based on a very small number of cancers harboring the signature (9 / 802, 1.1%; Fig. 5c).
[0268]
[0269]
[0270] Table 10. Trend enrichment of mutational signature prevalence with age of diagnosis in MSS colorectal cancers.
[0271] Colibactin mutagenesis is an early event in colorectal carcinogenesis
[0272] To time the imprinting of SBS88 and ID18, mutations were categorized as early clonal, late clonal, or subclonal during the development of each cancer and the contribution of each mutational signature to each category was determined. SBS88 and ID18 were both enriched in early clonal compared to late clonal mutations (q=4.2x10-4 and q=6.1x10-5; Fig. 6a), as well as a similar trend in clonal comparedto subclonal mutations (q=0.138 and q=0.058; Fig. 10a), consistent with the presence of these mutational signatures in normal colorectal epithelium (Lee-Six et al., 2019). This enrichment in earlier evolutionary stages was similar to the one observed for other well-known clock-like signatures like SBS1, SBS5, or ID1 (Fig. 6a-b), as previously shown in tumors (Dentro et al., 2021; Gerstung et al., 2020) and normal tissues (Lee-Six et al., 2019), and in contrast to signatures known to preferentially generate late clonal and subclonal mutations, such as SBS17a / b38. The enrichment of colibactin signatures in early clonal mutations was observed for both early-onset (q=0.004 for SBS88 and q=2.0*10-4 for ID18) and late-onset colorectal cancer cases (q=0.020 and q=0.024; Fig. 10b).
[0273] Since colibactin is produced by bacteria carrying the pks pathogenicity island, the inventors investigated whether colorectal cancer cases with SBS88 or ID18 harbored pks+ bacteria based on sequencing reads from the cancer sample that did not map to the human genome but mapped to the pks locus. There was no association between the presence of SBS88 or ID18 and that of pks+ bacteria (Fig. 6b; Fig. 11). Moreover, the inventors observed a younger age of diagnosis for cases with SBS88 or ID18 but without an identified pks+ bacteria (p=1.3x10-7; Fig. 6c-d). These results are consistent with the presence of pks+ bacteria and consequent imprinting of SBS88 and ID18 on the colorectal epithelium restricted to an early period of life, with plasticity of the microbiome over multiple decades resulting in loss and gain of pks+ bacteria.
[0274] Colibactin exposure and driver mutations
[0275] Using the IntOGen framework (Martinez-Jimenez et al., 2020), 46 genes under positive selection were identified, with eight mutated in more than 10% of cancers: APC, TP53, KRAS, FBXW7, SMAD4, PIK3CA, TCFL2, and SOX9 (Fig. 7a; Table 11). Forty-three of the 46 genes have been previously reported as colorectal cancer driver genes (Cornish et al., 2024; Martinez-Jimenez et al., 2020), two in other cancer types (MED12, NCOR1) (Martinez-Jimenez et al., 2020), and a putative novel colorectal cancer driver gene, CCR4, was identified with mutations indicating inactivation of the encoded protein. Mutations affecting these 46 cancer driver genes were annotated as driver mutations using a multi-step process based on the mutation type and the mode of action of the gene. An elevation in the total number of driver mutations was observed in late-onset compared to early-onset cases (FC=1.21, p=5.4x10-5;
[0276] Fig. 7b). In addition, an enrichment in APC driver mutation carriers was also found for late-onset cases (OR=2.7, q=0.027; Fig. 7c-d; Table 12), as previously reported (Kim et al., 2021), whereas no hotspot driver mutation (defined as those affecting the same genomic position in at least 10 cases) was associated with age of onset (q>0.05; Table 13). No statistically significant differences across countries were found for driver mutations within cancer driver genes or for hotspot driver mutations (q>0.05).
[0277] The contributions of SBS88 and ID18 to driver mutations were assessed using probabilistic assignment of signatures to individual mutations (Diaz-Gay et al., 2023). SBS88 accounted for 64.3% of the colibactin-induced (Terlouw et al., 2020) APC splicing variant c.835-8A>G in colibactin-exposed samples, compared to only 3.9% and 3.8% of driver substitutions in APC or other cancer genes (Fig.
[0278] 7e). Similarly, ID18 accounted for 25.3% of APC driver indels and 16.9% of other driver indels incolibactin-exposed cases (Fig. 7f). Overall, SBS88 and ID18 accounted for 8.3% of all SBS and ID driver mutations, and 15.5% of all APC driver mutations in colibactin positive cancers. Nevertheless, no differences were observed between early-onset and late-onset colibactin positive colorectal cancer in the proportion of driver mutations assigned to specific mutational signatures (Fig. 12).
[0279]
[0280]
[0281] Table 11. Driver genes detected in MSS colorectal cancers. LoF=loss of function. Act=activation.
[0282]
[0283]
[0284] Table 12. Enrichment of driver mutations in cancer driver genes in late-onset compared to early-onset MSS colorectal cancers. FBLRL= Firth's Bias-Reduced Logistic Regression. RLR= Regular Logistic Regression
[0285] > > > > > > > > > > > > > > > > >
[0286]
[0287] > > > > > > > > > > > > > > > > >
[0288]
[0289] Table 13. Enrichment of hotspot driver mutations in late-onset compared to early-onset MSS colorectal cancers. MS=mutated samples. OR=odds ratio. FBLRL= Firth's Bias-Reduced Logistic Regression. RLR= Regular Logistic Regression
[0290] Methods
[0291] Recruitment of patients. The inclusion criteria for patients were >18 years of age (ranging from 18 to 95, with a mean of 64 and a standard deviation of 12), confirmed diagnosis of primary colorectal cancer, and no prior treatment for colorectal cancer. Patients were excluded if they had any condition that could interfere with their ability to provide informed consent or if there were no means of obtaining adequate tissues or associated data as per the protocol requirements.
[0292] Bio-samples and data collection. Dedicated standard operating procedures, following guidelines from the International Cancer Genome Consortium (ICGC), were designed by IARC / WHO to select appropriate case series with complete biological samples and exposure information as described previously (Senkin et al. 2024; Moody et al., 2021; Perdomo et al. 2024). In brief, for all case series included, anthropometric measures were taken, together with relevant information regarding medical and familial history. All biological samples from retrospective cohorts were collected using rigorous, standardized protocols and fulfilled the required standards of sample collection defined by the IARC / WHO for sequencing and analysis. Potential limitations of using retrospective clinical data collected using different protocols from different populations were addressed by a central data harmonization to ensure a comparable group of exposure variables. All patient-related data were pseudonymized locally through the use of a dedicated alpha-numerical identifier system before being transferred to the IARC / WHO central database.Expert pathology review. Original diagnostic pathology departments provided diagnostic histological details of contributing cases through standard abstract forms, together with a representative hematoxylin-eosin-stained slide of formalin-fixed paraffin-embedded tumor tissues whenever possible. A centralized digital pathology examination of the frozen tumor tissues collected for the study as well as formalin-fixed paraffin-embedded sections when available, was coordinated via a web-based approach and dedicated expert panel, following standardized procedures as described previously (Senkin et al., 2024; Moody et al., 2021). Frozen tumor tissues were first examined to confirm the morphological type and the percentage of viable tumor cells. A random selection of tumor tissues was independently evaluated by a second pathologist. Enrichment of tumor component was performed by dissection of the non-tumoral part, if necessary. A minimum of 50% viable tumor cells were considered eligibile for whole-genome sequencing.
[0293] DNA extraction. Extraction of DNA from fresh frozen tumor and matched blood / normal tissue samples was conducted following a standardized DNA extraction procedure. Germline DNA was extracted from whole blood (n=957) or from adjacent normal tissue (n=24) using previously described protocols and methods (Senkin et al., 2024; Moody et al., 2021).
[0294] Whole-genome sequencing. A total of 1 ,977 colorectal cancer patients were enrolled into the study, 906 (45.8%) were excluded due to insufficient viable tumor cells (pathology level) or inadequate DNA (tumor or germline). DNA from 1,071 cases underwent whole-genome sequencing. Fluidigm SNP genotyping with a custom panel was performed to ensure that each pair of tumor and matched normal samples originated from the same individual. Whole-genome sequencing (150 bp paired-end) was performed on the Illumina NovaSeq 6000 platform with a target coverage of 40x for tumors and 20x for matched normal tissues. All sequencing reads were aligned to the GRCh38 human reference genome using the Burrows-Wheeler Aligner MEM (BWA-MEM; v0.7.16a and vO.7.17) (Li and Durbin, 2009). Postsequencing quality control metrics were applied for total coverage, evenness of coverage, and contamination. Cases were excluded if coverage was below 30x for tumor or 15x for normal tissue. For evenness of coverage, the median over mean coverage (MoM) score was calculated. Tumors with MoM scores outside the range of values determined by Whalley et al. (2020) to be appropriate for wholegenome sequencing (0.92-1.09) were excluded. Conpair (Bergmann et al., 2016) was used to detect contamination and cases were excluded if the result was greater than 3% (Whalley et al., 2020). A total of 981 pairs of colorectal cancer and matched-normal tissue passed all criteria.
[0295] Germline variant calling. Germline single nucleotide variants (SNV) and indels were derived from wholegenome sequencing from the normal paired material for each individual using Strelka2 with appropriate quality-control criteria (Kim et al., 2018). Variant calls were derived into genotypes for each individual and annotated using SnpEff (Cingolani et al., 2012).
[0296] Somatic variant calling. Variant calling was performed using the CASM IT pipeline (github.com / cancerit). Copy number profiles were determined using ASCAT (Van Loo et al., 2010) andBATTENBERG (Nik-Zainal et al., 2012b) when tumor purity allowed. SNVs were called with cgpCaVEMan (Jones et al., 2016), indels were called with cgpPINDEL (Raine et al., 2015), and structural rearrangements were called using BRASS (github.com / cancerit / BRASS). CaVEMan and BRASS were run using the copy number profile and purity values determined from ASCAT when possible (complete pipeline, n=916). When tumor purity was insufficient to determine an accurate copy number profile (partial pipeline, n=31 ) CaVEMan and BRASS were run using copy number defaults and an estimate of purity obtained from ASCAT. Finally, for a subset of cases which had no large copy number alterations (copy number normal pipeline, n=34), CaVEMan and BRASS were run using copy number defaults and an estimate of purity calculated by the median variant allele frequency (VAF) of indels multiplied by two. For SNVs, additional filters (ASRD>140 and CLPM=0) were applied in addition to the standard PASS filter to remove potential false positive calls. To further exclude the possibility of caller-specific artifacts being included in the analysis, a second variant caller, Strelka2 (Kim etal., 2018), was run for SNVs and indels. Only variants called by both variant calling pipelines were included in subsequent analysis.
[0297] Generation of mutational matrices. Mutational matrices for single base substitutions (SBS), indels (ID), doublet base substitutions (DBS), copy number alterations (CN), and structural variants (SV) were generated using SigProfilerMatrixGenerator with default options (v1.2.0) (Khandekar et al. 2023; Bergstrom et al. 2019).
[0298] Microsatellite instability validation. The presence of microsatellite instability (MSI) in colorectal cancers was validated using the QX200 Droplet Digital PCR System (Bio-Rad, Hercules, CA, USA) for the detection of five microsatellite markers (BAT25, BAT26, NR21, NR24, and Mono27) commercially pooled in three primer-probe mix assays, as previously described (Gilson et al. 2020). Samples were tested in duplicate, and each reaction comprised 1* ddPCR Multiplex Supermix for probes (Bio-Rad), 1X primer-probe mix, and 10 ng of extracted tumor DNA, in a total volume of 22 pl. MSI-positive, negative, and no-template (nuclease-free water) controls were included in each experiment. Droplet generation and plate preparation for thermal cycling amplification were performed using the QX200 AutoDG Droplet Digital PCR System (Bio-Rad). The following PCR protocol was applied on a C1000 Touch Thermal Cycler (Bio-Rad): 37 °C for 30 min, 95 °C for 10 min, followed by 40 cycles of denaturation at 94 °C for 30 s, annealing at 55 °C for 1 min, with a final extension at 98 °C for 10 min. Following PCR amplification, fluorescence signals were quantified using the QX200 Droplet Reader (Bio-Rad), and data were analyzed with QuantaSoft Analysis Pro v.1.0.596.0525 (Bio-Rad) software. Positive and negative controls served as guides to call markers and delineate clusters. For each assay, the cluster at the bottom left of the x-y plot was designated as the negative population. Clusters located vertically and horizontally from the negative cluster were identified as the mutant population, while clusters located diagonally from the negative cluster represented the wild-type population. Tumors were characterized for the MSI phenotype by analyzing the results for all five markers using the following criteria: MSI positive if two or more mutant microsatellite markers were observed, and microsatellite stable (MSS; i.e., MSI negative) when none or only one of the microsatellite markers was altered.Extraction and decomposition of mutational signatures
[0299] Mutational signatures were primarily extracted using SigProfilerExtractor (Islam et al., 2022), based on nonnegative matrix factorization, and validated by mSigHdp (Liu et al., 2023), based on hierarchical Dirichlet process mixture models. For SigProfilerExtractor (v1.1.21), de novo mutational signatures were extracted from SBS, DBS, and ID mutational matrices using 500 NMF replicates (nmf_replicates=500), nndsvd_min initialization (nmf_init=”nndvsd_min”), and default parameters. Extractions were performed separately on the subsets of 802 MSS and 153 MSI cases. De novo SBS mutational signatures were extracted for both SBS-288 and SBS-1536 contexts, which, beyond the common SBS-96 trinucleotide context using the mutated base and the 5’ and 3’ adjacent nucleotides (Alexandrov et al., 2013b; Bergstrom et al., 2019), also consider the transcriptional strand bias and the pentanucleotide context (two 5’ and 3’ adjacent nucleotides), respectively (Bergstrom et al., 2019). The results were largely concordant, with the SBS-288 de novo signatures allowing additional separation of mutational processes. Therefore, the SBS-288 de novo signatures were taken forward for further analysis (Table 4). Previously established mutational contexts DBS-78 and ID-83 (Alexandrov et al., 2020; Bergstrom et al., 2019) were used for the extraction of DBS and ID signatures (Table 5).
[0300] Copy number signatures were extracted de novo using SigProfilerExtractor with default parameters and following an updated context definition benefitting from WGS data (CN-68). which allowed to further characterize CN segments below 100kbp in length (in contrast to current COSMICv3.4 reference signatures using the CN-48 context, which were based on SNP6 microarray data and therefore without the resolution to characterize short CN segments) (Steele et al. 2022). SV signatures were extracted using a similarly refined context, with an in-depth characterization of short SV alterations below 1kbp (SV-38 context, in contrast to current COSMICv3.4 signatures based on the SV-32 context (Everall et al. 2023)).
[0301] After de novo extraction was completed, SigProfilerAssignment (Diaz-Gaye et al., 2023) v0.0.29 was used to decompose the de novo extracted SBS, ID, DBS, CN, and SV mutational signatures into COSMICv3.4 reference signatures based on the GRCh38 reference genome (Sondka et al., 2024) (Table 6). When possible, SigProfilerAssignment matched each de novo extracted mutational signature to a set of previously identified COSMICv3.4 signatures. For the SBS-288, CN-68, and SV-38 signatures, this required collapsing the high-definition classifications into the standard SBS-96, CN-48, and SV-32 mutational classifications, respectively. Four of the de novo extracted MSS SBS signatures did not match any previous COSMICv3.4 signatures, with three of them (SBS_F, SBS_H, and SBS_M) showing a strong similarity with previously reported signatures in the UK population (Degasperi et al., 2022) (cosine similarity > 0.93), and one (SBS_O) reflecting a cleaner version of a previously reported COSMICv3.4 signature (SBS41). To validate the latter, the inventors performed a decomposition of the current mutational profile of signature SBS41 using the decomposed signatures from their analysis, obtaining a confirmation that SBS41 can be reconstructed by a linear combination of SBS_O (contributing 19.00% of the mutational profile), SBS93 (62.54%), and SBS34 (12.60%), and SBS5(5.86%) with a cosine similarity of 0.91. For the MSS cohort, one ID (ID_J), one CN (CN_F), and two SV signatures (SV_B and SV_D) were additionally not decomposed into previously known signatures, and therefore considered as novel. The novel SV signature SV_D, identified in the MSS cohort, was also considered for the decomposition of de novo SV signatures extracted in the MSI cohort. In the MSI cohort, four of the de novo extracted SBS signatures (SBS_I_MSI, SBS_M_MSI, M, SBS_N_MSI, and SBS_O_MSI) as well as one de novo DBS signature (DBS_B_MSI) did not match COSMICv3.4 signatures, with SBS_M_MSI showing a strong similarity with a previously reported signature in the UK population (Degasperi et al., 2022) (cosine similarity=0.89), and the other four signatures considered as novel.
[0302] mSigHdp (Liu et al., 2023) extraction of SBS-96 and ID-83 signatures was performed on the 802 MSS subset using the suggested parameters and using the country of origin to construct the hierarchy. SigProfilerAssignment was subsequently used to match mSigHdp de novo signatures to previously identified COSMIC signatures. A comparison of the signatures extracted from mSigHdp and SigProfilerExtractor can be found below.
[0303] Extraction of de novo mutational signatures with SigProfilerExtractor in microsatellite stable tumors. Single base substitutions (SBS) extractions were performed using SigProfilerExtractor21with both SBS-288 and SBS-1536 contexts. SBS-288 extends the SBS-96 contexts by classifying mutations into transcribed, untranscribed, or intergenic non-transcribed regions, whereas SBS-1536 considers the two flanking bases on either side of the mutated base to form a pentanucleotide context (Bergstrom et al.
[0304] 2019). These extractions extracted 16 and 14 signatures, respectively. In order to calculate the cosine similarity, the SBS-288 and SBS-1536 signatures were both collapsed into the SBS-96 mutational context. 14 signatures were extracted in both formats with cosine similarity >0.9 (Table 14). The two signatures extracted only in the SBS-288 format were SBS_E and SBS_K. The former is a flat signature that decomposed to SBS1, SBS3 and SBS5, and the latter represents an incomplete separation of SBS17a and SBS17b. The SBS-288 results were used for the final analysis as this format allowed the extraction of these additional signatures. Notably, all four signatures where decomposition was rejected (SBS_F, SBS_H, SBS_M, and SBS_O) could be reproduced in both SBS-288 and SBS-1536 formats with a cosine similarity >0.95 (Table 14).
[0305] Extraction of de novo mutational signatures with mSigHdp in microsatellite stable tumors
[0306] In order to validate the mutational signatures obtained using SigProfilerExtractor, extractions were also performed with a second algorithm, mSigHdp23. SBS extraction was performed using the SBS-96 mutational context only, as mSigHdp has not been benchmarked for extended contexts. In total, mSigHdp extracted 17 signatures of which 15 were very close matches (cosine similarity >0.90) to the SBS-288 results, whereas signatures hdp.14 and hdp.17 were unique to mSigHdp (Table 14). The first unique signature, hdp.14, was an additional APOBEC containing signature, whereas hdp.17 was a combination of the new signature hdp.10 (SBS_H) and SBS93 (hdp.7 I SBS_I). The latter was confirmed by performing a decomposition using the panel of COSMICv3.4 and the new colorectalcancer signatures as the signature database. Again, all four signatures where decomposition was rejected (SBS_F, SBS_H, SBS_M, and SBS_O) can be reproduced in mSigHdp with a cosine similarity >0.95 (Table 14), indicating that these signatures are highly reproducible using independent methodology.
[0307] mSigHdp extracted 8 indel signatures in comparison to the 10 extracted from SigProfilerExtractor. Of these, 6 mSigHdp signatures could be matched directly to those from SigProfilerExtractor (Table 15).
[0308] Of the remaining 2 mSigHdp signatures, hdp.6 appears to be a combination of the SigProfilerExtractor de novo signatures ID_G and ID_H (but missing the ID1 component of both of those de novo signatures), whereas hdp.8 is a combination of ID2, ID5, and ID9. ID5 is not included in the final panel of signatures decomposed from the SigProfilerExtractor and was not included in the final panel of signatures used for signature assignment. Notably, the novel signature ID_J was reproducible in the mSigHdp but with low cosine similarity (0.64; Table 15). However, the low cosine similarity is explained by the lack of ID1 contamination in the signature extracted using mSigHdp.
[0309]
[0310] Table 14. Comparison of single base substitution signatures extracted by SigProfilerExtractor and mSigHdp in MSS colorectal cancer cases.
[0311]
[0312]
[0313] Table 15. Comparison of indel signatures extracted by SigProfilerExtractor and mSigHdp in MSS colorectal cancer cases.
[0314] Decomposition to reference signatures
[0315] The extracted de novo mutational signatures were decomposed into the COSMICv3.4 reference signatures using SigProfilerAssignment (Diaz-Gay et al. 2023). As the COSMICv3.4 reference signatures are only available for the SBS-96 mutational context, the de novo SBS-288 signatures were collapsed into the SBS-96 context. As a result of the loss of information from extended contexts and also due to the large number of COSMICv3.4 reference mutational signatures, it is likely that the decompositions on default settings will include signatures that are implausible given the cancer type. For the MSS signatures using the default decomposition settings, only SBS_M was not decomposed. In order to optimize the decomposition, the following signature subgroups were excluded (using the exclude_signature_subgroups parameter of SigProfilerAssignment): artifact signatures, ultraviolet signatures, lymphoid signatures, mismatch repair deficiency signatures, polymerase deficiency signatures, base excision repair deficiency signatures, and treatment signatures. In addition, the new_signature_threshold was set to 0.90. Using these settings, SBS_H, SBS_O, and ID_J , in addition to SBS_M were not decomposed. SBS_F was still decomposed into SBS1 , SBS5, SBS19 and SBS40a in the optimized decomposition. However, this was rejected on the basis that SBS_F had been previously extracted in an independent cohort10and the lack of individual spectra that supported the presence of SBS19 (the other reference signatures all appeared in other decompositions).
[0316] For the MSI extractions, using default settings for SigProfilerAssignment, only SBS_M_MSI and DBS_B_MSI were not decomposed. In order to optimize the decomposition, the following signature subgroups were excluded (using the exclude_signature_subgroups parameter): artifact signatures, ultraviolet signatures, lymphoid signatures, polymerase deficiency signatures, base excision repair deficiency signatures, homologous repair deficiency signatures, and treatment signatures. Optimizing did not change the number of signatures that were not decomposed. However, the decompositions for SBS_I_MSI, SBS_N_MSI, and SBS_O_MSI were subsequently rejected on the basis that individual spectra existed that strongly support these signatures being the result of distinct mutational processes.Attribution of mutational signatures to individual samples. As explained above, known COSMIC signatures and de novo signatures that were not decomposed into COSMIC signatures were attributed for each sample using MSA (Senkin, 2021) (v2.0) for SBS, ID, and DBS, whereas SigProfilerAssignment (Diaz-Gay et al., 2023) was used for CN and SV. A conservative approach was used for MSA attributions utilizing the (params. no_CI_for_penalties=False) option for the calculation of optimum penalties. Pruned attributions were used for the final analysis, where confidence intervals were applied to each attributed mutational signature and any signature activity with a lower confidence limit equal to 0 was removed.
[0317] Attribution of mutational signatures to individual somatic mutations. SBS and ID mutational signatures were probabilistically attributed to individual somatic mutations using the MSA activities per sample, based on Bayes’ rule and the specific mutational context for the mutation, as previously described (Diaz-Gay et al. 2023). Briefly, to calculate the probability of a specific mutational signature being responsible for a mutation in a given mutational context and in a particular sample, the general probability of the signature causing mutations in a specific mutational context (obtained from the mutational signature profile) was multiplied by the activity of the signature in the sample (obtained from the signature activities), and then this value was normalized by dividing by the total number of mutations corresponding to the specific mutational context (obtained from the reconstructed mutational profile of the sample). The signature with the maximum likelihood estimation was assigned to each individual somatic mutation.
[0318] Driver gene analysis. Consensus de novo driver gene identification was performed by IntOGen (Martinez-Jimenez et al., 2020). The genes identified as drivers with a combination q-value<0.10 were classified according to their mode of action in tumorigenesis ( / .e., tumor suppressor genes or oncogenes) based on the relationship between the excess of observed nonsynonymous and truncating mutations computed by dNdScv (Martincorena et al., 2017) and their annotations in the Cancer Gene Census (Sondka et al., 2018).
[0319] To identify potential driver mutations, the inventors selected SBS or ID mutations that fulfilled any of the following criteria: mutations classified as “Oncogenic” or “Likely Oncogenic” by OncoKB (Chakravarty et al., 2017); mutations classified as drivers in the TCGA MC3 drivers study (Bailey et al., 2018); truncating mutations in driver genes annotated as tumor suppressors; recurrent missense mutations (seen in at least three cases); mutations classified as “Likely Drivers” by boostDM (score >0.50) (Muinos et al., 2021); or missense mutations classified as “Likely Pathogenic” by AlphaMissense (Cheng et al., 2023) in driver genes annotated as tumor suppressors. Six of the IntOGen-identified driver genes did not carry any potential driver mutations according to these criteria and were therefore excluded from subsequent analysis. 60 driver genes were identified (46 and 31 for MSS and MSI cases, respectively).
[0320] Evolutionary analysis. DPCIust (Nik-Zainal et al., 2012) was run on all complete pipeline MSS samples with Battenberg data (n=774) to identify clonal structure in each sample. The DPCIust output was used in running MutationTimeR (Gerstung et al., 2020) to annotate somatic mutations as early clonal, lateclonal, subclonal, or NA clonal. Samples with at least 256 early clonal and late clonal SBSs or 100 early clonal and late clonal IDs were retained and split into separate VCF files (n=574 for SBS; n=430 for ID). MSA (Senkin, 2021) was run on the resulting VCF files to identify the active mutational signatures in the early clonal and late clonal mutations. Signatures that were found to generate early clonal somatic mutations in fewer than 50 samples and also generate late clonal somatic mutations in fewer than 20 samples were excluded from the analysis. The Wilcoxon signed-rank test was used to assess the differences in the relative activity of each signature between the early clonal and late clonal mutations. P-values were corrected across signatures using the Benjamini-Hochberg method (Benjamini and Hochberg, 1995), and adjusted q-values were reported. This process was repeated with the same thresholds for SBSs and IDs to also assess the difference in the relative activity of each signature between clonal and subclonal mutations (n=133 for SBS; n=64 for ID), but due to the lower numbers, signatures that were found to generate clonal somatic mutations in fewer than 10 samples and also generate clonal somatic mutations in fewer than 10 samples were excluded from the analysis.
[0321] Motif analysis. MutaGene (Goncearenco et al., 2017) was used to find the number of mutations with the WAWW[T>N]W motif, previously associated with colibactin mutagenesis (Pleguezuelos-Manzano et al., 2020), in each sample, regardless of the DNA strand. This value was then divided by the total number of W[T>N]W mutations per sample to identify the percentage of W[T>N]W mutations with the colibactin mutational motif.
[0322] Microbiome analysis. To identify microbial reads that map to the pks island (pks), non-human reads were aligned to the IHE3034 genome (RefSeq assembly: GCF_000025745.1) using Bowtie2 (Langmead and Salzberg, 2012). IHE3034 is a pks E. coli strain that contains the pks island with all 19 clb genes in the clbA-cibS gene cluster. Prior to alignment, poor quality reads were filtered using fastp (Chen et al., 2018), and the remaining human reads were removed by excluding those that mapped to GRCh38, T2T-CHM13v2.0, and the 47 pangenomes (Liao et al., 2023). A sample was considered pks+ if it had at least one read across at least 8 out of the 19 genes in the clbA-cibS gene cluster.
[0323] Regressions. To compare the mutation burden of different variant types, a linear regression of the mutation burden logarithm (base 10) was considered, using age, sex, tumor subsite, country, and tumor purity as independent variables. For mutational signature-based analyses, signature attributions were dichotomized into presence and absence using confidence intervals, with presence defined as both lower and upper limits being positive and absence as the lower limit being zero. If a signature was present in at least 70% of cases (SBS1, SBS5, SBS18, ID1, ID2, ID14 and CN2 for MSS cases; ID1, ID2, DBS_B_MSI, CN1, and SV_D for MSI cases), it was dichotomized into above and below the median of attributed mutation counts. The binary attributions served as dependent variables in logistic regressions, and relevant risk factors were used as factorized independent variables. Regressions with variables presenting complete orquasi-complete separation (Mansournia et al., 2018) were performed using Firth's bias-reduced logistic regressions based on the logistf R package. To adjust for confounding factors, sex, age of diagnosis, tumor subsite, country, and tumor purity were added as covariates in allregressions. The tumor subsite variable was categorized as proximal colon (ICD-10-CM codes C18.0, C18.2, C18.3, and C18.4), distal colon (C18.5, C18.6, and C18.7), or rectum (C19 and C20), unless otherwise specified. One MSI tumor from an unspecified subsite was removed for the multivariable regression models in MSI cases. The age of diagnosis variable was generally considered as a numerical variable, or categorized into two (early-onset, <50 years old; and late-onset, >50) or five subgroups (0-39, 40-49, 50-59, 60-69, >70), depending on the analysis performed, with specific indications in the corresponding figure legends. Similarly, regressions for cancer genes and hotspot driver mutations (present in at least 10 cases) were done using the same logistic regression models but replacing signature by driver mutation prevalence across samples.
[0324] Regressions with colorectal cancer incidence were performed as linear regressions with signature attributions with confidence intervals not consistent with zero as a dependent variable, and age-standardized rates (ASR) of colorectal cancer (and independent ASR of colon and rectal cancer) obtained from the Global Cancer Observatory (GLOBOCAN) (Bray et al., 2024), sex, age of diagnosis, tumor subsite, and tumor purity as independent variables. Regressions were performed on a sample basis. Regressions with colibactin presence (based on genomic and / or microbiome-derived detection) were performed as linear regressions with age of diagnosis as the dependent variable, and sex, tumor subsite, country, and tumor purity as independent variables.
[0325] Sensitivity analysis for the detection of colibactin signatures in microsatellite stable tumors. To access the resolution for detecting colibactin signatures in MSS cases, the inventors performed simulations for both SBSs and IDs. Specifically, synthetic SBS88 and ID18 mutations were injected at different average levels in each sample (scenarios 1%, 5%, 10%, 15% and 20% for SBS88 and 6%, 10%, 15%, 20% and 25% for ID18) for all MSS cases where the two colibactin-associated signatures were not originally detected. Mutations were injected according to a Gaussian distribution where the mean was equal to a percentage of a sample's total mutational burden, and the standard deviation was equal to 10% of the mean. Importantly, the overall mutational burden for each sample was kept the same by randomly subtracting the same number of mutations that were injected into the sample, while ensuring all mutation counts were still non-negative. Mutational signatures were re-extracted as done for the original data, and MSA attributions were performed using the same penalties applied for the original data. In the SBS context, our analysis indicates that among the non-SBS88 positive cases in the original data (693 / 802), SBS88 was attributed to samples if it contributes at least 2.5% of mutations (median of 0.025 relative proportion at the level of 1% injection). Indeed, MSA attributed SBS88 in about 90.6% of simulation trials at the 1% injection level, approximately 99.7% at the 5% injection level, and 100% at the 10%, 15%, and 20% injection levels. For the ID context, the results indicate that among the non-ID18 positive cases in the original data (649 / 802), ID18 was attributed to samples if it contributed at least 6.6% of mutations. MSA attributed ID18 in about 99.8% of simulation trials at the 6% injection level, and 100% at 10%, 15%, 20%, and 25% injection levels. These results suggest that our analyses are unlikely to have overlooked SBS88 and ID18 in the examined set of MSS colorectal cancers, assuming they contribute at least 1% and 6% of mutations per sample, respectively.Additional statistical analyses. For regressions of signatures, cancer genes, or hotspot driver mutations, p-values were adjusted for multiple comparisons based on the total number of decomposed reference mutational signatures considered per variant type ( / .e., 19 SBS, 7 DBS, 11 ID, 9 CN, and 11 SV signatures for MSS cases, and 18 SBS, 10 DBS, 2 ID, 4 CN, and 4 SV for MSI cases), cancer genes (46 for MSS and 31 for MSI), or hotspot driver mutations (38 for MSS and 14 for MSI) using the Benjamini-Hochberg method (Benjamini and Hochberg, 1995). For country enrichment analyses, the mutation burdens and binary attributions of mutational signatures were compared for each country against all others. Therefore, p-values were also adjusted for multiple comparisons based on the total number of countries assessed (a total of 11 countries). Adjusted q-values were reported, with q-values<0.05 considered statistically significant. For age of diagnosis-based regressions of colibactin presence across tumor subsites, p-values were adjusted based on the total number of tumor subsites assessed (a total of 3 tumor subsites). For the age of diagnosis trend enrichment analysis of signatures, p-trends were reported, with p-trends<0.05 considered statistically significant. For evidence of cooccurrence or mutual exclusivity of two signatures, two-sided Fisher’s exact tests were used, and p-values were reported, with p-values<0.05 considered statistically significant.
[0326] Discussion
[0327] Over the last seven decades, colorectal cancer incidence rates have shown complex changes with marked international variation. Notably, while many high-income countries have seen decreases in overall incidence rates, there has been an increase amongst adults under the age of 50. If these trends continue into older age groups, they could reverse the currently overall positive trajectory for colorectal cancer incidence. In the work described in the present example, whole-genome sequences of 981 colorectal cancers from 11 countries revealed evidence of geographic and age-related variation in their landscapes of somatic mutation which may contribute to explaining these global trends. These variations were mainly found in the 802 microsatellite-stable colorectal cancers with limited geographic or age-related differences detected for colorectal cancers with MSI and no differences for other DNA repair deficiencies, possibly due to the smaller sample size and the dominance of somatic mutations from failed DNA repair processes.
[0328] The prevalence of certain mutational signatures, notably SBS89 / DBS8 / ID_J in Argentina and SBS94 / SBS_F / DBS6 in Colombia, was higher in these countries compared to all others. Although such geographic variation could, in principle, be due to differences in population-specific inheritance, it is more plausible that these are due to differences in exogenous environmental or lifestyle mutagenic exposures. The natures of the putative exposures underlying SBS89 / DBS8 / ID_J and SBS94 / SBS_F / DBS6 are currently unknown. However, SBS89 shares several features with colibactin-induced signatures SBS88 and ID18. SBS89 has been previously found in normal colorectal crypts (Lee-Six et al., 2019) but not in other normal cells. In individuals with SBS89, some crypts have these mutations while others do not. SBS89 appears to be imprinted on the normal colorectal epithelium early in life, with mutagenesis ceasing thereafter (Lee-Six et al., 2019). Moreover, SBS89 mutations show transcriptional strand bias (Lee-Six et al., 2019), a common trait of mutations caused by exogenousmutagenic exposures that form bulky covalent DNA adducts. Thus, SBS89 may also be caused by a mutagen originating from the colorectal microbiome and it is conceivable that multiple microbiome-derived mutagens may contribute to the mutation burden of the colorectal epithelium. The correlations between colorectal cancer ASR and signatures SBS88 and ID18 suggest that microbiome-derived colibactin exposure may influence colorectal cancer incidence rates.
[0329] The evidence for enrichment of SBS88 and ID18 mutation burdens in early-onset colorectal cancers may indicate a role for colibactin exposure in the increase in early-onset colorectal cancer incidence over the last 20 years. Prior studies have indicated that mutagenesis due to colibactin exposure can occur within the first decade of life and then ceases (Lee-Six et al., 2019). In some instances, the mutation burden caused by this early-life mutation burst can endow affected colorectal crypts with the equivalent of decades of mutation accumulation and, thus, this ‘head start’ could plausibly result in an increased risk of early-onset cancers. One mechanism by which colibactin-induced mutagenesis might contribute to colorectal neoplastic change is by somatically inactivating one copy of APC through the generation of protein-truncating driver mutations. Since APC mutations usually occur early in the sequence of driver mutations leading to colorectal cancer (Fearon and Vogelstein, 1990; Gerstung et al., 2020), a first-hit inactivating mutation in APC during early life could put an individual several decades ahead for developing colorectal cancer and resulting in a higher likelihood of early-onset colorectal cancer. The mutation profile of SBS88, with its preponderance of T>C substitutions, is intrinsically ineffective in generating translation termination codons and SBS88 accounts for only a small proportion of APC driver base substitutions. However, colibactin mutagenesis entails a relatively high proportion of ID mutations, with the characteristic profile of ID18, almost all of which will introduce translational frameshifts in coding sequences. ID18 accounts for approximately one quarter of APC indel drivers in colibactin positive cancers and is elevated amongst APC indel drivers compared to indel drivers in other cancer genes such as TP53, which occur later in the multistep process of colorectal carcinogenesis (Carethers and Jung, 2015). Thus colibactin-induced indel driver mutations in APC may account for a substantial proportion of any putative impact colibactin exposure has on colorectal carcinogenesis. Conversely, the unexpected increase in driver mutations observed in late-onset colorectal cancers might suggest that there are additional driver mutational events in early-onset cases, owing to additional effects of colibactin or other mutagenic exposures.
[0330] Future studies should examine the SBS88 (at least, preferably also SBS89 and SBS_M) and ID18 mutation burdens of normal colorectal crypts from individuals with early-onset colorectal cancer (cases) and age-matched healthy individuals (controls) with the expectation of an enrichment in cancer cases if colibactin mutagenesis is causally implicated. If so, the increase in early-onset colorectal cancer over the last 30 years would indicate that an increased exposure to colibactin in affected populations occurred during the middle third of the 20thcentury, and genome sequences of appropriately selected colorectal cancers and normal colorectal tissues would inform on this historical flux. These studies could be supported by international and, if possible, retrospective studies of the prevalence of colibactin-producing pks+ bacteria in the colorectal microbiome. Finally, definitive evidence of a causal role forcolibactin in early-onset colorectal carcinogenesis would be provided by prevention of early-life exposure to colibactin-producing bacteria reducing cancer incidence.
[0331] In summary, mutational epidemiology reveals country-specific and age-specific variations in the prevalence of certain mutational signatures. The results also highlight the potential role of the large intestine microbiome as an early-life mutagenic factor in the development of colorectal cancer.
[0332] Example 2 - Detection of risk of colorectal cancer development from stool and cytobrush samples
[0333] It would be highly beneficial to develop a pre-early detection enrichment test that identifies young individuals at high risk for early-onset colorectal cancer based on stool samples positive for colibactin exposure (such as by testing for the presence of SBS88, ID18, or APC hotspot mutation, which have been shown in Example 1 above to be associated with early onset colorectal cancer, and to be markers of early age exposure to colibactin). In an ideal scenario, everyone would undergo colonoscopy starting at age 18. Currently, this is the standard for individuals over the age of 45 in the United States, following the guideline change in 2018 from 50 to 45 years. However, implementing such a screening protocol for those aged 18 to 45 is impractical for several reasons.
[0334] Firstly, the cost of universal colonoscopy in the age group 18 to 45 would be prohibitively high relative to the expected benefit. Secondly, the adverse effects associated with colonoscopies, such as perforation and bleeding, would outweigh the potential gains in detecting colorectal cancer in a low-risk population (albeit, this may change if the early-onset risk continues to increase). Notwithstanding any changes in guidelines, if we could implement a stool test that significantly enriches for individuals at risk — by a factor of at least 3x, but ideally 5x or more — targeted screening for this high-risk group would become a realistic option.
[0335] In this proposed process, healthy young individuals aged 18 to 45 would provide stool samples or anal swab (e.g. cytobrush) samples as part of a screening program. These samples would be analyzed for mutations and mutational signatures associated with colibactin exposure as identified in Example 1. Based on the presence and quantity of these biomarkers, individuals would be classified as either high-risk or low-risk for developing early-onset colorectal cancer. This last point relies on detection of colibactin signature / mutations from those types of samples as well as establishing accurate risk categories through subsequent clinical trials to define appropriate cutoffs for levels of signatures / mutations. The present inventors conducted proof of concept experiments to show that stool sample and anal swab sample analysis is effective and enables detection of the signatures of interest in healthy subjects (who are by virtue of this presumed to be at higher risk of developing early onset colorectal cancer), enabling subsequent clinical trial development. The data obtained demonstrated that the presence of SBS88 could be detected in anal swab (cytobrush) samples from cancer free subjects (i.e. satisfying quality data could be obtained from these samples, and a subset of subjects tested showed evidence of the presence of the signature - where presence in only a subset of subjects isexpected given the evidence in Example 1 that this may be associated with increased risk of early onset colorectal cancer). The process that was followed to process these samples is described below. Briefly, the stool samples and cytobrush samples were processed in a similar manner (host DNA extraction followed by single-molecule duplex sequencing), except that the stool samples were subject to an additional enrichment step to preferentially isolate host DNA over microbial DNA (which is more abundant in this type of samples than in the cytobrush samples). Note that although the UDseq protocol (Nandi et al. 2025) was used for this, other single-molecule duplex sequencing methods can be used, such as e.g. NanoSeqv2 (Lawson et al. 2025), or other low error rate I error-corrected sequencing strategies. However, the inventors believe that it is particularly beneficial in the context of the present disclosure to use a protocol that achieves an “ultra-low” error rate, such as below one error per 100 million sequenced base pairs. Examples of such ultra-low error rates protocols include the UDseq protocol (Nandi et al. 2025) and the NanoSeqv2 (Lawson et al. 2024) protocol.
[0336] Anal cytobrush samples collection and DNA extraction. All human biospecimens were collected in accordance with institutional ethical guidelines, with written informed consent obtained from all research participants or their legal guardians prior to sample collection. Anal cytobrush samples were received in 12 mL PreservCyt™ solution (Hologic™) and either processed immediately upon receipt or stored at 4°C for short-term storage until processing.
[0337] Upon processing, samples were centrifuged at 3,000 relative centrifugal force (RCF) for 7 minutes at 4°C to pellet cellular material. The supernatant was carefully aspirated to avoid disturbing the pellet, and the remaining cell pellet was resuspended as needed and subjected to genomic DNA extraction using the QIAamp DNA Mini Kit (QIAGEN; Cat. No. 51304, Valencia, CA), following the manufacturer’s protocol. DNA was eluted in nuclease-free buffer and quantified prior to downstream applications. Stool Sample Collection and DNA Extraction. A stool sample is collected using a standard stool collection kit, where the participant deposits a fresh stool sample into a sterile container. Using a collection tool, a small portion of the stool is transferred into a tube containing a DNA stabilization buffer. The tube is securely closed, labeled, and stored at -80°C until further processing.
[0338] Processing of stool samples for host DNA extraction. Host DNA was enriched from stool samples using a selective lysis approach followed by column-based DNA purification, designed to preferentially release host DNA while minimizing bacterial DNA contamination. Briefly, 1 mL of stool sample suspended in phosphate-buffered saline (PBS; 50:50 w / v) was combined with 500 pL Buffer AHL (QIAamp DNA Microbiome Kit) in a 2 mL microcentrifuge tube. Samples were incubated for 30 minutes at room temperature with end-over-end rotation to ensure uniform exposure to lysis buffer. For non-viscous samples, incubation was alternatively performed in a thermomixer at 600 rpm to maintain adequate mixing.
[0339] Following selective lysis, samples were centrifuged at 10,000 x g for 10 minutes at room temperature to pellet cellular debris and insoluble material. The resulting supernatant (-1,000-1,200 pL), containinghost DNA released during lysis, was carefully transferred to a fresh tube and processed for DNA extraction using a QIAamp Mini spin column (Qiagen) according to a modified protocol.
[0340] For protein digestion and nucleic acid preparation, up to 1,200 pL of the Buffer AHL supernatant was supplemented with 20 pL proteinase K and 4 pL RNase A, followed by the addition of 400 pL Buffer AL. Samples were vortexed briefly (15 s) to ensure complete homogenization. When indicated, an additional lysis step was performed by incubating samples at 56 °C for 10 minutes, with brief vortexing once during incubation to enhance protein digestion.
[0341] To promote DNA binding, 800 pL of 96-100% ethanol was added to the lysate and mixed thoroughly by vortexing. The sample was then loaded onto a QIAamp Mini spin column in successive aliquots of up to 700 pL per centrifugation step. Columns were centrifuged at 6,000 x g (8,000 rpm) for 1 minute per loading, with flow-through discarded after each step while reusing the same collection tube.
[0342] Bound DNA was washed sequentially with 500 pL Buffer AW1 (6,000 x g for 1 min) and 500 pL Buffer AW2 (14,000 x g for 3 min). To ensure complete removal of residual wash buffer, an additional dry spin was performed at 14,000 x g for 1 minute. DNA was eluted by placing the column in a clean 1.5 mL microcentrifuge tube and adding 200 pL Buffer AE directly to the membrane, followed by a 1 -minute incubation at room temperature and centrifugation at 6,000 x g for 1 minute.
[0343] For samples requiring enhanced removal of residual contaminants due to higher loading volumes or increased sample complexity, additional wash steps were performed or wash buffers were incubated on the column membrane for 2-3 minutes prior to centrifugation. Eluted DNA was quantified and stored under appropriate conditions prior to downstream analyses.
[0344] DNA quality assurance. Extracted DNA was quantified using a Qubit 4 Fluorometer (Thermo Fisher Scientific) with the dsDNA High Sensitivity assay. Genomic DNA integrity and fragment size distribution were assessed using the Agilent Genomic DNA ScreenTape assay on the Agilent TapeStation system. Samples with high-quality DNA, defined by DNA Integrity Number (DIN) values greater than 8.5 and predominant fragment sizes exceeding 50 kb, were processed for UDSeq library preparation following the standard protocol. For samples exhibiting lower DIN values, the end-repair step during UDSeq library preparation was omitted to avoid errors associated with end-repair.
[0345] Duplex Sequencing, Library Preparation, and Sequencing. Whole-genome and panel duplex sequencing of host genomic DNA was performed to enable highly accurate detection of low-frequency somatic mutations through error correction using complementary DNA strands. This approach can be implemented using several related protocols, including Nanorate sequencing (NanoSeqV2; Lawson et al. 2025) and universal duplex sequencing (UDSeq; Nandi et al. 2025). Here, whole-genome duplex sequencing of host DNA was carried out using the UDSeq protocol, as detailed below.
[0346] UDSeq libraries were prepared as described in Supplementary Note 1 of Nandi et al. 2025 (available a www.biorxiv.org / content / 10.1101 / 2025.09.14.676103v1.supplementary-material), which is incorporated herein by reference. Briefly, high-molecular-weight genomic DNA was enzymatically fragmented using either NEBNext dsDNA Fragmentase (M0348S, New England Biolabs) or UltraShear (M7634L, Covaris) to achieve a mean fragment size of approximately 350 bp. Fragmentation conditionswere optimized for each specimen type; specifically, cytobrush-derived DNA was fragmented for 15 minutes, whereas stool-derived DNA was fragmented for 10 minutes.
[0347] Fragmented DNA was subjected to ligation of unique molecular identifier (UMI)-containing adapters using thexGen™ cfDNA& FFPE DNA Library Preparation Kit (Integrated DNA Technologies). All library preparation steps were performed on magnetic beads to minimize DNA loss and improve library conversion efficiency. Library amplification was carried out using 0.2 fmol of adapter- ligated DNA and 15 PCR cycles, yielding an average of approximately 90* whole-genome duplex coverage.
[0348] For selected samples, particularly stool-derived libraries, targeted hybrid capture was performed following the protocol described in Supplementary Note 1 of the UDSeq method (Nandi et al. 2025). Capture reactions were conducted with six multiplexed libraries per reaction using 1 ,000 ng of adapter-ligated DNA per capture.
[0349] Final libraries were sequenced on Illumina NovaSeq 6000 and NovaSeq X platforms using 150-bp paired-end sequencing chemistry. Somatic mutation calling and estimation of mutational burden from UDSeq data with matched normal samples were performed using DupCaller (v1.0.6) (Cheng Y. et al.
[0350] 2025).
[0351] Mutational Profile and Signature Analysis. The derived somatic mutations are analyzed to identify specific mutational signatures, such as colibactin-associated SBS88 and ID18 (and likely also SBS89 and SBS_M). The analysis also screens for the APC hotspot mutations. If APC hotspot mutation is present and / or SBS88 and ID18 are present above certain levels, the collected stool sample is considered positive and the patient at high risk.
[0352] Specifically, the variant call format files (VCFs) from DupCaller were used for mutational profiles and signatures assignment. Analysis of mutational profiles was performed using the inventors’ previously established methodology with the SigProfiler suite of tools. Briefly, mutational matrices for single base substitutions (SBSs), doublet base substitutions (DBSs), and small insertions and deletions (IDs / indels) were constructed using SigProfilerMatrixGenerator (Version 1.2.16) (Bergstrom et al. 2019). Plotting of each mutational profile was done with SigProfilerPlotting (Version 1.3.13) (Bergstrom et al. 2019). Assistant of mutational signatures to samples was done with SigProfilerAssignment (Diaz-Gay et al.
[0353] 2023).
[0354] Using the methods described in this example, the proof of concept can be expanded to a targeted population of healthy young adults between 18 to 45 years of age (or any age from which local standard of care guidelines recommend regular colonoscopy or other type of screening for early-onset colorectal cancer - 45 being the current age from which regular colonoscopy is recommended in the USA). Indeed, the methods described herein can identify subjects that are below this age but that would nonetheless benefit from the standard of care practice that applies to subject above that age by virtue of being at higher risk of developing early onset colorectal cancer.REFERENCES
[0355] Alexandrov, L. B. et al. Signatures of mutational processes in human cancer. Nature 500, 415-421 (2013a).
[0356] Alexandrov, L. B., Nik-Zainal, S., Wedge, D. C., Campbell, P. J. & Stratton, M. R. Deciphering signatures of mutational processes operative in human cancer. Cell Rep 3, 246-259 (2013b).
[0357] Alexandrov, L. B. et al. Clock-like mutational processes in human somatic cells. Nat Genet 47, 1402-1407 (2015).
[0358] Alexandrov, L. B. et al. The repertoire of mutational signatures in human cancer. Nature 578, 94-101 (2020).
[0359] Ames, B. N., Durston, W. E., Yamasaki, E. & Lee, F. D. Carcinogens are mutagens: a simple test system combining liver homogenates for activation and bacteria for detection. Proc Natl Acad Sci U S 4 70, 2281-2285 (1973).
[0360] Bailey, M. H. et al. Comprehensive Characterization of Cancer Driver Genes and Mutations. Cell 173, 371-385 e318 (2018).
[0361] Benjamini, Y. & Hochberg, Y. Controlling the False Discovery Rate - a Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society Series B-Statistical Methodology 57, 289-300 (1995).
[0362] Bergmann, E. A., Chen, B. J., Arora, K., Vacic, V. & Zody, M. C. Conpair: concordance and contamination estimator for matched tumor-normal pairs. Bioinformatics 32, 3196-3198 (2016).
[0363] Bergstrom, E. N. etal. SigProfilerMatrixGenerator: a tool for visualizing and exploring patterns of small mutational events. BMC Genomics 20, 685 (2019).
[0364] Bray, F. et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin 74, 229-263 (2024).
[0365] Brennan, P. & Davey-Smith, G. Identifying Novel Causes of Cancers to Enhance Cancer Prevention: New Strategies Are Needed. J Natl Cancer Inst 114, 353-360 (2022).
[0366] Carethers, J. M. & Jung, B. H. Genetics and Genetic Biomarkers in Sporadic Colorectal Cancer. Gastroenterology 149, 1177-1190 e1173 (2015).
[0367] Chakravarty, D. et al. OncoKB: A Precision Oncology Knowledge Base. JCO Precis Oncol 2017, 1-16 (2017).
[0368] Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor.
[0369] Bioinformatics 34, i884-i890 (2018).
[0370] Cheng, J. et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381, eadg7492 (2023).
[0371] Cheng, Y. et al. Improved Mutation Detection in Duplex Sequencing Data with Sample-Specific Error Profiles. bioRxiv, 2025.2007. 2013.664565 (2025).
[0372] Cingolani, P. et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly (Austin) 6, 80-92 (2012).
[0373] Cornish, A. J. et al. The genomic landscape of 2,023 colorectal cancers. Nature (2024).
[0374] Degasperi, A. et al. Substitution mutational signatures in whole-genome-sequenced cancers in the UK population. Science 376, abl9283 (2022).Degasperi, A., Amarante, T.D., Czarnecki, J. et al. A practical framework and online tool for mutational signature analyses show intertissue variation and driver dependencies. Nat Cancer1], 249-263 (2020).
[0375] Dentro, S. C. et al. Characterizing genetic intra-tumor heterogeneity across 2,658 human cancer genomes. Cell M, 2239-2254 (2021).
[0376] Diaz-Gay, M. et al. Assigning mutational signatures to individual samples and individual somatic mutations with SigProfilerAssignment. Bioinformatics 39, btad756 (2023).
[0377] Everall, A. et al. Comprehensive repertoire of the chromosomal alteration and mutational signatures across 16 cancer types from 10,983 cancer patients. medRxiv, 2023.2006.2007.23290970 (2023). Fearon, E. R. &Vogelstein, B. A genetic model for colorectal tumorigenesis. Ce / / 61, 759-767 (1990). Gerstung, M. et al. The evolutionary history of 2,658 cancers. Nature 578, 122-128 (2020).
[0378] Gilson, P. et al. Evaluation of 3 molecular-based assays for microsatellite instability detection in formalin- fixed tissues of patients with endometrial and colorectal cancers. Sci Rep 10, 16386 (2020). Goncearenco, A. et al. Exploring background mutational processes to decipher cancer genetic heterogeneity. Nucleic Acids Res 45, W514-W522 (2017).
[0379] Islam, S. M. A. et al. Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor. Cell Genom 2, 100179 (2022).
[0380] Jones, D. et al. cgpCaVEManWrapper: Simple Execution of CaVEMan in Order to Detect Somatic Single Nucleotide Variants in NGS Data. Curr Protoc Bioinformatics 56, 1510 11-1510 18 (2016). Khandekar, A. et al. Visualizing and exploring patterns of large mutational events with SigProfilerMatrixGenerator. BMC Genomics 24, 469 (2023).
[0381] Kim, S. et al. Strelka2: fast and accurate calling of germline and somatic variants. Nat Methods 15, 591-594 (2018).
[0382] Kim, J. E. et al. High prevalence of TP53 loss and whole-genome doubling in early-onset colorectal cancer. Exp Mol Med 53, 446-456 (2021).
[0383] Kucab, J. E. et al. A Compendium of Mutational Signatures of Environmental Agents. Cell 177, 821-836 e816 (2019).
[0384] Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat Methods 9, 357-359 (2012).
[0385] Lawson, A.R.J., Abascal, F., Nicola, P.A. et al. Somatic mutation and selection at population scale. Nature 647, 411-420 (2025). doi.org / 10.1038 / s41586-025-09584-wLee-Six, H. et al. The landscape of somatic mutation in normal colorectal epithelial cells. Nature 574, 532-537 (2019).
[0386] Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform.
[0387] Bioinformatics 25, 1754-1760 (2009).
[0388] Li H. (2013) Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv:1303.3997v2 [q-bio.GN], 16 Mar 2013
[0389] Liao, W. W. et al. A draft human pangenome reference. Nature 617, 312-324 (2023).
[0390] Liu, M., Wu, Y., Jiang, N., Boot, A. & Rozen, S. G. mSigHdp: hierarchical Dirichlet process mixture modeling for mutational signature discovery. NAR Genom Bioinform 5, lqad005 (2023).
[0391] Mansournia, M. A., Geroldinger, A., Greenland, S. & Heinze, G. Separation in Logistic Regression: Causes, Consequences, and Control. Am J Epidemiol 187, 864-870 (2018).Martincorena, I. et al. Universal Patterns of Selection in Cancer and Somatic Tissues. Ce / / 171, 1029-1041 (2017).
[0392] Martinez-Jimenez, F. et al. A compendium of mutational cancer driver genes. Nat Rev Cancer 20, 555-572 (2020).
[0393] Moody, S. et al. Mutational signatures in esophageal squamous cell carcinoma from eight countries with varying incidence. Nat Genet 53, 1553-1563 (2021).
[0394] Muinos, F., Martinez-Jimenez, F., Pich, O., Gonzalez-Perez, A. & Lopez-Bigas, N. In silico saturation mutagenesis of cancer genes. Nature 596, 428-432 (2021).
[0395] Nandi SP, et al. A Universal Duplex Sequencing Approach for Accurate Detection of Somatic Mutations. bioRxiv [Preprint], 2025 Sep 16:2025.09.14.676103. doi: 10.1101 / 2025.09.14.676103. Nik-Zainal, S. et al. Mutational processes molding the genomes of 21 breast cancers. Cell 149, 979-993 (2012a).
[0396] Nik-Zainal, S. et al. The life history of 21 breast cancers. Cell 149, 994-1007 (2012b).
[0397] Nunes, L. et al. Prognostic genome and transcriptome signatures in colorectal cancers. Nature (2024).
[0398] Patel, S. G., Karlitz, J. J., Yen, T., Lieu, C. H. & Boland, C. R. The rising tide of early-onset colorectal cancer: a comprehensive review of epidemiology, clinical features, biology, risk factors, prevention, and early detection. Lancet Gastroenterol Hepatol 7, 262-274 (2022).
[0399] Pleguezuelos-Manzano, C. et al. Mutational signature in colorectal cancer caused by genotoxic pks(+) E. coli. Nature 580, 269-273 (2020).
[0400] Raine, K. M. et al. cgpPindel: Identifying Somatically Acquired Insertion and Deletion Events from Paired End Sequencing. Curr Protoc Bioinformatics 52, 15 17 11-1517 12 (2015).
[0401] Rosendahl Huber, A. et al. Improved detection of colibactin-induced mutations by genotoxic E. coli in organoids and colorectal cancer. Cancer Cell 42, 487-496 (2024).
[0402] Senkin, S. etal. Geographic variation of mutagenic exposures in kidney cancer genomes. Nature (2024).
[0403] Senkin, S. MSA: reproducible mutational signature attribution with confidence based on simulations. BMC Bioinformatics 22, 540 (2021).
[0404] Siegel, R. L. et al. Global patterns and trends in colorectal cancer incidence in young adults. Gut 68, 2179-2185 (2019).
[0405] Siegel, R. L., Jemal, A. & Ward, E. M. Increase in incidence of colorectal cancer among young men and women in the United States. Cancer Epidemiol Biomarkers Prev 18, 1695-1698 (2009).
[0406] Sinicrope, F. A. Increasing Incidence of Early-Onset Colorectal Cancer. N Engl J Med 386, 1547-1558 (2022).
[0407] Sondka, Z. et al. COSMIC: a curated database of somatic variants and clinical data for cancer.
[0408] Nucleic Acids Res 52, D1210-D1217 (2024).
[0409] Sondka, Z. et al. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat Rev Cancer 18, 696-705 (2018).
[0410] Spaander, M. C. W. et al. Young-onset colorectal cancer. Nat Rev Dis Primers 9, 21 (2023).
[0411] Steele, C. D. et al. Signatures of copy number alterations in human cancer. Nature 606, 984-991 (2022).Stigliano, V., Sanchez-Mete, L., Martayan, A. & Anti, M. Early-onset colorectal cancer: a sporadic or inherited disease? World J Gastroenterol 20, 12420-12430 (2014).
[0412] Terlouw, D. et al. Recurrent APC Splice Variant c.835-8A>G in Patients With Unexplained Colorectal Polyposis Fulfilling the Colibactin Mutational Signature. Gastroenterology 159, 1612-1614 e1615 (2020).
[0413] Van Loo, P. et al. Allele-specific copy number analysis of tumors. Proc Natl Acad Sci U SA 107, 16910-16915 (2010).
[0414] Venugopal, A. & Carethers, J. M. Epidemiology and biology of early onset colorectal cancer. EXCLI J 21, 162-182 (2022).
[0415] Vuik, F. E. et al. Increasing incidence of colorectal cancer in young adults in Europe over the last 25 years. Gut 68, 1820-1826 (2019).
[0416] Whalley, J. P. et al. Framework for quality assessment of whole genome cancer sequences. Nat Common 11, 5040 (2020).
[0417] You, Y. N., Xing, Y., Feig, B. W., Chang, G. J. & Cormier, J. N. Young-onset colorectal cancer: is it time to pay attention? Arch Intern Med 172, 287-289 (2012).
[0418] Zhang, T. et al. Genomic and evolutionary classification of lung cancer in never smokers. Nature Genetics 53, 1348-1359 (2021).
[0419] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.
Claims
CLAIMS1. A method of determining whether a subject is at risk of developing colorectal cancer, the method comprising:obtaining nucleic acid sequence data from a stool sample or an anal swab sample from said subject,quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound, anddetermining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer.
2. The method of claim 1 , wherein the microbial genotoxic compound is colibactin.
3. The method of any preceding claim, wherein the one or more metrics are selected from:metrics that quantify the presence or activity of a mutational signature associated with exposure to a microbial genotoxic compound, optionally colibactin, andmetrics that quantify the presence, number or proportion of mutations associated with exposure to a microbial genotoxic compound, optionally colibactin.
4. The method of any preceding claim, wherein quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound comprises determining the presence of one or more mutational signatures, the number of mutations associated with one or more mutational signatures, the proportion of mutations associated with one or more mutational signatures, and / or the activity of one or more mutational signatures, wherein the one or more mutational signature are mutational signatures associated with exposure to colibactin, andwherein a mutational signature comprises a plurality of proportions for respective categories of mutations, optionally wherein the plurality of categories of mutations are categories of single base substitutions or categories of indels.
5. The method of claim 3 or claim 4, wherein the one or more mutational signatures are selected from:(i) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS88) is characterised by: (a) a higher proportion of T>C mutations in ATA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in ATA context, T>C mutations in ATT context, T>C mutations in TTT context, and T>G mutations in TTT context than any other categories of single base substitutions;(ii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS89) ischaracterised by a higher proportion of T>G mutations in GTG context and C>T mutations in ACA context than any other categories of single base substitutions;(iii) a single base substitution signature comprising proportions for each of a plurality of categories of single base substitutions each defined by a unique combination of substitution type and trinucleotide context of the substitution, wherein the single base substitution signature (SBS_M) is characterised by: (a) a higher proportion of T>C mutations in CTA context than any other categories of single base substitutions, or (b) a higher proportion of T>C mutations in CTA context, T>C mutations in TTA context, and T>A mutations in CTA context than any other categories of single base substitutions; and(iv) an indel signature comprising proportions for each of a plurality of categories of indels each defined by a unique combination of whether the indel is a deletion or insertion, whether the indel is a 1 bp indel or longer, for 1 pb indel whether the indel is a C or T and the length of the mononucleotide repeat tract in which they occur, for longer indels whether the indel is an insertion at a repeat region and the deletion length and number of repeat units, whether the indel is a deletion at a repeat region and the insertion length and number of repeat units, and whether the indel is a deletion with microhomology at the deletion boundary and the deletion length and microhomology length, wherein the indel signature (ID18) is characterised by higher proportions of 1bp T deletions than any other categories of indels.
6. The method of any of claims 3 to 5, wherein the mutational signature is selected from: SBS88, SBS89, SBS_M and ID18, or corresponding signatures obtained by: extracting one or more signatures from a cohort of samples using the same plurality of categories, and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS88, SBS89 and / or SBS_M, or ID18; optionally wherein the mutational signature is SBS88, ora corresponding signature obtained by: extracting one or more signatures from a cohort of samples using the same plurality of categories, and mapping the extracted signatures to a catalogue of previously obtained signatures comprising SBS88.
7. The method of any preceding claim, wherein quantifying, using said sequence data, one or more metrics indicative of the presence of mutations associated with exposure to a microbial genotoxic compound comprises determining:the percentage or proportion of total W[T>N]W mutations that are within a WAWW[T>N]W motif, where W is A or T, and N is A, C, G or T; and / orthe presence of a splicing variant c.835-8A>G in the APC gene; and / orthe presence of one or more protein-truncating mutations in the APC gene; and / or the presence of indels in the APC gene associated with 1 bp T or A deletions and / or associated with indel signature ID18.
8. The method of any preceding claim, wherein the subject is a human subject, wherein the subject is a human subject under the age of 50, or wherein the subject is a human subject between the ages of 18 and 50.
9. The method of any preceding claim, wherein the subject is a healthy subject, and / or wherein the subject is a subject who has not been diagnosed with Lynch syndrome, and / or wherein the subject is a subject who has not been identified as having one or more germline mutations associated with DNA mismatch repair deficiency.
10. The method of any preceding claim, wherein the sample is a stool sample, and / or wherein a subject who is determined to be at risk of developing colorectal cancer is recommended or selected for a further colorectal cancer screening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy.
11. The method of any preceding claim, wherein the sequence data comprises DNA sequencing reads, or a list of mutations or a mutational profile derived therefrom, and / or wherein the sequence data comprises DNA sequencing reads obtained from bulk sequencing or error-corrected duplex sequencing, or information derived therefrom; and / or wherein the sequence data comprises DNA sequencing read obtained using a sequencing methodology associated with a sequencing error rate below one error per 10 million sequenced base pairs, or information derived therefrom; and / or wherein the sequence data comprises DNA sequencing read obtained using a single molecule duplex sequencing methodology, or information derived therefrom.
12. The method of claim 11 , wherein the sequence data comprises DNA sequencing reads from whole genome sequencing, whole exome sequencing or targeted panel sequencing, or information derived therefrom; optionally wherein the sample is a stool sample and the sequence data comprises DNA sequencing reads from a library derived from the sample and subject to a target sequence capture, optionally using an exome capture panel, or information derived therefrom.
13. The method of any preceding claim, wherein the colorectal cancer is rectal cancer and / or wherein the colorectal cancer is early-onset colorectal cancer.
14. The method of any preceding claim, wherein the sample is a stool sample that has been previously obtained from the subject, optionally fresh frozen, and from which DNA has been extracted and processed for sequencing, or wherein obtaining the sequence data comprises extracting DNA from a stool sample and / or processing extracted DNA from a stool sample for sequencing and / or sequencing the processed extracted DNA from the stool sample, optionally wherein DNA has been or is extracted from the sample using selective lysis of the subject’s cells over microbial cells followed by DNA purification, optionally column-based DNA purification, orwherein the sample is an anal swab sample that has been previously obtained from the subject, optionally stored in a cytology preservation solution, and from which DNA has been extracted and processed for sequencing, or wherein obtaining the sequence data comprises extracting DNA from an anal sample and / or processing extracted DNA from an anal swab sample for sequencing and / or sequencing the processed extracted DNA from the anal swab sample, orwherein the method comprises one or more of: receiving a stool sample previously obtained from the subject, optionally in a sterile container; storing the stool sample at -80C; extracting DNA from the sample from the subject; extracting DNA from the sample from the subject using a protocol comprising a step of selective lysis of the subject’s cells over microbial cells followed by DNA purification; testing the extracted DNA for DNA integrity and optionally performing an end-repair step when a DNA integrity number for the extracted DNA is above a predetermined threshold; obtaining a DNA sequencing library from DNA extracted from the sample from the subject, optionally including adding unique molecular identifiers to the DNA fragments in the sample and / or performing enrichment of sequences comprising predetermined sequences using target sequence capture; and sequencing a DNA library obtained from the sample from the subject, optionally using singe molecule duplex sequencing, orwherein the method comprises one or more of: receiving an anal swab sample previously obtained from the subject, optionally in a cytology preservation solution; storing the anal swab sample at 4C; extracting DNA from the sample from the subject; extracting DNA from the sample from the subject, optionally using a standard DNA extraction protocol; testing the extracted DNA for DNA integrity and optionally performing an end-repair step when a DNA integrity number for the extracted DNA is above a predetermined threshold; obtaining a DNA sequencing library from DNA extracted from the sample from the subject, optionally including adding unique molecular identifiers to the DNA fragments in the sample; and sequencing a DNA library obtained from the sample from the subject, optionally using singe molecule duplex sequencing.
15. The method of any preceding claim, wherein the sequence data comprises DNA sequencing reads and the method comprises:processing said DNA sequencing reads to identify single base substitutions and / or indels present in the sequence data, andoptionally obtaining a mutational profile from said identified substitutions and / or indels by classifying identified mutations in a predetermined plurality of categories.
16. The method of any preceding claim, wherein determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer comprises classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer.
17. The method of claim 16, wherein classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lowerrisk of developing colorectal cancer comprises comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein values of the metrics or score at or above the respective predetermined thresholds is indicative of a higher risk of developing colorectal cancerthan values of the metrics or score below the respective predetermined thresholds.
18. The method of claim 17, wherein the respective predetermined thresholds are thresholds that have been identified using a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer.
19. The method of claim 17 or claim 18, wherein classifying, using said one or more metrics, the subject between a first class that has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer comprises:comparing each of said one or more metrics or a score derived therefrom to respective predetermined thresholds, wherein a subject is:classified in the first class when said score is at or above a predetermined threshold, and classified in the second class otherwise;classified in the first class when each of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise; orclassified in the first class when a predetermined number or combination of said one or more metrics is at or above a respective predetermined threshold, and classified in the second class otherwise.
20. The method of any preceding claim, wherein determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer comprises:inputting said one or more metrics into a machine learning classifier trained to classify subjects between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer using training data comprising the one or more metrics for a cohort of subjects comprising a plurality of subjects that have developed early onset colorectal cancer and a plurality of subjects that have not developed early onset colorectal cancer.
21. The method of any preceding claim, further comprising:(i) selecting the subject for a further colorectal cancer screening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy, when the subject is determined to be at risk of developing colorectal cancer; and / or(ii) providing a report comprising an output of the step of determining, based on said one or more metrics, whether the subject is at risk of developing colorectal cancer, or information derived therefrom selected from: an indication of whether the subject was determined to be at risk of developing colorectal cancer, a classification between a first class has a higher risk of developing colorectal cancer, and a second class that has a lower risk of developing colorectal cancer, arecommendation for a patient monitoring regimen comprising one or more further colorectal cancer screening tests.
22. The method of any preceding claim, further comprising performing a further colorectal cancer screening test, optionally an invasive colorectal cancer screening test, optionally a colonoscopy, on the subject when the subject is determined to be at risk of developing colorectal cancer.
23. A method of selecting a subject for invasive screening for early onset colorectal cancer, the method comprising:determining whether the subject is at risk of developing colorectal cancer using the method of any of claims 1 to 20, andselecting the subject for invasive screening for early onset colorectal cancer when the subject is determined to be at risk of developing colorectal cancer.
24. A method of treating a subject for early onset colorectal cancer, the method comprising:determining whether the subject is at risk of developing colorectal cancer using the method of any of claims 1 to 20, andperforming invasive screening for early onset colorectal cancer when the subject is determined to be at risk of developing colorectal cancer.
25. A system comprising:a processor; anda non-transitory computer readable medium comprising instructions that, when executed by the processor, cause the processor to implement the method of any of claims 1 to 20.
26. A non-transitory computer readable medium or media comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any of claims 1 to 20.