Prediction of gene amplification and deletion
A predictive model using tumour DNA sequencing data and various genomic features effectively forecasts future gene amplifications and deletions, addressing the limitations of current methods and enabling early intervention in tumour management.
Patent Information
- Application Number
- PCT/EP2024/082564
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods are unable to accurately predict future gene amplification or tumour suppressor gene deletion in tumours, which limits early intervention and effective disease management.
A computer-implemented method that uses tumour DNA sequencing data to develop a predictive model. This model incorporates pan-cancer chromosomal instability signatures, local mutational processes, DNA topology features, and oncogene fitness effects to forecast future oncogene amplification and tumour suppressor gene deletion.
The method enables accurate prediction of gene amplifications and deletions, allowing for early intervention and improved disease management by identifying tumours at high risk of acquiring these alterations.
Smart Images

Figure EP2024082564_22052025_PF_FP_ABST
Abstract
Description
[0001] PREDICTION OF GENE AMPLIFICATION AND DELETION
[0002] FIELD OF THE DISCLOSURE
[0003] The present invention relates to methods of predicting whether a subject will acquire a gene amplification, and in particular to methods of predicting the likelihood of future gene amplification in a tumour genome based on copy number signature exposure in a sample comprising DNA from said tumour.
[0004] BACKGROUND
[0005] Genes that regulate normal cell growth can become constitutively active oncogenes that drive uncontrolled tumour growth via amplification of the number of copies or other modifications to the tumour DNA. Amplification of certain oncogenes can be observed at high frequency in specific tumour types, for example, ERBB2 amplification in 15-20% of breast cancers. The tissue-type specific nature of recurrent amplifications implies that each tumour type has its own set of unique selective pressures and amplification fitness effects.
[0006] The detection of an oncogene amplification in a tumour can serve as a prognostic or diagnostic biomarker. Therapeutic targeting of oncogenes enabled by companion diagnostic testing has been a major focus in precision oncology with highly effective therapies approved for use against EGFR, HER2 and FGFR2, among others. Oncogene amplification can also serve as a prognostic biomarker for improved disease management, for example, intensified treatment in MYCN amplified neuroblastomas. However, these interventions require that the amplification is detected via a diagnostic test, at which point the amplification has been clonally selected and is likely to be a key driver of tumour growth.
[0007] The presence of oncogene amplification in a tumour often represents a critical point in tumour evolution where the tumour is more aggressive and potentially difficult to treat. Identification of tumours at an earlier evolutionary time point that don’t yet have an amplification but are at high-risk of acquiring one in the future, offers the opportunity to realise the benefit of oncogene-based interventions earlier, potentially with preventive intent. However, there is no method that is able to do this accurately for a range of oncogenes at present.
[0008] SUMMARY
[0009] Activation of an oncogene via DNA amplification can underpin tumour initiation, progression and treatment resistance. The ability to predict future amplification using a genomic test offers new opportunities for improved disease management and treatment. The present disclosure provides a method to forecast future oncogene amplification and / or tumour suppressor gene deletion that only relies on tumour DNA sequencing data. The present inventors hypothesised that predicting oncogene amplification or tumour suppressor gene deletion would be possible by developing a predictive model that includes both mutational processes - represented by a compendium of pan-cancer chromosomal instability (CIN) signatures that they previously published - and selection, and a predictive model based on this model and that further extends it to uses mutational processes activity estimated from the input tumour DNA copy number profile - represented by a compendium of pan-cancer chromosomal instability (CIN) signatures previously published, combined with local mutational processes, DNA topology features and oncogene fitness effects learnt from a large cohort of tumours. They verified this hypothesis through extensive investigation of the mechanisms that underpin oncogene amplification and its relationship with signatures of chromosomal instability. Based on this, they developed methods to forecast oncogene amplification and tumour suppressor deletion for individual patients. They trained predictive models for each oncogene and tumour suppressor using the different types of chromosomal instability found underlying oncogenic amplifications across cohorts of patients. They validated the approach using longitudinal pairs of glioma, lung cancer, and prostate cancers. To demonstrate clinical utility they forecasted poor prognosis in gliomas via CDK4 amplification and CDKN2A deletion, and osimertinib resistance in lung cancers via MET amplification. They further provided guidelines and recommendations for clinical implementation paving the way for a new class of biomarker wherein selective pressures and mutation-generating processes are harnessed to anticipate future genomic alterations and thereby forecast tumour evolution.
[0010] Thus, according to a first aspect, there is provided a computer-implemented method of predicting whether a tumour is likely to acquire a gene copy number alteration, the method comprising: obtaining a tumour copy number profile for the tumour; determining, for each of one or more signatures of chromosomal instability, a metric indicative of the presence of the signature in the tumour copy number profile, wherein a signature of chromosomal instability comprises a set of weights for each of a plurality of components each associated with a copy number feature of a set of copy number features quantified for a plurality of segments in the copy number profile, wherein the weight of a component is indicative of the probability that a mutational process associated with the signature generates a segment having characteristics defined by the component; and determining whether the tumour is likely to acquire a gene copy number alteration based on the output of a model trained to take as input the values of said metrics and produce as output a score indicative of the likelihood that the tumour will acquire a gene copy number alteration, wherein the model has been trained using copy number profiles for a plurality of tumours comprising tumours that have the gene copy number alteration and tumours that do not have the gene copy number alteration, wherein the copy number alteration is an amplification or homozygous deletion.
[0011] The method may have any one or more of the following optional features.
[0012] The tumour copy number profile for the sample may identify one or more genomic segments, wherein each segment is associated with a copy number that differs from the copy number of the neighbouring segment(s).
[0013] The gene may be an oncogene. In such embodiments the copy number alteration may be an amplification. The gene may be an oncogene selected from the genes listed in Tables 7 (a and b) and 8 (a and b) and on Figure 9. The model may be specific to a single gene. Alternatively, the model may be predictive for a plurality of genes. The plurality of genes may be oncogenes, such as oncogenes selected from the genes listed in Tables 7 (a and b) and 8 (a and b) and on Figure 9. The gene may be a tumour suppressor gene. In such embodiments, the copy number alteration may be a homozygous deletion. The gene may be a tumour suppressor gene selected from the genes listed in Tables 7c, 8c, 9c and on Figure 45. The model may be specific to a single gene. Alternatively, the model may be predictive for a plurality of genes. The model may be specific to a single gene and tumour type. Alternatively, the model may be predictive for a plurality of tumour types.
[0014] The model may be a model that has been trained to take as input the values of said metrics for each of a plurality of signatures of chromosomal instability and to produce as output a score indicative of the likelihood of acquiring a gene copy number alteration. The model may have been trained to provide as output a score indicative of the likelihood of acquiring a gene copy number alteration at a single specific gene. The model may have been trained to provide as output a score indicative of the likelihood of acquiring a gene copy number alteration at any of a plurality of genes. The plurality of genes may each be oncogenes. The plurality of genes may each be tumour suppressor genes. The plurality of genes may comprise one or more or all oncogenes selected from Tables 7 (Table 7a, 7b) and 8 (Table 8a, Table 8b). The plurality of tumour suppressor genes may comprise one or more or all tumour suppressor genes selected from those listed in Tables 7c, 8c, 9c, or Figure 45.
[0015] A gene amplification may refer to the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number >=6 or a copy number >=4 above tumour ploidy. A gene amplification may refer to the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number >=8. A gene amplification may refer to the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number >=8, if present in an autosome, or >=6 if present in a sex chromosome.
[0016] The copy number features may comprise or consist of: segment size, copy number change-point, and breakpoint count in one or both flanking regions of a predetermined length. The breakpoint count may be a summarised value of the breakpoint count in the upstream and downstream flanking regions of a predetermined length around the segment. The summarised value may be the maximum. The predetermined length may be between 2 and 10 Mb, between 3 and 8 Mb, or about 5 Mb. Alternatively, the copy number features may comprise or consist of: segment size, copy number change-point, breakpoint count per predetermined length of sequence, breakpoint count per chromosome arm, and number of segments with oscillating copy number, optionally wherein the predetermined length of sequence is 10Mb.
[0017] Determining a metric indicative of the presence of a signature in the tumour copy number profile may comprise quantifying the value of each copy number feature for each of one or more segments in the copy number profile. Each of the one or more segments may be a non-diploid segment. Quantifying the set of copy number features may comprise using unrounded copy number segments. Obtaining a copy number profile may comprise collapsing and merging near diploid segments to a diploid state, wherein near diploid segments are segments that have a copy number within a predetermined distance from 2. The predetermined distance may be 0.1.
[0018] Determining a metric indicative of the presence of a signature in the tumour copy number profile may comprises obtaining, for each of a plurality of segments in the copy number profile, a value for each of the plurality of components for the segment. Each component may define characteristics of a segment as a predetermined distribution of values of a copy number feature. The value determined for each segment may be the probability of a quantified copy number feature of the segment belonging to each of a set of predetermined distribution associated with respective components of the plurality of components, for the copy number feature. Thus, the probability of the feature value from a segment belonging to the predetermined distribution associated with the component may be determined. This may be performed for each component, thereby obtaining a an “Event-by-component” probabilities vector for each segment, or an event-by-component probability matrix comprising respective vectors for each of a plurality of segments from the same copy number profile. These probabilities may be referred to as “posterior probabilities”.
[0019] The components may define a respective distribution of values of the associated copy number feature. The components may be selected from those in Table 1 or those in Table 1 1 , or corresponding distributions obtaining by fitting mixture models to values for the set of copy number features obtained from copy number profiles from a plurality of tumour samples. Each copy number feature may be associated with a plurality of components, and each component may define a characteristic of copy number events using a respective predetermined distribution, wherein the predetermined distributions for a copy number feature that is quasi-continuous (e.g. segment size, copy number changepoint) is a set of Gaussian distributions, and / or wherein the predetermined distributions for any count feature (e.g. breakpoint count in one or both flanking regions of a predetermined length) is a set of Poisson distributions. The predetermined distributions associated with the components may be the distributions defined by the parameters in Table 1 , or corresponding distributions obtained by fitting mixture models to values for the set of copy number features obtained from copy number profiles from a plurality of tumour samples. The predetermined distributions may comprise between 15 and 25 Gaussian distributions for the segment size, between 5 and 15 Gaussian distributions for the copy number changepoint, and / or between 5 and 15 Poisson distributions for the breakpoint count in one or both flanking regions of a predetermined length.
[0020] The set of copy number features may comprise a copy number change-point, wherein the copy number changepoint of a current segment is determined by reference to the upstream segment of the current segment when the copy number changepoint relative to the upstream segment is at or above a predetermined threshold, by reference to the downstream segment of the current segment when the copy number changepoint relative to the upstream segment is below the predetermined threshold and the copy number changepoint relative to the downstream segment is at or above the predetermined threshold, and by reference to a reference copy number value when the copy number changepoints relative to the upstream and downstream segments are below the predetermined threshold. The reference copy number value may be 2. The predetermined threshold may be -3.
[0021] The metric indicative of the presence of a signature in the tumour copy number profile may be the activity of the signature in the tumour copy number profile, optionally wherein the activity of one or more signatures in a tumour copy number profile is determined by: obtaining, for each component, the sum of probabilities associated with the respective component for each of a plurality of segments in the copy number profile; and determining a linear combination of one or more signatures comprising the signature that satisfies the equation PbC = E x SbC (Equation 1) where E is a vector of size n comprising coefficients Einwhere Ei is the activity of signature i; PbC is a vector of size c, each element in the vector representing a sum of probabilities associated with a respective component; and SbC is a matrix of size c by n, each value representing the weight of a component in a signature / , and the probability associated with a component for a segment is the probability of a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component.
[0022] The signatures may be selected from those in Table 3 or those in Table 12. The signatures may be selected from CX1 , CX2, CX3, CX4, 5, CX8, CX9, CX10, and CX13 defined in Table 3 or Table 12 or corresponding signatures obtained by quantifying the set of copy number features in a plurality of tumour samples, and identifying one or more signatures likely to result in the copy number profiles of the plurality of tumour samples by nonnegative matrix factorisation.
[0023] The model may be specific to a tumour type that is the same tumour type as that of the tumour that is being analysed. Such a model may be obtained by training the model using a reference cohort of samples comprising samples of said tumour type.
[0024] The trained model may be a model that produces a score indicative of the likelihood of acquiring an gene copy number alteration (e.g. oncogene amplification or tumour suppressor deletion) as a weighted combination of the values of the metrics indicative of the presence of the one or more signatures in the tumour copy number profile. The combination may be a linear combination. The model may produce a score according to the following equation: where CXmis the activity of signature m in the tumour copy number profile (also referred to herein as tumour mutational processes term), wmis a signature weight for signature m for the gene, and n is the total number of signatures considered in the model. The signature weight wmmay be obtained as a product of terms comprising a fitness effect term (foG.t ; frsG.t) and a local mutational processes term & TSG,t,m)- For example, the model may produce a score according to the following equation for an oncogene: or the following equation for a tumour suppressor gene:
[0025] The signature weight wmmay be obtained as a product of terms comprising a fitness effect term (foG.t), a local mutational processes term (<voG,t,m) and a DNA topology term (XoG.(.m). Thus, the model may produce a score according to the following equation: Alternatively, the signature weight wmmay be obtained as a local mutational processes term ( oG,t,m ; MTSG.t.m) , i.e. the model may produce a score according to the following equation for an oncogene: scoreOG t— 2m=iCXm* Woc.t.m or the following equation for a tumour suppressor gene:
[0026] SCOreTSG t= 2m=l CXm* TSG,t,m
[0027] A DNA topology term may advantageously be used in the context of predicting oncogene amplification. A DNA topology term may not be used in the context of predicting tumour suppressor gene homozygous deletion. A local mutational processes term may be estimated in a tumour type specific manner or across multiple tumour types. This may be used to obtain a signature weight as a local mutational processes term. Aternatively, fitness effect terms (foG.t ; frsG.t) and local mutational processes terms (MDG.t.m &>TSG,t,m) estimated across multiple individual tumour types may be aggregated to obtain a signature weight wmacross the multiple individual tumour types. The aggregation may comprise averaging the products of the fitness effect terms (f c.j ; frsG.t) and local mutational processes terms MrsG.t.m) , across the multiple individual tumour types. The multiple individual tumour types may be tumour types where the copy number alteration of the specific gene is under positive selection. The aggregation may comprise calculating a signature weight (also referred to as pan-cancer or pan-tumour signature weight or amplification score or deletion score) as: types, such as all tumour types in which the gene deletion / amplification is under positive selection.
[0028] The weights of the model may be determined in a tumour type-specific manner using a reference cohort of samples by determining the local mutational processes term as, for each signature, a ratio of (i) an averaged signature score of all amplified (or homozygously deleted) segments spanning the gene in samples of the specific tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment. Said metric may be the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum is the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures.
[0029] The weights of the model may be determined in a tumour type-agnostic manner using a reference cohort of samples by determining the local mutational processes term as, for each signature, a ratio of (i) a weighted average signature score of all amplified segments spanning the gene in samples of a specific tumour type in the reference cohort, wherein the average signature scores in the specific tumour type are weighted by the normalised empirical frequency of the gene amplification (or homozygous deletion, in the case of deletion prediction) in the respective tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment. Said metric may be the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum is the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures.
[0030] QOG.t.m
[0031] ^OG,t,m = -
[0032] In embodiments, the weights (com) are determined as(Im where:
[0033] &)oG,t, m is the weight for gene OG for signature m in samples of tumour type t in a reference cohort,
[0034] E .OG a=lea, m QOG,t,m — 4 qoG,t,m is given by:At,OG where ea,mis the activity-corrected score signature
[0035] (m) for a particular amplified event (a) (of homozygous deletion event, when predicting TSG deletion) encompassing gene OG in samples of tumour type t in the reference cohort; At OG is the number of amplified segments (or homozygously deleted segments) spanning gene OG in the samples of the tumour type (f) in the reference cohort, and qmis given bymP where ep mis the activity-corrected score of signature m for a particular copy number event (p) in a copy number profile of the reference cohort; and P is the total number of copy number events observed in the copy number profiles of the reference cohort. Such a model may be referred to as a tumour-type specific model.
[0036] QOG,m
[0037] ^OG,m
[0038] In embodiments, the weights (c m) are determined as Qm where:
[0039] C OG, m is the weight for gene OG for signature m in the reference cohort, qoG,m, is given by where qoG,t,m is given by: where ea,mis the activity-corrected score signature (m) for a particular amplified event (a) (or deletion event, when predicting deletions) encompassing gene OG in samples of tumour type t in the reference cohort; A.OG is the number of amplified segments (or deletion segments, when predicting deletions) spanning gene OG in the samples of the tumour type (f) in the reference cohort, qmis given by P where ep mis the activity-corrected score of signature m for a particular copy number event (p) in a copy number profile of the reference cohort; and P is the total number of copy number events observed in the copy number profiles of the reference cohort, and foG.t is the empirical frequency of the amplification of gene OG (or homozygous deletion of gene OG, when predicting deletions) in a tumour type f in the reference cohort; and Tis the number of tumour types in the reference cohort. Such a model may be referred to as a pancancer model.
[0040] In embodiments, the weights (&)m) are determined as &)0G>t m* f0G t(for an oncogene) or a)TSGit,m * frsc,t (for a tumour suppressor) or a)0G,t,m * foc,t *xoG,t,m (foranoncogene). The notations w0G t mand (^TSG.t.m refer to a local mutational processes term. The notations f0G tand fTSG trefer to a fitness effect term. The notation X0G t mrefers to a DNA topology term.
[0041] A local mutational process term ( oG,t,m; < rsG,t,m) reflects the propensity of plurality of CIN signature underlying the copy number alteration of the gene in a given tumour type (t). The local mutational processes term may be estimated by quantifying the enrichment of signatures assigned to copy number altered segments harbouring a given gene over a signature background. The local mutational processes term can be estimated for a particular tumour type t using a cohort of samples of the tumour type, comprising samples with a copy number alteration of the gene and samples without a copy number alteration of the gene. The local mutational processes term may be a vector of length equal to the number of signatures (n). For example, when using the signatures in Table 3 the term may be a vector of length 9. The local mutational process term for a copy number alteration may be estimated as, for each signature m of a set of n signatures: where G is a specific gene, which can be either an oncogene OG or a tumour suppressor gene TSG, and the numerator is the relative contribution of signature m to the copy number event of the gene across the cohort of the tumour type t divided by the relative contribution of signature m to any copy number event in the cohort. The denominator is the sum of the numerator term across the n signatures, i,e, a normalisation term across signatures. The term qt,m is the relative contribution of a signature m to the acquisition of any copy number event in tumour type t, computed by averaging the activity- corrected signature scores of all copy number events found in the tumour-specific training cohort. It can be calculated as: where £p mis the activity-corrected score of signature m for a particular copy number event (p), and Et represents the total number of copy number events observed in the training cohort of a given tumour type (f). The term qc,t,m is the relative contribution of signature m to the copy number alterations of gene G for tumour type t. It may be calculated by averaging the activity-corrected signature scores for signature m of all copy number events (amplification or homozygous deletion, depending on which is being predicted) spanning a given gene G in a specific tumour type according to the following formula: qG
[0042] 't’m- AtG where £a mrepresents the activity- corrected score of signature m for a particular copy number event of gene G, and At,G represents the number of segments spanning a particular gene G (OG or TSG) in a given tumour type (t) that are amplified or deleted.
[0043] The activity corrected score of signature m for an event (denoted e below, which can be a or p above) can be calculated as: the sum is shown over n=9 but should be read to encompass any set of signatures n, <D(S)is an event (e) by component (c) probabilities matrix for sample s (i.e. indexing on e makes this a vector of probabilities for event e across components c) , with elements <^e.c, ct is a signature (m) by component (c) definition matrix (i.e. the weights defining the signatures) e.g. as indicated in Table 3 or Table 12 or any other signature definition, with elements oc,m, and CXm<s)are sample-level signature activities.
[0044] When computing these activity corrected scores for homozygous deletions (e.g. in the context of calculating a local mutational process term for a tumour suppressor gene), the probabilities for the changepoint feature for 0-copy segments containing the deletion can be modified such that a changepoint feature component representing 1 -copy change is assigned a probability of 1 , and all other components are assigned probabilities of 0. 0-copy segments may be segments that have an absolute copy number below a predetermined threshold, such as e.g. 0.5.
[0045] For each tumour type, a “local mutational processes” matrix containing the weights of the n signatures for any one or more genes G (e.g. oncogenes OG and / or tumour suppressors TSG) may be constructed. To enhance predictive capacity for forecasting oncogenes in a given tumour type, signature weights learnt from multiple cohorts may be integrated, such as e.g. publicly available cohorts such as the TCGA, the PCAWG and the HMF dataset. In the integration process, the predictive capacity of the framework in the multiple cohorts using a retrospective approach may be initially assessed. Subsequently, for each tumour type, the gene-specific signature weights from the cohort with the highest predictive capacity based on Youden’s index may be selected. In cases where the copy number alteration of a gene is putatively not under selection in a specific tumour type (and therefore not tested for predictive capacity), the gene-specific signature weights from the cohort with at least >5 samples displaying the copy number alteration may be selected. Alternatively, the weights may be estimated using a single cohort.
[0046] A DNA topology term (XoG.(.m) reflects the effects of DNA topology constraints on the configuration of amplifications spanning the gene of interest (e.g. oncogene OG) in a given tumour type (t). The DNA topology term may be a vector of length equal to the number of signatures (n). For example, when using the signatures in Table 3 the term may be a vector of length 9. The DNA topology term may be estimated for a particular tumour type t using a cohort of samples of the tumour type comprising an amplification of the gene of interest. The DNA topology term may be estimated by determining the ratio of (i) a metric quantifying the local bias for each of the components associated with copy number features of the set of copy number features for which the signatures are defined, in segments encompassing amplifications of the gene of interest in the cohort; and (ii) a corresponding metric quantified for segments corresponding amplification events at a plurality of genes that are not oncogenes. The metrics may be the product of: (1) the average across the segments considered, of the posterior probability of the copy number features quantified for the segment belonging to each of the components; and (2) the weights of the signatures. The DNA topology term can be calculated as, for each signature m: where is the average of the posterior probability component vector across oncogene amplified segments in the cohort (for the oncogene OG), <p is the average of the posterior probability component vector across passenger (non-oncogene) amplified segments in the samples of the cohort that comprise the oncogene amplification, and a is the signature by component matrix for the signatures m (i.e. the weights defining the signatures m=1 to n). Any signature associated with mitotic error or any process not affected by local chromatin structure may be assigned a DNA topology weight of 1 . The values of the terms for all the signatures can then be normalise to sum to 1 across the plurality of signatures.
[0047] A fitness effect term (foG.t ; frsG.t)) may represent the selective advantage conferred by the copy number alteration in a given tumour type. The fitness effect term can be estimated for a particular tumour type t using a cohort of samples of the tumour type, comprising samples with a copy number alteration of the gene and samples without a copy number alteration of the gene. The fitness effect term may be a vector of length equal to the number of signatures (n). For example, when using the signatures in Table 3 the term may be a vector of length 9. A fitness effect may be estimated from a cohort comprising at least 5 samples with the copy number alteration of the gene. A fitness effect term may be determined as the frequency of samples in the cohort that carry a given copy number alteration of the gene. A fitness effect term for an oncogene may be calculated as: where AoG.t is the number of samples with the amplification of a given oncogene (OG) and tumour type (t), and Sf represents the total number of samples for a given tumour type (t) in the cohort. Frequency values for oncogenes not found to confer a selective advantage may be set to zero. This effectively sets the tumour-specific score to 0 for a particular gene in tumour types where there is no evidence of amplification of the gene. Similarly, a fitness effect term for a tumour suppressor gene may be calculated as: where Arsa.t is the number of samples with the homozygous deletion of the given tumour suppressor gene (TSG) in the tumour type (t), and Sf represents the total number of samples for the given tumour type (t). The frequency values of a tumour suppressor gene that does not confer a selective advantage in the tumour type may be set to zero. As above, this effectively sets the tumour-specific score to 0 for a particular gene in tumour types where there is no evidence of deletion of the gene. Whether a gene copy number alteration confers a selective advantage may be determined as a gene copy number alteration that is significantly amplified or deleted in the tumour type in a cohort of samples compared to background. For example, this may be a gene that is subject to focal amplification or deletion in the tumour type.
[0048] The weights determined in a tumour type-specific manner may be used when the reference cohort comprises more than a threshold number of samples of the specific tumour type or the frequency of amplification of the gene (or deletion of the gene, when prediction deletions) in the specific tumour type in the reference cohort is higher than a threshold frequency. The threshold number of sample may be 5 and the threshold frequency may be 5%. The weights determined in a tumour type-agnostic manner may be used otherwise.
[0049] The tumour may be a primary and / or untreated tumour. The tumour may be a treated and / or advanced tumour and the weights of the model are determined in a tumour type-agnostic manner using for (i) only tumour types in which the gene is recurrently amplified in the reference cohort, optionally wherein a gene is recurrently amplified in a tumour type if it is amplified in more than 5% of the samples of the tumour type in the reference cohort.
[0050] The output of model may be a score given by any one of: where m refers to one of CX1 , CX2, CX3, CX4, 5, CX8, CX9, CX10, and CX13 defined in Table 3 or
[0051] Table 12 or corresponding signatures obtained by quantifying the set of copy number features in a plurality of tumour samples, and identifying one or more signatures likely to result in the copy number profiles of the plurality of tumour samples by nonnegative matrix factorisation, CX represents the signature activities in the tumour copy number profile, and wmrepresents the signature weights computed for a given gene, including a single term as described above (m= or a plurality of terms including a local mutational processes term, a fitness effect term and optionally a DNA topology term, as described above. The score for a prediction of a risk of deletion of a gene in a tumour may be obtained as described above then divided by the number of copies observed in the tumour genome at the locus of the gene.
[0052] The method may further comprise classifying the tumour between a first class that has a high likelihood of acquiring the gene copy number alteration and a second class that has a low likelihood of acquiring the gene copy number alteration, wherein the tumour is classified in the first class when the output of the model is above a predetermined threshold. The predetermined threshold may have been determined by obtaining the output of the model for each of a plurality of samples in a reference cohort with a known expected % of samples with the gene copy number alteration, and selecting the threshold that is such that a % of samples corresponding to the expected % of samples are classified in the first class. The predetermined threshold may have been determined by obtaining the output of the model for each of a plurality of samples in a reference cohort with known gene copy number alteration status, and selecting the threshold that satisfies one or more predetermined criteria that apply to the sensitivity, selectivity or Youden’s index when used to classify the samples between the first and second classes.
[0053] The gene may be selected form those in Table 9 (Table 9a or Table 9b). The gene may be selected from MET, AR, EWSR1 , JUN, RBFOX2, POLD1 , MDM2, CDK4, HIST1 H3B, and ERBB2. The gene may be selected from Table 9 (Table 9a or Table 9b) and the model may be a corresponding linear model with signature-specific weights specified in Table 7 (Table 7a or Table 7b) or Table 8 (Table 8a or Table 8b), or equivalent model trained using a different reference cohort.
[0054] The model may have been trained using a reference cohort comprising samples that are one or more of: the same tumour type, the tumour grade / stage, undergone the same treatment, have the same driver mutations, have been analysed with the same sequencing technology, as the tumour to be analysed. The tumour may be selected from sarcoma, breast cancer, lung adenocarcinoma, ovarian cancer, bladder carcinoma, glioblastoma, head and neck squamous cell carcinoma, renal adenocarcinoma, lung squamous cell carcinoma, adenocarcinoma, stomach adenocarcinoma, esophagal carcinoma, liver hepatocellular carcinoma, Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma, Uterine Corpus Endometrial Carcinoma, Colon Adenocarcinoma, Brain Lower Grade Glioma.
[0055] Also described herein according to a second aspect is a computer-implemented method of predicting whether a tumour is likely to respond to a therapy, the method comprising: predicting whether the tumour is likely to acquire a gene copy number alteration (amplification or deletion) using the method of any embodiment of the first aspect, wherein the gene copy number alteration is associated with resistance to the therapy, and whether a tumour predicted as likely to acquire the gene copy number alteration is unlikely to respond to the therapy. The gene may be MET, the copy number alteration may be an amplification, and the therapy may be an EGFR inhibitor. The method may further comprise recommending a subject with a tumour that is unlikely to respond to the therapy for upfront treatment with both EGFR and MET inhibitors. The gene may be BRAF, the copy number alteration may be an amplification, and the therapy may be a MEK inhibitor. The gene may be androgen receptor (AR), the tumour may be a prostate tumour and the therapy may be androgen deprivation therapy.
[0056] Also described herein according to a third aspect is a computer-implemented method of providing a prognosis for a subject with a tumour, the method comprising: predicting whether the tumour is likely to acquire a gene copy number alteration using the method of any embodiment of the first aspect, wherein the gene amplification is associated with poor prognosis. The gene may be CDK4, he copy number alteration may be an amplification, and the tumour may be a low grade glioma.
[0057] Also described herein according to a fourth aspect is a computer-implemented of providing a model for predicting whether a tumour is likely to acquire a gene copy number alteration, the method comprising: obtaining, a plurality of tumour copy number profiles for a reference cohort of samples, the plurality of tumour copy number profiles comprising copy number profiles that comprise the gene copy number alteration and copy number profiles that do not comprise the gene copy number alteration; and determining, for each of one or more signatures of chromosomal instability, a weight that is a ratio of (i) an average signature score of all copy number altered segments spanning the gene in samples of a specific tumour type in the reference cohort, optionally wherein the average signature scores in the specific tumour type are weighted by the normalised empirical frequency of the gene copy number alteration in the respective tumour type in the reference cohort and (ii) an average signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment. The copy number alteration may be an amplification and amplified segments spanning the gene may be considered. The copy number alteration may be a homozygous deletion and homozygous deleted segments spanning the gene may be considered, Said metric may be the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures.
[0058] The method according to the current aspect may have any of the features described in relation to the first aspect.
[0059] According to any embodiment of any aspect, a tumour copy number profile or sample may have been obtained from a subject who has been diagnosed as having cancer. A tumour copy number profile may have been obtained by analysing sequence data. The sequence data may have been obtained using sequencing. The sequencing may be selected from WGS, sWGS, single cell WGS, WES or panel sequencing, or a copy number array. A tumour copy number profile may identify one or more genomic segments, wherein each segment is associated with a copy number that differs from the copy number of the neighbouring segment(s). A set of copy number features may be quantified for each of the segments, or one or more of the segments identified in the copy number profile (e.g. all segments identified on one chromosome or sets of chromosomes, etc.). The methods are not limited in this regard and even single events (segments) may be analysed productively. The terms “segment” and “genomic segment” are used interchangeably.
[0060] The method may further comprise determining that the copy number profile comprises a number of nondiploid segments above a predetermined threshold. The copy number profile may comprise at least 15 non-diploid segments. Samples with less than 20 non-diploid segments may automatically be classified as low risk of amplification. Samples with less than 15 non-diploid segments may be classified as low risk of homozygous deletion of a tumour suppressor gene. Samples with less than 15 non-diploid segments may be excluded from a training cohort used to train a model to predict a gene deletion. This may exclude samples that already have the homozygous deletion of a specific tumour suppressor gene or an amplification of the specific oncogene.
[0061] Obtaining the tumour copy number profile may comprise determining the tumour copy number profile from sequence data. The sequence data may be next-generation sequencing data. Determining the tumour copy number profile may comprise determining a copy number for each of a plurality of genomic bins of a predetermined size. The predetermined size may be about 100kb, at most 100kb, at most 50kb, about 50kb, or about 30kb. The method may comprise obtaining a tumour copy number profile from DNA sequence data comprising a plurality of sequence read, wherein the tumour copy number profile comprises a copy number estimate for each of a plurality of genomic bins comprising a number of mapped reads above a predetermined threshold. The predetermined threshold may be 15 reads per genomic bin per chromosome copy.
[0062] Methods of the disclosure also encompass treating a subject who has been diagnosed with cancer, the method comprising determining whether the subject is likely to respond to a therapy as described herein and administering the therapy to a subject who has been determined as likely to respond to the therapy and / or administering a treatment that does not comprise the therapy or that comprises the therapy in addition to a further therapy to a subject who has been determined as unlikely to respond to the therapy.
[0063] According to a further aspect, there is provided one or more non-transitory computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any embodiment of any preceding aspect.
[0064] According to a further aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform the steps of any method described herein, such as a method according to any embodiment of any of the first to fourth aspects.
[0065] According to a further aspect, there is provided system comprising: a processor; and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any embodiment of any of the first to fourth aspects.
[0066] BRIEF DESCRIPTION OF THE FIGURES
[0067] Figure 1 is a flowchart illustrating schematically a method of predicting whether a tumour in a subject is likely to acquire a gene amplification (A), and a method of obtaining a model for predicting whether a tumour in a subject is likely to acquire a gene amplification (B).
[0068] Figure 2 is a flowchart illustrating schematically methods of providing a prognosis, identifying a therapy or treating a subject.
[0069] Figure 3 shows an embodiment of a system for performing methods of the disclosure.
[0070] Figure 4 shows a schematic overview of copy number features and CIN signature identification according to embodiments of the disclosure, a) Schematic of feature encoding design and derivation of feature components. Mixture modelling was applied to split the 3 feature distributions into smaller components by either Variational Bayes Gaussian mixture models (for segment size and change point) or Finite Poisson mixture models (for breakpoints in flanking 5Mb). b) For all samples and features, the probability of each copy number event (no diploid segments) to belong to each mixture component is computed. Then, component probabilities of all events are summed to obtain a 41 -dimensional feature vector per sample. In the illustrated examples, the activities of CIN signatures according to Drews et al. 2022 were already known for all 6,335 samples of the TCGA cohort (see Example 2 below). Therefore, linear combination decomposition was applied for deriving signature definitions in the new feature space.
[0071] Figure 5 shows data investigating the impact of modifying the changepoint feature encoding in CIN signatures. The heatmap shows Pearson correlation coefficients between CIN signature activities derived by using the feature encoding (x-axis) described in Drews et al. 2022 and those derived by extracting the 5 original features but using the modified encoding version of the changepoint feature described in Example 1. Red colour denotes a high positive correlation between activities, while blue colour denotes a high negative correlation.
[0072] Figure 6 displays histograms and fitted density curves showing the distribution of the breakpoints per side feature in the TCGA cohort (see Example 1). For visualisation purposes, the data are split in four subsets according to the number of breakpoints within a 5 Mb window per segment side. The combined distribution is used for deriving the different feature components capturing the signal of this novel feature.
[0073] Figure 7 shows a comparison of CIN signatures encoded using the original features (Drews et al. 2022) and the new features of the present disclosure. The heatmap shows Pearson's r correlation coefficients between activities of original (x-axis) and novel (y-axis) CIN signatures computed in 6,335 TCGA samples. Red colour denotes a high positive correlation between activities, while blue colour denotes a high negative correlation.
[0074] Figure 8 shows data investigating the acquisition, selection and expansion of focal amplifications in triple-negative breast cancer. A Schematic showing the ability to identify unique and clonally expanded focal amplifications in single cell and bulk DNA sequencing data. B Histograms showing the expected number of oncogene amplifications in single cells (left) and in the TCGA-TNBC cohort (right) based on 1 ,000 iterations with lists of random genes. The red dashed line highlights the observed number of oncogenes amplified. Statistical significance was measured by calculating the empirical p-value. C Scatter plot comparing the number of unique focal amplifications detected across all single cells with the number of clonally expanded focal amplifications seen at pseudobulk resolution in 8 triple-negative breast cancer (TNBC) samples and 4 TNBC cell lines. D. Density plot showing the relative observed distance between adjacent unique amplifications observed across 16,178 cells from 12 TNBC samples. The red dashed line highlights the relative expected distance between adjacent amplifications under random acquisition. P-value is estimated by one-sample Wilcoxon’s two-sided test. E. This shows an overview of simulations of oncogene amplification acquisition and selection. Schematics of birth-death simulations of tumour growth (fitness advantage: s, death rate: a, birth rate: p, amplification rate: ju). Tumour cells without amplification die with rate a, divide yielding two cells without amplification with rate j?(1-ju) and divide giving rise to one cell with amplification and another without with rate j?ju (left box). Tumour cells with amplifications (right box) have a fitness advantage and as such die with rate a(1-s) and divide with rate j?(1 +s), giving rise to two cells with amplification. F. Mean fraction of cells harbouring an oncogene amplification at the end of 20 replicate birth-death tumour growth simulations using different combinations of selection coefficient (x-axis) and amplification rate (y-axis).
[0075] Figure 9 shows a summary of the different oncogenes positively selected to be amplified per tumour type. An oncogene was considered a putatively amplified oncogene in a tumour type when it was amplified in > 5 samples in the TCGA tumour-specific cohort, and it was previously identified as significantly amplified in the given tumour type by GISTIC.
[0076] Figure 10 shows a heatmap highlighting the enrichment of CIN signatures (signature weights) in amplified segments harbouring a specific oncogene across TCGA. Signature weights were learnt from the TCGA cohort by computing the ratio of the oncogene-specific signature score (i.e. the average of the signatures assigned to each amplified copy number event harbouring the specific oncogene) to the background signature score (i.e. the average of the signatures assigned to all copy number segments).
[0077] Figure 11A shows a heatmap representing the signature enrichment in amplified segments mapping to different oncogenes in TCGA-PRAD. Only those oncogenes whose amplification is positively selected in prostate cancer are included.
[0078] Figure 11 B shows a heatmap representing the signature enrichment in amplified segments mapping EGFR across tumour types. Only tumour types in which EGFR putatively confers a fitness advantage are included. Figure 12A-C illustrates schematically a process of building a signature-based predictor of focal amplification according to embodiments of the disclosure, a) Rationale for building a chromosomal instability signature-based approach for predicting focal amplifications, b) Schematic for computing the weight of a CIN signature to cause the amplification of a given oncogene. Signature weights were derived by evaluating 271 oncogenes and using 6,335 TCGA tumours of 33 cancer types. CXarepresents mean signature scores of amplified segments spanning a given oncogene. fampsrepresents the fraction of samples from a tumour type out of the total number of samples across tumour types harbouring the amplification of the specific oncogene; note that fampsis equal to 1 when calculating tumour-type specific weights, and CXbackgroundrepresents mean signature scores of all copy number events (see Methods for further details), c) Schematic of the signature-based predictor application. The amplification score of a sample is computed by summing up the result of multiplying activities of each CIN signature by its relative weight to the acquisition of an oncogene amplification.
[0079] Figure 12D illustrates schematically the different factors contributing to the amplification of a given oncogene.
[0080] Figure 13 illustrates schematically a workflow for applying a signature-based predictor according to the disclosure to a new clinical sample.
[0081] Figure 14 shows results of analyses assessing the ability to predict the acquisition of amplifications of specific oncogenes in TRACERx Lung (126 pairs, lung adenocarcionma and non-small cell lung cancer) and the Hartwig Medical Foundation cohorts (227 pairs, differenttumourtypes). Plots show the matched oncogene-specific model (i.e. HIST1 H3B in a), AR in b) and ERBB2 in c)) in the y-axis, the acquisition of the amplification in the x axis and a threshold to categorise patients based on the score into high risk and low risk of amplification groups, a) Acquisition of HIST1 H3B amplification in TRACERx Lung, b) Acquisition of AR amplification in subset of prostate cancers with paired data in the Hartwig Medical Foundation cohort and c) Acquisition of ERBB2 amplification in all patients with paired data in the Hartwig Medical Foundation cohort.
[0082] Figure 15 shows principles and results of analyses testing signature-based predictors of focal amplification according to embodiments of the disclosure, a-b) Pan-cancer assessment of predictor performance in the TCGA (training cohort) and PCAWG (held-out cohort). Violin plots showing differences in amplification scores in tumours with and without an amplification in at least one oncogene across TCGA (left) and PCAWG (right). Only oncogenes amplified in a minimum of 10 samples were included, c-d) Testing the ability to predict future oncogene amplifications using paired longitudinal samples from the GLASS cohort. Signature-based scores of acquiring an oncogene amplification were computed in the early biopsy sample, and box plots in d) show comparison between late biopsy samples with and without oncogene amplification acquired. All p-values were estimated by Wilcoxon two-sided rank test.
[0083] Figure 16 shows ROC curves showing performance of the signature-based predictor in the TCGA and PCAWG cohorts. The ROC curve showing the average performance across all individual oncogenes is represented in purple. Only oncogenes amplified in a minimum of 10 samples (grey lines) were included. An oncogene was considered amplified when its absolute copy number was > 8.
[0084] Figure 17 shows the rationale for an alternative procedure to train the predictor of oncogene amplification using the latest biopsy from a cohort of patients with paired biopsies.
[0085] Figure 18 shows results of gene-specific assessment of retrospective performance of predictors of the disclosure in the TCGA cohort. Box plots showing amplification scores (y-axis) in TCGA samples with and without amplification in a specific oncogene. A subset of 10 key oncogenes are illustrated. P-values estimated by Wilcoxon rank test.
[0086] Figure 19 shows results of gene-specific assessment of retrospective performance of predictors of the disclosure in the PCAWG cohort. Box plots showing amplification scores (y-axis) in PCAWG samples with and without amplification in a specific oncogene. A subset of 10 key oncogenes are illustrated. P- values estimated by Wilcoxon rank test.
[0087] Figure 20 shows results of a tumour-specific assessment of performance of predictors of the disclosure. Bar plots showing the average oncogene-specific performance (AUC values) by tumour type in a) the TCGA and b) the PCAWG cohorts. Only oncogenes amplified in a minimum of 10 samples per tumour type were included.
[0088] Figure 21 shows results of an assessment of performance of predictors of the disclosure in the PCAWG cohort after removing segments connected to focal amplifications by complex / simple structural variants (SVs). Box plots showing amplification scores (y-axis) in PCAWG samples with and without focal amplifications. Scores of samples with amplifications were calculated after removing or not segments connected to amplifications by SVs.
[0089] Figure 22 shows data investigating the optimal threshold for defining focal amplifications(A) AUROC and p-value of a Wilcoxon two-sided test (comparing amplification scores between samples with and without oncogene amplifications) of a pan-cancer pan-oncogene model for forecasting oncogene amplifications for each of a plurality of a threshold values for calling focal amplifications, using absolute copy number values (ranging from >6 to >10 copies). (B) AUROC and p-value of a Wilcoxon two-sided test (comparing amplification scores between samples with and without oncogene amplifications) of a pan-cancer pan-oncogene model for forecasting oncogene amplifications for each of a plurality of a threshold values for calling focal amplifications, using copy number values relative to tumour ploidy (ranging from >4 to >8 copies respect to tumour’s ploidy).
[0090] Figure 23 shows results of methods of predicting EGFR tyrosine kinase resistance according to embodiments of the disclosure, a) Experimental validation of predicting EGFR-TKI resistance via MET amplification using a panel of NSCLC cell lines treated with EGFR-TKI until the emergence of treatment resistance. Signature-based MET amplification scores were computed in parental cell lines, and the presence of MET amplification in resistant clones was tested via ddPCR. Amplification scores were normalised by the total number of CNAs in each parental cell line. Red colour indicates resistance driven by MET amplification; b) Experimental validation of prediction of EGFR-TKI resistance via MET amplification by associating signature-based MET amplification scores of 3 different parental HCC827 clones (y-axis) with the frequency of resistance clones acquiring MET amplification in culture (x-axis). MET amplification was determined via ddPCR.
[0091] Figure 24 shows results of methods of predicting prognosis in low grade gliomas according to embodiments of the disclosure, a) Kaplan-Meier curves showing overall survival for low grade glioma (LGG) patients from TCGA with CDK4 amplification (red), without CDK4 amplification but high probability to acquire it in the future (orange), and without CDK4 amplification and unlikely to acquire it (green). P-value was estimated by the log-rank test, b) Cox proportional hazards model for LGG-TCGA patients without CDK4 amplification. Within each molecular subtype, patients are classified as high or low risk of acquiring CDK4 amplification in the future by a signature-based classifier as described herein. Model was corrected for the fraction of genome altered and age at diagnosis.
[0092] Figure 25 shows results of a cox proportional hazards model for LGG-TCGA patients with and without CDK4 amplification. Model was corrected for the fraction of genome altered and age at diagnosis. Age at diagnosis was categorised into four groups by dividing cohort distribution into quantiles.
[0093] Figure 26 shows results of a cox proportional hazards model for LGG-TCGA patients without CDK4 amplification. Patients are classified as high or low risk of acquiring CDK4 amplification in the future by a signature-based classifier as described herein. Model was corrected forthe fraction of genome altered and age at diagnosis. Age at diagnosis was categorised into four groups by dividing cohort distribution into quantiles.
[0094] Figure 27 shows results of a cox proportional hazards model for LGG-TCGA patients with and without CDK4 amplified but at risk of acquiring it in the future according to a signature-based classifier as described herein. Model was corrected for the fraction of genome altered and age at diagnosis. Age at diagnosis was categorised into four groups by dividing cohort distribution into quantiles.
[0095] Figure 28 shows a framework for forecasting oncogene amplifications. A schematic representation of our forecasting framework. This framework takes as input a copy number profile, and optionally, the tumour type and an oncogene of interest. It then outputs whether the tumour has a high or low risk of acquiring amplification of an oncogene. This framework combines four main terms impacting on the emergence of oncogene amplifications: b) the activity of DNA damage processes, which is estimated from the patient sample; c) the fitness effect of the oncogene amplification, which is learnt from a large cohort of tumours; and d) the local mutational processes that act to amplify an oncogene and e) the DNA topology constraints, which are estimated by combining. The product of these three terms yields a score of amplification risk, which is then thresholded to forecast oncogene amplification risk.
[0096] Figure 29 Forecasting performance, a-b) ROC curves showing performance of the forecasting framework for a) AR amplification in post-treatment prostate cancer samples, and b) HIST1 H3B amplification in metastatic lung cancer samples. The forecasting framework was applied in the pretreated or primary tumour samples, respectively. The black dot represents the threshold learnt from a reference cohort.
[0097] Figure 30 Forecasting performance of three genomic features across longitudinal data, a-b) ROC curves showing performance of four genomic features (presence of CIN, number of CNAs, presence of amplifications, and number of amplifications) for forecasting the acquisition of a) AR amplification in post-treated prostate cancer samples, and b) HIST1H3B amplification in metastatic lung cancer samples. The predictions were performed in the early time point sample. The black dot highlights the effectiveness of the predictions by applying the optimal threshold derived by the Youden’s index.
[0098] Figure 31 shows data investigating the stability of predictor tools as described herein across different copy number profiling technologies. Scatter plot showing amplification scores derived from SNP6 (x- axis) and sWGS (y-axis) data from the same set of 478 tumours. Pearson correlation test used for comparing amplification scores between technologies.
[0099] Figure 32 shows correlation between amplification scores computed using the full or the universal set of CIN signatures in the TCGA cohort. Scatter plot showing amplification scores derived from the TCGA dataset using all 16 CIN signatures (x-axis) and those obtained using only the 9 universal signatures (y-axis). Pearson correlation test used for comparing amplification scores between signature sets. For comparison purposes, we averaged the oncogene-specific amplification scores per sample.
[0100] Figure 33 shows correlation between amplification scores computed using the breast-specific or the universal set of CIN signatures in the TCGA-BRCA cohort. Scatter plot showing amplification scores derived from the TCGA-BRCA dataset using the signatures specifically extracted in the cohort (x-axis) and those obtained using the 9 universal signatures (y-axis). Pearson correlation test used for comparing amplification scores between signature sets. For comparison purposes, we averaged the oncogene-specific amplification scores per sample.
[0101] Figure 34 shows results of an analysis of retrospective performance of our forecasting framework across a range of thresholds for defining focal amplifications, a) Results from the ploidy-independent strategy, b) Results from the ploidy-aware strategy. Plots show the median and interquartile range across oncogenes in each amplification threshold. (A) AUROC and p-value of a Wilcoxon two-sided test (comparing amplification scores between samples with and without oncogene amplifications) of a pan-cancer pan-oncogene model for forecasting oncogene amplifications for each of a plurality of a threshold values for calling focal amplifications, using absolute copy number values (ranging from >6 to >10 copies). (B) AUROC and p-value of a Wilcoxon two-sided test (comparing amplification scores between samples with and without oncogene amplifications) of a pan-cancer pan-oncogene model for forecasting oncogene amplifications for each of a plurality of a threshold values for calling focal amplifications, using copy number values relative to tumour ploidy (ranging from >4 to >8 copies respect to tumour’s ploidy).
[0102] Figure 35 shows results of methods of predicting EGFR tyrosine kinase resistance according to embodiments of the disclosure, a) Experimental validation of predicting EGFR tyrosine kinase inhibitor (EGFR-TKI) resistance via MET amplification using a panel of non-small cell lung cancer (NSCLC) cell lines treated with EGFR-TKI until the emergence of treatment resistance. MET amplification scores were computed in parental cell lines, and the presence of MET amplification in resistant clones was tested via digital droplet PCR (ddPCR). Red colour indicates resistance driven by MET amplification.; b) Contingency table showing the number of cells acquiring and not acquiring MET amplification after treating the HCC827 parental cell line with EGFR-TKI alone or in combination with a MET inhibitor (METi). P-value was estimated by Fisher’s exact test.
[0103] Figure 36. Acquisition of MET amplification in EGFR-TKI resistant clones. Resistant clones were derived from 4 NSCLC cell lines after long-term culture with EGFR-TKI. Red colour denotes clones acquiring MET amplification as the resistant mechanism, while green colour highlights clones not acquiring MET amplification. Droplet digital PCR was used to detect MET amplification, and a 2-fold increase from the parental cell line was considered as acquiring a MET amplification.
[0104] Figure 37. Forecasting MET amplification-driven resistance to erlotinib in lung cancer cell line cultures. Experimental validation of predicting MET amplification using a panel of NSCLC cell lines treated with erlotinib until the emergence of treatment resistance 5. Scores of MET amplification risk were computed in parental cell lines, and the mechanism of treatment resistance acquired in resistant cultures was evaluated using different experimental techniques as previously reported 5. Red colour indicates resistance driven by MET amplification.
[0105] Figure 38. Single-cell DNA sequencing of treatment-resistant HCC827 cells. Boxplot showing the absolute copy number value of MET in HCC827 single cells untreated and after 1 -month treatment with EGFR-TKI or EGFR-TKI and METi. The box bounds the interquartile range divided by the median. Red colour denotes cells with at least 8 copies of MET, while green colour denotes cells without focal amplification of MET. Horizontal dashed line indicates the absolute copy number value of MET in the parental cell line at bulk level. P-value was estimated by Kruskal-Wallis test.
[0106] Figure 39 is a REMARK diagram. Flow diagram summarising the quality control filtering of patients to get the curated cohort for forecasting the acquisition of MET amplification as a resistance to osimertinib treatment.
[0107] Figure 40 shows Kaplan-Meier curves showing progression-free survival for NSCLC patients predicted as having low (green) or high (red) risk of acquiring MET amplification as the resistant mechanism to osimertinib. P-value was estimated by the log-rank test. Cox proportional hazards regression models showing progression-free survival for NSCLC patients predicted as having low (green) or high (red) risk of acquiring MET amplification as the resistant mechanism to osimertinib. The model was corrected for the fraction of genome altered, the presence of metastasis in the central nervous system (CNS) at treatment initiation, and the physical status based on the Eastern Cooperative Oncology Group (ECOG) scale at treatment initiation. The threshold for classifying the patients was obtained from the in vitro study in Figure 35a.
[0108] Figure 41 shows Kaplan-Meier curves showing overall survival for NSCLC patients predicted as having low (green) or high (red) risk of acquiring MET amplification as the resistant mechanism to osimertinib. P-value was estimated by the log-rank test. Cox proportional hazards regression models showing overall survival for NSCLC patients predicted as having low (green) or high (red) risk of acquiring MET amplification as the resistant mechanism to osimertinib. The model was corrected for the fraction of genome altered, the presence of metastasis in the central nervous system (CNS) at treatment initiation, and the physical status based on the Eastern Cooperative Oncology Group (ECOG) scale at treatment initiation. The threshold for classifying the patients was obtained from the in vitro study in Figure 35a.
[0109] Figure 42 shows results of a cox proportional hazards model for LGG patients from TCGA and GLASS cohorts with and without CDK4 amplification and CDKN2A homozygous deletions. Model was corrected for the fraction of genome altered and age at diagnosis. Age at diagnosis was categorised into four groups by dividing cohort distribution into quantiles. Figure 43 shows results of methods of predicting prognosis in low grade gliomas according to embodiments of the disclosure. Cox proportional hazards model for LGG patients from TCGA and GLASS cohorts with and without CDK4 amplification and CDKN2A deletions. Within each molecular subtype, patients are classified as high or low risk of acquiring CDK4 amplification and / or CDKN2A deletion in the future by the forecasting framework as described herein. Model was corrected for the fraction of genome altered and age at diagnosis.
[0110] Figure 44 shows results of methods of predicting prognosis in low grade gliomas according to embodiments of the disclosure. Cox proportional hazards model for LGG patients from TCGA and GLASS cohorts with and without CDK4 amplification and CDKN2A deletions. All LGG patients are classified as high or low risk of acquiring CDK4 amplification and / or CDKN2A deletion in the future by the forecasting framework as described herein. Model was corrected for the fraction of genome altered and age at diagnosis.
[0111] Figure 45 shows a summary of the different tumour suppressor genes positively selected to be deleted per tumour type. An TSG was considered a putatively deleted TSG in a tumour type when it was deleted in > 5 samples in the TCGA tumour-specific cohort, and it was previously identified as significantly deleted in the given tumour type by GISTIC.
[0112] Figure 46 is a heatmap showing the mean of the sum-of-posterior probability component vectors of all copy number segments harbouring a particular TSG homozygous deletion (TSG-specific feature components weights) - . Weights are normalised per feature for ease in denoting the differences in terms of amplicon size, amplitude and density of flanking breakpoints across TSGs. Mixture components (22 components for the segment size feature, 10 components for the change point feature and 9 components for the flanking breakpoints feature) were derived from feature distributions. Top: Samples without detectable CIN (<20 CNAs). Bottom: Samples with detectable CIN (>=20 CNAs).
[0113] Figure 47 is a heatmap showing the mean of the sum-of-posterior probability component vectors of all copy number segments harbouring a particular TSG single-copy deletion (TSG-specific feature components weights). Weights are normalised per feature for ease in denoting the differences in terms of amplicon size, amplitude and density of flanking breakpoints across TSGs. Mixture components (22 components for the segment size feature, 10 components for the change point feature and 9 components for the flanking breakpoints feature) were derived from feature distributions. Top: Samples without detectable CIN (<20 CNAs). Bottom: Samples with detectable CIN (>=20 CNAs).
[0114] Figure 48 shows violin plots showing the distribution of copy number values in segments flanking homozygous deletions. Each dot represents a segment. Red colour indicates diploid tumours, while blue indicates tetrapioid tumours (ploidy > 2.7).
[0115] DETAILED DESCRIPTION
[0116] In the present disclosure, the following terms will be employed, and are intended to be defined as indicated below.
[0117] Different types of chromosomal instability (CIN) such as mitotic defects, replication stress or impaired homologous recombination leave distinct patterns of copy number change in the genome. The inventors previously established a CIN signature framework that can identify these patterns and quantify the activity of each corresponding type of CIN in a genome (see Drews et al. 2022 and WO 2023 / 057392). In this work, the inventors used this framework to obtain models that are predictive of future oncogene amplification in a sample, using activity of signatures of chromosomal instability in the sample.
[0118] The present disclosure relates broadly to the characterisation of DNA samples, particularly tumour samples, in terms of their copy number profiles and signatures of chromosomal instability that are identified to be present in these samples.
[0119] A ‘‘copy number profile” refers to the quantification of the number of copies for each of a plurality of portions of a genomic sequence. In the context of the present disclosure, a copy number profile is preferably a genome-wide copy number profile. A copy number profile is obtained by analysing sequencing data. For example, a copy number profile may be obtained by obtaining sequence data from a sample of genomic DNA (or a DNA library derived therefrom, as explained further below, including a sample of DNA derived from genomic DNA by fragmentation, such as e.g. cell free DNA), and quantifying the number of copies per portion (e.g. per bin, where a bin is a region of fixed size typically between 20kb and 100kb, preferably between 30kb and 50kb, depending on the data available - for example for data comprising sequencing reads a predetermined minimum number of reads per bin is preferred such as e.g. 15 reads, which limits the bin size resolution that can be used) of the genomic sequence, as known in the art. The methods of the present disclosure are not limited in relation to the size of the bins used. Smaller or larger bins (i.e. higher or lower resolution) can be used (e.g. depending on the data available), and signatures and / or signature activities can be extracted at any such resolution. Further, signatures extracted at a first (lower) resolution can be mapped to signatures extracted at a second (higher) resolution, for example by identifying signatures at the second resolution that are most highly correlated (e.g. using Spearman’s Rho or any other correlation metric known in the art) with a signature at the first resolution, in terms of activity of the signatures across a plurality of samples (where samples can be individual cells - e.g. when using single cell sequencing data, or plurality of cells or tissues). A “tumour copy number profile” refers to a copy number profile that is associated with a tumour genome. A tumour copy number profile may be obtained by sequencing a sample comprising tumour genomic DNA, as explained further below. In embodiments, a tumour copy number profile is a copy number profile that has been obtained by sequencing a sample of genomic DNA derived from tumour cells. Thus, a copy number profile is typically in the form of an estimate, for each of a plurality of genomic segments, of the number of copies of the segment in the sample or a subset of the sample (e.g. in the case of samples comprising cells with different genotypes or genetic material derived therefrom, such as e.g. samples comprising a mixture of tumour and normal cells or tumour and normal DNA). In other words, for samples comprising tumour cells and non-tumour cells, a tumour copy number profile can be estimated for the tumour genome using methods known in the art. Unless specified otherwise, in the context of the present disclosure a copy number profile refers to a copy number profile for a tumour genome regardless of whether it is obtained from a sample also comprising genetic material from normal cells. Copy number profiles can be obtained by segmentation and copy number estimation from read count data using methods known in the art. For example, methods such as circular binary segmentation may be used for segmentation. Copy number fitting may be performed as described in the examples below. Alternatively, methods such as ASCAT (Raine et al. 2016), ichorCNA (Adalsteinsson et al. 2017) and others may be used to perform segmentation and copy number estimation (and also % tumour DNA estimation, in the case of ASCATt and ichorCNA). A copy number profile may be a genome-wide copy number profile. A genome-wide copy number profile may be obtained from a sample using WGS, sWGS or single cell WGS.
[0120] A “segment” in a copy number profile refers to a portion of a sequence represented in a copy number profile which is associated with a consistent absolute copy number. The consistent copy number is different from that associated with the sequence directly upstream (if such a sequence is present and associated with a copy number estimate) and the sequence directly downstream (if such a sequence is present and associated with a copy number) of said portion. In other words, a segment refers to a portion of sequence that has a copy number associated with it, where the copy number associated with the segment differs from the copy number associated with its immediate neighbouring segment(s). The copy number associated with a segment may differ from the copy number associated with its immediate neighbouring segment(s) because the segment(s) that surround the segment are associated with a different copy number, because the segment(s) that surround the segment are not associated with a copy number (e.g. because data for the segment(s) is missing, or of insufficient quality), or a combination of both (e.g. a segment may be surrounded by a segment that is associated with a different copy number on one side, and a segment that is not associated with a copy number on the other side). In other words, segments refer to the longest continuous portion of a copy number profile that are each associated with a single copy number. A segment may be associated with a set of coordinates, e.g. genomic coordinates, which define the boundaries of the segment. Each boundary may be associated with a copy number changepoint (also referred to as “changepoint”). The copy number of a segment may in practice be a copy number estimate obtained using a method for determining tumour copy number profiles, such as e.g. ASCAT [Van loo et al. ,2010] or Sequenza [Favero et al., 2015]. According to the present disclosure a “copy number event” refers to a single segment in a copy number profile. In particular, a copy number event may refer to a segment with a copy number that deviates from the expected copy number in a normal genome (e.g. a non-diploid segment). Copy number events / segments are characterised by a plurality of copy number features, further defined below. Thus, the terms “copy number event” and “segment” may be used interchangeably. Note that this is not necessarily the case in the prior art, such as e.g. in WO 2023 / 057392 where copy number features are determined for copy number events that encompass multiple segments. In embodiments of the present disclosure, all features are determined in relation to a specific copy number event / segment, as will be described further below.
[0121] Chromosomal instability (CIN) is the process of accumulating numerical and structural changes in DNA. A signature of chromosomal instability (CIN) (also referred to herein as copy number signature or “signature”) is a signature (typically in the form of a set of weights associated with respective categories of copy number features, also referred to as “components”) representing the genome-wide imprint of distinct putative mutational processes (where a “mutational process” as used herein refers to any process that can cause chromosomal instability). CIN signatures are described in Macintyre et al. 2018, Drews et al. (2022) and in WO 2023 / 057392, which are incorporated herein by reference. The presence or absence of a CIN signature in a sample may be determined by calculating the exposure (also referred to as “activity”) of the sample to the signature. A “CIN signature” is a set of weights associated with each of a plurality of components. The plurality of weights may sum to 1 and may each represent the probability that a mutational process associated with the signature generates a copy number event with a particular characteristic defined by the component. In other words, the plurality of weights of a signature may each represent the probability that a mutational process associated with the signature generates a particular type of copy number event. Each type of copy number event is referred to as a “component”. Each component is defined by a predetermined distribution for possible values of a copy number feature.
[0122] The term “copy number (CN) feature” refers to properties of copy number events observable in a copy number profile. Copy number features typically include a plurality of features selected from: segment size (also referred to as “segment length” typically expressed in number of bases), changepoint copy number (the absolute difference in copy number between a segment and an adjacent / neighbouring segment in the copy number profile), number of breakpoints per flanking region (also referred to as “breakpoints per side”), breakpoint count per x MB (the number of changepoints appearing in a sliding windows across the copy number profile, where x can be e.g. 10 Mb), breakpoint count per chromosome arm (the number of changepoints occurring per chromosome arm), number of segments with oscillating copy number (sometimes referred to as “length of segments with oscillating copy number”; number of continuous segments alternating between two copy number states, rounded to the nearest integer copy- number state; also referred to as length of chain of oscillating copy number states), and absolute copy number. The signatures in Macintyre et al. 2018 use the following features: segment size, changepoint, breakpoint count per 10 Mb, breakpoint count per chromosome arm, number of segments with oscillating copy number and absolute copy number. The signatures in Drews et al. 2022 and WO 2023 / 057392 use the following features: segment size, changepoint, breakpoint count per 10 Mb, breakpoint count per chromosome arm, and number of segments with oscillating copy number (i.e. they exclude the "absolute copy number” feature from MacIntyre et al. 2018). The signatures used in the examples of the present disclosure use the following features: segment size, changepoint and number of breakpoints per flanking region (i.e. they exclude the breakpoint count per 10 Mb, breakpoint count per chromosome arm, and number of segments with oscillating copy number of Drews et al. 2022 and replace them with a single feature that can be quantified for individual segments). Any of these schemes may be used for determining an exposure to a signature for a sample. Only the latter can be used for determining an association between a signature and a single copy number event. The latter is used to identify models that can be used to predict the likelihood of gene amplification from sample level exposures in embodiments of the disclosure, and will therefore be described in further detail below. However, note that the other sets of features and corresponding signatures may also be used in other embodiments, where a machine learning model is fitted to predict a likelihood of amplification of a gene from sample level exposures to signatures, using training data comprising a plurality of copy number profiles (or quantified features and / or quantified exposures) and gene amplification ground truth labels (i.e. known presence or absence of a gene amplification). Further, the signatures used in the Examples herein (see Table 3) are derived from those in Drews et al. 2022 (see Table 7 of WO 2023 / 057392) and therefore exposure to these signatures at the sample level is interchangeable with exposure to the corresponding signatures in Drews et al. 2022.
[0123] Thus, in embodiments of the present disclosure, copy number features include: segment size (also referred to as "segment length” typically expressed in number of bases), change-point copy number (the absolute difference in copy number between a segment and an adjacent / neighbouring segment in the copy number profile, which may be defined relative to the upstream segment, the downstream neighbouring segment, or a reference copy number state, typically 2 when analysing normally diploid genomes), and number of breakpoints per flanking region (also referred to as “breakpoints per side”, the number of breakpoints, i.e. changes in copy number / junction between individual segments in a flanking region of the segment, which can be upstream or downstream of the segment). The flanking region used to determine the “breakpoints per side” feature may have a predetermined length between 2 and 10 Megabases (Mb), preferably between 3 and 8Mb, or about 5 Mb. The number of breakpoints per flanking region may be determined for one or both flanking regions, and a summarised metric may be obtained when both flanking regions are used. The summarised metric may be the maximum or average number of breakpoints observed in the upstream and downstream flanking regions. In embodiments, the breakpoint per side feature for a segment is the maximum number of breakpoints observed in the upstream or downstream flanking region of the segment. The copy number features according to such embodiments of the disclosure do not include any feature that is not uniquely associated with individual events. In other words, the copy number features according to these embodiments do not include any copy number feature that is quantified over a genomic region that is not defined by reference to a segment. In particular, the copy number features according to these embodiments of the disclosure do not use any of: breakpoint count per x MB (the number of changepoints appearing in a sliding window of a predetermined size x across the copy number profile), breakpoint count per chromosome arm (the number of changepoints occurring per chromosome arm), and number of segments with oscillating copy number (sometimes referred to as “length of segments with oscillating copy number”; number of continuous segments alternating between two copy number states, rounded to the nearest integer copy-number state; also referred to as length of chain of oscillating copy number states). Such a set of copy number features can be used to associate signatures with individual copy number events, and are used in the Examples of the present disclosure to identify parameters of models used to predict the likelihood of gene amplification in a sample (by comparing association between respective signatures for gene amplification events and other events in a cohort of samples), as will be described further below.
[0124] In embodiments, the copy number (CN) features used according to the present disclosure do not include the segment copy number (the observed absolute copy number state of each segment, also referred to herein as “copy number” or “absolute copy number”).
[0125] Each CN feature is associated with one or more components. Each component is defined by a distribution. CN features are observed for individual copy number events. Typically, CN features are observed for each event on a genome-wide basis for a sample or collection of samples, these are used to quantify a metric for each component that is the probability that the observed value of the feature has been drawn from the distribution defined by the component. Summarised metrics can then be obtained across events in a sample for each component, such as by summing the probabilities obtained. These summarised metrics are in turn used to determine signature exposures for the sample. According to the present disclosure, signature exposures may be determined for individual events, as well as for a sample (i.e. for a whole copy number profile).
[0126] Methods for determining the exposure to a signature, from a copy number profile (i.e. at a sample level), as well as identifying signatures from a cohort of samples are known in the art (see e.g. Macintyre et al., 2018, Drews et al. (2022)). In particular, the determination of the exposure of a sample to one or more CN signatures (i.e. activity of the one or more signatures in the sample) may be performed by identifying the matrix E that satisfies C~PEwhere C is a mutational catalogue for one or more samples for which exposure is to be determined (also referred to as patient-by-component, PbC matrix), P is a signature matrix comprising the one or more CN signatures for which exposure is to be determined (also referred to as signature-by-component, SbC matrix), and E is an exposure matrix (also referred to as patient-by-signature, PbS). For example, exposure to a signature may be determined by performing matrix decomposition to identify the value (or vector of values) that satisfies the equation: PbC « PbS x SbC (or, in practice PbC = PbS x SbC + E, where s is a residual term to be minimised), where SbC is the row of the signature-by-component matrix that corresponds to a particular signature of a set of signatures, and PbS is the patient-by-signature value (or vector of values). Exposure to multiple copy number signatures can be calculated using the corresponding rows of the signature-by- component matrix. In embodiments, exposure (E) to a copy number signature / as described herein (from a set of n signatures, where n can be e.g. 16 or 17, where the signatures can include any or all of the signatures disclosed herein or corresponding signatures) is the value E, that satisfies the equation:
[0127] PbC = E x SbC (Equation 1) where: E is a vector of size n comprising coefficients Einwhere Ei is the exposure to signature / ; PbC is a vector of size c, each value representing the sum-of-posterior probabilities of each copy number event in the copy number profile belonging to a component C, where each component C is a distribution of values for a copy number feature; and SbC is a matrix of size c by n, each value representing the weight of a component C in a copy number signature / . This is equivalent to solving the minimisation problem min(||E*SbC-PbC||) with constraints on nonnegativity of E, with SbC and PbC being known. For example, linear combination decomposition as implemented in R package YAPSA (Hubschmann, D. et al. 2021) may be used for this purpose. A similar approach can be used to identify signatures using a cohort of samples and nonnegative matrix factorisation (see e.g. Drews et al. 2022) to identify both an exposure matrix E and a signature matrix SbC. The present disclosure provides, in addition to the determination of exposure of samples to signatures, the determination of associations between signatures and individual segments (in particular, copy number events) in a copy number profile. This will be explained further below.
[0128] In the context of the present disclosure, each component represents a type of copy number abnormality defined by an individual component of a mixture model fitted to the distribution of values obtained for each of 3 copy number features described herein over a cohort of samples from the Cancer Genome Atlas (TCGA). Note that any other cohort of cancer samples can be used for this purpose and the disclosure is not limited in any way in relation to the cohort of samples from which the components (and corresponding signatures) are extracted. This follows a process as described in Drews et al. 2022, except that in embodiments of the disclosure three features in Drews et al. 2022 (breakpoints per chromosome arm, breakpoints per 10 Mb, number of segments with oscillating copy number) were removed and replaced by a new feature as described herein (breakpoints per side). The components in Drews et al. 2022 (see Table 6 of WO 2023 / 057392) were used for the segment size and changepoint features, and new components were identified for the new feature. Copy number signatures (i.e. signatures of copy number alteration processes) are obtained using these, which capture the probability of copy number alteration processes causing copy number events that are distributed according to each of the copy number feature components. Thus, the term “exposure” (also referred to as “signature activity” or “activity”) in this context captures the strength of evidence for the presence of copy number alteration events attributable to the signature. As previously mentioned, in embodiments, the components and associated signatures in Drews et al. 2022 and WO 2023 / 057392 may be used (see Tables 6 and 7 of WO 2023 / 057392 which provide, respectively, component definitions and signature definitions identified using a cohort of samples comprising a plurality of tumour types including the TCGA cohort), or corresponding components and signatures identified using a different cohort of samples. The term “genome-wide” refers to data that covers a whole genome or substantial parts thereof such as e.g. one or more entire chromosomes. As the skilled person understands, sequencing data or a copy number profile derived therefrom may be referred to as “genome wide” even though it may not contain information (i.e. reads or a copy number estimate) for each and every single position in the genome. Indeed, even in the context of sWGS there may be regions of the genome that have no mapping reads due to e.g. low mappability, sequencing difficulty, or the sampling nature of next generation sequencing. Thus, data may be referred to as “genome-wide” if it contains data for at least 70%, at least 80% or at least 90%, preferably at least 90%, of one or more chromosomes or all chromosomes present in a reference genome.
[0129] A “sample” as used herein may be a cell or tissue sample, a biological fluid, an extract (e.g. a DNA extract obtained from a cell, tissue or fluid sample), from which genomic material can be obtained for genomic analysis, such as genomic sequencing (e.g. whole genome sequencing, or targeted sequencing such as whole exome sequencing or targeted panel sequencing). The sample may be a cell, tissue or biological fluid sample obtained from a subject (e.g. a biopsy). Such samples may be referred to as “subject samples”. The sample may be any sample comprising genomic DNA or cell free DNA. In particular, the sample may be a tumour sample, a biological fluid sample containing DNA or cells, a blood sample (including plasma or serum sample), a urine sample, a cervical smear, an ascites fluid sample, or a sample derived therefrom (e.g. after DNA purification). It has been found that urine, ascites fluid and cervical smears contains cells, and so may provide a suitable sample for use in accordance with the present invention. Other sample types suitable for use in accordance with the present invention include fine needle aspirates, lymph nodes samples (e.g. aspirates or biopsies), surgical margins, bone marrow or other tissue from a tumour microenvironment, where traces of tumour DNA may be found or expected to be found. The sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to genomic / transcriptomic analysis (e.g. frozen, fixed or subjected to one or more purification, enrichment or extraction steps). The sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells or genomic material derived therefrom, whether from a biological sample obtained from a subject, or from a sample obtained from e.g. a cell line. In embodiments, the sample is a sample obtained from a subject, such as a human subject. The sample is preferably from a mammalian (such as e.g. a mammalian cell sample or a sample from a mammalian subject, such as a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig), preferably from a human (such as e.g. a human cell sample or a sample from a human subject). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the sequence data acquisition (e.g. sequencing) location, and / or any computer-implemented method steps described herein may take place at a location remote from the sample collection location and / or remote from the sequence data acquisition (e.g. sequencing) location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider). A sample may be a tissue biopsy, such as a sample of fresh frozen tissue or a sample of formalin-fixed paraffin embedded (FFPE) tissue, or a sample comprising circulating tumour DNA (ctDNA) such as e.g. a sample of biological fluid (sometimes referred to as liquid biopsy). The samples used in methods of the present disclosure are typically samples comprising tumour cells (e.g. a tumour sample or sample comprising circulating tumour cells) or genetic material derived from tumour cells (such as e.g. cell free DNA, or DNA extracted from a sample comprising tumour cells - e.g. from a cell line or tissue sample - or circulating tumour cells). A sample may be a “mixed” sample comprising cells with different genotypes or genetic material derived therefrom. For example, a sample may be a sample comprising tumour cells and normal cells, or DNA derived therefrom, such as e.g. in the context of cell free DNA (cfDNA) samples comprising circulating tumour DNA (ctDNA). Copy number profiles may be determined separately for cells that have a genotype comprising somatic copy number alterations, in a sample comprising both cells that have the genotype comprising somatic copy number alterations (e.g. tumour cells) and cells that do not contain such alterations (e.g. normal cells), using methods known in the art. For example, sequence analysis processes that deconvolute tumour and germline genomes from copy number profiles include ASCAT (Raine et al., 2016), ABSOLUTE (Carter et al., 2012), or, in the context of cfDNA, ichorCNA (Adalsteinsson et aL, 2017). A “tumour sample” refers to a sample that contains tumour cells or genetic material derived therefrom. The tumour sample may be a cell or tissue sample (e.g. a biopsy) obtained directly from a tumour. A tumour sample may be a sample that comprises tumour cell or genetic material derived therefrom, that has not been obtained directly from a tumour. For example, a tumour sample may be a sample comprising circulating tumour cells or circulating tumour DNA. Thus, a tumour sample may also be a biological fluid (e.g. a liquid biopsy such as a blood, urine, or cerebrospinal fluid biopsy). A sample comprising a mixture of tumour cells and other cells (or material genetic derived therefrom) may be subject to one or more processing steps, whether prior to or subsequent to the acquisition of sequence data, in order to identify sequence data that is representative of the genetic material from the tumour. For example, a sample comprising cells may be subject to one or more cell purification steps which selectively enrich the sample for tumour cells. As another example, a sample of genetic material may be subject to one or more capture and / or size selection steps to selectively enrich the sample for tumour-derived genetic material. Protocols for doing this are known in the art. As another example, sequence data may be subject to one or more filtering steps (e.g. based on fragment length in the context of cell-free DNA samples) to enrich the data for information that relates to tumour-derived genetic material. Protocols for doing this are known in the art. In embodiments, the sample is a sample comprising tumour cells. Preferably, such a sample has a tumour purity (where tumour purity can be quantified as the proportion of cells in the sample that are tumour cells) of at least 30%, at least 35%, at least 40%, at least 45%, or at least 50%. Advantageously, the sample has a tumour purity of at least 40%. Without wishing to be bound by theory, it is believed that the copy number profiles generated from samples that have lower tumour purity may be less suitable as the signal corresponding to the tumour genome may be lost amongst the signal from the genomes of other cells. A “normal sample” (also referred to as “germline sample”) refers to a sample that contains non-tumour or non-modified cells or genetic material derived therefrom. A normal sample may be matched to a particular tumour or modified sample in the sense that it is obtained from the same biological source (subject or cell line) as the tumour or modified sample. A matched normal sample is a sample that is expected to be representative of the genome of a tumour sample in the absence of somatic abnormalities. This is typically a sample obtained from the same subject as the tumour sample. A normal sample may be used in the context of the present invention as a control to analyse a tumour sample. For example, a tumour sample and a (typically matched) normal sample may be analysed together in order to obtain a copy number profile for the tumour. Methods to obtain tumour copy number profile from a pair of normal and tumour samples are known in the art. For example, these include ASCAT [Van Loo et al. 2010] and Sequenza [Favero et al. 2015]. Such methods may provide a tumour copy number profile as well as an estimate of the purity (proportion of tumour cells) of the tumour sample. The purity estimate may be used to exclude tumour samples that have low purity and hence could lead to low quality copy number estimates.
[0130] The term “sequence data” refers to information that is indicative of the presence of genetic material in a sample that has a particular sequence. Such information may be obtained using sequencing technologies, such as e.g. next generation sequencing (NGS), for example whole exome sequencing (WES), whole genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing), or using array technologies, such as e.g. SNP arrays, or other molecular counting assays. In embodiments, the sequence data is obtained by DNA sequencing, and particularly next generation sequencing. In such embodiments, the sequence data comprises sequencing reads, or information derived therefrom such as the count of the number of sequencing reads that have a particular sequence. When non-digital technologies are used such as array technology, the sequence data may comprise a signal (e.g. an intensity value) that is indicative of the number of sequences in the sample that have a particular sequence, for example by comparison to an appropriate control. Information such as the count of sequencing reads that have a particular sequence can be derived from sequencing reads by mapping the sequence data to a reference sequence, for example a reference genome, using methods known in the art (such as e.g. Bowtie (Langmead et al., 2009) or BWA (Li and Durbin 2009)). This may result in aligned sequencing reads, for example in the form of a SAM or BAM file. Thus, sequence data may be associated with a particular genomic location (where the “genomic location” refers to a location in the reference genome to which the sequence data was mapped). Sequence data may be filtered as part of the methods described herein or may have been filtered prior to application of the methods described herein, for example to remove low quality reads or remove reads that map to regions that have “low mappability”. Methods for filtering low quality reads are known in the art, and include e.g. removing duplicate reads, removing supplementary alignments, removing reads with low sequencing quality scores (reads below Q20 or Q30), removing reads with low mapping quality score (e.g. MAPQ<37), etc. Methods to filter reads that map to regions that have low mappability are known in the art and include e.g. Umap (Karimzadeh et al. 2018), and GenMap (Pockrandt et al. 2020).
[0131] A composition as described herein may be a pharmaceutical composition which additionally comprises a pharmaceutically acceptable carrier, diluent or excipient. The pharmaceutical composition may optionally comprise one or more further pharmaceutically active polypeptides and / or compounds. Such a formulation may, for example, be in a form suitable for intravenous infusion.
[0132] As used herein "treatment" refers to reducing, alleviating or eliminating one or more symptoms of the disease which is being treated, relative to the symptoms prior to treatment. “Patient” as used herein in accordance with any aspect of the present invention is intended to be equivalent to “subject” and specifically includes both healthy individuals and individuals having a disease or disorder (e.g. a proliferative disorder such as a cancer). Preferably, the patient is a human patient. In some cases, the patient is a human patient who has been diagnosed with, is suspected of having or has been classified as at risk of developing, a cancer.
[0133] As used herein, the terms “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above described embodiments. For example, a computer system may comprise a central processing unit (CPU) and / or a graphic processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. A computer system may comprise a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other non-transitory computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.
[0134] As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.
[0135] Predicting gene copy number alteration
[0136] The present disclosure relates at least in part to the prediction of future gene amplification (e.g. oncogene amplification) or deletion (e.g. tumour suppressor gene deletion). In other words, the present disclosure provides methods of determining whether a tumour is likely to acquire one or more gene copy number alterations (i.e. amplification events or homozygous deletion events that are not currently present in the tumour genome but likely to emerge in the future). The gene copy number alterations according to the disclosure are preferably amplification events that encompass all or part of the coding region of an oncogene, or deletion events that encompass all or part of the coding region of a tumour suppressor gene. An oncogene may also be referred to herein as “driver”. An oncogene or driver gene is a gene that is known to cause or support abnormal cellular proliferation and / or resistance to cytotoxic therapy. A list of known oncogenes is provided in Tables 7a and 8a (column “oncogene”), and Tables 7b and 8b. The present disclosure encompasses prediction of the likelihood of acquisition of an amplification of any of the genes in Tables 7a, 7b and 8a, 8b, but is not limited to this specific list of genes as new oncogenes are still being discovered and the same methodology described herein can be applied to any gene with oncogenic potential. A tumour suppressor gene (TSG) deletion event may be referred to as a driver event, in that deletion of a TSG is known to cause or support abnormal cellular proliferation and / or resistance to cytotoxic therapy. A list of known tumour suppressor genes is provided in Tables 7c, 8c, 9c, and Figure 45.
[0137] An oncogene may be selected from: RPL22, MTOR, MACF1 , NRAS, SETDB1 , MDM4, H3F3A, KAT6B, FGFR2, HRAS, SF1 , RELA, CCND1 , NUMA1 , KDM5A, CCND2, CHD4, KRAS, ERBB3, CDK4, MDM2, PTPN11, NCOR2, MAX, HSP90AA1 , AKT1 , MAP2K1 , SIN3A, IDH2, PRKCB, CBFB, SUZ12, ERBB2, RARA, TOP2A, SPOP, KAT7, H3F3B, SETBP1, BCL2, GNA11 , KEAP1 , CACNA1 A, PRKACA, JAK3, PIK3R2, CCNE1, AKT2, GRIN2D, POLD1, PPP2R1A, ALK, SOS1, XPO1, PCBP1, NFE2L2, PMS1, SF3B1, IDH1, RQCD1, SRC, U2AF1, MAPK1, PLXNB2, RAF1, CTNNB1, RHOA, PIK3CB, MECOM, PIK3CA, FGFR3, KIT, IL7R, MSH3, NSD1, HIST1H3B, HIST1H4I, CCND3, HSP90AB1, EEF1A1, ESR1, UNCX, RAC1, EGFR, GTF2I, HGF, PIK3CG, BRAF, CUL1, FGFR1, SOX17, NCOA2, MYC, PSIP1, GNAQ, SET, ABL1, NUP214, ABCB1, ABL2, ACKR3, ACSL3, AFF1, ALB, ARHGEF12, BCL11A, BCL11B, BCL6, BCL7A, BCL9, BCL9L, BCLAF1, BCR, BIRC3, BRD4, CARD11, CBFA2T3, CBL, CD79B, CDH10, CDX2, CIITA, CLTC, CLTCL1 , CNOT3, COL1A1, CPEB3, CRNKL1, CSF3R, CXCR4, CYSLTR2, DCSTAMP, DHX9, EIF3E, EPAS1, EPS15, ERG, ESRRA, ETV5, ETV6, EWSR1, FAM135B, FAT2, FBLN1, FGD5, FGFR4, FLT3, FOXL2, FOXO1, FOXP1, GATA2, GLI1, GNAS, GRIN2A, HIP1, HNRNPA2B1, HOXC13, IL6ST, KDR, KEL, KIF5B, KIFC1, KLF4, LDB1, LIFR, LPP, LTB, MET, MLLT1, MLLT3, MYCN, MYD88, MYH11, MYH9, NKTR, NRG1, NTRK1, NTRK3, NXF1, PAK2, PDCD1LG2, PDE4DIP, PDGFRA, PDGFRB, PEG3, PLAG1, PLCB4, PLCG1, PPT2, PREX2, PRRX1, PTPRC, PTPRD, RALGDS, RAP1GDS1 , RBFOX1 , RBFOX2, RBM39, RET, RGL3, RNF213, RRAGC, RRAS2, SALL4, SATB1, Sep-09, SFMBT2, SMO, SOHLH2, SRGAP3, SRSF2, STAT6, SUSD2, TCF4, TCF7L2, TLL1, TRIM24, TRIM49C, TRIP11, U2AF2, UBE2D2, UGT2B17, USP6, VAV1, ZBTB20, ZBTB7B, ZNF148, ZNF521 , ZNF780A, ZNF814, A1CF, CALR, CCR4, CD28, CD79A, CDH17, CSF1R, CTNNA2, CTNND2, DDR2, GRM3, HIF1A, JUN, KCNJ5, KNSTRN, MACC1, MAP2K2, MITF, MPL, MUC16, MUC4, MYCL, MYOD1, NT5C2, REL, SIX2, SKI, SOX2, TSHR, ARAF, AR, BTK, DCAF12L2, ELF4, GATA1, MED12, MTCP1, NRK, P2RY8, SMC1A, TMSB4X. An oncogene may be selected depending on the tumour type (here indicated as prefix to the gene name): GBM- CCND2, GBM-CDK4, GBM-EGFR, GBM-GLI1, GBM-MDM2, GBM-MDM4, GBM-MET, GBM-MYCN, GBM-PDGFRA, GBM-SMO, GBM-SOX2, GBM-TRIM24, OV-ALB, OV-BRD4, OV-CCND3, OV- CCNE1, OV-FGFR3, OV-KDM5A, OV-KRAS, OV-MACF1, OV-MUC4, OV-MYC, OV-NCOA2, OV- PAK2, OV-PREX2, OV-SALL4, LUAD-CCND1, LUAD-CCND3, LUAD-CDK4, LUAD-DCSTAMP, LUAD-DDR2, LUAD-DHX9, LUAD-EGFR, LUAD-EIF3E, LUAD-FGFR1, LUAD-GNAS, LUAD- HSP90AB1, LUAD-IL7R, LUAD-KRAS, LUAD-LIFR, LUAD-MDM2, LUAD-MET, LUAD-MYC, LUAD- SALL4, LUAD-ZBTB7B, LUAD-ZNF521, LUSC-ALB, LUSC-BCL11A, LUSC-CCND1, LUSC-CCNE1, LUSC-CTNND2, LUSC-DCSTAMP, LUSC-EGFR, LUSC-EIF3E, LUSC-EPAS1 , LUSC-ERBB2, LUSC- IL7R, LUSC-KDM5A, LUSC-KRAS, LUSC-LIFR, LUSC-MDM2, LUSC-MECOM, LUSC-MYC, LUSC- MYCL, LUSC-NFE2L2, LUSC-PAK2, LUSC-PDCD1 LG2, LUSC-REL, LUSC-RRAS2, LUSC-SOX2, PRAD-AR, PRAD-DCSTAMP, PRAD-MYC, PRAD-NCOA2, PRAD-PREX2, UCEC-BCL6, UCEC- CACNA1A, UCEC-CCNE1, UCEC-ERBB2, UCEC-FGFR1, UCEC-KRAS, UCEC-MECOM, UCEC- MYC, UCEC-NCOA2, UCEC-SOX17, BLCA-ALB, BLCA-BIRC3, BLCA-CCND1 , BLCA-CCNE1 , BLCA- CD79B, BLCA-CLTC, BLCA-EGFR, BLCA-ERBB2, BLCA-EWSR1 , BLCA-KRAS, BLCA-MDM2, BLCA-MYC, BLCA-MYCL, BLCA-PDCD1 LG2, BLCA-PEG3, BLCA-PIK3CA, BLCA-SALL4, BLCA- S0X2, BLCA-SRGAP3, TGCT-CCND2, TGCT-CHD4, TGCT-KDM5A, TGCT-KDR, TGCT-KIT, TGCT- KRAS, TGCT-NC0A2, TGCT-PDGFRA, TGCT-PLAG1 , TGCT-PREX2, TGCT-S0X17, ESCA-CCNE1 , ESCA-EGFR, ESCA-ERBB2, ESCA-KRAS, ESCA-LPP, ESCA-MUC4, ESCA-PAK2, ESCA-SALL4, ESCA-S0X2, LIHC-CCND1 , LIHC-CCNE1 , LIHC-CDH17, LIHC-FAM135B, LIHC-MET, LIHC-MTCP1 , LIHC-MYC, LIHC-NC0A2, LIHC-SALL4, LIHC-ZBTB7B, SARC-CCND3, SARC-CCNE1 , SARC-CDK4, SARC-FGFR1 , SARC-JUN, SARC-LIFR, SARC-MDM2, SARC-MTCP1 , SARC-PRRX1 , SARC-SALL4, BRCA-ALB, BRCA-BRD4, BRCA-CCND1 , BRCA-CCND3, BRCA-CCNE1 , BRCA-CDH17, BRCA- CDK4, BRCA-EGFR, BRCA-ERBB2, BRCA-FGD5, BRCA-FGFR2, BRCA-IDH2, BRCA-KAT6B, BRCA-KDM5A, BRCA-MDM4, BRCA-MECOM, BRCA-MUC4, BRCA-MYC, BRCA-PAK2, BRCA- PEG3, BRCA-PIK3CA, BRCA-SALL4, BRCA-U2AF2, BRCA-ZBTB7B, BRCA-ZNF814, STAD-CCND1 , STAD-CCNE1 , STAD-EGFR, STAD-ERBB2, STAD-FOXO1 , STAD-GNAS, STAD-IL7R, STAD-KIF5B, STAD-KRAS, STAD-MET, STAD-MTCP1 , STAD-SALL4, STAD-SOHLH2, SKCM-BCL9, SKCM-BRAF, SKCM-CARD11 , SKCM-CCND1 , SKCM-CDH17, SKCM-CDK4, SKCM-CTNND2, SKCM-CUL1 , SKCM-DDR2, SKCM-FAM135B, SKCM-IDH2, SKCM-LIFR, SKCM-MDM4, SKCM-MITF, SKCM-MYC, SKCM-NTRK1 , SKCM-PDE4DIP, SKCM-PRRX1 , SKCM-RAC1 , SKCM-SALL4, SKCM-SMO, SKCM- TRIM24, SKCM-UNCX, SKCM-ZBTB7B, CESC-BIRC3, CESC-CIITA, CESC-ERBB2, CESC-GRIN2A, CESC-MECOM, CESC-PDCD1 LG2, CESC-SALL4, COAD-CCND2, COAD-CDX2, COAD-EGFR, COAD-ERBB2, COAD-FAM135B, COAD-GNAS, COAD-MTCP1 , COAD-NCOA2, COAD-PLAG1 , COAD-PREX2, COAD-RARA, COAD-SALL4, COAD-SOX17, READ-CDH17, READ-CDX2, READ- ERBB2, READ-FAM135B, READ-FLT3, READ-PLCG1 , READ-PREX2, READ-RBM39, READ-SALL4, READ-SRC, HNSC-BIRC3, HNSC-EGFR, HNSC-ERBB2, HNSC-FAM135B, HNSC-FGFR1 , HNSC- KDM5A, HNSC-MDM2, HNSC-MTCP1 , HNSC-PDCD1 LG2, HNSC-SOX2, LGG-CDK4, LGG-CUL1 , LGG-EGFR, LGG-KRAS, LGG-MYC, LGG-PDGFRA, LGG-TRIM24, UCS-MYC, UCS-SOX17, ACC- CDK4, ACC-MDM2.
[0138] A tumour suppressor gene may be selected from: EPHA2, ARID1A, THRAP3, MUTYH, CDKN2C, FUBP1 , RPL5, ELF3, IRF6, GATA3, PTEN, MGMT, FANCF, WT1 , DDB2, CTNND1 , FEN1 , INPPL1 , EED, PGR, ATM, CDKN1 B, ATF7IP, ARID2, KMT2D, ACVR1 B, TBX3, SETD1 B, CLIP1 , BRCA2, RB1 , ERCC5, CHD8, AJUBA, FOXA1 , ATXN3, TRAF3, BUB1 B, MGA, B2M, TCF12, RNF111 , BLM, AXIN1 , NTHL1 , CREBBP, ERCC4, PALB2, BRD7, CTCF, CDH1 , FANCA, GPS2, TP53, CHD3, MAP2K4, NCOR1 , NF1 , CDK12, SMARCE1 , KRT222, BRCA1 , KANSL1 , RNF43, BRIP1 , AXIN2, SOX9, ZNF750, SMAD2, SMAD4, STK11 , DAZAP1 , DOT1 L, SMARCA4, KMT2B, CIC, ERCC2, ARHGAP35, DNMT3A, ASXL2, ZFP36L2, MSH2, MSH6, RANBP2, ERCC3, ACVR2A, CASP8, BARD1 , CUL3, PTMA, FOXA2, ASXL1 , SCAF4, SMARCB1 , CHEK2, NF2, EP300, FANCD2, VHL, XPC, TGFBR2, MLH1 , SETD2, BAP1 , PBRM1 , STAG1 , ATR, TBL1XR1 , FBXW7, IRF2, FAT1 , NIPBL, MAP3K1 , PIK3R1 , RAD17, RASA1 , APC, FANCE, CDKN1A, ARID1 B, PMS2, EZH2, KMT2C, WRN, NBN, RAD21 , CDKN2A, FANCG, FANCC, PTCH1 , XPA, TSC1 , ARHGAP5, ARHGEF10L, ATG7, BMPR1A, BMPR2, CAMTA1 , CASZ1 , CCDC6, CCR7, CD58, CDC73, CHD2, CMTR2, CYLD, CYP2C8, DAXX, DCC, DGCR8, DICER1 , DIS3, DNAJB1 , DROSHA, DUSP16, ERBB4, ETV4, EXT2, FAM46C, FAS, FAT4, FBXO11 , FH, FLCN, FN1 , FOXD4L1 , GMPS, GNA13, HNF1A, HOXD13, HTRA2, ID3, IFNAR1 , IFNGR1 , ING1 , IRF1 , KLHL6, KMT2A, LATS1 , LATS2, LOX, LRP1 B, LZTR1 , MAF, MALT1 , MAML2, MAP2K7, MB21 D2, MEF2B, MEN1 , NFKBIA, NIN, NPEPPS, NPRL2, NRP1 , NT5C3A, PCDH17, PCMTD1 , PIM1 , PML, PPM1 D, PPP3CA, PPP6C, PRDM1 , PRF1 , PRKAR1A, PRKCD, PRR14, PTPN13, PTPN14, PTPN6, PTPRK, PTPRT, OKI, RABEP1 , RASA2, RBM15, RECQL4, RGS7, RNF6, RXRA, SDC4, SIRPA, SIX1 , SLC45A3, SMAD3, SOCS1 , SOX21 , TCIRG1 , TEC, TET2, TFAP4, TFG, TGIF1 , TNFAIP3, TNFRSF14, TOP1 , TSC2, USP44, ZCRB1 , ZEB1 , ZFHX3, ZFP36L1 , ZNF165, ZNF429, ZNF93, ZNRF3, BAX, BAZ1A, BTG2, CASP3, CASP9, CBLB, CCNC, CNTNAP2, CSMD3, DNM2, ETNK1 , FBLN2, GPC5, IGF2BP2, KLF6, LARP4B, LEPROTL1 , N4BP2, PHOX2B, POLG, RFWD3, ROBO2, SBDS, SDHA, SDHAF2, SDHB, SDHC, SDHD, SFRP4, SUFU, TMEM127, TRAF7, EIF1AX, BCOR, USP9X, DDX3X, KDM6A, KDM5C, AMER1 , ATRX, STAG2, SMARCA1 , BCORL1 , IRAKI , IRS4, LPAR4, NONO, PHF6, RBM10, RPL10, RPS6KA3, UBE2A, WAS, WDR45, ZFX, ZRSR2, ZXDB, ATP2B3, ZMYM3. A tumour suppressor gene may be selected depending on the tumour type (here indicated as prefix to the gene name): GBM-CDKN2A, GBM-CDKN2C, GBM-PTEN, OV- CDKN2A, OV-MAP2K4, OV-NF1 , OV-PTEN, OV-RB1 , LUAD-CDKN2A, LUSC-CDKN2A, LUSC-PTEN, PRAD-LRP1 B, PRAD-MAP3K1 , PRAD-PTEN, UCEC-PTEN, BLCA-CDKN2A, BLCA-PTEN, ESCA- CDKN2A, PAAD-CDKN2A, LIHC-CDKN2A, KIRP-CDKN2A, SARC-CDKN2A, SARC-PTEN, BRCA- BRCA2, BRCA-CDKN2A, BRCA-MAP2K4, BRCA-PTEN, BRCA-RB1 , MESO-CDKN2A, STAD- CDKN2A, STAD-MAP2K4, STAD-PTEN, STAD-SMAD4, SKCM-CDKN2A, SKCM-PTEN, COAD- MAP2K4, COAD-PTEN, COAD-SMAD3, COAD-SMAD4, HNSC-CDKN2A, LGG-CDKN2A, DLBC- CDKN2A, ACC-ZNRF3.
[0139] An amplification, also referred to herein as “amplification event”, is the presence of an abnormal elevated copy number at a genomic region. In embodiments, any region with a copy number higher than 2 (diploid genome) may be considered an amplification event. Preferably, an amplification event refers to a genomic region with at least 6 copies. In embodiments, any region with an absolute copy number value >6 may be considered an amplification event. In embodiments, any region with a copy number value relative to tumour ploidy >4 may be considered an amplification event. For example, when the tumour ploidy is 2 (i.e. no whole genome duplication), any region with at least 6 copies may be considered amplified, whereas when the tumour ploidy is 4 (i.e. the tumour genome has undergone a whole genome duplication), any region with at least 8 copies may be considered amplified. A gene amplification may refer to the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number >=8, if present in an autosome, or >=6 if present in a sex chromosome.
[0140] Figure 1A is a flow diagram showing, in schematic form, a method of predicting whether a tumour is likely to acquire a gene copy number alteration using methods described herein.
[0141] At optional step 10, a DNA sample is obtained from a tumour of a subject. Optionally, a matched normal sample may also be obtained from the subject. At optional step 12, sequence data is obtained from the tumour (and optionally the matched normal) DNA sample(s), for example using sequencing or array technologies. At step 14, a tumour copy number profile for the sample is obtained, for example by analysing the sequence data obtained at step 12 or by receiving the copy number profile from a user, database etc. Thus, steps 10-12 are optional because methods described herein may start from a previously obtained tumour copy number profile for the sample. The data may be bulk sequence data (from a bulk DNA sequencing technology or an array technology) or single cell sequencing data (from a single cell DNA sequencing technology). When single cell sequencing data is used, the process described by reference to steps 12-20 may be performed individually for each of a plurality of single cells. Step 14 of obtaining a tumour CN profile may comprise one or more of: aligning a plurality of sequence reads to a reference genome, determining that the copy number profile comprises a number of non-diploid segments above a predetermined threshold (e.g. at least 15 or at least 20 events in the copy number profile), determining a copy number for each of a plurality of genomic bins of a predetermined size, excluding genomic bins comprising a number of mapped reads below a predetermined threshold, and / or performing absolute copy number fitting. The predetermined bin size may be at most 10Okb, at most 50kb, or about 30kb. The predetermined threshold on number of mapped reads may be 15 reads. A tumour copy number profile may identify one or more genomic segments, wherein each segment is associated with a copy number that differs from the copy number of the neighbouring segment(s). A set of copy number features may be quantified for each of the segments, or one or more of the segments identified in the copy number profile (e.g. all segments identified on one chromosome or sets of chromosomes, etc.). The methods are not limited in this regard and even single events (segments) may be analysed productively. The terms “segment” and “genomic segment” are used interchangeably.
[0142] Step 16 comprises determining, for each of one or more signatures of chromosomal instability, a metric indicative of the presence of the signature in the tumour copy number profile. As explained above, a signature of chromosomal instability comprises a set of weights for each of a plurality of components each associated with a copy number feature of a set of copy number features quantified for a plurality of segments in the copy number profile, wherein the weights are indicative of the probability that a mutational process associated with the signature generates a segment having characteristics defined by the component. For example, in the illustrated embodiment a sample level exposure to one or more signatures of chromosomal instability is determined (also referred to as activity of a signature in the sample). Step 16 may comprise step 16A of quantifying a set of copy number features, for each of one or more segments in the copy number profile. This step may use unrounded copy number segments. This step may comprise collapsing and merging near diploid segments to a diploid state, wherein near diploid segments are segments that have a copy number within a predetermined distance from 2. The predetermined distance may be 0.1. Thus, quantifying the set of copy number features may comprise assigning a copy number of 2 (i.e. collapsing) to any segment that has a copy number within predetermined boundaries (such as e.g. above 1.9 and below 2.1), and merging any contiguous segments that have the same copy number as a result of the assigning. This may advantageously avoid including signal from segments that are likely normal diploid segment and which could act as noise in the process. Unrounded copy number segments may be segments whose copy number has not been rounded to the near integer. Thus, the method preferably uses copy number profiles that have not been rounded and / or wherein the only rounding that is performed relates to the collapsing or near diploid segments. This uses more of the information present in copy number profiles (in terms of number of segments that would otherwise be merged but also in terms of the copy number information that is contained in said profile for each segment). Rounding of copy numbers for segments may be performed in order to remove noise in a copy number profile, at the cost of loss of information. It has been previously shown that the additional noise that is associated with the use of unrounded copy number segments could be compensated at least in part, with minimal loss of information relevant to CIN by merging segments with minor deviations from the “normal” copy number state, and the present inventors have shown that this also applies to the new methods described herein. The features may not be quantified for diploid segments. In other words, the methods of the present disclosure may comprise quantifying a plurality of copy number features for one or more non-diploid segment, such as e.g. each non-diploid segment in a copy number profile.
[0143] Step 16 may further comprise step 16B, where a metric indicative of association between one or more signatures of chromosomal instability and each segment is determined using the quantified features for the segment. A signature of chromosomal instability comprises a set of weights for each of a plurality of components each associated with a copy number feature of the set of copy number features, wherein the weights are indicative of the probability that a mutational process associated with the signature generates a segment having characteristics defined by the component. Thus, step 16B may comprise obtaining, for each segment, a value for each of the plurality of components for the segment. Each component defines characteristics of a segment as a predetermined distribution of values of a copy number feature. Determining a metric indicative of association between a signature of chromosomal instability and a segment may comprise determining the probability of the copy number features of the segment belonging to each of a set of predetermined distribution associated with respective components of the plurality of components. Thus, the probability of the feature value from a segment belonging to the predetermined distribution associated with the component may be determined. This may be performed for each component, thereby obtaining an “Event-by-component” probabilities vector for each segment, or an event-by-component probability matrix comprising respective vectors for each of a plurality of segments from the same copy number profile. These probabilities may be referred to as “posterior probabilities”. Thus, step 16B may comprise quantifying the posterior probabilities of each feature value belonging to each of a set of predetermined distributions, also referred to as “components”. The components may have been identified empirically by quantifying the copy number features in a plurality of tumour samples. Each copy number feature may be associated with a plurality of components. A posterior probability may be obtained for each component and for each event. Further, a summarised measure across the copy number profile may be obtained for each such component, such as e.g. by summing the posterior probabilities obtained for each event for the component. The use of empirically identified components means that the features are truly reflective of biological processes rather than arbitrary categories that while convenient to manipulate, may not be biologically relevant. For example, a set of predetermined distributions may be identified by applying mixture modelling on a data set comprises the quantified set of copy number features for a plurality of tumour samples. This enables the identification of the true states of the copy number features present in tumours, and hence the quantification of the evidence for the presence of these states in a new sample to be analysed. The use of posterior probabilities of each feature value belonging to each of a set of predetermined distributions advantageously enables the method to deal with uncertainty in the assignment of a segment to a state (i.e. uncertainty in whether the segment provides evidence for the presence of a particular category of CIN events characterised by one of the distributions for a copy number feature). This means that the method is by design able to deal with inevitable noise in the data, resulting in a more accurate characterisation of the sample. The set of predetermined distributions for a copy number feature may be a set of Gaussian distributions for any feature that is quasi-continuous, such as e.g. the segment size or changepoint feature. The set of predetermined distributions may be a set of Poisson distributions for any count feature, such as e.g. the breakpoint per side feature. A set of predetermined distributions may have been obtained using a mixture modelling technique. A quasi- continuous feature may be any feature that is not a count feature, such as e.g. segment size and copy number changepoint. For example, the segment size feature may be quantified for an event by obtaining posterior probabilities (or for a copy number profile as sum-of-posterior probabilities summed across copy number events in the copy number profile) for each of a plurality of Gaussian distributions, such as e.g. between 20 and 25 Gaussian distributions (e.g. 22 Gaussian distributions). As another example, the copy number changepoint feature may be quantified by obtaining posterior probabilities (or for a copy number profile as sum-of-posterior probabilities summed across copy number events in the copy number profile) for each of a plurality of Gaussian distributions, such as e.g. between 5 and 15 Gaussian distributions (e.g. 10 Gaussian distributions). As another example, the breakpoint per side feature may be quantified by obtaining posterior probabilities (or for a copy number profile as sum-of- posterior probabilities summed across copy number events in the copy number profile) for each of a plurality of Poisson distributions, such as e.g. between 5 and 15 Poisson distributions (e.g. 9 Poisson distributions). The precise number of distributions used may depend at least in part on the number and diversity of samples used to obtain the signatures. In the examples provided herein, the inventors used a very large number of samples from a wide variety of tumour types, enabling them to provide a nuanced picture of the behaviours of copy number features in cancer, resulting in a higher number of components identified than e.g. in Macintyre et al., 2018. The parameters of each of the predetermined distributions (such as e.g. mean and variance for a Gaussian distribution, A for a Poisson distribution) may have been determined as part of the process of obtaining the signatures of chromosomal instability. The parameters of each of the predetermined distributions (such as e.g. mean and variance for a Gaussian distribution, A for a Poisson distribution) may have been determined by fitting mixture models to the quantified set of copy number features in the plurality of tumour samples from which the signatures of chromosomal instability have been obtained. In the examples below, these parameters were obtained by analysis of over 6000 samples from the Cancer Genome Atlas (TCGA). As the skilled person understands, slightly different parameters may be obtained when using different data. The parameters of each of a set of Gaussian distributions may have been obtained using a Variational Bayes Gaussian mixture model. In particular, the parameters of Gaussian distributions (for example for features such as segment size and changepoint) may have been obtained by fitting Dirichlet-Process Gaussian mixture models using variational inference. The parameters of each of a set of Poisson distributions (for examples the breakpoint per side feature) may have been obtained using a Poisson mixture model and / or using predetermined parameters. The parameters of the components used in the examples described herein are provided in Table 1. Thus, in embodiments, step 16B uses the components defined in Table 1 or corresponding components defined using a different cohort of samples. Corresponding components refer to components defined by distributions that correspond to the distributions of the components in Table 1 , but where the exact parameters of each distribution may vary. Alternatively, the signatures in Drews et al. 2022 and WO 2023 / 057392 may be used, with components defined in Table 6 of WO 2023 / 057392 and signatures defined in Table 7 of WO 2023 / 057392. The component parameters for the changepoint and segment size features in Table 1 are as in Drews et al. 2022 and WO 2023 / 057392. The breakpoints per side feature is new to the present disclosure. In embodiments, the set of signatures used herein include one or more of the signatures in Drews et al. 2022 (e.g. signatures selected from CX1 to CX17, optionally excluding CX6). Aetiologies associated with these signatures are provided in Table 2. The full definitions of these signatures are provided in Drews et al. 2022 and WO 2023 / 057392 (see Table 7 of WO 2023 / 057392, reproduced herein as Table 12, with components specified in Table 11). In embodiments, signatures corresponding to those in WO 2023 / 057392 (see Table 7 of WO 2023 / 057392) in the new feature space described herein are obtained by applying linear combination decomposition using a cohort of samples with known exposures to the signatures. In other words, signature definitions can be identified as those that solve the minimisation problem min(||E*SbC-PbC||) with constraints on nonnegativity of SbC, with E and PbC being known. For example, linear combination decomposition as implemented in R package YAPSA (Hubschmann, D. et al. 2021) may be used for this purpose. Corresponding signature definitions according to the present disclosure are provided in Table 3. As mentioned above, the process of defining the signatures in Drews et al. 2022 in the new feature space may be performed using one or more or all of CX1 to CX17, excluding CX6. In embodiments, a subset of the signatures in Drews et al. 2022 and / or their equivalent in the feature space described herein (Table 3) are used. For example, a subset of the signatures that is found across a plurality of technologies (i.e. a plurality of methods for obtaining sequence data suitable for the generation of a copy number profile) may be used. These may be referred to as “universal signatures”. For example, CX1 -5, CX8-10, and CX13 were found by the present inventors to show robust signals across technologies (including at least array data such as SNP6 arrays and shallow whole genome sequencing data, sWGS), and were classified as universal signatures. Other signatures in Table 3 were classified as SNP6-specific signatures (CX7, CX1 1 -12, and CX14-17). CX14 was not considered universal because its signal was captured by CX1 in sWGS data.
[0144] Step 16 may further comprise step 16C where a sample level exposure to one or more of the signatures is determined (i.e. the activity of the signatures in the tumour copy number profile is determined). This may comprise obtaining, for each component, the sum of probabilities associated with the respective component for each of a plurality of segments in the copy number profile; and determining a linear combination of one or more signatures comprising the signature that satisfies the equation: PbCaE x SbC (Equation 1) where E is a vector of size n comprising coefficients Einwhere Ei is the exposure to signature i; PbC is a vector of size c, each element in the vector representing a sum of probabilities associated with a respective component; and SbC is a matrix of size c by n, each value representing the weight of a component in a signature / . The resulting exposures may further be normalised by sample such that they sum to 1 (i.e. dividing the exposure for each signature in a sample by the sum of exposures for all signatures in the sample). At step 18, the metrics obtained at step 16 (e.g. activity of the one or more signatures in the profile) are used as input to a model trained to take these metrics as inputs and produce as output a score indicative of the likelihood that the tumour will acquire a gene amplification. The model may be a linear model that obtains a score as a weighted combination of the metrics obtained at step 16 (e.g. signature activities). The weights may have been determined as will be described below by reference to Figure 1 B. At step 20, the sample / tumour may be classified between a first class that has a high likelihood of acquiring the gene copy number alteration and a second class that has a low likelihood of acquiring the gene copy number alteration, wherein the tumour is classified in the first class when the output of the model is above a predetermined threshold. At step 22, the results of any of the preceding steps may be provided to a user, for example through a user interface, in the form of a report, etc., or to a computing device or memory.
[0145] Identifying signatures associated with oncogene amplification or tumour suppressor gene deletion
[0146] The present disclosure also provides methods of identifying signatures associated with oncogene amplification or tumour suppressor gene deletion. Such methods can be used to define models that are used to determine the likelihood of future amplification / deletion in a sample as described by reference to Figure 1 A. In embodiments, these make use of methods of identifying signatures associated with individual copy number events (such as oncogene amplifications or tumour suppressor gene deletion). A method of identifying signatures of chromosomal instability associated with individual copy number alterations in a sample will now be described by reference to Figure 1B.
[0147] In particular, Figure 1 B is a flow diagram showing, in schematic form, a method of identifying a model for predicting the likelihood of acquisition of a gene amplification event or gene deletion event according to embodiments of the disclosure. At optional step 1 10, a plurality of reference tumour samples are obtained which together form a reference cohort. At optional step 1 12, sequence data is obtained from the samples, for example using sequencing or array technologies. At step 1 14, a tumour copy number profile is obtained for each sample in the reference cohort, for example by analysing the sequence data obtained at step 112 or by receiving the copy number profile from a user, database etc. Thus, steps I Q- 12 are optional because methods described herein may start from a previously obtained tumour copy number profile for each sample in the cohort of samples or from previously acquired sequence data for each sample. Steps 110-114 may have any of the features described above by reference to steps I Q- 14 of Figure 1A.
[0148] At step 116, a set of copy number features are quantified, for each of one or more segments in the copy number profiles. This step may have any of the features described above by reference to step 16A of Figure 1 A. The set of copy number features consists of copy number features that are associated with respective segments (i.e. copy number features that are quantified in relation to each individual segment). In other words, the set of copy number features may not include any copy number feature that is defined for a section of the genome that is not a segment or a region defined by reference to a segment. As such, each copy number feature is defined for a specific segment, and quantifies a property of the segment and / or a region defined by reference to a segment (such as e.g. a flanking region of a predetermined length, a region of a predetermined length centred on the segment, etc.). The set of copy number features quantified for each segment at step 1 16 may comprise a segment size feature, a copy number changepoint feature, and a breakpoint per side feature (breakpoint count in one or both flanking regions of a predetermined length). The copy number changepoint feature may be quantified for the first segment of any chromosome by as the copy number change relative to the diploid state (i.e. 2 is subtracted from the copy number value of the segment). In other words, the changepoint feature may be quantified by subtracting 2 from the absolute copy number of the first segment of any chromosome if said segment is not a normal segment. This assumes that the (non-existing) previous segment was a normal segment (copy number 2) if the segment is not a normal segment. The copynumber changepoint feature may be quantified for a current segment relative to the upstream or downstream segment of the current segment. In particular, the copy-number changepoint may be quantified relative to the upstream (left) neighbouring segment when the changepoint value relative to the upstream segment is at or higher than a predetermined threshold (e.g. -3 or more, such as e.g. -3, -2, -1 , 1 , 2, 3, etc.). The copy-number changepoint may be quantified relative to the downstream (right) neighbouring segment when the changepoint value relative to the upstream segment is lower than the predetermined threshold (e.g. less than -3, such as e.g. -4, -5, etc.). When the copy number changepoint relative to the downstream neighbouring segment is also lower than -3, the copy number changepoint may be calculated by removing two from the copy number value of the current segment (i.e. calculating the difference between (current segment copy number) and (diploid state)). In any embodiment, the copy number changepoint may be quantified as an absolute value. In other words, a relative copy number changepoint between a current segment and a neighbouring segment may be calculated and the absolute value of this may then be taken as the copy number changepoint for the current segment. This approach differs from the approach in Drews et al. 2022 in which the changepoint captured the difference in absolute copy number compared to the left neighbouring segment without taking into account whether the relative change was a positive or a negative value. The approach proposed herein advantageously avoids linking an event next to a focal amplification to a replication stress signature, as was the case in the previous approach since segments with low copy number value next to a focal amplification were defined by having high copy number change. In other words, the left segment may be used by default unless it is determined that the left segment is likely a segment with focal amplification. In this case, the right segment may be used unless it is determined that the right segment is also likely a segment with focal amplification. In such cases the copy number changepoint may be calculated relative to a diploid state.
[0149] At step 118, a metric indicative of association between one or more signatures of chromosomal instability and each segment is determined using the quantified features for the segment. Step 1 18 may have any of the features described above by reference to step 16B of Figure 1A. Step 1 18 (determining a metric indicative of association between a signature of chromosomal instability and a segment using the quantified features for the segment) may comprise multiplying probabilities associated with each of the plurality of components for the segment by the corresponding weight in the signature, and summing up the resulting values. As explained above, the probability associated with a component for the segment are typically the probability of a segment with the quantified feature having been drawn from the distribution associated with the component. The resulting summed up values may be referred to as an "exposure” of the event to the respective signatures, or “signature score”.
[0150] At step 120, a sample level exposure to one or more of the signatures may be determined. This may be performed as explained above by reference to step 16C of Figure 1A. At step 122, a sample- corrected signature score (also referred to herein as activity corrected signature score) is obtained for each segment, by multiplying the signature score obtained at step 1 18 by the sample level exposure obtained at step 120. Instead or in addition to this, step 122 may further comprise normalising the signature exposures for a segment by dividing each exposure (optionally sample-corrected) by the sum of exposures (optionally sample-corrected) across the plurality of signatures for the segment. Note that at this point, the results of steps 1 18 and optionally 120 may be used to determine the activity of signatures (and correspondingly, the activity of CIN causing processes associated with the signatures). This is not part of the process of generating a model as described herein. This may comprise the step of assigning a signature of chromosomal instability to a segment of the one or more segments as the signature of the one or more signatures of chromosomal instability with the highest value of the metric for the segment. For example, the metric for a signature may be the exposure of the segment to the signature and a segment may be associated with the signature with the highest exposure for the segment. Alternatively, the metric for a signature may be a distance between the vector of probabilities associated with the plurality of components for the segment and the corresponding weights in the signature, and a segment may be associated with the signature with the lowest value of the metric for the segment (when the distance is a true distance metric such as Euclidian distance) or the highest value of the of the metric for the segment (when the distance is a similarity metric such as cosine similarity). Alternatively when the metric for a signature is a probability that the segment is associated with the signature (exposure), an event may be associated with a plurality of signatures, each with respective probabilities. As explained above, the one or more signatures of chromosomal instability may be selected from those defined in Table 3 or corresponding signatures obtained by: (i) quantifying the set of copy number features in a plurality of tumour samples, and (ii) identifying one or more signatures likely to result in the copy number profiles of the plurality of tumour samples by nonnegative matrix factorisation. Examples of signature and their associations are provided in Table 2. Thus, determining the presence or absence of a signature may be equivalent to determining whether a corresponding CIN process is or has been active in the sample, or in other words determining whether the copy number alterations in the copy number profile have been generated at least in part by the CIN process. Thus, the one or more signatures comprise one or more signatures selected from: one or more signatures associated with chromosome missegregation, optionally chromosome missegregation via defective mitosis and / or telomere dysfunction, one or more signatures associated with impaired homologous recombination, one or more signatures associated with tolerance of whole genome duplication, one or more signatures associated with impaired non-homologous end joining, one or more signatures associated with replication stress, and one or more signature associated with impaired DNA damage sensing. Further, the presence / absence of a signature or corresponding process may also be determined at the sample level, by comparing the exposures obtained at step 120 to a respective predetermined threshold. The predetermined thresholds may have been obtained by performing simulations as the threshold such that a predetermined proportion (e.g. 95%) of simulated samples in which a signature is truly not present (but where the signature exposure may nonetheless be non-zero due to the presence of noise in the simulated data) are below the threshold. Examples thresholds for the signatures in Table 3 are provided in Table 4. Other thresholds may be used depending on the desired stringency in identifying the presence of a signature.
[0151] Turning back to a method of providing a model as described herein, at step 124, a weight may be obtained for each signature and for one or more gene amplifications (individually or collectively). In embodiments, weights (and therefore models comprising said weights) are obtained for individual genes. This will be described in further detail below, although the same principles can be applied when obtaining a model for a plurality of genes jointly. The weights may be determined in a tumour type specific manner, or in a tumour type-agnostic manner. When the weights of the model are determined in a tumour type-specific manner, a weight may be obtained for each signature as a ratio of (i) an averaged signature score of all amplified segments spanning the gene (when predicting amplification of the gene) or homozygously deleted segments (when predicting deletion of the gene) in samples of the specific tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort. As explained above, the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment, such as the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature (obtained at step 118), the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected (i.e. multiplied by the value obtained at step 120) and / or normalised version of said sum (obtained at step 122). As explained above, an activity corrected version of said sum may be the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures. Alternatively, the weights of the model may be determined in a tumour type-agnostic manner for each signature, a ratio of (i) a weighted average signature score of all amplified segments spanning the gene in samples of a specific tumour type in the reference cohort, wherein the average signature scores in the specific tumour type are weighted by the normalised empirical frequency of the gene amplification in the respective tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort. All tumour types present in the reference cohort may be used. Alternatively, only tumour types where the gene is recurrently amplified / deleted may be considered. The former may be particularly suitable for primary and / or untreated tumours. The latter may be particularly suitable for advanced and / or treated tumours.
[0152] Determining the weights at step 124 may comprise one or more of determining a fitness effect term, determining a local mutational process term, and a DNA topology term, as described herein.
[0153] At step 126, a predetermined threshold is identified applicable to the output of the model comprising the weights obtained at step 124. The predetermined threshold may be such that samples with a score from the model above the threshold are likely to acquire the gene amplification, and samples with a score from the model below the threshold are likely to acquire the gene amplification. Thus, the score and predetermined thresholds may be used to obtain a binary classifier that can classify samples between a first class that has a high likelihood of acquiring the gene amplification and a second class that has a low likelihood of acquiring the gene amplification, wherein the tumour is classified in the first class when the output of the model is above a predetermined threshold. The threshold may be identified using prior knowledge. For example, when there is a known expected % of samples with the gene amplification in the cohort or a subset of the cohort (e.g. a specific tumour type), the threshold may be selected such that a % of samples corresponding to the expected % of samples in the cohort or subset are classified in the first class (i.e. have scores above the threshold). Alternatively, the threshold may be selected using one or more criteria applying to the sensitivity, selectivity or Youden’s index of the binary classification (where a “positive” sample is one classified in the first class). For example, a threshold that maximises the Youden’s index may be used. It is advantageous (although not necessary) for the reference cohort to be selected to match one or more characteristics of a patient sample to be analysed. For example, the reference cohort may comprise or consist of samples that are one or more of: the same tumour type, the tumour grade / stage, undergone the same treatment, have the same driver mutations, have been analysed with the same sequencing technology, as the sample to be analysed. At optional step 128, the results of one or more of the preceding steps (e.g. one or more of the tumour copy number profile, quantified features for segments, quantified metrics for segments, parameters of the components used, sample exposure, weights, etc) are provided to a user, for example through a user interface, or to a computing system memory or database. In particular, the weights identified at step 124 may be provided for use in a method as described in relation to Figure 1A.
[0154] Table 1. Definition of the 43 mixture components used. Feature=One of the five fundamental features of copy number aberrations; Mean = If SD is present, then this is the mean of the Gaussian component, otherwise mean of the Poisson component; SD = If present, standard deviation of the Gaussian component; Number = Number of component used as identifiers.
[0155]
[0156] Table 2. Aetiologies of CIN signatures in Drews et al. 2022. Sign=signature. Note CX6 is not used in the present work.
[0157] Table 3. Signature by component definition matrix. CIN signatures are defined in the new feature space, which is composed by 41 components from 3 copy number features. Components definitions are provided in Table 1 (in the same order as below, i.e. segsize1=first segment size component listed, etc.).
[0158] Table 3 (continued). Signature by component definition matrix. CIN signatures are defined in the new feature space, which is composed by 41 components from 3 copy number features. Components definitions are provided in Table 1 (in the same order as below, i.e. segsize1=first segment size component listed, etc.).
[0159] Table 3 (continued). Signature by component definition matrix. CIN signatures are defined in the new feature space, which is composed by 41 components from 3 copy number features. Components definitions are provided in Table 1 (in the same order as below, i.e. segsize1=first segment size component listed, etc.).
[0160] Clinical Applications
[0161] The above methods find applications in the context of predicting whether a subject's tumour is likely to acquire a gene amplification, such as an oncogene amplification, or a gene deletion, such as a tumour suppressor gene deletion. Oncogene amplification and tumour suppressor gene deletion are associated with treatment response and prognosis. Thus, the disclosure also provides a method of determining a prognosis and / or predicting a treatment response for a subject. Further, the disclosure also provides a method of treating cancer in a subject, wherein the method comprises administering or recommending a subject for administration of a particular therapy, depending on the signature(s) of chromosomal instability that have been found to be active in the sample (and in particular depending on whether the subject is predicted as likely to acquire an oncogene amplification or TSG deletion associated with response to a therapy and / or prognosis).
[0162] Figure 2 illustrates a method of providing a prognosis, identifying a therapy or treating a subject according to an embodiment of the present disclosure.
[0163] At step 200, one or more tumour samples may be obtained. These may have been previously obtained from a subject. At step 202, a tumour copy number profile derived from the sample is analysed using a method as described herein, such as e.g. by reference to Figure 1A, to determine whether the tumour is likely to acquire an oncogene amplification or TSG deletion. At step 204, the subject may be identified as likely to respond to a therapy or to acquire resistance to the therapy (i.e. fail to respond to the therapy), based on the results of step 202. In particular the oncogene amplification or TSG deletion may be one that is known to be associated with resistance to the therapy. Instead or in addition to this, at step 205, the subject may be identified as having a poor or good prognosis based on the results of step 202. In particular the oncogene amplification may be one that is known to be associated with poor prognosis. At step 206, a treatment may be identified based on the results of step 204 and / or step 205. For example, a subject identified as likely to acquire resistance to a therapy through acquisition of the gene amplification or TSG deletion may be selected for treatment with an alternative or additional therapy. Conversely, a subject identified as unlikely to acquire resistance to the therapy may be selected for treatment with the therapy as a first line of therapy. As another example, a subject identified as having poor prognosis may be treated with more aggressive treatment and / or with a treatment suitable for advanced tumours. At optional step 208, the subject may be treated with the therapy(ies) identified at step 206.
[0164] Oncogene amplification and TSG deletion events have been shown to be associated with different response to therapy, and in particularto acquisition of resistance to therapy. Thus, also described herein are methods of determining whether a subject that has been diagnosed as having a cancer is likely to acquire resistance to a therapy, the method comprising determining whether the subject is likely to acquire a gene copy number alteration (e.g. oncogene amplification or TSG deletion) using a method as described herein. The method may further comprise classifying the subject between a group that is likely to acquire resistance to the therapy (i.e. unlikely to respond to the therapy), and a group that is unlikely to acquire resistance to the therapy (i.e. likely to respond to the therapy). For example, the method may comprise determining whether a sample from a tumour of the subject has a high or low likelihood of acquiring a gene amplification that has been identified to be associated with resistance to therapy (as explained above). A subject may then be classified in the group that is likely to acquire resistance to the therapy if the sample is determined to have a high likelihood of acquiring the gene amplification, and in a group that is unlikely to acquire resistance to the therapy otherwise.
[0165] Any treatment described herein may be used alone or in combination with another treatment. For example, any treatment with a drug may be used in combination with one or more chemotherapies, one or more courses of radiation therapy, and / or one or more surgical interventions. In particular, any treatment described herein may be used in combination with a treatment for which the subject has been identified as likely to be responsive.
[0166] Additionally, oncogene amplification and TSG deletion events have been shown to be associated with different prognosis. Thus, also described herein are methods of providing a prognosis for a subject that has been diagnosed as having a cancer, the method comprising determining whether the subject is likely to acquire a gene copy number alteration (e.g. oncogene amplification or TSG deletion) using a method as described herein. The method may further comprise classifying the subject between a group that has good prognosis, and a group that has poor prognosis. For example, the method may comprise determining whether a sample from a tumour of the subject has a high or low likelihood of acquiring a gene copy number alteration that has been identified to be associated with prognosis (as explained above). A subject may then be classified in the group that has poor prognosis if the sample is determined to have a high likelihood of acquiring the gene copy number alteration, and in a group that has good prognosis otherwise.
[0167] Whether a prognosis is considered good or poor may vary between cancers and stage of disease. In general terms a good prognosis is one where the overall survival (OS), disease free survival (DFS) and / or progression-free survival (PFS) is longer than that of a comparative group or value, such as e.g. the average for that stage and cancer type. A prognosis may be considered poor if OS, DFS and / or PFS is lower than that of a comparative group or value, such as e.g. the average for that stage and type of cancer. Thus, in general terms, a “good prognosis” is one where survival (OS, DFS and / or PFS) and / or disease stage of an individual patient can be favourably compared to what is expected in a population of patients within a comparable disease setting. Similarly, a ‘‘poor prognosis” is one where survival (OS, DFS and / or PFS) of an individual patient is lower (or disease stage worse) than what is expected in a population of patients within a comparable disease setting.
[0168] The subject is preferably a human patient. The subject may be a patient with a cancer selected from any cancer type present in the TCGA and / or PCAWG data. The cancer may be selected from: ovarian cancer, breast cancer, endometrial cancer, kidney cancer, lung cancer, pancreatic cancer, liver cancer, oesophagus cancer, stomach cancer, head and neck cancer, brain cancer, colon cancer, pancreatic cancer, prostate cancer, bladder cancer, cervical cancer, leukemia, lymphoma, testicular cancer, thyroid cancer, melanoma, adrenal cancer, bowel cancer, sarcoma, thymoma, neuroendocrine tumour, and bile duct cancer.
[0169] Systems
[0170] Figure 3 shows an embodiment of a system for implementing any method according to the present disclosure. The system comprises a computing device 1 , which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network 6, to sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 storing sequence data. The one or more databases may additionally store other types of information that may be used by the computing device 1 , such as e.g. reference sequences, parameters, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network such as e.g. over the public internet or over WiFi. The sequence data acquisition 3 means may be in wired connection with the computing device 1 and / or the database 2, or may be able to communicate through a wireless connection, such as e.g. through a network 6, as illustrated. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 3 are configured to acquire sequence data from nucleic acid samples, for example genomic DNA samples extracted from cells and / or tissue samples. The sequence data acquisition means is preferably a next generation sequencer, but can also be an array reader. The sequence data acquisition means 3 may be in direct or indirect connection with one or more databases 2, on which sequence data (raw or partially processed) may be stored. The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.
[0171] EXAMPLES
[0172] Introduction
[0173] Mutation, selection and genetic drift ultimately determine the structure of a tumour genome. Therefore, to accurately predict future somatic copy number alterations (CNAs), we need an understanding of how tumours acquire new CNAs and the selective pressures that act on them.
[0174] The inventors have recently described that tumours acquire somatic CNAs through different mechanisms, reflected in a compendium of pan-cancer chromosomal instability (CIN) signatures (Drews et al. 2022). These signatures provide a framework to read out the different types of CIN present in a tumour. Given that the probability of fixation of a CNA in a tumour can be linked to the rate at which the CNA emerges, the inventors hypothesised that the mutational mechanisms within a tumour influence the next most likely step in its progression. For example, tumours with breakage-fusion-bridge cycle errors tend to have higher numbers of focal amplifications. Therefore, the inventors hypothesised that it should be possible to use CIN signatures to identify the mutational processes active in a tumour and use their presence to forecast amplification.
[0175] Example 1 describes a new approach to CIN signature feature encoding to enable CIN signature quantification for individual copy number events. Example 2 provides an investigation of the mutation generating processes and selection in cancer demonstrating the theoretical feasibility of associating signatures of chromosomal instability with oncogene amplification. Example 3 explains a method of identifying CIN processes underlying oncogene amplification using the framework in Example 1 , in particular associating signature specific weights to oncogene amplifications. Example 4 demonstrates the use of the weights obtained in Example 3 to predict future oncogene amplification. Example 5 extensively evaluates the performance of the predictor developed in Examples 3 and 4. Example 6 provides an adapted version of the process described in Example 3, which uses weights identified as described in Example 3 but using a cohort of samples selected as later time points from longitudinal data. Example 7 evaluates the robustness of the predictor developed in Examples 3 and 4. Example 8 demonstrates the clinical utility of the predictor of Examples 3 and 4 to predict future treatment resistance and prognosis. Example 9 describes a variant of the approach in examples 3 and 4, using a pan-oncogene model. Example 10 describes a variant of the model described in Example 4 which includes a term that models the effect of DNA topology. Example 11 describes how the framework of Examples 4 and 10 can be adapted to forecast homozygous deletions. Example 12 extensively evaluates the performance of the predictor of Example 10. Example 13 provides an adapted version of the process described in Examples 3 and 10, which uses weights identified as described in Example 3 but using a cohort of samples selected as later time points from longitudinal data. Example 14 evaluates the robustness of the predictor developed in Example 10. Example 15 demonstrates the clinical utility of the predictor of Example 10 to predict future treatment resistance and prognosis. Data & Methods
[0176] Data Sources. All data used in the present examples are publicly accessible, except for the Hartwig Medical Foundation data (WGS). The access to the Hartwig data can be obtained by a Data Access Request through the Hartwig Medical Foundation https: / / www.hartwigmedicalfoundation.nl / en / data / data-access-request / . The TCGA copy number profiles used here were from Drews et al., 2022 (files hosted on github.com / VanLoo-lab / ascat). TCGA clinical data is available at TCGA Research Network: www.cancer.gov / tcga. PCAWG copy number profiles were from Gerstung, et al. 2020; Dentro, et al. 2021 (Files hosted on dcc.icgc.org / releases / PCAWG). Clinical annotation for PCAWG data is available from Pan-Cancer Analysis of Whole Genomes Consortium (2020). Files hosted on dcc.icgc.org / releases / PCAWG / . Triple-negative status for TCGA breast cancer samples was from Lehmann, et al., 2016. PCAWG SV amplicons annotation was from Kim et al., 2020. GLASS copy number profiles were from Barthel et al., 2019 (files hosted on synapse.org / glass under the project ID syn17038081). GLASS clinical data was from Barthel et al., 2019 (files hosted on synapse.org / glass under the project ID syn17038081). TRACERx Lung copy number profiles and clinical data from Bakir et al, 2023 are publically available at zenodo.org / records / 7649257. Raw single cell DNA sequencing data from TNBC tumours and cell lines was from Minussi et al., 2021 (files hosted on the NCBI Sequence Read Archive under the accession number PRJNA629885 www.ncbi.nlm.nih.gov / sra7terrrpPRJNA629885). Driver gene classification was from Bailey et al., 2018; Jassal et al., 2020; Martinez-Jimenez et al., 2020; Sondka et al., 2018. List of non-cancer genes was from Ensembl GRCh37 genome browser: grch37.ensembl.org / index.html. Genomic data of parental and resistant lung cancer cell lines was generated in this project (see section “Experimental data”).
[0177] Software. If not otherwise stated, all analyses were performed using the free statistical software R (v>4.0). If functions outside the standard packages were used, they are mentioned and cited throughout the text. The following packages were used: ggradar v.0.2 (Bion, R. ggradar: Create radar charts using ggplot2. (2023)); showtext v.0.9-6 (Qiu, Y. & details., A. of T. I. S. S. F. A. F. showtext: Using Fonts More Easily in R Graphs. Preprint at CRAN.R-project.org / package=showtext (2023)); showtextdb v.3.0 (Qiu, Y. & details., A. of T. I. F. S. F. A. F. showtextdb: Font Files for the 'showtext’ Package. CRAN.R- project.org / package=showtextdb (2020)); sysfonts v.0.8.8 (Qiu, Y. & details., A. of T. I. F. S. F. A. F. sysfonts: Loading Fonts into R. CRAN.R-project.org / package=sysfonts (2022)); QDNAseqmod v.1.21.0 (github.com / gmaci / QDNAseqmod); randomcoloR v.1 .1.0.1 (Ammar, R. randomcoloR: Generate Attractive Random Colors. CRAN.R-project.org / package=randomcoloR (2019)); ROCit v.2.1 .1 (Khan, M. R. A. & Brandenburger, T. ROCit: Performance Assessment of Binary Classifier with Visualization. CRAN.R-project.org / package=ROCit (2020)); CNpare v.0.99.0 (Chaves-Urbano, B., Hernando, B., Garcia, M. J. & Macintyre, G. CNpare: matching DNA copy number profiles. Bioinformatics 38, 3638-3641 (2022)); forcats v.1.0.0 (Wickham, H. forcats: Tools for Working with Categorical Variables (Factors). Preprint at https: / / CRAN.R-project.org / package=forcats (2023)); purrr v.1.0.1 (Wickham, H. & Henry, L. purrr: Functional Programming Tools. Preprint at CRAN.R- project.org / package=purrr (2023)); tidyr v.1.3.0 (Wickham, H., Vaughan, D. & Girlich, M. tidyr: Tidy Messy Data. Preprint at CRAN.R-project.org / package=tidyr (2023)); tidyverse v.2.0.0 (Wickham, H. et al. Welcome to the tidyverse. Journal of Open Source Software vol. 4 1686 doi.org / 10.21105 / joss.01686 (2019)); knitr v.1.42 (Xie, Y. Dynamic Documents with R and Knitr, 2nd Edition. (2015)); cowplot v.1 .1 .1 (Wilke, C. O. cowplot: Streamlined Plot Theme and Plot Annotations for ‘ggplot2’. CRAN.R-project.org / package=cowplot (2020)); lattice v.0.20-45 (Sarkar, D. Lattice: Multivariate Data Visualization with R. Preprint at lmdvr.r-forge.r-project.org (2008)); patchwork v.1 .1 .2 (Pedersen, T. L. patchwork: The Composer of Plots. Preprint at CRAN.R- project.org / package=patchwork (2022)); mclust v.6.0.0 (mclust: Gaussian Mixture Modelling for Model- Based Clustering, Classification, and Density Estimation; cran.r- project.org / web / packages / mclust / index.html); Isa v.0.73.3 (Wild, F. Isa: Latent Semantic Analysis. CRAN.R-project.org / package=lsa (2022); qgraph v.1.9.4 (Epskamp, S., Cramer, A. O. J., Waldorp, L. J., Schmittmann, V. D. & Borsboom, D. qgraph: Network Visualizations of Relationships in Psychometric Data. Journal of Statistical Software vol. 48 1-18 (2012)); RColorBrewer v.1.1-3 (Neuwirth, E. RColorBrewer: ColorBrewer Palettes. Preprint at CRAN.R- project.org / package=RColorBrewer (2022)); ggthemes v.4.2.4 (Arnold, J. B. ggthemes: Extra Themes, Scales and Geoms for ‘ggplot2’. CRAN.R-project.org / package=ggthemes (2021)); ComplexHeatmap v.2.12.1 (Gu, Z., Eils, R. & Schlesner, M. Complex heatmaps reveal patterns and correlations in multidimensional genomic data. Bioinformatics 32, 2847-2849 (2016)); ggplot2 v.3.4.2 (Wickham, H. ggplot2: Elegant Graphics for Data Analysis. Preprint at ggplot2.tidyverse.org (2016)); iterators v.1.0.14 (Analytics, R. & Weston, S. iterators: Provides Iterator Construct. CRAN.R- project.org / package=iterators (2022)); GenomicRanges v.1 .50.0 (Lawrence, M. et al. Software for Computing and Annotating Genomic Ranges. PLoS Computational Biology vol. 9 doi. org / 10.1371 / journal.pcbi.10031 18 (2013)); IRanges v.2.32.0 (Lawrence, M. et al. Software for Computing and Annotating Genomic Ranges. PLoS Computational Biology vol. 9 doi.org / 10.1371 / journal.pcbi.10031 18 (2013)); S4Vectors v.0.36.0 (Pages, H., Lawrence, M. & Aboyoun, P. S4Vectors: Foundation of vector-like and list-like containers in Bioconductor. bioconductor.org / packages / S4Vectors (2022)); data.table v.1 .14.8 (Dowle, M. & Srinivasan, A. data.table: Extension of 'data.frame'. CRAN.R-project.org / package=data.table (2023)); Biobase v.2.58.0 (Huber, W. et al. Orchestrating high-throughput genomic analysis with Bioconductor. Nature Methods vol. 12 115-121 nature.com / nmeth / journal / v12 / n2 / full / nmeth.3252.html (2015)); MEDICC2 v.1 .0.0 (Kaufmann, T. L. et al. MEDICC2: whole-genome doubling aware copy-number phylogenies for cancer evolution. Genome Biol. 23, 241 (2022)); survminer v.0.4.9 (Kassambara, A., Kosinski, M., Biecek, P. & Fabian, S. Drawing Survival Curves using ‘ggplot2’. (2021 )); survival v.3.5-5 (Therneau, T. M. A package for survival analysis in R. (2023)); PRROC v.1.3.1 (Keilwagen, J., Grosse, I. & Grau, J. Area under precision-recall curves for weighted and unweighted data. PLoS One 9, e92209 (2014)); gridExtra v.2.3 (Auguie, B. gridExtra: Miscellaneous Functions for ‘Grid’ Graphics. Preprint at CRAN.R- project.org / package=gridExtra (2017)); ggbeeswarm v.0.7.1 (Clarke, E., Sherrill-Mix, S. & Dawson, C. ggbeeswarm: Categorical Scatter (Violin Point) Plots. CRAN.R-project.org / package=ggbeeswarm (2022)); pROC v.1 .18.0 (Robin, X. et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics vol. 12 77 (2011)); biomaRt v.2.52.0 (Durinck, S., Spellman, P. T., Birney, E. & Huber, W. Mapping identifiers for the integration of genomic datasets with the R / Bioconductor package biomaRt. Nat. Protoc. 4, 1 184-1191 (2009)); lubridate v.1 .9.2 (Grolemund, G. & Wickham, H. Dates and Times Made Easy with lubridate. Journal of Statistical Software vol. 40 1- 25 www.jstatsoft.org / v40 / i03 / (201 1)); stringrv.1 .5.0 (Wickham, H. stringr: Simple, Consistent Wrappers for Common String Operations. CRAN.R-project.org / package=stringr (2022)); readr v.2.1.4 (Wickham, H., Hester, J. & Bryan, J. readr: Read Rectangular Text Data. CRAN.R- project.org / package=readr (2023)); tibble v.3.2.1 (Muller, K. & Wickham, H. tibble: Simple Data Frames. CRAN.R-project.org / package=tibble (2023)); rstudioapi v.0.14 (Ushey, K., Allaire, J. J., Wickham, H. & Ritchie, G. rstudioapi: Safely Access the RStudio API. CRAN.R-project.org / package=rstudioapi (2022)); circlize v.0.4.15 (Gu, Z., Gu, L., Eils, R., Schlesner, M. & Brors, B. circlize implements and enhances circular visualization in R. Bioinformatics vol. 30 2811 -2812 (2014)); flexmix v.2.3-19 (Grun, B. & Leisch, F., 2008); ggrepel v.0.9.3 (Sarkar, D. Lattice: Multivariate Data Visualization with R. Imdvr.r- forge.r-project.org (2008)); reshape2 v.1 .4.4 (Wickham, H. Reshaping Data with the reshape Package. Journal of Statistical Software vol. 21 1-20 www.jstatsoft.org / v21 / i12 / (2007)); dplyr 1.1.1 (Wickham, H., Francois, R., Henry, L., Muller, K. & Vaughan, D. dplyr: A Grammar of Data Manipulation. CRAN.R-project.org / package=dplyr (2023)); SnowballC v. 0.7.0 (Bouchet-Valat, M. SnowballC: Snowball Stemmers Based on the C ‘libstemmer’ UTF-8 Library. CRAN.R- project.org / package=SnowballC (2020)); igraph v. 1 .4.1 (Csardi, G. & Nepusz, T. The igraph software package for complex network research. InterJournal vol. Complex Systems 1695 igraph.org (2006)); lemon v. 0.4.6(Edwards, S. M. lemon: Freshing Up your ‘ggplot2’ Plots. Preprint at CRAN.R- project.org / package=lemon (2022)); ggpubr v.0.6.0 (Kassambara, A. ggpubr: ‘ggplot2’ Based Publication Ready Plots. Preprint at CRAN.R-project.org / package=ggpubr (2023)); YAPSA v1.22.0 (Hubschmann, D. et al. 2021); doMC v. 1 .3.8 (Analytics, R. & Weston, S. doMC: Foreach Parallel Adaptor for ‘parallel’. CRAN.R-project.org / package=doMC (2022)); foreach v.1.5.2 (foreach: Provides Foreach Looping Construct, cran.r-project.org / web / packages / foreach / index.html); GenomelnfoDb v.1.34.8 (Arora, S., Morgan, M., Carlson, M. & Pages, H. GenomelnfoDb: Utilities for manipulating chromosome names, including modifying them to follow a particular naming style. bioconductor.org / packages / GenomelnfoDb (2023).); DescTools v.0.99.49(DescTools: Tools for Descriptive Statistics; cran.r-project.org / package=DescTools); QDNAseq v.1.21.0 (Scheinin, I. et al. DNA copy number analysis of fresh and formalin-fixed specimens by shallow whole-genome sequencing with identification and exclusion of problematic regions in the genome assembly. Genome Res. 24, 2022-2032 (2014)); this. path v.1 .2.0 (Simmons, A. this. path: Get Executing Script’s Path, from ‘Rgui’, ‘RStudio’, 'VSCode', ‘source()’, and ‘Rscript’ (Shells Including Windows Command Line / / Unix Terminal). CRAN.R-project. org / package=this. path (2023)); BiocGenerics v.0.44.0 (Huber, W. et al. Orchestrating high-throughput genomic analysis with Bioconductor. Nature Methods vol. 12 115-121 Preprint at www.nature.com / nmeth / journal / v12 / n2 / full / nmeth.3252.html (2015)).
[0178] Kaplan-Meier estimates (function survfit) and Cox proportional hazard models (function coxph) were performed using the survival and survminer R packages. Experimental data
[0179] Human cancer cell lines. The 4 EGFR-mutant NSCLC cell lines (HCC827, HCC4006, PC9 and H1975) were obtained from the American Type Culture Collection (Manassas, VA). All cell lines were maintained in Roswell Park Memorial Institute (RPMI) 1640 medium (GIBCO, Carlsbad, CA) with 10% fetal bovine serum (FBS), penicillin (100 U / mL), and streptomycin (50 mg / mL) (Complete Medium) in a humidified CO2incubator at 37°C. All cells were passaged for less than 3 months before being renewed with frozen, early-passage stocks. Cells were regularly screened for mycoplasma using a MycoAlert Mycoplasma Detection Kit (Lonza). Before performing the experiments, cell line identities were confirmed by STR profiling and EGFR mutation was confirmed by droplet digital PCR (ddPCR) using the following ddPCR™ Mutation Detection Assays (Bio-Rad Laboratories): ddPCR EGFR Exon 19 Deletions Screening Kit (12002392), EGFR p.L858R (dHsaMDS463105111), and EGFR p.T790M (dHsaMDS759157834).
[0180] Resistant clones were generated by long-term culture of individual cell lines with high concentration of a EGFR tyrosine kinase inhibitor (EGFR-TKI), as previously described (Ramiraz et al. 2016). Briefly, a total of 1x106cells were seeded into 10-cm dishes and allowed to adhere overnight. The following day, the medium was replaced with fresh complete medium containing 2 mM erlotinib for HCC827, 10 mM erlotinib for HCC4006 and PC9, or 0.7 mM osimertinib for H1975. After approximately 9 days, most cells perished, leaving behind a small number of isolated survivor cells. Clearly separated colonies were isolated and transferred to 96-well plates between 8 and 12 weeks of drug treatment. From there, colonies were expanded successively from 96-well plates to 24-well plates, then to 6-well plates, until finally reaching confluency in a 10-cm plate, thereby establishing individual resistant clones. During the whole process and after generation of resistant clones, medium was replaced every 2-3 days with fresh drug at the same concentration used initially. Cell lines were periodically examined and found negative for mycoplasma contamination during the course of this work. Canonical resistant mutations, such as T790M (dHsaMDS759157834) and C797S (dHsaMDS759157834), were assessed by ddPCR and tested negative for all resistant clones.
[0181] DNA extraction. Formalin-fixed, paraffin-embedded (FFPE) tissue blocks were cut as 5pm sections. Microdissection was not used for recovering tumour-enriched regions. DNA was extracted from 10-25 sections using QIAamp DNA FFPE Tissue Kit from Qiagen (Cat. No. 56404) following manufacturer’s instructions. DNA extraction from cell pellets of about five million cells per sample was performed using the DNeasy® Blood & Tissue Kit from Qiagen (Cat. No. 69504). Cell pellets were equilibrated at room temperature for 15 minutes and then resuspended in PBS as recommended by the manufacturer. DNA was finally eluted in Tris-HCI (pH=8, 10 mM) instead of using Qiagen’s Buffer AE. This buffer was not used for elution because it contains EDTA, which can inhibit some needed enzymatic reactions during Library Preparation. DNA was quantified using the ADNds Quant-iT™ PicoGreen™ kit from Invitrogen™ (Cat. No. P7589).
[0182] DNA sequencing - Shallow whole genome sequencing (sWGS). For sample preparation, DNA at 5 ng / pL was required. 50 pL of sample was dispensed into a Covaris microTUBE AFA Fiber Pre-Slit Snap-Cap 6x16mm (Cat. No. 520045) and DNA was mechanically fragmented using the E220evolution focused-ultrasonicator from Covaris. 150 seconds of shearing for cell pellets and 300 seconds for FFPE samples was needed for obtaining fragments of about 240 bp. DNA profile was checked with the LabChip DNA High Sensitivity Reagent kit from PerkinElmer (Cat. No. CLS760672) using the LabChip GX Touch Nucleic Acid Analyzer (Cat. No. CLS137031) from PerkinElmer. For library preparation, ThruPLEX® DNA-Seq Kit (Cat. No. R400674) was used following manufacturer's instructions for 50 ng of input DNA sample, but performing 5 cycles for sample amplification. We used indexes from the DNA HT Dual Index Kit - 24 N from Takara (Cat. No. R400664) for cell pellet DNA samples, and from the Unique Dual Index Kit - 48U from Takara (Cat. No. R400744) for FFPE DNA samples. Library cleanup was performed with Agencourt AMPure XP Reagent (Cat. No. A63881) following manufacturer's instructions. Final quality control of libraries was assessed via sample quantification using the ADNds Quant-iT™ PicoGreen™ kit, and via library profiles using the LabChip DNA High Sensitivity Reagent kit. Libraries were sequenced by the Genomics Unit of the Spanish National Cancer Research Centre (CNIO) with the NextSeq™ 550 system from Illumina.
[0183] DNA sequencing - Whole exome sequencing (WES). For sample preparation, a total of 200 ng of DNA was required. WES was conducted by a commercial service (Macrogen Inc., Seoul, Korea). Libraries were constructed using the SureSelect Human All Exon v6 (Agilent Technologies, Santa Clara, CA, USA). Sequencing of pair-end reads of 151 bp length was performed with the Novaseq system from Illumina, obtaining an average total reads across samples of 350 million reads (range 98-510 million reads).
[0184] Sequence read alignment. Reads were aligned as single-end against the human genome assembly GRCh37 using BWA-MEM (v.0.7.17)(Li and Durbin, 2009). Duplicate reads were then identified and marked using samtools-markdup (v1 .15)(Li et al. 2009).
[0185] Generating copy number profiles. After alignment, absolute copy number profiles were fitted from sequencing data generated across different high-throughput technologies (method applied for each dataset is detailed below). Following absolute copy number fitting, samples were rated using a star system as previously described (Macintyre, G. et al. 2018; Madrid, L. et al. 2023). Briefly, samples with noisy or flat copy number profiles where absolute copy number cannot be estimated are given 1 -star, samples with fittable absolute copy number but with some noise are given 2-star, well fit samples without noise are given 3-star. Those with 1-star were discarded for downstream analyses. In addition, samples without the appropriate read counts (15 number of reads per bin per chromosome copy) for supporting the fitting were also discarded.
[0186] Absolute copy number fitting from sl / l / GS. The inventors used the QDNAseq R package (Scheinin, I. et al. 2014) to count reads within 30 kb bins (resolution threshold determined to accurately call copy number signatures in Drews et al. 2022). Bins mapped to centromeres and regions of undefined sequence in the reference genome hg19 were excluded. Read counts were then corrected for sequence mappability and GC content, followed by copy number segmentation. After the profile was segmented, we inferred absolute copy numbers for combinations of purity and ploidy values (0.05<purity<1 , and 1 .5<ploidy<8) as follows: where purity is the fraction of tumour cells in the sample, r is the read counts, and d is a constant proportional to the read depth and the average absolute copy number of tumour cells in the sample
[0187] The optimal mathematical solution was considered the one with the lowest root mean squared deviation (RMSD) between the non-rounded and the rounded copy numbers for all bins. All solutions were manually inspected to either confirm or correct for more appropriate fits. For pure cancer models (i.e. cell lines, single cells), a range of purities from 0.9 to 1 for fitting absolute copy numbers can be used instead.
[0188] Absolute copy number fitting from WES. Paired-end raw reads were aligned to the human hg19 reference genome. The resulting alignments were split into equally-sized bins of 50 kb in size. Bins were annotated with replication timing and a set of metadata inherited from QDNAseq: gc, mappability, residuals, bases and blacklist. The annotated bins were then interrogated for overlaps with target regions (bed file) and divided in two groups: 1) off-target bins, with a zero overlap with targeted regions. In this case, reads were counted normally in the entire bin. 2) on-target bins, those with at least 1 overlap with a target region. In this case every overlap was deleted and the remaining off-target regions were pasted together forming a new, smaller bin, ready for read counting. GC correction was performed using LOESS. To correct for artificially high off-target read counts in parts of the genome with high sequence similarity to the target regions we generated a score per-bin that quantified the magnitude of this bias and used it for a single LOESS fit and correction. We segmented these data using a modified version of the circular binary segmentation (CBS) from the DNAcopy R package. Absolute copy numbers were computed as described above in the “Inferring absolute copy number from sWGS” section.
[0189] Absolute copy number fitting from single cell l / l / GS. To infer copy number profiles from single cell DNA sequencing data, the inventors defined 100 kb windows along the genome and applied the same methodology described in the “Inferring absolute copy number from sWGS” section.
[0190] Absolute copy number fitting from pseudobulk l / l / GS. The inventors reconstructed bulk copy number profiles by concatenating reads of all single cells sequenced from a specific sample. They then downsampled pseudobulk bams to appropriate read counts based on the bin size (30 kb), purity and ploidy. Absolute copy number fitting can then be performed as described in the “Inferring absolute copy number from sWGS” section.
[0191] Detection of focal amplifications - bulk samples. Focal amplifications in unpaired WGS, WES and SNP6 bulk data were defined as genomic regions with absolute copy number higher or equal to 8. This threshold was selected after an extensive analysis using a range of threshold values (see section “Assessing the threshold for defining focal amplifications”). The same threshold was applied for detecting clonal amplifications in pseudobulk profiles. In longitudinal analyses, the acquisition of a new focal amplification in the late / relapse / metastasis biopsies was defined as the gain of at least 6 copies of a genomic region with respect to the early time point sample.
[0192] Detection of focal amplifications single cells. Unique focal amplifications were defined as those genomic regions with absolute copy numbers higher or equal to 8 present in a specific cell. If those amplifications were present in more than one cell, the inventors defined them as shared focal amplifications. Given the uncertainty in the location of the breakpoints during segmentation of single cell copy number profiles, the inventors took a heuristic approach to determine whether an amplification was truly present in only one or more cells. In brief, they extended each of the putative unique amplifications by a 1 Mb window per segment side, and looked for overlaps with the amplifications present in other cells. If they found overlaps, the amplification was classified as shared, and otherwise as unique.
[0193] Detection of focal deletions - bulk samples. Focal deletions in unpaired WGS, WES and SNP6 bulk data were defined as genomic regions with absolute copy number lower or equal to 0.5.
[0194] Exploring acquisition of focal amplifications. The inventors first explored how amplifications arise in the absence of selection pressures. In this regard, they used publicly available single cell sequencing data from 8 triple negative breast cancer (TNBC) patients and 4 TNBC cell lines (Minussi, D. C. et al. 2021). Focal amplifications present in one single cell (unique amplifications) are randomly caused by a mutational process and their acquisition is therefore not subjected to selection pressures. Unique amplifications cannot be detected at bulk level. On the contrary, only positively selected amplifications (which are therefore shared across a high number of cells in the tumour, i.e. clonal amplifications) can be detected at bulk level. The inventors used a binomial test (function binom.test from the R package stats) to compare the fraction of cells carrying at least one unique focal amplification with the fraction of clonal amplifications observed across the 12 matched pseudobulk samples and the 145 TCGA-TNBC samples. In addition, they tested if having higher amplification rates is required for acquiring amplifications targeting oncogenes. Jonckheere’s trend test (function JonckheereTerpstraTest from the DescTools package in R) was used to test an association between the number of amplified segments that span and that do not span TCGA-BRCA oncogenes. Finally, they compared the number of unique focal amplifications detected across all single cells with the number of clonally-expanded focal amplifications seen at pseudobulk resolution (Figure 8C).
[0195] Exploring selection of focal amplifications. The inventors tested the role of selection in sculpting the landscape of focal amplifications in a tumour. Briefly, they compared the number of focal amplifications spanning oncogenes with a background distribution derived by using a Monte Carlo approach. First, they empirically selected the 10 oncogenes most frequently amplified in the TCGA-TNBC cohort and counted the number of focal amplifications mapping to those 10 cancer-specific oncogenes. A background distribution was then derived by counting the number of focal amplifications mapping 10 non-cancer genes which were randomly selected from a list of 53,989 total non-cancer genes annotated in ensembl GRCh37 genome assembly. In total, the inventors ran this procedure 1 ,000 times. Statistical significance was measured by calculating the empirical p-value. Empirical p-values were computed as the fraction of iterations where the number of amplifications within the 10 non-cancer genes (background) was greater than the number of amplifications observed within the 10 TNBC-specific oncogenes (Figure 8B).
[0196] Distribution of unique amplifications across the genome. The inventors tested whether the acquisition of amplifications was uniform across the genome. Specifically, they explored the genomic distribution of unique amplifications in single cells from 12 TNBC samples. For each sample, they initially mapped all observed unique amplifications to their respective genomic regions, arranging them in sequential order from the start to the end of the chromosome. They next calculated the distance of each unique amplification to its nearest adjacent unique amplification. These distances were then normalised by the total length of the chromosome. Only unique amplifications located in autosomal chromosomes were considered in this analysis. Next, for each autosomal chromosome, the inventors estimated the expected distance between adjacent unique amplifications, assuming a homogeneous amplification rate across the genome. This distance was computed by dividing the chromosome length by the total number of unique amplifications observed on it. They normalised this distance by the chromosome length to facilitate the comparison across chromosomes, resulting in the expected distance being equal to 1 / (number unique amplifications + 1). Finally, they transformed those distances into a relative measure by subtracting the expected distances to the observed distances. This relative distance facilitates interpretation: a distribution centred at 0 implies an even distribution of unique amplifications across the genome. The inventors determined if the observed distance distribution significantly deviated from this expectation by performing a one-sample Wilcoxon two-sided test, with the null hypothesis assuming a mean distance equal to 0 and the alternative hypothesis suggesting a mean distance different from 0 (Figure 8D).
[0197] Exploring acquisition and selection of focal amplifications - simulations. The inventors performed simulations to model the acquisition of an oncogenic amplification with varied amplification rates and selection coefficients. They modelled cancer evolution as a stochastic system of birth (occurring at rate?) and death (occurring at rate a) events. In this process, cells without oncogene amplification can acquire it in the process of cell division at rate ju (Figure 8E). This means that the acquisition of oncogene amplification occurs at overall rate j?ju. In our simulations, we conferred to cells with oncogene amplification a fitness advantage (s), which increases the division rate by (1 + s) and decreases death rate by (1 -s). Overall, this means that tumour cells without amplification die with rate a, divide yielding two cells without amplification with rate j?(1-ju) and divide giving rise to one cell with amplification and another without with rate j?ju. Meanwhile, tumour cells with amplifications die with rate a(1-s) and divide with rate j?(1 +s), giving rise to two cells with amplification.
[0198] The inventors used the Gillespie algorithm (Lopez, S. et al. Interplay between whole-genome doubling and the accumulation of deleterious alterations in cancer evolution. Nat. Genet. 52, 283-293 (2020)) to implement the simulations in practice. For all simulations, we used the following fixed parameters: starting population of 1 ,000 cells, cell division rate of 0.3, cell death rate of 0.15, and maximum number of amplified drivers equal to 1. In the different simulations, we tested a range of amplification rates (from 5e-6 to 0.1) and selection coefficients (from 0 (neutral selection) to 1). Simulations of tumour growth were run until the population size reached 200,000 cells or until 20,000 generations. In total, the inventors ran 20 different simulations for each of the combinations of amplification rate and selection coefficient. For each simulation, they calculated the fraction of cells that acquired the oncogene amplification and computed the average across simulations with the same parameters (Figures 8E, F).
[0199] Example 1 - Mapping signatures to individual copy number events
[0200] An approach to evaluate the origin and extent of chromosomal instability at a genome-wide level has been previously described in Drews, R. M. et al. 2022. This approach was used to identify a compendium of 17 copy number signatures. As part of this method, the copy number changes observed in a tumour were embedded into a feature space consisting of 5 fundamental copy number features (the segment size, the difference in copy number between adjacent segments, the lengths of oscillating copy number segment chains, the breakpoint count per 10 Mb, and the breakpoint count per chromosome arm). These 5 fundamental features were capable of capturing well-known copy number patterns and thus were able to distinguish different CIN types at genome-wide level.
[0201] However, this feature encoding cannot be directly applied to analyse individual copy number changes in a genome. In the original feature encoding the 3 copy number features used to model the distribution of copy number alterations across the genome (the lengths of oscillating copy number segment chains, the breakpoint count per 10 Mb, and the breakpoint count per chromosome arm) were encoded per genomic region rather than per individual copy number event. Thus, they cannot be used to map CIN signatures to each individual copy number event in a tumour.
[0202] To overcome this limitation, in the present example, the inventors designed a single new feature to replace these 3 features (Figure 4a). This new feature, called “the breakpoint per side” feature, is described below. This feature, alongside the segment size and changepoint features, can be used to recapitulate the original CIN signatures and facilitates mapping of signatures to individual copy number events. In addition to this new feature, the inventors also made slight modifications to the changepoint feature to enable single event mapping.
[0203] New “breakpoints per side” feature
[0204] When defining this new feature, the inventors focused on an encoding that not only captured the key properties of the 3 features that it was replacing, but also facilitated recovery of the CIN signatures with high weights for these features. Furthermore, they aimed to maintain a distinction between mutational processes leading to both clustered (i.e. replication stress) and isolated (i.e. mitotic errors) copy number alterations. In order to construct an event-based encoding, one must start with a specific copy number event in the genome. For ease in explanation, we call this event the currentCN.
[0205] To encode clustered copy number changes as previously represented by the breakpoint counts per 10 Mb and chromosome arm features, the inventors devised an encoding that took a window around the currentCN and counted the number of breaks occurring either side. The side with the maximum number of breaks was considered the count of the feature. The inventors took the maximum to account for the case where the currentCN is located at the telomere or centromere. For example, for segments located within e.g. 5Mb of the telomere or centromere, the counts would be expected to be underrepresented in the flanking region proximal to the centromere / telomere. This encoding has a high count if other copy number events are clustered with the currentCN and 0 if there are no neighbouring copy number events. The inventors opted for a 5 Mb window flanking each side of the currentCN. While this doesn’t completely represent the breakpoint per chromosome arm feature, they reasoned that it was a good compromise between the two original breakpoint features (the breakpoint count per 10 Mb, and the breakpoint count per chromosome arm). This feature is able to capture more information than just the breakpoint count per 10Mb feature (although it does also capture the information capture by the latter) because if the currentCN is large then the counts effectively span a much larger region than 10Mb. For example, if the currentCN is 40MB, taking +5MB flanking regions means that the total region being interrogated is 50MB (close to the size of a chromosome arm).
[0206] Fortunately, this encoding also managed to capture information encoded by the previous length of oscillating chains feature: neighbouring copy number events of the same size that are part of an oscillating chain will have the same number of breaks in the 5 Mb flanking region, and will therefore cause an overrepresentation of a specific count value in the new feature. Hence, the encoding of the breakpoints per size feature also captures oscillating chains.
[0207] As was also done in Drews et al. 2022, diploid segments were not considered and thus the breakpoints per size feature was only extracted for non-diploid copy number events.
[0208] To test the robustness of this new feature encoding, the inventors rederived signatures and compared their activities to those obtained by using the original encoding (see below).
[0209] Modifications to the changepoint feature
[0210] This feature was also included in the feature space in Drews et al. 2022, however the feature encoding has been modified in order to avoid linking an event next to a focal amplification to a replication stress signature. In the original feature encoding, the changepoint captured the difference in absolute copy number compared to the left neighbouring segment without taking into account whether the relative change was a positive or a negative value. Therefore, segments with low copy number value next to a focal amplification were defined by having high copy number change, and therefore were assigned to one of the amplification-related signatures.
[0211] For non-diploid segments, the relative changepoint value is now calculated with respect to the segments on both sides (even if the neighbouring segments are diploid). The changepoint value is assigned with respect to the left neighbouring segment when this value is equal or higher than -3. In case this condition is not fulfilled, the changepoint value is assigned with respect to the right neighbouring segment. Then, for those segments with relative changepoint values lower than -3 (i.e. segments flanked by focal amplifications on both sides), the changepoint value is recalculated by subtracting 2 (normal copy number) from the copy number value of this segment. In other words, this subtracts the normal state from the current copy number (as if the segment was flanked by normal segments), to avoid assigning segments that are between two focal amplifications to a replication stress signature. Finally, relative changepoint values are transformed to absolute changepoint values. Making these changes improves the precise mapping of each event to a specific signature. This is not as important when signatures are evaluated at the genome wide level, but improves the accuracy at the single event level.
[0212] As in Drews et al. 2022, no changepoint value is extracted for diploid segments. For non-diploid segments located at the beginning of the chromosome, the changepoint value is calculated with respect to the normal copy number value (i.e. 2 is subtracted from the copy number value).
[0213] This change had overall little impact on the signature activities quantified for the TCGA samples using the original feature encoding (Figure 5, see Example 2 for methods of TCGA data analysis). The major difference was a merge of mitotic signatures CX1 and CX6 (the activities of these signatures are highly correlated and it is therefore possible that these two signatures are in fact the same signature artificially split during the NMF process due to limited data input). For that reason, signature CX6 is excluded from any downstream analyses.
[0214] Absolute copy number feature
[0215] As explained in Drews et al. 2022, the inventors made a deliberate decision to exclude the absolute copy number of a segment as a feature for decoding chromosomal instability. This choice was motivated by the need to ensure robustness in the presence of whole-genome doubling (WGD). Through an extensive examination of the ovarian copy number signatures (Macintyre et al. 2018), the inventors discovered that the inclusion of the copy number feature could artificially split a signature based on ploidy status, even when the underlying mutational process was the same. For instance, a single copy loss on a whole genome duplication (WGD) background would result in a copy number state of 3, while the same deletion on a non-WGD background would yield a copy number state of 1. Consequently, a segment loss could be encoded by two distinct signatures, despite originating from the same mutational process that generates deletion patterns. As in Drews et al. 2022, the new feature encoding described herein exclusively incorporates the changepoint feature, which quantifies changes in copy number relative to adjacent segments. By using simulated genomes, the inventors explicitly confirmed that the new feature encoding remains resilient to the influence of WGD effects (see section “Assessing feature encoding performance”). Furthermore, the approach to assign CIN signatures to individual copy number events was also robust to the WGD effects.
[0216] Feature distribution and modelling
[0217] Each feature was split into categories or components by applying mixture modelling to the feature distributions (see Figure 4a). For the segment size and changepoint feature, the same mixture components used in Drews et al. 2022 were used (see Table 1). For the new feature (bpwindow), the inventors estimated Poisson mixture models with the flexmix package in R (Grun, B. & Leisch, F. 2008), and used the Bayesian Information Criterion (BIC) to decide on the number of components. The algorithm was initially run using all data with the number of components ranging from 1 to 10. This only yielded 2 components which was not sufficient to represent range of values observed (Figure 6). Therefore, given that 96% of the data presented less than 6 breakpoints at 5 Mb window per segment side, the inventors decided to manually define 7 components representing from 0 to 6 counts (see Table 1). Then, they applied mixture modelling to the remaining 4% of the data and derived 2 additional components, resulting in a total of 9 components for the breakpoint counts at 5 Mb window per segment size feature (see Table 1). The Lambda / Mean values of each of these distributions are provided in Table 1 . The bpwindow components with lambda values from 0 to 6 were manually predetermined. For the last two components, the inventors run flemix for getting lambda values with the following parameters: MINCOMP=1 , MAXCOMP=10, FITDIST="pois", PRIOR=0.001 , NITER=1000, DECISION="BIC", SEED=77777. Note that other parameters may be used to identify mixtures of Poisson distributions for events that present more than 6 breakpoints per 5Mb, as well as different cohorts of samples, which would result in Poisson distributions with slightly different lambda values.
[0218] In total, this resulted in a feature space defined by 41 mixture components (22 from the segment size feature, 10 from the changepoint feature, and 9 from the breakpoints per side feature).
[0219] Event-by-component probabilities matrix
[0220] Once components were derived, the inventors computed the probability for each copy number event to belong to each component per feature, resulting in a 1x41 dimensional vector of posterior probabilities. For a given sample, all probability vectors were combined in an event-by-component probabilities matrix, which served as the input matrix for mapping each individual copy number event to a specific CIN type.
[0221] Event-by-signature scores matrix
[0222] For each signature and copy number event, the inventors multiplied the component probabilities vector by the signature definition (signature definition in the new feature space, see Table 3 - obtained by linear combination decomposition given a sample level input matrix in the new feature space and the signature definitions in Drews et al. 2022 in the original feature space). Then, they summed up the resulting vector to calculate a signature score per event. For the purpose of oncogene amplification forecasting (se Examples 2 onwards), the resulting event-by-signature matrix was multiplied by the sample-level signature activities (see section “Sample-level signature quantification”) to compute activity-corrected signature scores for each copy number event. This is not performed when computing signature scores for the purpose of assigning signatures to individual events.
[0223] Signature scores per event were finally normalised per signature. This is done per event so that the signature scores sum to 1 , i.e. the score for each signature for a specific event is divided by the sum of scores for all signatures for the event.
[0224] Deriving novel signatures
[0225] The inventors aimed at defining a compendium of CIN signatures in the new feature space (see Figure 4b). In the present examples, the inventors used the signatures in Drews et al. 2022 in the new feature encoding. However, it is also possible to use the new feature encoding to derive signatures from scratch (e.g. using NMF) as described in Drews et al. 2022. To derive signature definitions, they used the 6,335 TCGA samples with sufficient evidence of chromosomal instability (>20 CNAs) according to the threshold estimated in Drews et al. 2022. In the context that the activities of the original copy number signatures are already known for all 6,335 samples, the inventors used the LCD function from YAPSA (Hubschmann, D. et al. 2021) for deriving signature definitions given the new feature space input matrix (6,335 patients by 41 components sum-of-posterior matrix).
[0226] However, the novel feature encoding introduces changes in the input matrix that need to be taken into account for following this procedure. As stated before (see section “Copy number feature encoding: “Modifications to the changepoint feature”), the inventors have modified the original encoding for the changepoint feature, which limits their ability to distinguish between mitotic signatures CX1 and CX6. This precludes the use of the original signature activities in Drews et al. 2022 as the gold-standard activities for deriving event-level signatures.
[0227] After excluding CX6 from the signature definition matrix, the final signature activities for all 6,335 TCGA samples with detectable chromosomal instability were computed using the linear combination decomposition (LCD) function from YAPSA on the input matrix. This input matrix included the sum-of- posterior probabilities for components of the 5 original features but applying the new changepoint encoding.
[0228] As described in Drews et al. 2022, the inventors computed and applied signature-specific thresholds to clarify small signature activities. The activity matrix was finally normalised per sample to obtain the new gold-standard activities for all TCGA samples. This gold-standard sample-by-activities matrix was used for deriving the definitions of the novel signatures, as explained above, using linear combination decomposition. Novel signatures (defined in Table 3) captured the same signal of different CIN types as the original compendium of copy number signatures in Drews et al. 2022 (Figure 7).
[0229] Sample-level signature quantification
[0230] After defining CIN signatures that can be mapped to individual events, the final signature activities for all 6,335 TCGA samples were computed using the linear combination decomposition function from YAPSA on the input matrix (sample-by-component sum-of-posterior matrix). This input matrix was derived by summing up the event probability vectors per sample (i.e. this is a matrix comprising, for each sample, a vector that has activities that are the sum of a row of the event-by-signature matrix (i.e. sum of all event specific activities for a single signature) prior to any adjusting based on sample-level signature and normalising by signature). Indeed, as the linear decomposition is not exact, the sample level activities are re-computed given the signatures defined in the new feature space. For new samples, the component vector (new encoding) is computed, then the new signature definitions (Table 3 or any equivalent signature definitions using the new feature encoding) are used to estimate the activity using linear value decomposition (e.g. as implemented in YAPSA).
[0231] Following the same procedure as in our Drews et al. 2022, the inventors identified signature-specific thresholds to be confident in small signature activities. Briefly, the inventors performed 1000 simulations whereby noise was added to the input components. For signatures with a true zero value, the non-zero simulated values were fitted with a single Gaussian distribution using Mclust from the mclust R package, v5.4.6. A threshold was estimated at 95% of this distribution and activities were set to zero if they were below the signature-specific threshold (Table 4). The sample-by-activity matrix was finally normalised per sample (sample-level activity matrix from here) - dividing by the sum of all signatures for that sample so the signature activity vector sums to one.
[0232] Table 4. Signature-specific thresholds for setting zero activity values.
[0233] The inventors evaluated the performance of the new signature mapping method described above in terms of its ability to capture types of CIN samples in simulated data, correctly assign these types of CIN to individual events, and perform well across a range of technologies and extent of CIN.
[0234] Assessing feature encoding performance
[0235] To evaluate the robustness of the new feature encoding described in Example 1 , the inventors used a dataset of 240 simulated genomes which were generated in Drews et al. 2022. Briefly, they simulated the activity of five CIN types by introducing copy number changes into diploid genomes. These CIN types included: chromosome missegregation via mitotic errors (CHR), large-scale state transitions via homologous recombination deficiency (LST), focal amplification via ecDNA circularisation and amplification (ecDNA), early whole-genome duplication via cytokinesis failure (WGD early), and late whole-genome duplication via endoreduplication (WGD late). For any given sample, three of the five mutational processes were active. Half of the samples had one dominant signature and the other half had two. Each mutational process was designed to be dominant in a specific subset of samples (N=100 for CHR, ecDNA and LST; N=96 for WGD early; N=24 for WGD late).
[0236] The inventors applied the new feature encoding described above to the simulated genomes to generate a sample-by-component sum-of-posteriors matrix. Then, the NMF package in R (with the Brunet algorithm specification) was used to deconvolute the sample-by-component sum-of-posteriors matrix into a sample-level activity matrix and a signature definition matrix. A signature search range of 3 to 6 was applied, running the NMF algorithm 1 ,000 times with different random seeds. The optimal number of signatures was based on achieving stability in cophenetic, dispersion, and silhouette coefficients. These evaluation metrics were computed for the definition matrix (basis), the activity matrix (coefficients), and the connectivity matrix (consensus). This consensus matrix represented sample clusters based on their predominant signature across the 1 ,000 runs. The optimal solution was defined as the smallest number of signatures that achieved stability in the cophenetic, dispersion, and silhouette coefficients; but also maximised sparsity without surpassing the sparsity observed in the randomly permuted matrices. This process yielded a total of 5 signatures.
[0237] Next, the inventors assessed whether the identified signatures effectively captured the simulated CIN types. They observed a strong alignment between the definitions of the identified signatures and the expected definitions of the simulated CIN types within the feature space (average cosine similarity of 0.66; Table 5). The expected definitions were generated by adding weight to those components that reflect the mutational process (adding 1 to those components that define a simulated CIN type, and 0 to those components that do not define it, then normalising the definition vector to sum to 1). The signature activities largely agreed with the simulated distributions across samples, resulting in an overall accuracy of 70.5% (296 / 420) in identifying samples with a dominant signature. The inventors also compared activities of identified signatures by using the original and the new feature encoding. This analysis revealed that both feature spaces were equally proficient in extracting CIN signals, as indicated by an average Pearson's r of 0.71. Overall, these results demonstrate that the new feature encoding described herein effectively accomplishes both identifying signatures of mutational processes and quantifying their corresponding activities.
[0238] Table 5. Similarity between expected and observed definitions of identified signatures.
[0239] Assessing signature assignment to individual events
[0240] The new approach described above for mapping CIN signatures to individual copy number events outputs a score that reflects the likelihood of an event being attributed to each signature. This flexibility accommodates situations where copy number events exhibit features resembling more than one signature, and also avoids assigning a specific CIN type to a copy number event which has been produced by an uncertain or unrepresented mutational process. Note that there are multiple options to map signatures to events, including: 1) assign to the event a probability of being caused by a signature (as described in Example 1), and 2) assign a specific signature to the event by selecting the signature with the highest probability. The process of computing the signature probability or score of an event is performed by multiplying "event by component" vector to the "signature by component" vector. These scores can be but do not need to be corrected by the sample-level activity. Another possibility to map signatures to events would be to assign to an event the signature with the most similar definition compared to the "event by component" vector (e.g. using cosine similarity, Euclidean distance, etc).
[0241] In an attempt to assess the reliability of the proposed mapping approach, the inventors allocated each copy number event to the signature with the highest associated probability. Subsequently, they assessed the accuracy of these assignments by calculating the fraction of events displaying similar features to the definition of the signature assigned. The signature with the highest cosine similarity in comparing the event-by-component vector and the signature-by-component vector was considered as the most similar. This analysis revealed an overall accuracy of 72.2% (479,565 out of 647,630 events), denoting a substantial alignment between the mapped signature and features defining an individual copy number event. Table 6 summarises mapping accuracy for each CIN signature.
[0242] Table 6. Fraction of events correctly assigned to the CIN signature with the most similar definition.
[0243] Assessing number of individual events required for sample-level signature quantification
[0244] The inventors performed an analysis to determine the minimum number of copy number events required within a sample to accurately replicate its signature composition. To achieve this, they progressively increased the number of events of a sample profile via random selection. They then quantified samplelevel signature activities by using the linear combination decomposition function from YAPSA on the input matrix (sample-by-component sum-of-posterior matrix). In each iteration of this decomposition, they expanded the information within the input matrix by progressively introducing an increasing number of events, ranging from 1 to 100. They then evaluated whether a sample’s signature composition (defined by its three dominant signatures) was correctly captured depending on the number of events. This procedure was repeated 1000 times with 1000 TCGA samples being randomly selected for each run. As little as 15 copy number events were sufficient to correctly decompose >90% cases. This procedure was also expanded to encompass individual signatures. For each specific signature, the inventors performed random selection of copy number events from the subset of 200 TCGA samples exhibiting the highest signature activity. Then, they quantified sample-level signature activities and evaluated the success rate of detecting the signature. This procedure was done 100 times per signature. The results indicated that an average of 7 copy number events was sufficient for calling an individual signature with >90% of accuracy.
[0245] Sample-level signature stability across technologies
[0246] CIN signatures were derived using TCGA SNP 6.0 array data because this represents the largest pancancer collection of high-resolution copy number profiled tumours to date. In Drews et al. 2022, the inventors evaluated the robustness of the approach across different high-throughput sequencing technologies by comparing signature definitions and activities. As expected, they saw a drop in similarity scores for the activities and signature definitions between SNP6 and sWGS data (Drews et al. 2022).
[0247] Here, the inventors extended this analysis by means of testing the ability to capture the compendium of signatures in sWGS. For this analysis, the inventors used 478 TCGA samples with detectable levels of CIN that were also part of the PCAWG project. WGS data from the 478 TCGA / PCAWG samples was downsampled to 15 number of reads per bin per chromosome copy (according to the ploidy, purity and bin size) for mimicking sWGS. The inventors fitted absolute copy number profiles (see section “Absolute copy number fittings from sWGS”), extracted the 5 original copy number features but using the new changepoint encoding (see section “Copy number feature encoding: Changepoint”), quantified activities for the 16 compendium signatures (CX6 excluded) using YAPSA, and compared them to the SNP6- derived signature activities using Kendall's tau. This looks at the effect of the changes to the changepoint feature.
[0248] This analysis showed some deterioration of some signatures across technologies. These deteriorations were not due to methodology but rather due to differences in how the copy number is measured by these technologies. All signatures would be detectable given a sufficient number of training samples for that technology. Fortunately, this deterioration was only observed for tumour type-enriched signatures, with the remaining signatures showing robust signals across technologies. Signatures showing robust signals across technologies were classified as universal signatures (CX1 -5, CX8-10, and CX13) whereas others were classified as SNP6-specific signatures (CX7, CX1 1-12, and CX14-17). CX14 was not considered universal because its signal was captured by CX1 in sWGS data. It is important to note here, that SNP6-specific signatures can also be detected in sWGS, it is just that the copy number changes captured by the signatures manifest differently across the two technologies.
[0249] Sample-level signature stability across bin resolutions
[0250] Given that the gold-standard bin size for quantifying the compendium of CIN signatures in Drews et al. 2022 is 30 kb, and the bin size used for deriving copy number profiles from WES data is 50 kb, the inventors explored the robustness of the novel signatures across bin resolutions by comparing activities.
[0251] Here, they segmented copy number profiles from the 478 TCGA / PCAWG samples using 8 different bin sizes (30, 40, 50, 60, 70, 80, 90 and 100 kb). Once absolute copy number profiles were fitted, the inventors quantified the novel signatures developed in this work per sample. They then computed cosine similarities of SNP6-derived signature activities with those obtained at each bin resolution (all of which had very similar distributions with a single peak around cosine similarity ~0.9). In addition, they compared signature activities derived at 30 kb (gold-standard) and 50 kb (resolution used for WES data) using Kendall’s tau (resulting in Kendall's tau between the same signature at the two resolutions between 0.710 and 0.910 for all signature apart from CX10 which had a slightly lower tau of 0.610). These comparisons indicated that the novel signatures are stable across bin resolutions.
[0252] Example 2 - Mutation generating processes and selection
[0253] The ultimate fixation of an amplification is dependent on selective pressures, however selection is challenging to model directly given the diversity of pressures found within any given patient. Across a population of tumours, the frequency of a CNA observed can act as a proxy for selection (Zack, T. I. et al., 2013) and thus can be used to forecast. However, in a single tumour, there needs to be sufficient passenger CNAs present to detect the signature of the mutational process which will ultimately give rise to the future CNA. Contrary to single-base substitutions, where passenger mutations are abundant and have limited functional consequences (Stratton et al. 2009), it is unclear how frequent passenger CNAs are outside the context of whole-genome duplication (Lopez, S. et al. 2020). This potentially creates a challenge for forecasting.
[0254] Classical tumour evolutionary theory states that mutation and selection (plus genetic drift) underpin the acquisition of a future oncogene amplification. In simple terms, individual cells within a tumour acquire new amplifications via various DNA damage mechanisms, generating cell-to-cell karyotypic variation. Selection then favours the expansion of cells with an amplification that confers a fitness advantage.
[0255] Single-cell whole genome DNA sequencing data offers an opportunity to observe this process in action. Amplifications observed exclusively in single cells provide a detailed view of the process of karyotypic diversification in a tumour (Fig. 8A). Meanwhile, amplifications detectable in bulk (or pseudo-bulk) DNA sequencing data indicate those that are present in multiple cells, which means that they were present in a cell that underwent a clonal expansion in the past, potentially driven by the fitness advantage conferred by the amplification itself (Fig. 8A). The inventors compared the number of clonally-expanded and unique amplifications in 16,178 single cells from 8 human triple-negative breast cancers (TNBC) and 4 cell lines. The inventors observed that the number of unique amplifications detected at single cell resolution was consistently higher than the number of clonally expanded amplifications at bulk resolution (Figure 8B). In addition, the fraction of samples with at least one detected amplification at the single cell level was significantly higher than the fraction seen across 12 matched pseudobulk samples and 145 unmatched triple-negative breast cancer samples as part of The Cancer Genome Atlas (TCGA) project (p-value < 2.2e-16, Binomial test; Figure 8C). The inventors observed unique amplifications generated during recent karyotypic diversification in 1 1 out of 12 samples (Figure 8C) and clonally-expanded amplifications in 7 out of these 1 1 samples. This suggests that amplifications are actively generated in the evolution of most TNBC samples. However, the rate at which new amplifications emerged widely varied across TNBC samples, ranging from 2.7 unique amplifications per 1 ,000 cells in sample TN3 to 13.8 unique amplifications per 1 ,000 cells in TN4. This indicates significant variation in the ongoing activity of amplification-generating mechanisms across TNBC samples. One sample did not show any unique amplifications, suggesting that mechanisms generating amplifications were inoperative in the evolutionary history of this tumour. Overall, these findings highlight that the acquisition of amplifications is an ongoing consequence of karyotypic diversification in most TNBCs, occurring at widely different rates.
[0256] Selection ultimately determines whether these amplifications expand to dominate the tumour population. Notably, 4 samples did not show a clonally expanded amplification despite showing a moderate number of unique amplifications (mean of 7.15 and range from 2.72 to 9.93 unique amplifications per 1 ,000 cells) (Fig. 8C), suggesting that the amplifications had not yet been positively selected. Indeed, if these mechanisms of amplification were active during the previous evolution of the tumour, they must have generated amplifications during karyotypic diversification. Their absence at the bulk level suggests that, in such a scenario, the generated amplifications did not confer a positive advantage or were indeed negatively selected.
[0257] To formally assess the extent to which selection drives the expansion of amplifications generated during karyotypic diversification, the inventors designed a positive selection test for amplifications (see Methods). They reasoned that positively selected amplifications should predominantly amplify oncogenes, contrary to passenger amplifications. They applied this test to unique amplifications, finding that they did not enrich breast cancer oncogenes (Figure 8B). By contrast, clonal amplifications were enriched in breast cancer oncogenes, suggesting that they were the product of positive selection (Figure 8B). This underscores the role of selection in constraining the initial pool of amplifications generated during karyotypic diversification, resulting in a landscape of clonally -expanded amplifications enriched in oncogenes. The absence of detectable positive selection in unique amplifications enabled the inventors to exclude selection as a confounder when studying the genome-wide activity of mechanisms of amplification. The inventors reasoned that if these mechanisms were to operate uniformly across the genome, then we would see an even distribution of unique amplifications across a chromosome (i.e. nearest unique amplifications would be at a distance equal to the length of the chromosome divided by the total number of unique amplifications observed in the chromosome). However, across all 12 samples they observed that unique amplifications clustered together, with distances between adjacent unique amplifications being significantly shorter than expected (Fig. 8D). This suggests that amplification rate varies across genomic regions, potentially depending on the active mechanisms of amplification within a sample.
[0258] To understand how these evolutionary factors are intertwined in determining how likely a specific oncogene amplification is to arise, the inventors simulated tumour evolution using different selection coefficients for the oncogene amplification and rates at which it is acquired in the tumour population (Figure 8E). The latter depends both on the activity of the mechanisms of amplification and how frequently these mechanisms operate on the locus harbouring the oncogene. Our simulations revealed that the widely varying rates of amplification acquisition observed during karyotypic diversification in TNBCs could significantly affect the likelihood of a new amplification emerging and clonally expanding (Figure 8F). This implies that tumours with driver amplifications should have higher numbers of passenger amplifications. Indeed, TCGA triple-negative breast cancers with a greater number of amplified oncogenes showed a higher number of passenger amplifications (p-value = 0.001 , Jonckheere’s trend test; Figure 8F). Selection also had a notable effect, with high selection coefficients promoting the clonal expansion of amplifications even when their rate of acquisition was lower (Figure 8F). These findings emphasise the benefits of jointly modelling the varied selective coefficients and amplification rates to accurately forecast oncogene amplification.
[0259] Taken together, these results support a model whereby certain types of CIN can generate amplifications at random across the genome and only when one of these amplifications spans an oncogene does it become fixed in the population. Thus, the data indicates that it should be feasible to forecast the likely acquisition of an oncogene using the CIN types found in a tumour. Example 3 - Determining the CIN processes underlying oncogene amplification
[0260] The inventors computed the contribution of each CIN signature to the amplification of each oncogene by comparing the signature scores assigned to the amplified segment containing the oncogene to the signature scores for all other segments (described in detail below). This, in effect, provides a way to determine which CIN types likely caused the amplification of a given oncogene. This was performed using 6335 genome-wide copy number profiles from 33 cancer types derived from SNP6 array data from the TCGA. To create a list of oncogenes, the inventors integrated and curated the annotations provided by Bailey et al., 2018; Jassal et al., 2020; Martinez-Jimenez et al., 2020; and Sondka et al., 2018. This was performed for single tumour types when sufficient data was available, and across multiple tumour types otherwise.
[0261] Here, the inventors aim was to estimate the mutational processes underpinning oncogene amplifications. To this end, the inventors used the novel copy number encoding feature space capable of discerning local activity of CIN-related mutational processes (see Example 1).
[0262] After establishing the new feature encoding the inventors aimed to use this encoding to observe biases in the size, changepoint and breakpoint density of oncogene amplifications observed in TCGA. To do this, they computed the posterior probability 1x41 component vector for each oncogene amplification segment. Then, they combined probability vectors from all samples carrying the oncogene amplification in an oncogene-by-component probabilities matrix, which was then normalised per feature. They finally averaged feature component probabilities to compute an oncogene-specific vector of feature component probabilities per tumour type. All oncogene-specific vectors were then combined to generate the “oncogene by copy number feature” matrix per tumour type.
[0263] The inventors then combined the “oncogene by copy number features” and the “CIN signature definition map” (see Example 1 ) to derive the “local mutational processes” term. This combination procedure essentially quantifies the enrichment of signatures assigned to amplified segments harbouring a given oncogene over a signature background (described in detail below). Therefore, to compute this signature background, they dissected the copy number configurations, i.e. the posterior probability 1x41 component vector, for all copy number events (e) in the tumour-specific TCGA cohort.
[0264] For each sample (s), the inventors defined the event (e) by component (c) probabilities matrix <D(S), with elements <^e,c. The matrix d>(s)was then multiplied by the signature (m) by component (c) definition matrix ct, with elements oc,m (see Example 1). The resulting event-by-signature matrix was then multiplied by the sample-level signature activities CXm(s)(see Example 1) to compute activity-corrected signature scores for each copy number event (e), which was finally normalised per signature (m). Activity-corrected event-by-signature scores of all samples were finally joined into a single matrix E: Equation (8)
[0265] Within a tumour-type: In a first approach, an oncogene was considered a putatively amplified oncogene in a tumour type when 1 ) it was amplified in > 5 samples in the tumour-specific cohort, 2) its amplification frequency in the tumour-specific cohort was > 5%, and 3) it was previously identified as significantly amplified in the tumour type by GISTIC (Mermel, C. H. et al. 2011). Given an amplified oncogene, the inventors calculated the relative contribution of a Cl N signature to the oncogene amplification for a given tumour type (goG.tm) by averaging the activity-corrected CIN signature score of all amplified events spanning a given oncogene in a specific tumour type as follows: Equation (2) where ea>mrepresents the activity-corrected score of one of the 9 universal CIN signatures (m) for a particular amplified event (a) (see section “Event-by-signature scores matrix” and Equation (8) above, where ea>mis the value of £e,mfor a particular amplification event (a) and signature (m)); and .OG represents the number of amplified segments spanning a particular oncogene (OG) in a given tumour type (f).
[0266] Next, the inventors computed the relative contribution of a CIN signature to the acquisition of any copy number event in any tumour type (qt ,m) by averaging the activity-corrected CIN signature scores of all copy number events found in the entire TCGA cohort as follows: Equation (3a) where ep mrepresents the activity-corrected score of one of the 9 universal CIN signatures (m) for a particular copy number event (p) (see section “Event-by-signature scores matrix”); and P represents the total number of copy number events observed in TCGA. Finally, the inventors computed the tumourspecific oncogene-specific CIN signature weights (<woG,tm) by normalising the relative contribution of a CIN signature to the oncogene amplification in a given tumour type (qoG .m) to the relative contribution of a CIN signature to the acquisition of any copy number event in any tumour type (gm). Note that he normalisation term qmis obtained across all tumour types in TCGA whereas the denominator qoa.t.m is tumour type specific. The inventors found that this approach better captures the expected contribution of each signature to a copy-number alteration, hence being a better representation of the “background” to use for correction. These weights represent the relative contribution of each CIN signature to an oncogene amplification. Equation (4a)
[0267] In total, the inventors carried out this process across 19 tumour types and 40 different oncogenes resulting in 86 tumour-specific oncogene-specific weight vectors. These are listed in Table 7a.
[0268] In a slightly modified approach (used in Examples 10 onwards), the inventors computed the relative contribution of a CIN signature to the acquisition of any copy number event in a specific tumour type (qt,m) by averaging the activity-corrected CIN signature scores of all copy number events found in the tumour-specific training cohort as follows: Equation (3b) where £p mrepresents the activity-corrected score of one of the 9 CIN signatures (m) for a particular copy number event (p), and Et represents the total number of copy number events observed in the training cohort of a given tumour type (t). They then generated the “local mutational processes” term by computing the tumour-specific oncogene-specific CIN signature weights (<woG,tm). These weights were derived by normalising the relative contribution of a CIN signature to the oncogene amplification in a given tumour type (qoG,t,m) to the relative contribution of a CIN signature to the acquisition of any copy number event in any tumour type (qt,m), which were finally normalised per signature (m) as follows: Equation (4b)
[0269] For each tumour type, the inventors constructed a “local mutational processes” matrix containing the weights of the 9 universal CIN signatures for 271 oncogenes (<woG,tm). To enhance predictive capacity for forecasting oncogenes in a given tumour type, the inventors integrated signature weights learnt from three different publicly available cohorts: the TCGA, the PCAWG and the HMF dataset. This integration helps to mitigate for the influence of the training cohort on learning framework parameters. In the integration process, the inventors initially assessed the predictive capacity of the framework in the three cohorts using a retrospective approach (see section “Testing the predictive capacity of the forecasting framework”). Subsequently, for each tumour type, they selected the oncogene-specific signature weights from the cohort with the highest predictive capacity based on Youden’s index. In cases where the amplification of an oncogene was not putatively under selection in a specific tumour type (and therefore not tested for predictive capacity), they selected the oncogene-specific signature weights from the cohort that has at least > 5 samples displaying the oncogene amplification.
[0270] Computing the enrichment of CIN signatures in oncogene amplifications (signature weights) allowed the inventors to link the observed oncogene amplifications directly back to putative causal CIN signatures. They applied the procedure for computing the oncogene-specific signature weights to 6,335 genome-wide copy number profiles from 33 cancer types derived from SNP6 array data from the TCGA. From a total of 271 oncogenes included in their list, they limited this exploration analysis to those oncogenes that are putatively selected to be amplified in a given tumour type. An oncogene was considered a putatively amplified oncogene in a tumour type when: 1 ) it was amplified in > 5 samples in the TCGA tumour-specific cohort; and 2) it was previously identified as significantly amplified in the tumour type by GISTIC2.0 In total, they computed signature weight vectors for 257 oncogene / tumour type pairs - see Table 7b for examples (note the signature vectors in Table 7b also include the fitness effect and local DNA topology effects calculated as explained in Example 10). The inventors used the Heatmap function from ComplexHeatmap to illustrate the contribution of CIN signatures to oncogene amplifications across 21 tumour types and 95 different oncogenes (Figure 10 and Figure 11). Hierarchical clustering (using the hclust function from the stats package in R) was used for grouping tumour types without forcing the number of clusters. The inventors observed some oncogenes with similar CIN signatures weights pan-cancer (i.e. CDK4 or CCND1), while others showed more diverse patterns across tumour types (i.e. MYC or PDGFRA), likely due to cell-of-origin differences in replicative, transcriptional, repair and chromatin dynamics. Interestingly, they observed both shared and variable patterns of signature enrichment across tumour types for given oncogenes (Table 7b). For example, most tumours with EGFR amplification showed an enrichment of CX13; however, lung adenocarcinomas showed a distinct enrichment of CX8, a signature putatively linked with R-loop accumulation (Figure 11B). They also observed shared and variable enrichment patterns across genes within a tumour type (Table 7b). For instance, in prostate cancer, amplification of AR and NCOA2 were enriched for CX8. As the amplification of both oncogenes is linked to resistance to ADT, the mechanisms underpinning CX8 may enable multiple paths to resistance (Figure 11A). In contrast, amplification of MYC in prostate cancer, also known to modulate AR signalling, was enriched for CX9 (Table 7b), suggesting multiple DNA damage mechanisms can drive resistance in the disease. These results show that the DNA topology maps may be useful for exploring amplification aetiology independent of their use in the forecasting framework.
[0271] Table 7a. Tumour type specific signature weights (Equation (4a)).
[0272]
[0273] Table 7a (continued). Tumour type specific signature weights (Equation (4a)).
[0274] Table 7b. Tumour type specific signature weights (Equation (4b)) for selected oncogenes. Weights can be calculated as explained above from the freely accessible data in TCGA for all other cancer types and oncogenes. One row per oncogene, data provided as CX1 I CX2 I CX3 I CX4 I CX5 I CX8 I CX9 I CX10 / CX13 / ttype / oncogene. Ill separates two tumour types.
[0275] Across tumour types: if there are insufficient numbers of samples with an oncogene amplification for a given tumour type (i.e. less than 5 samples with the oncogene amplification in that tumour type or a frequency of oncogene amplification lower than 5% in that tumour type), it is possible to determine the
[0276] CIN signatures contributing to an oncogene amplification across tumour types. Here, the inventors followed a similar procedure as above but the training dataset consisted of samples from different tumour types. Only tumour types in which the oncogene demonstrated recurrent amplification (>5 samples with the oncogene amplification) were included to calculate oncogene-specific weights. Unlike in the context of a specific tumour, estimating pan-cancer CIN signature weights must capture features impacting on the acquisition (i.e. genomic architecture) and expansion (i.e. selective advantage) of oncogene amplifications across tumour types. Hence, the inventors included a factor to assign greater weight to the tumour types where the amplification is more frequently observed. The relative contribution of a CIN signature to the oncogene amplification pan-cancer (goG.m) was therefore computed as follow: Equation (5a) where qoa,t,m represents the relative contribution of a CIN signature (m) to the amplification of a given oncogene (OG) in a tumour type (t) (equation (2)); foG.t represents the empirical frequency of the oncogene amplification (OG) in each tumour type (t); and T represents all tumourtypes used for building the pan-cancer model. Thus, qoG.m is a weighted sum of the tumour type specific factors qoG.t.m, where the weight is proportional to the frequency at which the oncogene amplification is observed in the cohort in the tumour type. Finally, the pan-cancer CIN signature weight per oncogene (6>oG,m) was computed by normalising the pan-cancer relative CIN signature contribution to the oncogene amplification (goG.m) by the relative contribution of a CIN signature to the acquisition of any copy number event in any tumour type (gm, equation (3)): Equation (6)
[0277] The inventors estimated pan-cancer weights using equations (5a)-(6) for a total of 187 oncogenes. The oncogenes and signature weights are provided below in Table 8a.
[0278]
[0279] Table 8a. Pancancer signature weights.
[0280] Table 8a (continued). Pancancer signature weights.
[0281] The inventors estimated pan-cancer weights using the 4 factors model Equation (5b) for a total of 96 oncogenes. The oncogenes and signature weights are provided below in Table 8b. The weights in Table 8a / weights obtained using Equations (5a)-(6)) are trained using all samples across tumour types. The weights in Table 8b / weights obtained using equation (5b) are obtained by training for each tumour type individually and taking an average.
[0282] Example 4 - Forecasting focal amplifications
[0283] The most significant hurdle in precision medicine lies in the dynamic nature of tumours, evolving from benign to malignant, acquiring metastatic potential, and ultimately developing resistance to specific therapies. The amplification of oncogenes can play a role in all of these tumour progression steps and therefore forecasting the amplification of a relevant oncogene in an individual tumour can allow us to design tailored intervention strategies.
[0284] Classical evolutionary theory, together with the findings presented in this work (see Example 2), suggest that in many instances the acquisition of an amplification in a particular oncogene requires that a single cell acquires the oncogene amplification, which, if positively selected or hitchhiking a clonal expansion, will eventually dominate the tumour population. By dissecting the evolutionary forces of mutation and selection, the inventors identified three different major factors usually required for the amplification of an oncogene (Figure 12D). First, an amplification generating process must be active in a tumour (process specific factor); second, this amplification generating process must be operating on the genomic region containing the oncogene of interest; finally, the amplification of the oncogene of interest must confer a selective advantage (oncogene amplification specific factors). Thus, by jointly modelling these 3 different factors, the inventors hypothesised that it should be feasible to forecast an oncogene amplification. The inventors developed a model to predict the risk of acquiring a future amplification in a tumour sample (referred to herein as FOCAS - Forecasting OnCogenic Amplifications). This model is made up of two major components, which capture the process and oncogene amplification specific factors discussed above (Figure 12A, B, C):
[0285] 1) the CIN signature activities for a given tumour (computed as described in Example 1 , section “Sample-level signature quantification”), which represents the different CIN types operating in a tumour;
[0286] 2) the oncogene amplification CIN signature weights (computed as described in Example 3), which represents the amplification generating process that is operating on the genomic region containing the oncogene of interest. (Note; these weights also encode a degree of positive selection pressure which is discussed further below).
[0287] In this model, the inventors only use the 9 universal signatures (CX1 -5, CX8-10 and CX13) in order to be able to apply the model across all sample types and technologies. Specifically, the signature-based predictor combines the two types of factors captured by the activities and oncogene amplification CIN signature weights within a linear model to output a score reflecting the likelihood of acquiring the amplification of the given oncogene in the future: Equation (7a) where m refers to one of the 9 universal CIN signatures, CX represents the sample-level signature activities (i.e. CXmis the sample-level activity of signature m in a sample), and com represents the signature weight for signature m computed for a given oncogene in a pan-cancer or tumour-type context (see section “Determining the CIN processes underlying oncogene amplification”). When tumour specific weights available for a particular sample type and oncogene, these are used for prediction. A pan-cancer model is used when no tumour-specific model is available.
[0288] By applying a predetermined threshold, this score can be transformed into a binary classifier to explicitly forecast whether a tumour sample is at high risk or low risk of acquiring the amplification of the oncogene of interest. Given that each oncogene has a different prior probability of being amplified and selected depending on the disease context, each oncogene and cohort is associated with a respective appropriate threshold matched to the clinical and / or biological context. The procedure for setting the optimal threshold for a given oncogene in a specific context will also depend on the availability of a reference cohort with similar features than the sample to be forecasted. This is discussed further below.
[0289] While this prediction model does not explicitly contain a term representing selection pressures, the approach for determining the oncogene signature weights (see Example 3) indirectly captures selection pressure signals: by using only examples of recurrently amplified oncogenes within tumour types for estimating the M term in the model, the resulting weights encapsulated only amplification properties which have already been subjected to positive selection.
[0290] To further ensure that the forecasting model for an oncogene is only used in a context where a similar selective pressure is present, the inventors further designed a set of guidelines for setting the optimal predetermined threshold and setting the clinical context of interest. Indeed, accurate forecasting of an oncogene is dependent on applying the forecasting model to the correct patient cohort. Ideally, the patient cohort should be homogenous with respect to selection pressures applied to the oncogene. Selecting the optimal threshold for the amplification score is dependent on the prior probability of acquiring an amplification, which is typically cohort specific.
[0291] Therefore the inventors developed a set of guidelines for applying the oncogene forecasting models to patient samples, summarised in Figure 13. Note that these guidelines are applicable to any of the forecasting methods described herein, including e.g. the forecasting of oncogene amplification using the methods described in Example 10 and the forecasting of tumour suppressor deletions described in Example 11 . Thus, references to “oncogene” and “amplification” below should be taken to refer equally to “tumour suppressor” and “deletion” unless context indicates otherwise. To apply a signature-based model for forecasting oncogene amplification (of tumour suppressor deletion, see Example 11 below) to a new tumour sample when a reference cohort is available, the inventors recommended to adhere to the following guidelines:
[0292] 1. Define a matched reference cohort: An adequately matched reference cohort is ideal for accurately estimating oncogene signature weights MOG and defining the optimal prediction threshold. The new sample preferably has similar clinical and pathological traits to the reference cohort (such as e.g. a cohort of samples with the same tumour type, similar tumour grade / stage, treated with the same treatment, having the same driver mutations, etc.). The sequencing technology employed for the new sample preferably matches that used for the reference cohort, and the copy number profiles preferably have been derived using the same algorithm and resolution.
[0293] 2. Estimate oncogene signature weights OG- The reference cohort is used to estimate the oncogene signature weights.
[0294] 3. Define an optimal threshold: The threshold on the amplification score for determining whether a sample has a high or low risk of future amplification is dependent on the clinical context. If prior knowledge is available, this is advantageously used to prioritise sensitivity or specificity and an appropriate threshold selected. For example, if the % of samples that ultimately acquire the amplification in this population is known, then the threshold ca be set such that the top x% of samples with highest score (where x is the % of samples that are expected to acquire the amplification in the population) are classified as positives (likely to acquire the amplification). If no prior knowledge is available, then the Youden’s index in the reference cohort can be used as a threshold. In particular, the Youden’s index can be calculated to identify a threshold that has good sensitivity and specificity in identifying samples from the reference cohort that have an amplification (score above the threshold) vs those that do not (score below the threshold). The Youden’s index is J=(TP / (TP+FN))+(TN / (TN+FP))-1 , where TP=true positives, FN=false negatives, TN=true negatives, FP=false positives, where positives are samples identified as likely to have the amplification, and negatives are samples identified as not having the amplification. Note that other approaches such as maximising sensitivity at a chosen specificity or vice-versa, may be used depending on the circumstances. 4. Calculate amplification score and threshold for new sample: After acquiring the CIN signature activities of the new sample of interest, calculate the amplification score, and then apply the optimal threshold.
[0295] When a comparable reference cohort is not available, the recommendations for forecasting oncogene amplification can depend on the clinical context (i.e. whether the sample is a primary untreated tumour or a treated tumour).
[0296] For primary, untreated tumours:
[0297] 1 . Estimate oncogene signature weights MOG'- Use a suitable default cohort (e.g. TCGA / PCAWG) data to estimate the oncogene signature weights. Compute these oncogene amplification CIN signature weights for the specific tumour type only. As these weights are established using TCGA / PCAWG data, which primarily encompasses primary tumours, the prediction models are best suited to forecast oncogene amplification linked to the initiation and progression of a specific tumour type.
[0298] 2. Define an optimal threshold: The threshold on the amplification score for determining whether a sample has a high or low risk of future amplification is dependent on the clinical context. If prior knowledge is available, this should be used to prioritise sensitivity or specificity and an appropriate threshold selected. For example, if it is known that approximately 30% of patients in the cohort will likely acquire an amplification, the threshold can be set to classify the 30% of samples with the higher scores as “positives” (i.e. samples likely to acquire the amplification). If no prior knowledge is available, then the frequency at which the amplification is found in the default cohort (e.g. TCGA / PCAWG cohort) should be used as the threshold. In other words, if the % of samples likely to acquire the amplification is not known for the particular population to which the sample belongs, the reference cohort is used to set the threshold, in a similar way as explained above.
[0299] 3. Calculate amplification score and threshold for new sample: After acquiring the CIN signature activities of the new sample of interest, calculate the amplification score, and then apply the optimal threshold.
[0300] For treated and / or advanced tumours:
[0301] 1. Estimate oncogene signature weights MGG'- In this scenario, the oncogene of interest is not expected to be recurrently amplified in the TCGA / PCAWG tumour-specific cohort as it may not be a critical event during the initial stages of tumour development. Ideally, for training oncogene-specific models, the training dataset would consist of samples at the same tumour stage and / or treated with the same therapy. The absence of such a comparable cohort can potentially limit the ability of the model to capture the specific selection pressure present in the samples. To address this limitation, an alternative approach is to compute oncogene amplification CIN signature weights on a pan-cancer level, but only considering those tumour types in which this oncogene demonstrates recurrent amplification. A default cohort (e.g .TCGA / PCAWG data) can be used to estimate the oncogene signature weights, in particular for all tumour types where an amplification is observed. 2. Define an optimal threshold: The threshold on the amplification score for determining whether a sample has a high or low risk of future amplification is dependent on the clinical context. If prior knowledge is available, this can be used to prioritise sensitivity or specificity and an appropriate threshold selected. If no prior knowledge is available, then the frequency at which the amplification is found in the TCGA / PCAWG cohort can be used as the threshold.
[0302] 3. Calculate amplification score and threshold for new sample: After acquiring the CIN signature activities of the new sample of interest, calculate the amplification score, and then apply the optimal threshold.
[0303] This process is the same as that explained above for primary untreated tumours except that the oncogene signature weights are estimated from the default reference cohort (e.g. TCGA / PCAWG) using a plurality of tumour types in which the oncogene is recurrently amplified (rather than a single tumour type matched to the patient). This is because the default reference cohort used here is composed primarily of untreated samples in which the amplification frequency is likely to be lower than in a treated / advanced tumour population.
[0304] Example 5 - Predictor Performance Assessment
[0305] Testing the predictive capacity of each oncogene forecasting model
[0306] The inventors assessed the predictive capacity of all 271 forecasting models (86 tumour-specific and 185 pan-cancer) by applying them back to the training cohort after removing the oncogene amplifications and adjacent segments prior to computing the sample-level CIN signature activities from each copy number profile. They used a 10-fold cross-validation strategy with stratified sampling, selecting 80% of samples to determine the optimal threshold for transforming the amplification score in a binary classifier, then applying that threshold to forecast amplifications in the remaining 20%. The optimal threshold was selected using the Youden’s index which is defined as the cut point on the receiver operating characteristic (ROC) curve which maximises the sensitivity and specificity. After applying the classifier to the remaining 20%, the Youden’s index was also computed for assessing the predictive capacity of the models across each fold. The inventors ranked all models according to the average of Youden’s index across runs. Out of all prediction models, 60 (29 tumour-specific and 31 pan-cancer) showed a low prediction capacity (Youden’s index < 0.25) and therefore they were excluded from our collection of signature-based prediction models. These tended to be cases where there are few examples of tumours with amplifications. Further, it is possible that some amplifications simply may not arise due a specific mutational process, rather they are more stochastic or arise via mechanisms not captured by the CIN signatures. It is therefore possible that predictive models for any of these oncogenes could be obtained using the methods described herein in specific and / or larger cohorts of samples. Nevertheless, the data shows that even with these limitations it was possible to obtain a model with medium or high predictive capacity for the vast majority of oncogenes (211 models of the 271 models tested had medium or high predictive capacity. These are listed in Table 9a. The parameters of the models are listed in Tables 7a and 8a, although notes that specific parameters to be applied to a sample may vary since as explained above, a reference cohort specifically matched to the sample is preferably used. Indeed, among the remaining 211 models with predictive signal, 6 of them (4 pan-cancer - P0LD1 , RBF0X2, JUN and EWSR1 - and 2 tumour-specific - CDK4 and MDM2 amplifications in sarcoma) exhibited an excellent prediction capacity, with Youden's indexes higher than 0.75. Table 9a.
[0307] Models with medium and high predictive capacity (high predictive capacity indicated with *). Pan=pan cancer.
[0308] SARC=sarcoma, BRCA=breast cancer, LUAD=lung adenocarcinoma, OV=ovarian cancer,
[0309] BLCA=bladder carcinoma, GBM=glioblastoma, HNSC=head and neck squamous cell carcinoma, READ=renal adenocarcinoma, LUSC=lung squamous cell carcinoma, ACC=adenocarcinoma, STAD=stomach adenocarcinoma, ESCA=esophagal carcinoma, LIHC=liver hepatocellular cariconma, CESC= Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma, UCEC= Uterine Corpus Endometrial Carcinoma, COAD= Colon Adenocarcinoma, LGG= Brain Lower Grade Glioma.
[0310] Prospective assessment of predictor performance
[0311] To further assess performance in a prospective setting, the inventors used data from two different cohorts: i) TRACERx Lung, including 126 patients with lung adenocarcinoma and / or non-small cell lung cancer with paired primary and metastatic samples, and ii) Hartwig Medical Foundation (HMF), an extensive cohort that includes a subset of 227 patients with different tumour types that have biopsies obtained at different time points. To also illustrate how thresholding would perform in these examples, the inventors set the threshold to binarise scores i) such that it would yield the same number of predicted positives as in the training (TCGA) cohort in TRACERx Lung and ii) based on the frequency of amplifications in patients that do not have paired data in the Hartwig Medical Foundation cohort. Bakir et al (2023) previously described in TRACERx Lung the acquisition of new amplifications in the oncogene HIST 1H3B in metastases that were not present in the matched primary tissue. The inventors posited that the HIST 1H3B model constructed in TCGA should enable to predict which primary tumours went on to acquire the amplification of this oncogene later on in the matched metastases. HIST1H3B model was only available pan-cancer (see Table 8a), due to the limited number of samples with the amplification in primary lung tumours. This pan-cancer model, when applied to primary tumours in TRACERx Lung, accurately predicted which patients went on to acquire this amplification (Figure 14A)). Because multiple primary and metastatic samples were available for some patients, the inventors randomly selected only one sample of each type per patient.
[0312] Acquisition of an amplification of the oncogene AR is a well-established driver of aggressive disease and androgen deprivation therapy in prostate cancer. The inventors posited that the acquisition of AR could be forecasted using the model trained in TCGA. Again, only a pan-cancer model was available for AR oncogene (Table 8a) due to the low number of samples with acquired AR in early primary prostate tumours; a frequency that substantially increases upon treatment and metastasis. The inventors applied this pan-cancer model of AR amplification to the first sample of 48 pairs of metastatic samples from prostate cancer patients in the Hartwig Medical Foundation cohort. The results showed that patients acquiring AR amplification had significantly higher scores according to their model than patients that did not acquire such amplification (Figure 14B).
[0313] Finally, ERBB2 amplification drives tumorigenesis in multiple tumour types, including breast and oesophageal cancer. Because of its relevance across multiple tumour types, the inventors tested the pan-cancer ERBB2 model constructed in TCGA (Table 8a) in 227 pairs from the Hartwig Medical Foundation. Again, the results demonstrated that patients acquiring ERBB2 amplification had significantly higher scores according to their model than patients that did not acquire such amplification (Figure 14C).
[0314] To further assess performance, the inventors used the Glioma Longitudinal Analysis Consortium (GLASS) cohort consisting of 154 diffuse glioma patients that were biopsied at two time points.
[0315] In particular, the inventors downloaded copy number profiles from a total of 504 diffuse glioma samples included in the Glioma Longitudinal Analysis Consortium (GLASS) dataset. After filtering samples with no detectable CIN (<20 CNAs), the dataset included 154 patients with at least one pair of time- separated tumour samples which were used to test the predictor. Only one sample pair per patient was considered in the analysis. If a patient had multiple sample pairs, we prioritised the pair where the early biopsy was obtained from the primary tumour. If this was not feasible, the inventors randomly selected one of the available sample pairs. Copy number profiles were derived from whole genome sequencing data for 36 patients (72 samples, 36 pairs) and from whole exome sequencing data for 118 patients (236 samples, 1 18 pairs) as described in Barthel et al. 2019. For each patient, the inventors computed amplification scores in the early time point biopsy and observed the presence of a new focal amplification in the later time point biopsy (>6 copies with respect to the early sample). Wilcoxon two- sided rank test was used to compare amplification scores between patients with and without a new focal amplification in the later relapse biopsy. Spearman’s rho was also computed to assess the correlation of amplification scores and the number of amplifications acquired in the relapse.
[0316] Thus, the inventors applied the method to the early biopsy and observed the emergence of any amplification in the later relapse biopsy (Figure 15c). Tumour pairs with new amplifications at the second time point had significantly higher scores than pairs without amplifications (Figure 15d). The amplification scores also correlated with the total number of amplifications acquired (data not shown).
[0317] Taken together these results show that the method described herein can predict oncogene amplification across different tumour types.
[0318] Example 6 - Training of the oncogene amplification model using paired longitudinal datasets
[0319] The inventors also demonstrated the potential to use a different framework to train and develop models and test them when pairs of samples from the same patient are available. This procedure consists of i) defining the first and second sample within a pair from a given patient, ii) identifying the mechanisms of CIN causing the amplification of a specific oncogene in the second biopsy (as described above in TCGA, but in this case using the amplifications observed in the second biopsies) and Hi) applying the model to the first biopsy and test whether the predictions match with the observed acquisition of amplification (Figure 17).
[0320] The inventors used this approach with data from two different cohorts: i) TRACERx Lung, and ii) HMF (as described above in Example 5). Using TRACERx Lung pairs, the inventors applied this procedure to metastatic samples to train models of HIST1 H3B, SETBD1 and MDM2. All of these models, when applied to primary samples in TRACERx Lung, predicted with high accuracy patients that went on to acquire the amplification of these oncogenes (TRACERx, 126 pairs: HIST1 H3B AUC=0.89, SETDB1 AUC=0.98, MDM2 AUC=0.83). Using Hartwig Medical Foundation pairs from different tumour types, the inventors applied this procedure to the latest biopsy from each patient to train models of AR and ERBB2 amplification. The application of these models to the earliest biopsy from each patient identified with high accuracy patients that went on to acquire the amplification of these oncogenes (Hartwig, 227 pairs, AR AUC=0.76; ERBB2 AUC=0.94).
[0321] Taken together, these results show that training models of acquisition of amplification using the latest biopsy within each of the pairs in a cohort yields accurate results, hence serving as an alternative accurate framework to obtain these models. Thus, in practice a patient for which a reference cohort with longitudinal data exists, where the patient matches the characteristics of an early time point in the longitudinal, models can be used with parameters estimated using later time points for the reference cohort.
[0322] Example 7 - Robustness Assessment
[0323] Robustness of the signature-based predictor across technologies
[0324] To test the robustness of the signature-based predictor across sequencing technologies, the inventors computed amplification scores in the SNP6- and sWGS-derived copy number profiles from the 478 samples included in both the TCGA and the PCAWG project (see Example 1 , section ‘‘Sample-level signature stability across technologies”). They used a Pearson correlation to compare sample pairs and showed a high correlation of scores obtained from different technologies (r = 0.82, p-value < 2.2e-16).
[0325] Robustness of using the universal set of signatures
[0326] The use of a signature subset for decoding mutational processes operating in tumours could erroneously attribute alterations caused by unrepresented signatures to the signature set. This presents a fundamental challenge in the mutational signatures field, potentially affecting our ability to predict amplifications driven by these unrepresented signatures. Previous attempts to mitigate this effect have involved de novo signature discovery within the cohort of interest and subsequently mapping them back to a compendium, which may identify unmapped signatures. However, this approach is only feasible when there is a sufficiently large cohort for de novo discovery. Unfortunately, this condition is not met for many of the cohorts with oncogene amplifications used here. Therefore, in order to strike a balance and facilitate technology-agnostic forecasting, the inventors decided to exclusively use the universal signatures.
[0327] To evaluate the extent of copy number events that may be erroneously assigned to the set of universal signatures, the inventors applied the mapping approach described in Example 1 using the whole set of CIN signatures and allocated each copy number event to the signature with the highest associated probability. Subsequently, they estimated the rate of events assigned to a CIN signature not included in the universal set. This analysis revealed a misassignment rate of 22.5% (145,554 out of 647,630 events), with most of them (17%) assigned to a signature linked to subclonal changes with unknown aetiology (CX7). The average misassignment rate per sample was 23.3±14.4%.
[0328] To further warrant the use of universal signatures to build the signature-based predictor, the inventors explored the impact of the presence of these misassigned events within copy number profiles on the signature quantification task. They compared activities of universal signatures quantified in TCGA profiles including and excluding copy number events not mapped to the universal signature set using both Pearson’s correlation (average Pearson’s r across signatures of 0.84; Table 10) and cosine similarities (average cosine similarity across signatures of 0.91 ; Table 3).
[0329] Table 10. Concordance of sample-level activities after excluding misassigned events In addition, the inventors tested whether quantifying the universal set of CIN signatures was sufficient for predicting amplifications. To do this, they performed a Pearson’s correlation to compare amplification scores computed by including all 16 signatures (CX6 was excluded since it’s not captured with the new feature encoding) and those obtained using only the 9 universal signatures. Results demonstrated a high correlation between the scores obtained with both signature sets in the TCGA dataset (r = 0.86, p- value < 2.2e-16).
[0330] Finally, as the mutational processes active in a specific cohort may be different from those captured by the 9 universal signatures, the inventors also conducted a comparison between the amplification scores computed using the universal signature set to those derived from a cohort-specific set of signatures. To undertake this analysis, they leveraged the TCGA-BRCA cohort, which has the largest sample size within the TCGA dataset (n=730), as the de novo signature extraction process requires a substantial cohort. This extraction was performed using non-negative matrix factorisation (NMF) on the input matrix as detailed in Drews et al. 2022. A total of 6 breast-specific signatures were identified, all of which were mapped with one of the signatures included in our compendium of CIN signatures (CX2, CX3, CX4, CX7, CX9, and CX14). Results demonstrated a high correlation between the scores obtained using both signature sets in the TCGA-BRCA dataset (r = 0.91 , p-value < 2.2e-16).
[0331] Altogether, these findings show that there is low impact of applying a fixed signature repertoire for sample-level quantification and for forecasting focal amplifications.
[0332] Assessing the threshold for defining focal amplifications
[0333] To select the optimal threshold for defining focal amplifications, the inventors conducted an extensive analysis of the predictor's performance across a range of thresholds following two strategies: 1) they used absolute copy number values (ranging from >6 to >10 copies), and 2) they used copy number values relative to tumour ploidy (ranging from >4 to >8 copies respect to tumour’s ploidy). For each threshold value, the inventors built a pan-cancer, pan-oncogene model for forecasting oncogene amplifications. They then assessed its performance by applying it back to the training cohort using a retrospective strategy, where signature activities were quantified after removing amplified segments harbouring an oncogene. Performance was evaluated by comparing amplification scores between samples with and without oncogene amplifications using Wilcoxon two-sided test, and by computing the area under the receiver operating characteristic curve (AUC) at pan-cancer and pan-oncogene level (Figure 22a, b). Although the inventors observed marginal differences in performance across these different threshold values, higher thresholds tended to show improved performance. However, higher threshold values resulted in fewer amplifications available to train the model, limiting our ability to predict across as many oncogenes. The inventors obtained similar positive performance using the ploidy-aware strategy, which indirectly may take into account whether a whole-genome doubling event has occurred. They therefore chose a threshold of >8 (absolute copy number) as it represented a balance between performance and inclusivity. Example 8 - Demonstrating clinical utility
[0334] To demonstrate the clinical utility of forecasting oncogene amplification the inventors applied the method to two use-cases: predicting probable MET-amplification driven resistance to EGFR inhibitor treatment; and identifying low-grade gliomas likely to become aggressive via CDK4 amplification.
[0335] Predicting amplification-driven drug resistance
[0336] Acquired resistance to EGFR tyrosine kinase inhibitors (EGFR-TKI) ( / .e. osimertinib, erlotinib, gefitinib or afatinib) via MET amplification occurs in approximately 10-30% of patients with NSCLC (Oxnard, G. R. et al., 2018).The inventors applied the signature-based approach to compute the MET amplification scores in a cohort of primary EGFR-mutant non-small cell lung cancers (NSCLC) treated with Osimertinib, using the / WET-specific signature weights for computing the predictions (data not shown). The top 30% of samples based on amplification scores were set as likely to acquire MET amplification driven resistance to EGFR-TKI. This threshold is set based on the expected percentage of MET amplification after EGFR-TKI resistance described in the literature.
[0337] As relapse biopsies were not available in this cohort to confirm acquisition of MET amplification, the inventors carried out long term culture of 4 lung cancer cell lines in the presence of EGFR-TKI and tracked the emergence of MET amplifications in resistant clones.
[0338] The inventors computed MET amplification scores (using the pancancer MET model with weights provided in Table 8a - row “pancan - MET”) in four parental EGER-mutant NSCLC cell lines, including HCC827, HCC4006, PC9 and H1975, and observed the acquisition of MET amplification in the resistant clones emerged after long-term culture in the presence of EGFR-TKI (see section “Experimental data: Human cancer cell lines” for further details). Amplification scores were normalized by the total number of CNAs per cell line. Droplet digital PCR was used to detect MET amplification in both the parental and resistant cell lines, using the ddPCR™ Copy Number Assay: MET, Human (dHsaCP1000038). We used the ddPCR™ Copy Number Assay: AGO1 , Human (dHsaCP2500349) as a reference. Copy number was estimated as the ratio of the concentration of MET with respect to AGO1 using the Bio-Rad QuantaSoft software. MET amplification was absent in the parental culture. An acquisition of MET amplification-driven resistance was considered when resistant clones showed 2-fold increase of the parental value.
[0339] None of the EGFR-mutant NSCLC cell lines (HCC827, HCC4006, PC9 and H1975) had initial MET amplifications. Canonical mechanisms of EGFR-TKI resistance were analysed in each resistant subclone, confirming that none of them acquired EGFR T790M or C797S mutations. MET amplification- driven resistance was observed in 1 1 / 12 resistant clones derived from HCC827, while no MET amplifications were observed in resistant clones from the other lines. The signature-based approach assigned the highest score to the HCC827 parental cell line (Figure 23a).
[0340] Similar results were observed when analysed data from an independent study (Jia et al. 2013), in which 3 out of the 4 EGFR-mutant NSCLC cell lines (HCC827, HCC4006 and PC9) were long-term cultured in the presence of erlotinib until the emergence of resistant clones. The highest amplification score was assigned to the two HCC827 parental cell lines, which were in fact the unique ones acquiring MET amplification as a resistant mechanism to erlotinib.
[0341] The inventors then also explored if the signature-based predictor can dissect heterogenous development of resistance mechanisms. To this, they generated single clones from the HCC827 parental cell line and carried out long term culture in the presence of EGFR-TKI to generate resistant clones. They then evaluated the number of resistant clones that acquired MET amplification and correlated them with the amplification scores computed by the signature-based predictor (Figure 23B). The parental clone with the lowest number of / WET-amplified resistant clones (HCC827_MC2) presented the lowest amplification score. This indicates that the signature-based predictor is able to distinguish subclonal heterogeneity even in a cell line that is well-known and fully-described in the literature to resist EGFR-TKIs via MET amplification.
[0342] Predicting prognosis
[0343] Oncogene amplification plays a fundamental role in cancer initiation and progression. Therefore, forecasting oncogene amplification linked to aggressiveness can help identify early-stage patients with likely poor outcomes. To test this, we focused on CDK4 amplification in low-grade glioma (LGG) patients from the TCGA, where amplification is linked to worse outcomes (Forst et al. 2014; Mirchia et al. 2019).
[0344] The inventors applied their signature-based approach, and more specifically, the low-grade glioma specific model of CDK4 amplification (model weights provided in Table 7a, row ‘‘LGG - CDK4”) to compute the CDK4 amplification scores in 273 low-grade gliomas (LGG) from the TCGA project. The inventors applied a univariate (Kaplan-Meier estimates) and a multivariate method (Cox proportional hazard model) for comparing overall survival between TCGA-LGG patients with and without CDK4 amplification, and also between non-CDK4 amplified patients at high and low risk of acquiring CDK4 amplification in the future based on the signature-based approach. The inventors used the CDK4- specific signature weights for computing amplification scores, and the median value in the CDK4 amplified group was used as threshold for classifying those patients without CDK4 amplification but with different risks of acquiring it in the future. For the Cox proportional hazard models, the inventors included as covariates the fraction of genome altered and age at diagnosis. Age at diagnosis was categorised into four groups by dividing cohort distribution into quantiles.
[0345] As expected, TCGA-LGG patients with CDK4 amplification survived shorter regardless of the fraction of genome altered (a genomic feature previously liked to poor prognosis) and age at diagnosis (HR = 5.48, p-value = 4.81 e-07; Figure 24a, Figure 25 - the latter shows that the amplification scores is associated with the largest hazard ratio, which was also the case when testing variables independently). However, the inventors hypothesised that the non-CDK4 amplified group may in fact include two subsets of patients: one with worse prognosis where CDK4 amplification will likely emerge in the future, and another with longer survival rates where CDK4 amplification is unlikely to be acquired. Classification of the 253 TCGA-LGG patients without CDK4 amplification into high and low-risk according to the signature-based amplification scores yielded a significant difference in overall survival after controlling for the fraction of genome altered and age at diagnosis (HR = 2.56, p-value = 0.017; Figure 24a, Figure 26). In fact, the TCGA-LGG patients without CDK4 amplification that were predicted to likely acquire CDK4 amplification presented similar overall survival to those who already had the amplification (HR = 1.96, p-value = 0.17; Figure 27).
[0346] Considering the limited progress in improving survival outcomes in glioma patients, the inventors then proceeded to explore the potential of integrating the signature-based predictor into the brain tumour risk stratification algorithm (Sepulveda-Sanchez, J. M. et al., 2018; Louis et al. 2016) (i.e. applying both the method described herein and the brain tumour risk stratification algorithm). This algorithm currently classifies patients into three subgroups: i) IDH mutant, ii) IDH wildtype without 1 p / 19q codeletion and iii) IDH wildtype with 1 p / 19q codeletion, from lower to higher risk groups. Importantly, this algorithm does not currently include CDK4 amplification as a molecular feature for classifying patients, probably due to its presence in only a fraction of the rapidly progressing gliomas at time of diagnosis (Richardson et al. 2018). In the analysis the inventors performed in TCGA-LGG, none of the patients included in the subtype with lowest risk (IDH-mutation with 1 p / 19q codeletion) was classified as having high risk of acquiring CDK4 amplification based on the signature-based predictor. In the IDH-mutant subtype, which generally exhibits a more favourable prognosis compared to its IDH-wildtype counterparts, the inventors observed substantial differences in survival rates across TCGA-LGG patients without CDK4 amplification classified as high and low-risk according to our signature-based predictor (HR = 4.62, p- value = 0.058; Figure 24b). Similarly, classification of IDH-wildtype TCGA-LGG patients without CDK4 amplification as likely and unlikely to acquire the oncogene amplification yielded a significant separation in overall survival (HR = 8.80, p-value = 0.021 ; Figure 24b). All these survival analyses were corrected by the fraction of genome altered and age of diagnosis via multivariate Cox proportional hazards model (i.e. adding these variables in the Cox model as independent variables - fraction of genome altered was added as a continuous / numeric variable, and age was included as categorical variable splitting samples in 4 groups according to quantile distribution.) to account for confounding factors. Hence, this predictorbased subclassification provides an opportunity to define tailored follow-up and treatment strategies for glioma patients, even within a particular molecular subtype.
[0347] Example 9 - Pan-oncogene model
[0348] The inventors also constructed a model capturing the potential for amplification of a tumour, not constrained to any particular oncogene. Following the same rationale as the oncogene-specific tumourspecific models and the pan-cancer oncogene-specific models, the predictive model of amplification is a linear model of copy-number signature activities - i.e. the result of adding up a set of multiplications of copy-number signature activities by matched weights. For this model, the weights are the average vector of all the pan-cancer oncogene-specific model weights woG.m , above defined.
[0349] Retrospective assessment of predictor performance
[0350] The inventors first tested the performance of the signature-based predictor of amplification using retrospective data from two different pan-cancer cohorts (TCGA and PCAWG) with copy number profiles derived from different sequencing platforms (SNP6 arrays and WGS, respectively). They hypothesized that amplification scores should be higher in samples with an oncogene amplification. To avoid the circularity of predicting an event already present in the genome with a metric that depends on its structure, the inventors removed amplified segments covering one of the 271 oncogenes included in the list (see Example 3) and both of its adjacent segments.
[0351] Thus, the inventors first applied the predictor back to the training cohort (Figure 15a) computing patient CIN signature activities after removing the signal corresponding to the copy number segment harbouring the oncogene. The inventors compared the amplification scores between samples with and without at least an oncogene amplification using Wilcoxon two-sided rank test. In addition, we also assessed the sensitivity and specificity of our signature-based approach to predict the amplification of a specific oncogene via ROC curves and evaluation of the AUC. This analysis was done only for those oncogenes amplified in a minimum of 10 samples, and the average performance across all oncogenes was also computed. The inventors observed significantly higher scores for samples with an amplification than for those without across 254 oncogenes (p-value < 2.2e-16, Wilcoxon two-sided test; Figure 15b). The mean of the area under the receiver operating characteristic curve (AUC) for specific oncogenes was 0.79 [95%CI: 0.78-0.80] (Figure 16).
[0352] Performance assessment using the same strategy across an independent cohort of 1 ,932 tumours from the Pan-Cancer Analysis of Whole Genomes (PCAWG) project also showed significantly higher scores for samples with an amplification across 168 oncogenes (p-value < 2.2e-16, Wilcoxon two-sided rank test; Figure 15b) and a mean AUC of 0.74 [95% Cl: 0.73-0.75] (Figure 16).
[0353] To illustrate the pan-cancer performance of predicting the amplification of a specific oncogene, the inventors computed the distribution of pan-cancer amplification scores (from the pan-cancer, panoncogene model) in samples with and without an amplification in 10 key oncogenes (FGFR1 , CDK4, EGFR, MYC, ERBB2, CCNE1 , MET, CCND1 , KRAS and SOX17) both in the TCGA (Figure 18) and the PCAWG cohorts (Figure 19). Across both cohorts, they observed significant differences between amplification scores in samples with and without amplification. This supports the accuracy of the model predicting the amplification of a specific oncogene in a pan-cancer manner.
[0354] Furthermore, to illustrate that the accuracy of predicting oncogene amplifications was not driven by differences in amplification rates and CIN signatures across tumour types, the inventors also performed a cancer-specific assessment of the signature-based predictor by averaging prediction performance (AUC values) across all oncogenes by tumour type in both the TCGA (Figure 20A) and the PCAWG cohorts (Figure 20B). Only oncogenes amplified in a minimum of 10 samples per tumour type were included in this analysis. The data show that the area under the ROC curve is consistently well above 0.5 when predicting the amplification of oncogenes in an individual tumour type. This suggests that the method can be applied to predict the amplification of specific oncogenes in a specific tumour type. As an example, MYC ROC curves in all specific tumour types in TCGA and PCAWG consistently show high sensitivities and specificities.
[0355] The presence of large-scale rearrangements in a tumour sample may lead to the amplification of distant genomic regions linked to the amplified oncogene. For that reason, besides removing the amplified segments covering one of the 271 oncogenes and its adjacent segments, the inventors also removed segments connected to them via complex / simple structural variants (SV). These segments connected to the regions with amplification were detected in the PCAWG cohort by Kim et al, Nature Genetics 2020, and available in Supplementary Table S1 of the Kim et al. article. This is to confirm that the performance of the approach in this retrospective scenario is not entirely driven by the amplification of other non-adjacent segments (results on Figure 21). Wilcoxon two-sided rank test showed that the amplification scores were significantly higher even after removing segments connected to amplifications via simple / complex SV. This shows that the signal that enables accurate performance in the retrospective scenario comes from copy-number alterations unrelated to the amplifications that are predicted. This analysis was done only in the subset of 980 PCAWG samples with available amplicon annotations.
[0356] To assess the performance of the pan-oncogene model, the inventors used the Glioma Longitudinal Analysis Consortium (GLASS) cohort consisting of 154 diffuse glioma patients that were biopsied at two time points. This cohort includes multiple patients that acquire many amplifications but that are not restricted to specific oncogenes, hence making it a suitable cohort for the evaluation of the panoncogene model.
[0357] In particular, the inventors downloaded copy number profiles from a total of 504 diffuse glioma samples included in the Glioma Longitudinal Analysis Consortium (GLASS) dataset. After filtering samples with no detectable CIN (<20 CNAs), the dataset included 154 patients with at least one pair of time- separated tumour samples which were used to test the predictor. Only one sample pair per patient was considered in the analysis. If a patient had multiple sample pairs, we prioritised the pair where the early biopsy was obtained from the primary tumour. If this was not feasible, the inventors randomly selected one of the available sample pairs. Copy number profiles were derived from whole genome sequencing data for 36 patients (72 samples, 36 pairs) and from whole exome sequencing data for 118 patients (236 samples, 1 18 pairs) as described in Barthel et al. 2019. For each patient, the inventors computed amplification scores in the early time point biopsy and observed the presence of a new focal amplification in the later time point biopsy (>6 copies with respect to the early sample). Wilcoxon two- sided rank test was used to compare amplification scores between patients with and without a new focal amplification in the later relapse biopsy. Spearman’s rho was also computed to assess the correlation of amplification scores and the number of amplifications acquired in the relapse.
[0358] Thus, the inventors applied the method to the early biopsy and observed the emergence of any amplification in the later relapse biopsy (Figure 15c). Tumour pairs with new amplifications at the second time point had significantly higher scores than pairs without amplifications (Figure 15d). The amplification scores also correlated with the total number of amplifications acquired.
[0359] Example 10 - Forecasting focal amplifications - 4 factors model
[0360] The emergence of oncogene amplification in a tumour depends on a complex interplay between tumour and local DNA damage activity, fitness effects, and DNA topology constraints (see Example 2). Here, the inventors present a predictive framework containing four key terms representing these constraints (Figure 28). Estimating oncogene fitness effects
[0361] To estimate the fitness effect an oncogene amplification may confer on a given tumour type, the inventors relied on the concept that the frequency of recurrent oncogene amplifications observed at a population level can be used as a proxy for fitness effects.
[0362] As above, thehe inventors initially created a list of oncogenes by integrating and curating the annotations provided by Bailey et al. 2018, Jassal et al. 2020, Martinez-Jimenez et al. 2020, and Sondka et al. 2018. This yielded a total of 271 oncogenes. For each tumour type, they selected the oncogenes that were amplified in at least 5 samples of the TCGA tumour-specific cohort. This resulted in a total of 929 oncogene / tumour type pairs, including 22 different tumour types and 188 oncogenes. Then, they confirmed that the amplification of those oncogenes confers a selective advantage in the tumour type by leveraging results obtained by GISTIC2.0 (Mermel, C. H. et al. 201 1 ) applied to 10,844 tumours from the TCGA cohort (portals.broadinstitute.org / tcga / home, data version: “2015-06-01 stddata 2015_04_02 regular peel-off’). From their list of oncogene / tumour type pairs, they only retained those oncogenes located in an amplification peak that were significantly amplified in a specific tumour type (q-value < 0.25 - see Mermel et al. 201 1 ). This resulted in a total of 257 oncogene / tumour type pairs where the amplification may be under positive selection.
[0363] The inventors then generated the elements of a fitness effects matrix f (271 oncogenes by 33 tumour types matrix) by computing the frequency of samples carrying the amplification of a given oncogene per tumour type, as follows: Equation (9a) where foG.t represents the fitness effect of a given oncogene (OG) amplification for a tumour type (f), AoG.t is the number of samples with the amplification of a given oncogene (OG) in the tumour type (t), and Sj represents the total number of samples for a given tumour type (t). The fitness effect of those oncogenes that are not included in the oncogene / tumour type combinations conferring selective advantage was set to 0.
[0364] Estimating oncogene “local mutational processes” effects
[0365] This term reflects the propensity of the different CIN types underlying the amplification of the oncogene of interest in a given tumour type. The estimation of this term is detailed in Example 3 (Equations (2), (3b), (4b)).
[0366] Estimating oncogene DNA topology constraints
[0367] The impact of the local chromatin architecture can affect the configuration of amplifications in a particular genomic location. To explore this, the inventors assumed that the configuration of amplifications caused by an amplification-generating mutational process in a particular genomic location may be influenced by the local chromatin structure. Therefore, through a comprehensive analysis of the landscape of amplifications across multiple samples, the inventors hypothesised they could infer how oncogene amplifications are constrained by DNA topology. To this end, they used their novel copy number encoding feature space to observe biases in the size, changepoint and breakpoint density of oncogene amplifications over a null distribution (described in detail below). For each tumour type (t) and oncogene (OG), this null distribution was estimated by dissecting configurations of non-oncogene (“passenger”) amplifications (posterior probability 1x41 component vector) observed in tumour samples that harbour the amplification of the given oncogene. The inventors averaged the passenger amplification-by-component probabilities matrix over all passenger events to compute a passenger amplification vector <p of feature component probabilities specific for each oncogene / tumour type pair (<POG,(,C).
[0368] For each sample (s) of a given tumour type (t) harbouring the amplification of a specific oncogene (OG), the inventors computed the posterior probability 1x41 component vector of the oncogene amplification segment. They then combined probability vectors from all samples harbouring the oncogene amplification (averaging by component over all oncogene amplified segments in a tumour type for a specific oncogene) in a oncogene-by-component probabilities matrix, which was finally averaged to compute a vector of feature component probabilities per oncogene / tumour type pair (<^OG,(,C). Vectors were then combined to generate the “oncogene by copy number feature” matrix per tumour type (horizontal concatenation).
[0369] The inventors generated the “DNA topology” term by dividing the regional biases constraining the configuration of oncogene amplifications (^OG,(,C) by the regional biases constraining the configuration of passenger amplifications (<POG, (,<?)■ To integrate this term into the forecasting framework, the inventors expressed the matrices in the signature space by multiplying this by the signature-by-component definition matrix (ct), to obtain an oncogene-by-signature matrix as follow: Equation (10)
[0370] Since CX1 is associated with mitotic errors and is likely unaffected by local chromatin structure, the inventors assigned a weight of 1 to CX1 in the resulting oncogene-by-signature matrix and then renormalized per oncogene to sum to...
Claims
CLAIMS1 . A computer-implemented method of predicting whether a tumour is likely to acquire a copy number alteration of a gene, the method comprising: obtaining a tumour copy number profile for the tumour; determining, for each of one or more signatures of chromosomal instability, a metric indicative of the presence of the signature in the tumour copy number profile, wherein a signature of chromosomal instability comprises a set of weights for each of a plurality of components each associated with a copy number feature of a set of copy number features quantified for a plurality of segments in the copy number profile, wherein the weight of a component is indicative of the probability that a mutational process associated with the signature generates a segment having characteristics defined by the component; and determining whether the tumour is likely to acquire the gene copy number alteration based on the output of a model trained to take as input the values of said metrics and produce as output a score indicative of the likelihood that the tumour will acquire the copy number alteration of the gene, wherein the model has been trained using copy number profiles for a plurality of tumours comprising tumours that have the gene copy number alteration and tumours that do not have the gene copy number alteration.
2. The method of claim 1 , wherein the gene is an oncogene, and the copy number alteration is an amplification, or wherein the gene is a tumour suppressor gene, and the copy number alteration is a homozygous deletion.
3. The method of any preceding claim, wherein the model is specific to a single gene, and / or wherein a copy number alteration is an amplification and an amplification is the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number >=6 or a copy number >=4 above tumour ploidy, optionally where the segment has an absolute copy number >=8 if present in an autosome or >=6 if present in a sex chromosome, or wherein a copy number alteration is a homozygous deletion and a gene homozygous deletion is the presence of a segment encompassing the gene in a tumour copy number profile, where the segment has an absolute copy number <=0.5.
4. The method of any preceding claim, wherein the copy number features comprises or consists of: (a) segment size, copy number change-point, and breakpoint count in one or both flanking regions of a predetermined length, optionally wherein the breakpoint count is a summarised value of the breakpoint count in the upstream and downstream flanking regions of a predetermined length around the segment; or (b) segment size, copy number change-point, breakpoint count per predetermined length of sequence, breakpoint count per chromosome arm, and number of segments with oscillating copy number, optionally wherein the predetermined length of sequence is 10Mb, wherein the components define a respective distribution of values of the associated copy number feature, optionally wherein the components are selected from those in Table 1 or those in Table 1 1 , orcorresponding distributions obtaining by fitting mixture models to values for the set of copy number features obtained from copy number profiles from a plurality of tumour samples.
5. The method of any preceding claim, wherein the metric indicative of the presence of a signature in the tumour copy number profile is the activity of the signature in the tumour copy number profile, optionally wherein the activity of one or more signatures in a tumour copy number profile is determined by: obtaining, for each component, the sum of probabilities associated with the respective component for each of a plurality of segments in the copy number profile; and determining a linear combination of one or more signatures comprising the signature that satisfies the equation: PbC ~ E x SbC where E is a vector of size n comprising coefficients Einwhere Ei is the activity of signature i; PbC is a vector of size c, each element in the vector representing a sum of probabilities associated with a respective component; and SbC is a matrix of size c by n, each value representing the weight of a component in a signature / , and the probability associated with a component for a segment is the probability of a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component.
6. The method of any preceding claims, wherein the signatures are selected from those in Table 3 or those in Table 12, optionally wherein the signatures are selected from CX1 , CX2, CX3, CX4, 5, CX8, CX9, CX10, and CX13 defined in Table 3 or Table 12 or corresponding signatures obtained by quantifying the set of copy number features in a plurality of tumour samples, and identifying one or more signatures likely to result in the copy number profiles of the plurality of tumour samples by nonnegative matrix factorisation.
7. The method of any preceding claim, wherein the trained model is a model that produces a score indicative of the likelihood of acquiring the copy number alteration of the gene as a weighted combination of the values of the metrics indicative of the presence of the one or more signatures in the tumour copy number profile, optionally wherein the combination is a linear combination and / or wherein the model produces a score according to the equation: score = 2m=i CXma)mwhere CXmis the activity of signature m in the tumour copy number profile, wmis a signature weight for signature m for the gene, and n is the total number of signatures considered in the model.
8. The method of claim 7, wherein the signature weight &)mis obtained as a product of terms comprising a fitness effect term (foG.t ; frsG.t) and a local mutational processes term (<uoG,t,m; &>rsG,t,m), as a product of terms comprising a fitness effect term (JOG.I), a local mutational processes term ( woG,(,m) and a DNA topology term (XoG.tm)or as a local mutational processes term ( oG.t.m ;< rsG,t,m), or wherein the trained model produces a score according to the following equation for an oncogene amplification prediction:or the following equation for a tumour suppressor gene deletion prediction:ZnCXm* >rsG,t,m * frsG.t m=lof the following equation for an oncogene amplification prediction:or the following equation for an oncogene amplification prediction:or the following equation for a tumour suppressor gene deletion prediction:
9. The method of claim 7 or claim 8, wherein the weights of the model are determined in a tumour typespecific manner using a reference cohort of samples as, for each signature, a ratio of (i) an averaged signature score of all amplified segments spanning the gene in samples of the specific tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment, optionally wherein said metric is the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum is the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures.
10. The method of claim 7 or claim 8, wherein the weights of the model are determined in a tumour type-agnostic manner using a reference cohort of samples as, for each signature, a ratio of (i) a weighted average signature score of all amplified segments spanning the gene in samples of a specific tumour type in the reference cohort, wherein the average signature scores in the specific tumour type are weighted by the normalised empirical frequency of the gene amplification in the respective tumour type in the reference cohort and (ii) an averaged signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment, optionally wherein said metric is the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum is the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures.11 . The method of any of claims 7 to 10, wherein weights determined in a tumour type-specific manner are used when the reference cohort comprises more than a threshold number of samples of the specifictumour type or the frequency of amplification of the gene in the specific tumour type in the reference cohort is higher than a threshold frequency, optionally wherein the threshold number of sample is 5 and the threshold frequency is 5%, and weights determined in a tumour type-agnostic manner are used otherwise.
12. The method of claim 8, wherein the fitness effect terms (foG.t ; frsG.t) and local mutational processes terms (<woG,(,m; <wsG,t,m) estimated across multiple individual tumour types are aggregated to obtain a signature weight wmacross the multiple individual tumour types, optionally wherein aggregation comprise averaging the products of the fitness effect terms (foG.t ; frsG.t) and local mutational processes terms ( woG,(,m; <wreG.(.m), across the multiple individual tumour types by calculating a weight for each signature m as:ypes.
13. The method of any of claims 8 or 12, wherein the local mutational process term (<4>oG (m < rsG,t,m) for a signature m, gene G=OG or TSG and tumour type t estimated by quantifying the enrichment of signature m assigned to copy number altered segments harbouring gene G over a signature background for signature m, optionally wherein the local mutational process term for a copy number alteration is estimated as, for each signature m of a set of n signatures:where G is the specific gene, and the numerator is the relative contribution of signature m to the copy number event of the gene across the cohort of the tumour type t divided by the relative contribution of signature m to any copy number event in the cohort, and the denominator is the sum of the numerator term across the n signatures, optionally wherein term qt,m is the relative contribution of a signature m to the acquisition of any copy number event in tumour type t, computed by averaging the activity-corrected signature scores of all copy number events found in the tumour-specific training cohort and the term qc,t,m is the relative contribution of signature m to the copy number alterations of gene G for tumour type t.
14. The method of any of claims 8, 12 or 13, wherein the DNA topology term (XoG.tm) is estimated for a signature m by determining the ratio of (i) a metric quantifying the local bias for each of the components associated with copy number features of the set of copy number features for which the signature is defined, in segments encompassing amplifications of the gene OG the cohort; and (ii) a corresponding metric quantified for segments corresponding amplification events at a plurality of genes that are not oncogenes, optionally wherein the metrics are the product of: (1) the average across the segments considered, of the posterior probability of the copy number features quantified for the segment belonging to each of the components; and (2) the weights of the signatures.
15. The method of any of claims 8, 12, 13 or 14, wherein the fitness effect term foG.t ; frsG.t ) is estimated for a tumour type t using a cohort of samples of the tumour type, comprising samples with a copy number alteration of the gene and samples without a copy number alteration of the gene, as the frequency of samples in the cohort that carry the copy number alteration of the gene.
16. The method of any preceding claim, wherein the tumour is a primary and / or untreated tumour, or wherein the tumour is a treated and / or advanced tumour and the weights of the model are determined in a tumour type-agnostic manner using for (i) only tumour types in which the gene is recurrently amplified or deleted in the reference cohort, optionally wherein a gene is recurrently amplified or deleted in a tumour type if it is amplified, respectively deleted, in more than 5% of the samples of the tumour type in the reference cohort.
17. The method of any preceding claim, further comprising classifying the tumour between a first class that has a high likelihood of acquiring the gene copy number alteration and a second class that has a low likelihood of acquiring the gene copy number alteration, wherein the tumour is classified in the first class when the output of the model is above a predetermined threshold, wherein the predetermined threshold has been determined by obtaining the output of the model for each of a plurality of samples in a reference cohort with a known expected % of samples with the gene copy number alteration, and selecting the threshold that is such that a % of samples corresponding to the expected % of samples are classified in the first class, or wherein the predetermined threshold has been determined by obtaining the output of the model for each of a plurality of samples in a reference cohort with known gene copy number alteration status, and selecting the threshold that satisfies one or more predetermined criteria that apply to the sensitivity, selectivity or Youden’s index when used to classify the samples between the first and second classes.
18. The method of any preceding claim, wherein the model has been trained using a reference cohort comprising samples that are one or more of: the same tumour type, the tumour grade / stage, undergone the same treatment, have the same driver mutations, have been analysed with the same sequencing technology, as the tumour to be analysed; and / or wherein the tumour is selected from sarcoma, breast cancer, lung adenocarcinoma, ovarian cancer, bladder carcinoma, glioblastoma, head and neck squamous cell carcinoma, renal adenocarcinoma, lung squamous cell carcinoma, adenocarcinoma, stomach adenocarcinoma, esophagal carcinoma, liver hepatocellular carcinoma, Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma, Uterine Corpus Endometrial Carcinoma, Colon Adenocarcinoma, Brain Lower Grade Glioma.
19. A computer-implemented method of predicting whether a tumour is likely to respond to a therapy, the method comprising: predicting whether the tumour is likely to acquire a gene copy number alteration using the method of any preceding claim, wherein the gene copy number alteration is associated with resistance to the therapy, and whether a tumour predicted as likely to acquire the gene copy number alteration is unlikely to respond to the therapy.
20. The method of claim 19, wherein the gene is MET and the therapy is an EGFR inhibitor, optionally wherein the method further comprises recommending a subject with a tumour that is unlikely to respond to the therapy for upfront treatment with both EGFR and MET inhibitors; the gene is BRAF and the therapy is a MEK inhibitor; the gene is androgen receptor (AR), the tumour is a prostate tumour and the therapy is androgen deprivation therapy.21 . A computer-implemented method of providing a prognosis for a subject with a tumour, the method comprising: predicting whether the tumour is likely to acquire a gene copy number alteration using the method of any of claims 1 to 22, wherein the gene copy number alteration is associated with poor prognosis, optionally wherein the gene is CDK4 and the tumour is a low grade glioma.
22. A computer-implemented of providing a model for predicting whether a tumour is likely to acquire a gene copy number alteration, the method comprising: obtaining, a plurality of tumour copy number profiles for a reference cohort of samples, the plurality of tumour copy number profiles comprising copy number profiles that comprise the gene copy number alteration and copy number profiles that do not comprise the gene copy number alteration; and determining, for each of one or more signatures of chromosomal instability: a weight that is a ratio of (i) an average signature score of all amplified segments spanning the gene in samples of a specific tumour type in the reference cohort, optionally wherein the average signature scores in the specific tumour type are weighted by the normalised empirical frequency of the gene amplification in the respective tumour type in the reference cohort and (ii) an average signature scores of all non-diploid segments found in the reference cohort, wherein the signature score for a segment and signature is a metric indicative of association between the signatures of chromosomal instability and the segment, optionally wherein said metric is the sum of the product of probabilities associated with each of the plurality of components of the signature for the segment by the corresponding weight in the signature, the probability for a segment being the probability that a segment with a copy number feature value equal to the value for the segment having been drawn from a distribution associated with the component, or an activity corrected and / or normalised version of said sum, wherein an activity corrected version of said sum the value of the sum multiplied by the activity of the signature in the copy number profile and a normalised version of said sum or activity corrected sum is a value obtaining by dividing the value by the sum of corresponding values for each of the one or more signatures; or the product of (i) a local mutational process term that quantifies the enrichment of signatures assigned to copy number altered segments comprising the gene over a signature background and (ii) a fitness effect term that represents the frequency of plurality of tumour copy number profiles that include the gene copy number alteration in the plurality of tumour copy number profiles.
Citation Information
Patent Citations
Method of characterising a DNA sample
WO2023057392A1
Cited By
Application method of donkey whole genome 60K liquid phase chip
CN120796453A