Genomic sequencing analysis

EP4747880A1Pending Publication Date: 2026-05-27TAILOR BIO LTD +2

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
TAILOR BIO LTD
Filing Date
2024-07-19
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Current methods for analyzing genomic sequencing data from targeted sequencing panels are limited in generating robust genome-wide copy number profiles due to biases in read counts, particularly capture bias and GC content bias, which lead to noise and gaps in the data.

Method used

A computer-implemented method that corrects for biases in read counts by estimating the departure between observed and expected read distributions in genomic bins, using a capture bias score and GC content as variables, without removing off-target reads, thereby generating robust genome-wide copy number profiles.

Benefits of technology

The method provides comparable results to shallow whole-genome sequencing and improves upon existing computational methods by generating high-resolution copy number profiles without the need for a matched normal tissue or noisy data filtering, making it more widely applicable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024070580_30012025_PF_FP_ABST
    Figure EP2024070580_30012025_PF_FP_ABST
Patent Text Reader

Abstract

Methods of analysing targeted genomic sequencing data comprising a plurality of sequencing reads are provided The methods comprise determining an observed read count, from the genomic sequence data, for each of a plurality of genomic bins; and determining a corrected read count for one or more of the plurality of genomic bins, the corrected read count based on the observed read count and one or more estimates of biases in read counts as a function of a variable associated with said respective biases. The biases comprise capture bias and the variable associated with said bias is a capture bias score that is indicative of a difference between an observed read location metric distribution in the bin and an expected read location distribution metric in the bin. Related methods and products are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GENOMIC SEQUENCING ANALYSIS

[0002] FIELD OF THE DISCLOSURE

[0003] The present disclosure relates to methods for analysing genomic sequencing data, and in particular genomic sequence data from targeted sequencing. The present disclosure also relates to methods of obtaining copy number profiles derived from such untargeted sequencing data.

[0004] BACKGROUND

[0005] Genome-wide copy number profiles are essential in cancer genomics as they provide information on the levels of chromosomal instability (CIN) of a tumour, a hallmark of cancer. Since most computational methods for CIN quantification depend on genome-wide copy number profiles, it is essential not only to know the copy number status of certain driver genes, but also the rest of the genome.

[0006] However, whole genome sequencing data is expensive to obtain and therefore capture-based DNA targeted panels are routinely used for small variant detection and for a very limited genelevel copy number profiling in cancer clinical care. This makes use of the information from sequence reads that map to captured regions (on target reads), and typically disregards off- target sequence reads. Kuilman et al. (2015) proposed a method that makes use of off target reads to perform copy number detection, by detecting and removing genomic regions enriched for sequence reads using a peak detection algorithm developed for ChlPseq data. However, the method still leaves undesirable noise in the data and discards data thereby creating gaps in the genome which makes higher resolution copy number calling difficult.

[0007] Thus, there is a need for improved methods to analyse targeted sequencing data.

[0008] SUMMARY

[0009] The present inventors have recognised that capture-based targeted sequencing panels could represent a valuable source of data for analyses that typically rely on genome-wide copy number profiles, because this data also contain information spread over the entire genome. They therefore set out to design a method that can process this data to generate robust genome-wide copy number profiles. Briefly, the method corrects for biases introduced in read counts in off-target regions by the capture process. The correction does not remove reads in off target regions but instead estimates a departure between an observed distribution of the location of reads in a genomic bin and the expected distribution of the location of reads in the genomic bin in the absence of capture bias. The present inventors showed that the new method provides comparable results with shallow whole-genome sequencing (sWGS) samples, the current gold standard, when applied to a variety of next generation sequencing (NGS) protocols. They further benchmark the tool with comparative computation methods and show that the new method improves on these. The method further adds the benefit that it does not require a matched normal tissue or a noisy data filtering step, thereby being significantly more widely applicable and avoiding steps that discard potentially informative data (i.e. preserving more of the information content in the data).

[0010] Thus, according to a first aspect, there is provided a computer-implemented method of analysing genomic sequencing data, the method comprising: obtaining genomic sequence data from one or more samples, comprising a plurality of DNA sequencing reads; determining an observed read count, from the genomic sequence data, for each of a plurality of genomic bins; and determining a corrected read count for one or more of the plurality of genomic bins, the corrected read count based on the observed read count and one or more estimates of biases in read counts as a function of a variable associated with said respective biases. The genomic sequencing data is targeted genomic sequencing data, the biases comprise capture bias and the variable associated with said bias is a capture bias score that is indicative of a difference between an observed read location metric distribution in the bin and an expected read location distribution metric in the bin.

[0011] Also described herein according to a second aspect is a computer-implemented method of analysing genomic sequencing data, the method comprising: obtaining genomic sequence data from one or more samples, comprising a plurality of DNA sequencing reads; determining a raw read count, from the genomic sequence data, for each of a plurality of genomic bins; and determining a corrected read count for at least one of the plurality of genomic bins, the corrected read count based on the raw read count and an estimate of a bias in read counts as a function of a variable associated with said bias. The bias is a bias associated with GO content and the variable associated with said bias is the GO content, and wherein the estimate of GO content bias in read counts as a function of GO content is obtained by: identifying a plurality of sets of bins where all bins in a set have the same estimated copy number, and fitting a model between (i) the observed read counts in one of the sets of bins comprising the bin for which a corrected read count is determined, and (ii) the value of GO content for the bins in the set. The methods according to the second aspect are applicable to any genomic sequencing data, whether from targeted or untargeted genomic sequencing data, and whether from bulk or single cell sequencing data. The methods according to the first and second aspects may have any one or more of the following optional features.

[0012] The method may comprise determining a corrected read count for each of the plurality of genomic bins. The method according to the second aspect may comprise obtaining corrected read counts according to the first aspect. The method according to the first aspect may comprise obtaining corrected read counts according to the second aspect. An estimate of a bias in read counts as a function of a variable associated with said bias may be estimated by fitting a model between (i) the observed read counts in the plurality of genomic bins, and (ii) the value of the bias variable for the plurality of genomic bins. The model may be a curvefitting model. The model may be a regression model, such as a local polynomial regression model. Determining a corrected read count for a bin may comprise subtracting from the observed read count a correction factor that is equal to the difference between: (i) the value of the model at the coordinate corresponding to the value of the bias variable for the bin, and (ii) a reference value. The reference value may be a summary value derived from the observed read counts for the plurality of bins.

[0013] An expected read location metric distribution in the bin may be a read location metric distribution that corresponds to a uniform distribution of locations of reads in the bin. The locations of reads in a bin may be represented by a metric of pairwise distance between pairs of reads in the bin. A capture bias score may be a variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin. A pairwise distance between reads may be a left-most pairwise distance, a right-most pairwise distance or a middle pairwise distance. An expected distribution of pairwise distances between reads in a bin may be a distribution that has the same number of observations arranged in a left triangular distribution with a mode of 0 or the minimum observed pairwise distance in the bin, a minimum of 0 or the minimum observed pairwise distance in the bin and a maximum of the maximum observed pairwise distance in the bin, the length of the bin, the sum of length of the bin and the length of a read (which may be taken as e.g. the average length of a read), or any value between the length of the bin and the sum of length of the bin and the length of a read. The capture bias score may be a variable that summarises across a plurality of windows of values of the read location metric, the absolute difference between the frequency for the window in the observed read location metric distribution of the bin and the frequency for the window in the expected read location metric distribution of the bin. The capture bias score may be a variable that summarises across a plurality of observed values of the read location metric in the bin, the probability that the observed read location metric is sampled from the expected read location metric distribution of the bin. The method may comprise obtaining the plurality of genomic bins by: applying a nonoverlapping fixed width window along genomic coordinates for one or more genomic regions that the reads map to, thereby obtaining a plurality of fixed width bins. The genomic sequence data may comprise on-target reads that are at least partially mapped to captured regions and off target reads that are not mapped to captured regions, and the observed read counts are obtained using the off-target reads. The method may comprise selecting reads that are not mapped to captured regions and obtaining the observed read counts using the selected reads. When using paired-end sequencing data, the off-target reads may be reads that do not map to a captured region, and reads that do not have a mate read that maps to a captured region. The method may comprise: obtaining a processed bin width for a bin by removing any region from a fixed width bin that is a captured region. The method may comprise obtaining a further corrected read count for a bin by adding to the corrected read count an adjustment factor that is proportional to the difference between the processed bin width and the width of a fixed width bin. The adjustment factor may be proportional to the corrected read count and the difference between the processed bin width and the width of a fixed width bin. Alternatively, the method may further comprise: determining a corrected read count for one or more of the plurality of genomic bins comprises determining a corrected read count for a bin comprising one or more on-target reads by: (i) obtaining a first corrected read count based on the observed off target read count for the bin and one or more estimates of biases in read counts as a function of a variable associated with said respective biases, and (ii) obtaining a second corrected reads count by adding a pseudocount to the first corrected read count, wherein the pseudocount is proportional to the first read count and the proportion of the bin that overlaps with a captured region. For example, a pseudocount may be calculated as (%overlap with the captured regions file)*(count for the bin after all corrections applied). The captured regions may be as defined in a BED file.

[0014] The method may further comprise: obtaining GC corrected read counts for the bins by fitting a model between the observed read counts, and (ii) the value of GC content for the bins. The method may further comprise using the GC corrected read counts to determine whether the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins, all bins in a set having the same estimated copy number and the estimate copy numbers associated with each of the plurality of sets being different from each other. Determining whether the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins may comprise: segmenting the GC corrected read counts for the bins to obtain a segment value associated with each bin; and clustering the segment values for the bins, wherein the presence of more than 1 cluster is indicative of the presence of a plurality of sets of bins. Determining whether the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins may comprise: segmenting the GC corrected read counts for the bins to obtain a segment value associated with each bin; and determining whether at least one segment has an estimated copy number different from other segments, wherein the presence of such a segment is indicative of the presence of a plurality of sets of bins. In embodiments of the first aspect, the biases may further comprise GC content bias and the method may further comprise obtaining an estimate of GC content bias in read counts as a function of GC content by: identifying a plurality of sets of bins, all bins in a set having the same estimated copy number, and fitting a model between (i) the observed read counts in one of the sets of bins comprising the bin for which a corrected read count is determined, and (ii) the value of GC content for the bins in the set. In embodiments of any aspect, identifying a plurality of sets of bins where all bins in a set have the same estimated copy number may comprise clustering the raw read counts, corrected read counts or segment read counts associated with the bins. The method may further comprise determining a corrected read count based on the observed read count (optionally already corrected for capture bias, or prior to correcting for capture bias) and an estimate of biases in read counts as a function of GC content). The estimate of GC content bias in read counts as a function of GC content is obtained by: identifying a plurality of sets of bins where all bins in a set have the same estimated copy number, and fitting a model between (i) the observed read counts in one of the sets of bins comprising the bin for which a corrected read count is determined, and (ii) the value of GC content for the bins in the set. This may be performed when it has been determined that the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins. The GC corrected read count for the plurality of bins obtained by fitting a model between the observed read counts (for all of the plurality of bins), and the value of GC content for the bins may be used when it has been determined that the observed read counts for each of the plurality of genomic bins are not indicative of the presence of a plurality of sets of bins.

[0015] Identifying a plurality of sets of bins where all bins in a set have the same estimated copy number may comprise: obtaining GC corrected read counts for the bins in the plurality of sets of bins by fitting a model between the observed read counts in all of the sets of bins together, and (ii) the value of GC content for the bins, and clustering the GC corrected read counts for the bins, or segment read counts or segment copy number associated with the bins obtained from the GC corrected read counts. Thus, the method may comprise segmenting the GC corrected read counts for the bins to obtain segment values. The segment values may be segment read counts or segment copy numbers. Identifying a plurality of sets of bins where all bins in a set have the same estimated copy number may comprise segmenting the (optionally GC corrected) read counts for the bins in the plurality of sets of bins, thereby identifying a plurality of genomic segments, each segment having the same copy number state, and determining a segment read count or copy number for every segment. Identifying a plurality of sets of bins where all bins in a set have the same estimated copy number may comprise clustering values associated with each of the bins in the plurality of sets of bins, wherein the values are, for each bin, a segment read count or segment copy number associated with a segment that overlaps or encompasses the bin. For bins that are overlapped by multiple segments, an average segment read count or copy number over the multiple segments, weighted average segment read count or copy number over the multiple segments (e.g. weighted by the proportion of the bin that is overlapped by each of the multiple segments) or selected segment read count or copy number (e.g. segment read count or copy number that represents the largest proportion of the bin) may be used. Clustering may be performed using any clustering algorithm known in the art, such as e.g. k-means, self-organising maps, hierarchical clustering, etc. The clustering may be performed using a predetermined maximum number of clusters. For example, a predetermined maximum number of clusters may be set to 25. The clustering may be performed using k-means, for example using a Bayesian information criterion to automatically select a number of clusters. The automatically selected number of clusters may be within a predetermined range, such as e.g. between 1 and 25. The clustering may comprise applying a clustering algorithm to identify an initial plurality of clusters, and iteratively removing one or more clusters using one or more criteria selected from: the number of data points in the clusters, and the distance between cluster centres. The process of iteratively removing one or more clusters may be performed by clustering the values associated with each of the bins in the plurality of sets of bins using a maximum number of clusters that is lower than the maximum number of clusters used at the previous iteration (or the initial maximum number of clusters, at the first iteration). The maximum number of clusters may be reduced by 1 when at last one cluster satisfies the one or more criteria. The maximum number of clusters may be reduced by the number of clusters that satisfy the one or more criteria. For example, clusters that have a number of data points below a predetermined threshold, such as e.g. 5, may be selected for combining with another cluster, such as e.g. the closest cluster identified. As another example, clusters that have centres that are less than a predetermined distance (e.g. Euclidian distance) apart may be selected and merged. The predetermined distance may be e.g. 2 in embodiments where the values associated with the bins being clustered are segment copy numbers. As another example, an initial set of clusters may be iteratively refined by: identifying a cluster that satisfies one or more criteria selected from: the number of data points in the cluster being below a predetermined threshold, and the distance between the centre of the cluster and the centre of another cluster being below a predetermined threshold; and if a cluster that satisfies the one or more criteria has been identified, clustering the values associated with each of the bins in the plurality of sets of bins using a maximum number of clusters that is lower than the initial maximum number of clusters used to obtain the initial set of clusters. The resulting set of clusters may then be refined by identifying a cluster that satisfies the one or more criteria, and if a cluster that satisfies the one or more criteria has been identified, clustering the values associated with each of the bins in the plurality of sets of bins using a maximum number of clusters that is lower than the maximum number of clusters used to obtain the set of clusters at the preceding iteration. The method may comprise fitting a plurality of models, each model between (i) the observed read counts in a respective set of bins, and (ii) the value of GC content for the bins in the set. The number of models fitted may be equal to a number of clusters identified by clustering the raw read counts, corrected read counts or segment read counts associated with the bins. Thus, the method may comprise identifying a plurality of clusters of bins by clustering raw read counts, corrected read counts or segment read counts associated with the bins, and fitting a model for each of the plurality of clusters or simultaneously fitting a plurality of models to the read counts for all bins in the plurality of sets of bins, wherein the number of models is equal to the number of clusters identified.

[0016] The targeted sequencing data may be capture-based sequencing data, e.g. from whole exome sequencing or targeted panel sequencing. The targeted sequencing data may be paired end sequencing data. A sample may be a tissue sample, a FFPE tissue sample, a fresh frozen tissue sample, or a liquid biopsy sample. A sample may be a sample comprising tumour cells or genetic material derived therefrom, such as a tumour biopsy or sample comprising cell free DNA. The one or more samples may not comprise a matched normal sample. The one or more samples may all be samples comprising tumour cells or genetic material derived therefrom.

[0017] The biases may further comprise replication timing bias. The variable associated with said bias is replication timing. The method may comprise determining whether the observed read counts for the plurality of bins are significantly associated with replication timing, and correcting the read counts for replication timing bias when the observed read counts for the plurality of bins are significantly associated with replication timing.

[0018] The method may comprise obtaining corrected read counts for each of the plurality of bins. The method may further comprise obtaining a copy number profile from the corrected read counts for the plurality of bins. A copy number profile may be obtained from the corrected read counts for the plurality of bins by segmenting and absolute copy number fitting. The method may further comprise obtaining information derived from the copy number profile. The information may comprise the presence of one or more signatures of chromosomal instability, and / or the percentage of tumour cells or tumour DNA in the sample. As the skilled person understands, the complexity of the operations described herein (due at least to the complexity of aligning sequencing data to reference sequences, and the amount of data that is typically generated by sequencing genetic material) are such that they are beyond the reach of a mental activity. Thus, unless context indicates otherwise (e.g. where sample preparation or acquisition steps are described), all steps of the methods described herein are computer implemented.

[0019] Also described according to a third aspect is a method comprising one or more of: sequencing one or more samples (e.g. using a targeted genomic sequencing assay) thereby obtaining genomic sequence data from the one or more samples, pre-processing the genomic sequence data (e.g. aligning reads, filtering low quality reads, obtaining bins, annotating bins etc.), obtaining the one or more samples (e.g. from a subject, cell line, etc.); and analysing the genomic sequencing data using the method of any embodiment of the preceding aspects.

[0020] The methods of any embodiment of any aspect may comprise providing to a user (e.g. through a user interface) the results of any one or more steps of the methods, such as e.g. one or more of: a corrected read count for one or more bins, the value of a bias variable for one or more bins, a copy number profile for a sample, and information derived from a copy number profile for a sample.

[0021] Also described according to a fourth aspect is a method of obtaining untargeted genomic sequencing data from a sample, the method comprising: obtaining genomic sequence data from the sample, comprising a plurality of DNA sequencing reads obtained using a targeted genomic sequencing assay; and analysing the data using the method of any of embodiment of the first aspect.

[0022] Also described according to a fifth aspect is a method of providing a copy number profile or genomic feature derived therefrom for a sample, the method comprising: obtaining genomic sequence data from the sample, comprising a plurality of DNA sequencing reads obtained using a targeted genomic sequencing assay; analysing the data using the method of any embodiment of any preceding aspect; and obtaining a copy number profile and / or a genomic feature derived therefrom, using the corrected read counts. The method may comprise obtaining a tumour copy number profile and a percentage of tumour DNA or tumour cells with the tumour copy number profile.

[0023] Also described according to a sixth aspect is a method of determining the percentage of tumour DNA or tumour cells in a sample, the method comprising: obtaining genomic sequence data from the sample, comprising a plurality of DNA sequencing reads obtained using a targeted genomic sequencing assay; analysing the data using the method of any embodiment of any preceding aspect; and obtaining a tumour copy number profile and a percentage of tumour DNA or tumour cells with the tumour copy number profile, using the corrected read counts.

[0024] According to a further aspect, there is provided one or more non-transitory computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any embodiment of any preceding aspect.

[0025] According to a further aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform the steps of any method described herein, such as a method according to any embodiment of any of the first to sixth aspects.

[0026] According to a further aspect, there is provided system comprising: a processor; and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any embodiment of any of the first to sixth aspects.

[0027] BRIEF DESCRIPTION OF THE FIGURES

[0028] Figure 1 is a flowchart illustrating schematically a method of obtaining untargeted genomic sequencing data.

[0029] Figure 2 is a flowchart illustrating schematically a method of providing and analysing a genome-wide copy number profile.

[0030] Figure 3 shows an embodiment of a system for analysing targeted DNA sequencing data and / or providing or analysing a genome-wide copy number profile.

[0031] Figure 4 illustrates a method for GC correction or sequencing reads by fitting of a LOESS curve to raw counts data as a function of GC content - GC ratio calculated as (G+C) / (A+T+G+C)%.

[0032] Figure 5 shows a result of a copy number aware GC content correction process as described herein. The data shows that in a tumour genome, the variable copy number states cause a shift in the read depth. This is properly accounted for in embodiments of the methods described herein.

[0033] Figure 6 illustrates the process of GC content correction in a conventional manner (top) and according to an embodiment of the present disclosure (copy number aware).

[0034] Figure 7 illustrates schematically an embodiment of a method according to the present disclosure, as used in the examples. Figure 8A illustrates the calculation of a score for capture bias correction that is a histogram based score.

[0035] Figure 8B illustrates the calculation of a score for capture bias correction that is a probability based score.

[0036] Figure 9A shows a copy number profile obtained using the method as described herein applied to whole exome sequencing (WES) data (top) overlaid with a corresponding copy number profile obtained by shallow whole genome sequencing (sWGS) from the same sample.

[0037] Figure 9B shows a copy number profile obtained using the method as described herein applied to TSO500 data overlaid with a corresponding copy number profile obtained by shallow whole genome sequencing (sWGS) from the same sample.

[0038] Figure 9C shows data illustrating the correction of counts using the methods as described herein and a statistical metric comparing two distributions. The top plot shows the GO corrected counts as a function of the statistical metric, and the bottom plot shows the corresponding counts after correction using the statistical metric.

[0039] Figure 9D shows data illustrating the effects of applying the correction in Figure 90: the resulting corrected copy number profile shows low capture bias effects.

[0040] Figure 10 shows the distribution of Pearson’s r, Manhattan distance, Euclidean distance and Cosine similarity values between copy number profiles obtained using methods of the disclosure applied to WES data from cancer cell lines with or without capture bias correction, and corresponding copy number profiles obtained by sWGS. The data show a clear increase in performance (higher Pearson correlation and cosine similarity, lower Manhattan and Euclidian distance) of the profiles when performing capture bias correction and GO correction using methods of the disclosure than when performing GO correction alone.

[0041] Figure 11 shows the distribution of Pearson’s r when correlating compared to sWGS gold standard of all FFPE clinical ovarian cancer samples included in Table 2. These correlations of the copy number segments against sWGS are performed using 2 different technologies: TS0500 and TS0500+HRD. In addition, it shows the results for TS0500 not only using CopyRight but also two other genome-wide copy number profiling methods publicly available for targeted sequencing panels (CopywriteR and CNVkit). CopyRight demonstrated a better performance than CopywriteR and CNVkit and comparable performance to TS0500+HRD.

[0042] Figure 12 shows a heatmap of CIN signature activities obtained with WES and sWGS from two lung adenocarcinoma cell line pairs. The samples are represented in rows and the CIN signatures in columns. Only signatures which have previously been deemed to be stable across profiling technologies are shown. The heatmap shows a hierarchical clustering applied to the rows based on Euclidean distance.

[0043] Figure 13A,B show comparisons of the performance of extracting CIN signatures from TSO500 (with CopyRight), TSO500+HRD and sWGS using all the clinical FFPE ovarian cancer samples detailed in Table 2. These results demonstrate that TSO500 (with copyRight) performs comparably to TSO500+HRD when extracting CIN signatures, with many samples having cosine similarities > 0.8 compared to sWGS.

[0044] Fig. 13A shows stacked barplots of the raw signature activities.

[0045] Fig. 13B shows the cosine similarities against sWGS of TSG500 and TSG500+HRD.

[0046] DETAILED DESCRIPTION

[0047] In the present disclosure, the following terms will be employed, and are intended to be defined as indicated below.

[0048] The present disclosure provides methods for obtaining untargeted genomic sequencing data, from targeted sequencing data for a sample.

[0049] The term “untargeted genomic sequencing data” refers to data that has been obtained using an untargeted sequencing technology such as whole genome sequencing or shallow whole genome sequencing, or data that is equivalent to that which would have been obtained using such technologies. The data may be in the form of read counts per genomic bin. A genomic bin is a region of the genome. A bin is typically defined by applying a fixed size window to a reference genome. A read count is a number of sequencing reads mapping to a bin. Raw read counts are count of reads as obtained from a sequencing assay, optionally after one or more filtering steps, but before any bias correction. A bias correction may be a GO bias correction, a capture bias correction and / or a replication timing bias correction, as will be explained further below. Thus, in particular the term “untargeted genomic sequencing data” may refer to read counts that have been obtained from targeted sequencing data corrected for at least capture bias. Such data may be equivalent to data that would have been obtained from the same sample using e.g. sWGS. A first and second sequencing datasets may be considered to be equivalent if they result in copy number profiles over a predetermined region (such as e.g. one or more chromosomes, or a whole genome) that have a Pearson’s r of at least 50%, at least 60%, at least 70%, at least 80% or at least 90%.

[0050] Targeted genomic sequencing data refers to data obtained from genomic sequencing assays that capture a plurality of predetermined regions of the genome (also referred to as ’’target regions” or “target loci”). Such assays may be referred to “panel” sequencing, particularly when specific genes of interest or parts thereof are targeted. The methods described herein are applicable to any capture-based targeted sequencing technologies. For example, the methods described herein may be applied to Whole Exome Sequencing (WES), and to gene panel sequencing assays such as oncology panel assays or any panel sequencing assays with custom panel design. Examples of gene panel sequencing assays include the TruSight Oncology 500 assay (Illumina, see Zhao et al. 2020), the TruSight Tumor 170 assay (Heydt et al. 2018), the Tempus xT assay (GTR Test ID: GTR000558436.10, see www.ncbi.nlm.nih.gov / gtr / tests / 558436 / ), FoundationOne CDx (F1CDx, GTR Test ID: GTR000593451.1 , see Milbury et al. 2022 and www.ncbi.nlm.nih.gov / gtr / tests / 593451 / ), FoundationOne Heme (GTR Test ID: GTR000527977.1 , www.ncbi.nlm.nih.gov / gtr / tests / 527977 / ), and FoundationOne (GTR Test ID: GTR000527976.2, www.ncbi.nlm.nih.gov / gtr / tests / 527976 / ). The methods described herein are applicable to any targeted genomic sequencing assay that generates off target reads. This may include any capture-based assay. This may exclude amplicon-based technologies, or any targeted sequencing technologies that do not generate off-target reads. The sequencing assay may use paired-end sequencing. Thus, the targeted genomic sequencing data may comprise paired end reads. Alternatively, the sequencing may use single-end sequencing. The methods described herein are believed to be applicable to both single and paired-end sequencing. Paired-end sequencing is relatively common in gene panel sequencing assays and advantageously reduces mappability issues.

[0051] On-target reads are reads obtained from a targeted genomic sequencing assay (i.e. DNA sequencing reads present in targeted genomic sequence data) that at least partially map to target regions. Off target reads are reads obtained from a targeted genomic sequencing assay that do not map to a target region. The methods described herein may remove on-target reads (or modify read counts accordingly, for example obtaining read counts for genomic bins from which any target region is excluded). The methods described herein may not remove off target reads other than based on quality control metrics inherent to the read rather than the distribution of reads. For example, the methods described herein may not remove off target reads that cause a peak in a read count profile, or entire regions of a read count profile that comprise such peaks. Instead, the methods described herein may correct the read counts in each of one or more genomic bins, based on a deviation between an observed and an expected distribution of a read location metric in the bin. Such a deviation may be referred to herein as a “capture bias score” as will be explained further below.

[0052] The methods described herein may be used to obtain a genome-wide copy number profile for a sample that has been sequenced using a targeted sequencing technology. In other words, the methods described herein may be used to obtain a genome-wide copy number profile from targeted genomic sequence data. The term “genome-wide” refers to data that covers a whole genomes or substantial parts thereof such as e.g. one or more entire chromosomes. As the skilled person understand, sequencing data or a copy number profile derived therefrom may be referred to as “genome wide” even though it may not contain information (i.e. reads or a copy number estimate) for each and every single position in the genome. Indeed, even in the context of sWGS there may be regions of the genome that have no mapping reads due to e.g. low mappability, sequencing difficulty, or the sampling nature of next generation sequencing. Thus, data may be referred to as “genome-wide” if it contains data for at least 70%, at least 80% or at least 90% of one or more chromosomes or all chromosomes present in a reference genome.

[0053] A copy number profile is typically in the form of an estimate, for each of a plurality of genomic segments, of the number of copies of the segment in the sample or a subset of the sample (e.g. in the case of samples comprising cells with different genotypes or genetic material derived therefrom, such as e.g. samples comprising a mixture of tumour and normal cells or tumour and normal DNA). Copy number profiles can be obtained by segmentation and copy number estimation from read count data using methods known in the art. For example, methods such as circular binary segmentation may be used for segmentation. Copy number fitting may be performed as described in the examples below. Alternatively, methods such as Ascat (Raine et al. 2016), ichorCNA (Adalsteinsson et al. 2017) and others may be used to perform segmentation and copy number estimation (And also % tumour DNA estimation, in the case of Ascat and ichorCNA). A copy number profile may be a genome-wide copy number profile. A genome-wide copy number profile obtained using methods described herein from targeted sequencing data may be equivalent to a copy number profile obtained for the same sample using sWGS.

[0054] A “segment” in a copy number profile refers to a portion of a sequence represented in a copy number profile which is associated with a consistent absolute copy number. The consistent copy number is different from that associated with the sequence directly upstream (if such a sequence is present and associated with a copy number estimate) and the sequence directly downstream (if such a sequence is present and associated with a copy number) of said portion. In other words, a segment refers to a portion of sequence that has a copy number associated with it, where the copy number associated with the segment differs from the copy number associated with its immediate neighbouring segment(s). The copy number associated with a segment may differ from the copy number associated with its immediate neighbouring segment(s) because the segment(s) that surround the segment are associated with a different copy number, because the segment(s) that surround the segment are not associated with a copy number (e.g. because data for the segment(s) is missing, or of insufficient quality), or a combination of both (e.g. a segment may be surrounded by a segment that is associated with a different copy number on one side, and a segment that is not associated with a copy number on the other side). In other words, segments refer to the longest continuous portion of a copy number profile that are each associated with a single copy number.

[0055] The methods described herein may comprise one or more steps of correcting a read count for a genomic bin, using on an estimate of a systematic bias in read counts as a function of a variable associated with said bias (referred to as “bias variable”). The bias variable may be selected from GC content, replication timing, and a capture bias score. The systematic bias may be estimated by fitting a model between (i) the observed read counts in a plurality of genomic bins, and (ii) the value of the bias variable for the bins. The model may be a curvefitting model. The model may be a continuous or piecewise model. The model may be a linear or non-linear model. The model may be a linear model or a polynomial model. The model may be a regression model or a generalised additive model. The model may be a non-linear regression model, such as a local polynomial regression model. For example, the model may be a LOESS (locally estimated scatterplot smoothing) or a LOWESS (locally weighted scatterplot smoothing) model. Alternatively, the model may be a piecewise linear model (e.g. a step-wise model) or a piecewise non-linear model (e.g. spline). The model may be a univariate model or a multivariate model. Thus, in embodiments where read counts are corrected for a plurality of systematic biases, successive rounds of correction may be applied, each correcting for a different systematic bias by fitting a univariate model between (i) the observed read counts in a plurality of genomic bins, and (ii) the value of the respective bias variable for the bins. Alternatively, one or more rounds of correction may be applied, each correcting for a plurality of systematic biases by fitting a multivariate model between (i) the observed read counts in a plurality of genomic bins, and (ii) the value of the respective plurality of bias variables for the bins. Multiple univariate models are advantageously simple to fit and were found to perform well across all tested embodiments. Correcting a read count using the estimate may comprise adjusting the observed read count value for a bin based on the value of the model at the coordinate corresponding to the value of the bias variable(s) for the bin. Thus, correcting a read count using the estimate may comprise obtaining a corrected read count that depends on the observed read count value for a bin and the value of the model at the coordinate corresponding to the value of the bias variable(s) for the bin. For example, correcting a read count using the estimate may comprise subtracting from the observed read count value for a bin a correction factor that is equal to the difference between the value of the model at the coordinate corresponding to the value of the bias variable(s) for the bin and a reference value. The reference value may be a summary value derived from the observed read counts for the plurality of bins used to fit the model. For example, the reference value may be the mean or median read count across the plurality of bins used to fit the model. This is equivalent to adjusting the observed reads counts such that they are distributed around the reference value at all values of the bias variable(s). Thus, correcting a read count for a genomic bin may refer to obtaining a first read count for a bin and determining a second read count for the bin based on the first read count and an estimate of a systematic bias in read counts as a function of a bias variable. The second read count may be referred to as an “adjusted” or “corrected” read count. The first read count may be referred to as a “raw” read count or “observed read count”, even though in some instances it may have already gone through one or more rounds of corrections (e.g. a read count that has been GC corrected may be referred to as a raw read count for the purpose of capture bias correction).

[0056] The GC content of a genomic region (e.g. a bin) refers to the percentage of G or C (guanine or cytosine) bases in the region, as a proportion of all bases in the region (e.g. (G+C) / (G+C+A+T)%). GC content is known to affect the likelihood that a region will be sequenced by next generation sequencing. Correcting for GC content have been proposed which fit a model to the observed read counts per bin as a function of GC content of the bin, and use the value of the model to adjust the observed read counts. In embodiments of the present disclosure, a plurality of models are fitted using a respective plurality of bins, where the bins in each of the respective plurality of bins are associated with the same estimated copy number state. Correction may then proceed for each plurality of bins using the respective model as explained above. Thus, in embodiments of the present disclosure, a copy number may be estimated for each bin (or for each of a plurality of genomic segments, from which bin copy numbers can be obtained), and this may be used to identify multiple sets of bin, each set of bins comprising bins with the same estimated copy number. GC bias correction may then be performed separately for at least one (or all) of the multiple sets of bins. This may be referred to herein as “copy number aware GC correction”. Note that the benefits of such an approach are applicable to all genomic sequencing data, whether from targeted or untargeted sequencing. Copy number aware GC correction may be particularly useful when analysing samples that comprise tumour genetic material. Indeed, such samples may contain abnormal copy number states that can result in poor GC correction if unaccounted for. Thus, copy number aware GC correction may be useful when analysing any sample that has copy number instability. A sample may be considered to have copy number instability if a tumour copy number profile obtained from the sample comprises at least 1 copy number change different from the diploid state. Copy number aware GC correction may be particularly useful when analysing samples that comprise tumour genetic material, where the tumour has high chromosomal instability (high-CIN). A tumour with a copy number profile comprising at least 1 copy number change different from a diploid state may be considered as having CIN. A tumour with a copy number profile comprising at least 5, 10, 15, 20, 25 or 30 copy number changes different from a diploid state may be considered as having “high-CIN”. The present inventors have found copy number aware GC correction to be advantageous for any sample comprising tumour genetic material, where the tumour has a copy number profile comprising at least 1 copy number changes different from a diploid state.

[0057] Replication timing refers to the order in which different regions of a chromosome are duplicated during replication. Replication timing profiles are available for many genomes, which provide the temporal order of replication of all segments in the genome. Replication timing is known to have an effect on the abundance of sequences in a sample, leading to low amplitude changes in DNA copy number (typically much lower than the stepwise changes observed in the presence of copy number variants or somatic copy number aberrations. Not all samples show a systematic bias in read counts as a function of replication timing. However, in embodiments the read count for a bin may be corrected using an estimate of a systematic bias in read counts as a function of replication timing. This may be performed by fitting a model to the observed read counts as a function of replication timing, as explained above. Replication timing information may be obtained from reference genome databases, such as e.g. the USCS genome browser. Replication timing may be obtained from repli-seq timing data for one or more samples. For example, replication timing information may be obtained based on wavelet smoothed repli-seq timing data from LICSC (UW Repli-seq). Replication timing may be obtained as a summarised metric summarising data for a plurality of samples (e.g. average avelet smoothed repli-seq timing data for a plurality of cell lines). Replication timing correction may be advantageous for samples where the observed bin read counts and corresponding replication timings are significantly correlated. For example, replication timing correction may be applied when a correlation coefficient (e.g. the Pearson correlation (r)) between the observed bin read count values and the replication timing is at or above a predetermined threshold, such as e.g. 0.1. Thus, the methods described herein may comprise determining a correlation coefficient (e.g. a Pearson correlation coefficient) between the observed read counts for the plurality of bins (optionally after GC content and / or capture bias correction), and correcting the read count for a bin (or all of the plurality of bins) using an estimate of a systematic bias in read counts as a function of replication timing, when the correlation coefficient is at or above a predetermined threshold.

[0058] Capture bias refers to read counts for one or more bins that are artificially inflated compared to other bins due to the use of a sequence specific capture step in the sequencing process used to obtain the read counts. According to embodiments of the present disclosure, a read count for a genomic bin is corrected using on an estimate of a systematic bias in read counts as a function of a variable associated with said bias. The variable may be a capture bias score. Thus, methods described herein may comprise, in embodiments, fitting a model between (i) the observed read counts in a plurality of genomic bins, and (ii) the value of a capture bias score for the bin, and correcting a read count for a bin by adjusting the observed read count value for the bin based on the value of the model at the coordinate corresponding to the value of the capture bias score for the bin. For example, correcting a read count may comprise subtracting from the observed read count value for a bin a correction factor that is equal to the difference between the value of the model at the coordinate corresponding to the value of the capture bias score for the bin and a reference value. The reference value may be the mean or median read count across the plurality of bins used to fit the model. Capture bias correction may be applied to genomic bins that only include off-target regions and / or only include off target reads. Such bins may be obtained by applying a binning process to the genome or part of the genome being analysed, using a fixed size window, and removing any target region from the resulting bins. This may therefore result in bins of variable sizes.

[0059] A capture bias score is a variable that is indicative of a difference between an observed read location metric distribution in a bin and an expected read location metric distribution in the bin. The expected read location metric distribution may be a read location metric distribution that is associated with a bin that is not affected by capture bias. The expected read location metric distribution may be a read location metric distribution that corresponds to a uniform distribution of reads across the bin. A capture bias score may be any variable that captures a departure between an observed distribution of locations of reads in a bin and a uniform distribution of locations of reads across the bin. The locations of reads in a bin may be represented by a metric of pairwise distance between pairs of reads in the bin. In other words, the read location metric may be a pairwise distance between pairs of reads. Thus, a capture bias score may be any variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin. The expected distribution of pairwise distances of reads across the bin may be a distribution of pairwise distances of reads in the bin corresponding to a uniform distribution of reads in the bin. A pairwise distance between reads may be a left-most pairwise distance (distance between the left most boundaries of two aligned reads), a right most pairwise distance (distance between the right most boundaries of two aligned reads), a middle pairwise distance (distance between the centres of two aligned reads) or any distance between predetermined locations of a read (e.g. the left most boundary, right most boundary or centre location of the read). A distance between reads may be expressed in units of a number of nucleotides. A distance refers to a distance in genomic coordinates of a reference sequence to which the two reads are aligned. The leftmost distance between two reads in a pair of reads may be the distance between the leftmost boundary of the first read and the leftmost boundary of the second read. A distance of 0 means that the two reads have the same leftmost end. An expected distribution of pairwise distances between reads for a bin may be a left triangular distribution with a mode of 0 or the minimum observed pairwise distance in the bin, a minimum of 0 or the minimum observed pairwise distance in the bin and a maximum equal to the maximum observed distance in the bin, the length of the bin, or the length of the bin + the length of a read, or any value between the length of the bin and the length of the bin + the length of a read. The length of a read is typically a known feature of the sequencing technology used. For example, when determining a score for a bin of 50kb and reads of 150 nt (nucleotides) long, the expected distribution may be a triangular distribution with minimum=0, mode=0 and maximum=50kb nt, 50150 nt or the maximum observed distance in the bin (typically within or near the range of 50000-50150 nt). The present inventors have found all such values for the maximum of the expected distribution to lead to similar results. An expected distribution may be approximated by sampling from such a left triangular distribution a number of observations corresponding to the number of observations from which an observed distribution is obtained (i.e. number of observed pairwise distances between reads). A variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin may be a variable that summarises (e.g. a sum, mean, average, etc.) across a plurality of windows of pairwise distances, the absolute difference between the observed frequency for the window and the expected frequency for the window. The plurality of windows may be nonoverlapping windows that together cover the range from the smallest observed distance to the largest observed distance. The plurality of windows may be non-overlapping windows that together cover the range of the expected distribution of distances. Thus, a variable that captures a departure between an observed and an expected distribution of pairwise distances between reads in a bin may be a variable that summarises the differences between values of a histogram of observed pairwise distances between reads in the bin (i.e. a set of estimated frequencies for respective windows of pairwise distances) and corresponding value from the expected distribution (e.g. left triangular distribution). The number of windows in the histogram may be predetermined or may be dynamically determined based on the number of observations in the distribution. For example, the variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise locations of reads across the bin may be the mean of the absolute value of the deviations between a histogram density function fitted on the observed pairwise location data and a corresponding theoretical left-triangular distribution. The number of windows is in principle not limited, such that when very large amount of data is available, a very high-resolution histogram (that could also be referred to as a density function) may be fitted. Thus, the, the variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise locations of reads across the bin may be the integral of the area between a density function fitted on the observed pairwise location data and a corresponding theoretical left-triangular distribution. A variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin may be a variable that summarises (e.g. a sum, mean, average, etc.) across a plurality of observed pairwise distances (e.g. across all observed pairwise distances in the bin), the probability that the observed pairwise distance is sampled from the expected distribution. The capture bias score may be obtained as 1 minus the value of such a variable, such that higher values are associated with larger departures from the expected distribution. A variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin may be a statistical metric that compares two samples to determine whether they come from the same distribution. Thus, a variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin may be a statistical metric associated with a hypothesis test that two samples are from the same distribution. The statistical metric may be selected from: a Kolmogorov-Smirnov statistic, a DTS statistic, and Anderson-Darling test statistic, a Cramer- von Mises test statistic, etc. In embodiments, the statistical metric is a DTS statistic. The statistical metric may be calculated between a first sample comprising the observed pairwise distances between reads and a sample of expected pairwise distances between reads drawn from the expected distribution of pairwise distances between reads (e.g. left triangular distribution).

[0060] A capture bias score may be obtained for any bin that contains at least a predetermined number of reads, such as e.g. 5, 10, 15 or 20 reads. A minimum value of 10 was found to strike a good balance between keeping as much raw data as possible and obtaining reliable scores (which depend on the availability of sufficient data for a bin). For bins that contain fewer than the predetermined number of reads, the score may be set to a default equal to the minimum capture bias score identified for the sample.

[0061] A “genomic bin” (also referred to herein as “bin”) is a region of a genome that has been obtained through a process comprising applying a fixed width window to a genomic region (e.g. a chromosome, plurality of chromosomes or whole genome). The fixed width window is typically applied in sliding, non-overlapping manner, creating a plurality of adjacent bins of equal size. The size of a bin may be set to a predetermined default, such as e.g. 20kb nt, 50 kb nt or 100 kb nt. In embodiments, a default of 50kb nt is used. Alternatively, the size of a bin may be determined based on the properties of a sample to be analysed, such as e.g. based on factors including tumour purity (when a tumour copy number profile is determined from a mixed sample), number of off-target reads and sequencing depth. All of these factors influence the density of informative reads mapped to the genome, and thus influence the size of bins that will contain enough informative reads. Increasing the size of the bins from a default of e.g. 50kb nt may be useful when tumour purity and / or the number of off-target reads is low. Conversely, when tumour purity is high and / or the number of off target reads is high, narrower bins can be used to obtain higher resolution data (e.g. a higher resolution copy number profile). It is generally within the ability of the skilled person’s to test plurality of bin width values (e.g. within predetermined ranges such as between 10kb and 1000kb nt, or between 50kb and 500kb nt) to identify a bin size that strikes a good balance between accuracy and resolution for a particular sample. See e.g. Macintyre, Ylstra & Brenton (2016) and Scheinin et al. 2014. For example, a bin width may be selected so as to obtain on average at least 15 reads per bin per tumour copy (i.e. average of 15 reads per absolute copy number of the tumour at the location of the bin). Bin sizes between 10kb and 500kb are believed to be particularly advantageous as they provide a good resolution for a range of downstream analyses. However, bin sizes of up to 1000kb are useful if such a resolution is necessary to ensure that a sufficient number of reads per bin is present (such as e.g. at least 15 reads per absolute copy number of the tumour at the location of the bin, on average). For targeted capture sequencing assays, sufficient number of reads are typically available with samples that have at least 10%, preferably at least 20% tumour purity. For targeted capture sequencing assays, sufficient number of reads are typically available with samples that have at least 200,000 reads. High resolution copy number profiles (e.g. 50kb bins) can typically be obtained for samples with about 2 million reads. As will be explained further below, a genomic bin may be further processed to remove one or more target regions, for the purpose of capture bias correction. This results in a processed bin width that is smaller than the original bin width. Therefore, for the purpose of capture bias correction, a plurality of bins may have respective processed lengths that differ from each other. This is despite all these bins having been obtained through a process comprising applying a fixed width window to a genomic region. The difference between the (original) bin width and the processed bin width may be used to adjust read counts for the bin after bias correction including capture bias correction. In particular, an adjusted read count may be obtained by adding, to a corrected read count for a bin, a read count proportional to the corrected read count in the bin and the size of the difference between the (original) bin width and the processed bin width. Thus, the processed bin width may be used for the purpose of adjusting the read count for a bin after capture bias correction using a capture bias score. Adjusting a read count in this context refers to adding a pseudo-count to a corrected count (which is a count obtained after at least capture bias correction, and also optionally after GC content and / or replication timing correction), where the pseudo-count is: pseudo-count=(%overlap previously removed from the bin)*(count forthe bin after corrections applied). Thus, the adjusted count in this embodiment will be: (count for the bin after corrections applied)* (1+(=(%overlap previously removed from the bin)). The determination of the capture bias score may use the original bin width as a parameter (e.g. for determining the expected distribution of a read location metric). Alternatively, the bin width may not be processed and the corrected read count for the bin may be obtained by obtaining a first corrected read count based on the observed off target read count for the bin and one or more estimates of biases in read counts as a function of a variable associated with said respective biases (i.e. obtaining a corrected read count for the bin as explained herein, simply removing on-target reads in the bin), and adding a pseudocount to this read count. The pseudocount in this context is proportional to the first read count and the proportion of the bin that overlaps with a captured region. For example, a pseudocount may be calculated as (%overlap with the captured regions file)*(count for the bin after all corrections applied). The captured regions may be as defined in a BED file.

[0062] A “sample” as used herein may be a cell or tissue sample, a biological fluid, an extract (e.g. a DNA extract obtained from a cell, tissue or fluid sample), from which genomic material can be obtained for genomic analysis, such as genomic sequencing (e.g. whole genome sequencing, or targeted sequencing such as whole exome sequencing or targeted panel sequencing). The sample may be a cell, tissue or biological fluid sample obtained from a subject (e.g. a biopsy). Such samples may be referred to as “subject samples”. In particular, the sample may be a blood sample, a urine sample, a cerebrospinal fluid sample, or a tumour sample (e.g. a tumour biopsy), or a sample derived therefrom. The sample may be any sample comprising genomic DNA or cell free DNA. The sample may be one which has been freshly obtained from a subject or may be one which has been processed and / or stored prior to genomic / transcriptomic analysis (e.g. frozen, fixed or subjected to one or more purification, enrichment or extraction steps). The sample may be a cell or tissue culture sample. As such, a sample as described herein may refer to any type of sample comprising cells or genomic material derived therefrom, whether from a biological sample obtained from a subject, or from a sample obtained from e.g. a cell line. In embodiments, the sample is a sample obtained from a subject, such as a human subject. The sample is preferably from a mammalian (such as e.g. a mammalian cell sample or a sample from a mammalian subject, such as a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit or guinea pig), preferably from a human (such as e.g. a human cell sample or a sample from a human subject). Further, the sample may be transported and / or stored, and collection may take place at a location remote from the sequence data acquisition (e.g. sequencing) location, and / or any computer-implemented method steps described herein may take place at a location remote from the sample collection location and / or remote from the sequence data acquisition (e.g. sequencing) location (e.g. the computer-implemented method steps may be performed by means of a networked computer, such as by means of a “cloud” provider). A sample may be a tissue biopsy, such as a sample of fresh frozen tissue or a sample of formalin-fixed paraffin embedded (FFPE) tissue, or a sample comprising circulating tumor DNA (ctDNA) such as e.g. a sample of biological fluid (sometimes referred to as liquid biopsy).

[0063] The samples used in methods of the present disclosure are typically samples comprising tumour cells (e.g. a tumour sample or sample comprising circulating tumour cells) or genetic material derived from tumour cells (such as e.g. cell free DNA, or DNA extracted from a sample comprising tumour cells - e.g. from a cell line or tissue sample - or circulating tumour cells). The methods described herein may be applicable to single samples. In particular, the methods described herein may enable to obtain a copy number profile for a sample including a sample comprising one or more copy number abnormalities (e.g. a tumour sample or sample comprising genetic material derived from tumour cells) in the absence of a matched normal sample. A matched normal sample is a sample that is expected to be representative of the genome of a tumour sample in the absence of somatic abnormalities. This is typically a sample obtained from the same subject as the tumour sample.

[0064] The term “sequence data” refers to information that is indicative of the presence of genetic material in a sample that has a particular sequence. Such information may be obtained using sequencing technologies, such as e.g. next generation sequencing (NGS), for example whole exome sequencing (WES), whole genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing). In the context of the present disclosure, the sequence data is typically obtained by DNA sequencing, and particularly next generation sequencing. Thus, the sequence data may comprise sequencing reads, or information derived therefrom. Information such as the count of sequencing reads that have a particular sequence, that cover a particular genomic location or that fall within a particular genomic region (e.g. genomic bin) can be derived from sequencing reads by mapping the sequence data to a reference sequence, for example a reference genome, using methods known in the art (such as e.g. Bowtie (Langmead et al., 2009) or BWA (Li and Durbin 2009)). This may result in aligned sequencing reads, for example in the form of a SAM or BAM file. Thus, sequencing may be associated with a particular genomic location (where the “genomic location” refers to a location in the reference genome to which the sequence data was mapped). Sequencing data may be filtered as part of the methods described herein or may have been filtered prior to application of the methods described herein, for example to remove low quality reads or remove reads that map to regions that have “low mappability”. Methods for filtering low quality reads are known in the art, and include e.g. removing duplicate reads, removing supplementary alignments, removing reads with low sequencing quality scores (reads below Q20 or Q30), removing reads with less than a predetermined % of characterised bases (e.g. reads with less than 100% characterised bases, where characterised bases can be identified as “non-NNNN” bases), removing reads with low mapping quality score (e.g. MAPQ<37), etc. Methods to filter reads that map to regions that have low mappability are known in the art and include e.g. llrnap (Karimzadeh et al. 2018), and GenMap (Pockrandt et al. 2020).

[0065] The present disclosure also provides methods that make use of the obtained untargeted genomic sequencing data, to generate copy number profiles or genomic features that are derived from copy number profiles, such as the presence or absence of one or more signatures of chromosomal instability (CIN), the tumour purity or percentage of circulating tumour DNA (ctDNA) in a sample, etc.

[0066] Chromosomal instability (CIN) is the process of accumulating numerical and structural changes in DNA. A signature of chromosomal instability (CIN) (also referred to as copy number signature) is a signature (typically in the form of a set of weights associated with respective categories of copy number features) representing the genome-wide imprint of distinct putative mutational processes (where a “mutational process” as used herein refers to any process that can cause chromosomal instability). CIN signatures are described in MacIntyre et al. 2018, Drews et al. (2022) and in PCT / EP2022 / 077473, which are incorporated herein by reference. The presence or absence of a CIN signature in a sample may be determined by calculating the exposure of the sample to the signature. Methods for determining the exposure to a signature, from a copy number profile, are known in the art (see e.g. MacIntyre et al., 2018, Drews et al. (2022)). The term “copy number (CN) features” refers to properties of copy number events observable in a copy number profile. Copy number features may include: segment absolute copy number, segment size (also referred to as “segment length” typically expressed in number of bases), breakpoint count per x MB (the number of changepoints appearing in a sliding windows across the copy number profile, where the window is preferably 10 MB and the copy number profile is preferably genome-wide), change-point copy number (the absolute difference in copy number between a segment and an adjacent / neighbouring segment in the copy number profile, which may be defined relative to the upstream or downstream neighbouring segment), breakpoint count per chromosome arm (the number of changepoints occurring per chromosome arm), and number of segments with oscillating copy number (sometimes referred to as “length of segments with oscillating copy number”; number of continuous segments alternating between two copy number states, rounded to the nearest integer copy-number state; also referred to as length of chain of oscillating copy number states). A sample may be a “mixed” sample comprising cells with different genotypes or genetic material derived therefrom. For example, a sample may be a sample comprising tumour cells and normal cells, or DNA derived therefrom, such as e.g. in the context of cell free DNA (cfDNA) samples comprising circulating tumour DNA (ctDNA). Copy number profiles may be used to determine the percentage of cells or DNA derived from cells that have a genotype comprising somatic copy number alterations, in a sample comprising both cells that have the genotype comprising somatic copy number alterations (e.g. tumour cells) and cells that do not contain such alterations (e.g. normal cells). The term “tumour fraction” (also sometimes referred to as “tumour purity” or simply “purity”, or aberrant cell fraction (ACF)) refers to the proportion of DNA containing cells within a mixed sample that are tumour cells, or to the equivalent proportion that is assumed to result in a particular mixture of genetic material from tumour and non-tumour cells in a sample. Thus, the methods described herein may be used to determine the tumour fraction in a sample comprising cells or genetic, or the percentage of ctDNA in a sample comprising cfDNA. A tumour fraction of % ctDNA may be estimated using sequence analysis processes that attempt to deconvolute tumour and germline genomes from copy number profiles, such as e.g. ASCAT (Raine et al., 2016), ABSOLUTE (Carter et al., 2012), or, in the context of cfDNA, ichorCNA (Adalsteinsson et al., 2017).

[0067] As used herein, the terms “computer system” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above described embodiments. For example, a computer system may comprise a central processing unit (CPU) and / or a graphic processing unit (GPU), input means, output means and data storage, which may be embodied as one or more connected computing devices. A computer system may comprise a display or comprises a computing device that has a display to provide a visual output display. The data storage may comprise RAM, disk drives or other non- transitory computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system may consist of or comprise a cloud computer.

[0068] As used herein, the term “computer readable media” includes, without limitation, any non- transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media. Obtaining untargeted genomic sequencing data

[0069] The present disclosure provides methods for analysing genomic sequence data comprising correcting one or more systematic bias in the data. In particular, the present disclosure provides methods for obtaining untargeted genomic sequencing data, from targeted sequencing data. An illustrative method will be described by reference to Figure 1 .

[0070] The method may comprise optional step 10 of obtaining one or more samples comprising cells or genetic material derived therefrom. Optionally, the samples may be sequenced at step 12, using a targeted panel DNA sequencing technology. Alternatively, the sequence data may have been previously obtained and may be received from a user interface, computing device or database. The sequence data may comprise sequence data from a targeted sequencing assay. The sequence data may comprise a plurality of sequencing reads. The seqnece data may further comprise coordinates of a plurality of captured regions. At optional step 14, the sequence data may be preprocessed. This may comprise aligning a plurality of reads in the sequence data to a reference genome. Alternatively, the sequence may have been previously aligned, and may be received in the form of one or more files comprising reads and aligned locations in a reference genome (e.g. a SAM or BAM file). Step 14 may comprise applying a non-overlapping fixed width window along genomic coordinates for one or more genomic regions that the reads map to, thereby obtaining a plurality of fixed width bins. Step 14 may further comprise annotating the plurality of bins, for example with replication timing information and / or GC content. At step 16, reads that are not mapped to captured regions are selected and used to determine observed counts for each of the plurality of bins. These may be referred to as “off-target reads” and observed read counts obtained using these may be used for all subsequent steps. When using paired-end sequencing data, the off-target reads may be reads that do not map to a captured region, and reads that do not have a mate read that maps to a captured region. Optionally, a processed bin width for a bin may also be obtained by removing any region from a fixed width bin that is a captured region. Thus, a processed bin width may be obtained for each bin that overlaps with a captured region. The processed bin widths may differ from each other and from the fixed width. Instead or in addition to this, a proportion or percentage of any fixed width bin that is a captured region may be computed.

[0071] At step 18, GC correction is performed. This may comprise obtaining GC corrected read counts for the bins by fitting a model between the observed read counts, and (ii) the value of GC content for the bins. This may be referred to as “single GC correction” as it corrects read counts for all bins using a single model. Step 18 may further comprise using the GC corrected read counts to determine whether the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins, all bins in a set having the same estimated copy number and the estimate copy numbers associated with each of the plurality of sets being different from each other. When it has been determined that the observed read counts for each of the plurality of genomic bins are indicative of the presence of a plurality of sets of bins, step 18 may further comprise determining a corrected read count based on the observed read count (optionally already corrected for capture bias, or prior to correcting for capture bias) and an estimate of biases in read counts as a function of GC content). The estimate of GC content bias in read counts as a function of GC content is obtained by: identifying a plurality of sets of bins where all bins in a set have the same estimated copy number, and fitting a model between (i) the observed read counts in one of the sets of bins comprising the bin for which a corrected read count is determined, and (ii) the value of GC content for the bins in the set. This may be referred to as “set specific GC correction” by reference to the fact that the correction uses a plurality of models each fitted to a specific set of bins. For both single and set-specific GC correction, the estimate of GC bias in read counts as a function of GC content is estimated by fitting a model between (i) the observed read counts in the plurality of genomic bins (or a selected set), and (ii) the value of the GC content for the plurality of genomic bins (or a selected set). A corrected read count is then determined for each bin by subtracting from the observed read count a correction factor that is equal to the difference between: (i) the value of the model at the coordinate corresponding to the value of the GC content for the bin, and (ii) a reference value, such as a summary value (e.g. median) derived from the observed read counts for the plurality of bins (or a selected set).

[0072] At step 20, capture bias correction is performed by determining a corrected read count for one or more of the plurality of genomic bins, the corrected read count based on the GC corrected observed read counts obtained at step 18 and an estimates of bias in read counts as a function of a variable associated with said bias, wherein the variable associated with said bias is a capture bias score that is indicative of a difference between an observed read location metric distribution in the bin and an expected read location distribution metric in the bin. The estimate of a bias in read counts as a function of the variable associated with said bias is estimated by fitting a model between (i) the GC corrected observed read counts in the plurality of genomic bins, and (ii) the value of the bias variable for the plurality of genomic bins. A corrected read count is determined for each bin by subtracting from the GC corrected observed read count a correction factor that is equal to the difference between: (i) the value of the model at the coordinate corresponding to the value of the bias variable for the bin, and (ii) a reference value, such as a summary value (e.g. median) derived from the GC corrected observed read counts for the plurality of bins. At step 22, the corrected counts obtained at step 20 may be further corrected for replication timing. This may comprise determining whether the observed read counts (typically already corrected, such as e.g. resulting from steps 18 or 20, in the illustrated embodiment) for the plurality of bins are significantly associated with replication timing, and further correcting the observed read counts (typically already corrected) for replication timing bias when the (typically corrected) observed read counts for the plurality of bins are significantly associated with replication timing. This may be performed by: fitting a model between the (typically corrected) observed read counts, and (ii) the value of replication timing for the bins, and determining a corrected read count is for each bin by subtracting from the (typically corrected) observed read count a correction factor that is equal to the difference between: (i) the value of the model at the coordinate corresponding to the value of the replication timing for the bin, and (ii) a reference value, such as a summary value (e.g. median) derived from the (typically corrected) observed read counts for the plurality of bins.

[0073] At optional step 24, a further corrected read count is obtained for each bin which overlaps with a captured region and / or in which on-target reads were removed at step 16. This can be performed by adding to the corrected read count (result of any one or more of steps 18, 20, 22) for a bin, an adjustment factor that is proportional to the difference between the processed bin width determined at step 16 for the bin and the width of a fixed width bin. The adjustment factor can be proportional to the corrected read count for the bin and the difference between the processed bin width for the bin and the width of a fixed width bin. Alternatively, this can be performed by adding to the corrected read count (result of any one or more of steps 18, 20, 22) for the bin an adjustment factor (also referred to as “pseudocount”) that is proportional to the corrected read count for the bin and the proportion or percentage of the bin that overlaps with a captured region. For example, the corrected read count for a bin that overlaps with a captured region may be calculated as (count for the bin after all corrections applied)+pseudocount, where pseudocount can be calculated as pseudocount=(%overlap with captured regions)*(count for the bin after all corrections applied).

[0074] At optional step 26, the results of one or more of the preceding steps are provided to a user, for example through a user interface.

[0075] Also described herein is a method for obtaining targeted and untargeted genomic sequence data from a sample using a targeted DNA sequencing protocol, the method comprising: obtaining targeted genomic sequence data from the sample, comprising a plurality of DNA sequencing reads obtained using a targeted genomic sequencing assay; and applying the method described above, such as e.g. the method of Figure 1 to the sequence data, thereby obtaining untargeted genomic sequence data for the sample.

[0076] Applications

[0077] The above methods find applications in the context of determining copy number profiles or the presence of one or more genomic features that are derived from copy number profiles, such as the presence or absence of one or more signatures of chromosomal instability, or the percentage of tumour or ctDNA in a sample.

[0078] Thus, also described herein are methods of providing a copy number profile for a sample or a genomic feature derived therefrom, the methods comprising obtaining untargeted genomic sequence data using a method as described herein. An example of such a method will be described by reference to Figure 2.

[0079] At optional step 210, one or more samples cells or genetic material derived therefrom are obtained. At optional step 212, genomic sequence data is obtained for the samples. At step 214, corrected read counts are obtained from the genomic sequence data for a sample using a method as described herein, such as e.g. by reference to Figure 1. At step 216, a copy number profile is obtained for the sample using the corrected counts obtained at step 216. This may comprise step 216A of segmenting the corrected read count profile along genomic coordinates, thereby identifying segments of the genome that have the same copy number state. Step 216 may comprise determining an absolute copy number for each segment at step 216B. Step 216 may comprise determining a tumour purity or percentage of ctDNA in a cfDNA sample, using the corrected read counts. Absolute copy number fitting and purity I %ctDNA estimation may in practice be performed simultaneously through a process of deconvolution a mixed copy number profile obtained from data comprising reads associated with tumour genetic material and reads associated with normal genetic material into a copy number profile associated with the tumour genomic material in the sample and a copy number profile associated with the normal genomic material in the sample (typically assumed to be diploid at all segments). At step 218, derived information may be obtained by analysing the copy number profile(s) obtained at step 216. This may comprise e.g. determining whether the copy number profile is indicative of the presence or one or more CIN signatures. This may comprise determining a prognostic and / or diagnostic for a subject from which the sample has been obtained. Step 218 may comprise determining the presence of one or more copy number abnormalities in the sample, such as e.g. the presence of a copy number abnormality at one or more locations of clinical relevance (e.g. oncogene amplification, tumour suppressor deletion, etc.) At step 220, the results of any preceding step may be provided to a user, for example through a user interface. This may take the form of e.g. a report indicating the presence of one or more copy number abnormalities, one or more CIN signatures, a prognostic or diagnostic metric, a tumour purity or % ctDNA, etc.

[0080] Also described herein are methods of determining the percentage of tumour cells or tumour DNA in a sample, the methods comprising obtaining untargeted genomic sequence data for the sample using a method as described herein (i.e. corrected read counts, whether obtained from targeted or untargeted genomic sequence data), and identifying a tumour copy number profile and percentage of tumour DNA in the sample from said data. This may comprise deconvoluting an observed read count profile between a contribution from normal cells with known expected ploidy and a contribution from tumour cells with unknown copy number profile. Methods for performing this are known in the art and include e.g. Ascat (Raine et al. 2016), ichorDNA (Adalsteinsson et al. 2017), Absolute (Carter et al. 2012), Sequenza (Favero et al. 2015), etc.

[0081] The sample may be a cell free DNA sample and the method may determine the percentage of circulating tumour DNA in the sample. For example, a method as described in Adalsteinsson et al. (2017) may be applied to untargeted genomic sequence data obtained as described herein (i.e. read counts associated with genomic coordinates, corrected for GC bias and / or capture bias).

[0082] The sample may be a sample comprising cells (e.g. a tumour biopsy) and the method may determine the percentage of tumour cells in the sample. For example, a method as described in Raine et al. (2016) or Carter et al. (2012) may be applied to untargeted genomic sequence data obtained as described herein (i.e. read counts associated with genomic coordinates, corrected for GC bias and / or capture bias).

[0083] The percentage of ctDNA in a cfDNA sample from a subject has been identified as an indicator of minimal residual disease (MRD), prognosis and response to therapy in many cancers. Thus, the present disclosure also provides method of providing prognosis, MRD diagnosis or indication of response to therapy for a subject that has been diagnosed as having or being likely to have cancer, the method comprising obtaining targeted sequencing data obtained from a sample of biological fluid from the subject and analysing the data using the methods described herein to determine the percentage of ctDNA in the sample, wherein the percentage of ctDNA in the sample is indicative of the subject’s prognosis, MRD status or response to therapy. Any method known in the art to determine a prognosis, MRD status or response to therapy from a percentage ctDNA in a liquid biopsy may be used for this purpose. Systems

[0084] Figure 3 shows an embodiment of a system for obtaining untargeted genomic sequencing data, and / or for providing a copy number profile or genomic feature derived therefrom, according to the present disclosure. The system comprises a computing device 1 , which comprises a processor 101 and computer readable memory 102. In the embodiment shown, the computing device 1 also comprises a user interface 103, which is illustrated as a screen but may include any other means of conveying information to a user such as e.g. through audible or visual signals. The computing device 1 is communicably connected, such as e.g. through a network 6, to sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 storing sequence data. The one or more databases may additionally store other types of information that may be used by the computing device 1 , such as e.g. reference sequences, parameters, etc. The computing device may be a smartphone, tablet, personal computer or other computing device. The computing device is configured to implement a method for obtaining untargeted genomic sequencing data, and / or for providing a copy number profile or genomic feature derived therefrom, as described herein. In alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which is itself configured to implement a method of obtaining untargeted genomic sequencing data, and / or for providing a copy number profile or genomic feature derived therefrom, as described herein. In such cases, the remote computing device may also be configured to send the result of the method to the computing device. Communication between the computing device 1 and the remote computing device may be through a wired or wireless connection, and may occur over a local or public network such as e.g. over the public internet or over WiFi. The sequence data acquisition 3 means may be in wired connection with the computing device 1 and / or the database 2, or may be able to communicate through a wireless connection, such as e.g. through a network 6, as illustrated. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (such as e.g. through a remote computer). The sequence data acquisition means 3 are configured to acquire sequence data from nucleic acid samples, for example genomic DNA samples extracted from cells and / or tissue samples. In some embodiments, the sample may have been subject to one or more preprocessing steps such as DNA purification, fragmentation, library preparation, target sequence capture (such as e.g. exon capture and / or panel sequence capture). The sequence data acquisition means is preferably a next generation sequencer. The sequence data acquisition means 3 may be in direct or indirect connection with one or more databases 2, on which sequence data (raw or partially processed) may be stored. The following is presented by way of example and is not to be construed as a limitation to the scope of the claims.

[0085] EXAMPLES

[0086] Introduction

[0087] Capture-based DNA targeted panels are routinely used for small variant detection and for a very limited gene-level copy number profiling in cancer clinical care. The capture efficiency of these panels is not optimal, resulting in off-target DNA being sequenced alongside on-target DNA. These off-target reads can range from 8 to 50% of the total number of reads and they are scattered across the genome, allowing for a genome-wide copy number profile to be derived. However, there are significant biases in the number of reads in these regions which makes a comprehensive copy number characterization very difficult. In the present examples, the present inventors describe a method to correct these biases and generate robust genomewide copy number profiles. They demonstrate the method by analysing a variety of NGS protocols and showing that after correcting for GC content and off-target read-depth bias, the two major sources of noise in this type of data, the proposed method could provide comparable results with shallow whole-genome sequencing (sWGS) samples, the current gold standard. Benchmarks with other computational methods further show an improvement in performance without the need of a matched normal tissue or filtering out noisy data, usual drawbacks for other methods in the cancer field.

[0088] Example 1- Calling genome-wide copy number from targeted capture panel sequencing data

[0089] Calling genome-wide copy number from targeted capture panel sequencing data is challenging due to two major factors which artificially alter the number of reads from the expected number, which should be proportional to the amount of DNA: GC content and capture bias.

[0090] GC content is a well-known source of bias that has been extensively addressed in previous work. The ability of next-generation sequencing approaches to amplify and sequence any given fragment of DNA is dependent on the ratio of guanines (G) or cytosines (G) to adenosine (A) and tyrosines (T) in the sequence. Generally, the higher the GC content of a DNA fragment, the more likely it will be sequenced. This causes artificially uneven read-depth when aligning to the reference genome. To account for this, the observed read-depth is usually corrected based on the GC content. Typically, a LOESS curve is fit between the raw read-depth and the GC content of the genomic region (see e.g. Figure 4) and this is used to adjust the readdepth, removing the bias caused by variable GC content. While generally successful for normal genomes, challenges arise when applied to tumour genomes due to the variable copy number states. In a tumour genome, the variable copy number states cause a shift in the read depth distribution (see Figure 5). A standard GC fitting procedure will fit to the centre of this split distribution and thus the correction procedure will subtly shrink the true separation between copy number states towards the average (see Figure 6, top plot). This has a detrimental impact on downstream applications such as estimation of absolute copy number. Recently, an approach developed for genome-wide copy number estimation from low-pass single cell DNA sequencing was described by Wang et al. (2020). In order to resolve this issue, this approach uses a Poisson latent factor model for read-depth normalization and models GC content bias using an expectation-maximization (EM) algorithm embedded in the Poisson regression models to account for the discretized copy-number states along the genome. However, no approach has been developed that can deal with this issue in the context of bulk DNA sequencing, offering an opportunity to improve GC correction, including from capture panel sequence data.

[0091] The second major source of read-depth bias is caused by off-target enrichment of DNA by the capture process. Any DNA sequence outside the designated capture regions which is sufficiently similar to the capture regions will be enriched and thus show higher read-depth. For example, regions which contain pseudogenes. One way to correct for this bias is to sequence a matched normal sample alongside the tumour sample where the expected coverage should be even across the genome and thus an appropriate correction model can be developed. However, in many clinical scenarios a matched normal sample is not available and thus these approaches are not applicable. Another approach is to determine which regions are likely to have artificially high read-depth due to capture bias and remove them from downstream analysis. This approach is only successful for low-resolution copy number calling and creates gaps in the genome which makes higher resolution copy number calling difficult. The present inventors have postulated that with a sufficient framework for scoring regions across the genome for their likelihood of containing sequence content similar to the capture regions, it may be possible to employ a GC correction style procedure and retain all regions across the genome.

[0092] Other potential sources of noise were also investigated. Mappability has historically been considered as a source of bias in many DNA sequencing experiments. The bias is particularly problematic in single-end sequencing experiments since the sequence lengths to be mapped are shorter compared to paired-end sequencing and thus less unique. Most targeted panels are designed for paired-end sequencing, so the negative impact caused by mappability is minimal in these cases. However, the inventors have created their own paired-end mappability annotations to address any issues that may be associated with this source of noise. With these annotations, bins with mappability below 95% (a vast minority) are blacklisted without requiring further corrections. On the other hand, extensive testing has revealed that replication timing may in some cases represent a source of bias. Indeed, bin read counts were found to show a high correlation with replication timing in certain specific cases, likely samples from aggressive tumors with a high replication rate. Nevertheless, it was found that replication timing only affects a minority of cases. Therefore, the present inventors also devised methods to automatically identify those samples before correction instead of correcting every sample by default.

[0093] A number of methods that are able to determine genomic copy number profiles sequencing data exist, but very few are designed to deal with targeted panel sequencing data and therefore many do not correct for biases associated with targeted panel sequencing data. Even those that do attempt to deal with the specific features of targeted panel sequencing data have limitations as will be apparent from the discussion below. Selected features of these methods are summarised in Table 1 and compared to the methods described herein (labelled “disclosure”). These include: (i) whether the methods are able to provide a genome wide copy number profile for a single sample even when the sample is a tumour sample (all prior methods require a reference sample in order to be able to correct for bias that may occur in tumour samples due to copy number alterations, if they are even able to make such corrections); (i) whether the methods use off-target reads only (thereby avoiding probe affinity bias and uneven distribution of on-target reads); (iii) whether the methods output high resolution genome-wide copy number profiles or are limited to gene-based copy number, (iv) whether the methods are capable of replication timing correction, and (v) whether the methods apply a correction that is GC aware.

[0094] Table 1. Comparison of methods for obtaining genome wide copy number profiles

[0095] CNVkit (Talevich et al. 2016) uses both on and off target reads in bin-level data, used to provide genome-wide copy number profiles. Normal samples or pools of tumour samples are used as a reference to address GC content bias, repeat-masked fraction bias, and target density bias. CopywriteR (Kuilman et al. 2015) provides genome-wide copy number profiles using off target reads only, and addresses off-target capture bias by applying the Model-based Analysis for ChlPseq (MACS) algorithm to detect peaks in germline samples enriched using various capture sets and removing those regions from the analysis, with or without a matched normal. Control-FREEC (Boeva et al. 2012) does not use off-target reads (it is based on the depth of coverage of captured exons) and does not provide genome-wide copy number profiles if the input is targeted panel. CODEX (Jiang et al. 2015) does not use off-target reads (it is based in the depth of coverage of captured exons) and does not provide genome-wide copy number profiles. ASCAT (Raine et al. 2016) uses both on and off-target reads to obtain bin-level data, from which genome-wide copy number profiles are obtained. Biases are addressed using matched normal samples. ADTEx (Amarasinghe et al. 2014) does not provide genome-wide copy number profiles, and uses on-target reads only. Any tumour related biases are addressed biases using matched normal samples. FACETS (Shen and Seshan, 2016) can provide genome wide genome-wide copy number profiles depending on the parameterization using only on-target or both on and off-target, but only addresses tumour related biases using matched normal samples. CNV Radar (Soong et al. 2020) does not provide genome-wide copy number profiles, focussing instead on reporting copy number status of cancer genes, using on-target reads only. Any tumour related biases are addressed using normal samples.

[0096] Of the methods discussed above, only CopywriteR is designed to only use off-target reads. Methods in the table below that use on-target reads are expected to have artefact peaks at target locations, therefore being unsuitable for providing genome wide copy number profiles. Further, many methods in the above table only provide the copy number status of genes included in the panel, or are optimized for whole exome, providing the copy number status of every gene and then this is used to indirectly estimate the copy number status of the entire genome. This is different from actually estimating the copy number for entire genomes (even intergenic sequences) from the data available. To the best of the inventors’ knowledge only two previous methods use off-target reads therefore actually interrogating every location in the genome, regardless of the input: CNVkit and CopywriteR. However, CNVkit does not use off-target reads only, and CopywriteR has limitations at least because it removes data associated with off-target reads by peak calling, resulting in gaps in the copy number profile. Further, none of these methods can do GC correction in an accurate manner for samples with high chromosomal instability, since all perform simple GC correction that fits a single model across all copy number states.

[0097] To improve copy number calling from targeted capture panels the present inventors created a new methodology that addresses at least four current deficiencies of existing methods in that the new method: 1) works on tumour-only samples (no need for matched normal); 2) corrects the noise in the data rather than filtering out data; 3) is comprehensive, i.e. provides the copy number status of the entire genome, not only certain genes; and 4) performs copy number aware GC correction.

[0098] Methods

[0099] Data. The method was demonstrated using a number of different capture-based targeted sequencing technologies. These include: Whole Exome Sequencing (WES), more specifically the SureSelect Human All Exon protocol (Agilent), the T ruSight Oncology 500 assay (Illumina), TruSight Tumor 170 assay, and non-targeted approaches such as sWGS. The method was tested on multiple different tumor types and different preservation methods including fresh frozen tissues, formalin-fixed paraffin embedded (FFPE) tissues, as well as circulating tumor DNA (ctDNA) liquid biopsies demonstrating that it could be used in a wide range of situations. All the samples used in the present examples are summarised in Table 2 below.

[0100] Table 2. Summary of samples used.

[0101] Data is only shown in these examples for the first two datasets (WES data from lung adenocarcinoma cell lines and TSO500 for high-grade serous ovarian cancer tumours).

[0102] Read alignment and quality control. Paired-end raw reads were aligned to the human hg19 reference genome. The method uses paired-end binary sequence alignment files (BAM file) and the coordinates of the target regions (BED file) as input, both using the plain vanilla “hg19” human reference genome assembly with no ALT contigs or decoy sequences. Duplicate reads, supplementary alignments, and reads not passing quality controls (mapping quality score <37) are discarded. Binning, annotating and read counting. The resulting alignments are split into equally-sized bins of 50, 100 or 500 kb in size (i.e. the hg19 reference genome is split into equally-sized windows of a size selected depending on the desired resolution). In the present examples, a resolution that leads to a minimum of 15 reads per bin per tumour copy (i.e. 15 reads per absolute copy number of the tumour at the location of the bin) on average was selected in all cases. Note that the method performed equally well for all sizes of bins tested, the only difference being that larger bins lead to lower resolution copy number profiles. Then, bins are annotated with replication timing, GC content, mappability and the % of characterised nucleotides within a bin (i.e. non NNNNN bases). GC content and the % of characterised nucleotides are annotations inherited from QDNAseq (Scheinin et al. 2014). Mappability tracks suitable for paired-end data (500mer, allowing for 2 mismatches) were generated using GenMap (v1.3.0). GenMap computes the uniqueness of k-mers for each position across the genome with a number of mismatches (i.e. calculating how unique a given sequence of length k in the genome is compared to all occurrences of k-mers in the genome). This measure of uniqueness provides an indication as to how easy it is to map a sequencing read to that location of the genome. Mappability tracks with 50-mers are suitable for establishing mappability for single-end sequencing, whereas tracks of 500-mers are suitable for paired- end sequencing. We generated custom paired-end mappabilities for the hg19 human genome reference using GenMap (v.1.3.0) (Pockgrandt et al. 2020, available at github.com / cpockrandt / genmap), bedGraphToBigWig (LICSC Genome Browser Utilities) and the calculateMappability function from the QDNAseq R package (Scheinin et al. 2014). We tested 50mer, 150mer, 300mer, 400mer, 500mer, 600mer and 700mer mappabilities, allowing for 2-4 mismatches. Given that no significant differences were observed between using 400- 700mer mappabilities, we chose 500mer with 2 mismatches as an universal solution for modern paired-end assays, as they tend to be of an average 150nt read length and a varying insert size - commonly around 200nt long. Replication timing was generated by computing the mean across 15 different cell-lines of read wavelet smoothed repli-seq timing data from UCSC (available at hgdownload.cse.ucsc.edu / goldenpath / hg19 / encodeDCC / wgEncodeUwRepliSeq / ). Other approaches can be used such as e.g. using replication timing data from a single cell line or control sample, such as e.g. a cell line that has high correlation with a sample to be analysed. The % of characterised nucleotides and the mappability annotations were used to apply a quality control step: any bin that met any of the following criteria was removed: GC content == NA, characterised bases != 100% or mappability < 95%. This is a pragmatic approach that removes less reliable bins and was found to leave enough data for high quality high resolution copy number profiles. Any one or more of these criteria could be omitted, and other bin quality filtering strategies could be used.

[0103] The annotated bins are then interrogated for overlaps with target regions using the BED file listing target regions, and divided in two groups: 1) off-target bins, with a zero overlap with targeted regions. In this case, reads are counted normally in the entire bin. 2) on-target bins, those with at least 1 overlap with a target region. In this case, reads anchored in the on-target regions are removed (i.e. removing every read in the BAM file that overlaps the regions specified in the BED file), generating a new BAM file containing only off-target reads. This is then scanned for counting. Thus, the width of the bins is not itself changed but signal from captured regions is removed. This process excludes pieces of information such as variant calling or gene-level copy number calling for target genes, which are seen as enriched regions in the genome that, if kept, would introduce very high levels of noise for the downstream analysis. Since most targeted panel designs are paired-end, on-target reads with a mate-pair falling outside of the target coordinates are excluded as well using the read names, which share a common part in both mates.

[0104] Finally, the “scanBam” function from Rsamtools is used to count the remaining piece of data (which are only off-target reads). Although the off-target reads are unevenly distributed, they are still scattered across the entire genome and they constitute the information from which genome-wide copy number profiles are extracted.

[0105] GC correction. In all instances, the method applies a local polynomial (LOESS) curve fitting and correction approach to correct for biases in GC content. Two different types of correction were implemented depending on CIN levels. Both types use GC LOESS fitting and corrections:

[0106] 1) Stable tumors (NO-CI N - samples with only one main copy number state): single LOESS on the entire dataset at once. All the data points are fit at once based on GC content using this formula: loess(raw counts ~ gc). Then each bin read count is corrected by subtracting the difference between the fitted value for the GC value of the bin and the median value. This subtracts the effect of GC content, removing all the correlation between counts and GC.

[0107] 2) Unstable tumors (CIN): copy number aware GC correction, multiple LOESS, one per copy number state. Samples with high CIN display notable differences in the copy number status across the genome. For that reason, the method performs a copy number aware GC correction, avoiding an undesired shrinking effect to the median value of the raw counts that would eliminate true biological variation if a single LOESS were applied. Thus, this method does a better job accounting for true biological variation when correcting GC content, which ultimately helps for a more robust absolute fitting as it leads to a better estimation of clonality, purity and ploidy. To the best of the inventors knowledge, no other bulk copy number method does this. Copy number aware correction can be done in SCOPE (single cell method). Copy number aware correction requires each data point to be labelled to their copy number state in order to perform a LOESS fit separately for each of the copy number states. Since the final true copy number state is not known at this stage, the method tries to infer it by running an initial GC correction followed by segmentation and clustering using the segment values assigned to each data point (bin) as input. More specifically, the once GC corrected read counts are used as input to a segmentation algorithm, then bins are assigned the value of the segment that they correspond to, and bins are clustered using their assigned segment values. The clustering part employs a modified version of k-means from the “Ckmeans.1d.dp" R package, which automatically selects the best number of clusters from a range of 1 to 25 using the Bayesian information criterion (BIC). Then, two constraints are applied in an iterative process until both are fulfilled: 1) every cluster should have at least a minimum of 5 data points so that the LOESS can be performed, and 2) the minimum difference between all cluster centres must be at least 2. At each iteration, if there is a pair of clusters such that the Euclidian distance between their centres is less than 2 and / or a cluster that has fewer than 5 data points, one cluster is removed by rerunning the clustering with a maximum number of clusters reduced by 1 compared to the previous iteration.

[0108] In the present examples, samples with less than 1 copy number changes different from a diploid state were considered “NO-CI N”. Samples with 1 or more copy number changes different from a diploid state were considered “CIN”. This was assessed by reviewing the results of the clustering and if there was more than 1 cluster identified, copy number aware GC correction was performed.

[0109] Off-target capture bias score correction. Off-target reads are assumed to be uniformly distributed across the genome, with peaks emerging as natural perturbators of this distribution, inherent to targeted sequencing data. The left-most pairwise distances between every pairs of read in each bin is computed. The distribution of left-most pairwise distances of every read in a bin generates a left-triangular distribution (where the minimum distance value, usually zero, equals the mode) with a consistent shape in normal, unbiased bins. The mode is usually the minimum distance value observed, which is usually zero, because sequencing data typically comprises many sets of reads starting from the same positions due to fragmentation bias. Indeed, genomic DNA is typically fragmented in the process of preparing a sequencing library, and cell free DNA is naturally fragmented, both fragmentation processes being nonrandom (i.e. biased). By contrast, off-target bins with an artificially high read count (peaks) display some perturbations that can be measured and are proportional to the read count (see Figure 8A). Three different approaches were developed to capture this and correct non- uniformly distributed reads across bins in capture-based DNA targeted sequencing assays.

[0110] 1) The histogram-based distance score (see Figure 8A). Here, we first compute the density function using the observed minimum and maximum bin distances (or bin size+average length of a read) as parameters from a left-triangular distribution using the dtri” function from the “triangulr” R package (cran.r-project.org / web / packages / triangulr / ). We then generate a histogram from the observed pairwise distances and compute a vector of midpoints of the histogram bins. Using the left triangular density we estimate the expected values for each of the histogram bin mid-points then subtracting the resulting vector from the observed histogram values we obtain a vector that reflects the observed deviations from the theoretical left- triangular distribution. Converting this to a vector of absolute values and calculating its mean, provides the final score that can be used for noise correction. 2) The probability-based distance score (see Figure 8B). In this case, we use the same min, max and mode parameters that are used in the histogram-based score (i.e. mode=observed minimum distance, min=observed minimum distance, max=observed maximum distance in the bin, or bin size + average length of a read) but using the matrix of pairwise distances directly as input to the “ptri” function to compute the distribution function, therefore calculating the probability of each pair in the bin to be sampled from a left-triangular distribution. We subtract 1 to the mean of the resulting vector of probabilities and that is the final probabilitybased score.

[0111] 3) Statistical metric of distribution difference: This score is generated for each bin by comparing the vector of all observed leftmost pairwise distances within the bin to the theoretical distances as if they were sampled from a right-skewed triangular distribution with min=0, mode=0 and max=binsize+average read length (max would be 50150 for a 50kb bin, as a result of 50000+150). This comparison can be done with any statistical metric that can compare two distributions, such as the Kolmogorov-Smirnov test, the DTS test, and many others. In the data shown on Figure 9C, the DTS statistic is used. This is a test statistic that compares two Empirical cumulative distribution functions (ECDFs) by looking at the reweighted Wasserstein distance between the two (see Dowd 2020). This is calculated here with the funcion “dts_stat” from the “twosamples” Rpackage (search. r- project.org / CRAN / refmans / twosamples / html / two_sample.html). This R package contains functions to run statistical tests capable of comparing two distributions - such as Kolmogorov- Smirnov test, DTS test, and many others. Any of these could be used but the DTS statistic is believed to be better powered, and was shown to show good results for the particular purpose at hand (see Fig. 9C). In this particular case, to obtain the DTS score the dts_stat function takes as input the vector of observed distances within the bin and a vector generated randomly using the “rtri” function from the “triangulr” R package (cran.r- project.org / web / packages / triangulr / triangulr.pdf) with the above mentioned min, max and mode parameters, being the theoretical vector equal to the length of the observed vector. As explained in Dowd 2020, the DTS test compares two empirical cumulative distribution functions (ECDFs) generated out of the two input vectors and then computes the reweighted Wasserstein distance between the two. Formally, if E is the ECDF of sample 1 , F is the ECDF of sample 2, and G is the ECDF of the combined sample, then:

[0112] The DTS statistic shows high correlation with GC corrected counts (see Fig. 9C), similar to the histogram-based score or the probability-based score, and therefore can be used to correct artificially high read count bins. Any of the scores described above can be used for another round of LOESS correction using already GC-corrected counts. In other words, a LOESS model can be fitted to the distribution of read counts per bin (optionally already GO corrected), as a function of the score. The observed value for each bin is then corrected by subtracting the difference between the fitted value for the score of the bin and the median value across bins (equivalent to bending the curve to the horizontal line at the raw per bin count median value). Both scores were shown to produce similar results. The histogram-based score is used by default as it showed an overall slightly better performance, but the option is given to the user for different parameterization. In samples with very low read counts, the probability-based score may be advantageous as it does not require fitting a density distribution I histogram to the data, which may be difficult when there is not much data. A histogram-based score or probability-based score may be obtained for any bin that contains at least 10 reads. For bins that contain fewer than 10 reads, the score may be set to a default equal to the minimum computed score for the sample.

[0113] Treatment of on-target bins. On-target bins are defined as those with any overlap with on- target regions. These are very few in small panels, around 1-10% percent of the total number of bins, but approximately half of the total number of bins in a WES sample. For these bins, for each on-target region where reads have been removed, a pseudocount is then added which is proportional to the percentage of target region overlap relative to the whole bin length. The pseudo count is estimated from the remaining reads in the bin after GO correction, capture bias correction and replication timing correction (if performed). Specifically, the pseudo-count is calculated as pseudo-count=(%overlap with the bed file)*(count for the bin after all corrections applied). Thus, the final count will be (count for the bin after all corrections applied)+pseudocount.

[0114] Replication timing correction. Since not all tumour samples require replication timing, only those that exhibit a Pearson correlation (r) between the bin read count values and replication timing score equal or greater to 0.15 are corrected. As for the other biases, the same correction procedure was applied: LOESS fitting and transformation.

[0115] Segmentation and absolute copy number fitting. Corrected bin read counts were subjected to segmentation, using a modified version of the circular binary segmentation (CBS) from the DNAcopy R package. In particular, the method uses the “segment” function from the DNAcopy R package, based on a circular binary segmentation approach. The following parameters are selected: transformFun- 'sqrt", alpha = 1e-10 , undo. splits = “sdundo” , undo.SD = 1. These parameters are believed to be optimal for a sequencing depth of at least 15 reads per bin per chromosome copy (see e.g. Scheinin et al. 2014). Once all possible corrections have been applied, the resulting relative copy number profile is then subjected to absolute fitting, a process in which we manually select the most likely purity-ploidy combination in the range 0.05<purity<1 and 1 ,5<ploidy<8, as follows: where purity is the fraction of tumour cells in the sample, r is the read counts, and d is a constant proportional to the read depth and the average absolute copy number of tumour cells in the sample (ploidy d = - (ploidyxpurity -+ -2x(l-purityy)

[0116] The optimal mathematical solution was considered the one with the lowest root mean squared deviation (RMSD) between the non-rounded and the rounded copy numbers for all bins, meaning the highest clonality.

[0117] Results

[0118] A new method to perform copy number calling from targeted capture panels is proposed, illustrated on Figure 7. Paired-end raw reads are aligned to a reference genome. The resulting alignments are split into equally-sized bins. Sizes of 50, 100 or 500 kb were tested with similar results. The bins are then annotated with information including GC content, mappability and replication timing. The annotated bins are then divided in two groups depending on whether they overlap with target regions: 1) off-target bins, with a zero overlap with targeted regions, and 2) on-target bins, those with at least 1 overlapping target region. For on target bins, pseudocounts are added as described above. The bins are then corrected for GC content using a single LOESS fitted on the entire dataset at once and then segmented. For samples with no chromosomal instability (i.e. mostly diploid samples) the samples will then proceed for capture bias score correction; for samples with high chromosomal instability clustering is performed and then a “copy number aware GC correction” is implemented where individual LOESS models are fitted to the data for each copy number state observed. Each copy number state is then corrected for GC content separately and segmented. This is illustrated on Figure 6. Following this, both CIN and NO-CI N have the read counts per bin corrected to account for biases due to off-target capture. Targeted sequencing assays generate artificially high off- target read counts in parts of the genome with high sequence similarity to the target regions. Thus, a new score per-bin is generated that quantifies the magnitude of this bias. This can be done using one of two methods (histogram-based distance score or probability-based distance score), as explained above and as illustrated on Figure 8. This is then used for a single LOESS fit and correction (i.e. correcting all bins at once). In some cases it was observed that replication timing introduced a systematic bias in the read count data. This is automatically identified and corrected in a similar way as GC content or capture bias correction (i.e. by fitting a LOESS model to the bin read count as a function of replication timing and adjusting the counts per bin to remove the bias estimated as the value of the LOESS model). Note that in the present workflow, individual univariate models are fitted to correct for GC content bias and capture bias (and replication timing if used). However, a multivariate model could be fitted to any combination of GC content bias, capture bias and replication timing bias. Finally, the counts for on-target bins are re-adjusted to account for the data that was removed. Conversion of on-target bins to off-target bins results in an artificially lower read count due to the gaps left by the overlaps exclusion. To solve this, a pseudo-count directly proportional to the size of the overlapping areas is added, essentially filling the gaps that the on-target areas left after their exclusion. For example, if 20% of the bin is covered by an on-target region, 0.2*read count will be added to the observed read count of the bin.

[0119] The above results in bin read counts that are essentially equivalent to whole genome sequencing ones. These can be considered trustworthy and can be analysed in a similar way as data resulting from whole genome sequencing. Therefore each sample is subjected to segmentation, and absolute copy number fitting (which aims to identify the most likely combination of purity and ploidy, with the lowest clonality), using methods known in the art.

[0120] Example 2 - Performance assessment

[0121] Since Shallow Whole Genome Sequencing (sWGS) is a current gold standard for low- resolution genome-wide copy number profiling, the method described in Example 1 was benchmarked by comparison to this assay, using a set of samples for which both sWGS and targeted sequencing data was available. For sWGS, Reads were aligned as single-end against the human genome assembly GRCh37 using BWA-MEM (vO.7.17, Li 2013). Duplicate reads were then identified and marked using samtools-markdup (v1.15). We used the QDNAseq R package to count reads within 50 kb bins. Bins mapped to centromeres and regions of undefined sequence in the reference genome hg19 were excluded. Read counts were then corrected for sequence mappability and GC content, followed by copy number segmentation. After the profile was segmented, we inferred absolute copy numbers as outlined in Example 1. Each sWGS and capture sequence copy number profile was compared in pairs by their rounded segment values, using four different metrics from the CNpare package (Chaves-Urbano et al. 2022): Pearson’s r, cosine similarity, Manhattan distance, Euclidean distance.

[0122] Examples are shown on Figure 9, which shows an overlaid sWGS- WES cancer cell line data (Fig. 9A) and an overlaid sWGS-TS0500 FFPE clinical sample data (Fig. 9B) processed according to the present methods for samples analysed in the dataset described in Table 2 (first two lines). The paired copy number profiles from the WES-sWGS lung cancer cell line and the TS0500-sWGS FFPE ovarian cancer sample differ in their genomes by 11.57% and 1 .39% respectively, indicating that the results obtained using the present methods applied to the capture data is highly similar to sWGS data.

[0123] Figure 10 shows results of applying the methods to WES data from 3 lung cancer cell lines, with and without the bias correction to demonstrate the benefit of bias correction. Across all metrics including Pearson’s r, Manhattan distance, Euclidean distance and Cosine similarity, the inclusion of bias correction yields a copy number profile derived from WES that is more similar to the sWGS than without bias correction.

[0124] Figure 11 shows the results of applying the methods disclosed, in addition to two other genome-wide copy number profiling methods publicly available for targeted sequencing panels (CopywriteR and CNVkit), to 26 FFPE ovarian cancer samples profiled using the TSG500 gene panel assay (see emea.illumina.com / content / dam / illumina / gcs / assembled- assets / marketing-literature / trusight-oncology-500-data-sheet-m-gl-00173 / trusight-oncology- 500-and-ht-data-sheet-m-gl-00173.pdf) and sWGS. In addition, these samples were also profiled using the TS0500+HRD panel (see www.illumina.com / content / dam / illumina / gcs / assembled-assets / marketing-literature / tso500- hrd-data-sheet-m-gl-00748 / tso500-hrd-data-sheet-m-gl-00748.pdf), which has approximately 25000 additional capture regions interrogating single-nucleotide polymorphisms (SNPs) such that the copy number profile can be determined using a SNP based approach (similar to ASCAT). Pearson’s r of each method’s output compared to sWGS gold standard, showed the results for TS0500 using CopyRight to have superior performance to CopywriteR and CNVkit and comparable performance to TSG500+HRD.

[0125] Example 3 - Application to chromosomal instability signatures

[0126] The methods described herein enable the quantification of signatures of chromosomal instability from samples that have been sequenced using targeted sequencing assays. This typically requires genome-wide copy number profiles at a resolution of 50kb at least, which was not available for this type of samples prior to the present invention.

[0127] After a copy number profile has been obtained as explained in Example 1 , signatures of chromosomal instability can be calculated as described in Macintyre et al., 2018 and in Drews et al. 2022. This was simply not possible previously. As explained in e.g. Drews et al. 2022, these signatures can be used to predict drug response, identify new drug targets and understand tumour aetiology. Figure 12 shows a heatmap of the raw CIN signature activities obtained with both WES and sWGS of the three lung adenocarcinoma cell lines. The samples are represented in rows and the CIN signatures in columns. The heatmap shows a hierarchical clustering applied to the rows which results in each pair grouped together, indicating similarity in the CIN signatures obtained with both approaches (WES and sWGS).

[0128] Figure 13 shows the comparisons of the performance of extracting CIN signatures from TSG500 (with CopyRight), TSG500+HRD and sWGS using all the clinical FFPE ovarian cancer samples detailed in Table 2. Fig. 13A shows stacked barplots of the raw signature activities. Fig. 13B shows the cosine similarities against sWGS of TSG500 and TSG500+HRD. These results demonstrate that TSG500 (with CopyRight) performs comparably to TSG500+HRD when extracting CIN signatures, with many of the samples showing high similarity (cosine > 0.8) with sWGS.

[0129] Example 4 - Discussion

[0130] The present examples present and demonstrate a new method to analyse data from capturebased targeted sequencing technologies. The method was demonstrated using a number of different capture-based targeted sequencing technologies. These include: Whole Exome Sequencing (WES), more specifically the SureSelect Human All Exon protocol (Agilent), widely used in academia for small variant discovery in any protein-coding exon, and the TruSight Oncology 500 assay (Illumina) which is regulatory approved in the cancer clinical setting and provides information for 523 genes and a set of biomarkers with clinical relevance such as TMB, MSI or HRD. Other used and tested protocols were the TruSight Tumor 170 assay, which targets 171 genes and custom panels designs. Amplicon-based technologies could not be used since they do not generate off-target reads.

[0131] The method is applicable to many different sample types (e.g. FFPE, fresh frozen, cfDNA).

[0132] In the context of liquid biopsies, the method might be useful not only to generate copy number profiles but also to estimate % ctDNA from panels, similarly to what ichorCNA algorithm would do in sWGS. This may be particularly useful as there are already targeted liquid biopsy assays that regulatory approved and used in the clinic, and tumor fraction is emerging as a strong indicator of response to therapy, prognosis, and relapse. Thus, a method of obtaining improved estimates of ctDNA from targeted sequencing assays in the context of cfDNA would represent an invaluable tool in the clinical cancer care. References

[0133] Talevich E, Shain AH, Botton T, Bastian BC (2016) CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12(4): e1004873.

[0134] Kuilman, T., Velds, A., Kemper, K. et al. CopywriteR: DNA copy number detection from off- target sequence data. Genome Biol 16, 49 (2015).

[0135] Boeva V, Popova T, Bleakley K, Chiche P, Cappo J, Schleiermacher G, Janoueix-Lerosey I, Delattre O, Barillot E. Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data. Bioinformatics. 2012 Feb 1 ;28(3):423-5.

[0136] Jiang Y, Oldridge DA, Diskin SJ, Zhang NR. CODEX: a normalization and copy number variation detection method for whole exome sequencing. Nucleic Acids Res. 2015 Mar 31 ;43(6):e39.

[0137] Raine KM, Van Loo P, Wedge DC, Jones D, Menzies A, Butler AP, Teague JW, Tarpey P, Nik-Zainal S, Campbell PJ. ascatNgs: Identifying Somatically Acquired Copy-Number Alterations from Whole-Genome Sequencing Data. Curr Protoc Bioinformatics. 2016 Dec 8;56:15.9.1-15.9.17.

[0138] Amarasinghe, K.C., Li, J., Hunter, S.M. et al. Inferring copy number and genotype in tumour exome data. BMC Genomics 15, 732 (2014).

[0139] Shen R, Seshan VE. FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing. Nucleic Acids Res. 2016 Sep 19;44(16):e131.

[0140] Soong, D., Stratford, J., Avet-Loiseau, H. et al. CNV Radar: an improved method for somatic copy number alteration characterization in oncology. BMC Bioinformatics 21 , 98 (2020).

[0141] Langmead, B., Trapnell, C., Pop, M. et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009).

[0142] Carter, S., Cibulskis, K., Helman, E. et al. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol 30, 413-421 (2012).

[0143] Favero F, Joshi T, Marquard AM, Birkbak NJ, Krzystanek M, Li Q, Szallasi Z, Eklund AC. Sequenza: allele-specific copy number and mutation profiles from tumor sequencing data. Ann Oncol. 2015 Jan;26(1):64-70.

[0144] Adalsteinsson, V.A., Ha, G., Freeman, S.S. et al. Scalable whole-exome sequencing of cell- free DNA reveals high concordance with metastatic tumors. Nat Commun 8, 1324 (2017).

[0145] Drews, R.M., Hernando, B., Tarabichi, M. et al. A pan-cancer compendium of chromosomal instability. Nature 606, 976-983 (2022).

[0146] Macintyre, G. et al. Copy number signatures and mutational processes in ovarian carcinoma. Nat. Genet. 50, 1262-1270 (2018).

[0147] Milbury CA, Creeden J, Yip WK, Smith DL, Pattani V, Maxwell K, Sawchyn B, Gjoerup O, Meng W, Skoletsky J, Concepcion AD, Tang Y, Bai X, Dewal N, Ma P, Bailey ST, Thornton J, Pavlick DC, Frampton GM, Lieber D, White J, Burns C, Vietz C. Clinical and analytical validation of FoundationOneOCDx, a comprehensive genomic profiling assay for solid tumors. PLoS One. 2022 Mar 16;17(3):e0264138. Chen Zhao, Tingting Jiang, Jin Hyun Ju, Shile Zhang, Jenhan Tao, Yao Fu, Jenn Lococo, Janel Dockter, Traci Pawlowski, Sven Bilke. TruSight Oncology 500: Enabling Comprehensive Genomic Profiling and Biomarker Reporting with Targeted Sequencing. bioRxiv 2020.10.21.349100

[0148] C. Heydt, R. Pappesch, K. Stecker, J. Neumann, R. Buettner, S. Merkel bach- Bruse. Evaluation of the TruSight Tumor 170 (TST170) assay and its value in clinical research. Annals of oncology, volume 29, supplement 6, vi7-vi8, September 2018.

[0149] Wang, R., Lin, D-Y, Jiang Y. SCOPE: A Normalization and Copy-Number Estimation Method for Single-Cell DNA Sequencing. Cell Systems Vol. 10, Issue 5, P445-452, May 20, 2020.

[0150] Christopher Pockrandt and others, GenMap: ultra-fast computation of genome mappability, Bioinformatics, Volume 36, Issue 12, June 2020, Pages 3687-3692.

[0151] Mehran Karimzadeh and others, llrnap and Bismap: quantifying genome and methylome mappability, Nucleic Acids Research, Volume 46, Issue 20, 16 November 2018, Page e120.

[0152] Scheinin I, Sie D, Bengtsson H, van de Wiel MA, Olshen AB, van Thuijl HF, van Essen HF, Eijk PP, Rustenburg F, Meijer GA, Reijneveld JC, Wesseling P, Pinkel D, Albertson DG, Ylstra B. DNA copy number analysis of fresh and formalin-fixed specimens by shallow wholegenome sequencing with identification and exclusion of problematic regions in the genome assembly. Genome Res. 2014 Dec;24(12):2022-32.

[0153] Macintyre G, Ylstra B, Brenton J.D. Sequencing Structural Variants in Cancer for Precision Therapeutics. Trends in Genetics. Vol. 32, Issue 9, P530-542, September 2016.

[0154] Li H, Durbin R. 2009. Fast and accurate short read alignment with Burrows- Wheeler transform. Bioinformatics 25: 1754-1760.

[0155] Li H. (2013) Aligning sequence reads, clone sequences and assembly contigs with BWA- MEM. arXiv:1303.3997v2

[0156] Blas Chaves-Urbano and others, CNpare: matching DNA copy number profiles, Bioinformatics, Volume 38, Issue 14, July 2022, Pages 3638-3641

[0157] Dowd, Connor, A New ECDF Two-Sample Test Statistic. arXiv:2007.01360v1. 2 July 2020.

[0158] All references cited herein are incorporated herein by reference in their entirety and for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0159] The specific embodiments described herein are offered by way of example, not by way of limitation. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Any sub-titles herein are included for convenience only and are not to be construed as limiting the disclosure in any way. Unless context dictates otherwise, the descriptions and definitions of the features set out above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments which are described. The methods of any embodiments described herein may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described above.

[0160] Throughout the specification and claims, the following terms take the meanings explicitly associated herein, unless the context clearly dictates otherwise. The phrase “in one embodiment” as used herein does not necessarily refer to the same embodiment, though it may. Furthermore, the phrase “in another embodiment” as used herein does not necessarily refer to a different embodiment, although it may. Thus, as described below, various embodiments of the invention may be readily combined, without departing from the scope or spirit of the invention. It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%. Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. Other aspects and embodiments of the invention provide the aspects and embodiments described above with the term “comprising” replaced by the term “consisting of” or “consisting essentially of”, unless the context dictates otherwise, “and / or” where used herein is to be taken as specific disclosure of each of the two specified features or components with or without the other. For example “A and / or B” is to be taken as specific disclosure of each of (i) A, (ii) B and (iii) A and B, just as if each is set out individually herein. The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof.

Claims

CLAIMS1. A computer implemented method of analysing genomic sequencing data, the method comprising: obtaining genomic sequence data from one or more samples, comprising a plurality of DNA sequencing reads; determining an observed read count, from the genomic sequence data, for each of a plurality of genomic bins; and determining a corrected read count for one or more of the plurality of genomic bins, the corrected read count based on the observed read count and one or more estimates of biases in read counts as a function of a variable associated with said respective biases, wherein the genomic sequencing data is targeted genomic sequencing data, the biases comprise capture bias and the variable associated with said bias is a capture bias score that is indicative of a difference between an observed read location metric distribution in the bin and an expected read location distribution metric in the bin.

2. The method of claim 1 , wherein an estimate of a bias in read counts as a function of a variable associated with said bias is estimated by fitting a model between (i) the observed read counts in the plurality of genomic bins, and (ii) the value of the bias variable for the plurality of genomic bins.

3. The method of claim 2, wherein the model is a curve-fitting model, optionally a regression model, preferably a local polynomial regression model.

4. The method of claim 2 or claim 3, wherein determining a corrected read count for a bin comprises subtracting from the observed read count a correction factor that is equal to the difference between: (i) the value of the model at the coordinate corresponding to the value of the bias variable for the bin, and (ii) a reference value, optionally wherein the reference value is a summary value derived from the observed read counts for the plurality of bins.

5. The method of any preceding claim, wherein an expected read location metric distribution in the bin is a read location metric distribution that corresponds to a uniform distribution of locations of reads in the bin.

6. The method of any preceding claim, wherein the locations of reads in a bin is represented by a metric of pairwise distance between pairs of reads in the bin, and / or wherein a capture bias score is a variable that captures a departure between an observed distribution of pairwise distances between reads in a bin and an expected distribution of pairwise distances between reads in the bin.

7. The method of claim 6, wherein a pairwise distance between reads is a left-most pairwise distance, a right-most pairwise distance or a middle pairwise distance.

8. The method of claim 6 or claim 7, wherein an expected distribution of pairwise distances between reads in a bin is a left triangular distribution with: a mode of 0 or the minimum observed pairwise distance in the bin, a minimum of 0 or the minimum observed pairwise distance in the bin, and a maximum of the maximum observed pairwise distance in the bin, the length of the bin, the sum of length of the bin and the length of a read, or any value between the length of the bin and the sum of length of the bin and the length of a read; optionally wherein the expected distribution of pairwise distances between reads in a bin is approximated by sampling from said left triangular distribution a number of observations corresponding to the number of observed pairwise distances between reads in the bin.

9. The method of any of claims, wherein the capture bias score is a variable that summarises across a plurality of windows of values of the read location metric, the absolute difference between the frequency for the window in the observed read location metric distribution of the bin and the frequency for the window in the expected read location metric distribution of the bin.

10. The method of any of claims 1 to 8, wherein the capture bias score is a variable that summarises across a plurality of observed values of the read location metric in the bin, the probability that the observed read location metric is sampled from the expected read location metric distribution of the bin.

11. The method of any of claims 1 to 8, wherein the capture bias score is a statistical metric that compares two samples to determine whether they come from the same distribution, optionally wherein the statistical metric is a DTS statistic.

12. The method of any preceding claim, wherein the genomic sequence data comprises on- target reads that are at least partially mapped to captured regions and off target reads that arenot mapped to captured regions, and the observed read counts are obtained using the off- target reads, and / or wherein the method comprises selecting reads that are not mapped to captured regions and obtaining the observed read counts using the selected reads.

13. The method of any preceding claim, wherein the method comprises obtaining the plurality of genomic bins by: applying a non-overlapping fixed width window along genomic coordinates for one or more genomic regions that the reads map to, thereby obtaining a plurality of fixed width bins.

14. The method of claim 12 or claim 13, wherein determining a corrected read count for one or more of the plurality of genomic bins comprises determining a corrected read count for a bin comprising one or more on-target reads by: obtaining a first corrected read count based on the observed off target read count for the bin and one or more estimates of biases in read counts as a function of a variable associated with said respective biases, and obtaining a second corrected reads count by adding a pseudocount to the first corrected read count, wherein the pseudocount is proportional to the first read count and the proportion of the bin that overlaps with a captured region.

15. The computer-implemented method of any preceding claim, wherein the biases further comprise GC content bias and the method further comprises obtaining an estimate of GC content bias in read counts as a function of GC content by: identifying a plurality of sets of bins, all bins in a set having the same estimated copy number, and fitting a model between (i) the observed read counts in one of the sets of bins comprising the bin for which a corrected read count is determined, and (ii) the value of GC content for the bins in the set.

16. The method of claim 15, wherein identifying a plurality of sets of bins where all bins in a set have the same estimated copy number comprises clustering the raw read counts, corrected read counts or segment read counts associated with the bins.

17. The computer-implemented method of any preceding claim, wherein the targeted sequencing data is capture-based sequencing data, optionally from whole exome sequencing or targeted panel sequencing.

18. The computer-implemented method of any preceding claim, wherein a sample is a tissue sample, a FFPE tissue sample, a fresh frozen tissue sample, or a liquid biopsy sample.

19. The computer-implemented method of any preceding claim, wherein a sample is a sample comprising tumour cells or genetic material derived therefrom, optionally a tumour biopsy or sample comprising cell free DNA.

20. The computer-implemented method of claim 19, wherein the one or more samples do not comprise a matched normal sample, or wherein the one or more samples are one or more samples comprising tumour cells or genetic material derived therefrom.

21. The computer-implemented method of any preceding claims, wherein the biases further comprise replication timing bias and the variable associated with said bias is a replication timing, optionally wherein the method comprises determining whether the observed read counts for the plurality of bins are significantly associated with replication timing, and correcting the read counts for replication timing bias when the observed read counts for the plurality of bins are significantly associated with replication timing.

22. The computer-method of any preceding claim, further comprising obtaining corrected read counts for each of the plurality of bins, and further comprising obtaining a copy number profile from the corrected read counts for the plurality of bins, optionally by segmenting and absolute copy number fitting, and / or further comprising obtaining information derived from a copy number profile obtained from the corrected read counts, optionally wherein the information comprises the presence of one or more signatures of chromosomal instability, and / or the percentage of tumour cells or tumour DNA in the sample.

23. A method of providing a copy number profile or genomic feature derived therefrom for a sample, the method comprising: obtaining genomic sequence data from the sample, comprising a plurality of DNA sequencing reads obtained using a targeted genomic sequencing assay; analysing the data using the method of any of claims 1 to 22; and obtaining a copy number profile and / or a genomic feature derived therefrom, using the corrected read counts, optionally wherein the method comprises obtaining a tumour copy number profile and a percentage of tumour DNA or tumour cells with the tumour copy number profile.

24. A system comprising: a processor; and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of any method described herein, such as a method according to any of claims 1 to 23.

25. One or more non-transitory computer readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of any method described herein, such as a method according to any of claims 1 to 23.