Genome sequencing analysis

By correcting for capture bias in off-target reads, the method enhances the use of off-target data in targeted sequencing, achieving accurate genome-wide copy number profiles comparable to shallow whole-genome sequencing.

JP2026525380APending Publication Date: 2026-07-29テイラー バイオ リミテッド +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
テイラー バイオ リミテッド
Filing Date
2024-07-19
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing methods for analyzing genomic sequence data, particularly from targeted sequencing, fail to effectively utilize off-target reads and introduce noise, leading to gaps in genome coverage and difficulty in achieving high-resolution copy number determinations.

Method used

A method that corrects for capture bias in off-target regions by estimating the discrepancy between observed and expected read distributions, allowing for the use of off-target reads to generate robust genome-wide copy number profiles without discarding potentially useful data.

Benefits of technology

The method produces results comparable to shallow whole-genome sequencing, improving data utilization and accuracy in copy number profiling, applicable to both targeted and non-targeted sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026525380000001_ABST
    Figure 2026525380000001_ABST
Patent Text Reader

Abstract

Analyzing target genome sequencing data containing multiple sequencing reads. [Solution] A method is provided which includes determining an observed read count for each of a plurality of genome bins from genome sequence data, and determining a corrected read count for one or more of the plurality of genome bins, wherein the corrected read count is determined based on the observed read count and one or more estimates of the bias in the read count as a function of variables associated with each of the biases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field of the present disclosure The present disclosure relates to methods for analyzing genomic sequencing data, particularly methods for analyzing genomic sequence data from targeted sequencing. The present disclosure also relates to methods for obtaining copy number profiles derived from such non-targeted sequencing data.

Background Art

[0002] Background Genome-wide copy number profiles are essential in cancer genomics to provide information on the level of chromosomal instability (CIN) of tumors, which is a characteristic of cancer. Since most computational methods for CIN quantification rely on genome-wide copy number profiles, it is essential to know not only the copy number status of specific driver genes but also the rest of the genome.

[0003] However, whole-genome sequencing data is expensive to obtain, so capture-based DNA target panels are routinely used for small variant detection and very limited gene-level copy number profiling in cancer clinical care. This utilizes information from sequence reads (on-target reads) mapped to the capture region and typically ignores off-target sequence reads. Kuilman et al. (2015) proposed a method for performing copy number detection using off-target reads by detecting and removing genomic regions where sequence reads are enriched using a peak detection algorithm developed for ChIPseq data. However, this method still leaves unwanted noise in the data, creates gaps in the genome by discarding the data, and makes it difficult to make higher-resolution copy number determinations.

[0004] Therefore, there is a need for an improved method for analyzing targeted sequencing data.

Summary of the Invention

[0005] overview The inventors recognize that capture-based target sequencing panels can be a valuable data source for analyses that typically rely on genome-wide copy number profiles, because this data includes information scattered across the entire genome. Therefore, the inventors embarked on designing a method that could process this data to generate robust genome-wide copy number profiles. Briefly, the method corrects for the bias introduced into read counts in off-target regions by the capture process. This correction does not remove off-target reads, but instead estimates the discrepancy between the observed distribution of read locations within the genome bin and the expected distribution of read locations within the genome bin in the absence of capture bias. The inventors have shown that this novel method yields results comparable to the current gold standard shallow whole-genome sequencing (sWGS) samples when applied to various next-generation sequencing (NGS) protocols. The inventors further benchmark this tool using comparative computation methods, demonstrating improvements in these areas. This method further has the advantage of not requiring a step to filter out corresponding normal tissue or noisy data, thereby making it significantly more widely applicable and avoiding the process of discarding potentially useful data (i.e., retaining more informational content in the data).

[0006] Accordingly, according to the first aspect, a computer method for analyzing genome sequencing data is provided, comprising: obtaining genome sequence data containing multiple DNA sequencing reads from one or more samples; determining observed read counts for each of a plurality of genome bins from the genome sequence data; and determining corrected read counts for one or more of the plurality of genome bins, wherein the corrected read counts are determined based on the observed read counts and one or more estimates of the bias in the read count as a function of a variable associated with each bias. The genome sequencing data is target genome sequencing data, the bias includes capture bias, and the variable associated with the bias is a capture bias score that indicates the difference between the observed read position metric distribution within a bin and the expected read position distribution metric within a bin.

[0007] The present invention also describes a computer method for performing genome sequencing data analysis, comprising: obtaining genome sequence data containing multiple DNA sequencing reads from one or more samples; determining raw read counts for each of a plurality of genome bins from the genome sequence data; and determining a corrected read count for at least one of the plurality of genome bins, wherein the corrected read count is determined based on the raw read count and an estimate of the bias in the read count as a function of a variable associated with the bias. The bias is a bias associated with GC content, the variable associated with the bias is GC content, and the estimate of the GC content bias in the read count as a function of GC content is obtained by identifying a plurality of bin sets in which all bins in the set have the same estimated copy number; and fitting a model between (i) the observed read count in one of the bin sets containing the bins for which the corrected read count is determined and (ii) the GC content values ​​for the bins in the set. The method according to the second embodiment is applicable to any genome sequencing data, regardless of whether it is from target genome sequencing data or non-target genome sequencing data, and regardless of whether it is from bulk cell sequencing data or single cell sequencing data.

[0008] The method according to the first or second embodiment may have one or more of the following features:

[0009] The method may include determining corrected read counts for each of several genome bins. The method according to a second embodiment may include obtaining corrected read counts according to the first embodiment. The method according to the first embodiment may include obtaining corrected read counts according to the second embodiment. An estimate of the bias in the read count as a function of a variable associated with the bias can be estimated by fitting a model between (i) the observed read counts for several genome bins and (ii) the values ​​of the bias variable for several genome bins. The model may be a curve-fitting model. The model may be a regression model, such as a local polynomial regression model. Determining corrected read counts for a bin may include subtracting from the observed read count a correction coefficient equal to the difference between (i) the model value at coordinates corresponding to the values ​​of the bias variable for the bin and (ii) a reference value. The reference value may be a summary value derived from the observed read counts for several bins.

[0010] The expected read position metric distribution within a bin may be a read position metric distribution corresponding to a uniform distribution of read positions within a bin. The position of a read within a bin may be represented by a metric of the pairwise distance between pairs of reads within a bin. The capture bias score may be a variable that captures the deviation between the observed distribution of pairwise distances between reads within a bin and the expected distribution of pairwise distances between reads within a bin. The pairwise distance between reads may be the left-end pairwise distance, the right-end pairwise distance, or the center pairwise distance. The expected distribution of pairwise distances between reads within a bin may be a distribution with the same number of observations arranged in a left triangular distribution, where the mode is 0 or the smallest observed pairwise distance within a bin, the minimum is 0 or the smallest observed pairwise distance within a bin, and the maximum is the largest observed pairwise distance within a bin, the length of the bin, the sum of the length of the bin and the length of the reads (which may be taken as, for example, the average length of the reads), or any value between the length of the bin and the sum of the length of the bin and the length of the reads. The capture bias score may be a variable that summarizes the absolute difference between the frequency of a window in the bin's observed read position metric distribution and the frequency of a window in the bin's expected read position metric distribution, across multiple windows of read position metric values. The capture bias score may also be a variable that summarizes the probability that an observed read position metric was sampled from the bin's expected read position metric distribution, across multiple observations of the read position metric within a bin.

[0011] This method may include obtaining multiple genomic bins by applying a non-overlapping fixed-width window along genomic coordinates for one or more genomic regions to which reads are mapped, thereby obtaining multiple fixed-width bins. Genomic sequence data may include on-target reads that are at least partially mapped to capture regions and off-target reads that are not mapped to capture regions, and the observed read count is obtained using the off-target reads. This method may include selecting reads that are not mapped to capture regions and obtaining the observed read count using the selected reads. When using paired-end sequencing data, off-target reads may be reads that are not mapped to capture regions and reads that do not have mate reads that are mapped to capture regions. This method may include obtaining the processed bin width for a bin by removing any region that is a capture region from the fixed-width bin. This method may include obtaining a further corrected read count for a bin by adding an adjustment factor proportional to the difference between the processed bin width and the fixed-width bin width to the corrected read count. The adjustment factor may be proportional to the corrected read count and the difference between the processed bin width and the fixed-width bin width. Alternatively, the method may further include determining the corrected read count for one or more genome bins by: i) obtaining a first corrected read count based on the observed off-target read count for the bin and one or more estimates of the bias in the read count as a function of the variables associated with each bias; and (ii) obtaining a second corrected read count by adding a pseudocount to the first corrected read count, wherein the pseudocount is proportional to the proportion of the first read count that the bin overlaps with the capture region. For example, the pseudocount may be calculated as (proportion of overlap with the capture region file) * (count for the bin after all corrections have been applied). The capture region may be as defined in the BED file.

[0012] The method may further include obtaining GC-corrected read counts for a bin by fitting a model between (ii) the observed read counts and (ii) the GC content values ​​for the bins. The method may further include using the GC-corrected read counts to determine whether the observed read counts for each of several genomic bins indicate the presence of multiple bin sets, where all bins in a set have the same estimated copy number and the estimated copy numbers associated with each of the multiple sets are different from one another. Determining whether the observed read counts for each of several genomic bins indicate the presence of multiple bin sets may include segmenting the GC-corrected read counts for the bins to obtain segment values ​​associated with each bin, and clustering the segment values ​​for the bins, where the presence of more than one cluster indicates the presence of multiple bin sets. Determining whether the observed read counts for each of several genomic bins indicate the presence of multiple bin sets may include segmenting the GC-corrected read counts for the bins to obtain segment values ​​associated with each bin, and determining whether at least one segment has a different estimated copy number than the other segments, where the presence of such a segment indicates the presence of multiple bin sets. In an embodiment of the first aspect, the bias may further include a GC content bias, and the method may further include obtaining an estimate of the GC content bias in the read count as a function of GC content by identifying a plurality of bin sets in which all bins in the set have the same estimated copy number, and by fitting a model between (i) the observed read count in one of the bin sets containing the bins in which the corrected read count is determined, and (ii) the GC content values ​​for the bins in the set. In an embodiment of any aspect, identifying a plurality of bin sets in which all bins in the set have the same estimated copy number may include clustering the raw read count, corrected read count, or segmented read count associated with the bins.This method may further include determining the corrected read count based on the observed read count (which has been selectively corrected for capture bias or is uncorrected for capture bias) and an estimate of the bias in the read count as a function of GC content. The estimate of the GC content bias in the read count as a function of GC content is obtained by identifying multiple bin sets in which all bins in the set have the same estimated copy number, and by fitting a model between (i) the observed read count in one of the bin sets containing the bins for which the corrected read count is determined, and (ii) the GC content values ​​for the bins in the set. This may be done when it is determined that the observed read count for each of the multiple genome bins indicates the existence of multiple bin sets. The GC-corrected read count for multiple bins, obtained by fitting a model between the observed read count (for all of the multiple bins) and the GC content values ​​for the bins, may be used when it is determined that the observed read count for each of the multiple genome bins does not indicate the existence of multiple bin sets.

[0013] Identifying multiple bin sets in which all bins have the same estimated copy number may include, together in all bin sets, obtaining GC-corrected read counts for bins in the multiple bin sets by fitting a model between (ii) the observed read count and the GC content value for the bin, and clustering the GC-corrected read counts for bins, or the segment read counts or segment copies associated with bins obtained from the GC-corrected read counts. Therefore, this method may include segmenting the GC-corrected read counts for bins and obtaining segment values. Segment values ​​may be segment read counts or segment copies. Identifying multiple bin sets in which all bins have the same estimated copy number may include segmenting the (optionally GC-corrected) read counts for bins in the multiple bin sets, thereby identifying multiple genomic segments in which each segment has the same copy number state, and determining the segment read count or copy number for each segment. Identifying multiple bin sets in which all bins have the same estimated copy number may involve clustering the values ​​associated with each bin in the multiple bin sets, where this value for each bin is the segment read count or segment copy number associated with the segments that overlap with or contain the bin. For bins that overlap with multiple segments, the average segment read count or copy number across the multiple segments, the weighted average segment read count or copy number across the multiple segments (e.g., each segment weighted by the proportion of overlapping bins), or a selected segment read count or copy number (e.g., a segment read count or copy number representing the largest proportion of bins) may be used. Clustering may be performed using any clustering algorithm known in the art, such as k-means, self-organizing maps, hierarchical clustering, etc. Clustering may be performed using a predetermined maximum number of clusters. For example, the predetermined maximum number of clusters may be set to 25.Clustering may be performed using k-means and, for example, a Bayesian information criterion, and the number of clusters may be automatically selected. The automatically selected number of clusters may be within a predetermined range, such as 1 to 25. Clustering may involve applying a clustering algorithm to identify a number of initial clusters and iteratively removing one or more clusters using one or more criteria selected from the number of data points in a cluster and the distance between cluster centers. The process of iteratively removing one or more clusters may be performed by clustering the values ​​associated with each bin in multiple bin sets using a maximum number of clusters smaller than the maximum number of clusters used in the previous iteration (or the initial maximum number of clusters in the first iteration). The maximum number of clusters may be reduced by 1 when, at last, one cluster satisfies one or more criteria. The maximum number of clusters may be reduced by the number of clusters that satisfy one or more criteria. For example, clusters with a number of data points below a predetermined threshold, such as 5, may be selected to combine with another cluster, such as the nearest identified cluster. As another example, clusters with centers separated by a predetermined distance (e.g., Euclidean distance) may be selected and merged. The predetermined distance may be, for example, 2 in an embodiment where the value associated with the bins being clustered is the segment copy number. As another example, the initial cluster set may be iteratively refined by identifying clusters that satisfy one or more criteria selected from the fact that the number of data points in a cluster is below a predetermined threshold and the distance between the center of one cluster and the center of another cluster is below a predetermined threshold, and if clusters that satisfy one or more criteria are identified, clustering the values ​​associated with each of the bins in multiple bin sets using a maximum cluster number smaller than the initial maximum cluster number used to obtain the initial cluster set.The resulting cluster set can then be refined by identifying clusters that satisfy one or more criteria, and if clusters that satisfy one or more criteria are identified, by clustering the values ​​associated with each bin in multiple bin sets using a maximum number of clusters smaller than the maximum number of clusters used to obtain the cluster set in a preceding iteration. This method may involve fitting multiple models between (i) the observed read counts in each bin set and (ii) the GC content values ​​for the bins in the set. The number of models fitted may be equal to the number of clusters identified by clustering the raw read counts, corrected read counts, or segmented read counts associated with the bins. Thus, this method may involve identifying multiple clusters of a bin by clustering the raw read counts, corrected read counts, or segmented read counts associated with the bins and fitting a model to each of the multiple clusters, or by simultaneously fitting multiple models to the read counts for all bins in multiple bin sets, with the number of models equal to the number of clusters identified.

[0014] Target sequencing data may be, for example, capture-based sequencing data from whole exome sequencing or targeted panel sequencing. Target sequencing data may be paired-end sequencing data. Samples may be tissue samples, FFPE tissue samples, fresh-frozen tissue samples, or liquid biopsy samples. Samples may contain tumor cells or genetic material derived therefrom, such as tumor biopsies or samples containing cell-free DNA. One or more samples may not contain corresponding normal samples. One or more samples may all contain tumor cells or genetic material derived therefrom.

[0015] The bias may further include replication timing bias. The variable associated with the above bias is replication timing. This method may include determining whether the observed read counts for multiple bins are significantly associated with replication timing, and correcting the read counts for replication timing bias if the observed read counts for multiple bins are significantly associated with replication timing.

[0016] This method may include obtaining corrected read counts for each of several bins. This method may further include obtaining a copy number profile from the corrected read counts for the several bins. The copy number profile may be obtained from the corrected read counts for the several bins by segmentation and absolute copy number fitting. This method may further include obtaining information derived from the copy number profile. This information may include the presence of one or more chromosomal instability signatures and / or the proportion of tumor cells or tumor DNA in the sample.

[0017] As those skilled in the art will understand, the complexity of the operations described herein (at least due to the complexity of aligning sequencing data to a reference sequence and the amount of data typically generated by sequencing genetic material) is beyond the realm of mental activity. Therefore, unless otherwise indicated in the context (for example, when sample preparation or acquisition steps are described), all steps of the methods described herein are performed by computer.

[0018] Furthermore, the third embodiment describes a method comprising one or more of the following: sequencing one or more samples (for example, using a target genome sequencing assay) to obtain genome sequence data from one or more samples; preprocessing the genome sequence data (for example, aligning reads, filtering low-quality reads, obtaining bins, annotating bins, etc.); obtaining one or more samples (for example, from subjects, cell lines, etc.); and analyzing the genome sequencing data using a method of any embodiment of the above-described embodiment.

[0019] Any embodiment of any aspect of the method may include providing to the user (e.g., via a user interface) the results of any one or more steps of the method, such as corrected read counts for one or more bins, values ​​of bias variables for one or more bins, copy number profiles for the samples, and information derived from the copy number profiles for the samples.

[0020] Furthermore, a fourth embodiment will be described as a method for obtaining non-target genome sequencing data from a sample, the method comprising: obtaining genome sequence data from a sample, which includes a plurality of DNA sequencing reads obtained using a target genome sequencing assay; and analyzing the data using any of the embodiments of the first embodiment.

[0021] Furthermore, a fifth embodiment will be described of a method for providing a copy number profile or genomic features derived from a copy number profile of a sample, the method comprising: obtaining genomic sequence data from a sample, including a plurality of DNA sequencing reads obtained using a target genome sequencing assay; analyzing the data using a method of any embodiment of any of the aforementioned embodiments; and obtaining a copy number profile and / or genomic features derived from a copy number profile using a corrected read count. The method may also include obtaining a tumor copy number profile and the proportion of tumor DNA or tumor cells having a tumor copy number profile.

[0022] Also described according to the sixth aspect is a method for determining the proportion of tumor DNA or tumor cells in a sample, the method comprising: obtaining genomic sequence data from a sample, including a plurality of DNA sequencing reads obtained using a target genome sequencing assay; analyzing the data using a method of any embodiment of any of the aforementioned aspects; and obtaining a tumor copy number profile and the proportion of tumor DNA or tumor cells having a tumor copy number profile using a corrected read count.

[0023] In a further embodiment, one or more non-temporary computer-readable media are provided, which, when executed by one or more processors, cause one or more processors to perform a step of any of the methods described herein, such as the method according to any embodiment of any of the aforementioned embodiments.

[0024] In a further embodiment, a computer program is provided which, when the code is executed on a computer, causes the computer to perform a step of any of the methods described herein, such as the method according to any embodiment of any of the first to sixth embodiments.

[0025] According to a further aspect, there is provided a system including a processor and a computer-readable medium including instructions that, when executed by the processor, cause the processor to perform the steps of any of the methods described herein, such as the method according to any one of the first to sixth aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] A flowchart schematically showing a method for obtaining non-target genome sequencing data. [Figure 2] A flowchart schematically showing a method for providing and analyzing a genome-wide copy number profile. [Figure 3] An embodiment of a system for analyzing target DNA sequencing data and / or providing or analyzing a genome-wide copy number profile is shown. [Figure 4] GC correction by fitting a LOESS curve to raw count data as a function of GC content, where the GC ratio is calculated as (G + C) / (A + T + G + C)%, or a method of sequencing reads, is shown. [Figure 5] The results of the copy number recognition GC content correction process described herein are shown. The data shows that in the tumor genome, variable copy number states cause a shift in read depth, which is appropriately explained in embodiments of the methods described herein. [Figure 6] The process of GC content correction by a conventional method (top) and an embodiment of the present disclosure (copy number recognition) is shown. [Figure 7] An embodiment of a method according to the present disclosure used in an example is schematically shown. [Figure 8A] Calculation of a score for capture bias correction, which is a histogram-based score, is shown. [Figure 8B] Calculation of a score for capture bias correction, which is a probability-based score, is shown. [Figure 9A]The image shows a copy number profile obtained using the method described herein, applied to whole exome sequencing (WES) data (above) overlaid with a corresponding copy number profile obtained from the same sample by shallow whole-genome sequencing (sWGS). [Figure 9B] The copy number profiles obtained using the method described herein are shown, applied to TSO500 data overlaid with corresponding copy number profiles obtained from the same sample by shallow whole-genome sequencing (sWGS). [Figure 9C] This specification shows data illustrating the method described herein and the correction of counts using a statistical metric to compare the two distributions. The upper plot shows the GC-corrected count as a function of the statistical metric, and the lower plot shows the corresponding count after correction using the statistical metric. [Figure 9D] Figure 9C shows data illustrating the effect of applying the correction, and the resulting corrected copy number profile exhibits a low capture bias effect. [Figure 10] The data shows the distribution of Pearson r, Manhattan distance, Euclidean distance, and cosine similarity values ​​between copy number profiles obtained using the method of this disclosure applied to WES data from cancer cell lines with or without capture bias correction, and the corresponding copy number profiles obtained by sWGS. The data show that profile performance is significantly improved (higher Pearson correlation and cosine similarity, lower Manhattan and Euclidean distances) when capture bias correction and GC correction are performed using the method of this disclosure compared to when GC correction is performed alone. [Figure 11]Table 2 shows the distribution of Pearson's r when correlated with the sWGS gold standard for all FFPE clinical ovarian cancer samples included. These correlations of copy number segments to sWGS were performed using two different techniques, namely TSO500 and TSO500+HRD. Furthermore, this figure shows the results for TSO500 using not only CopyRight but also two other genome-wide copy number profiling methods (CopywriteR and CNVkit) that are publicly available for targeted sequencing panels. CopyRight performed better than CopywriteR and CNVkit and was comparable to TSO500+HRD. [Figure 12] This heatmap shows CIN signature activity obtained by WES and sWGS from two pairs of lung adenocarcinoma cell lines. Samples are represented in rows, and CIN signatures are represented in columns. Only signatures previously considered stable across the entire profiling technique are shown. The heatmap shows hierarchical clustering applied to the rows based on Euclidean distance. [Figure 13A] Table 2 compares the performance of TSO500 (using copyright), TSO500+HRD, and sWGS in extracting CIN signatures using all clinical FFPE ovarian cancer samples detailed in Table 2. These results indicate that TSO500 (using copyright) performs comparably to TSO500+HRD when extracting CIN signatures, and that many samples exhibit a cosine similarity greater than 0.8 compared to sWGS. A stacked bar graph of raw signature activity is shown. [Figure 13B]Table 2 compares the performance of TSO500 (using copyright), TSO500+HRD, and sWGS in extracting CIN signatures using all clinical FFPE ovarian cancer samples detailed in Table 2. These results indicate that TSO500 (using copyright) performs comparably to TSO500+HRD when extracting CIN signatures, and that many samples have a cosine similarity of greater than 0.8 compared to sWGS. The cosine similarity of TSO500 and TSO500+HRD to sWGS is shown. [Modes for carrying out the invention]

[0027] Detailed explanation In this disclosure, the following terms are used and are intended to be defined as set forth below.

[0028] This disclosure provides a method for obtaining non-target genome sequencing data from target sequencing data relating to a sample.

[0029] The term "non-target genome sequencing data" refers to data obtained using non-target sequencing techniques such as whole-genome sequencing or shallow whole-genome sequencing, or data equivalent to data that would have been obtained using such techniques. The data may take the form of read counts per genome bin. A genome bin is a region of the genome. A bin is typically defined by applying a fixed-size window to a reference genome. A read count is the number of sequencing reads mapped to a bin. A raw read count is the count of reads obtained as they are from the sequencing assay, after one or more selective filtering steps but before any bias correction. Bias correction can be GC bias correction, capture bias correction, and / or replication timing bias correction, as further described below. Therefore, the term "non-target genome sequencing data" may, in particular, refer to read counts obtained from target sequencing data that has been corrected for at least capture bias. Such data may be equivalent to data that would have been obtained from the same sample using, for example, sWGS. A first sequencing dataset and a second sequencing dataset may be considered equivalent if they yield copy number profiles with at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% Pearson's r across a given region (e.g., one or more chromosomes or the entire genome).

[0030] Targeted genome sequencing data refers to data obtained from genome sequencing assays that capture multiple predetermined regions of the genome (also referred to as “target regions” or “target locus”). Such assays may be referred to as “panel” sequencing, particularly when a specific gene of interest or a portion thereof is targeted. The methods described herein are applicable to any capture-based targeted sequencing technology. For example, the methods described herein are applicable to whole exome sequencing (WES) and gene panel sequencing assays, such as oncology panel assays or any panel sequencing assay with a custom panel design. Examples of gene panel sequencing assays include the TruSight Oncology 500 assay (see Illumina, Zhao et al. 2020), the TruSight Tumor 170 assay (Heydt et al. 2018), the Tempus xT assay (GTR test ID: GTR000558436.10, see www.ncbi.nlm.nih.gov / gtr / tests / 558436 / ), FoundationOne CDx (F1CDx, GTR test ID: GTR000593451.1, see Milbury et al. 2022 and www.ncbi.nlm.nih.gov / gtr / tests / 593451 / ), and FoundationOne Heme (GTR test ID: Examples include GTR000527977.1 (www.ncbi.nlm.nih.gov / gtr / tests / 527977 / ) and FoundationOne (GTR test ID: GTR000527976.2 (www.ncbi.nlm.nih.gov / gtr / tests / 527976 / )). The methods described herein are applicable to any target genome sequencing assay that generates off-target reads. This may include any capture-based assay. This may exclude amplicon-based techniques or any target sequencing techniques that do not generate off-target reads. The sequencing assay may use paired-end sequencing.Therefore, target genome sequencing data may include paired-end reads. Alternatively, single-end sequencing may be used. The methods described herein are considered applicable to both single-end and paired-end sequencing. Paired-end sequencing is relatively common in gene panel sequencing assays and advantageously reduces the challenges of mapping.

[0031] On-target reads are reads obtained from a target genome sequencing assay that map at least partially to a target region (i.e., DNA sequencing reads present in the target genome sequence data). Off-target reads are reads obtained from a target genome sequencing assay that do not map to a target region. The methods described herein may remove on-target reads (or correct the read count accordingly, e.g., by obtaining read counts for genome bins from which any target region is excluded). The methods described herein may not remove off-target reads unless based on quality control metrics specific to the reads rather than the distribution of the reads. For example, the methods described herein may not remove off-target reads that cause a peak in the read count profile or entire regions of the read count profile containing such peaks. Instead, the methods described herein may correct the read count in each of one or more genome bins based on the difference between the observed and expected distributions of the read location metric within the bin. Such a difference may be referred to herein as the “capture bias score,” as further described below.

[0032] The methods described herein may be used to obtain genome-wide copy number profiles for samples sequenced using targeted sequencing techniques. In other words, the methods described herein may be used to obtain genome-wide copy number profiles from targeted genome sequence data. The term “genome-wide” refers to data covering the entire genome or a substantial portion thereof, such as one or more entire chromosomes. As those skilled in the art will understand, sequencing data or copy number profiles derived therefrom may be called “genome-wide” even though they may not contain information about any location in the genome (i.e., reads or copy number estimates). Indeed, even in the context of sWGS, there may be regions of the genome that do not have mapping reads due to, for example, low mapping potential, sequencing difficulty, or the sampling nature of next-generation sequencing. Thus, data may be called “genome-wide” if it contains data about one or more chromosomes or at least 70%, at least 80%, or at least 90% of all chromosomes present in the reference genome.

[0033] A copy number profile is typically an estimate of the number of copies of a segment in a sample or subset of a sample (e.g., a sample containing cells with different genotypes or genetic material derived therefrom, such as a sample containing tumor and normal cells or a mixture of tumor and normal DNA) for each of several genomic segments. Copy number profiles can be obtained by segmentation and copy number estimation from read count data using methods well known in the art. For example, methods such as circular binary segmentation may be used for segmentation. Copy number fitting may be performed as described in the following examples. Alternatively, methods such as Ascat (Raine et al. 2016), ichorCNA (Adalsteinsson et al. 2017), and others may be used to perform segmentation and copy number estimation (and tumor DNA proportion estimation in the case of Ascat and ichorCNA). The copy number profile may be a genome-wide copy number profile. Genome-wide copy number profiles obtained from target sequencing data using the methods described herein may be equivalent to copy number profiles obtained using sWGS for the same sample.

[0034] In copy number profiles, a "segment" refers to a portion of an array represented in a copy number profile associated with a consistent absolute copy number. The consistent copy number is different from that associated with the immediately upstream array (if such arrays exist and are associated with a copy number estimate) and the immediately downstream array (if such arrays exist and are associated with a copy number). In other words, a segment refers to a portion of an array that has a copy number associated with it, and the copy number associated with a segment is different from the copy number associated with its immediately adjacent segment. The copy number associated with a segment may be different from the copy number associated with its immediately adjacent segment because the segments surrounding it are associated with different copy numbers, the segments surrounding it are not associated with a copy number (e.g., because the data about the segment is missing or of poor quality), or a combination of both (e.g., a segment may be surrounded by segments associated with different copy numbers on one side and segments not associated with a copy number on the other side). In other words, a segment refers to the longest continuous portion of a copy number profile, each associated with a single copy number.

[0035] The methods described herein may include one or more steps of correcting the read count for a genome bin using an estimate of systematic bias in the read count as a function of a variable associated with the bias (referred to as the “bias variable”). The bias variable may be selected from GC content, replication timing, and capture bias score. Systematic bias may be estimated by fitting a model between (i) observed read counts in multiple genome bins and (ii) values ​​of the bias variable for the bins. The model may be a curve-fitting model. The model may be a continuous model or a piecewise model. The model may be a linear model or a nonlinear model. The model may be a linear model or a polynomial model. The model may be a regression model or a generalized additive model. The model may be a nonlinear regression model, such as a local polynomial regression model. For example, the model may be a LOESS (locally estimated scatter plot smoothing) or LOWESS (locally weighted scatter plot smoothing) model. Alternatively, the model may be a piecewise linear model (e.g., a stepwise model) or a piecewise nonlinear model (e.g., a spline). The model may be a univariate model or a multivariate model. Therefore, in embodiments in which read counts are corrected for multiple systematic biases, each may be subject to a series of correction rounds that correct for different systematic biases by fitting a univariate model between (i) the observed read counts in multiple genome bins and (ii) the values ​​of each bias variable for each bin. Alternatively, each may be subject to one or more correction rounds that correct for multiple systematic biases by fitting a multivariate model between (i) the observed read counts in multiple genome bins and (ii) the values ​​of each of the multiple bias variables for each bin. Multiple univariate models are advantageously easy to fit and have been found to work well across all embodiments tested. Correcting read counts with estimates may involve adjusting the observed read count values ​​for each bin based on the model values ​​in coordinates corresponding to the values ​​of the bias variables for each bin.Therefore, correcting the read count using estimates may involve obtaining a corrected read count that depends on the observed read count value for the bin and the model value at the coordinates corresponding to the bias variable value for the bin. For example, correcting the read count using estimates may involve subtracting a correction factor equal to the difference between the model value at the coordinates corresponding to the bias variable value for the bin and a reference value from the observed read count value for the bin. The reference value may be a summary value derived from the observed read counts for multiple bins used to fit the model. For example, the reference value may be the mean or median read count across multiple bins used to fit the model. This is equivalent to adjusting the observed read count so that it is distributed around the reference value for all values ​​of the bias variable. Therefore, correcting the read count for a genome bin may mean obtaining a first read count for the bin and determining a second read count for the bin based on the first read count and an estimate of systematic bias in the read count as a function of the bias variable. The second read count may be referred to as the "adjusted" or "corrected" read count. The first read count may, in some cases, be referred to as the "raw" read count or "observed read count," even if it has already undergone one or more correction rounds (for example, a GC-corrected read count may be referred to as the raw read count for the purpose of correcting for capture bias).

[0036] The GC content of a genomic region (e.g., a bin) refers to the percentage of G or C (guanine or cytosine) bases in the region as a proportion of the total bases in the region (e.g., (G+C) / (G+C+A+T)%). GC content is known to affect the likelihood of a region being sequenced by next-generation sequencing. GC content corrections have been proposed, which involve fitting a model to the observed read count per bin as a function of the bin's GC content and adjusting the observed read count using the model values. In embodiments of this disclosure, multiple models are fitted using multiple bins, each bin of each multiple bin is associated with the same estimated copy number state. The correction can then be performed with respect to each multiple bin using each model, as described above. Thus, in embodiments of this disclosure, the copy number can be estimated for each bin (or for each of multiple genomic segments from which bin copy numbers can be obtained), which can be used to identify multiple bin sets, each containing bins with the same estimated copy number. Subsequently, GC bias correction may be performed separately for at least one (or all) of the multiple binsets. This may be referred to herein as “copy number-recognized GC correction.” It should be noted that the advantage of such an approach is that it is applicable to all genome sequencing data, whether from targeted sequencing or non-targeted sequencing. Copy number-recognized GC correction may be particularly useful when analyzing samples containing tumor genetic material. Indeed, such samples may contain abnormal copy number states that, if not considered, would result in insufficient GC correction. Therefore, copy number-recognized GC correction may be useful when analyzing any sample with copy number instability. A sample may be considered to have copy number instability if the tumor copy number profile obtained from the sample contains at least one copy number change that differs from the diploid state. Copy number-recognized GC correction may be particularly useful when analyzing samples containing tumor genetic material in which the tumor has high chromosomal instability (high CIN).Tumors having a copy number profile that includes at least one copy number variation different from the diploid state may be considered to have CIN. Tumors having a copy number profile that includes at least 5, 10, 15, 20, 25, or 30 copy number variations different from the diploid state may be considered to have "high CIN". We have found that copy number recognition GC correction is advantageous for any sample containing tumor genetic material in which the tumor has a copy number profile that includes at least one copy number variation different from the diploid state.

[0037] Replication timing refers to the order in which different regions of a chromosome are replicated during replication. Replication timing profiles are available for many genomes and provide the temporal order of replication for all segments in the genome. Replication timing is known to affect the abundance of sequences in a sample, leading to low-amplitude changes in DNA copy number (typically much lower than the stepwise changes observed in the presence of copy number mutations or somatic copy number abnormalities). Not all samples exhibit systematic bias in read count as a function of replication timing. However, in embodiments, read counts for a bin can be corrected using estimates of systematic bias in read count as a function of replication timing. This can be done by fitting the model to read count as a function of observed replication timing, as described above. Replication timing information can be obtained from reference genome databases, such as the USCS genome browser. Replication timing can be obtained from repli-seq timing data for one or more samples. For example, replication timing information can be obtained from wavelet-smoothed repli-seq timing data from UCSC (UW Replication timing can be obtained based on Repli-seq. Replication timing can be obtained as a summary metric that summarizes data from multiple samples (e.g., average avelet-smoothed Repli-seq timing data from multiple cell lines). Replication timing correction may be advantageous in samples where the observed bin read count and the corresponding replication timing are significantly correlated. For example, replication timing correction may be applied when the correlation coefficient (e.g., Pearson correlation (r)) between the observed bin read count value and replication timing is greater than or equal to a predetermined threshold, such as 0.1.Therefore, the methods described herein may include determining a correlation coefficient (e.g., a Pearson correlation coefficient) between observed read counts (optionally corrected for GC content and / or capture bias) for a plurality of bins, and, if the correlation coefficient is above a predetermined threshold, correcting the read count for a bin (or all of the plurality of bins) using an estimate of systematic bias in the read count as a function of replication timing.

[0038] Capture bias refers to a read count for one or more bins that is artificially inflated compared to other bins due to the use of sequence-specific capture steps in the sequencing process used to obtain the read count. According to embodiments of this disclosure, the read count for a genomic bin is corrected using an estimate of systematic bias in the read count as a function of a variable associated with the bias. The variable may be the capture bias score. Thus, the methods described herein may, in embodiments, include (i) fitting a model between observed read counts in a plurality of genomic bins and (ii) a value of the capture bias score for the bin, and correcting the read count for a bin by adjusting the observed read count value for the bin based on a value of the model at a coordinate corresponding to the value of the capture bias score for the bin. For example, correcting the read count may include subtracting from the observed read count value for the bin a correction factor equal to the difference between the value of the model at a coordinate corresponding to the value of the capture bias score for the bin and a reference value. The reference value may be the mean or median read count across a plurality of bins used to fit the model. The capture bias correction may be applied to genomic bins that contain only off-target regions and / or only off-target reads. Such bins can be obtained by applying the binning process to the genome or a portion of the genome being analyzed using a fixed-size window, and then removing any target region from the resulting bins. This, therefore, can result in bins of variable size.

[0039] The capture bias score is a variable that represents the difference between the observed read position metric distribution within a bin and the expected read position metric distribution within a bin. The expected read position metric distribution may be the read position metric distribution associated with a bin that is unaffected by capture bias. The expected read position metric distribution may be the read position metric distribution corresponding to a uniform distribution of reads across the bins. The capture bias score may be any variable that captures the discrepancy between the observed distribution of read positions within a bin and the uniform distribution of read positions across the bins. The position of a read within a bin may be represented by a metric of the pairwise distance between pairs of reads within the bin. In other words, the read position metric may be the pairwise distance between pairs of reads. Therefore, the capture bias score may be any variable that captures the discrepancy between the observed distribution of pairwise distances between reads within a bin and the expected distribution of pairwise distances between reads within a bin. The expected distribution of pairwise distances between reads across the bins may be the distribution of pairwise distances between reads within a bin that corresponds to a uniform distribution of reads within the bins. The pairwise distance between reads can be the left-end pairwise distance (distance between the left-end boundaries of two aligned reads), the right-end pairwise distance (distance between the right-end boundaries of two aligned reads), the center pairwise distance (distance between the centers of two aligned reads), or any distance between a given position on the reads (e.g., the left-end boundary, right-end boundary, or center position of the reads). The distance between reads can be expressed in units of nucleotides. The distance refers to the distance in genomic coordinates of the reference sequence to which the two reads are aligned. In a pair of reads, the left-end distance between two reads can be the distance between the left-end boundary of the first read and the left-end boundary of the second read. A distance of 0 means that the two reads have the same left end. The expected distribution of pairwise distances between reads with respect to a bin can be a left triangular distribution where the mode is 0 or the smallest observed pairwise distance in the bin, the minimum is 0 or the smallest observed pairwise distance in the bin, and the maximum is equal to the largest observed distance in the bin, the bin length, or the bin length + read length, or any value between the bin length and the bin length + read length. The read length is typically a known feature of the sequencing technique used.For example, when determining scores for 50kb bins and 150nt (nucleotide) length reads, the expected distribution could be a triangular distribution with a minimum value = 0, a mode = 0, and a maximum value = 50kb nt, 50150nt, or the maximum observed distance within the bin (typically in the range of 50000 to 50150nt or its vicinity). The inventors have found that all such values ​​for the maximum value of the expected distribution yield similar results. The expected distribution can be approximated by sampling a number of observations from such a left triangular distribution that corresponds to the number of observations from which the observed distribution is obtained (i.e., the number of observed pairwise distances between reads). A variable that captures the discrepancy between the observed distribution of pairwise distances between reads within a bin and the expected distribution of pairwise distances between reads within a bin could be a variable (e.g., sum, mean, average, etc.) that summarizes the absolute difference between the observed frequency and the expected frequency across multiple windows of pairwise distances. Multiple windows may be non-overlapping windows that together cover the range from the minimum to the maximum observed distance. Multiple windows may be non-overlapping windows that together cover the range of the expected distribution of distance. Therefore, a variable that captures the discrepancy between the observed and expected distributions of pairwise distances between reads in a bin may be a variable that summarizes the difference between the values ​​of the histogram of observed pairwise distances between reads in a bin (i.e., the set of estimated frequencies for each window of pairwise distances) and the corresponding values ​​from the expected distribution (e.g., a left triangular distribution). The number of windows in the histogram may be predetermined or may be determined dynamically based on the number of observations in the distribution. For example, a variable that captures the discrepancy between the observed distribution of pairwise distances between reads in a bin and the expected distribution of pairwise positions of reads across the bin may be the mean of the absolute values ​​of the deviations between the histogram density function fitted to the observed pairwise position data and the corresponding theoretical left triangular distribution. The number of windows is not limited in principle, and if a very large amount of data is available, a very high-resolution histogram (which may also be called a density function) can be fitted.Therefore, a variable that captures the discrepancy between the observed distribution of pairwise distances between reads in a bin and the expected distribution of pairwise positions of reads across the bins may be the integral of the area between the density function fitted to the observed pairwise position data and the corresponding theoretical left triangular distribution. A variable that captures the discrepancy between the observed distribution of pairwise distances between reads in a bin and the expected distribution of pairwise distances between reads in a bin may be a variable (e.g., sum, mean, average, etc.) that summarizes the probability that the observed pairwise distances are sampled from the expected distribution over multiple observed pairwise distances (e.g., over all observed pairwise distances in a bin). A capture bias score may be obtained by subtracting the value of such a variable from 1, with higher values ​​associated with greater discrepancies from the expected distribution. A variable that captures the discrepancy between the observed distribution of pairwise distances between reads in a bin and the expected distribution of pairwise distances between reads in a bin may be a statistical metric that compares two samples to determine whether the two samples are from the same distribution. Therefore, a variable that captures the discrepancy between the observed distribution of pairwise distances between reads in a bin and the expected distribution of pairwise distances between reads in a bin may be a statistical metric associated with a hypothesis test that the two samples are from the same distribution. The statistical metric may be selected from the Kolmogorov-Smirnov statistic, the DTS statistic, and the Anderson-Darling test statistic, the Cramer-von Mises test statistic, etc. In this embodiment, the statistical metric is the DTS statistic. The statistical metric may be calculated between a first sample containing the observed pairwise distances between reads and a sample of the expected pairwise distances between reads drawn from the expected distribution of pairwise distances between reads (e.g., a left triangular distribution).

[0040] The capture bias score can be obtained for any bin containing at least a predetermined number of reads, such as 5, 10, 15, or 20 reads. A minimum value of 10 has been found to strike a good balance between retaining as much raw data as possible and obtaining a reliable score (depending on whether sufficient data is available for each bin). For bins containing fewer than a predetermined number of reads, the score can be set to a default value equal to the minimum capture bias score identified for the sample.

[0041] A “genome bin” (also referred to herein as a “bin”) is a region of the genome obtained through a process that includes applying a fixed-width window to a genomic region (e.g., a chromosome, multiple chromosomes, or the whole genome). The fixed-width window is typically applied by sliding it non-overlapping to create multiple adjacent bins of equal size. The size of the bin can be set to a predetermined default value, such as 20kb nt, 50kb nt, or 100kb nt. In embodiments, the default value of 50kb nt is used. Alternatively, the size of the bin can be determined based on the characteristics of the sample being analyzed, including factors such as tumor purity (if the tumor copy number profile is determined from a mixed sample), the number of off-target reads, and sequencing depth. All of these factors affect the density of useful reads mapped to the genome and therefore affect the size of the bin that contains enough useful reads. If tumor purity and / or the number of off-target reads are low, it may be beneficial to increase the bin size from the default value of, for example, 50kb nt. Conversely, when tumor purity is high and / or the number of off-target reads is high, narrower bins can be used to obtain higher resolution data (e.g., higher resolution copy number profiles). Generally, it is within the capabilities of those skilled in the art to test multiple bin width values ​​(e.g., within a predetermined range such as 10kb–1000kb nt or 50kb–500kb nt) to identify a bin size that provides a good balance between precision and resolution for a particular sample. See, for example, Macintyre, Ylstra & Brenton (2016) and Scheinin et al. 2014. For example, the bin width may be selected to obtain an average of at least 15 reads / bin per tumor copy (i.e., an average of 15 reads per absolute number of tumor copies at the bin location). Bin sizes of 10kb–500kb are considered particularly advantageous because they provide good resolution in the range of downstream analyses.However, bin sizes up to 1000kb are useful when such resolution is required to ensure that a sufficient number of reads / bins are present (e.g., on average, at least 15 reads per absolute copy number of tumor at the bin location). For targeted capture sequencing assays, a sufficient number of reads are typically available in samples with at least 10%, preferably at least 20%, tumor purity. For targeted capture sequencing assays, a sufficient number of reads are typically available in samples with at least 200,000 reads. High-resolution copy number profiles (e.g., 50kb bins) can typically be obtained for samples with approximately 2 million reads. As will be further discussed below, genomic bins may be further processed to remove one or more target regions for the purpose of correcting for capture bias. This results in a processed bin width smaller than the original bin width. Thus, for the purpose of correcting for capture bias, multiple bins may have different processed lengths from one another, even though all these bins were obtained through a process that involves applying a fixed-width window to genomic regions. After bias correction, including capture bias correction, the read count for a bin may be adjusted using the difference between the (original) bin width and the processed bin width. In particular, the adjusted read count may be obtained by adding to the corrected read count for the bin a read count proportional to the corrected read count in the bin and the size of the difference between the (original) bin width and the processed bin width. Thus, the processed bin width may be used for the purpose of adjusting the read count for a bin after capture bias correction using the capture bias score. Adjusting the read count in this context means adding a pseudocount to the corrected count (a count obtained at least after capture bias correction, and also optionally after GC content and / or replication timing correction), where the pseudocount is given by pseudocount = (percentage of duplicates previously removed from the bin) * (count for the bin after correction has been applied).Therefore, the adjusted count in this embodiment is (count for the bin after correction is applied) * (1 + (= (percentage of overlap previously removed from the bin)). The determination of the capture bias score may use the original bin width as a parameter (for example, to determine the expected distribution of the read location metric). Alternatively, the bin width may not be processed, and the corrected read count for a bin may be obtained by obtaining a first corrected read count (i.e., obtaining a corrected read count for the bin as described herein, simply removing on-target reads in the bin) based on the observed off-target read count for the bin and one or more estimates of the bias in the read count as a function of the variables associated with each bias, and then adding a pseudo-count to this read count. The pseudo-count in this context is proportional to the first read count and the percentage of overlap between the bin and the capture region. For example, the pseudo-count may be calculated as (percentage of overlap with the capture region file) * (count for the bin after all corrections are applied). The capture region may be as defined in the BED file.

[0042] As used herein, “sample” can be a cell or tissue sample, biological fluid, or extract (e.g., a DNA extract obtained from a cell, tissue, or bodily fluid sample) from which genomic material can be obtained for genome analysis, such as genome sequencing (e.g., whole genome sequencing, or targeted sequencing such as whole exome sequencing or targeted panel sequencing). A sample can be a cell, tissue, or biological fluid sample (e.g., a biopsy) obtained from a subject. Such a sample may be referred to as a “subject sample.” In particular, a sample can be a blood sample, urine sample, cerebrospinal fluid sample, or tumor sample (e.g., a tumor biopsy), or a sample derived therefrom. A sample can be any sample containing genomic DNA or cell-free DNA. A sample may be freshly obtained from a subject, or it may have been processed and / or stored prior to genome / transcriptome analysis (e.g., frozen, fixed, or subjected to one or more purification, concentration, or extraction steps). A sample may be a cell or tissue culture sample. Therefore, the samples described herein may refer to any type of sample containing cells or genomic material derived therefrom, whether from biological samples obtained from subjects or from samples obtained from, for example, cell lines. In embodiments, the sample is a sample obtained from a subject, such as a human subject. The sample is preferably from a mammal (e.g., mammalian cell samples, or samples from mammalian subjects, such as cats, dogs, horses, donkeys, sheep, pigs, goats, cattle, mice, rats, rabbits, or guinea pigs), and preferably from a human (e.g., human cell samples or samples from human subjects).Furthermore, samples may be transported and / or stored, collection may be performed at a location away from the location of sequence data acquisition (e.g., sequencing), and / or any computer-aided procedures described herein may be performed at a location away from the sample collection location and / or at a location away from the location of sequence data acquisition (e.g., sequencing) (for example, computer-aided procedures may be performed by a networked computer, for example, by a “cloud” provider). Samples may include circulating tumor DNA (ctDNA), such as tissue biopsies, such as fresh-frozen tissue samples, or formalin-fixed paraffin-embedded (FFPE) tissue samples, or biological fluid samples (sometimes referred to as liquid biopsies).

[0043] The samples used in the methods of this disclosure are typically tumor cells (e.g., a tumor sample or a sample containing circulating tumor cells), or samples containing genetic material derived from tumor cells (e.g., cell-free DNA, or DNA extracted from a sample containing tumor cells or circulating tumor cells, e.g., from a cell line or tissue sample). The methods described herein may be applicable to a single sample. In particular, the methods described herein may enable obtaining a copy number profile for a sample containing one or more copy number abnormalities (e.g., a tumor sample or a sample containing genetic material derived from tumor cells) in the absence of a corresponding normal sample. The corresponding normal sample is a sample expected to represent the genome of the tumor sample in the absence of somatic abnormalities. This is typically a sample obtained from the same subject as the tumor sample.

[0044] The term “sequence data” refers to information indicating the presence of genetic material in a sample having a specific sequence. Such information can be obtained using sequencing techniques such as next-generation sequencing (NGS), such as whole exome sequencing (WES), whole-genome sequencing (WGS), or sequencing of captured genomic loci (targeted or panel sequencing). In the context of this disclosure, sequence data is typically obtained by DNA sequencing, particularly next-generation sequencing. Therefore, sequence data may include sequencing reads or information derived therefrom. Information such as the count of sequencing reads having a specific sequence covering a specific genomic location or within a specific genomic region (e.g., a genomic bin) can be derived from sequencing reads by mapping sequence data to a reference sequence, such as a reference genome, using methods well known in the art (e.g., Bowtie (Langmead et al., 2009) or BWA (Li and Durbin 2009). This can result in aligned sequencing reads in the form of, for example, SAM or BAM files. Therefore, sequencing can be associated with specific genomic locations (where “genomic location” refers to a location in the reference genome to which the sequence data is mapped). Sequencing data may be filtered as part of the methods described herein, or filtered before applying the methods described herein to remove, for example, low-quality reads or reads that map to regions with “low mapping potential”. Methods for filtering low-quality reads are well known in the art and include, for example, the removal of duplicate reads, the removal of supplementary alignments, the removal of reads with low sequencing quality scores (reads less than Q20 or Q30), the removal of reads with less than a certain percentage of identified bases (e.g., reads with less than 100% identified bases, where identified bases may be identified as “non-NNNN” bases), and the removal of reads with low mapping quality scores (e.g., MAPQ<37).Methods for filtering reads that map to regions with low mapping potential are well known in the art, including, for example, Umap (Karimzadeh et al. 2018) and GenMap (Pockrandt et al. 2020).

[0045] This disclosure also provides a method for generating copy number profiles, or genomic features derived from copy number profiles, such as the presence or absence of one or more chromosomal instability (CIN) signatures, or the tumor purity or percentage of circulating tumor DNA (ctDNA) in a sample, using acquired non-targeted genome sequencing data.

[0046] Chromosomal instability (CIN) is a process of accumulating numerical and structural changes in DNA. A CIN signature (also referred to as a copy number signature) is a signature (typically in the form of a set of weights associated with each category of copy number features) that represents the genome-wide trace of different putative mutation processes (as used herein, “mutation process” refers to any process that can cause chromosomal instability). CIN signatures are described in MacIntyre et al. 2018, Drews et al. (2022), and PCT / EP2022 / 077473, which are incorporated herein by reference. The presence or absence of a CIN signature in a sample can be determined by calculating the sample's exposure to the signature. Methods for determining exposure to a signature from a copy number profile are well known in the art (see, for example, MacIntyre et al. 2018 and Drews et al. (2022)). The term “copy number (CN) feature” refers to the characteristics of copy number events observable in a copy number profile. Copy number features may include the absolute copy number of a segment, segment size (also referred to as “segment length,” typically expressed in terms of base pairs), breakpoints per xMB (the number of change points appearing in a sliding window across the copy number profile, where the window is preferably 10MB and the copy number profile is preferably genome-wide), changepoint copy number (the absolute difference in copy number between a segment and an adjacent / next-segment segment in the copy number profile, which can be defined with respect to upstream or downstream neighboring segments), breakpoints per chromosome arm (the number of change points occurring per chromosome arm), and the number of segments with oscillating copy numbers (sometimes referred to as “length of segments with oscillating copy numbers,” the number of consecutive segments alternating between two copy number states, rounded to the nearest integer copy number state, also referred to as the length of the chain of oscillating copy number states).

[0047] A sample may be a “mixed” sample containing cells with different genotypes or genetic material derived therefrom. For example, a sample may contain tumor cells and normal cells, or DNA derived therefrom, such as in the context of a cell-free DNA (cfDNA) sample containing circulating tumor DNA (ctDNA). Copy number profiles can be used to determine the percentage of cells with genotypes containing somatic copy number abnormalities or cell-derived DNA in a sample containing both cells with genotypes containing somatic copy number abnormalities (e.g., tumor cells) and cells without such abnormalities (e.g., normal cells). The term “tumor fraction” (sometimes referred to as “tumor purity” or simply “purity” or “abnormal cell fraction (ACF)”) refers to the proportion of DNA-containing cells that are tumor cells in a mixed sample, or an equivalent proportion that is expected to result in a particular mixture of genetic material derived from tumor and non-tumor cells in the sample. Accordingly, the methods described herein may be used to determine the tumor fraction in a sample containing cells or genetic material, or the percentage of ctDNA in a sample containing cfDNA. Tumor fractions of ctDNA proportions can be estimated using sequence analysis processes that attempt to deconvolve tumor and germline genomes from copy number profiles, such as ASCAT (Raine et al., 2016), ABSOLUTE (Carter et al., 2012), or, in the context of cfDNA, ichorCNA (Adalsteinsson et al., 2017).

[0048] As used herein, the term “computer system” includes hardware, software, and data storage devices for implementing or performing methods of the systems according to the embodiments described above. For example, a computer system may include a central processing unit (CPU) and / or graphics processing unit (GPU), input means, output means, and data storage devices, which may be implemented as one or more connected computing devices. A computer system may include a display, or a computing device having a display that provides a visual output display. Data storage devices may include RAM, disk drives, or other non-temporary computer-readable media. A computer system may include a plurality of computing devices connected by a network and capable of communicating with each other through this network. It is expressly assumed that a computer system may consist of a cloud computer, or may include a cloud computer.

[0049] As used herein, the term “computer-readable media” includes, but is not limited to, any non-temporary media that can be directly read and accessed by a computer or computer system. These media include, but are not limited to, magnetic storage media such as floppy disks, hard disks, and magnetic tapes; optical storage media such as optical disks or CD-ROMs; electrical storage media such as RAM, ROM, and flash memory; and hybrids and combinations thereof, such as magnetic / optical storage media.

[0050] Acquisition of non-target genome sequencing data This disclosure provides a method for analyzing genome sequence data, including correcting for one or more systematic biases in the data. In particular, this disclosure provides a method for obtaining non-target genome sequencing data from target sequencing data. An exemplary method is illustrated with reference to Figure 1.

[0051] The method may include an optional step 10 in which one or more samples containing cells or genetic material derived therefrom are obtained. Optionally, the samples may be sequenced in step 12 using a target panel DNA sequencing technique. Alternatively, the sequence data may be pre-obtained and received from a user interface, computing device, or database. The sequence data may include sequence data from a target sequencing assay. The sequence data may include multiple sequencing reads. The sequence (seqnece) data may further include coordinates of multiple capture regions. In an optional step 14, the sequence data may be pre-processed. This may include aligning multiple reads in the sequence data to a reference genome. Alternatively, the sequences may be pre-aligned and received in the form of one or more files (e.g., SAM or BAM files) containing the reads and alignment locations in the reference genome. Step 14 may include obtaining multiple genomic bins by applying a non-overlapping fixed-width window along the genomic coordinates with respect to one or more genomic regions mapping the reads, thereby obtaining multiple fixed-width bins. Step 14 may further include annotating multiple bins, for example, using replication timing information and / or GC content. In step 16, reads not mapped to capture regions are selected and used to determine the observation count for each of the multiple bins. These may be referred to as “off-target reads,” and the observation read counts obtained using them may be used in all subsequent steps. When using paired-end sequencing data, off-target reads may be reads not mapped to capture regions and reads that do not have mate reads mapped to capture regions. Optionally, the processed bin width for a bin may also be obtained by removing any region that is a capture region from a fixed-width bin. Thus, a processed bin width can be obtained for each bin that overlaps with a capture region. The processed bin widths may differ from each other and may also differ from the fixed width. Alternatively, or in addition to this, the proportion or percentage of any fixed-width bins that are capture regions may be calculated.

[0052] In step 18, GC correction is performed. This may further include obtaining GC-corrected read counts for a bin by fitting a model between the observed read counts and (ii) the GC content values ​​for the bins. This may be referred to as “single GC correction” because it corrects the read counts for all bins using a single model. Step 18 may further include using the GC-corrected read counts to determine whether the observed read counts for each of the multiple genomic bins indicate the presence of multiple bin sets, where all bins in a set have the same estimated copy number, and the estimated copy numbers associated with each of the multiple sets are different from one another. When it is determined that the observed read counts for each of the multiple genomic bins indicate the presence of multiple bin sets, step 18 may further include determining the corrected read counts based on the observed read counts (which are optionally already corrected for or before correcting for capture bias) and an estimate of the bias in the read counts as a function of GC content. An estimate of the GC content bias in read count as a function of GC content can be obtained by identifying multiple bin sets in which all bins in a set have the same estimated copy number, and by fitting a model between (i) the observed read count in one of the bin sets containing the bins for which the corrected read count is determined, and (ii) the GC content values ​​for the bins in the set. This may be called a “set-specific GC correction” given that the correction uses multiple models, each fitted to a specific bin set. For both single-GC corrections and set-specific GC corrections, an estimate of the GC bias in read count as a function of GC content is estimated by fitting a model between (i) the observed read count in multiple genome bins (or selected sets) and (ii) the GC content values ​​for multiple genome bins (or selected sets).Subsequently, the corrected read count for each bin is determined by subtracting from the observed read count a correction factor equal to the difference between (i) the model value at the coordinates corresponding to the GC content value for the bin and (ii) a reference value such as a summary value (e.g., median) derived from the observed read counts for multiple bins (or a selected set).

[0053] In step 20, capture bias correction is performed by determining a corrected read count for one or more of the genome bins, the corrected read count being based on the GC-corrected observed read count obtained in step 18 and an estimate of the bias in the read count as a function of a bias-associated variable, where the bias-associated variable is the capture bias score, which represents the difference between the observed read position metric distribution within the bin and the expected read position distribution metric within the bin. The estimate of the bias in the read count as a function of a bias-associated variable is estimated by fitting a model between (i) the GC-corrected observed read counts for the genome bins and (ii) the values ​​of the bias variable for the genome bins. For each bin, the corrected read count is determined by subtracting a correction factor from the GC-corrected observed read count that is equal to the difference between (i) the model value at the coordinates corresponding to the value of the bias variable for the bin and (ii) a reference value such as a summary value (e.g., median) derived from the GC-corrected observed read counts for the genome bins.

[0054] In step 22, the corrected count obtained in step 20 may be further corrected with respect to replication timing. This may include determining whether the observed read counts for multiple bins (typically already corrected, e.g., obtained from step 18 or step 20 in the illustrated embodiment) are significantly associated with replication timing, and further correcting the observed read counts (typically already corrected) with respect to replication timing bias if the observed read counts (typically corrected) for multiple bins are significantly associated with replication timing. This may be done by fitting a model between the (typically corrected) observed read counts and (ii) the replication timing values ​​for the bins, and determining the corrected read counts by subtracting a correction factor from the (typically corrected) observed read counts that is equal to the difference between (i) the model values ​​in the coordinates corresponding to the replication timing values ​​for the bins and (ii) a reference value such as a summary value (e.g., median) derived from the (typically corrected) observed read counts for multiple bins.

[0055] In any step 24, for each bin, an additional corrected lead count is obtained in which the bins overlapping with the capture area and / or on-target leads removed in step 16 are obtained. This can be done by adding an adjustment factor proportional to the difference between the processed bin width and the fixed-width bin width determined in step 16 for each bin to the corrected lead count for each bin (the result of one or more of steps 18, 20, and 22). The adjustment factor may be proportional to the corrected lead count for each bin and the difference between the processed bin width and the fixed-width bin width for each bin. Alternatively, this can be done by adding an adjustment factor (also referred to as a "pseudo-count") proportional to the corrected lead count for each bin and the percentage of bins overlapping with the capture area to the corrected lead count for each bin (the result of one or more of steps 18, 20, and 22). For example, the corrected read count for bins that overlap with the capture region can be calculated as (count for bins after all corrections have been applied) + pseudo-count, and the pseudo-count can be calculated as pseudo-count = (percentage of overlap with the capture region) * (count for bins after all corrections have been applied).

[0056] In any step 26, one or more results from a preceding step are provided to the user, for example, via a user interface.

[0057] Furthermore, this specification describes a method for obtaining target and non-target genome sequence data from a sample using a target DNA sequencing protocol, the method comprising obtaining target genome sequence data from a sample, which includes a plurality of DNA sequencing reads obtained using a target genome sequencing assay, and applying the above method, such as the method in Figure 1, to the sequence data to obtain non-target genome sequence data relating to the sample.

[0058] Application examples The above method is applicable in the context of determining copy number profiles, the presence or absence of one or more genomic features derived from copy number profiles such as the presence or absence of one or more chromosomal instability signatures, or the percentage of tumors or ctDNA in a sample.

[0059] Therefore, also described herein are methods for providing copy number profiles or genomic features derived from copy number profiles relating to a sample, the methods comprising obtaining non-target genomic sequence data using the methods described herein. An example of such a method is illustrated with reference to Figure 2.

[0060] In an optional step 210, one or more sample cells or genetic material derived therefrom are obtained. In an optional step 212, genomic sequence data for the sample is obtained. In step 214, a corrected read count is obtained from the genomic sequence data for the sample using a method described herein, for example, with reference to Figure 1. In step 216, a copy number profile for the sample is obtained using the corrected count obtained in step 216. This may include step 216A, which segments the corrected read count profile along genomic coordinates to identify genomic segments having the same copy number state. Step 216 may include determining the absolute copy number for each segment in step 216B. Step 216 may include determining the tumor purity or percentage of ctDNA in the cfDNA sample using the corrected read count. Absolute copy number fitting and purity / ctDNA ratio estimation can actually be performed simultaneously through a process of deconvolving a mixed copy number profile obtained from data including reads associated with tumor genetic material and reads associated with normal genetic material into a copy number profile associated with tumor genomic material in the sample and a copy number profile associated with normal genomic material in the sample (typically assumed to be diploid in all segments). In step 218, the derived information can be obtained by analyzing the copy number profile obtained in step 216. This may include, for example, determining whether the copy number profile is present or shows one or more CIN signatures. This may include determining the prognosis and / or diagnosis of the subject from whom the sample was obtained. Step 218 may also include determining the presence of one or more copy number abnormalities in the sample, such as the presence of one or more clinically relevant copy number abnormalities (e.g., oncogene amplification, tumor suppressor deletion, etc.).

[0061] In step 220, the results of any preceding steps may be provided to the user, for example, via a user interface. This may take the form of a report showing, for example, the presence of one or more copy number anomalies, one or more CIN signatures, prognostic or diagnostic metrics, tumor purity or ctDNA percentage, etc.

[0062] Also described herein is a method for determining the percentage of tumor cells or tumor DNA in a sample, the method comprising: obtaining non-target genomic sequence data (i.e., corrected read counts, whether obtained from target or non-target genomic sequence data) relating to the sample using the method described herein; and identifying the tumor copy number profile and the percentage of tumor DNA in the sample from the above data. This may include deconvolving the observed read count profile between contributions from normal cells with known expected ploidy and contributions from tumor cells with unknown copy number profiles. Methods for carrying this out are well known in the art and include, for example, Ascat (Raine et al. 2016), ichorDNA (Adalsteinsson et al. 2017), Absolute (Carter et al. 2012), Sequenza (Favero et al. 2015), etc.

[0063] The sample may be a cell-free DNA sample, and this method can determine the percentage of circulating tumor DNA in the sample. For example, the method described by Adalsteinsson et al. (2017) can be applied to non-targeted genomic sequence data obtained as described herein (i.e., read counts associated with genomic coordinates and corrected for GC bias and / or capture bias).

[0064] The sample may include cells (e.g., tumor biopsy), and this method can determine the percentage of tumor cells in the sample. For example, the method described in Raine et al. (2016) or Carter et al. (2012) can be applied to non-targeted genomic sequence data obtained as described herein (i.e., read counts associated with genomic coordinates and corrected for GC bias and / or capture bias).

[0065] The percentage of ctDNA in a cfDNA sample from a subject has been identified as an indicator of minimal residual disease (MRD), prognosis, and response to treatment in many cancers. Therefore, this disclosure also provides a method for providing an indicator of prognosis, MRD diagnosis, or response to treatment for a subject diagnosed with or likely to have cancer, the method comprising: obtaining targeted sequencing data from a biological fluid sample from the subject; analyzing the data using the method described herein to determine the percentage of ctDNA in the sample, the percentage of ctDNA in the sample indicating the subject's prognosis, MRD status, or response to treatment. For this purpose, any method well known in the art for determining prognosis, MRD status, or response to treatment from the percentage of ctDNA in a liquid biopsy may be used.

[0066] system Figure 3 shows one embodiment of a system for acquiring non-target genome sequencing data and / or providing copy number profiles or genomic features derived therefrom, according to the present disclosure. The system includes a computing device 1, which includes a processor 101 and computer-readable memory 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103. The user interface 103 is shown as a screen, but may include any other means of communicating information to the user, such as via audible or visual signals. The computing device 1 is communicably connected, for example, via a network 6, to sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 that store sequence data. The one or more databases may further store other types of information, such as reference sequences, parameters, etc., which may be used by the computing device 1. The computing device may be a smartphone, tablet, personal computer, or other computing device. The computing device is configured to carry out a method for acquiring non-target genome sequencing data and / or providing copy number profiles or genomic features derived therefrom, as described herein. In another embodiment, computing device 1 is configured to communicate with a remote computing device (not shown) which is configured to perform a method for acquiring non-target genome sequencing data and / or providing copy number profiles or genomic features derived therefrom, as described herein. In such a case, the remote computing device may also be configured to transmit the results of the method to computing device. Communication between computing device 1 and the remote computing device may be via a wired or wireless connection and may take place via a local or public network, such as the public internet or Wi-Fi.The sequence data acquisition means 3 may be wiredly connected to the computing device 1 and / or the database 2, or, as shown in the figure, may communicate via a wireless connection, for example, via a network 6. The connection between the computing device 1 and the sequence data acquisition means 3 may be direct or indirect (for example, via a remote computer). The sequence data acquisition means 3 is configured to acquire sequence data from nucleic acid samples extracted from cell and / or tissue samples, such as genomic DNA samples. In some embodiments, the samples may have undergone one or more preprocessing steps, such as DNA purification, fragmentation, library preparation, and target sequence capture (e.g., exon capture and / or panel sequence capture). The sequence data acquisition means is preferably a next-generation sequencer. The sequence data acquisition means 3 may be directly or indirectly connected to one or more databases 2 in which sequence data (raw or partially processed) may be stored.

[0067] The following are presented as examples only and should not be interpreted as limitations on the claims. [Examples]

[0068] Examples Introduction Capture-based DNA target panels are routinely used for small variant detection and very limited gene-level copy number profiling in cancer clinical care. The capture efficiency of these panels is not optimal, resulting in off-target DNA being sequenced alongside on-target DNA. These off-target reads can range from 8 to 50% of the total number of reads, and they are scattered throughout the genome, enabling the derivation of genome-wide copy number profiles. The number of reads in these regions presents a significant bias that makes comprehensive copy number characterization extremely difficult. In this embodiment, we describe a method for correcting these biases and generating robust genome-wide copy number profiles. We demonstrate the method by analyzing various NGS protocols and showing that, after correcting for two major noise sources in this type of data—GC content and off-target read depth bias—the proposed method can provide results comparable to the current gold standard shallow whole-genome sequencing (sWGS) samples. Benchmarks using alternative calculation methods further demonstrate performance improvements without the need for corresponding normal tissue or the exclusion of noisy data, which are common drawbacks of other methods in the field of cancer.

[0069] Example 1 - Determination of genome-wide copy number from target capture panel sequencing data Determining genome-wide copy number from target capture panel sequencing data is difficult due to two major factors that artificially alter the expected number of reads, which should be proportional to the amount of DNA: GC content and capture bias.

[0070] GC content is a well-known source of bias that has been widely addressed in past studies. The ability of next-generation sequencing approaches to amplify and sequence any given DNA fragment depends on the ratio of guanine (G) or cytosine (G) to adenine (A) and tyrosine (T) in the sequence. Generally, the higher the GC content of a DNA fragment, the more likely it is to be sequenced. This results in artificially heterogeneous read depths when aligned with a reference genome. To account for this, observed read depths are typically corrected based on GC content. Typically, a LOESS curve is fitted between the raw read depth and the GC content of a genomic region (see, e.g., Figure 4), and this is used to adjust the read depths and remove the bias caused by variable GC content. Generally, this is successful for normal genomes, but when applied to tumor genomes, challenges arise due to the variable copy number state. In tumor genomes, the variable copy number state causes a shift in the read depth distribution (see Figure 5). Standard GC fitting procedures fit to the center of this divided distribution, so the correction procedure slightly reduces the true separation between copy number states toward the mean (see plot above in Figure 6). This has detrimental effects on downstream applications such as absolute copy number estimation. Recently, an approach developed for genome-wide copy number estimation from low-pass single-cell DNA sequencing was described by Wang et al. (2020). To address this challenge, this approach uses a Poisson latent factor model for read depth normalization and models GC content bias using an expectation maximization (EM) algorithm incorporated into a Poisson regression model to account for discretized copy number states along the genome. However, no approach has been developed that can address this challenge in the context of bulk DNA sequencing, and there is an opportunity to improve GC correction, including from capture panel sequence data.

[0071] A second major source of read depth bias is caused by off-target enrichment of DNA due to the capture process. Any DNA sequences outside a designated capture region that are sufficiently similar to the capture region are enriched and therefore exhibit higher read depths. For example, regions containing pseudogenes. One approach to correct this bias is to sequence the tumor sample in parallel with the corresponding normal sample, in which case the expected coverage should be uniform across the genome, and thus a suitable correction model can be developed. However, in many clinical scenarios, the corresponding normal sample is not available, and therefore these approaches are not applicable. Another approach is to determine which regions may have artificially high read depths due to capture bias and remove them from downstream analysis. This approach is only successful for low-resolution copy number determination, creating gaps in the genome and making higher-resolution copy number determination difficult. We hypothesized that it may be possible to retain all regions across the genome using a GC-correction style procedure, provided there is a sufficient framework for scoring regions across the genome for their likelihood of containing sequence content similar to the capture region.

[0072] Other potential noise sources were also investigated. Mapping possibility has historically been considered a source of bias in many DNA sequencing experiments. This bias is particularly problematic in single-end sequencing experiments because the mapped sequence length is shorter and therefore less unique compared to paired-end sequencing. Since most target panels are designed for paired-end sequencing, the negative impact caused by mapping possibility is minimal in these cases. However, to address any issues that may be associated with this noise source, we have created our own paired-end mapping possibility annotation. With these annotations, bins with less than 95% mapping possibility (a very small number) are blacklisted and do not require further correction. On the other hand, extensive testing has revealed that replication timing can sometimes be a source of bias. In fact, bin read counts were found to correlate highly with replication timing in certain cases, possibly in samples from invasive tumors with high replication rates. Still, replication timing was found to have an effect in only a small number of cases. Therefore, we have also devised a method to automatically identify these samples before correction, rather than correcting all samples by default.

[0073] While many methods exist for determining genome copy number profiling sequencing data, very few are designed to address targeted panel sequencing data, and therefore many do not correct for biases associated with targeted panel sequencing data. Even those that attempt to address specific features of targeted panel sequencing data have limitations, as will become apparent from the discussion below. Selected features of these methods are summarized in Table 1 and compared with the methods described herein (labeled “Disclosure”). These include: (i) whether the method can provide a genome-wide copy number profile for a single sample, even if the sample is a tumor sample (all prior art methods, even if they could perform such corrections, require a reference sample to correct for biases that may arise in tumor samples due to copy number abnormalities); (iii) whether the method uses only off-target reads (thereby avoiding probe affinity bias and heterogeneity of on-target reads); (iii) whether the method outputs a high-resolution genome-wide copy number profile or is limited to gene-based copy number; (iv) whether the method is replication timing correctable; and (v) whether the method applies GC-recognized corrections.

[0074] [Table 1]

[0075] CNVkit (Talevich et al. 2016) is used to provide genome-wide copy number profiles using both on-target and off-target reads in bin-level data. A pool of normal or tumor samples is used as a reference to address GC content bias, repeat mask fractionation bias, and target density bias. Copywriter (Kuilman et al. 2015) provides genome-wide copy number profiles using only off-target reads and addresses off-target capture bias, with or without corresponding normal samples, by applying a model-based analysis algorithm for ChIPseq (MACS) to detect peaks in germline samples enriched with various capture sets and removing those regions from the analysis. Control-FREEC (Boeva ​​et al. 2012) does not use off-target reads (based on the depth of coverage of captured exons) and does not provide genome-wide copy number profiles if the input is a target panel. Codex (Jiang et al. 2015) does not use off-target reads (based on the depth of coverage of captured exons) and does not provide genome-wide copy number profiles. ASCAT (Raine et al. 2016) uses both on-target and off-target reads to acquire bin-level data and obtain genome-wide copy number profiles from them. Bias is addressed using corresponding normal samples. ADTEx (Amarasinghe et al. 2014) does not provide genome-wide copy number profiles and uses only on-target reads. Any tumor-related bias is addressed using corresponding normal samples. FACETS (Shen and Seshan, 2016), depending on parameterization, can provide genome-wide copy number profiles using only on-target or both on-target and off-target reads, but only addresses tumor-related bias using corresponding normal samples.The CNV Radar (Soong et al. 2020) does not provide genome-wide copy number profiles; instead, it focuses on reporting the copy number status of oncogenes using only on-target reads. Any tumor-related bias is addressed using normal samples.

[0076] Of the methods described above, only CopywriteR is designed to use only off-target reads. The methods in the table below that use on-target reads are expected to have artifact peaks at target locations and are therefore unsuitable for providing genome-wide copy number profiles. Furthermore, many of the methods in the table above either provide only the copy number state of genes included in the panel, or are optimized for whole exomes, providing the copy number state of all genes and then using this to indirectly estimate the copy number state of the entire genome. This is different from actually estimating the copy number of the entire genome (even inter-gene sequences) from the available data. To the best of our knowledge, only two of the aforementioned methods, namely CNVkit and CopywriteR, use off-target reads and therefore actually query all locations in the genome regardless of the input. However, CNVkit does not use only off-target reads, and CopywriteR has limitations, at least because it removes data associated with off-target reads by peak determination, creating gaps in the copy number profile. Furthermore, all of these methods perform a simple GC correction by fitting a single model across all copy number states, and therefore cannot accurately perform GC correction on samples with high chromosomal instability.

[0077] To improve copy number determination from target capture panels, the inventors have developed a novel methodology that addresses at least four current shortcomings of existing methods. This novel method 1) works with tumor-only samples (without requiring corresponding normal samples), 2) corrects noise in the data rather than filtering it out, 3) is comprehensive, i.e., provides copy number status for the entire genome as well as specific genes, and 4) performs copy number recognition GC correction.

[0078] method Data. This method was demonstrated using a number of different capture-based target sequencing techniques. These include whole exome sequencing (WES), more specifically the SureSelect Human All Exon protocol (Agilent), the Oncology 500 assay (Illumina), the Oncology 170 assay, and non-targeted approaches such as sWGS. The method was tested with multiple different tumor types and different preservation methods, including fresh-frozen tissue, formalin-fixed paraffin-embedded (FFPE) tissue, and circulating tumor DNA (ctDNA) liquid biopsy, demonstrating its applicability in a wide range of situations. All samples used in this example are summarized in Table 2 below.

[0079] [Table 2]

[0080] In these examples, data are shown only for the first two datasets (WES data from lung adenocarcinoma cell lines and TSO500 for high-grade serous ovarian cancer tumors).

[0081] Read alignment and quality control. Paired-end raw reads were aligned to the human hg19 reference genome. This method uses paired-end binary sequence alignment files (BAM files) and target region coordinates (BED files) as inputs, both using a standard "hg19" human reference genome assembly without ALT contigs or decoy sequences. Duplicate reads, supplementary alignments, and reads that do not pass quality control (mapping quality score less than 37) are discarded.

[0082] Binning, annotation, and read counting. The resulting alignment is divided into equal-sized bins of 50, 100, or 500 kb size (i.e., the hg19 reference genome is divided into equal-sized windows of a size selected according to the desired resolution). In this example, a resolution was selected that yielded, on average, a minimum of 15 reads / bin per tumor copy (i.e., 15 reads per absolute copy number of tumor at the bin location). Note that this method worked equally well for all bin sizes tested, with the only difference being that larger bins yielded lower-resolution copy number profiles. The bins are then annotated with replication timing, GC content, mapability, and the percentage of identified nucleotides within the bin (i.e., non-NNNNN bases). The GC content and percentage of identified nucleotides are annotations inherited from QDNAseq (Scheinin et al. 2014). A mapping feasibility track suitable for paired-end data (500 mers, allowing two mismatches) was generated using GenMap (v1.3.0). GenMap calculates the uniqueness of k-mers for each position across the genome, including the number of mismatches (i.e., it calculates how unique a given sequence of length k in the genome is compared to all occurrences of k-mers in the genome). This measure of uniqueness provides an indicator of how easy it is to map a sequencing read to that position in the genome. A mapping feasibility track with 50 mers is suitable for establishing mapping feasibility for single-end sequencing, while a 500-mer track is suitable for paired-end sequencing.The inventors generated custom paired-end mapping possibilities for the hg19 human genome reference using GenMap (v.1.3.0) (Pockgrandt et al. 2020, available at github.com / cpockrandt / genmap), bedGraphToBigWig (UCSC Genome Browser Utilities), and the calculateMappability function of the QDNAseq R package (Scheinin et al. 014). The inventors tested the mapping possibilities of 50mer, 150mer, 300mer, 400mer, 500mer, 600mer, and 700mer, tolerating 2–4 mismatches. Since no significant difference was observed between the use of 400–700mer mapping possibilities, the inventors selected 500mer with 2 mismatches as a general-purpose solution for modern paired-end assays because these tend to have an average read length of 150 nt and various insert sizes generally around 200 nt in length. Replication timing was generated by calculating the average across read wavelet-smoothed repli-seq timing data from 15 different cell lines from UCSC (available at hgdownload.cse.ucsc.edu / goldenpath / hg19 / encodeDCC / wgEncodeUwRepliSeq / ). Other approaches can also be used, such as using replication timing data from a single cell line or control sample, for example, a cell line that has a high correlation with the sample being analyzed. The percentage of identified nucleotides and mapping possibility annotation were used to apply a quality control process; i.e., any bin that met any of the following criteria, i.e., GC content == NA, identified base ≠ 100%, or mapping possibility < 95%, was removed. This has been found to be a practical approach to removing unreliable bins while leaving sufficient data for high-quality, high-resolution copy number profiles. Any one or more of these criteria can be omitted, and other bin quality filtering strategies can be used.

[0083] Next, the annotated bins are queried for overlap with target regions using a BED file that lists the target regions, and are divided into two groups: 1) Off-target bins with zero overlap with target regions. In this case, reads are typically counted across the entire bin. 2) On-target bins with at least one overlap with the target region. In this case, reads anchored to on-target regions are removed (i.e., all reads in the BAM file that overlap with the region specified in the BED file are removed), and a new BAM file containing only off-target reads is generated. This is then scanned for counting. Thus, the bin width itself is not changed, but signals from capture regions are removed. This process excludes information such as variant determination or gene-level copy number determination of target genes, which are considered enriched regions in the genome that, if retained, would introduce a very high level of noise into downstream analyses. Since most target panel designs are paired-end, on-target reads with mate pairs outside the target coordinates are also excluded using read names that share a common portion in both mates.

[0084] Finally, use the Rsamtools "scanBam" function to count the remaining data (off-target reads only). Although off-target reads are unevenly distributed, they are still scattered throughout the genome and constitute the information from which a genome-wide copy number profile can be extracted.

[0085] GC correction. In all cases, this method applies local polynomial (LOESS) curve fitting and correction approaches to correct for GC content bias. Two different types of correction were performed depending on the CIN level. In both types, GC LOESS fitting and correction were performed, i.e., 1) Stable tumors (NO-CIN - samples with only one major copy number state): Perform a single LOESS on the entire dataset at once. All data points are fitted at once based on GC content using the following formula: loess(raw counts~gc). Each bin read count is then corrected by subtracting the difference between the fitted value and the median for the bin's GC value. This subtracts the effect of GC content and removes all correlation between counts and GC. 2) Unstable Tumors (CIN): Copy number-recognized GC correction, multiple LOESS, one per copy number state. Samples with high CIN show significant differences in copy number states across the genome. For this reason, this method performs copy number-recognized GC correction to avoid an undesirable contraction effect on the median of raw counts, which would eliminate true biological variability if a single LOESS were applied. Thus, this method better accounts for true biological variability when correcting GC content, which ultimately leads to a more robust absolute fitting, as it results in a better estimate of clonality, purity, and ploidy. To the best of our knowledge, other bulk copy number methods do not do this. Copy number-recognized correction can be performed using SCOPE (single-cell method). Copy number-recognized correction requires labeling each data point with its copy number state in order to perform LOESS fitting separately for each copy number state. Since the final true copy number state is unknown at this stage, this method attempts to infer it by performing segmentation and clustering following the initial GC correction, using the segment values ​​assigned to each data point (bin) as input. More specifically, the read counts, once GC-corrected, are used as input to the segmentation algorithm, and then the bins are assigned the values ​​of the segments to which they correspond, and the bins are clustered using the segment values ​​assigned to them. The clustering part uses a modified version of k-means from the "Ckmeans.1d.dp" R package, which automatically selects the optimal number of clusters from the range of 1 to 25 using the Bayesian Information Criterion (BIC). Two constraints are then applied in the iterative process until both are satisfied: 1) every cluster should have at least 5 data points so that LOESS can be performed, and 2) the minimum difference between all cluster centers must be at least 2. In each iteration, if there are pairs of clusters and / or clusters with fewer than 5 data points where the Euclidean distance between centers is less than 2, one cluster is removed by re-running the clustering with the maximum number of clusters reduced by 1 compared to the previous iteration.

[0086] In this example, samples with a copy number change of less than 1 that differs from the diploid state were considered "NO-CIN". Samples with a copy number change of 1 or more that differs from the diploid state were considered "CIN". This was evaluated by scrutinizing the clustering results, and if more than 1 clusters were identified, copy number recognition GC correction was performed.

[0087] Off-target capture bias score correction. Off-target reads are assumed to be uniformly distributed across the genome and appear as peaks that are inherent to this distribution and are natural disruptors of the target sequencing data. The left-end pairwise distance between pairs of all reads in each bin is calculated. The distribution of left-end pairwise distances for all reads in a bin produces a left triangular distribution with a consistent shape in a normal, unbiased bin (where the minimum distance value, which is usually zero, is equal to the mode). The mode is usually the observed minimum distance value, which is usually zero, because sequencing data typically contains sets of many reads that start from the same location due to fragmentation bias. Indeed, genomic DNA is typically fragmented during the sequencing library preparation process, cell-free DNA is naturally fragmented, and both fragmentation processes are non-random (i.e., biased). In contrast, off-target bins with artificially high read counts (peaks) are measurable and exhibit some disruption proportional to the read count (see Figure 8A). To capture this and correct for heterogeneously distributed reads across bins in capture-based DNA target sequencing assays, three different approaches were developed. 1) Histogram-based distance score (see Figure 8A). Here, the inventors first use the "dtri" function from the "triangulr" R package (cran.r-project.org / web / packages / triangulr / ) to calculate the density function using the observed minimum and maximum bin distances (or bin size + average read length) as parameters from the left triangular distribution. Next, the inventors generate a histogram from the observed pairwise distances and calculate the midpoint vector of the histogram bins. Using the left triangular density, the inventors estimate the expected value of each histogram bin midpoint and then subtract the resulting vector from the observed histogram values ​​to obtain a vector that reflects the observed deviation from the theoretical left triangular distribution. By converting this to an absolute value vector and calculating its mean, a final score is provided that can be used for noise correction. 2) Probability-based distance score (see Figure 8B). In this case, the inventors use the same minimum, maximum, and mode parameters as used in the histogram-based score (i.e., mode = minimum observed distance, minimum = minimum observed distance, maximum = maximum observed distance within the bin, or bin size + average read length), but directly use the matrix of pairwise distances as input to the "ptri" function to compute the distribution function, and thus calculate the probability that each pair within the bin is sampled from a left triangular distribution. The inventors subtract 1 from the mean of the resulting probability vector. This is the final probability-based score. 3) Statistical Metric for Distribution Difference: This score is generated for each bin by comparing the vector of all observed left-end pairwise distances in the bin to a theoretical distance as if it were sampled from a right-skewed triangular distribution with minimum = 0, mode = 0, and maximum = bin size + mean read length (the maximum is 50150 for a 50kb bin, resulting from 50000 + 150). This comparison can be done using any statistical metric that can compare two distributions, such as the Kolmogorov-Smirnov test, the DTS test, and many others. In the data shown in Figure 9C, the DTS statistic is used. This is a test statistic that compares two empirical cumulative distribution functions (ECDFs) by looking at the reweighted Wasserstein distance between these two (see Dowd 2020). This is calculated here using the "dts_stat" function from the "twosamples" R package (search.r-project.org / CRAN / refmans / twosamples / html / two_sample.html). This R package contains functions for performing statistical tests that allow comparing two distributions, including the Kolmogorov-Smirnov test, the DTS test, and many others. While any of these can be used, the DTS statistic is considered to have higher power and has been shown to produce good results for the current particular purpose (see Figure 9C). In this particular case, to obtain the DTS score, the dts_stat function takes as input a vector of observed distances in bins and a vector whose theoretical vector is equal in length to the observed vector, randomly generated using the "rtri" function from the "triangulr" R package (cran.r-project.org / web / packages / triangulr / triangulr.pdf) with the minimum, maximum, and mode parameters described above. As described in Dowd 2020, the DTS test compares two empirical cumulative distribution functions (ECDFs) generated from the two input vectors and then calculates the reweighted Wasserstein distance between these two.Formally, if E is the ECDF of sample 1, F is the ECDF of sample 2, and G is the ECDF of the combined sample, then the following equation holds:

number

[0088] Any of the above scores can be used for LOESS correction in another round using already GC-corrected counts. In other words, the LOESS model can be fitted as a function of the score to a distribution of read counts per bin (optionally already GC-corrected). The observations for each bin are then corrected by subtracting the difference between the fitted value for the bin's score and the median across the bins (equivalent to bending the curve to a horizontal line at the raw median count per bin). Both scores have been shown to yield similar results. The histogram-based score performs slightly better overall and is therefore used by default, but the user is given the option for different parameterizations. For samples with very low read counts, the probability-based score may be advantageous as it does not require fitting a density distribution / histogram to the data, which can be difficult when there is not much data. Either the histogram-based or probability-based score can be obtained for any bin containing at least 10 reads. For bins containing fewer than 10 reads, the score can be set to a default value equal to the minimum score calculated for the sample.

[0089] Processing of On-Target Bins. On-target bins are defined as bins with any overlap with the on-target region. These are very few in small panels, about 1-10% of the total number of bins, but about half of the total number of bins in WES samples. For these bins, for each on-target region from which reads have been removed, a pseudo-count is then added, proportional to the percentage of target region overlap relative to the total bin length. The pseudo-count is estimated from the remaining reads in the bin after GC correction, capture bias correction, and replication timing correction (if performed). Specifically, the pseudo-count is calculated as pseudo-count = (percentage of overlap with bed file) * (count for bin after all corrections have been applied). Therefore, the final count is (count for bin after all corrections have been applied) + pseudo-count.

[0090] Replication timing correction. Since not all tumor samples require replication timing, only samples with a Pearson correlation (r) of 0.15 or higher between bin read count value and replication timing score are corrected. For other biases, the same correction procedure, namely LOESS fitting and transformation, was applied.

[0091] Segmentation and absolute copy number fitting. Corrected bin read counts were subjected to segmentation using a modified version of circular binary segmentation (CBS) from the DNAcopy R package. In particular, this method uses the “segment” function from the DNAcopy R package, based on a circular binary segmentation approach. The following parameters are selected: transformFun=“sqrt”, alpha=1e-10, undo.splits=“sdundo”, undo.SD=1. These parameters are considered optimal for sequencing depths of at least 15 reads / bin per chromosome copy (see, e.g., Scheinin et al. 2014). Once all possible corrections are applied, the resulting relative copy number profiles are then subjected to absolute fitting. This is a process of manually selecting the most likely purity-ploidy combination in the ranges 0.05 < purity < 1 and 1.5 < ploidy < 8, as shown in the following equation:

number

number

[0092] result A novel method for performing copy number determination from a target capture panel is proposed and shown in Figure 7. Paired-end raw reads are aligned to a reference genome. The resulting alignment is divided into bins of equal size. Sizes of 50, 100, or 500 kb were tested and similar results were obtained. The bins are then annotated with information including GC content, mapping potential, and replication timing. The annotated bins are then divided into two groups depending on whether they overlap with target regions: 1) off-target bins with zero overlap with target regions, and 2) on-target bins with at least one overlapping target region. For on-target bins, a pseudocount is added as described above. The bins are then corrected for GC content and segmented using a single LOESS fitted to the entire dataset at once. For samples without chromosomal instability (i.e., mostly diploid samples), the samples then proceed to capture bias score correction. For samples with high chromosomal instability, clustering is performed, followed by “copy number recognition GC correction,” where individual LOESS models are fitted to the data for each observed copy number state. Each copy number state is then corrected and segmented separately with respect to GC content, as shown in Figure 6. Subsequently, both CIN and NO-CIN are corrected for read counts per bin to account for bias due to off-target capture. Target sequencing assays artificially generate high off-target read counts in genomic regions with high sequence similarity to the target region. Therefore, a new score per bin is generated to quantify the magnitude of this bias. This can be done using one of two methods (histogram-based distance score or probability-based distance score), as described above and shown in Figure 8. This is then used for a single LOESS fitting and correction (i.e., correcting all bins at once).

[0093] In some cases, replication timing has been observed to introduce a systematic bias into the read count data. This is automatically identified and corrected using methods similar to those used for GC content or capture bias correction (i.e., by fitting a LOESS model to the bin read count as a function of replication timing and adjusting the count per bin to remove the bias estimated as a value in the LOESS model). Note that in this workflow, separate univariate models are fitted to correct for GC content bias and capture bias (and replication timing, if used). However, multivariate models can be fitted for any combination of GC content bias, capture bias, and replication timing bias. Finally, the counts for on-target bins are readjusted to account for the removed data. The conversion of on-target bins to off-target bins results in an artificially lower read count due to the gaps left by deduplication. To address this, a pseudo-count directly proportional to the size of the overlapping region is added, essentially filling the gaps left by the on-target region after their exclusion. For example, if 20% of a bin is covered by the on-target region, 0.2*read count is added to the observed read count for the bin.

[0094] The above results in bin read counts that are essentially equivalent to those obtained from whole-genome sequencing. These can be considered reliable and can be analyzed using similar methods to those used for data obtained from whole-genome sequencing. Therefore, each sample is subjected to segmentation and absolute copy number fitting (aimed at identifying the most likely combination of purity and ploidy with the lowest clonality) using methods well known in the art.

[0095] Example 2 - Performance Evaluation Since shallow whole-genome sequencing (sWGS) is the current gold standard for low-resolution genome-wide copy number profiling, the method described in Example 1 was benchmarked by comparison with this assay using a set of samples in which both sWGS and target sequencing data were available. For sWGS, reads were aligned as single-ended to the human genome assembly GRCh37 using BWA-MEM (v0.7.17, Li 2013). Duplicate reads were then identified and marked using samtools-markdup (v1.15). The inventors used the QDNAseq R package to count reads in 50kb bins. Bins mapped to centromere and undefined sequence regions in the reference genome hg19 were excluded. Read counts were then corrected for sequence mapping potential and GC content, followed by copy number segmentation. After the profile was segmented, the inventors estimated the absolute copy number as outlined in Example 1. Each sWGS and capture sequence copy number profile was compared pairwise using their rounded segment values ​​with four different metrics from the CNpare package (Chaves-Urbano et al. 2022): Pearson's r, cosine similarity, Manhattan distance, and Euclidean distance.

[0096] An example is shown in Figure 9, which displays superimposed sWGS-WES cancer cell line data (Figure 9A) and superimposed sWGS-TSO500 FFPE clinical sample data (Figure 9B) processed according to this method for samples analyzed in the datasets listed in Table 2 (first two rows). The pair copy number profiles from the WES-sWGS lung cancer cell line and TSO500-sWGS FFPE ovarian cancer samples show that the genomes differ by 11.57% and 1.39%, respectively, and that the results obtained using this method applied to the capture data are very similar to the sWGS data.

[0097] Figure 10 shows the results of applying this method to WES data from three lung cancer cell lines with and without bias correction to demonstrate the benefits of bias correction. Across all metrics, including Pearson r, Manhattan distance, Euclidean distance, and cosine similarity, including bias correction yields copy number profiles derived from WES that are more similar to sWGS than without bias correction.

[0098] Figure 11 shows the results of applying two other genome-wide copy number profiling methods (CopywriteR and CNVkit) that are publicly available for targeted sequencing panels, in addition to the disclosed methods, to 26 FFPE ovarian cancer samples profiled using the TSO500 gene panel assay (see emea.illumina.com / content / dam / illumina / gcs / assembled-assets / marketing-literature / trusight-oncology-500-data-sheet-m-gl-00173 / trusight-oncology-500-and-ht-data-sheet-m-gl-00173.pdf) and sWGS. In addition, these samples were also profiled using the TSO500+HRD panel (see www.illumina.com / content / dam / illumina / gcs / assembled-assets / marketing-literature / tso500-hrd-data-sheet-m-gl-00748 / tso500-hrd-data-sheet-m-gl-00748.pdf), which has approximately 25,000 additional capture regions to query for single nucleotide polymorphisms (SNPs) so that the copy number profile can be determined using an SNP-based approach (similar to ASCAT). Pearson's r of the output of each method compared to the sWGS gold standard showed that the TSO500 results with CopyRight performed better than CopywriteR and CNVkit and were comparable in performance to TSO500+HRD.

[0099] Example 3 - Application to chromosomal instability signatures The method described herein enables the quantification of chromosomal instability signatures from samples sequenced using a targeted sequencing assay. This typically requires a genome-wide copy number profile at a resolution of at least 50 kb, which was not available for this type of sample prior to the present invention.

[0100] After copy number profiles are obtained as described in Example 1, chromosomal instability signatures can be calculated as described in Macintyre et al., 2018 and Drews et al. 2022. This was previously simply impossible. For example, as described in Drews et al. 2022, these signatures can be used to predict drug responses, identify novel drug targets, and understand tumor pathogenesis.

[0101] Figure 12 shows a heatmap of raw CIN signature activity obtained from both WES and sWGS of three lung adenocarcinoma cell lines. Samples are represented in rows, and CIN signatures are represented in columns. The heatmap shows hierarchical clustering applied to the rows, which results in each pair being grouped and illustrating the similarity of CIN signatures obtained from both approaches (WES and sWGS).

[0102] Figure 13 shows a comparison of the performance of TSO500 (using Copyright), TSO500+HRD, and sWGS in extracting CIN signatures using all clinical FFPE ovarian cancer samples detailed in Table 2. Figure 13A shows a stacked bar graph of raw signature activity. Figure 13B shows the cosine similarity of TSO500 and TSO500+HRD to sWGS. These results indicate that TSO500 (using Copyright) performs comparably to TSO500+HRD when extracting CIN signatures, and that many samples show high similarity (cosine > 0.8) to sWGS.

[0103] Example 4 - Discussion This embodiment presents and demonstrates a novel method for analyzing data from capture-based targeted sequencing technologies. This method was demonstrated using several different capture-based targeted sequencing technologies. These include whole exome sequencing (WES), more specifically the SureSelect Human All Exon protocol (Agilent), widely used in academia for small variant discovery in any protein-coding exon, and the TruSight Oncology 500 assay (Illumina), which is regulatory-approved in a cancer clinical setting and provides information on 523 genes and a set of clinically relevant biomarkers such as TMB, MSI, or HRD. Other protocols used and tested were the TruSight Tumor 170 assay and custom panel designs targeting 171 genes. Amplicon-based technologies could not be used because they do not generate off-target reads.

[0104] This method is applicable to many different sample types (e.g., FFPE, fresh-frozen, cfDNA).

[0105] In the context of liquid biopsy, this method may be useful not only for generating copy number profiles but also for estimating ctDNA proportions from panels, similar to what the ichorCNA algorithm would do with sWGS. This may be particularly useful given the existence of regulatory-approved and clinically used targeted liquid biopsy assays and the emerging emergence of tumor fractions as powerful indicators of treatment response, prognosis, and recurrence. Therefore, a method for obtaining improved estimates of ctDNA from targeted sequencing assays in the context of cfDNA would be a valuable tool in clinical cancer care.

[0106] References Talevich E, Shain AH, Botton T, Bastian BC (2016) CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing. PLoS Comput Biol 12(4): e1004873. Kuilman, T., Velds, A., Kemper, K. et al. CopywriteR: DNA copy number detection from off-target sequence data. Genome Biol 16, 49 (2015). Boeva V, Popova T, Bleakley K, Chiche P, Cappo J, Schleiermacher G, Janoueix-Lerosey I, Delattre O, Barillot E. Control-FREEC: a tool for assessing copy number and allelic content using next-generation sequencing data. Bioinformatics. 2012 Feb 1;28(3):423-5. Jiang Y, Oldridge DA, Diskin SJ, Zhang NR. CODEX: a normalization and copy number variation detection method for whole exome sequencing. Nucleic Acids Res. 2015 Mar 31;43(6):e39. Raine KM, Van Loo P, Wedge DC, Jones D, Menzies A, Butler AP, Teague JW, Tarpey P, Nik-Zainal S, Campbell PJ. ascatNgs: Identifying Somatically Acquired Copy-Number Alterations from Whole-Genome Sequencing Data. Curr Protoc Bioinformatics. 2016 Dec 8;56:15.9.1-15.9.17. Amarasinghe, K.C., Li, J., Hunter, S.M. et al. Inferring copy number and genotype in tumour exome data. BMC Genomics 15, 732 (2014). Shen R, Seshan VE. FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing. Nucleic Acids Res. 2016 Sep 19;44(16):e131. Soong, D., Stratford, J., Avet-Loiseau, H. et al. CNV Radar: an improved method for somatic copy number alteration characterization in oncology. BMC Bioinformatics 21, 98 (2020). Langmead, B., Trapnell, C., Pop, M. et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009). Carter, S., Cibulskis, K., Helman, E. et al. Absolute quantification of somatic DNA alterations in human cancer. Nat Biotechnol 30, 413-421 (2012). Favero F, Joshi T, Marquard AM, Birkbak NJ, Krzystanek M, Li Q, Szallasi Z, Eklund AC. Sequenza: allele-specific copy number and mutation profiles from tumor sequencing data. Ann Oncol. 2015 Jan;26(1):64-70. Adalsteinsson, V.A., Ha, G., Freeman, S.S. et al. Scalable whole-exome sequencing of cell-free DNA reveals high concordance with metastatic tumors. Nat Commun 8, 1324 (2017). Drews, R.M., Hernando, B., Tarabichi, M. et al. A pan-cancer compendium of chromosomal instability. Nature 606, 976-983 (2022). Macintyre, G. et al. Copy number signatures and mutational processes in ovarian carcinoma. Nat. Genet. 50, 1262-1270 (2018). Milbury CA, Creeden J, Yip WK, Smith DL, Pattani V, Maxwell K, Sawchyn B, Gjoerup O, Meng W, Skoletsky J, Concepcion AD, Tang Y, Bai X, Dewal N, Ma P, Bailey ST, Thornton J, Pavlick DC, Frampton GM, Lieber D, White J, Burns C, Vietz C. Clinical and analytical validation of FoundationOne® CDx, a comprehensive genomic profiling assay for solid tumors. PLoS One. 2022 Mar 16;17(3):e0264138. Chen Zhao, Tingting Jiang, Jin Hyun Ju, Shile Zhang, Jenhan Tao, Yao Fu, Jenn Lococo, Janel Dockter, Traci Pawlowski, Sven Bilke. TruSight Oncology 500: Enabling Comprehensive Genomic Profiling and Biomarker Reporting with Targeted Sequencing. bioRxiv 2020.10.21.349100 C. Heydt, R. Pappesch, K. Stecker, J. Neumann, R. Buettner, S. Merkelbach-Bruse. Evaluation of the TruSight Tumor 170 (TST170) assay and its value in clinical research. Annals of oncology. volume 29, supplement 6, vi7-vi8, september 2018. Wang, R., Lin, D-Y, Jiang Y. SCOPE: A Normalization and Copy-Number Estimation Method for Single-Cell DNA Sequencing. Cell Systems Vol. 10, Issue 5, P445-452, May 20, 2020. Christopher Pockrandt and others, GenMap: ultra-fast computation of genome mappability, Bioinformatics, Volume 36, Issue 12, June 2020, Pages 3687-3692. Mehran Karimzadeh and others, Umap and Bismap: quantifying genome and methylome mappability, Nucleic Acids Research, Volume 46, Issue 20, 16 November 2018, Page e120. Scheinin I, Sie D, Bengtsson H, van de Wiel MA, Olshen AB, van Thuijl HF, van Essen HF, Eijk PP, Rustenburg F, Meijer GA, Reijneveld JC, Wesseling P, Pinkel D, Albertson DG, Ylstra B. DNA copy number analysis of fresh and formalin-fixed specimens by shallow whole-genome sequencing with identification and exclusion of problematic regions in the genome assembly. Genome Res. 2014 Dec;24(12):2022-32. Macintyre G, Ylstra B, Brenton JD Sequencing Structural Variants in Cancer for Precision Therapeutics. Trends in Genetics. Vol. 32, Issue 9, P530-542, September 2016. Li H, Durbin R. 2009. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25: 1754-1760. Li H. (2013) Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv:1303.3997v2 Blas Chaves-Urbano and others, CNpare: matching DNA copy number profiles, Bioinformatics, Volume 38, Issue 14, July 2022, Pages 3638-3641 Dowd, Connor, A New ECDF Two-Sample Test Statistic. arXiv:2007.01360v1.2 July 2020.

[0107] All references cited herein are incorporated herein by reference in whole for any purpose to the same extent that each individual publication or patent or patent application is specifically and individually indicated to be incorporated by whole.

[0108] The specific embodiments described herein are provided as examples, not as limitations. Various modifications and variations of the use of the compositions, methods, and techniques described herein will be apparent to those skilled in the art without departing from the scope and spirit of the techniques described herein. Any subtitles herein are included for convenience only and should not be construed as limiting the disclosure. Unless otherwise indicated in the context, the descriptions and definitions of features set forth above are not limited to any particular aspect or embodiment of the invention and apply equally to all aspects and embodiments described herein.

[0109] Any method of any embodiment described herein may be provided as a computer program, or as a computer program product or computer-readable medium holding a computer program configured to perform the above method when executed on a computer.

[0110] Throughout the specification and claims, unless the context clearly indicates otherwise, the following terms have the meanings expressly associated herein. Where used herein, the phrase “in one embodiment” may refer to the same embodiment, though not necessarily the same embodiment. Furthermore, where used herein, the phrase “in another embodiment” may refer to a different embodiment, though not necessarily a different embodiment. Thus, various embodiments of the present invention can be readily combined without departing from the scope or spirit of the invention, as described below. Where used herein and in the appended claims, the singular forms “a,” “an,” and “the” should be noted to include multiple references unless the context clearly indicates otherwise. Ranges may be expressed herein as “about” a particular value to and / or “about” another particular value. Where such ranges are expressed, another embodiment includes a range from and / or other particular values. Similarly, where a value is expressed as an approximation by the preceding use of “about,” it will be understood that a particular value constitutes another embodiment. The term “approximately” in relation to numerical values ​​is optional and means, for example, + / - 10%. Throughout this Spec., including the following claims, unless the context requires otherwise, the words “comprise” and “inclusive,” as well as variations such as “comprises,” “inclusive,” and “inclusive,” are understood to mean that they include the integer or process or group of integers or processes described, but do not exclude any other integer or process or group of integers or processes. Other aspects and embodiments of the Invention provide the above aspects and embodiments in which the term “comprising” is replaced with the terms “consisting of” or “consisting essentially of,” unless the context requires otherwise. As used herein, “and / or” should be interpreted as a specific disclosure of each of two particular features or components, with or without the other.For example, “A and / or B” should be interpreted as each of the specific disclosures of (i) A, (ii) B, and (iii) A and B, as if each were described separately herein. The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, are expressed in relation to means for performing the disclosed functions or methods or processes for obtaining the disclosed results, and may be used, separately or in any combination as needed, to realize the invention in its various forms.

Claims

1. A computer-based method for analyzing genome sequencing data, Obtaining genome sequence data containing multiple DNA sequencing reads from one or more samples, From the aforementioned genome sequence data, the observed read count for each of the multiple genome bins is determined, Determining a corrected read count for one or more of the plurality of genome bins, wherein the corrected read count is determined based on the observed read count and one or more estimates of the bias in the read count as a function of the variables associated with each bias. Includes, The genome sequencing data is target genome sequencing data, the bias includes capture bias, and the variable associated with the bias is a capture bias score that represents the difference between the observed read position metric distribution in the bin and the expected read position distribution metric in the bin.

2. The method according to claim 1, wherein an estimate of the bias in the read count as a function of a variable associated with the bias is estimated by fitting a model between (i) the observed read counts in the plurality of genome bins and (ii) the values ​​of the bias variable for the plurality of genome bins.

3. The method according to claim 2, wherein the model is a curve-fitting model, optionally a regression model, preferably a local polynomial regression model.

4. The method of claim 2 or 3, wherein determining a corrected read count for a bin includes subtracting from the observed read count a correction coefficient equal to the difference between (i) the model value at coordinates corresponding to the value of the bias variable for the bin and (ii) a reference value, which optionally is a summary value derived from the observed read counts for the plurality of bins.

5. The method according to any one of claims 1 to 4, wherein the expected read position metric distribution within the bin is a read position metric distribution corresponding to a uniform distribution of read positions within the bin.

6. The method according to any one of claims 1 to 5, wherein the position of the reads in the bin is represented by a metric of the pairwise distance between pairs of reads in the bin, and / or a capture bias score is a variable that captures the discrepancy between the observed distribution of the pairwise distance between reads in the bin and the expected distribution of the pairwise distance between reads in the bin.

7. The method according to claim 6, wherein the pairwise distance between leads is the left end pairwise distance, the right end pairwise distance, or the center pairwise distance.

8. The expected distribution of pairwise distances between reads in a bin is: The mode is 0 or the minimum observed pairwise distance within the bin. The minimum value is 0 or the minimum observed pairwise distance within the bin. The maximum value is the maximum observed pairwise distance within the bin, the length of the bin, the sum of the length of the bin and the length of the lead, or any value between the length of the bin and the sum of the length of the bin and the length of the lead. It is a left triangular distribution, Optionally, the expected distribution of pairwise distances between reads in a bin is approximated by sampling a number of observations corresponding to the number of observed pairwise distances between reads in the bin from the left triangular distribution. The method according to claim 6 or 7.

9. The method according to any one of claims 1 to 8, wherein the capture bias score is a variable that summarizes the absolute difference between the frequency of the window in the observed read position metric distribution for the bin and the frequency of the window in the expected read position metric distribution for the bin, across a plurality of windows of the read position metric values.

10. The method according to any one of claims 1 to 8, wherein the capture bias score is a variable that summarizes the probability that the observed read position metric was sampled from the expected read position metric distribution of the bin, across a plurality of observations of the read position metric within the bin.

11. The method according to any one of claims 1 to 8, wherein the capture bias score is a statistical metric that compares two samples to determine whether the two samples are from the same distribution, and optionally the statistical metric is a DTS statistic.

12. The method according to any one of claims 1 to 11, wherein the genome sequence data includes on-target reads that are at least partially mapped to a capture region and off-target reads that are not mapped to a capture region, the observed read count is obtained using the off-target reads, and / or the method includes selecting reads that are not mapped to a capture region and obtaining the observed read count using the selected reads.

13. The method according to any one of claims 1 to 12, comprising obtaining the plurality of genomic bins by applying a non-overlapping fixed-width window along the genomic coordinates with respect to one or more genomic regions that map the read, thereby obtaining the plurality of fixed-width bins.

14. Determining the corrected read count for one or more of the aforementioned genome bins is equivalent to determining the corrected read count for a bin containing one or more on-target reads. A first corrected read count is obtained based on the observed off-target read count for the bin and one or more estimates of the bias in the read count as a function of the variables associated with each bias. A second corrected read count is obtained by adding a pseudo-count to the first corrected read count, wherein the pseudo-count is proportional to the ratio of the overlap between the first read count and the bin with the capture region. The method according to claim 12 or 13, which includes determining by

15. The bias further includes a GC content bias, and the method provides an estimate of the GC content bias in the read count as a function of GC content. Identifying multiple bin sets in which all bins within a set have the same estimated number of copies, (i) fitting a model between the observed read count in one of the bin sets, which includes the bin in which the corrected read count is determined, and (ii) the GC content value for the bin in the set, A computer implementation method according to any one of claims 1 to 14, further comprising obtaining by

16. The method according to claim 15, wherein identifying multiple sets of bins in which all bins have the same estimated copy number includes clustering the raw read count, corrected read count, or segment read count associated with the bins.

17. The computer implementation method according to any one of claims 1 to 16, wherein the target sequencing data is capture-based sequencing data and optionally comes from whole exome sequencing or target panel sequencing.

18. The computer-aided method according to any one of claims 1 to 17, wherein the sample is a tissue sample, an FFPE tissue sample, a fresh frozen tissue sample, or a liquid biopsy sample.

19. The computer-aided method according to any one of claims 1 to 18, wherein the sample is a sample containing tumor cells or genetic material derived therefrom, or optionally, a sample containing tumor biopsy or cell-free DNA.

20. The computer-aided method according to claim 19, wherein the one or more samples do not include a corresponding normal sample, or the one or more samples include tumor cells or genetic material derived therefrom.

21. The computer implementation method according to any one of claims 1 to 20, wherein the bias further includes a replication timing bias, the variable associated with the bias being replication timing, and optionally the method comprises determining whether the observed read counts for the plurality of bins are significantly associated with replication timing, and correcting the read counts with respect to replication timing bias if the observed read counts for the plurality of bins are significantly associated with replication timing.

22. A computer-aided method according to any one of claims 1 to 21, further comprising obtaining corrected read counts for each of the plurality of bins, optionally obtaining copy number profiles from the corrected read counts for the plurality of bins by segmentation and absolute copy number fitting, and / or obtaining information derived from the copy number profiles obtained from the corrected read counts, optionally including the presence of one or more chromosomal instability signatures and / or the proportion of tumor cells or tumor DNA in the sample.

23. A method for providing a copy number profile of a sample or genomic features derived from a copy number profile, Obtaining genome sequence data from the sample, including multiple DNA sequencing reads obtained using a target genome sequencing assay, Analyzing the data using the method described in any one of claims 1 to 22, Using the corrected read count, obtain copy number profiles and / or genomic features derived from copy number profiles, A method comprising, optionally, obtaining a tumor copy number profile and the proportion of tumor DNA or tumor cells having the tumor copy number profile.

24. A system comprising: a processor; and a computer-readable medium containing instructions that, when executed by the processor, cause the processor to perform a step of any of the methods described herein, such as any of claims 1 to 23.

25. One or more non-temporary computer-readable media containing instructions that, when executed by one or more processors, cause one or more processors to perform a step of any of the methods described herein, such as any of claims 1 to 23.