Methods and systems for allele-specific copy number calling

The method addresses the challenges of ASCN calling in SBX chemistry by employing data processing techniques to enhance precision and reliability, particularly in cancer samples, through normalization and statistical modeling.

WO2026050541A1PCT designated stage Publication Date: 2026-03-05ROCHE SEQUENCING SOLUTIONS INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/044005
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-08-28
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Traditional methods for allele-specific copy number (ASCN) calling are inadequate for sequencing data collected by sequencing-by-expansion (SBX) chemistry due to chemistry differences, leading to inferior performance and challenges in estimating tumor purity and ploidy, particularly in complex configurations like cancer samples.

Method used

A method for ASCN calling tailored to SBX chemistry, involving data processing steps such as coverage ratio calculation, filtering, segmentation, and estimation of tumor purity and ploidy using B-allele frequency (BAF) values, with normalization for GC content bias and application of statistical models like Hidden Markov Models.

Benefits of technology

Enhances the reliability and precision of ASCN calling for SBX-based sequencing data, supporting various configurations and disease contexts, and improves the accuracy of tumor purity and ploidy estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044005_05032026_PF_FP_ABST
    Figure US2025044005_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are systems and methods for allele-specific copy number (ASCN) calling using sequencing data. The presently described systems and methods can support various configurations based on different sequencing assays and disease contexts. The techniques include segmentation of target regions of interest based on similarity of coverage ratio and B-allele frequency (BAF) values for adjacent target regions, and estimation of an allele-specific copy number for each segment based at least one the observed coverage ratio, an expected coverage ratio, an observer BAF value, and an expected BAF value for each segment, wherein the expected coverage ratio and / or the expected BAF are calculated based on tumor purity, tumor ploidy, and total copy number for each segment.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Client Reference No.: P39388-WO-1 INTERNATIONAL PATENT APPLICATION Title: METHODS AND SYSTEMS FOR ALLELE-SPECIFIC COPY NUMBER CALLING Inventors: Fong Chun Chan, a Canadian citizen, resident of Richmond, BC, Canada Seyed Parsa Eskandar, a U.S. citizen, resident of Santa Cruz, CA Mahdi Golkaram, a U.S. citizen, resident of San Diego, CA Taher Kasim Mun, a U.S. citizen, resident of West Chester, PA Fan Song, a citizen of China, resident of San Diego, CA Assignee: Roche Sequencing Solutions, Inc. 4300 Hacienda Drive Pleasanton, CA 94588 United States of America Entity: LargePATENT Client Reference No.: P39388-WO-1 METHODS AND SYSTEMS FOR ALLELE-SPECIFIC COPY NUMBER CALLING CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims priority to United States Provisional Patent Application No.63 / 688,074, filed on August 28, 2024, which is herein incorporated by reference in its entirety. BACKGROUND

[0002] In the field of genomics, accurate detection and quantification of allele-specific copy number (ASCN) variations are crucial for understanding the genetic alterations associated with various diseases, such as cancer. Traditional methods for ASCN calling are well-established for sequencing data collected using sequencer instruments from Illumina, Inc., but face challenges when applied to sequencing data collected by other types of sequencer instruments, particular sequencer instruments that rely on sequencing-by- expansion (SBX) chemistry. This is due to chemistry differences resulting in SBX data having different error profiles, stronger GC bias, lower uniformity of coverage, lower and variable insert sizes, and so forth. These differences can result in methods designed for sequencing data collected by Illumina sequencer instruments having inferior performance when instead applied to sequencing data collected on sequencer instruments optimized for SBX chemistry. Thus, there exists a need for a robust, accessible, and accurate method of ASCN calling tailored to SBX-based sequencing data, especially in complex configurations like cancer samples where tumor purity and ploidy also need to be estimated. SUMMARY

[0003] The embodiments described herein relate to systems and methods for secondary analysis of sequencing data. More particularly, the embodiments described herein related to allele-specific copy number calling based on sequencing data produced using sequencing by expansion (SBX) chemistry.

[0004] The present disclosure addresses the challenges set forth above by introducing a method for ASCN calling tailored specifically to sequencing data collected using SBX chemistry. The techniques disclosed herein are capable of supporting multiple configurations,PATENT Client Reference No.: P39388-WO-1 adapting to various sequencing assays and disease contexts, thereby enhancing the reliability and precision of ASCN calling for SBX-based sequencing data.

[0005] In accordance with a first aspect of the present disclosure, a computer- implemented method for allele-specific copy number (ASCN) calling is provided. The method includes: receiving sequencing data comprising a plurality of target regions of interest derived from a sample; calculating a coverage ratio for each target region of the plurality of target regions based on the sequencing data; receiving variant call information derived from the sequencing data; filtering the variant call information according to one or more filter criteria; segmenting the target regions to define a plurality of segments based at least in part on the coverage ratio for each target region and at least one B-allele frequency (BAF) value for one or more filtered variants in the variant call information within the target region; and estimating an allele-specific copy number (ASCN) state for a segment of the plurality of segments based at least in part on a total copy number, a tumor purity, and a tumor ploidy for the segment. The tumor purity and tumor ploidy are estimated using the coverage ratio, an expected coverage ratio, the BAF value, and an expected BAF value for each segment.

[0006] In at least some embodiments of the first aspect, the method further includes estimating the total copy number for the segment based at least in part on the coverage ratio, the tumor purity, and the tumor ploidy.

[0007] In at least some embodiments of the first aspect, the segmenting comprises comparing an ASCN state of adjacent target regions and defining one or more segments where each segment comprises two or more adjacent target regions having similar ASCN state. The ASCN state for a target region comprises at least the coverage ratio and a BAF value for the target region.

[0008] In at least some embodiments of the first aspect, estimating the expected BAF value for each segment comprises estimating a minor allele number based on the total copy number and the tumor purity.

[0009] In at least some embodiments of the first aspect, calculating the coverage ratio for each target region comprises: normalizing raw coverage data for the target region; selecting aPATENT Client Reference No.: P39388-WO-1 reference based on a configuration; and calculating the coverage ratio based on the normalized raw coverage data and the selected reference.

[0010] In at least some embodiments of the first aspect, normalizing the raw coverage data comprises correcting the raw coverage data for GC content bias.

[0011] In at least some embodiments of the first aspect, correcting the raw coverage data for GC content bias comprises: generating a correction factor for each GC content value using a locally estimated scatterplot smoothing (LOESS) curve; applying the correction factor to the raw coverage data to obtain a GC-corrected coverage value; and comparing the GC-corrected coverage value against the reference to compute the coverage ratio.

[0012] In at least some embodiments of the first aspect, the configuration is selected from one of: ASCN calling in a target enrichment (TE) assay; ASCN calling in whole genome sequencing (WGS) assay; or copy number variation (CNV) calling in WGS assay.

[0013] In at least some embodiments of the first aspect, the reference is selected from one of: a panel of normal samples; a matched normal sample; or the sample that acts as self- normalizing.

[0014] In at least some embodiments of the first aspect, the filtering comprises filtering for heterozygous SNPs, and one or more of the following: filtering for non-phased variant calls; filtering for single nucleotide variants; filtering for SNPs that do not overlap with a deletion; and / or filtering for SNPs that are not in regions with high variant density.

[0015] In at least some embodiments of the first aspect, the expected coverage ratio for a segment is calculated according to the following formula:where is tumor purity, is tumor ploidy, and is the total copy number. In some embodiments, the coverage ratio and expected coverage ratio are scaled via a base two logarithm.

[0016] In at least some embodiments of the first aspect, the expected BAF for a segment is calculated according to the following formula:PATENT Client Reference No.: P39388-WO-1where is tumor purity, is minor allele number, and is the total copy number.

[0017] In at least some embodiments of the first aspect, the method further comprises performing a two-dimensional grid search to estimate the tumor purity and the tumor ploidy for the sample. The grid search comprising: generating a plurality of pairs of candidate values for the tumor purity and the tumor ploidy; calculating the expected coverage ratio and the expected BAF for each pair of candidate values for each segment in the plurality of segments; comparing the observed coverage ratio to the expected coverage ratio in each segment and comparing the observed BAF to the expected BAF in each segment; and selecting the pair of candidate values that minimizes the differences between the observed coverage ratios and the expected coverage ratios and between the observed BAF and the expected BAF across all segments in the plurality of segments.

[0018] In at least some embodiments of the first aspect, the method further comprises identifying off-target genomic regions in a TE assay and processing the sequencing data in the off-target genomic regions to improve the estimates for tumor purity and tumor ploidy in the on-target regions.

[0019] In at least some embodiments of the first aspect, the sequencing data is generated by one of: Sanger sequencing, Next-Generation sequencing (NGS), nanopore sequencing, ion torrent sequencing, or single molecule real-time sequencing (SMRT).

[0020] In at least some embodiments of the first aspect, the nanopore sequencing comprises sequencing-by-expansion (SBX) chemistry.

[0021] In at least some embodiments of the first aspect, segmenting the target regions comprises applying a Hidden Markov Model (HMM) to the coverage ratios and BAF values to infer the ASCN state of each segment.

[0022] In accordance with a second aspect of the present disclosure, a system is provided for ASCN calling. The system includes: a sequencing module configured to identify a plurality of target regions of interest in sequencing data derived from a sample; a data processing module configured to calculate a coverage ratio for each target region of thePATENT Client Reference No.: P39388-WO-1 plurality of the target regions; a filtering module configured to filter variant call information; a segmentation module configured to define a plurality of segments based at least in part on the coverage ratio for each target region and at least one B-allele frequency (BAF) value for one or more filtered variants in the variant call information within the target region; and an estimation module configured to estimate an allele-specific copy number (ASCN) state for a segment of the plurality of segments based at least in part on a total copy number, a tumor purity, and a tumor ploidy for the segment / The tumor purity and tumor ploidy are estimated using the coverage ratio, an expected coverage ratio, the BAF value, and an expected BAF value for each segment.

[0023] In at least some embodiments of the second aspect, the data processing module generates a corrected coverage value to correct for GC bias using a LOESS curve.

[0024] In at least some embodiments of the second aspect, the data processing module denoises the sequencing data using singular value decomposition.

[0025] In at least some embodiments of the second aspect, the system further includes a display device for visualization of the estimated ASCN state for one or more segments.

[0026] In accordance with a third aspect of the present disclosure, a non-transitory computer-readable medium is provided for performing ASCN calling. The computer-readable medium stores instructions that, upon execution by one or more processors, cause a computing device to: receive sequencing data comprising a plurality of target regions of interest derived from a sample; calculate a coverage ratio for each target region of the plurality of target regions based on the sequencing data; receive variant call information derived from the sequencing data; filter the variant call information according to one or more filter criteria; segment the target regions to define a plurality of segments based at least in part on the coverage ratio for each target region and at least one B-allele frequency (BAF) value for one or more filtered variants in the variant call information within the target region; and estimate an allele-specific copy number (ASCN) state for a segment of the plurality of segments based at least in part on a total copy number, a tumor purity, and a tumor ploidy for the segment. The tumor purity and tumor ploidy are estimated using the coverage ratio, an expected coverage ratio, the BAF value, and an expected BAF value for each segment.PATENT Client Reference No.: P39388-WO-1 BRIEF DESCRIPTION OF THE DRAWINGS

[0027] FIG.1 sets forth an illustrative system including a sequencing device communicatively coupled to a computing system, in accordance with at least some embodiments of the present disclosure.

[0028] FIG.2 is a flowchart of a method for performing ASCN calling based on sequencing data, in accordance with at least some embodiments of the present disclosure.

[0029] FIG.3 is a block diagram of a system that is configured to implement at least a portion of the method of FIG.2, in accordance with at least some embodiments of the present disclosure.

[0030] FIG.4 is a flowchart of a method for performing a 2D grid search, in accordance with at least some embodiments of the present disclosure.

[0031] FIG.5 is a block diagram illustrating a computing system consistent with implementations of the current subject matter

[0032] FIG.6A is a graph illustrating tumor purity correlation between the presently described method and PureCN, in accordance with at least some embodiments of the present disclosure.

[0033] FIG.6B is a graph illustrating tumor ploidy correlation between the presently described method and PureCN, in accordance with at least some embodiments of the present disclosure.

[0034] FIG.7 is a graph illustrating gene-level ASCN calling concordance between the presently described method and PureCN, for samples > 25% purity, in accordance with at least some embodiments of the present disclosure.

[0035] FIG.8 illustrates a conceptual flow chart of a unified copy number caller, in accordance with at least some embodiments of the present disclosure.

[0036] FIGS.9A & 9B illustrate an HMM for performing segmentation, in accordance with at least some embodiments of the present disclosure.PATENT Client Reference No.: P39388-WO-1 DETAILED DESCRIPTION

[0037] The present disclosure pertains to methods and systems for allele-specific copy number (ASCN) calling using sequencing data, designed to address technical challenges inherent in genomic analysis. One advantage of the techniques disclosed herein is that ASCN calling can support various configurations based on different sequencing assays and disease contexts. In general, sequencing assays provide information about a nucleic acid sequence of interest, and a sequencing assay includes experimental workflows or protocols designed to provide information related to a nucleic acid sequence using sequencing technologies. The workflows or protocols include sample processing and wet chemistry assay steps required to prepare the sample and conduct the diagnostics assay, as well as bioinformatic steps that analyze the data resulting from the assay.

[0038] The following descriptions and examples illustrate embodiments of the present disclosure in detail. Although the present disclosure has been described in some details by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications can be practiced within the scope of the appended claims.

[0039] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0040] Although various features of the disclosure can be described in the context of a single embodiment, the features can also be provided separately or in any suitable combination. Conversely, although the present disclosure can be described herein in the context of separate embodiments for clarity, the present disclosure can also be implemented in a single embodiment. It is to be understood that the present disclosure is not limited to the particular embodiments described herein and as such can vary. Those of skill in the art will recognize that there are variations and modifications of the present disclosure, which are encompassed within its scope. Definitions

[0041] All terms are intended to be understood as they would be understood by a person skilled in the art. Unless defined otherwise, all technical and scientific terms used herein havePATENT Client Reference No.: P39388-WO-1 the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains.

[0042] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0043] An “allele” refers to a variant of the sequence of nucleotides at a particular location, or locus, on a DNA molecule. Alleles can differ at a single position through single nucleotide polymorphisms (“SNPs”), but an allele can also include a variant of a nucleic acid sequence resulting from an insertion or deletion of up to several thousand base pairs relative to the target nucleic acid sequence. In genetics, an allele refers to one of two or more versions of a DNA sequence (a single base or a segment of bases) at a particular region on a chromosome. In this example, an individual inherits two alleles for each gene, one from each parent, for any given genomic location where such variation exists. If the two alleles are the same, the individual is homozygous for that allele. If the alleles are different, the individual is heterozygous for that allele.

[0044] Copy number variation (CNV) is a phenomenon in which sections of the genome are repeated and / or removed. Copy number variation is a type of structural variation in which a duplication or deletion event can affect one or more base pairs. Somatic copy number alterations (SCNAs) are a type of copy number alteration (CNA) that occur in somatic tissue, such as a tumor, after conception. SCNAs are common in cancer and involve alterations in a large portion of the cancer genome. As used herein, identifying SCNAs may be referred to as ASCN calling in the cancer context, and identifying CNVs may be referred to as ASCN calling in the non-cancer context.

[0045] A “nucleic acid” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. The term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, or non-naturally occurring, which have similar bindingPATENT Client Reference No.: P39388-WO-1 properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation, phosphorothioates, phosphoramidites, methyl phosphonates, chiral-methyl phosphonates, 2- O-methyl ribonucleotides, and peptide-nucleic acids (PNAs). Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (see, e.g., Batzer et al., “Enhanced evolutionary PCR using Oligonucleotides with inosine at the 3’-terminus,” Nucleic Acid Res.19:5081 (1991); Ohtsuka et al., “An alternative approach to deoxyoligonucleotides as hybridization probes by insertion of deoxyinosine at ambiguous codon positions,” J. Biol. Chem.260:2605-2608 (1985); Rossolini et al., “Use of deoxyinosine – containing primers vs degenerate primers for polymerase chain reaction based on ambiguous sequence information,” Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid can be used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.

[0046] The term “nucleotide,” in addition to referring to the naturally occurring ribonucleotide or deoxyribonucleotide monomers, can be understood to refer to related structural variants thereof, including derivatives and analogs, that are functionally equivalent with respect to the particular context in which the nucleotide is being used (e.g., hybridization to a complementary base), unless the context clearly indicates otherwise.

[0047] In a specific embodiment, the methods described herein can be used to evaluate data generated by nanopore-based sequencing methods, such as sequencing-by-synthesis (SBS) or sequencing-by-expansion (SBX) chemistries. Nanopore-based sequencing-by- synthesis (“SBS”) uses a polymerase (or other strand-extending enzyme) covalently linked to a nanopore to synthesize a DNA strand complementary to a target sequence template (i.e., a copy strand). The nanopore embedded in a membrane in an electrochemical cell is used to concurrently detect the identity of each nucleotide monomer as it is added to that growing strand. (see, e.g., US Pat. Publ. Nos.2013 / 0244340 A1, 2013 / 0264207 A1, 2014 / 0134616 A1, 2015 / 0368710 A1, and 2018 / 0057870 A1, and published International Application WOPATENT Client Reference No.: P39388-WO-1 2019 / 166457 A1). Each added nucleotide monomer is detected by monitoring signals due to changes in ion flow through the nanopore as a tag moiety attached to each added nucleotide monomer enters the nanopore and alters the ion flow. For optimal performance, the tag moiety should reside in the nanopore for a sufficient amount of time to provide for a detectable, identifiable, and reproducible signal associated with altering ion flow through the nanopore (relative to the baseline “open current” flow), such that the specific nucleotide associated with the tag can be distinguished unambiguously from the other tagged nucleotides in the SBS solution.

[0048] In the SBS technique, a template nucleic acid to be sequenced (e.g., a nucleic acid molecule or another analyte of interest) and a primer can be introduced into the sample chamber of the nanopore cell. As examples, templates can be circular or linear. A nucleic acid primer can be hybridized to a portion of template to which four differently polymer- tagged nucleotides can be added. In some embodiments, an enzyme (e.g., a polymerase, such as a DNA polymerase) is associated with the nanopore for use in the synthesis of a complementary strand derived from the template. The polymerase can catalyze the incorporation of nucleotides onto the primer using a single stranded nucleic acid molecule as the template. The nucleotides can comprise tag species (“tags”) with the nucleotide being one of four different types: A, T, G, or C. When a tagged nucleotide is correctly complexed with polymerase, the tag can be pulled (e.g., loaded) into the nanopore by an electrical force, such as a force generated in the presence of an electric field generated by a voltage applied across a lipid bilayer membrane and / or the nanopore disposed therein. The tail of the tag can be positioned in the barrel of the nanopore. The tag held in the barrel of the nanopore can generate a unique ionic blockade signal due to the tag’s distinct chemical structure and / or size, thereby electronically identifying the added base to which the tag is attached.

[0049] Alternatively, the methods described herein may be used to evaluate sequencing data generated by SBX-based techniques, a nanopore-based nucleic acid sequencing method that uses a biochemical process to transcribe the sequence of DNA onto a measurable polymer molecule referred to as an Xpandomer (which may sometimes be referred to as an “XP molecule” or “XP”). (see, e.g., U.S. Pat. No.7,939,259, and International Application Publication No. WO2020236526A1) In the SBX-based sequencing process, a target nucleic acid sequence is encoded along a backbone XP sequence with reporter constructs that arePATENT Client Reference No.: P39388-WO-1 separated by, e.g., ~10 nm that are designed to provide high signal-to-noise ratio (SNR), well-differentiated response signals during nanopore translocation. The enhanced SNR provided by the different response signals provides significantly increased sequence read efficiency and accuracy of XPs relative to native nucleic acid molecules.

[0050] SBX chemistry sequences nucleic acids by creating a derivative Xpandomer molecule from a target nucleic acid template. This is achieved by encoding the nucleic acid information on a surrogate polymer of extended length which is easier to detect during nanopore-based sequencing. The surrogate polymer, i.e., XP, is formed by template directed synthesis, which preserves the original genetic information of the target nucleic acid while also increasing linear separation of the individual elements of the sequence data.

[0051] Target enrichment is a pre-sequencing DNA preparation step that isolates and separates specific regions of the genome for analysis by next generation sequencing (NGS). One example of this method includes whole exome sequencing, which uses target enrichment techniques to capture the exomes of interest before sequencing. Whole exome sequencing can be used to determine the nucleotide sequence primarily of the exonic (or protein-coding) regions of an individual’s genome and related sequences, representing approximately 1% of the complete DNA sequence. In contrast, whole genome sequencing can be used to evaluate nearly all of the approximately 3 billion nucleotides of an individual’s complete DNA sequence, including non-coding sequences.

[0052] It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features of the present disclosure, which are, for brevity, described in the context of a single embodiment, can also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the present disclosure are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.PATENT Client Reference No.: P39388-WO-1

[0053] It is intended that every maximum numerical limitation given throughout this specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification will include every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every numerical range given throughout this specification will include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein. Systems and Devices for Allele-Specific Copy Number Calling

[0054] The present disclosure pertains to systems and methods for allele-specific copy number (ASCN) calling, particularly as applied using computer-implemented techniques for processing and / or analyzing sequencing data. ASCN calling refers to methods used to determine specific alleles in different genomic regions. The presently described methods involve preprocessing steps and a statistical framework to predict an ASCN for one or more genomic segments, the ASCN caller supporting various configurations based on the sequencing assay and disease context.

[0055] FIG.1 sets forth an illustrative system 100 including a sequencing device 110 (which may be alternately referred to as a sequencer instrument) communicatively coupled to a computing system 102. Sequencing device 110 can be coupled to computing system 102 either directly (e.g., through one or more communication cables) or through network 130, which may be the Internet or any other combination of wide-area, local area, wired, and / or wireless networks. In some embodiments, computing system 102 may be included in or integrated with the sequencing device 110. In some embodiments, sequencing device 110 may sequence (e.g., perform a biochemical assay) a sample containing genetic material and produce resulting sequencing data. The sequencing data can be sent to computing system 102 (e.g., through network 130) or stored on a storage device and at a later stage transferred to computing system 102 (e.g., through network 130). In some embodiments, computing system 102 may or may not include a display 108 and one or more input devices (not illustrated) for receiving commands from a user or operator (e.g. a technician or a geneticist). In some embodiments, computing system 102 and / or sequencing device 110 can be accessed by users or other devices remotely through network 130. Thus, in some embodiments various methods discussed herein may be run remotely on computing system 102.PATENT Client Reference No.: P39388-WO-1

[0056] Computing system 102 may include one computing device or a combination of a number of computing devices of any type, such as personal computers, laptops, network servers (e.g., local servers or servers included on a public / private / hybrid cloud), mobile devices, etc., where some or all of the devices can be interconnected. Computing system 102 may include one or more processors (not illustrated), each of which can have one or more logic cores. In some embodiments, computing system 102 can include one or more general- purpose processors (e.g., CPUs), special-purpose processors such as graphics processors (GPUs), digital signal processors, or any combination of these and other types of processors. In some embodiments, some or all processors in computing system can be implemented using customized or customizable circuitry, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Computing system 102 can also in some embodiments retrieve and execute non-transitory computer-readable instructions stored in one or more memories or storage devices (not illustrated) integrated into or otherwise communicatively coupled to computing system 102. The memory / storage devices can include any combination of non-transitory computer readable storage media including semiconductor memory chips of various types (e.g., dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), flash memory, programmable read-only memory, etc.) and so on. Magnetic and / or optical disks can also be used. The memories / storage devices can also include removable storage media that can be readable and / or writeable; examples of such media include compact disc (CD), read-only digital versatile disc (e.g., DVD-ROM, dual-layer DVD-ROM), read-only and recordable Blu-ray® disks, ultra-density optical disks, flash memory cards (e.g., SD cards, mini-SD cards, micro-SD cards, etc.), and so on. In some embodiments, data and other information (e.g. sequencing data) can be stored in one or more remote locations, e.g., cloud storage, and synchronized with the other components of system 100.

[0057] In some embodiments, the sequencing device 110 can generate sequencing data by a sequencing by expansion (SBX) process. Examples of the SBX process include those described in U.S. Patent Application No.17 / 456,342 (U.S. Publication No. US20220411458A1), entitled “Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing,” filed November 23, 2021, which is herein incorporated by reference in its entirety. During library preparation in thePATENT Client Reference No.: P39388-WO-1 SBX process, a number of surrogate molecules are derived from and characterize nucleic acid material provided in a sample.

[0058] More particularly, the SBX process may translate a sequence of DNA into a measurable surrogate molecule called an Xpandomer. Xpandomer synthesis based on the natural function of DNA replication uses expandable nucleotide triphosphates (X-NTPs) that act as substrates for template-dependent, polymerase-based replication. These Xpandomer molecules are then processed by a sequencer instrument (e.g., the sequencing device 110) to measure the sequence in the original DNA template. As the Xpandomer molecule transits through nanometer-sized openings (e.g., a nanopore) in an electrode-resistant membrane, each nanopore corresponding to a selective channel, a distinct electrical signal is generated for each base reporter and identifiable to enable highly accurate and high throughput nanopore-based nucleic acid sequencing (also referred to generally as nanopore sequencing).

[0059] Sequencing device 110 can generate a plurality of sequence reads corresponding to a genetic sample (e.g., a sample comprising a patient’s DNA or RNA material). For example, a sequence may be identified by processing a sample, which may include (for example) a blood, saliva, or tissue biopsy collected from a subject. Sequence reads can be obtained either directly from sequencing device 110, or from one or more local or remote volatile or non-volatile memories, storage devices, or databases communicatively coupled to computing system 102. Sequence reads can be pre-processed (e.g., pre-aligned) or they can be “raw,” in which case a downstream method may include a preprocessing (e.g., pre- aligning) step. Also, while in some embodiments, entire sequence reads (as generated by sequencing device 110) can be obtained, in other embodiments only sections of the sequence reads can be obtained. Thus, “obtaining a sequence read,” as used herein, refers generally to obtaining one or more sections of one or more (e.g., adjacent) sequence reads. Method for performing ASCN calling based on Sequencing Data

[0060] FIG.2 illustrates a flow chart of a method 200 for performing ASCN calling based on sequencing data, in accordance with some embodiments. The method 200 may be performed, at least in part, by one or more processors (e.g., of computing system 102). The one or more processors may execute computer program code (e.g., a set of instructions) that cause the processor(s) to perform one or more steps of the method 200. The instructions mayPATENT Client Reference No.: P39388-WO-1 process one or more input files and generate one or more output files, which may be stored in one or more computer-readable storage media (e.g., DRAM, HDD, SDD, etc.). In some embodiments, the method 200 may be implemented by one or more logic circuits, e.g., as implemented in an ASIC or FPGA. It will be appreciated that the method 200 may be performed in any combination of hardware, firmware, or software, as executed by a general purpose processor.

[0061] The method 200 includes, at 202, identifying a plurality of target regions of interest in sequencing data derived from a sample. The sequencing data may be derived using an assay consisting of whole genome sequencing, whole exome sequencing, targeted enrichment (TE) sequencing, RNA sequencing, or metagenomic sequencing. These target regions may be dynamically adjusted for optimal sensitivity and accuracy.

[0062] At 204, the method 200 further includes calculating a coverage ratio for each target region. Calculating a coverage ratio involves normalizing raw coverage data for the target region, selecting a reference based on the configuration, and calculating the coverage ratio based on the normalized raw coverage data and the selected reference. The reference may vary based on the configuration, such as a panel of normal samples for TE assays or a matched normal sample for WGS assays. In some embodiments, this operation further comprises normalizing the raw coverage data for GC content bias.

[0063] At 206, the method 200 further includes filtering input variant calls to handle SBX-specific noise. This involves applying various filters to variant information provided by a variant caller to ensure the accuracy of the B-allele frequency (BAF) values required for ASCN calling. The variant filters may include filtering for non-phased variant calls, single nucleotide variants with a single reference and a single alternative base, SNPs with allele frequency greater than a threshold and / or lower than a second threshold (e.g., within a range of allele frequencies as indicated in a population database), heterozygous SNPs, SNPs that do not overlap with deletions, and SNPs that are not in regions with high variant density. In tumor-only data, the variant calls can be filtered for SNPs with 0.03 < BAF < 0.97. In matched tumor-normal data, the variant calls can be filtered for SNPs having greater than 30X coverage and 0.4 < BAF < 0.6 in the matched normal. The filtering step can help alleviate certain SBX-specific noise inherent in SBX-based sequencing data.PATENT Client Reference No.: P39388-WO-1

[0064] At 208, the method 200 further includes segmenting the sequencing data based at least on the calculated coverage ratios and B-allele frequency (BAF) values to generate a plurality of segments. In an embodiment, the segmentation step includes grouping adjacent target regions based on the similarity of the coverage ratios to generate a plurality of segments. These segments may then be further subdivided into smaller segments based on the similarity of the BAF values within each segment. The BAF values may be based at least in part on the single nucleotide polymorphism (SNP) variants filtered from the variant information. These segments represent genomic regions that harbor the same ASCN state, which will be predicted as discussed below.

[0065] In an embodiment, the segmentation algorithm comprises a recursive algorithm. In the first step, a most divergent sub-segment is differentiated from the rest of the possible sub- segments by comparing the similarity statistics between the signals within the sub-segment to all the signals external to the sub-segment. This similarity statistic can be a t-test or any other statistic that can be used to compare two groups of observations. If the similarity statistic of the most divergent sub-segment is significant, then this step is recursively applied to each of the sub-segments formed from splitting the original segment by breakpoints defined by the most divergent sub-segment. The recursion continues until no new statistically significant sub-segments are found. Finally, to address the issue of over-segmentation, segments with similar coverage ratios and BAF are clustered together using hierarchical clustering. Consecutive segments in the same cluster are merged and assigned with the mean coverage ratio and BAF across all segments being merged.

[0066] At 210, the method 200 further includes estimating an allele-specific copy number (ASCN) state for each segment of the plurality of segments based at least in part on the total copy number, tumor purity, and tumor ploidy. A statistical framework is used to perform the estimation of the ASCN state of each segment. The statistical framework is based on the observation that each segment has associated therewith an observed coverage ratio and one or more BAF values. These values are generated from summarizing the coverage ratios of each target region and the BAF values for the filtered SNPs that overlap with a segment.

[0067] The expected coverage ratio for a segment may be calculated as follows:PATENT Client Reference No.: P39388-WO-1 where is the tumor purity, is the tumor ploidy, C is the total copy number. As usedherein, refers to the set of parameterized values ( , , C). The difference between theobserved coverage ratio, , and the expected coverage ratio, can be measured for different parameterizations of . The set of values of that results in an with the least difference from , over all segments, can be considered the optimal set of values of to use. In some embodiments, the coverage ratio may be calculated as shown above, but on a logarithmic scale. In other words, the calculated coverage ratio is given as Equation 1 taken to the base 2 logarithm.

[0068] By modeling the coverage ratios as a probability distribution, the difference between the observed coverage ratio and the expected coverage ratio can be statistically quantified according to the corresponding likelihood function of the probability distribution, as follows: , (Eq.2) where refers to the set of parameterized values that pertains to segment . Assuming known values for and , a coverage ratio likelihood can be calculated for each parameterization of C, as follows: , (Eq.3) where . The value of that gives the maximum likelihood is the optimal C value that best fits the observed coverage ratio.

[0069] A similar concept can then be applied to the BAF values. The expected BAF value of a segment may be calculated as follows: (Eq.4)where is the tumor purity, is the minor allele number, and is the total copy number (which is known from calculations performed for the coverage ratios described above. Similar to the coverage ratio data, the difference between the expected BAF value and the observed BAF value can be statistically quantified using a likelihood function of the probability distribution:PATENT Client Reference No.: P39388-WO-1 , (Eq.5) Assuming known values for and C, a BAF value likelihood can be calculated for each parameterization of , as follows: , (Eq.6) . The highest BAF value likelihood represents the value of the minor allele number that fits the observed data. When the configuration is in a CNV context (i.e., non-cancer), then we can assume the following values for tumor purity and tumor ploidy: equal to 1 and equal to 2. Such an assumption can make the estimation of the ASCN state as simple as calculating the above likelihood values for the different possible values of total copy number and minor allele number for each segment. However, in an SCNA context (e.g., tumor-only sample or matched tumor-normal sample), then the process above involves joint estimation of the values of .

[0070] In an embodiment, joint estimation of the values of comprises conducting a 2D grid search over all possible values of and , which generates a plurality of pairs of candidate values for the tumor purity and the tumor ploidy, and selects a pair of the candidate values for the tumor purity and the tumor ploidy that minimizes the difference between the observed coverage ratio and the expected coverage ratio and / or between the observed BAF value and the expected BAF value, across all segments of the sample. The selected tumor ploidy and tumor purity values are then used to predict the ASCN state of each segment, in accordance with the statistical methods set forth above.

[0071] In some embodiments, the method 200 supports various configurations based on different sequencing assays and contexts, including dynamic interval sizing and mappability file selection based on SBX data characteristics. The flexibility in defining target regions, dynamic interval sizing, and mappability file selection distinguish the approach described herein from traditional methods. The process may also identify off-target genomic regions in the TE assay and process the off-target data for the ASCN calling, applying bias correction for the off-target data by identifying technical biases unique to the off-target capture process and applying a denoising algorithm tuned for the off-target genomic regions to correct the bias. The corrected and denoised off-target coverage ratios may be used for segmentation,PATENT Client Reference No.: P39388-WO-1 integrating segmented off-target regions with on-target regions to form comprehensive genomic segments for the ASCN calling. This integration extends copy number analysis across a broader portion of the genome, enhancing detection of copy number variations (CNVs) by providing additional data points from the off-target regions. The off-target data may also be used to estimate the tumor purity and tumor ploidy alongside the on-target data, integrating the off-target data from multiple sequencing runs to increase overall coverage depth.

[0072] In some embodiments, the sequencing data is generated by Sanger sequencing, Next-Generation sequencing (NGS), nanopore sequencing, sequencing by synthesis (SBS), ion torrent sequencing, or single molecule real-time sequencing (SMRT). In an exemplary embodiment, the nanopore sequencing may be associated with an SBX chemistry.

[0073] FIG.3 is a block diagram of a system 300 that is configured to implement at least a portion of the method 200, in accordance with at least some embodiments. The system 300 for ASCN calling may comprise a sequencing module 302 configured to identify a plurality of target regions of interest in sequencing data, a data processing module 304 configured to calculate a coverage ratio for a target region of the plurality of target regions and to denoise the sequencing data, a segmentation module 306 configured to group genomic regions based at least on the coverage ratio and BAF value, and an estimation module 310 configured to estimate allele-specific copy numbers, total copy numbers, tumor purity, and tumor ploidy, wherein the tumor purity and the tumor ploidy are estimated using the coverage ratio, expected coverage ratio, the BAF value, and an expected BAF value. The system 300 may also include a filter module 306 for applying one or more filters to the called variants. The data processing module 304 may correct GC bias using a LOESS regression model and denoise the sequencing data using singular value decomposition. The system 300 may also include an automated workflow to streamline sequencing, data processing, segmentation, and ASCN estimation, with the data processing module 304 being cloud-based and a user interface for visualization.

[0074] The system 300 for ASCN calling may comprise at least one programmable processor and a non-transient computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising identifying a plurality of target regions of interest in sequencing data, calculating a coverage ratio for aPATENT Client Reference No.: P39388-WO-1 target region of the plurality of target regions, segmenting the sequencing data based at least on the coverage ratio and a B-allele frequency (BAF) value, and estimating an ASCN state for each segment based at least in part on total copy number, tumor purity, and tumor ploidy, wherein the total copy number is estimated based at least in part on the calculated coverage ratio, and wherein the tumor purity and tumor ploidy are estimated using the coverage ratios, expected coverage ratios, the BAF values, and expected BAF values. The operations may further comprise estimating a minor allele copy number based at least in part on the total copy number and the tumor purity. Calculating a coverage ratio for a target region of the plurality of target regions may comprise correcting raw coverage values for GC content bias, generating a normalizing factor for each GC content value using a locally estimated scatterplot smoothing (LOESS) curve to correct the raw coverage values, and comparing the corrected raw coverage values against a reference to compute the coverage ratios. I. Identifying Target Regions

[0075] First, target regions of interest can be identified by the sequencing module 302 to generate a target map 312. This step can involve determining regions of the genome that can be utilized for estimating allele-specific copy numbers. For target enrichment (TE) assays, specific areas of the genome can be captured using probes, which can then be sequenced. In some embodiments, this capture process can be imperfect, resulting in the presence of off- target DNA within the sequenced library. While off-target DNA is not suitable for variant calling, it can still be processed, albeit with different handling compared to on-target regions. In contrast, for whole genome sequencing (WGS) assays, the concept of on-target and off- target regions does not apply. Instead, all genomic regions can be considered equally without distinction. This broader approach in WGS can simplify the identification of target regions as it can eliminate the need to account for capture biases inherent in TE assays. Thus, the identification step in WGS can focus on comprehensive genomic coverage, ensuring that all regions can be included in the subsequent analysis for ASCN estimation.

[0076] For WGS data, the entire genome can be regarded as the target region of interest. Consequently, all targets can be considered on-target, rendering the concept of off-target regions inapplicable. To establish on-target regions for WGS, non-overlapping intervals can be tiled across the genome. Unlike probe designs with fixed lengths, the size of these intervals can be dynamically adjusted. This dynamic adjustment can be advantageous in twoPATENT Client Reference No.: P39388-WO-1 primary ways. Firstly, certain genomic loci can be known to harbor small SCNA events. By reducing the target size in these regions, the sensitivity in detecting these subtle SCNA events can be enhanced. Secondly, SBX-based sequencing data can exhibit lower uniformity of coverage compared to other types of sequencing data based on other chemistries (e.g., sequencing data from Illumina sequencers). In areas where coverage is less uniform, increasing the target size can help average out the noise, thereby improving data quality.

[0077] Another critical aspect in the creation of target regions can be the selection of regions with high mappability, which can determine how uniquely mappable a locus is. Illumina-based ASCN callers assign mappability files to each region using an externally provided file tailored to a specific read length unique to Illumina-based sequencing systems. This approach is compatible with sequencing data generated by Illumina sequencers, given their consistent read lengths. However, due to the lower and variable insert sizes of SBX- based sequencing data, the ASCN caller tailored for Illumina sequencers can result in suboptimal ASCN calling performance when applied to the SBX-based sequencing data. In the presently described methods, the ASCN caller can dynamically determine the median read length of the SBX data and can select the most appropriate mappability file for the input data.

[0078] The flexibility in defining target regions, along with dynamic interval sizing and tailored mappability file selection, distinguishes the presently described methods from Illumina-based ASCN callers. II. Calculating Coverage Ratio

[0079] Second, coverage ratios can be calculated by the data processing module 304 to generate target coverage data 314. The data processing module 304 determines, using sequencing data 301, which may be provided in FASTQ files or BAM files (if the sequencing reads have been aligned to a reference sequence), the raw coverage for each target region and corrects the raw coverage for GC content using, for example, a locally estimated scatterplot smoothing (LOESS) curve. These corrected values can be compared against a reference 320 to create coverage ratios, with an additional denoising step applied for TE assays to account for technical biases. The reference 320 can be derived from a panel of normal samples 322PATENT Client Reference No.: P39388-WO-1 from a group of cohorts, a matched normal sample 324 from the same patient, or a normalized sample 326.

[0080] The term “coverage” refers to the number of times a particular segment of the genome is sequenced, which can provide a measure of the abundance of that segment in the sample. The coverage ratio is an important metric in the ASCN calling process that can quantify the relative coverage of genomic regions in a sample compared to a reference. Calculating the coverage ratio is important in ASCN calling because it provides a means to normalize the amount of DNA sequenced from a specific genomic region relative to a reference.

[0081] This normalization allows for the detection of copy number variations by comparing the observed amount of DNA in a sample to what is expected under normal conditions. Without this comparison, it would be challenging to determine whether a region of the genome has been duplicated, deleted, or remains unchanged, as raw sequencing data can vary significantly due to technical factors. The coverage ratio also helps to correct for biases introduced by the sequencing process itself. Factors such as GC content can cause certain regions of the genome to be over-represented or under-represented in sequencing data. By calculating the coverage ratio and applying appropriate corrections, these technical biases can be mitigated. Moreover, the coverage ratio can be important for distinguishing between different genomic states, especially in complex samples like those from cancer tissues. Tumor samples often contain a mixture of normal and cancerous cells, making it difficult to discern true genomic alterations without a normalized measure. The coverage ratio can enable the separation of signals arising from normal cells and those from cancer cells, facilitating a more precise analysis of tumor-specific changes. Using coverage ratios also allows for the comparison of sequencing data across different samples and experiments. By normalizing the data, it becomes possible to identify genuine biological differences rather than artifacts introduced by varying sequencing conditions.

[0082] The coverage ratio can be calculated by first determining a raw coverage value for each target region of interest identified in the target region map 312. This raw coverage value can then be corrected for variations in GC content, which can affect sequencing efficiency and lead to biases. The corrected coverage values can provide a more accurate representation of the actual genomic content.PATENT Client Reference No.: P39388-WO-1

[0083] GC content, which refers to the proportion of guanine (G) and cytosine (C) bases in a DNA sequence, can significantly influence sequencing efficiency. Regions with extreme GC content, either very high or very low, tend to be underrepresented or overrepresented in sequencing data due to the biases introduced during the amplification and sequencing processes. To correct for GC bias, the raw coverage data must first be analyzed to determine how GC content can affect coverage. This can be done by plotting the raw coverage against the GC content for each target region. In some embodiments, a LOESS (locally estimated scatterplot smoothing) curve can then be fitted to this data. The LOESS curve is a non- parametric method that can create a smooth line through the data points, capturing the general trend of how coverage can vary with GC content. This step is crucial to deal with the stronger GC bias in SBX data.

[0084] The fitted LOESS curve can provide a normalization factor for each GC content value, which can be used to adjust the raw coverage data. By applying these normalization factors, the coverage values can be corrected for GC bias, effectively leveling the playing field across regions with different GC contents. This correction can ensure that the coverage can more accurately reflect the true genomic content, reducing the influence of sequencing biases.

[0085] After GC bias correction, the coverage data can become more reliable and can be compared against a reference 320 to calculate the coverage ratio. Once the coverage has been corrected, the coverage ratio can be obtained by comparing the sample’s coverage to that of a reference 320. The choice of reference varies depending on the configuration, such as ASCN calling in TE assays, ASCN calling in WGS assays, or CNV calling in WGS assays. In some embodiments, the reference 320 can be a panel of normal samples 322 for ASCN calling in TE assays. In some embodiments, the reference 320 can be a matched normal sample 324 for ASCN calling in WGS assays. In some embodiments, the reference 320 can be the sample 326 itself using (normalized) regions known to be copy number stable for CNV calling in WGS assays. The coverage ratio can be calculated by dividing the corrected coverage value of the sample by the corresponding coverage of the reference 320. This ratio can reflect the relative abundance of DNA in the target regions of the sample 301 compared to the reference 320, allowing for the detection of copy number variations.PATENT Client Reference No.: P39388-WO-1

[0086] The presently described methods for ASCN calling can support various configurations, which are determined by the combination of the sequencing assay and the disease context. These configurations can influence the specific steps and approaches used in the process, ensuring that the method can be adaptable to different experimental setups and clinical scenarios. In some embodiments, the configuration can be ASCN calling in TE assays. In some embodiments, the configuration can be ASCN calling in WGS assays. In some embodiments, the configuration can be CNV calling in WGS assays. Each of these exemplary, non-limiting configurations are described in detail infra.

[0087] Accordingly, in any of the computer-implemented methods described herein, calculating the coverage ratio can further include normalizing the raw coverage data for GC content bias. In some embodiments, the methods can further include: generating a correction factor for each GC content value using a locally estimated scatterplot smoothing (LOESS) curve; applying the correction factor to the raw coverage data to obtain a GC-corrected coverage value; and comparing the GC-corrected coverage value against the coverage value for the reference to compute the coverage ratio. In some embodiments, the GC-corrected coverage value can be used to correct the GC content bias. In some embodiments, the reference 320 can vary based on the configuration. A. ASCN calling in target enrichment (TE) assays

[0088] In an embodiment, the configuration can be ASCN calling in a TE assay, on a single tumor sample. In this setup, probes can be used to capture specific regions of the genome, known as on-target regions. However, due to the imperfect nature of the capture process, off-target DNA can still be present in the sequenced library. These off-target regions can be accounted for differently from on-target regions and can be processed to ensure accurate ASCN calling.

[0089] In TE assays, an additional denoising step can be necessary to account for technical biases introduced during the enrichment process. These biases can arise from factors such as uneven capture efficiency, variability in sequencing depth, and GC content biases. These biases can distort the true relationship between coverage and copy number. In some embodiments, singular value decomposition (SVD) can be employed to identify and to correct these biases, particularly in target enrichment assays.PATENT Client Reference No.: P39388-WO-1

[0090] The SVD process can begin by organizing the coverage data from the panel of normal samples into a matrix. Each column of the matrix can represent the coverage values of a normal sample in the panel across various genomic targets. SVD can then be performed on this matrix, decomposing it into three matrices: , , VT. For any given matrix A, A = VT, where U is an orthogonal matrix, representing the left singular vectors of A, is a diagonal matrix containing the singular values of A, and VTis another orthogonal matrix, representing the right singular vectors of A. By examining the singular values in , the significant components that capture the systematic technical noise can be identified. The less significant components, which are more likely to represent random noise rather than systematic bias, can be discarded. This step effectively can isolate the main patterns of technical noise present in the normal samples. The denoising process can then be applied to the tumor sample data 301. The systematic noise patterns identified from the panel of normal 322 can be subtracted from the tumor sample’s coverage data. This correction can help to normalize the data, reducing the impact of technical artifacts.

[0091] On-target refers to the specific regions of DNA intended to be sequenced in a TE assay. Off-target regions refer to genomic areas that are not specifically targeted by the probes but still end up being sequenced due to the imperfect nature of the capture process. These off-target regions, while not intended to be captured, can still provide valuable information for ASCN calling if processed correctly.

[0092] The first step in accounting for off-target regions can involve identifying these regions within the sequenced data. Although the coverage in off-target regions can generally be lower than in on-target regions, it can still be sufficient for ASCN calling. The off-target reads can be distinguished from the on-target reads by mapping the sequenced reads back to the reference genome and determining which regions are intended targets based on the probe design.

[0093] Once the off-target regions are identified, their coverage can be calculated similarly to the on-target regions. The raw coverage of the off-target regions can be measured and corrected for GC content to mitigate any biases that might distort the true representation of genomic content. In some embodiments, a LOESS curve can be used to normalize the coverage values of the off-target regions. After normalization, the corrected coverage values of off-target regions can be integrated into the overall analysis.PATENT Client Reference No.: P39388-WO-1

[0094] Incorporating off-target regions can increase the overall coverage of the genome. TE assays focus on specific genomic regions, leaving large portions of the genome either sparsely covered or not covered at all. By including off-target regions, even if they are covered at a lower depth than on-target regions, the process can gain additional data points that can fill gaps and provide a more continuous and complete view of the genome. This increased coverage data can particularly be useful for detecting copy number variations (CNVs) that can occur outside the targeted regions.

[0095] In some embodiments, the use of off-target regions can enhance the statistical power of the ASCN calling process. More data points from off-target regions can contribute to a larger dataset, which can improve the reliability and robustness of statistical analyses. This can lead to more accurate segmentation and better estimation of allele-specific copy numbers. The additional data from off-target regions can help in identifying subtle copy number changes that may be missed if only on-target data are used.

[0096] In some embodiments, incorporating off-target regions can also improve the detection of structural variations and complex genomic rearrangements. These events can span large genomic regions and may not be fully captured by on-target data alone. Off-target regions can provide supplementary information that can aid in reconstructing these complex events, leading to a more accurate and detailed understanding of the genomic alterations present in the sample.

[0097] In some embodiments, including off-target regions can help in the normalization and correction processes. Since off-target regions can be randomly distributed across the genome, they can serve as a valuable reference for correcting systematic biases and technical artifacts introduced during the sequencing and enrichment processes. This can improve the overall quality and accuracy of the coverage data, leading to more precise ASCN calling.

[0098] In some embodiments, leveraging off-target regions can maximize the utility of the sequencing data obtained from TE assays. Rather than discarding these regions as noise, the present methods incorporate them into valuable sources of information, increasing the efficiency and cost-effectiveness of the sequencing procedure. The presently described methods extract as much useful data as possible from the sequencing reads, enhancing the overall value of the genomic analysis.PATENT Client Reference No.: P39388-WO-1

[0099] Accordingly, the configuration can be ASCN calling in target enrichment (TE) assay. In some embodiments, the ASCN calling in the TE assay can be a whole exome sequencing. In some embodiments, the sample can be derived from a liquid biopsy or solid tumor tissue biopsy. In some embodiments, the reference 320 can be a panel of normal samples 322. In some embodiments, the calculating a coverage ratio step can further include a denoising step using singular value decomposition (SVD).

[0100] In some embodiments, any of methods of ASCN calling in TE assay as described herein can further include identifying off-target genomic regions in the TE assay and processing the off-target data for the ASCN calling. In some embodiments, bias correction for the off-target data can include: identifying technical biases unique to the off-target capture process; and applying a denoising algorithm tuned for the off-target genomic regions to correct the bias. In some embodiments, the corrected and denoised off-target coverage ratios can be used for segmentation by: grouping adjacent off-target regions based on a similarity of the coverage ratios; and integrating segmented off-target regions with on-target regions to form comprehensive genomic segments for the ASCN calling. In some embodiments, the integrating of the off-target data can extend copy number analysis across a broader portion of the genome. In some embodiments, the off-target data can enhance detection of copy number variations (CNVs) by providing additional data points from the off-target regions. In some embodiments, the off-target data can be used to estimate the tumor purity and tumor ploidy alongside the on-target data. In some embodiments, any of methods of ASCN calling in TE assay as described herein can further include integrating the off-target data from multiple sequencing runs to increase overall coverage depth. B. ASCN calling in WGS assays

[0101] In another embodiment, the configuration can be ASCN calling in whole genome sequencing (WGS) assays, where the entire genome can be sequenced without distinguishing between on-target and off-target regions. This approach can simplify the identification of target regions but can require comprehensive data handling to manage the vast amount of genomic information generated. In some embodiments, GC bias can be corrected by fitting e.g., a LOESS curve to the coverage data. The LOESS curve can normalize the coverage values based on GC content, reducing the bias and ensuring that the coverage data can more accurately reflect the true genomic content.PATENT Client Reference No.: P39388-WO-1

[0102] The corrected coverage values can then be compared against a reference to create coverage ratios. In the context of ASCN calling in WGS assays, a matched normal sample 324 can typically be used as a reference 320. The matched normal sample 324, which can be sequenced under the same conditions as the tumor sample 301, can provide a baseline for comparison, allowing for the identification of regions with altered copy numbers in the tumor sample.

[0103] The comprehensive coverage of WGS assays in ASCN calling can ensure that all genomic regions can be analyzed. This can be important for detecting complex rearrangements and structural variations that can be missed with targeted approaches. Furthermore, the use of a matched normal sample 324 as a reference 320 can provide a baseline for identifying copy number changes. WGS can enable the detection of both large- scale alterations and focal changes, offering a detailed and comprehensive understanding of the genomic alterations in a disease context, such as cancer.

[0104] Accordingly, the configuration can be ASCN calling in a whole genome sequencing (WGS) assay. In some embodiments, the sample can be derived from a liquid biopsy or solid tumor tissue biopsy. In some embodiments, the reference 320 can be a panel of matched normal samples 324. In some embodiments, the coverage ratios can be calculated by dividing the sample’s coverage by the reference’s coverage. In some embodiments, the plurality of target regions of interest can be an entire genome. In some embodiments, on- target regions for WGS can be created by tiling non-overlapping intervals across the entire genome. In some embodiments, the size of the non-overlapping intervals can be adjusted dynamically along the entire genome without being fixed to a certain length. In some embodiments, any of the methods described herein can further include selecting target regions of the plurality of target regions with high mappability. C. CNV calling in WGS assays

[0105] In yet another embodiment, the configuration can be copy number variation (CNV) calling in a non-cancer context, using WGS assays. This configuration can detect variations in the number of copies of genomic regions, which can include deletions, duplications, and more complex rearrangements.PATENT Client Reference No.: P39388-WO-1

[0106] WGS assays generate a large number of reads, covering the entire genome. The reads can be aligned to a reference genome using a software that can map each read to its corresponding location in the reference genome. In some embodiments, the software can be Burrows-Wheeler aligner (BWA). In other embodiments, the software can be Bowtie2; HISAT2 (hierarchical indexing for spliced alignment of transcript 2); STAR (spliced transcripts alignment to a reference); Minimap2; Novoalign; GEM (genome multitool); SNAP (scalable nucleotide alignment program); or the like.

[0107] The coverage can be calculated for each genomic region once the reads are aligned, creating a coverage profile for the entire genome. High coverage can indicate a higher number of reads aligning to that region, while low coverage can indicate fewer reads.

[0108] GC content bias can be corrected using, for example, a LOESS curve by fitting it to the coverage data. The LOESS curve can model the relationship between GC content and observed coverage, providing normalization factors that can adjust the raw coverage values.

[0109] In this configuration, the sample 326 of interest can act as its own reference 320, a process known as self-normalizing. This approach can use a set of genomic regions known to be copy number stable to create a reference 320 for calculating coverage ratios. This self- normalization technique can be particularly useful when matched normal samples 324 or panels of normals 322 are not available. The normalized coverage values can be compared against a reference to identify deviations that indicate CNVs. The comparison can involve calculating coverage ratios, where the normalized coverage of the sample can be divided by the coverage of the reference. This step can identify regions with abnormal copy numbers, such as duplications or deletions. The identified CNVs can be validated using additional methods, such as quantitative PCR, array comparative genomic hybridization (aCGH), or other sequencing technologies. Annotating the CNVs can involve mapping them to, e.g., known genes, regulatory elements, and other functional genomic regions.

[0110] Accordingly, the configuration can be copy number variation (CNV) calling in WGS assay. In some embodiments, the context can be non-cancer. In some embodiments, the reference 320 can be the sample 326 that acts as self-normalizing. In some embodiments, the coverage ratios can be calculated with a set of genomic regions that are copy number stable. In some embodiments, the plurality of target regions of interest can be an entirePATENT Client Reference No.: P39388-WO-1 genome. In some embodiments, on-target regions for WGS can be created by tiling non- overlapping intervals across the entire genome. In some embodiments, the size of the non- overlapping intervals can be adjusted dynamically along the entire genome without being fixed to a certain length. In some embodiments, any of the methods described herein can further include selecting target regions of the plurality of target regions with high mappability. III. Filtering of SBX Variant Calls

[0111] Third, input variant calls 303 (e.g., as provided in a variant call format (VCF) file) are filtered down by a filter module 306 for ASCN calling usage due to the unique error profiles of SBX data. The VCF file 303 can be generated by processing aligned sequence data 301 (e.g., as provided in a BAM file) by an existing variant caller. As set forth above, the filtering can include, but is not limited to, filtering for non-phased variant calls, single nucleotide variants with a single reference and a single alternative base, SNPs with allele frequency greater than a threshold, heterozygous SNPs, SNPs that do not overlap with deletions, and SNPs that are not in regions with high variant density. The filtering step can help alleviate certain SBX-specific noise inherent in SBX-based sequencing data.

[0112] The step of filtering SBX variant calls can involve a series of criteria to ensure the reliability and accuracy of the ASCN calling. Initially, non-phased variant calls can be filtered to retain only those variants with clear phasing information. Single nucleotide variants (SNVs) can then be identified and selected based on the presence of a single alternative base compared to the reference sequence at that loci, ensuring the focus remains on clear and distinct genetic variations. Additionally, variants can be filtered based on allele frequency and heterozygosity. Specifically, single nucleotide polymorphisms (SNPs) with an allele frequency greater than 0.001% from a population genetics database can be selected, and only heterozygous SNPs can be considered to enhance the accuracy of B-allele frequency (BAF) calculations. For tumor-only data, SNPs with a BAF between 0.03 and 0.97 can be retained, while in tumor-normal matched data, SNPs must have greater than 30X coverage and a BAF between 0.4 and 0.6 in the matched normal sample. Further, SNPs overlapping with deletions or located in regions with high variant density can be excluded to avoid potential biases and inaccuracies in the ASCN calling process.PATENT Client Reference No.: P39388-WO-1

[0113] In some embodiments, a blacklist can be used to filter out SNPs in certain problematic regions. The blacklist can be used to force some regions to be off-target and a number of different sources can be used to identify regions for the blacklist. One example may be to use the easy regions identified in Li, Heng, “Finding easy regions for short-read variant calling from pangenome data,” arXiv:2507.03718v1 (July 4, 2025), which is herein incorporated by reference in its entirety.

[0114] Accordingly, any of the presently described methods for the ASCN calling can further include filtering the input variant calls for the ASCN calling. In some embodiments, the filtering can include: filtering for non-phased variant calls; filtering for single nucleotide variants identified by variants with only a single alternative base; filtering for single polymorphisms (SNPs); filtering for heterozygous SNPs; filtering for SNPs that do not overlap with a deletion; and filtering for SNPs that are not in regions with high variant density. In some embodiments, the filtering for heterozygous SNPs can include filtering for SNPs with 0.03 < BAF < 0.97 in tumor-only data. In some embodiments, the filtering for heterozygous SNPs can include filtering for SNPs using the matched normal with greater than 30X coverage and where the SNP has a BAF value between 0.4 < BAF < 0.6. IV. Segmentation

[0115] Fourth, segmentation can be performed by the segmentation module 308 to group adjacent target regions and single nucleotide polymorphisms (SNPs) into segments, stored in segment map 316, based on the coverage ratios for adjacent target regions and B-allele frequency (BAF) values for each SNP. More specifically, the genome can be divided into contiguous regions, or segments, that share similar copy number states. The segmentation module 308 can identify regions of the genome that have undergone copy number variations, such as deletions or duplications, which can provide insights on genomic alterations in both cancer and non-cancer contexts.

[0116] In an embodiment, the segmentation module 308 groups adjacent genomic target regions based on the similarity of their coverage ratios. This grouping can ensure homogeneity within segments, aiding in the identification of contiguous genomic regions that have undergone the same copy number changes. It can also ensure heterogeneity betweenPATENT Client Reference No.: P39388-WO-1 segments by separating regions with different copy number states, which allows for more accurate and specific analysis of each segment.

[0117] Once segments have been defined based on coverage ratios, segments can be further subdivided based on BAF values associated with SNPs within the segments. The BAF values within the segment can be analyzed via a statistical analysis. If all BAF values are similar according to the statistical analysis, then the segment is not further sub-divided. However, if at least one BAF value differs significantly from the other BAF values within the segment, then the segment can be subdivided according to the most divergent BAF value. In other words, the segment may be subdivided into two sub-segments at the location of the SNP corresponding to the most divergent BAF value. The process can be iterated recursively for the two new sub-segments until all BAF values within the sub-segments are statistically similar.

[0118] Once a segment has been defined, a segment is associated with a coverage ratio and BAF value that is aggregated from the overlapping target regions and filtered SNPs. In an embodiment, the coverage ratio for a segment is calculated as the mean value of all coverage ratios for target regions that overlap the segment. Similarly, the BAF value for a segment is calculated as the mean BAF value for all filtered SNPs that overlap the segment.

[0119] By aggregating data from multiple target regions and SNPs within a segment, the noise inherent in individual measurements can be reduced. This reduction in noise can enhance the ability to detect true copy number alterations, thereby improving the accuracy of the analysis. Random variations can be averaged out within each segment, providing more reliable estimates of the ASCN for each genomic region. Copy number alterations often affect contiguous regions of the genome, and segmenting the genome aligns with the biological reality that such alterations span multiple adjacent targets and can include a large number of SNPs.

[0120] Segmentation allows for the use of statistical models to estimate the copy number state for each segment by comparing the observed coverage ratios and the observed BAFs against the expected coverage ratios and expected BAFs, which are described in detail infra, under different copy number scenarios. This process enables the joint estimation of tumor purity and ploidy by evaluating these parameters across segmented regions.PATENT Client Reference No.: P39388-WO-1

[0121] Segmentation can begin by analyzing the coverage ratios and B-allele frequency (BAF) values derived from sequencing data. Calculating coverage ratios is described in detail supra. BAF values represent the proportion of one allele relative to the other alleles at single nucleotide polymorphism (SNP) sites. Coverage ratios and BAF values can provide complementary information about the genomic content and allelic balance in a sample.

[0122] In some embodiments, segmentation can be adapted to improve allosome handling. The reproductive chromosomes (X and Y) exhibit different rates of variation compared to the other autosomal chromosomes (1 to 22). There can also be significant differences between sexes as male differences between the X and Y chromosomes can result in false positive indications of LOH. Consequently, the segmentation algorithm can include improved allosome handling. The X and Y chromosomes have X- or Y-specific genes located between pseudo-autosomal regions (PARs) at the ends of the allosomes. The segmentation algorithm can be adapted to treat all SNPs identified between PARs as false positives. In addition, all Y chromosome PAR regions are masked (i.e., base calls are set to N) in the reference sequencing data to force all Y chromosome reads in the PAR regions to be aligned to the X chromosome in the reference.

[0123] In some embodiments, adjacent genomic regions with similar coverage ratios values and BAF values can be grouped together to form segments. This grouping can be based on statistical algorithms that can identify breakpoints where significant changes in these values occur. Various algorithms, such as, but not limited to, Hidden Markov Models (HMMs), Circular Binary Segmentation (CBS), or other statistical methods, can be used to detect these breakpoints and define the boundaries of each segment.

[0124] In an exemplary embodiment, an HMM is used to analyze coverage ratio values for the plurality of target regions and BAF values for a plurality of filtered SNPs. The HMM generates a plurality of segments based on the coverage ratios and BAF values.

[0125] Once the segments are defined, each segment can be analyzed to determine its allele-specific copy number state. These segments represent genomic regions that harbor the same ASCN state. The ASCN state can be defined by the total copy number (C) and the minor allele copy number (nb) for each segment. The segmentation process allows for the detection of both large-scale copy number changes and smaller, more localized alterations.PATENT Client Reference No.: P39388-WO-1

[0126] Accordingly, the BAF values can be based at least in part on SNP variants. In some embodiments, segmentation can include: grouping adjacent target regions based at least in part on similarity of the coverage ratios, and grouping adjacent SNPs based at least in part on similarity of the BAF values; and combining the grouped target regions and the grouped adjacent SNPs into segments with the same ASCN state. In some embodiments, any of the methods described herein can further include defining boundaries of each segment by maximizing a likelihood that the coverage ratios and the BAF values within each segment are similar.

[0127] In some embodiments, hyper segmentation can be mitigated by determining whether to merge some segments after segmentation. More specifically, the coverage ratio signal can be analyzed for all segments to determine a variance in the coverage ratio signal. Then, if the difference between the mean coverage ratio for two adjacent segments is less than the variance of the coverage ratio signal over all segments, then the two segments can be merged. A. B-Allele Frequency (BAF)

[0128] A B-allele refers to one of the two possible alleles at a specific single nucleotide polymorphism (SNP) site in the genome. For a given SNP, there are typically two possible alleles, which can be designated as the A-allele and the B-allele. The designation of A and B is arbitrary and is used simply to distinguish between the two different nucleotide sequences that can occur at that SNP position.

[0129] B-allele frequency (BAF) refers to the proportion of B-allele relative to A-allele at a specific SNP site. BAF can provide information about the allelic composition at a particular genomic location for CNV and ASCN analysis. BAF can be calculated from sequencing or genotyping data that can identify the two alleles present at an SNP site. In a normal diploid genome, where there are two copies of each chromosome, the BAF typically has a value of approximately 0.5 at heterozygous SNP sites, indicating that both alleles are present in equal proportions. If the SNP site is homozygous, the BAF has a value close to 0 (if only the A allele is present) or close to 1 (if only the B allele is present). The calculation of BAF can involve determining the number of reads supporting each allele at the SNP site. For example, if there are 100 reads covering a particular SNP, with 50 reads corresponding toPATENT Client Reference No.: P39388-WO-1 the A allele and 50 reads to the B allele, the BAF would be 0.5. If there are 80 reads for the A allele and 20 reads for the B allele, the BAF would be 0.2, indicating an imbalance in the allelic representation. BAF can be useful in detecting and interpreting copy number alterations and loss of heterozygosity (LOH). In cases of allele-specific copy number gains or losses, the BAF values can deviate from the expected 0.5 at heterozygous sites. In this instance, the total copy number can remain neutral but BAF can shift because one allele can be lost while the remaining allele can be gained. For example, if there is a duplication of one allele (resulting in three copies of the A allele and one copy of the B allele), the BAF can shift to values closer to 0.33 or 0.67. In cases of LOH, where one allele is completely lost, the BAF can shift to 0 or 1, depending on which allele is retained.

[0130] BAF can allow for more precise segmentation because it can add an additional layer of information about the genetic variation within the segments compared to using just coverage ratios in the segmentation process. This can lead to more accurate identification of breakpoints and the definition of segments with consistent copy number states.

[0131] In some embodiments, BAF can help the segmentation module 308 detect allelic imbalances that can occur due to copy number variations (CNVs) such as deletions, duplications, or more complex alterations. For example, in a diploid genome, deviations from the expected BAF value of around 0.5 at heterozygous SNP sites can indicate an allelic imbalance. In another example, a BAF value of 0 or 1 can suggest LOH, while values like 0.33 or 0.67 can indicate the presence of three copies, with two copies of one allele and one of the other.

[0132] In some embodiments, BAF can provide information that can complement the coverage ratio data. While the coverage ratio measures the overall amount of DNA in a region, it does not distinguish between contributions from different alleles. By incorporating BAF, ASCN calling can more accurately determine whether changes in coverage are due to gains or losses of specific alleles. Thus, the combined approach as described herein can distinguish the differences between total copy number changes and allele-specific alterations.

[0133] In some embodiments, a mirrored BAF value can be used. In segments where there is an allelic imbalance, the density function of BAF values within the segment may have a bimodal distribution. By mirroring the BAF values around the mid-point of 0.5, wePATENT Client Reference No.: P39388-WO-1 can change the distribution back into a unimodal distribution. A mirrored BAF value is simply one minus the BAF value if the BAF value is greater than 0.5, or the BAF value if the BAF value is less than 0.5.

[0134] In some embodiments, BAF can be useful for identifying regions with copy- neutral LOH, where the total copy number remains the same but one allele is lost and the other is duplicated. In such cases, the coverage ratio will not show a change, but the BAF will shift towards 0 or 1, indicating the allelic imbalance. This type of alteration can be important in cancer genomics, as it can lead to the unmasking of recessive mutations or the duplication of oncogenic alleles. In cancer genomics, BAF can provide insights into tumor heterogeneity by revealing subclonal populations with different allelic compositions. Tumors often consist of multiple cell populations with distinct genetic profiles. By analyzing BAF along with coverage data, the presently described ASCN calling methods can identify and characterize these subclonal populations, aiding in the understanding of tumor evolution and progression. B. Total Copy Number

[0135] Total copy number (C) refers to a total number of copies of a specific genomic region present in a cell. In a normal diploid cell, the total copy number for most autosomal regions is typically 2, representing one copy inherited from each parent. On the other hand, allosome regions can be one, depending on the sex. However, in cancer cells and other contexts with genomic alterations, the total copy number can vary significantly, with regions being duplicated (increased copy number) or deleted (decreased copy number).

[0136] To determine a total copy number, a coverage ratio can be calculated by comparing the sequencing read depth of a sample to a reference. This read depth indicates how much DNA from each region is present in the sample, normalized against either a reference (e.g., matched normal sample 324 or a panel of normal samples 322) to correct for any biases in sequencing. GC content correction can be performed using e.g., LOESS smoothing, which can adjust the coverage data to account for these variations. A segmentation process can be applied once the corrected coverage ratios are obtained by grouping adjacent genomic regions that show similar coverage ratios and B-allele frequenciesPATENT Client Reference No.: P39388-WO-1 (BAFs) into segments, each representing a contiguous region with a consistent copy number state.

[0137] For each segment, the observed coverage ratio (CR) for a segment can then be compared against the expected coverage ratio under different copy number scenarios. These comparisons can help identify regions where the copy number deviates from the normal 2 copies. The difference can be quantified using a likelihood function p(CR | E[CR | ]), where refers to the set of values pertaining to a particular segment. For each possible copy number, a likelihood can be calculated using a likelihood function p(CR | E[CR | , , Ck]), where is tumor purity, is tumor ploidy, C is total copy number, and k {0, …, 7}. The C value that can maximize this likelihood can be selected as an optimal total copy number for the segment.

[0138] In the presently described methods, the total copy number can be used to refine the ASCN estimates further. Once the total copy number for a segment is determined, B- allele frequencies (BAFs) can be used to distinguish between the contributions of different alleles within that total. This can identify whether the observed changes are due to alterations in one allele (e.g., one copy being duplicated) or both alleles (e.g., both copies being deleted or duplicated).

[0139] This process can be iterative and involve evaluating a range of candidate values for tumor purity and ploidy, using a 2D grid search. For each combination of these parameters, the likelihood of observing the given coverage ratio can be calculated, and the combination that can provide the highest total likelihood across all segments can be chosen. C. Minor Allele Copy Number

[0140] Minor allele copy number (nb) refers to a number of copies of the less frequent allele at a specific genomic locus in a sample. In diploid organisms, each individual has two alleles for each locus, one inherited from each parent. The minor allele is the one present in a smaller number of copies compared to the major allele, which is a more frequent allele.

[0141] Determining the minor allele copy number involves analyzing the B-allele frequency (BAF), which can represent the proportion of sequence reads supporting the minor allele at a given SNP. In an embodiment, if the total copy number for a region is 3 based onPATENT Client Reference No.: P39388-WO-1 the observed coverage ratio, and the BAF for B-allele at a SNP in this region is observed to be 0.33, it suggests that one out of three copies is the minor allele (nb= 0.33 × 3 = 1).

[0142] The minor allele copy number can be used alongside the total copy number (C) to distinguish between different types of genomic events. For example, a region with an elevated total copy number and a high minor allele copy number may indicate a duplication event involving both alleles. In another example, a region with a high total copy number but a low minor allele copy number may indicate a copy-neutral loss of heterozygosity (LOH), where one allele is duplicated while the other is lost. In some embodiments, the minor allele copy number can aid in the estimation of tumor purity and ploidy. V. Estimation of Segment ASCN

[0143] Lastly, for ASCN calling, the ASCN state, tumor purity, and tumor ploidy can be estimated by the estimation module 310 using a statistical framework that can maximize the likelihood of observed data given expected values. Tumor purity and ploidy 342 can be inferred by evaluating a 2D grid of candidate values and selecting the combination that can yield the highest total likelihood. The ASCN state 344 can then be estimated based on the total copy number, tumor purity, and tumor ploidy. The ASCN state refers to the specific configuration of copy numbers for each allele in a genomic segment, including the total copy number as well as the distribution of these copies between the different alleles (typically A- allele and B-allele).

[0144] In non-ASCN calling contexts, such as CNV calling in non-cancerous cells, the estimation of segment ASCN can be simplified due to the assumptions that the cells are diploid and the purity is essentially 100% (i.e., a sample is not mixed with normal cells andcancer cells). In such case, tumor purity ( ) can be set to 1, and tumor ploidy ( ) can be setto 2, reflecting the diploid nature of normal human cells.

[0145] In the context of ASCN calling, which pertains to cancer genomics, tumor purity and tumor ploidy are required for accurately estimating the ASCN state. These parameters account for the presence of normal cells and the abnormal chromosome numbers typical in cancer cells. Tumor purity and tumor ploidy estimations are explained in detail infra.PATENT Client Reference No.: P39388-WO-1

[0146] Accordingly, the total copy number can be estimated by determining the expected coverage ratios and comparing the observed coverage ratios with the expected coverage ratios. A. Expected Coverage Ratio

[0147] The expected coverage ratio ( ) can be calculated using Equation 1, setforth above, wherein represents the parameters of tumor purity ( ), tumor ploidy ( ), andtotal copy number (C) for segment s.

[0148] The observed coverage ratio (CRs) for a segment can then be compared against the expected coverage ratio. The difference can be quantified using a likelihood function p(CRs| ). For each possible copy number, a likelihood can be calculated using a likelihood function p(CRs| ), where k {0, …, 7}. The value that can maximize this likelihood can be selected as an optimal total copy number for the segment s.

[0149] In some embodiments, calculating the expected coverage ratios of a segment can further include determining the values of the tumor purity, the tumor ploidy, and the total copy number to minimize the difference between the coverage ratios and the expected coverage ratios. B. Expected B-Allele Frequency

[0150] The expected BAF ( ) can be calculated using the fixed purity and thenumber of minor alleles (nb), as shown in Equation 2, set forth above, wherein represents tumor purity, nb represents the number of minor alleles, C represents total copy number, which comes from the C value that can provide the maximum likelihood on the coverage ratios described supra.

[0151] The observed BAF (BAFs) can be compared against the expected BAF, and the difference can be quantified using a likelihood function, p(BAFs |, where {0, …, C / 2}. Thevalue that can maximize this likelihood can be selected as the optimal number of minor alleles for the segment s. The highest BAF likelihood represents the value of nb that can fit the observed data the best.PATENT Client Reference No.: P39388-WO-1

[0152] In some embodiments, the minor allele numbercan be estimated based at least in part on the total copy number C and the tumor purity . In some embodiments, the estimated minor allele number can make the BAF values most probable given the expected BAF values. C. Tumor Purity and Ploidy Estimation

[0153] Tumor purity refers to the proportion of cancer cells in a sample. Tumor ploidy refers to the number of sets of chromosomes present in the tumor cells. Tumor purity and tumor ploidy must be jointly estimated to accurately determine the genomic alterations present in the cancer cells in the context of ASCN calling.

[0154] For estimation of tumor purity ( ) and tumor ploidy ( ), a 2D grid over a range ofcandidate and values can be created. For tumor purity , values between 0.01 and 1 can be considered, representing the fraction of the sample that is made up of tumor cells. For tumor ploidy, values between 1 and 6 can be used, reflecting the possible sets of chromosomes in the cancer cells. In some embodiments, the 2D grid can consist of a total of 4,950 purity / ploidy combinations to consider.

[0155] Each combination of and can then be used to calculate the expected coverageratio ( ),) for each segment. The observed coverage ratio for segment s (CRs)can be compared to the expected value, and the difference can be quantified using alikelihood function. The expected BAF ( ) can be calculated for eachcombination of and . The observed BAF for segment s (BAFs) can be compared to this expected value, and the difference can also be quantified using a likelihood function.

[0156] The total likelihood can be calculated by summarizing across all segments. The joint likelihood of each segment can be calculated by combining the likelihood values from the coverage ratio and BAF value. The total likelihood across all segments can then be summarized to identify the combination of tumor purity and ploidy values that can maximize this total likelihood. The combination with the highest total likelihood can be identified as the optimal estimate of tumor purity and ploidy for the sample. These optimal values can then be used to generate the final ASCN predictions for each segment.PATENT Client Reference No.: P39388-WO-1

[0157] In some embodiments, a known issue with tumor purity / ploidy estimation can be the tendency for high ploidy solutions to be favored, a common problem in many Illumina- based ASCN callers. This issue can be addressed by setting a higher prior on the copy number neutral state (i.e., 2 for total copy number, abbreviated as 2N) for each segment. This approach can be defensible as the baseline copy number in normal cells is 2. Consequently, the higher prior on the 2N state can necessitate significant evidence to classify a segment as an ASCN event. This method naturally downweighs high ploidy solutions, as the baseline copy number in such solutions can exceeds 2. Additionally, somatic single nucleotide variants (SSNV) can be incorporated to assist in selecting the correct tumor purity / ploidy solution.

[0158] Another embodiment can exclude segments with high coverage ratio values from the purity / ploidy estimation. In the scenario with few SCNA events, the highly amplified segments can bias the model towards high purity scenarios. In order to mitigate this effect, a threshold can be set to exclude certain segments for the purposes of estimating tumor purity and ploidy. For example, a threshold of 1 for log coverage ratio values (indicating a corrected coverage ratio of 2) can be used to exclude those segments that have higher coverage counts from total likelihood calculations. This threshold corresponds to a total copy number of 5+ at greater than 75% tumor purity and 6+ at greater than 50% tumor purity, which would likely exclude only a small subset of all segments and is only relevant in samples with few SCNA events.

[0159] Accordingly, any of the methods described herein can further include: generating a plurality of pairs of candidate values for the tumor purity and the tumor ploidy; calculating the expected coverage ratios and the expected BAF values for each pair of the candidate values for the tumor purity and the tumor ploidy using the Equations 1 & 2 set forth above; selecting a subset of pairs of the candidate values for the tumor purity and the tumor ploidy that provides minimum combined differences between (i) the coverage ratio and the expected coverage ratio and (ii) the BAF values and the expected BAF values.

[0160] In some embodiments, any of the methods provided herein can further include: determining the tumor purity and tumor ploidy values based on the selected subset of candidate pairs; and using the determined tumor purity and tumor ploidy values to refine the estimation of ASCN states for each segment by minimizing the combined differencesPATENT Client Reference No.: P39388-WO-1 between the observed coverage ratios and the expected coverage ratios, and between the observed BAF values and the expected BAF values.

[0161] In some embodiments, any of the methods provided herein can further include setting a higher prior on a copy number neutral state for each segment. In some embodiments, the copy number neutral state is 2. In some embodiments, the higher prior can downweigh a baseline copy number in a high ploidy solution. In some embodiments, the copy number neutral state of the high ploidy solution can be as greater than 2.

[0162] The presently described method of ASCN calling is superior to existing methods due to its comprehensive approach and enhanced accuracy in estimating allele-specific copy numbers, tumor purity, and ploidy. The presently described methods integrate both coverage ratio and B-allele fraction (BAF) data, providing a more detailed and nuanced analysis of genomic alterations, particularly in cancer samples.

[0163] For example, compared to existing calling methods like PureCN, the presently described method offers several advantages. PureCN is a purity and ploidy aware copy number caller for cancer samples that was designed for hybrid capture sequencing data (see, e.g., Riester et al., “PureCN: copy number calling and SNV classification using targeted short read sequencing,” Source Code Biol Med 11, 13 (2016)). While PureCN is effective in certain contexts, it may not perform as well with whole genome sequencing data. The presently described method, however, is adaptable to different sequencing configurations, including both target enrichment and whole genome sequencing, thereby offering broader applicability.

[0164] Another existing method, CNVkit, provides tools to infer and visualize copy number from sequencing data. CNVkit is designed for use with hybrid capture, including whole-exome and custom target panels, and short-read sequencing platforms, such as Illumina and Ion Torrent (see, e.g., Talevich et al., “CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted DNA Sequencing,” PLOS Computational Biology (April 21, 2016)). However, CNVkit does not integrate B-allele frequency (BAF) data as effectively as the presently described method. B-allele frequency (BAF) data sequencing is a method for calculating the frequency of variant alleles in a sequenced sample. CNVkit’s approach is primarily based on coverage data, which can limit its accuracy in distinguishingPATENT Client Reference No.: P39388-WO-1 between different allele-specific events, particularly in regions with complex genetic alterations.

[0165] FACETS is another tool used for copy number analysis and tumor purity estimation. FACETS is an ASCN tool that can be used with whole-genome and whole- exome sequencing assays, as well as targeted panel sequencing platforms. It is a fully integrated, stand-alone pipeline that includes sequencing BAM file post-processing, joint segmentation of total- and allele-specific read counts, and integer copy number calls corrected for tumor purity, ploidy and clonal heterogeneity, with comprehensive output and integrated visualization (see, e.g., Shen et al., “FACETS: allele-specific copy number and clonal heterogeneity analysis tool for high-throughput DNA sequencing,” Nucleic Acids Res. 19;44(16):e131 (2016)). While FACETS also utilizes both coverage and BAF data, it often requires more extensive computational resources and time to produce results. The presently described method, by optimizing the grid search over candidate values for tumor purity and ploidy, can achieve similar or better accuracy with improved efficiency.

[0166] The presently described methods of ASCN calling utilize a 2D grid over candidate values for tumor purity and ploidy to allow for a systematic exploration of possible solutions, ensuring that the optimal values can be identified with high confidence. This comprehensive grid search can enhance robustness and accuracy, particularly in samples with mixed cell populations and varying levels of tumor purity.

[0167] As demonstrated in Example 1 infra, the presently described methods of ASCN calling demonstrate high concordance with PureCN in terms of tumor purity and ploidy estimation. The integration of coverage ratio and BAF data, combined with the thorough grid search approach, makes the presently described methods of ASCN calling a more accurate, efficient, and versatile tool for genomic analysis. These improvements over existing methods like PureCN, CNVkit, and FACETS demonstrate its superiority in providing detailed insights into genomic alterations, particularly in the context of cancer research and diagnostics. 2D Grid Search

[0168] FIG.4 illustrates a flowchart of a method 400 for performing a 2D grid search, in accordance with at least some embodiments of the present disclosure. The method 400 mayPATENT Client Reference No.: P39388-WO-1 be performed as part of step 210 of the method 200. The grid search may be performed, at least in part, by one or more processors executing instructions stored in a computer-readable medium. Of course, any system capable of performing the steps of the method 400 is within the scope of the present disclosure.

[0169] At 402, a number of values for a first independent variable are defined. The number of values can be selected within a range of possible values for the first independent variable. In an embodiment, the first independent variable is tumor purity , and the range of values may be between 0.01 and 1, inclusive. In an embodiment, the values can be evenly distributed within the range, and a number of values within the range can be configured statically (e.g., pre-set) or dynamically (e.g., based on some parameters such as desired accuracy and / or execution time). For example, the number of values within the range can be fixed (e.g., 50, 100, 1000, etc.) and evenly distributed. In other embodiments, the values may be selected randomly or pseudo-randomly according to a target distribution (e.g., as defined by a mean and standard deviation, for example).

[0170] At 404, a number of values for a second independent variable are defined. The number of values for the second independent variable can be selected within a range of values for the second variable, which may be different from the first range. In an embodiment, the second independent variable is tumor ploidy , and the range of values for tumor ploidy can be, e.g., between 1 and 5, 6, 7, or more to represent the number of sets of chromosomes in the tumor tissue for a given segment. In an exemplary embodiment, the range of values for tumor ploidy can be between 1 and 6, inclusive.

[0171] At 406, each combination of values for the first and second independent variable are used to calculate a dependent variable determined by application of a formula of the first independent variable and the second independent variable. Each combination of and can then be used to calculate the expected coverage ratio ) for each segment, and theobserved coverage ratio ( ) for segment can be compared to the expected value for thatparticular combination, where the difference between the observed and expected values can be quantified using a likelihood function. In an embodiment, each combination of values for the first and second independent variable can also be used to calculate a dependent variable determined by application of a second formula. Each combination of and can then bePATENT Client Reference No.: P39388-WO-1 used to calculate the expected coverage ratio ) for each segment, and theobserved BAF ( ) for segment can be compared to the expected value for thatparticular combination, where the difference between the observed and expected values can also be quantified using a likelihood function.

[0172] At 408, select a particular combination of values for the two independent variables that maximizes a total likelihood of dependent variable(s) across a plurality of segments. A total likelihood value can be calculated by summarizing the likelihood values across all segments . The joint likelihood of each segment can be calculated by combining the likelihood values from the observed coverage ratio and BAF values given the expected values for these dependent variables. The total likelihood across all segments can then be summarized to identify the combination of tumor purity and ploidy values that can maximize this total likelihood. The combination of values for tumor purity and tumor ploidy with the highest total likelihood can be identified as the optimal estimate of tumor purity and tumor ploidy for the sample. These optimal values can then be used to generate the final ASCN predictions for each segment. Unified Copy Number Caller

[0173] FIG.8 illustrates a conceptual flow chart of a unified copy number caller, in accordance with at least some embodiments. As discussed above, the ASCN caller can be adapted for different configurations. In an embodiment, the configurations include CNV calling for germline WGS, ASCN calling for somatic tumor / normal WGS, and / or ASCN calling for targeted enrichment (TE) assays. The unified copy number caller shown in FIG.8 includes different workflows for each of these types of assays, but all three of the workflows eventually result in the calculation of (log) coverage ratios, which are then used as an input to the segmentation algorithm 855, which is common to all three workflows.

[0174] As shown in FIG.8, a first workflow 810 may be configured to ASCN calling in tumor / normal WGS assays. The first workflow 810 starts with receiving a BAM file containing aligned SBX sequencing data for a normal sample. As described above, the sequence data is subdivided into a plurality of target regions of interest and coverage values are calculated for each region. The coverage values are corrected for GC-bias using, e.g., LOESS curve fitting, to generate a set of corrected normal counts for the coverage values. InPATENT Client Reference No.: P39388-WO-1 parallel, corresponding aligned SBX sequencing data for a tumor sample is processed in parallel (or sequentially), and coverage values are calculated and corrected for GC-bias to generate a corresponding set of tumor counts for the coverage values. A log coverage ratio value is then calculated for each region by comparing the normal counts for the normal sample to the tumor counts for the tumor sample, the ratio being scaled by a base two logarithm to generate log coverage ratio values 850 for each of the regions of interest in the sequence.

[0175] A second workflow 820 may be configured for CNV calling in germline WGS assays. The first few steps of the second workflow are similar to the first workflow in that a BAM file containing aligned SBX sequencing data for a sample is received, coverage values are calculated and corrected for GC-bias to generate counts for corrected coverage values for each of the regions of interest. In the case of germline WGS assays, the reference used to calculate the log coverage ratio is a self-normalization technique where coverage values from within the same sample in a number of low CNV portions of the genome are averaged to generate a mean coverage value for the sample. The counts for the corrected coverage values are then compared against the mean coverage value for the sample, the ratio being scaled by the base two logarithm to generate the coverage ratio values 850.

[0176] A third workflow 830 may be configured for ASCN calling in TE assays. First, aligned SBX sequencing data is received in a BAM file. Coverage values are calculated for each of the regions and corrected for GC-bias. Information specifying the augmented baits are provided in a TSV file, which is used to distinguish the on-target regions for the TE assay from the off-target regions. In the third workflow 830, a panel of normal is used as a reference to calculate log coverage ratio values by comparing the tumor counts with average counts for coverage values associated with the panel of normals. The log coverage ratio values can also be denoised by subtracting the projection of the tumor log-ratio values onto the singular value decomposition (SVD) of the counts from the panel of normal which removes technical variation introduced by the sequencing technology.

[0177] The segmentation algorithm 855 then combines regions of interest having similar log coverage ratio values. Optionally, germline heterozygous variants information is received at 840 in a VCF file. BAF values for SNPs can be extracted from the VCF file and provided to the segmentation algorithm 855. The segmentation algorithm 855 generates a number ofPATENT Client Reference No.: P39388-WO-1 segments 860 corresponding to similar coverage ratio values and, optionally, BAF values. Purity / Ploidy values 865 are estimated using the techniques described above, and a copy number and / or minor allele number 870 as well as corresponding likelihood values are estimated for each segment, which is output as the ASCN state 875. In addition to the techniques described above using statistical analysis, in an embodiment, the segmentation algorithm 855 may utilize a Hidden Markov Model for estimating the copy number state for each segment based on the calculated log coverage ratio values and / or BAF values calculated for each segment. Hidden Markov Model (HMM) for Segmentation

[0178] FIGS.9A & 9B illustrate a HMM 900 for performing segmentation, in accordance with at least some embodiments of the present disclosure. In an embodiment, as illustrated in FIG.9A, when performing ASCN calling for germline WGS, a HMM 900 is provided for inferring the copy number of each target region given the calculated coverage signals.

[0179] In an embodiment, each of the hidden states 910 represents the set of total copy numbers associated with a region of the genome associated with the sequencing data. The emission values 920 represents the calculated coverage ratios . In an embodiment, the coverage ratio is a log coverage ratio, which is similar to the coverage ratio calculated above taken to the base 2 logarithm. The use of the log coverage ratio instead of the coverage ratio can restrict the range of emission values that can be generated by the HMM 900.

[0180] Furthermore, while the coverage ratio is defined within a continuous range of values, the log coverage ratios can be quantized such as by rounding down to the nearest whole integer. Thus, the emission values for the calculated log coverage ratios may be limited to ranges of calculated coverage ratios between powers of two, which may make estimation of specific copy numbers for each target region easier to calculate. Of course, any suitable quantization may be appropriate, such as quantizing the log coverage ratio to 0.1 increments, which will give more possible discrete emission values 920 in the HMM 900.

[0181] In an embodiment, the number of hidden states can be limited to a fixed number of possible copy numbers (e.g., 7) that are expected to be encountered in genetic informationPATENT Client Reference No.: P39388-WO-1 of a population. For example, in a human diploid organism, the most probable copy number in a normal region of the genome is 2. However, aneuploidy in tumor tissue may result in copy numbers higher than 2 and deletions may result in copy numbers less than 2 for particular regions of the genome. By limiting the number of hidden states to a reasonable number, the complexity of the HMM 900 is reduced. Further, the state associated with the highest copy number , for example, can be used to represent any copy number greater than or equal to 7. Similarly, the emission values can similarly be limited to a reasonablerange with , for example, covering coverage ratios up to 32 (. Thispractical limitation of the number of possible hidden states and the range of expected coverage ratios can be adapted based on the type of sequencing being performed. In practice, there is a practical limit to the number of reads that are generated for a given region of the genome in a germline WGS assay. Given this practical limit and the coverage ratios in the reference (e.g., self-normalized sample), the calculated coverage ratio from sequencing data is likewise limited to a particular range. While this range could be increased by changing the parameters of the sequencing process, the architecture of the HMM can be adapted to match that procedure.

[0182] The HMM 900 can be trained by using sequencing data with calculated coverage ratios using other methods. Given a plurality of samples in a training data set, sequencing data can be generated for each of the samples. Coverage ratios can then be calculated for each of the target regions in the sequencing data. The plurality of series of coverages ratios for each of the samples in the training data set can then be used to train the model to find the transition probabilities between hidden states and the emission probabilities for observing each emission value given the current hidden state. In the HMM 900 the emission probabilities from each hidden state are modeled as Gaussian probability density functions (PDF), with the mean of the PDF calculated as the expected log coverage ratio value, given estimates for tumor purity, tumor ploidy, and the copy number for the corresponding hidden state. The PDF may be given as follows:where is the copy number from the corresponding hidden state, is the tumor ploidy, is the tumor ploidy, and is a predefined variance of the log coverage ratio.PATENT Client Reference No.: P39388-WO-1

[0183] Instead of observation values being sampled over time (i.e., a time series), here, the observation values are the calculated coverage ratio information over the different regions in the sequence. Intuitively, this makes sense as the copy number for one region is likely to be similar to the copy number for adjacent regions having the same ASCN state. However, when two adjacent regions do not share the same ASCN state, then that indicates a transition between hidden states associated with the two adjacent regions in the genome. Thus, by iterating over the plurality of target regions and their calculated log coverage ratio values, the HMM 900 can be used to infer the most likely sequence of copy number states for each of the regions. This inference can be used to segment the plurality of target regions into corresponding segments that share the same copy number value.

[0184] In the case of germline WGS assays, the HMM 900 can rely on, as input, the set of log coverage ratio values. However, in the case of tumor / normal WGS or TE assays, a different embodiment of the HMM may be used that relies on both the log coverage ratio values as well as mirror BAF values.

[0185] As shown in FIG.9B, the architecture of the HMM 950 can be expanded to include two element vectors for each of the hidden states 960 and emission values 970. The hidden states 960 include states for all possible combinations of both the total copy number and the minor allele number. It will be appreciated that the minor allele number is limited to be half of the total copy number and, thus, the total number of hidden states is therefore limited to all possible combinations of both the total copy number and the minor allele number, which will be equal to. The emission values 970 for HMM 950 also take the form of two element vectors that include all combinations of both the log coverage ratio CR and BAF values. Like the log coverage ratio values in HMM 900, both he CR and BAF values may be quantized to reduce the total possible combinations of emission values to some reasonable number of discrete combinations. For example, in an embodiment, BAF values are quantized to be between 0 and 1 in 0.1 increments. In other embodiments, BAF values are quantized to be between 0 and 1 in 0.01 increments, exclusive, yielding 99 possible BAF values.

[0186] The HMM 950 is trained and used for inference in a similar way to HMM 900, except that both total copy number and minor allele number are estimated for eachPATENT Client Reference No.: P39388-WO-1 target region. Segmentation is then performed by combining adjacent target regions that share the same total copy number and minor allele number .

[0187] In the HMM 950 the emission probabilities from each hidden state are modeled as a bivariate Gaussian probability density functions (PDF), which is dependent on both the expected values for the log coverage ratio (which depends on estimates for tumor purity / ploidy as well as the copy number for the corresponding hidden state) and the expected values for the BAF (which depends on estimates for tumor purity as well as the copy number and minor allele number for the corresponding hidden state). The bivariate PDF may be given as follows: , (Eq.8) where is the copy number from the corresponding hidden state, is the tumor ploidy, is the tumor ploidy, is the minor allele number, and and are predefined variances of the log coverage ratio and BAF, respectively. Computing System

[0188] FIG.5 depicts a block diagram illustrating a computing system 500 consistent with implementations of the current subject matter. As shown in FIG.5, the computing system 500 can include a processor 510, a memory 520, a storage device 530, and input / output interface 540. The processor 510, the memory 520, the storage device 530, and the input / output interface 540 can be interconnected via a system bus 501. The computing system 500 may additionally or alternatively include a graphic processing unit (GPU) 550, such as for image processing, and / or an associated memory for the GPU. The GPU 550 and / or the associated memory for the GPU may be interconnected via the system bus 501 with the processor 510, the memory 520, the storage device 530, and the input / output interface 540. The memory associated with the GPU may store one or more images described herein, and the GPU may process one or more of the images described herein. In an embodiment, the GPU may be coupled to and / or form a part of the processor 510. The processor 510 is capable of processing instructions for execution within the computing system 500. In some implementations of the current subject matter, the processor 510 can be a single-threaded processor. Alternately, the processor 510 can be a multi-threaded processor.PATENT Client Reference No.: P39388-WO-1 The processor 510 is capable of processing instructions stored in the memory 520 and / or on the storage device 530 to display graphical information for a user interface provided via the input / output interface 540.

[0189] The memory 520 is a computer readable medium such as volatile or non-volatile memory that stores information within the computing system 500. The memory 520 can store data structures representing configuration object databases, for example. The storage device 530 is capable of providing persistent storage for the computing system 500. The storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, or other suitable persistent storage means. The input / output device 540 provides input / output operations for the computing system 500. In some implementations of the current subject matter, the input / output interface 540 includes an interface for a keyboard and / or pointing device. In various implementations, the computing system 500 includes a display device 555 for displaying graphical user interfaces. The display device 555 may receive data from the GPU 550 either through the system bus 501 or through another communication interface (e.g., PCIe, HDMI, SDP, etc.) associated with or directly connected to the GPU 550.

[0190] According to some implementations of the current subject matter, the input / output interface 540 can provide input / output operations for a network device. For example, the input / output interface 540 can include Ethernet ports or other networking ports to communicate with one or more wired and / or wireless networks (e.g., a local area network (LAN), a wide area network (WAN), the Internet).

[0191] In some implementations of the current subject matter, the computing system 500 can be used to execute various interactive computer software applications that can be used for organization, analysis and / or storage of data in various (e.g., tabular) format (e.g., Microsoft Excel®, and / or any other type of software). Alternatively, the computing system 500 can be used to execute any type of software applications. These applications can be used to perform various functionalities, e.g., planning functionalities (e.g., generating, managing, editing of spreadsheet documents, word processing documents, and / or any other objects, etc.), computing functionalities, communications functionalities, etc. The applications can include various add-in functionalities or can be standalone computing products and / or functionalities. Upon activation within the applications, the functionalities can be used to generate the userPATENT Client Reference No.: P39388-WO-1 interface provided via the input / output interface 540. The user interface can be generated and presented to a user by the computing system 500 (e.g., on a computer screen monitor, etc.).

[0192] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed framework specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0193] These computer programs, which can also be referred to as programs, software, software frameworks, frameworks, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural language, an object-oriented programming language, a functional programming language, a logical programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs), used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine- readable medium can store such machine instructions non-transitory, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example as would a processor cache or other random access memory associated with one or more physical processor cores.PATENT Client Reference No.: P39388-WO-1

[0194] To provide for interaction with a user, one or more aspects or features of the subject matter described herein can be implemented on a computer having a display device 555, such as for example a cathode ray tube (CRT) or a liquid crystal display (LCD) or a light emitting diode (LED) monitor for displaying information to the user and a keyboard and a pointing device, such as for example a mouse or a trackball, by which the user may provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, such as for example visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including, but not limited to, acoustic, speech, or tactile input. Other possible input devices include, but are not limited to, touch screens or other touch-sensitive devices such as single or multi-point resistive or capacitive trackpads, voice recognition hardware and software, optical scanners, optical pointers, digital image capture devices and associated interpretation software, and the like. Examples

[0195] These examples are provided for illustrative purposes only and not to limit the scope of the claims provided herein. Example 1. Tumor Purity and Ploidy Estimation and Gene-Level ASCN Calling

[0196] An experiment was performed to evaluate the performance of the presently described ASCN calling techniques. A comparative analysis was conducted against PureCN, an existing ASCN caller designed for target enrichment (TE) assays.

[0197] A dataset comprising 25 tumor samples, along with a reference panel of 8 normal samples, was utilized. The tumor purity and tumor ploidy for each tumor sample was estimated as described herein and using PureCN, and gene-level ASCN calling was performed. The results were then compared with those from PureCN.

[0198] The correlation between the tumor purity and tumor ploidy estimates of the presently described method and those obtained using PureCN was assessed. The presently described method demonstrated a high degree of concordance with PureCN, with a purity correlation of 0.939 and a ploidy correlation of 0.934, as illustrated in FIGS.6A & 6B, respectively.

Claims

PATENT Client Reference No.: P39388-WO-1 [0199] Further analysis was performed to determine the concordance of gene-level ASCN calling between the presently described method and PureCN. Focusing on samples with a tumor purity greater than 25%, the presently described method exhibited a gene-level ASCN calling concordance of 95% with PureCN, as illustrated in FIG.

7. [0200] These results demonstrates that the presently described method provides reliable and accurate estimates of tumor purity, ploidy, and gene-level ASCN calls, closely matching the performance of PureCNPATENT Client Reference No.: P39388-WO-1 CLAIMS 1. A computer-implemented method for allele-specific copy number (ASCN) calling, the method comprising: receiving sequencing data comprising a plurality of target regions of interest derived from a sample; calculating a coverage ratio for each target region of the plurality of target regions based on the sequencing data; receiving variant call information derived from the sequencing data; filtering the variant call information according to one or more filter criteria; segmenting the target regions to define a plurality of segments based at least in part on the coverage ratio for each target region and at least one B-allele frequency (BAF) value for one or more filtered variants in the variant call information within the target region; and estimating an allele-specific copy number (ASCN) state for a segment of the plurality of segments based at least in part on a total copy number, a tumor purity, and a tumor ploidy for the segment, wherein the tumor purity and tumor ploidy are estimated using the coverage ratio, an expected coverage ratio, the BAF value, and an expected BAF value for each segment.

2. The method of claim 1, further comprising estimating the total copy number for the segment based at least in part on the coverage ratio, the tumor purity, and the tumor ploidy.

3. The method of claim 1, wherein the segmenting comprises comparing an ASCN state of adjacent target regions and defining one or more segments where each segment comprises two or more adjacent target regions having similar ASCN state, wherein the ASCN state for a target region comprises at least the coverage ratio and a BAF value for the target region.

4. The method of claim 1, wherein estimating the expected BAF value for each segment comprises estimating a minor allele number based on the total copy number and the tumor purity.PATENT Client Reference No.: P39388-WO-1 5. The method of claim 1, wherein calculating the coverage ratio for each target region comprises: normalizing raw coverage data for the target region; selecting a reference based on a configuration; and calculating the coverage ratio based on the normalized raw coverage data and the selected reference.

6. The method of claim 5, wherein normalizing the raw coverage data comprises correcting the raw coverage data for GC content bias.

7. The method of claim 6, wherein correcting the raw coverage data for GC content bias comprises: generating a correction factor for each GC content value using a locally estimated scatterplot smoothing (LOESS) curve; applying the correction factor to the raw coverage data to obtain a GC-corrected coverage value; and comparing the GC-corrected coverage value against the reference to compute the coverage ratio.

8. The method of claim 5, wherein the configuration is selected from one of: ASCN calling in a target enrichment (TE) assay; ASCN calling in whole genome sequencing (WGS) assay; or copy number variation (CNV) calling in WGS assay; 9. The method of claim 5, wherein the reference is selected from one of: a panel of normal samples; a matched normal sample; or the sample that acts as self-normalizing.

10. The method of claim 1, wherein the filtering comprises filtering for heterozygous SNPs, and one or more of the following: filtering for non-phased variant calls;PATENT Client Reference No.: P39388-WO-1 filtering for single nucleotide variants; filtering for SNPs that do not overlap with a deletion; and filtering for SNPs that are not in regions with high variant density.

11. The method of claim 1, wherein the expected coverage ratio for a segment is calculated according to the following formula:where is tumor purity, is tumor ploidy, and is the total copy number.

12. The method of claim 11, wherein the expected BAF for a segment is calculated according to the following formula:where is tumor purity, is minor allele number, and is the total copy number.

13. The method of claim 12, further comprising performing a two-dimensional grid search to estimate the tumor purity and the tumor ploidy for the sample, the grid search comprising: generating a plurality of pairs of candidate values for the tumor purity and the tumor ploidy; calculating the expected coverage ratio and the expected BAF for each pair of candidate values for each segment in the plurality of segments; comparing the observed coverage ratio to the expected coverage ratio in each segment and comparing the observed BAF to the expected BAF in each segment; and selecting the pair of candidate values that minimizes the differences between the observed coverage ratios and the expected coverage ratios and between the observed BAF and the expected BAF across at least some of segments in the plurality of segments.

14. The method of claim 1, further comprising identifying off-target genomic regions in a TE assay and processing the sequencing data in the off-target genomic regions to improve the estimates for tumor purity and tumor ploidy in the on-target regions.PATENT Client Reference No.: P39388-WO-1 15. The method of claim 1, wherein the sequencing data is generated by one of: Sanger sequencing, Next-Generation sequencing (NGS), nanopore sequencing, ion torrent sequencing, or single molecule real-time sequencing (SMRT).

16. The method of claim 15, wherein the nanopore sequencing comprises sequencing-by- expansion (SBX) chemistry.

17. The method of claim 1, wherein segmenting the target regions comprises applying a Hidden Markov Model to the coverage ratios and BAF values.

18. A system for ASCN calling, the system comprising: a sequencing module configured to identify a plurality of target regions of interest in sequencing data derived from a sample; a data processing module configured to calculate a coverage ratio for each target region of the plurality of the target regions; a filtering module configured to filter variant call information; a segmentation module configured to define a plurality of segments based at least in part on the coverage ratio for each target region and at least one B-allele frequency (BAF) value for one or more filtered variants in the variant call information within the target region; and an estimation module configured to estimate an allele-specific copy number (ASCN) state for a segment of the plurality of segments based at least in part on a total copy number, a tumor purity, and a tumor ploidy for the segment, wherein the tumor purity and tumor ploidy are estimated using the coverage ratio, an expected coverage ratio, the BAF value, and an expected BAF value for each segment.

19. The system of claim 18, wherein the data processing module generates a corrected coverage value to correct for GC bias using a LOESS curve.

20. The system of claim 18 wherein the data processing module denoises the sequencing data using singular value decomposition.

Citation Information

Patent Citations

  • Nanopore Based Molecular Detection and Sequencing

    US20130244340A1

  • DNA sequencing by synthesis using modified nucleotides and nanopore detection

    US20130264207A1

  • Nucleic acid sequencing using tags

    US20140134616A1

  • Chemical methods for producing tagged nucleotides

    US20150368710A1

  • Tagged nucleotides useful for nanopore detection

    US20180057870A1