Targeted variant calling using target enrichment sequencing data

The targeted copy number variant calling method using hybrid capture and two-step normalization corrects for probe variability and noise, enabling accurate detection of copy number variants in target enrichment sequencing, matching whole genome sequencing performance.

WO2026073129A1PCT designated stage Publication Date: 2026-04-02ILLUMINA INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing targeted sequencing methods face challenges in accurately detecting copy number variants due to variability in capture efficiencies of probes and systematic noise, making it difficult to identify variants in regions of interest, especially in whole exome sequencing data.

Method used

A targeted copy number variant calling approach using a hybrid capture workflow with probes specific for copy number variable regions, combined with a two-step normalization process involving self-normalization and normalization against a panel of normals, to correct for probe capture variability and systematic noise.

Benefits of technology

This approach enables accurate detection of copy number variants in target enrichment sequencing data, reducing resource consumption and maintaining high concordance with whole genome sequencing results, even at lower coverage levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025048483_02042026_PF_FP_ABST
    Figure US2025048483_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The presently described techniques provide methods for identifying copy number variants using target enrichment sequencing data. In an embodiment, the normalization approach is a two-step normalization that uses a self-normalization and a normalization against target enrichment sequencing data of other samples that make up a panel of normals, which themselves are self-normalized. The self-normalization may be against a set of preselected regions already represented in the target enrichment sequencing data and that pass a set of normalization quality criteria. The normalization using the panel of normals may include a clustering technique that identifies a likely copy number neutral cluster for copy number variable region of interest in the panel of normals. In an embodiment, the normalization approach can be used in real time to normalize a sample against other samples in a sequencing run that were prepared using the same sample preparation conditions.
Need to check novelty before this filing date? Find Prior Art

Description

IP-2867-PCT|ILUM:0195PCTTARGETED VARIANT CALLING USING TARGET ENRICHMENT SEQUENCING DATATECHNICAL FIELD

[0001] The disclosed technology relates generally to sequence data assessment techniques to identify copy number variants of a genomic region, such as a gene or genes of interest.BACKGROUND

[0002] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, a problem mentioned in this section or associated with the subject matter provided as background should not be assumed to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which in and of themselves can also correspond to implementations of the claimed technology.

[0003] Genetic sequencing has become an increasingly important area of genetic research, promising future uses in diagnostic and other applications. In general, genetic sequencing involves determining the order of nucleotides for a nucleic acid such as a fragment of RNA or DNA. Some techniques involve whole genome sequencing, which involves a comprehensive method of analyzing a genome. Other techniques involve targeted sequencing of a subset of genes or regions of the genome. Targeted sequencing focuses on regions of interest, generating a smaller and more compact data set. Further, targeted sequencing reduces sequencing costs and data analysis burdens while also allowing deep sequencing at high coverage levels for detection of variants in the regions of interest.

[0004] In many cases, specific genes and / or critical diagnostic markers have been identified in portions of the genome that are present at abnormal copy numbers. For example, deletion or multiplication of copies of whole chromosomes or chromosomal segments, and higher level amplifications of specific regions of the genome, are common occurrences in certain diseases. Detection of variants may provide cliniciansIP-2867-PCT|ILUM:0195PCT with information about disease likelihood or susceptibility. Accordingly, there is a need for improved detection of variants in sequencing data.SUMMARY

[0005] In one embodiment, the present disclosure provides a normalization method. The method includes generating target enrichment sequencing data from a plurality of samples in a multiplexed sequencing run, wherein the target enrichment sequencing data is generated from each sample using sample fragments enriched using a target enrichment sequencing panel having probes specific for copy number variable regions; generating alignments to a plurality of target regions for each sample of the target enrichment sequencing data; determining a self-normalized number of sequence reads aligned to each of the plurality of target regions for each sample of the plurality of samples, wherein the self-normalized number is determined based on normalizing to alignments to a subset of target regions of the plurality of target regions; generating a panel of normals from the target enrichment sequencing data from the plurality of samples and identifying a copy number variant in an individual sample of the plurality of samples using the panel of normal. The panel of normals is generated by clustering the self-normalized number of sequence reads for each target region of the plurality of target regions across the plurality of samples; and selecting one cluster for each target region as being copy number neutral to generate the panel of normal.

[0006] In one embodiment, the present disclosure provides a normalization method. The method includes generating sequencing data from an individual sample, wherein the target enrichment sequencing data is generated from using sample fragments enriched using a target enrichment sequencing panel having probes specific for a plurality7of target regions, wherein a first subset of the target regions comprise copy number variable regions; generating alignments to the plurality of target regions using the sequencing data; determining a self-normalized number of sequence reads aligned to each of the plurality7of target regions, wherein the self-normalized number is determined based on normalizing to alignments to a second subset of target regions of the plurality of target regions; normalizing the self-normalized number of sequence reads for each individual target region of the first subset plurality of target regions usingIP-2867-PCT|ILUM:0195PCT a panel of normal, the panel of normals comprising self-normalized sequence reads for each individual target region of the first subset for one or more different samples; and identifying a copy number variant associated with an individual target region in the first subset in the individual sample based on the normalizing.

[0007] In one embodiment, the present disclosure provides a target enrichment sequencing normalization method. The method includes providing nucleic acid fragments generated from a sample; denaturing the nucleic acid fragments to generate single-stranded nucleic acid fragments; contacting the single-stranded nucleic acid fragments with a plurality of probes under conditions that permit a subset of the singlestranded nucleic acid fragments to hybridize to the plurality’ of probes, wherein the plurality of probes comprise single-stranded oligonucleotides configured to hybridize to respective different target sequences in the sample, wherein the target sequences comprise one or more sequences associated with a copy number variable gene; separating the hybridized subset of single-stranded nucleic acid fragments from an unhybridized subset of the single-stranded nucleic acid fragments to generate an enriched group of single-stranded nucleic acid fragments from the separated hybridized subset of single-stranded nucleic acid fragments; generating sequencing data comprising a plurality of sequence reads from the enriched group of single-stranded nucleic acid fragments; determining a first self-normalized number of sequence reads of the plurality of sequence reads aligned to the copy number variable gene, wherein the normalized number is normalized relative to sequence reads aligned to a set of selected sequences within the target sequences; normalizing the self-normalized number of sequence reads of the plurality of sequence reads aligned to the copy number variable gene using a second self-normalized number of sequence reads, the second self-normalized number of sequence reads being aligned to the copy number variable gene in a different sample to generate a ratio; and determining a copy number of the copy number variable gene based on the ratio.

[0008] The various methods and techniques discussed herein, may be embodied and implemented as processor-executable code or stored routine, such as may be stored on a processor-accessible memory and executed via one or more hardware or virtual processors or dedicated circuitry. As such, it should be understood that actions or stepsIP-2867-PCT|ILUM:0195PCT described a part of a method or process herein may be implemented as processor executable code, routines, or algorithms stored on one or more tangible computer- readable media.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1 shows coverage variability for whole genome sequencing (WGS) data and whole exome sequencing (WES) data.

[0010] FIG. 2 shows an example target enrichment sequencing workflow, in accordance with aspects of the present disclosure.

[0011] FIG. 3 shows an example method of targeted copy number variant calling, in accordance with aspects of the present disclosure.

[0012] FIG. 4 shows an example targeted copy number variant calling workflow for target enrichment sequencing data, in accordance with aspects of the present disclosure.

[0013] FIG. 5 shows the distribution of the number of regions by their fraction GC content for WGS vs. WES, in accordance with aspects of the present disclosure.

[0014] FIG. 6 shows concordance shown between WES genotype / copy number and WGS or Coriell genotype / copy number results.

[0015] FIG. 7 shows concordance between WES genoty pe / copy number results and WGS or Coriell genotype / copy number results for HBA.

[0016] FIG. 8 shows concordance between WES copy number results for SMN1 / SMN2.

[0017] FIG. 9 shows concordance across different coverage levels between WES genotype / copy number results and WGS or Coriell genoty pe / copy number results.

[0018] FIG. 10 show s HBA normalized depth distribution within a potential panel of normals.IP-2867-PCT|ILUM:0195PCT

[0019] FIG. 11 shows a flow diagram that illustrates the QC checks performed when running a case sample.

[0020] FIG. 12 shows concordance between WES genotype / copy number results and WGS or Coriell genoty pe / copy number results using an enhanced panel of normals.

[0021] FIG. 13 shows example copy number clusters.

[0022] FIG. 14 is a schematic illustration of polymorphic site heuristics.

[0023] FIG. 15 is a schematic illustration of polymorphic site heuristics.

[0024] FIG. 16 is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.

[0025] FIG. 17 is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 16.

[0026] FIG. 18 depicts a schematic view of an example of a system that may be used to provide biological or chemical analysis, in accordance with aspects of the present disclosure.DETAILED DESCRIPTION

[0027] The present disclosure provides a novel approach for detection of copy number variations in a biological sample. As provided herein, copy number variations (CNVs) are genomic alterations that result in an abnormal number of copies of one or more genomic regions. Structural genomic rearrangements such as duplications, multiplications, deletions, translocations, and inversions can cause CNVs. Like singlenucleotide polymorphisms (SNPs), certain CNVs have been associated with disease susceptibility.

[0028] Copy number variants, e.g., variants of HBA1 / 2, SMN1 / 2, CYP2D6 and CYP2B6, have been detected from whole genome sequencing (WGS) data. However, many research facilities are limited to or prefer whole exome sequencing (WES) or custom target enrichment panels as they provide a more affordable, scalable alternativeIP-2867-PCT|ILUM:0195PCT to WGS, requiring a smaller infrastructure (sequencers, compute, storage, etc.) to deliver the same volume of samples. A targeted variant calling solution for WES and custom target enrichment panels would facilitate copy number variant carrier screening for a broader set of customers. However, copy number variants are more difficult to identify in WES. For example, the coverage for WES is much more variable relative to WGS, which makes copy number detection challenging.

[0029] FIG. 1 shows WES data using the exome plus spike-in panel with coverage in non-exonic regions) compared to WGS data for a region on chromosome 16. As shown, the coverage for WES (top panel) shows greater variability across the region, while the WGS coverage is more even across the same region. This variability is consistent across samples prepared with the same conditions. In WES or other target enrichment sequencing, the sample nucleic acid undergoes an enrichment step that is typically not performed in WGS. The enrichment may include using capture probes to separate fragments including regions of interest in the sample from other fragments. Thus, only some of the nucleic acid in the sample is sequenced. However, capture probes introduce variability based on relative capture efficiencies of the individual probes. More efficient probes do a better job of pulling out their regions of interest relative to less efficient probes. Thus, the enriched library has an additional layer of variability in the fragments based on capture efficiencies of the various probes. This in turn results in differences in capture efficiency potentially swamping coverage differences that are related to copy number, which may lead to inaccurate copy number variant calling for target enrichment sequencing data. In addition, for copy number variants, duplicated sequences, and / or gene / pseudogenes, it may be challenging to provide capture probes to these sequences in appropriate concentrations. A probe that targets a gene having a pseudogene can hybridize to a region of both the gene and pseudogene having high sequence correspondence. Copy number variable regions may have ambiguity of mapping the short reads to the high homology regions.

[0030] Provided herein is a novel targeted copy number variant calling approach for detecting copy number variants, which may include variants in HBA1 / 2, SMN1 / 2, CYP2D6, and CYP2B6 among others, from target enrichment workflows. This approach differs from WGS-based approaches because the probe enrichment panel inIP-2867-PCT|ILUM:0195PCT use includes probes that target a set of predefined genomic regions, or copy number variable regions, that facilitate detection. The disclosed approaches permit accurate copy number variant detection using a more efficient and less data-intensive sequencing approach, which in turn results in more efficient distribution of reagents and computing resources. By eliminating a separate copy number variant caller for samples in a target enrichment workflow, the disclosed techniques save resources. That is, for each sample, rather than running the sample through multiple separate assays to yield sequencing data and copy number variant calling, the disclosed techniques provide an accurate all-in-one technique to generate both sequencing data and copy number variant calling in a target enrichment workflow.

[0031] While WGS-based approaches include all regions of the genome, the disclosed targeted sequencing copy number variant caller adds or spikes in probes during library preparation and / or as part of the target enrichment probe set that hybridize to preselected portions of copy number variable regions that pass certain quality filters. Accordingly, full-length copy number variable regions of the genome may not be sequenced as in WGS. The disclosed techniques for copy number variant calling in target enrichment sequencing data also include a normalization approach that corrects for variability introduced by different capture efficiencies of the capture probes and corrects for systematic noise from the protocol. In an embodiment, the normalization approach is a two-step normalization that uses a self-normalization and a normalization against target enrichment sequencing data of other samples that make up a panel of normals, which themselves are self-normalized. The self-normalization may be against a set of preselected regions already represented in the target enrichment sequencing data and that pass a set of normalization quality criteria. The normalization using the panel of normals may include a clustering technique that identifies a likely copy number neutral cluster for the copy number variable region of interest in the panel of normals. In an embodiment, the normalization approach can be used in real time to normalize a sample against other samples in a sequencing run that were prepared using the same sample preparation conditions.

[0032] FIG. 2 illustrates the exemplary steps of ahybrid capture workflow 10 of library preparation for target enrichment sequencing that may be used in conjunction with theIP-2867-PCT|ILUM:0195PCT targeted caller disclosed herein. Sequencing methodology of next-generation sequencing (NGS) platforms typically makes use of nucleic acid fragment libraries. In targeted sequencing techniques, a subset of fragments containing genes or regions of interest of the genome are isolated from the nucleic acid library and sequenced. Targeted approaches using NGS allow researchers to focus time, expenses, and data analysis on specific areas of interest. Such targeted analysis can include the exome (the protein-coding portion of the genome), specific genes of interest (custom content), targets within genes, or mitochondrial DNA. Targeted approaches contrast with whole genome sequencing approaches that are more comprehensive, but that also involve sequencing regions of the genome that may not be of interest to all users.

[0033] In one example of a targeted sequencing technique, hybrid capture methods use a panel or set of probes that hybridize to target sequences in the nucleic acid library. Hybridization of the probes to the target sequences allows these sequences to be separated from the rest of the fragments in the library for sequencing. By targeting only a portion of the nucleic acid library, hybrid capture methods avoid sequencing nucleic acid fragments that do not contain sequences of interest.

[0034] The workflow includes preparation of a nucleic acid library’ 20 formed from a plurality of nucleic acid fragments 22 from the sample 12, such as a sample including genomic DNA (e.g., a human genome, an animal genome, a bacterial genome) or other nucleic acids. The nucleic acid library’ 20 includes fragments 22 having sequences from regions of interest that include target sequences 24 . It should be understood that fragments 22 that include target sequences 24 may be formed entirely from regions of interest or may include other regions that are not of interest. The target-specific hybridization probes 30 are designed to be complementary' to one or more target sequences 24 on a fragment 22. Accordingly, under hybridization conditions, one or more target-specific hybridization probes 30 will bind to the complementary’ target sequences 24. This facilitates separation of fragments 22 that have the target sequences 24 from the fragments 22 that do not have the target sequences 24 to create a target-enriched sample for sequencing.IP-2867-PCT|ILUM:0195PCT

[0035] As provided herein, a target sequence 24 is a nucleic acid sequence present in a nucleic acid library that is complementary to a target-specific hybridization probe 30. Depending on the desired sequencing outcome, the target sequences 24 may be exonic sequences as part of exome sequencing. Accordingly, in some embodiments, the targetspecific hybridization probes 30 are directed to target sequences 24 of exons, e.g., all exons for a WES panel. In another embodiment, the target sequences 24 may be custom sequences, or disease or allele-specific sequences. The target sequence 24 may be part of a region of interest in a nucleic acid sample, and the target-specific hybridization probe 30 may be designed based on various metrics to be specific for a portion of the region of interest. The probes 30 may include a subset of probes that target copynumber variable regions to enrich the library for sequences associated with copy number variants. Thus, the sequenced library includes sequencing data for the copy number variable regions that can be used to make copy number variant calls.

[0036] As provided herein, a probe (e.g., a target-specific hybridization probe 30) is an oligonucleotide, such as a single-stranded nucleic acid molecule. The target-specific hybridization probe 30 may be part of a set or panel of target-specific hybridization probes 30. The target-specific hybridization probes 30 may be 80-120 bases in length, 80-150 bases, 80-100 bases in length, 90-110 bases in length, 100-120 bases in length, etc. In certain embodiments, if the target-specific hybridization probe 30 is 80-120 bases in length, at least 30-50 of the bases of the target-specific hybridization probe are complementary- to the target sequence 24.

[0037] It should be understood that a hybrid capture sequencing reaction may be performed using a set of target-specific hybridization probes 30, wherein different probes are representative of different target sequences 24 in the nucleic acid library722. For example, the set of target-specific hybridization probes 30 may be representative of at least 2000 different target sequences 24, at least 5000 different target sequences 24, at least 10,000 different target sequences 24, and so on. The probes 30 may include probes that capture 15-50 Mb of coding content for an exome panel. Within the probes 30 may be a set of probes representing 100-500 kb (e.g., 350 kb) of coding content for copy number variable regions. In embodiments, the disclosed techniques may include probes to enhance or spike in to an existing target enrichmentIP-2867-PCT|ILUM:0195PCT panel for I i bran- preparation to permit use for targeted variant calling. In embodiments, the disclosed techniques may include kits, compositions, and / or reagents of a preprepared enrichment panel with copy number variable probes already provided. The use of a hybrid capture workflow that includes probes that capture one or more copy number variable regions permits targeted variant calling in the context of target enrichment sequencing as disclosed herein.

[0038] In certain embodiments, the target-specific hybridization probes 30 may have modifications that facilitate separation of bound fragments 22 from the unbound fragments 22. Such modifications may include biotinylation of the probe to facilitate selection via streptavidin (e.g., streptavidin beads). However, it should be understood that the probes as provided herein may be coupled to an affinity binding molecule that is part of a binding pair. For example, biotin and streptavidin, biotin and avidin, or digoxigenin and a specific antibody that binds digoxigenin are examples of specific binding pairs. The affinity’ binding molecule may be an antibody ligand capable of being conjugated to a nucleotide. In certain embodiments, the modification is provided at the 5' or the 3' end of the probe. Further, in other embodiments, the probes may be unmodified. The target-specific hybridization probes 20 may also include unique barcodes or sequences that facilitate identification. Such sequences may be part of a region of the probe 20 that is non-complementary to the target sequence 14. The targetspecific hybridization probes 30 may be in solution or immobilized on a solid support (e.g., an array).

[0039] In the illustrated example, the fragments are generated via bead- linked transposome complexes 50 affixed to a solid surface (such as a bead) through an affinity element. In one example, two populations of annealed first and second transposons (oligonucleotides comprising an annealed double stranded region comprising transposon end sequence and a single stranded region), where in each population the first transposon includes one of two adaptor sequences and an affinity element, such as biotin, at the 5' end. For example, a plurality of biotinylated first and second transposons comprising the two adaptor sequences (e.g., primer sequences such as A14 and Bl 5) are generated. The transposome complexes generate double-stranded fragments 52 that include adaptor ends, and that can be further modified viaIP-2867-PCT|ILUM:0195PCT amplification to generate a pre-enriched indexed library 56 of fragments 22. The preenriched library 56 is enriched using the probes 30 and separated from unbound fragments 22 lacking target sequences 24. After enrichment, the enriched library 58. shown as including adaptor ends, is ready for sequencing using the targeted variant caller as discussed herein.

[0040] FIG. 3 shows an example method 70 of targeted variant calling using the disclosed targeted copy number variant caller for use in conjunction with target enrichment sequencing data. As with the WGS callers, this targeted variant caller approach uses copy number estimations of distinct target regions (e.g., a target region in a copy number variable region) within or around each gene (and pseudogene, when applicable) to predict the gene’s copy number variants and genotypes. The method 70 includes a step of generating target enrichment sequencing data (block 72). The target enrichment sequencing data may be generated using a library prepared using target enrichment probes with a set of spike-in probes that target copy number variable regions (FIG. 2). The target enrichment sequencing data is used to generated alignments to a plurality of target regions (block 74). For example, the alignments may be relative to a reference sequence, such as hg38. After alignments are generated, the depth of coverage, or a number od sequencing reads aligned to each of the target regions is determined. This number is self-normalized to the length of the region to generate the self-normalized number aligned to each region (block 76). The self-normalized number is normalized a self-normalized panel of normal samples to identify a copy number variant for certain copy number variable regions (block 78).

[0041] The method 70 includes a two-step normalization for target enrichment sequencing data of a sample of interest. The two-step normalization includes 1) a selfnormalization (performed separately for the sample of interest and sample(s) in a panel of normal that proceeds along a parallel path) and 2) a normalization of the selfnormalized sample of interest using a self-normalized panel of normal samples. The output of the two-step normalization of the sample of interest can be used for targeted copy number variant calling. Generation of the panel of normal and the generation of the self-normalization data may occur in series and / or in parallel.IP-2867-PCT|ILUM:0195PCT

[0042] As illustrated in FIG. 4, inputs to the workflow include target enrichment sequencing data 100 from a clinical sample of interest 102. The target enrichment sequencing data 100 may be generated using the workflow of FIG. 2, by way of example. The target enrichment sequencing data 100 includes target regions of interest, such as genes of interest. The target regions of interest also include copy number variable regions having differentiating sites that differentiate between a gene and pseudogene in one example. The targeted variant caller first counts the number of reads uniquely aligned to each target region and normalizes the read count by the region's length. Then, the targeted variant caller repeats the same process for a set of normalization regions found within target regions captured by the target enrichment panel referenced in FIG. 2.

[0043] To use the data 100 to generate copy number variant calls, the data includes sequencing reads that align to one or more copy number variable regions. These regions may be associated with genes that, in some individuals, have copy number variability. Alignment to copy number variable regions may include identification in the sequence data 100 of distinguishing nucleotide differences between a gene and gene paralog to accurately distribute or assign reads that have potential to align to both. Accordingly, the capture probes 30 may include a set of probes designed to capture copy number variable regions adjacent to these nucleotide differences such that the fragments pulled down wall include these differences. In certain embodiments, the probes may be exonic and / or intronic. For a particular copy number variable region the probes 30 may tile the region and may overlap to achieve higher depth.

[0044] In one embodiment, the probes may be designed such that the probe sequence does not hybridize directly to regions with characterized nucleotide differences for a given gene / paralog combination. In an embodiment, the probes 30 include probes that hybridize to sequences within 10-100 bases of a characterized nucleotide difference. In some cases, the probes may be designed to hybridize to regions outside of the gene of interest that are also part of the copy number variable region associated with that gene. The probes may target regions within 100, 500, or 1000 bases of a gene start or terminus.IP-2867-PCT|ILUM:0195PCT

[0045] The disclosed techniques may assign highly similar reads of copy number variable regions between a gene and gene paralog (e.g., between SMN1 and SMN2) as a step of targeted variant calling. Accordingly, the disclosed self-normalized depth of coverage may refer to depth of coverage for reads that are assigned to a gene and not the gene paralog. In certain embodiments, the assignment of highly similar sequence reads between the gene and the gene paralog may be as disclosed in US20200087723, US20210166781A1, or US20230053523A1, which are hereby incorporated by reference in their entireties herein.

[0046] The length-normalized read counts for each target region are pooled together with the length-normalized read counts for the normalization regions to correct for sample-specific coverage bias via self-normalization. In one example, normalization may occur as generally discussed in U.S. Publication No. 20210166781 or W02024010812A2, which are hereby incorporated by reference in their entireties herein for all purposes. For example, the length-normalized read counts can be normalized against the normalization regions (e.g., 1000-3000 genomic regions selected as disclosed herein and expected to be consistently diploid across populations). The length-normalized read counts for the normalization regions can be pooled to generate a normalization mean, and a ratio of length-normalized reads aligned to each region of interest relative to the length-normalized mean read count of the normalization regions can provide a self-normalized depth for each region of interest in the sample 102.

[0047] After determining the self-normalized depth for data 100 of the sample 102, Panel of Normals (PON) normalization is performed. PON is a reference-based normalization algorithm that uses additional “normal” samples to determine a baseline level from which to call CN. The PON approach uses normal samples to adjust for variable depth at probe sites (e.g., adjust for probe capture variability) and, in embodiments, converts the read count values into a copy ratio value that can be used for targeted variant calling. A “normal” sample is a sample free from known genetic disease. The PON samples 110 may include one or more samples from one or more different individuals relative to the sample 102 (e.g., that are not the same individual from which the sample 102 is taken). The PON samples 110 may be from individualsIP-2867-PCT|ILUM:0195PCT not related to the individual from whom the sample 102 was obtained. Normal samples 110 may refer to samples that are pre-identified as or estimated to be free from copy number variants in the copy number variable region(s) of interest. Normal samples 110 may have undergone the same library prep workflow, sequencing workflow to generate sequencing data 112, and bioinformatics analysis workflow of the data 112 as the case sample 102. Accordingly, the samples 102, 110 may have been prepared using a same capture probe panel for target enrichment. For example, the PON samples 110 may have undergone a WES workflow if the sample 102 is a WES sample. The PON samples 110 may have undergone sample preparation as part of a same batch and sequencing run as the sample 102 in an embodiment. Alternatively or additionally, the PON samples 110 may refer to stored data and / or normalization information from a previous sample preparation batch and run. In certain cases, the PON analysis may include both real-time (e.g., concurrent) samples and stored samples in an expanded PON set, as discussed below.

[0048] Each individual PON sample 110 may undergo self-normalization as discussed with respect to the sample 102 using the same regions. The self-normalized PON data can then be pooled to generate self-normalized depths for each target region.

[0049] The self-normalized depth for both the target and normalization regions can be pre-computed for all PON samples 110 using the same method as the sample 102. The case sample self-normalized read count value (depth) for each region is normalized by the corresponding PON median read count value (depth) to determine the copy ratio value in haploid space. In one example, the copy ratio value is: case_sample_target / median (pon_sample_l_target, pon_sample_2_target, ... , pon_sample_N_target) where N is the number of samples in the PON (that are identified with clustering as CN neutral)For example, for a case sample with a self normalized depth of 2 divided by the PON median self normalized depth of 2 = 1. What this means is that relative to the normalized regions (for both the sample and the PON), the target region has 2x reads. The PON represents the enrichment bias (i.e. that in CN neutral cases, there are 2x reads). 2 / 2=1IP-2867-PCT|ILUM:0195PCT means that in haploid space we have 1 copy, but in diploid we have 2 (which in this example we assume CN 2 is neutral)

[0050] A region may represent a contiguous genome region in or near a gene of interest, such as a copy number variable gene. In certain cases, the region may be targeted by multiple probes, and different sequencing reads can be aligned to the region.

[0051] Once PON normalization has been applied to the case sample regions to adjust for probe variability, an optional final step of GC correction is performed using the copy ratios across all regions as input to a loess regression model. This step corrects for bias in sequencing coverage due to variable GC content among different regions, although some GC correction is inherent to PON normalization. That is, PON normalization may address GC bias. Accordingly, the illustrated GC correction may be an optional step.

[0052] In an embodiment, a GC bias metric is computed as follows:1. Calculate GC content using a 100 bp wide, per-base rolling window over all regions in the reference genome. Windows containing more than four masked (N) bases in the reference are discarded.2. Calculate the average coverage for each window, excluding any non-PF, duplicate, secondary, and supplementary reads.3. Calculate the average global coverage across the whole panel.4. Group valid windows based on the percentage of GC content, both at individual percentages and five 20% ranges as summary.5. Calculate the normalized coverage for each group by dividing the average coverage for the bin by the global average coverage. Values below 1.0 indicate a lower than expected coverage at the given GC percent or range. Coverages significantly deviating from 1.0 at greater GC values are an expected result.IP-2867-PCT|ILUM:0195PCT6. Calculate dropout metrics as the sum of all positive values of (percentage of windows at GC X-percentage aligned reads at GC X) for each GC < 50% and > 50% for AT and GC dropout.

[0053] GC correction has been described in Benjamini, Y, et al., Summarizing and correcting the GC content bias in high-throughput sequencing, Nucl. Acids Res., 2012, 40 (10): e72. doi: 10.1093 / nar / gks001, and Miller, C A, et al., ReadDepth: A Parallel R Package for Detecting Copy Number Alterations from Short Sequencing Reads, PLoS One., 2011, 6: el6327. doi: 10.1371 / joumal. pone.0016327; the content of each is incorporated by reference in its entirety. In another embodiment, the GC correction process involves both normalization and GC correction, and the output of this process is the normalized and GC corrected depth. A regression model can be used to model the relationship between GC and normalized depth. Then, the correction is applied based on the loess regression prediction value for the target region and the observed target region value.

[0054] This final self-normalized, PON normalized and optionally GC-corrected copy ratio value can then be used for determining final copy numbers (CNs) for the target regions by either rounding to the nearest integer copy number value or using a Gaussian mixture model (GMM) with pre-defined parameters (shift, prior, mean, and sd) or as generally discussed in U.S. Publication No. 20210166781. These parameters are derived from the PON or prior knowledge about frequency of CNV events at relevant sites in the population. In one example, the technique may be used to generate copy number genotypes of HBA1 / 2. SMN1 / 2, CYP2D6, and CYP2B6 in the sample 102.

[0055] The main step in the prototy pe that differs between a WGS workflow and the disclosed PON / WES workflows is the total CN estimation for the target regions. The WGS mode takes the joint pileup depth of the target regions and performs only selfnormalization and gc correction. However, for the case of targeted callers for hybrid capture or target enrichment w orkflows, the use of a panel of normals corrects for the biases introduced by probe capture. This step in not necessary' in WGS, because no probes are used, and coverage in WGS is much more uniform than WES. PON mode used in systems disclosed herein can be used to generate the panel(s) of normals, whichIP-2867-PCT|ILUM:0195PCT is a panel of per-target self-normalized joint pileup depths (normalization region depths, and polymorphic site allele frequencies) across several CN-neutral samples. The PON could also be the raw joint pileup depths and self-normalization can occur when processing the case sample. Note that not all samples in a PON will be CN-neutral for every target, so a CN neutral panel fde can be used (indicating which samples are CN neutral for which targets). This information can be computed via clustering techniques or provided by the user if known in advance. The WES mode takes the joint pileup depth of the target regions for a case sample and performs self-normalization, then PON normalization (and optionally gc correction). The PON normalization step may involve target-factor normalization (dividing the case target self-normalized depth by the target PON median self-normalized depth) and singular value decomposition (SVD)+ logistical regression (LR). In an embodiment, the LR is a QR decomposition of the top 5 components of SVD. This is done per target and predicted values are computed from only the CN neutral samples in the PON. The observed / case value is then divided by the predicted value. Once the total CN estimation is computed, the adjusted depths are processed further by the per-gene targeted calling components (i.e. using differentiating sites to determine final CNs and genotypes).

[0056] FIG. 5 show s a comparison between a set of 3000 copy number constant 2000bp normalization regions used in WGS (left) compared to 2743 normalization regions used in a WES panel (right). To identify the 2743 normalization regions, the total regions of the WES panel with copy number variable spike-ins (middle) w ere assessed for potential for normalization. The normalization regions were selected from the target regions of an enrichment panel as follows: regions that are greater than lOOObp and that are not in the following categories:1. in chrX, chrY2. in the UCSC segdups track (https: / / genome.ucsc.edu / cgi- bin / hgTrackUi?g=genomicSuperDups)3. in masked analysis regions (e.g. associated with false duplications in the reference genome, e.g., hg38, used for alignments)IP-2867-PCT|ILUM:0195PCT4. in regions where the hg38 reference FASTA base pair includes an N base (representing any of A or G or C or T)As illustrated, the WES 2743 normalization regions have a more distributed GC content relative to the WGS regions, which skew towards lower GC.

[0057] FIG. 5 shows results of the targeted copy number variant caller from one run having 100% concordance for most genes analyzed and >98% concordance across all genes. The concordance shown is between WES genotype / copy number results and WGS or Coriell genotype / copy number results. Any samples with no call from WGS or no information provided from Coriell were omitted from the evaluation on a per gene basis. For CYP2D6, one sample made a discordant star allele genotype call, however, the copy number as determined by the methodology' described herein was concordant.

[0058] FIGS. 7-8 show copy number results for HBA and SMN using the targeted copy number variant call with concordance between WES genotype / copy number results (“Targeted Caller Copy Number” on the x-axis) and WGS or Coriell genotype / copy number results (“Truth Copy Number” on the y-axis). The results shown are for an in- run normalization with a predefined PON (known CN neutral information provided). FIG. 6 shows 100% concordance across 63 samples tested in one run. Not all genotypes are represented, but most interpretations are included except for Hb Bart’s which is lethal and HbH which is rare in the population. Samples tested were in the IkG 3202 cohort. FIG. 7 shows 100% concordance across 82 samples tested in one run. Concordance shown for SMN1 and SMN2 separately. Samples tested were in either the IkG 3202 cohort or from Coriell.

[0059] FIG. 9 shows concordance shown for downsampled coverage levels between WES genotype / copy number results and WGS or Coriell genotype / copy number results. Any samples with no call from WGS or information provided from Coriell w ere omitted from the evaluation on a per gene basis. Concordance is maintained across all genes at >=70x coverage. The discordant samples made differing star allele genotype calls, however, the copy number calls as determined by the methodology described herein were concordant. Initial validation data was at high coverage (average coverage >IP-2867-PCT|ILUM:0195PCT300x), so BAMs were downsampled to test the limit of detection. WES data may be generated at 80-100x coverage, and the WES targeted copy number variant caller maintains its performance at this range of coverage.

[0060] While the disclosed methods may be used in conjunction with predefined PONs with characterized copy numbers for each target of interest, PON normalization may also occur with uncharacterized samples. For example, normalization may be supported against in-run samples or customer-generated samples that are not 100% copy number normal for each target. FIG. 10 shows an example analysis for a PON that includes noncopy number normal samples showing outliers for certain intergenic HBA regions indicating both high and low normalized depth.

[0061] In some embodiments, potential PON samples may be evaluated in real-time, e.g., during a sequencing run, for suitability for inclusion in the PON. That is, for a multiplexed sample run in which multiple different samples, prepared under a same sample preparation workflow, are generating sequencing data, each sample may be normalized against at least some of the other samples in the run. Thus, each individual sample can be normalized against itself as a first step in the targeted caller and against at least some of the other samples as a second step to adjust each target region in each sample for variable depth at probe sites and convert the read count values for each target region into a copy ratio value from which copy number variants can be identified.

[0062] FIG. 11 is a flow diagram show ing an example method illustrate the QC checks performed when running a case sample. The illustrated method may be used to identify samples for in-run normalization. The method may be applied to each target region for each sample to identify suitable clusters representative of a copy number normal region. For each case sample, a concordance correlation coefficient (CCC) is computed vs. the same region in each different PON sample using their self-normalized depth of the normalization regions only. The median of all pairwise comparison CCCs of each PON-to-case pairwise comparison (of self normalized depth across all normalization regions) for the target region is used to determine whether the PON matches the case sample.IP-2867-PCT|ILUM:0195PCT

[0063] In some cases, one sample in a multi-sample run may fail or not be matched to the PONs generated from the run. However, other samples in the run may be well- matched. Thus, an individual sample can fail.

[0064] After passing the initial screen, DPGMM or GMM modeling is used to cluster PON self-normalized depths, and heuristics can be used to select a CN neutral cluster in the PONs. In one example, the largest cluster is used first as CN neutral. Then, the selected cluster is confirmed or reselected based on a heuristic that utilizes b-allele. After cluster selection, each case sample self-normalized target depth is normalized against the selected PON cluster median for that target in target factor normalization to generate a ratio of the sample to baseline (median normalized sequencing depth / self- normalized read count within the cluster).

[0065] The disclosed techniques also permit cross-run normalization that use single run PON for future, independent runs. In-run normalizations may include predefined PON samples in same run as clinical / case samples to help identify CN neutral clusters. Additionally or alternatively, the methods may be used to identify CN normal samples within a run to use for normalization of all samples in that run

[0066] In certain cases, cross-run and in-run normalization may be combined for increased accuracy. Thus, the PON may be an enhanced PON of multiple-run selfnormalized depths per target. A case-run “PON’' of self-normalized depth is created, and target-factor (TF) normalization is performed on each run, on per-target basis. Clustering (e.g., hierarchical clustering, Dirichlet process Gaussian mixture model (DPGMM) ) can be used for CN-neutral identification, and / or a user-provided CN- neutral input can be used. SVD+LR is computed using enhanced PON + case-run “PON’' TF normalized depth to generate predicted values and the case TF normalized depth as observed values. To get final per-target CN estimation, each case observed value is divided by the SVD+LR predicted value for every target. Once many runs / batches are TF normalized, the batch-enrichment bias has been factored out which enables combining cross-run / batch data for SVD+LR. The SVD+LR step then takes all PON data and for each target using only the previously identified CN neutral samples, computes the SVD, takes the top 5 components, and performs the QR decompositionIP-2867-PCT|ILUM:0195PCT to get the predicted valuesFIG. 12 shows results for internal runs normalized against an enhanced PON.

[0067] As discussed, for each target, clustering, such as DGMM clustering, can be used to group each sample’s self-normalized depth into clusters. Each cluster is expected to be composed of samples with the same total copy number. For each target, determine which cluster of samples appears to be copy number neutral. Start by looking at the largest cluster. Compare the allele frequencies (BAFs) at polymorphic sites in each cluster to the frequencies that are expected for the neutral CN. “Peaks'’ every 1 / N steps, where N = neutral copy number, may be observed. The largest cluster that also best fits the expected frequencies is chosen as the CN neutral cluster.

[0068] FIG. 13 shows example total copy number data (from 1KGP data) for different genes showing different copy number clusters for copy number variable genes and other gene regions. For example, copy number variable genes cyp2d6, cyp2b6, smn are all shown as having the neutral copy number cluster (CN4) as the largest cluster in a random sampling of the population. Other gene regions are shown with dominant clusters of CN2, such as upstream or regions of hba.

[0069] FIGS. 14-15 show different examples of cluster heuristics. At each site, (count allele 1) / (count allele 1 + count allele 2), where allele 1 is the most common, allele 2 is the second most common. This results in allele frequencies greater than 0.5. The polymorphic site heuristic is as follows: Assign each BAF to the closest “peak"’. For each peak, find the median absolute difference between all BAFs and the peak. Sum all median absolute differences together and divide by the number of peaks. Do this for the second largest cluster if its cluster size >= 25% of samples. If the largest cluster is CN 3 (and the second largest is CN 4), trying to fit the CN 3 data to the CN 4 model yields a higher heuristic value. Comparing the heuristics of the largest and second largest clusters, the method will choose the largest cluster as CN neutral. The heuristic enforces a sample cluster size limitation because small clusters are more likely to be clustered incorrectly by the DPGMM.

[0070] Embodiments of Sequencing SystemsIP-2867-PCT|ILUM:0195PCT

[0071] FIG. 16 illustrates a diagram of an environment in which a targeted copy number variant caller system can operate in accordance with one or more implementations. The following paragraphs describe the targeted copy number variant caller system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 16 illustrates a schematic diagram of a computing system 300 in which a targeted copy number variant caller system 316 operates in accordance with one or more implementations. As illustrated, the computing system 300 includes one or more server device(s) 302 connected to a user client device 308, a local device 318, and a sequencing device 314 via a network 312. The network 312 can comprise any suitable network over which computing devices can communicate.

[0072] As shown in FIG. 16, the computing system 300 includes the server device(s) 312, In various implementations, the server device(s) 302 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 302 receive various data from the sequencing device 314, such as data from a sample genome and / or sequence reads. The server device(s) 302 may also communicate with the user client device 308. In particular, the server device(s) 302 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 308.

[0073] As shown, the server device(s) 302 includes a sequencing application 310. In general, the sequencing application 310 analyzes the data (such as call data) received from the sequencing device 314 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 310 can receive raw data from the sequencing device 314 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 310 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.

[0074] As also shown, the sequencing application 310 includes the targeted copynumber variant caller system 316. As described below, the targeted copy number variant caller system 316. can identify copy number variants from target enrichmentIP-2867-PCT|ILUM:0195PCT sequencing data of a nucleic acid sample. For example, in some embodiments, the targeted copy number variant caller system 316 receives sequence reads obtained from a nucleic acid sample. The targeted copy number variant caller system 316 further counts sequence reads which align to diploid regions in a human genome within the nucleic acid sample. The targeted copy number variant caller system 316 further counts sequence reads which align to a target region of one or more target regions adjacent or within a copy number variable region. The targeted copy number variant caller system 316 can determine a copy number variant genotype based on the count of the sequence reads which align to a target region of the one or more target regions.

[0075] Moreover, while the targeted copy number variant caller system 316 is described being implemented on the server device(s) 302, as part of the sequencing application 310, in some implementations, the targeted copy number variant caller system 316 is implemented by (such as located entirely or in part) on the user client device 308, the sequencing device 314, and / or the local device 318. As mentioned, in some implementations, the targeted copy number variant caller system 316 is implemented by one or more other components of the computing system 300, such as the sequencing device 314. In particular, the targeted copy number variant caller system 316 can be implemented in a variety of different ways across the server device(s) 302, the network 312. the user client device 308, the local device 318, and the sequencing device 314.

[0076] As further shown in FIG. 16, the computing system 300 includes the user client device 308. In various implementations, the user client device 308 can generate, store, receive, and send digital data. In particular, the user client device 308 can receive the data from the sequencing device 314. As further illustrated, the user client device 308 includes a sequencing application 310 . The sequencing application 310 may be a web application or a native application stored and executed on the user client device 308 (e.g., a mobile application, desktop application, or web application). The sequencing application 310 can receive data from the sequencing application 310 and / or Targeted copy number variant caller system 316. For example, the user client device 308 can receive variant call files and / or alignment files from the sequencing application 310 .IP-2867-PCT|ILUM:0195PCT

[0077] The sequencing application 310 can also include instructions that (when executed) cause the user client device 308 to receive data from the Targeted copy number variant caller system 316 and present data from the sequencing device 314 and / or the server device(s) 302. Furthermore, the sequencing application 310 can instruct the user client device 308 to display data for variant calls, such as nucleobase calls or an indication of a copy number variant. Indeed, the user client device 308 can display nucleobase call results for a genome sample and / or an indication of a predicted copy number variant.

[0078] As further shown in FIG. 16, the computing system 300 includes the sequencing device 314. In various implementations, the sequencing device 314 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 314 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 314, More particularly, the sequencing device 314 receives and analyzes, within nucleotide- sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 314 utilizes SBS to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 3112, in some implementations, the sequencing device 314 bypasses the network 3112 and communicates directly with the user client device 308.

[0079] As further depicted in FIG. 16, in some implementations, the server device(s) 302 includes a distributed collection of servers, where the server device(s) 302 include several server devices distributed across the network 312 and located in the same or different physical locations. For instance, the server device(s) 302 can be implemented, in whole or in part, on the local device 318. To illustrate, the local device 318 may implement the sequencing application 310 and / or the copy number detection system 316. Further, the server device(s) 302 and / or the local device 318 can include a content server, an application server, a communication server, a web-hosting server, or another ty pe of server.IP-2867-PCT|ILUM:0195PCT

[0080] The user client device 308 illustrated in FIG. 16 can include various ty pes of client devices. For example, in some implementations, the user client device 308 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 308 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0081] Though FIG. 16 illustrates the components of the computing system 300 communicating via the network 312, in certain implementations, the components of computing system 300 can also communicate directly with each other, bypassing the network 312. For instance, in some implementations, the user client device 308 communicates directly with the sequencing device 314. Additionally, in some implementations, the user client device 308 communicates directly with the copy number detection system 316 and / or the server device(s) 302. In some implementations, the user client device 308 communicates directly with the local device 318. Moreover, the copy number detection system 316 can access one or more databases housed on or accessed by the server device(s) 302 or elsewhere in the computing system 300.

[0082] FIG. 17 is a block diagram of an exemplary7server device 302 that may be used in connection with the illustrative sequencing system 300 of FIG. 16. The server device 302 may be configured to determine a copy number variant genotype in a nucleic acid sample. The general architecture of the server device 302 depicted in FIG. 17 includes an arrangement of computer hardware and software components. The server device 302 may include many more (or fewer) elements than those shown in FIG. 17. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 302 includes a processing unit 319, a network interface 320, a computer readable medium drive 330, an input / output device interface 340, a display 350, and an input device 360, ah of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity7to one or more networks or computing systems. The processing unit 310 may7thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350 via the input, / output device interface 340. The input / output deviceIP-2867-PCT|ILUM:0195PCT interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0083] The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the server device 302. The memory 370 may store a reference genome 373, such as for use by the sequencing application 310. The memory 370 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0084] For example, in one embodiment, the memory 370 includes a sequencing application 310, which may include copy number detection system 316. The system 316 can perform the methods disclosed herein. In addition, memory 370 may include or communicate with the data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining a copy number variant genotype in a nucleic acid sample of the present disclosure, such the sequencing reads, the estimated copy number(s), and the variant call (for example, the detection of a copy number variant) determined.

[0085] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloudIP-2867-PCT|ILUM:0195PCT computing environment facilitates modification or annotation of sequence data byusers. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0086] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory' stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0087] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0088] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0089] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory' comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assayIP-2867-PCT|ILUM:0195PCT instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0090] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to. an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0091] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, dataIP-2867-PCT|ILUM:0195PCT for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or serv er may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0092] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0093] In some embodiments, a hardware platform for providing a computational environment comprises a processor (i.e. , CPU) wherein processor time and memory layout such as random access memory (i.e., RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer netw ork.IP-2867-PCT|ILUM:0195PCT

[0094] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (i. e.. grid technology) which may run a variety of operating systems m a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary' of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.

[0095] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Overview of System for Biological or Chemical Analysis

[0096] With the preceding discussion in mind, examples and embodiments described herein may be used in various biological or chemical processes and systems for academic analysis, commercial analysis, or other analysis. More specifically, examples described herein may be used in various processes and systems where it is desired to detect an event, property, quality, or characteristic that is indicative of a designated reaction. Bioassay systems such as those described herein may be configured to perform a plurality of designated reactions that may be detected individually or collectively. For example, bioassay systems may be used to sequence a dense array of nucleic acid features through iterative cycles of enzymatic manipulation and image acquisition. In some examples, nucleic acids can be attached to a surface and amplified. Examples of such amplification are described in U.S. Pat. No. 7,741,463, entitled “Method of Preparing Libraries of Template Polynucleotides,” issued June 22, 2010, the disclosure of which is incorporated by reference herein, in its entirety; and / or U.S. Pat. No. 7,270,981, entitled “Recombinase Polymerase Amplification,” issued September 18, 2007, the disclosure of which is incorporated by reference herein, in its entirety.IP-2867-PCT|ILUM:0195PCT

[0097] Components that are used in the bioassay systems may include one or more microfluidic channels that deliver reagents or other reaction components to a reaction site. The reaction sites may be randomly distributed across a substantially planar surface; or may be patterned across a substantially planar surface. Each of the reaction sites may be imaged to detect light from the reaction site. The signals indicating photons emitted from the reaction sites and detected by image sensors may provide illumination values. These illumination values may be combined into an image indicating photons as detected from the reaction sites. These images may be further analyzed to identify compositions, reactions, conditions, etc., at each reaction site.Examples of Fluidics Devices and Fluid Flow Paths

[0098] With the preceding in mind, FIG. 18 illustrates a schematic diagram of an example of a system 1100 that may be used to perform an analysis on one or more samples of interest. A controller 1114 of the present example includes a user interface 1206, a communication interface 1208, one or more processors 1210, and a memory 1212 storing instructions executable by the one or more processors 1210 to perform various functions including the disclosed implementations. User interface 1206, communication interface 1208. and memory 1212 are electrically and / or communicatively coupled to the one or more processors 1210. User interface 1206 may be adapted to receive input from a user and to provide information to the user associated with the operation of system 1100 and / or an analysis taking place. User interface 1206 may include a touch screen, a display, a keyboard, a speakers, a mouse, a track ball, and / or a voice recognition system.

[0099] Communication interface 1208 is adapted to enable communication between system 1100 and a remote systems e.g., computers via a networks e.g., the Internet, an intranet, a local-area network LAN. a wide-area network WAN. a coaxial-cable network, a wireless network, a wired network, a satellite network, a digital subscriber line DSL network, a cellular network, a Bluetooth connection, a near field communication NFC connection, etc. Some of the communications provided to the remote system may be associated with analysis results, imaging data, etc. generated or otherwise obtained by system 1100. Some of the communications provided to systemIP-2867-PCT|ILUM:0195PCT1100 may be associated with a fluidics analysis operation, patient records, and / or a protocols to be executed by system 1100.

[0100] The one or more processors 1210 and / or system 1100 may include one or more of a processor-based system(s) or a microprocessor-based system(s). In some implementations, the one or more processors 1210 and / or system 1100 includes one or more of a programmable processor, a programmable controller, a microprocessor, a microcontroller, a graphics processing unit (GPU), a digital signal processor (DSP), a reduced-instruction set computer (RISC), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a field programmable logic device (FPLD), a logic circuit, and / or another logic-based device executing various functions including the ones described herein.

[0101] Memory 1212 may include one or more of a semiconductor memory', a magnetically readable memory, an optical memory, a hard disk drive (HDD), an optical storage drive, a solid-state storage device, a solid-state drive (SSD), a flash memory, a read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), a random-access memory (RAM), a non-volatile RAM (NVRAM) memory’, a compact disc (CD), a compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a Blu-ray disk, a redundant array of independent disks (RAID) system, a cache and / or any other storage device or storage disk in yvhich information is stored for any duration (e.g., permanently, temporarily, for extended periods of time, for buffering, for caching).

[0102] In some implementations, the sample may include one or more clusters of nucleotides (e.g., DNA) that have been linearized to form a single stranded DNA (sstDNA). In the implementation shoyvn, system 1100 is configured to receive a flow cell cartridge assembly 1102 including a flow cell assembly 1103 and a sample cartridge 1104. System 1100 includes a flow cell receptacle 1122 that receives flow cell cartridge assembly 1102, a vacuum chuck 1124 that supports floyv cell assembly 1103, and a floyv cell interface 1126 that is used to establish a fluidic coupling between system 1100 and flow cell assembly 1103. Flow cell interface 1126 may include one or more manifolds. System 1100 further includes a sipper manifold assembly 1106. aIP-2867-PCT|ILUM:0195PCT sample loading manifold assembly 1108, and a pump manifold assembly 1110. System 1100 also includes a drive assembly 1112, the controller 1114, an imaging system 1116, and a waste reservoir 1118. Controller 1114 is electrically and / or communicatively coupled to drive assembly 1112 and to imaging system 1116; and is configured to cause drive assembly 1112 and / or the imaging system 1116 to perform various functions for performing the techniques disclosed herein.

[0103] In the present example, flow cell assembly 1103 includes a flow cell 1 128 having a channel 1130 and defining a plurality of first openings 1132, which are fluidically coupled to the channel 1130 and arranged on a first side 1134 of the channel 1130. Flow cell 1128 further includes a plurality of second openings 1136 fluidically coupled to the channel 1130 and arranged on a second side 1 138 of the channel 1130. Fluid may thus flow through flow cell 1128 via the channel 1130. While the flow cell 1128 is shown including one channel 1130, flow cell 1128 may include two or more channels 1130. Flow cell assembly 1103 also includes a flow cell manifold assembly 1140 coupled to flow cell 1 128 and having a first manifold fluidic line 1142 and a second manifold fluidic line 1144. Flow cell manifold assembly 1140 may be in the form of a laminate including a plurality of layers as discussed in more detail below.

[0104] In the implementation shown, first manifold fluidic line 1142 has a first fluidic line opening 1146 and is fluidically coupled to each of the first openings 1132 of flow cell 1128; and second manifold fluidic line 1144 has a second fluidic line opening 1148 and is fluidically coupled to each of the second openings 1136. As shown, flow cell assembly 1103 includes gaskets 1150 coupled to flow cell manifold assembly 1140 and fluidically coupled to fluidic line openings 1146, 1148. In some implementations where flow cell 1128 includes a plurality of channels 1130, flow cell manifold assembly 1140 may include additional fluidic lines 1152 that couple first fluidic line openings 1146 to a single manifold port 1154. In such implementations, a single gasket 1150 may be coupled to flow cell manifold assembly 1140 that surrounds the manifold port 1154 and is in fluidic communication with a plurality of channels 1130. In operation, flow cell interface 1126 engages with corresponding gaskets 1150 to establish a fluidic coupling between system 1100 and flow cell 1128. TheIP-2867-PCT|ILUM:0195PCT engagement between flow cell interface 1126 and gaskets 1150 reduces or eliminates fluid leakage between flow cell interface 1126 and flow cell 1128.

[0105] In the implementation shown, first manifold fluidic line 1142 has a portion 1156 that is substantially parallel to a longitudinal axis 1158 of channel 1130; and second manifold fluidic line 1144 has a portion 1160 that is substantially parallel to longitudinal axis 1158 of channel 1130. Additionally, first manifold fluidic line 1142 is shown being at least partially adjacent a first end 1162 of flow cell 1128 and spaced from a second end 1164 of flow cell 1128; and second manifold fluidic line 1144 is shown being at least partially adjacent second end 1164 of flow cell 1128 and spaced from first end 1162. Other arrangements of manifold fluidic lines 1142, 1144 may prove suitable, however.

[0106] In the implementation shown, system 1100 includes a sample cartridge receptacle 1166 that receives sample cartridge 1104 that carries one or more samples of interest e.g., an analyte. System 1100 also includes a sample cartridge interface 1 168 that establishes a fluidic connection with sample cartridge 1104. Sample loading manifold assembly 1108 includes one or more sample valves 1170. Pump manifold assembly 1110 includes one or more pumps 1172. one or more pump valves 1174, and a cache 1176. Valves 1170, 1174 and pumps 1172 may take any suitable form. Cache 1176 may include a serpentine cache and may temporarily store one or more reaction components during, for example, bypass manipulations of the system 1100. While cache 1176 is shown being included in pump manifold assembly 1110, cache 1176 may alternatively be located elsewhere (e.g., in sipper manifold assembly 1106 or in another manifold downstream of a bypass fluidic line 1178, etc.).

[0107] Sample loading manifold assembly 1108 and pump manifold assembly 1110 flow one or more samples of interest from sample cartridge 1104 through a fluidic line 1180 toward flow cell cartridge assembly 1102. In some implementations, sample loading manifold assembly 1108 may individually load or address each channel 1130 of flow cell 1128 with a respective sample of interest. The process of loading channel 1130 with a sample of interest may occur automatically using system 1100. As shown in FIG. 17. sample cartridge 1104 and sample loading manifold assembly 1108 areIP-2867-PCT|ILUM:0195PCT positioned downstream of flow cell cartridge assembly 1102. In the implementation shown, sample loading manifold assembly 1108 is coupled between flow cell cartridge assembly 1 102 and pump manifold assembly 1110. To draw a sample of interest from sample cartridge 1104 and toward pump manifold assembly 1110, sample valves 1170, pump valves 1174, and / or pumps 1172 may be selectively actuated to urge the sample of interest toward pump manifold assembly 1110. Sample cartridge 1104 may include a plurality of sample reservoirs that are selectively fluidically accessible via the corresponding sample valves 1170. To individually flow the sample of interest toward channel 1130 of flow cell 1128 and away from pump manifold assembly 1110, sample valves 1170, pump valves 1174, and / or pumps 1172 may be selectively actuated to urge the sample of interest toward flow cell cartridge assembly 1102 and into respective channels 1130 of flow cell 1128.

[0108] Drive assembly 1112 interfaces with sipper manifold assembly 1106 and pump manifold assembly 1110 to flow one or more reagents that interact with the sample within flow cell 1128. In some scenarios, a reversible terminator is attached to the reagent to allow a single nucleotide to be incorporated onto a growing DNA strand. In some such implementations, one or more of the nucleotides has a unique fluorescent label that emits a color when excited. The color (or absence thereof) is used to detect the corresponding nucleotide. In the implementation shown, imaging system 1116 excites one or more of the identifiable labels (e.g., a fluorescent label) and thereafter obtains image data for the identifiable labels. The labels may be excited by incident light and / or a laser and the image data may include one or more colors emitted by the respective labels in response to the excitation. The image data (e.g., detection data) may be analyzed by system 1100.

[0109] After image data is obtained, drive assembly 1112 interfaces with sipper manifold assembly 1106 and pump manifold assembly 1110 to flow another reaction component e.g., a reagent through flow cell 1128 that is thereafter received by waste reservoir 1118 via a primary waste fluidic tine 1182 and / or otherwise exhausted by system 1100. Some reaction components may perform a flushing operation that chemically cleaves the fluorescent label and the reversible terminator from the sstDNA. The sstDNA may then be ready for another cycle.IP-2867-PCT|ILUM:0195PCT

[0110] The primary' waste fluidic line 1182 is coupled between pump manifold assembly 1110 and waste reservoir 1118. In some implementations, pumps 1172 and / or pump valves 1174 of pump manifold assembly 1110 selectively flow the reaction components from flow cell cartridge assembly 1102, through fluidic line 1180 and sample loading manifold assembly 1108 to primary' waste fluidic line 1182. Flow cell cartridge assembly 1102 is coupled to a central valve 1184 via flow cell interface 1126. Central valve 1184 is coupled with flow cell interface 1126 via a fluidic line 1185. An auxiliary waste fluidic line 1186 is coupled to central valve 1184 and to waste reservoir 1118. In some implementations, auxiliary waste fluidic line 1186 receives excess fluid of a sample of interest from flow cell cartridge assembly 1102, via central valve 1184, and flows the excess fluid of the sample of interest to waste reservoir 1118 when back loading the sample of interest into flow cell 1128.

[0111] Sipper manifold assembly 1106 includes a shared line valve 1188 and a bypass valve 1190. Shared line valve 1188 may be referred to as a reagent selector valve. Central valve 1184 and the valves 1 188, 1190 of sipper manifold assembly 1 106 may be selectively actuated to control the flow of fluid through fluidic lines 1192, 1194, 1196. Sipper manifold assembly 1106 may be coupled to a corresponding number of reagent reservoirs 1198 via reagent sippers 1200. Reagent reservoirs 1198 may contain fluid e.g.. reagent and / or another reaction component. In some implementations, sipper manifold assembly 1106 includes a plurality of ports. Each port of sipper manifold assembly 1106 may receive one of the reagent sippers 1200. Reagent sippers 1200 may be referred to as fluidic lines. Some forms of reagent sippers 1200 may include an array of sipper tubes extending downwardly along the z-dimension from ports in the body of sipper manifold assembly 1106. Reagent reservoirs 1 198 may be provided in a cartridge, and the tubes of reagent sippers 1200 may be configured to be inserted into corresponding reagent reservoirs 1198 in the reagent cartridge so that liquid reagent may be drawn from each reagent reservoir 1198 into the sipper manifold assembly 1106.

[0112] Shared line valve 1188 of sipper manifold assembly 1106 is coupled to central valve 1184 via shared reagent fluidic line 1192. Different reagents may flow through shared reagent fluidic line 1192 at different times. In some versions, whenIP-2867-PCT|ILUM:0195PCT performing a flushing operation before changing between one reagent and another, pump manifold assembly 1110 may draw wash buffer through shared reagent fluidic line 1 192, central valve 1184, and flow cell cartridge assembly 1102.

[0113] Bypass valve 1190 of sipper manifold assembly 1106 is coupled to central valve 1184 via dedicated reagent fluidic lines 1194, 196. Each of the dedicated reagent fluidic lines 1194, 1196 may be associated with a single reagent. The fluids that may flow through dedicated reagent fluidic lines 1194, 1196 may be used during sequencing operations and may include a cleave reagent, an incorporation reagent, a scan reagent, a cleave wash, and / or a wash buffer.

[0114] Bypass valve 1 190 is also coupled to cache 1 176 of pump manifold assembly 1110 via bypass fluidic line 1178. One or more reagent priming operations, hydration operations, mixing operations, and / or transfer operations may be performed using bypass fluidic line 1178. The priming operations, the hydration operations, the mixing operations, and / or the transfer operations may be performed independent of flow cell cartridge assembly 1102. Thus, the operations using bypass fluidic line 1178 may occur during, for example, incubation of one or more samples of interest within flow cell cartridge assembly 1102. That is, shared line valve 1188 may be utilized independently of bypass valve 1190 such that bypass valve 1190 may utilize bypass fluidic hne 1178 and / or cache 1 176 to perform one or more operations while shared line valve 1 188 and / or central valve 1184 simultaneously, substantially simultaneously, or offset synchronously perform other operations.

[0115] Drive assembly 1112 includes a pump drive assembly 1202 and a valve drive assembly 1204. Pump drive assembly 1202 may be adapted to interface with one or more pumps 1172 to pump fluid through flow cell 1128 and / or to load one or more samples of interest into flow cell 1128. Valve drive assembly 1204 may be adapted to interface with one or more of the valves (1170, 1174, 1184, 1 188, 1190) to control the position of the corresponding valves (1170, 1174, 1184, 1188, 1190).

[0116] While the foregoing examples are provided in the context of a system 1100 that may be used in nucleotide sequencing processes, the teachings herein may also beIP-2867-PCT|ILUM:0195PCT readily applied in other contexts, including in systems that perform other processes (i. e. , other than nucleotide sequencing procedures). The teachings herein are thus not necessarily limited to systems that are used to perform nucleotide sequencing processes.

[0117] In an embodiment, notifications related to CNV calling are provided on the display 534 or communicated via the communications circuitry 527 to a remote device or a cloud server.

[0118] The term copy number variation, copy number variant, or CNV may refer to variation in the number of copies of a nucleic acid sequence present in a test sample in comparison with the copy number of the nucleic acid sequence present in a reference sample. In certain embodiments, the nucleic acid sequence is 1 kb or larger. In some cases, the nucleic acid sequence is a whole chromosome or significant portion thereof. A copy number variant may refer to a sequence in which copy -number differences are found by comparison of reads (e.g., a number of reads, a self-normalized or normalized number of reads) of a region a nucleic acid sequence of interest in test sample with a copy number neutral number of reads of the region of the nucleic acid sequence of interest. For example, the level of the nucleic acid sequence of interest in the test sample is compared to that present in a qualified sample. Copy number variants / variations include deletions, including microdeletions, insertions, including microinsertions, duplications, multiplications, and translocations. CNVs encompass chromosomal aneuploidies and partial aneuploidies.

[0119] A CNV call may be based on a variation from a threshold. A threshold or threshold value may refer to a number that is used as a cutoff to characterize a sample such as a test sample containing genetic material from an organism suspected of having a medical condition. The threshold may be compared to a parameter value to determine whether a sample giving rise to such parameter value suggests that the organism has the medical condition. In certain embodiments, a qualified threshold value is calculated using a qualifying data set and serves as a limit of diagnosis of a SNV or CNV. If a threshold is exceeded by results obtained from methods disclosed herein, a subject can be diagnosed with a SNV or CNV.IP-2867-PCT|ILUM:0195PCT

[0120] A sequencing coverage refers to a percentage of bases in a reference sequence covered by the mapped or aligned sequence reads. A set of sequence reads are mapped to a reference genome at the various genomic regions. The total percentage of target bases within the reference genome to which sequenced reads are mapped is quantified as the coverage of the genome. The average depth of sequencing coverage is the ratio of the number of reads, e.g., scaled scaled by read length, to the total referenced genome length. The read depth may be normalized to a mean depth across a genome. Coverage may be determined using the Lander / Waterman equation:C = LN / GC stands for coverageG is the haploid genome lengthL is the read lengthN is the number of readsA Poisson distribution can be used to model any discrete occurrence given an average number of occurrences. The probability7function is the following: P(Y=y) = (Cy x e- C) / y! y is the number of times a base is readC stands for coveragePoisson distribution is used to compute the probability' of a base being sequenced a certain number of times. SNP callers may require at least four calls at a base position to call SNPs.

[0121] A sequence read may refer to a sequence obtained from a fragment of a nucleic acid sample. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T. C, or G) of the sample portion. It may be stored in a memory' device and processed as appropriate to determine whether itIP-2867-PCT|ILUM:0195PCT matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a sequence read is a contiguous DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region, e.g., that can be aligned and specifically assigned to a chromosome or genomic region or gene. In some embodiments, the plurality of sequence reads comprises sequence reads that are about 100 base pairs to about 1000 base pairs in length each. The plurality of sequence reads can comprise paired-end sequence reads and / or single-end sequence reads. The plurality of sequence reads may be generated by targeted sequencing techniques, such as WES. Sequence data may include at least a thousand or at least a million individual sequence reads.

[0122] It should be understood that the biological sample sequencing data may be in the form of raw data, base call data, or data that has gone through primary or secondary analysis. Further, it should be understood that CNVs may be identified as being part of a gene, an intragenic region, etc. It should also be understood that CNV detection may be associated with duplicate or deleted sequences. Accordingly, CNV detection may represent duplicate copies of a nucleic acid region, such as a region including one or more genes. In one embodiment, CNVs are duplicate or deleted genomic regions of at least Ikb in size.

[0123] An alignment may refer to comparing a sequence read to a reference sequence and thereby determining whether the reference sequence contains the read sequence. If the reference sequence contains the read sequence, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment to generate aligned reads is implemented by a computer algorithm, as it would be impossible to align the thousands to millions of reads generated by an individual sample (e.g., sample 102) in a sequencing run in a reasonable time period for implementing the methods disclosed herein. Non limiting examples of alignment methods include global alignments (such as Needleman- Wunsch algorithm), local alignments, dynamic programming (such as Smith-WatermanIP-2867-PCT|ILUM:0195PCT algorithm), heuristic algorithms or probabilistic methods, progressive methods, iterative methods, motif finding or profile analysis, genetic algorithms, simulated annealing, pairwise alignments, multiple sequence alignments. In some cases, the alignment may generate an alignment score. The alignment method may include Burrows-Wheeler Aligner (BWA), 1SAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW. CUSHAW2, CUSHAW2-GPU, drFAST, ELAND. ERNE. GNUMAP, GEM, GensearchNGS. GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh &NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX. QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3- dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and / or ZOOM.

[0124] An alignment score is a score indicating a similarity of two sequences determined using an alignment method. In some implementations, an alignment score accounts for number of edits (e.g., deletions, insertions, and substitutions of characters in the string). In some implementations, an alignment score accounts for a number of matches. In some implementations, an alignment score accounts for both the number of matches and a number of edits. In some implementations, the number of matches and edits are equally weighted for the alignment score. As provided herein, highly similar regions may have alignment scores above a threshold, e.g., above 90% similarity. In one embodiment, a threshold percent identity or identity score (e.g.. above 85% identity, above 90% identity based on alignment) is used to select two regions as being similar and / or as having corresponding bins. In one embodiment, two highly similar regions may be similar over a minimum length (e.g., at least 500 or at least 1000 bases).

[0125] The output from the alignment module is a Binary Alignment Map (BAM, e.g., binary version of a Sequence Alignment Map (SAM)) file along with a mapping quality score (MAP A), which quality score reflects the confidence that the predicted and aligned location of the read to the reference is actually where the read is derived. BAM files contain a header section and an alignment section. The header contains information about the entire file, such as sample name, sample length, andIP-2867-PCT|ILUM:0195PCT alignment method. Alignments in the alignments section are associated with specific information in the header section. The alignments contain read name, read sequence, read quality, alignment information, and custom tags. The read name includes the chromosome, start coordinate, alignment quality, and the match descriptor string. The alignments section includes one or more of Read group, which indicates the number of reads for a specific sample, Barcode tag, which indicates the demultiplexed sample ID associated with the read. Single-end alignment quality. Paired-end alignment quality. Edit distance tag, which records the Levenshtein distance between the read and the reference, and Amplicon name tag, which records the amplicon tile ID associated with the read.

[0126] Output from variant calling is a Variant Call Format (VCF). The VCF file header includes the VCF file format version and the variant caller version and lists the annotations used in the remainder of the file. The VCF header also includes the reference genome file and BAM file. The last line in the header contains the column headings for the data lines. Each of the VCF file data lines contains information about 1 variant. VCF files may include the following information:The chromosome of the reference genome.The single-base position of the variant in the reference chromosome. For SNVs, this position is the reference base with the variant. For indels, this position is the reference base immediately preceding the variant.The rs number for the SNP obtained from dbSNP.txt, if applicable. If multiple rs numbers exist at this location, the list is delimited by semicolons. If a dbSNP entry’ does not exist at this position, a missing value marker ('.') is used.The reference genotype. For example, a deletion of a single T is represented as reference TT and alternate T. An A to T single nucleotide variant is represented as reference A and alternate T.IP-2867-PCT|ILUM:0195PCTThe alleles that differ from the reference read. For example, an insertion of a single T is represented as reference A and alternate AT. An A to T single nucleotide variant is represented as reference A and alternate T.A Phred-scaled quality' score assigned by the variant caller. Higher scores indicate higher confidence in the variant and lower probability of errors. For a quality score of Q, the estimated probability of an error is 10-(Q / l 0). For example, the set of Q30 calls has a 0. 1% error rate. Many variant callers assign quality scores based on their statistical models, which are high in relation to the error rate observed.

[0127] A reference sequence refers to any particular known genome sequence, whether partial or complete, of any organism or yarns which may be used to reference identified sequences from a subject. In one example, the reference sequence is that of a full length human genome. Such sequences may be referred to as genomic reference sequences. In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hgl9. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as yvell as chromosomes, sub- chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.

[0128] A sample, target sample, or biological sample may refer to a sample typically derived from a biological fluid, cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for SNVs or CNVs. In certain embodiments the sample comprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (e g., surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (e.g., patient), the assays can be used to SNVs orIP-2867-PCT|ILUM:0195PCTCNVs in samples from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, fdtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are ty pically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (e.g., namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein.

[0129] The sample can comprise cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. The sample can be obtained directly from a subject. The sample can be generated from another sample obtained from a subject. The other sample can be obtained directly from the subject or the other sample can be generated from another sample obtained from the subject. The computing system can store the plurality of sequence reads in its memory.Copy number variable regions

[0130] Segmental duplications are hotspots for structural variants and gene recombinant variants (for example, gene conversion). Segmental duplications can occur for genes with highly homologous gene family members or pseudogenes. The term “copy number variation”, “copy number variable”, or “copy number variant” herein may refer to variation in the number of copies of a nucleic acid sequence present in a test sample of interest in comparison with an expected copy number. For example, for certain human genes, the expected copy number of autosome sequences (and X chromosome sequences in females) is two. Other organisms may have different expected copy numbers according to their genomic structure. Copy number variationIP-2867-PCT|ILUM:0195PCT may be the result of duplication or deletion. In certain embodiments, copy number variants refer to sequences of at least Ikb that are duplicated or deleted. In one embodiment, copy number variants may be at least a single gene in size. In another embodiment, copy number variants may be at least 140bp, 140-280bp, or at least 500bp. The sequencing data generated by sequencing the biological sample may analyzed to characterize any copy number variation after being normalized as generally discussed herein.

[0131] It should be understood that any copy number variant may be identified by the present techniques. The disclosed techniques use target enrichment sequencing data, e.g., data captured from samples that have undergone a hybrid capture sample preparation and sequencing library preparation workflow. In embodiments that use capture probes, at least a set of the capture probes are designed to hybridize to sequences within copy number variable regions. As disclosed herein, a copy number variable region refers to a region that includes exons and / or introns of a gene associated with copy number variability.

[0132] In some embodiments, copy number variable regions may include regions of the genome that encompass the following genes: CYP21A2. ABCC6, ABCD1, ACTB, ACTG1, ACTN4. ADAMTSL2, ADIPOR1. AFG3L2. AGK. ALG1, ALMS1. AMY1A, AMY1B, AMY1CANKRD11, ANOS1, AP4S1, ARMC4, ARSE, ASNS, ATAD3A, B3GAT3, BCAP31, BDP1, BMPR1A, BRAF, BRCA1, C2, CACNA1C, CALM1, CD46, CEP290, CFH, CFH, CFH, CHEK2, CISD2, CLCNKA, CLCNKB, CORO1A, COXIO, CP, CRYBB2, CSF2RA. CUBN, CUBN, CYCS, CYP11B1, CYP21A2, CYP2B, CYP2D, DCLRE1C, DHFR, DICER1, DIS3L2, DNAH1 1, DNAH11, DNM1, DSE, DU0X2, DUX4, EGLN1, ELK1, ELMO2, ERCC6, ESPN, EYS, F8, FANCD2. FANCD2, FARE FHL1, FLG, FLNC, FOXD4, FXN, GBA, GBAP1, GH1, GJA1, GK, GLUD1, GLUD1. GOSR2, GUSB, HBA1, HBA2, HNRNPA1, HPS1, HSPD1, HYDIN, IDS, IFT122, IGLL1, KANSL1, KCTD1, KIF1C, KRAS, KRT14, KRT16, KRT17, KRT6A, KRT6B, KRT6C, LEFTY2, LRP5, LRP5, MAT2A, MIDI, MOCS1, MSN, MSX2, MY05B, NCF1, NEB, NECAP1, NEFH, NF1. NF1, NF1, NOTCH2, NXF5, OCLN, OTOA, PARN, PBX1, PIGA, PIGN, PIK3CA. PIK3CD, PKD1. PKP2. PMS2, PMS2, PMS2, PNPT1, POLH.IP-2867-PCT|ILUM:0195PCTPRODH, PRODH, PROS1, PRPS1, PRSS1, PTEN, RAD21, RBM8A, RBPJ, RDX, RMND1. RNF216, RNF216, RPL15, RPS17, SALL1. SBDS, SDHA. SHOX, SLC25A15, SLC25A15, SLC33A1. SLC6A8, SMN1, SMN2. SOX2, SPTLC1. SRD5A3, SRP72, STAT5B, STRC, SYT14, TARDBP, TBL1XR1, TBX20, TIMM8A, TPM3, TPMT, TRAPPC2, TRIP11, TTN, TUBA1A, TUBB2A, TUBB2B, TUBB3, TUBB4A, TUBG1, TYR, UBA5, UBE3A, UNC93B1, USP18, VPS35, VWF, WRN, XI AP, ZEB2, or ZNF341.

[0133] The disclosed techniques may permit CNV calling for genes with gene paralogs, such as SMN, HBA, CYP2D6, or CYP2B6. Alignments to these are identified from a set of alignments performed on thousands to millions of sequence reads that include other sequences for other genes or genome regions. Sequence differences relative to a reference sequence may be used to distinguish between reads that align to SMN 1 or SMN2 for normalized copy number determination of 0, 1, 2, 3 etc., and, for SMNl / SMN2-aligned reads (and, in embodiments, ambiguously aligned reads) that have a probability of being aligned to both of SMN 1 or SMN2. Thus, a computational alignment may be performed to generate aligned reads using a programmed computer with alignment parameters that are set according to a particular alignment technique. However, even reads that are aligned to a reference sequence may not be identical. In one example, the sequence differences represent a nucleotide difference at corresponding or aligned positions in the alignments between a sequence read of a sample of interest and the reference sequence.

[0134] For example, alpha globin coded for by two genes (a-globin genes. HBA1 and HBA2) on chromosome 16. Each person needs four functional HBA genes (two from each parent) to make enough a-globin for the body's hemoglobin to work normally. Different forms of a-thalassemia occur if one or more of these genes are defective. If one gene is defective, then a person is a "‘silent’" carrier of the a- thalassemia trait and usually has no signs or symptoms. If two genes are defective, then a person has a-thalassemia trait (also called alpha thalassemia minor) and may have mild anemia. If three genes are defective, then a person has hemoglobin H disease. This can cause moderate to severe anemia. If all four genes are missing, then a person has a-thalassemia major (also called hemoglobin Bart's or hydrops fetalis). This is the mostIP-2867-PCT|ILUM:0195PCT severe type of a-thalassemia. A fetus with this disorder will usually die in the womb or the baby will die soon after birth because the child is unable to make normal hemoglobin to carry oxygen throughout the body.

[0135] More than 90% of a-thalassemia results from the deletion of two or more copies of the a-globin genes (HBA1 and HBA2) on chromosome 16. The HBA1 and HBA2 genes are located within an ~30 kb a-globin gene cluster on chromosome 16, that includes the following alpha globin genes and (pseudogenes) from telomere to centromere in this order: HBZ, (HBZP1) HBM, (HBAP1), HBA2, HBA1, HBQ1. The coding sequences of HBA1 and HBA2 are identical with divergent sequences located in the introns and 5'- and 3'-untranslated regions. In addition, the deletion of the HS-40 major hypersensitive site, which is located 40 kb upstream of the HBZ gene in the promoter region, affects RNA expression of both HBA1 and HBA2, thereby causing an a-thalassemia trait in heterozygotes. The disclosed techniques may be used to distinguish between a normal genotype and silent carrier or minor trait genotypes based on copy number. In an embodiment, the disclosed techniques use probes that hybridize to a plurality of target regions around the HBA1 and HBA2 genes to generate data for copy number variant calling of HBA1.

[0136] Spinal muscular atrophy (SMA) is an autosomal recessive, neuromuscular disorder characterized by loss of motor neurons and progressive muscle wasting, often leading to early death. The disorder is caused by a genetic defect in the SMN1 gene, which encodes survival of motor neuron (SMN) protein, a protein expressed in all eukaryotic cells and necessary for the survival of motor neurons. Lower levels of the protein result in loss of function of neuronal cells in the anterior horn of the spinal cord and subsequent system-wdde muscle wasting (atrophy). A person is affected with SMA if the person only has defective copies of the SMN 1 gene. A person is a carrier of SMA if the person has one chromosome containing at least one normal copy of the SMN1 gene and at least one chromosome containing no normal copies of the SMN 1 gene (i.e., either no copies of SMN1 or only defective copies of SMN1). A small amount of SMN protein can be produced from a gene similar to SMN1 called SMN2. Several different versions of the SMN protein are produced from the SMN2 gene, but only one version (called isoform d) is full size and fully functional. The other versions are smaller andIP-2867-PCT|ILUM:0195PCT may be easily broken down. The full-size protein made from the SMN2 gene is identical to the protein made from SMN1; however, much less full-size SMN protein is produced from the SMN2 gene compared with the SMN1 gene. SMN1 and SMN2 genes are nearly identical and encode the same protein. The sequence difference between the two is a single nucleotide in exon 7 which is thought to be an exon splice enhancer. It is thought that gene conversion events may involve the two genes, leading to exchanges of sequence between SMN1 and SMN2.

[0137] Sequencing coverage describes the average number of sequencing read counts that align to. or "cover," known reference bases. The coverage level often determines whether variant discovery can be made with a certain degree of confidence at particular base positions. At higher levels of coverage, each base is covered by a greater number of aligned sequence reads, so base calls can be made with a higher degree of confidence. Reads are not distributed evenly over an entire genome, simply because the reads will sample the genome in a random and independent manner. Therefore many bases will be covered by fewer reads than the average coverage, while other bases will be covered by more reads than average. This is expressed by the coverage metric, which is the number of times a genome has been sequenced (the depth of sequencing). The disclosed embodiments address noise in sequencing coverage due to bias.

[0138] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. For example, “a processor” can include distributed processing between multiple processors.

[0139] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof.IP-2867-PCT|ILUM:0195PCTAny listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2. 3, 4, or 5 articles, and so forth.

[0140] This written description uses examples to enable any person skilled in the art to practice the disclosed embodiments, including making and using any devices or systems and performing any incorporated methods. The patentable scope is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal languages of the claims.

[0141] It is to be understood that the subject matter described herein is not limited in its application to the details of construction and the arrangement of components set forth in the description herein or illustrated in the drawings hereof. The subject matter described herein is capable of other implementations and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. As used herein, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is explicitly stated. Furthermore, references to “one example” are not intended to be interpreted as excluding the existence of additional examples that also incorporate the recited features. The use of “including,”IP-2867-PCT|ILUM:0195PCT“comprising,” or “having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0142] When used in the claims, the term “set” should be understood as one or more things which are grouped together. Similarly, when used in the claims “based on” should be understood as indicating that one thing is determined at least in part by what it is specified as being “based on.” Where one thing is required to be exclusively determined by another thing, then that thing will be referred to as being “exclusively based on” that which it is determined by.

[0143] Unless specified or limited otherwise, the terms “mounted.” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, “connected” and “coupled” are not restricted to physical or mechanical connections or couplings. Also, it is to be understood that phraseology and terminology used herein with reference to device or element orientation (such as, for example, terms like “above,” “below,” “front,” “rear,” “distal,” “proximal,” and the like) are only used to simplify description of one or more examples described herein, and do not alone indicate or imply that the device or element referred to must have a particular orientation. In addition, terms such as “outer” and “inner” are used herein for purposes of description and are not intended to indicate or imply relative importance or significance.

[0144] It is to be understood that the above description is intended to be illustrative, and not restrictive. For example, the above-described examples (and / or aspects thereof) may be used in combination with each other. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the presently described subject matter without departing from its scope. While the dimensions, types of materials and coatings described herein are intended to define the parameters of the disclosed subject matter, they are by no means limiting and instead illustrations. Many further examples will be apparent to those of skill in the art upon reviewing the above description. The scope of the disclosed subject matter should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to whichIP-2867-PCT|ILUM:0195PCT such claims are entitled. In the appended claims, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Moreover, in the following claims, the terms “first,” “second.” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects. Further, the limitations of the following claims are not written in means — plus-function format and are not intended to be interpreted based on 35 U.S.C. §112(f) paragraph, unless and until such claim limitations expressly use the phrase “means for” followed by a statement of function void of further structure.

[0145] The following claims recite aspects of certain examples of the disclosed subject matter and are considered to be part of the above disclosure. These aspects may be combined with one another.

Claims

IP-2867-PCT|ILUM:0195PCTWHAT IS CLAIMED IS;1. A normalization method, comprising: generating target enrichment sequencing data from a plurality of samples in a multiplexed sequencing run, wherein the target enrichment sequencing data is generated from each sample using sample fragments enriched using a target enrichment sequencing panel having probes specific for copy number variable regions; generating alignments to a plurality of target regions for each sample of the target enrichment sequencing data; determining a self-normalized number of sequence reads aligned to each of the plurality of target regions for each sample of the plurality of samples, wherein the self-normalized number is determined based on normalizing to alignments to a subset of target regions of the plurality of target regions; generating a panel of normals from the target enrichment sequencing data from the plurality of samples by: clustering the self-normalized number of sequence reads for each target region of the plurality of target regions across the plurality of samples; and selecting one cluster for each target region as being copy number neutral to generate the panel of normals; and identifying a copy number variant in an individual sample of the plurality of samples using the panel of normals.

2. The method of claims 1, wherein the plurality of samples are from different individuals.

3. The method of claims 1-2, wherein the plurality of samples have undergone a same sample preparation using a same exome sequencing panel composition relative to one another.IP-2867-PCT|ILUM:0195PCT4. The method of claims 1-3, wherein the copy number variant in the individual sample is identified based on deviation from a baseline established by copy number neutral cluster in a copy number variable target region.

5. The method of claims 1-4, wherein the subset of target regions of the plurality of target regions are not in the copy number variable regions.

6. The method of claims 1-5, comprising determining that the panel of normals is matched to the individual sample based on a correlation coefficient before using the panel of normals.

7. The method of claims 1-6, wherein the copy number variable region is in HBA. SMN, GBA, or CYP2B6, or CYP2D6.

8. A normalization method, comprising: generating sequencing data from an individual sample, wherein the target enrichment sequencing data is generated from using sample fragments enriched using a target enrichment sequencing panel having probes specific for a plurality' of target regions, wherein a first subset of the target regions comprise copy number variable regions; generating alignments to the plurality of target regions using the sequencing data; determining a self-normalized number of sequence reads aligned to each of the plurality of target regions, wherein the selfnormalized number is determined based on normalizing to alignments to a second subset of target regions of the plurality of target regions; normalizing the self-normalized number of sequence reads for each individual target region of the first subset, the panel of normals comprising self-normalized sequence reads for each individual target region of the first subset for one or more different samples; andIP-2867-PCT|ILUM:0195PCT identifying a copy number variant associated with an individual target region in the first subset in the individual sample based on the normalizing.

9. The method of claim 8, wherein the one or more different samples are multiple samples from different individuals.

10. The method of claims 8-9, wherein the one or more different samples have undergone a same sample preparation using a same panel of probes relative to one another.

11. The method of claims 8-11, wherein the copy number variant associated with an individual target region is identified based on a ratio of self-normalized sequence reads aligned to the individual target region in the individual sample relative to the panel of normals.

12. The method of claims 8-11, wherein the self-normalized sequence reads aligned to the individual target region are aligned to a bin that covers at least a portion of a copy number variable gene.

13. The method of claims 8-11, wherein the first subset is nonoverlapping with the second subset.

14. The method of claims 8-11, wherein the panel of normal is at least in part from a different 11 bran preparation batch than the sequencing data.

15. A target enrichment sequencing normalization method, comprising: providing nucleic acid fragments generated from a sample; denaturing the nucleic acid fragments to generate singlestranded nucleic acid fragments; contacting the single-stranded nucleic acid fragments with a plurality of probes under conditions that permit a subset of the singlestranded nucleic acid fragments to hybridize to the plurality of probes,IP-2867-PCT|ILUM:0195PCT wherein the plurality7of probes comprise single-stranded oligonucleotides configured to hybridize to respective different target sequences in the sample, wherein the target sequences comprise one or more sequences associated with a copy number variable gene; separating a hybridized subset of single-stranded nucleic acid fragments from an unhybridized subset of the single-stranded nucleic acid fragments to generate an enriched group of single-stranded nucleic acid fragments from the separated hybridized subset of singlestranded nucleic acid fragments; generating sequencing data comprising a plurality7of sequence reads from the enriched group of single-stranded nucleic acid fragments; determining a first self-normalized number of sequence reads of the plurality7of sequence reads aligned to the copy number variable gene, wherein the first self-normalized number is normalized relative to sequence reads aligned to a set of selected sequences within the target sequences; normalizing the first self-normalized number of sequence reads of the plurality of sequence reads aligned to the copy number variable gene using a second self-normalized number of sequence reads, the second self-normalized number of sequence reads being aligned to the copy number variable gene in a different sample to generate a ratio; and determining a copy number of the copy number variable gene based on the ratio.

16. The method of claim 15, wherein the plurality7of probes comprises a whole exome sequencing panel.

17. The method of claims 15-16, wherein the plurality of probes comprises one or more probes configured to hybridize to the one or more sequences within theIP-2867-PCT|ILUM:0195PCT copy number variable gene and to one or more sequences within a paralog of the copy number variable gene.

18. The method of claim 17, wherein determining the first self-normalized number of sequence reads of the plurality of sequence reads aligned to the copy number variable gene comprises: identifying sequence reads aligned to the copy number variable gene and / or to the paralog of the copy number variable gene; identifying sequence reads supporting one or more differentiating nucleotides between the copy number variable gene and the paralog of the copy number variable gene within the aligned sequence reads; and assigning the sequence reads to the copy number variable gene or the paralog of the copy number variable gene based on the differentiating nucleotides; and normalizing a number of sequence reads assigned to the copy number variable gene.

19. The method of claims 15-18, wherein the set of selected sequences within the target sequences are at least 1000 bases in length.

20. The method of claims 15-18, wherein the plurality of probes comprises a plurality' of copy number variant calling probes configured to hybridize to respective different copy number variant target sequences in a sample, wherein an individual copy number variant calling probe of the plurality is capable of hybridizing to a gene region and a corresponding paralog gene region.

Citation Information

Patent Citations

  • Methods and systems for determining paralogs

    US20200087723A1

  • Methods and systems for diagnosing from whole genome sequencing data

    US20210166781A1

  • Methods and systems for identifying recombinant variants

    US20230053523A1

  • Recombinase polymerase amplification

    US7270981B2

  • Method of preparing libraries of template polynucleotides

    US7741463B2