Method and system for determining copy number polymorphism genotypes

A method for determining HBA1/2 copy number polymorphisms using diploid and target region alignment with a Gaussian model improves accuracy and sensitivity in detecting HBA1/2 variants.

JP2025524474APending Publication Date: 2025-07-30ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575797
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-07
Filing Date
2023-07-05
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing methods struggle to accurately determine HBA1/2 copy number polymorphisms and detect deletions such as α3.7 and α4.2 due to high sequence similarity between the HBA1 and HBA2 genes, leading to ambiguous read alignment and inaccurate genotype determination.

Method used

A computer-implemented method that aligns sequence reads to diploid and target regions of the human genome, applies a mixture Gaussian model to estimate integer copy numbers, and normalizes read counts to improve genotype determination.

Benefits of technology

Enhances the specificity and sensitivity of HBA1/2 copy number polymorphism detection, increasing true positive detection of variants by 20% to 100% and improving variant call accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524474000001_ABST
    Figure 2025524474000001_ABST
Patent Text Reader

Abstract

Disclosed herein are systems, devices, and methods for identifying recombinant variants (such as deletions or duplication variants) of genes such as the HBA1 gene and the HBA2 gene, copy numbers of HBA1 and / or HBA2, and copy number polymorphism genotypes. Further disclosed herein are systems, devices, and methods for detecting one or more single nucleotide variants or indels in the HBA1 / 2 region in a nucleic acid sample.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Citation by reference to any priority application) In the application data sheet filed together with this application, any application in which a foreign or domestic priority claim is identified is incorporated herein by reference pursuant to 37 CFR 1.57.

[0002] This application was filed on July 7, 2022, claims the priority of U.S. Provisional Application No. 63 / 367,888, entitled "METHODS AND SYSTEMS FOR IDENTIFYING GENOTYPES", and incorporates the same herein by reference in its entirety.

[0003] (Field of the Invention) The technology of the present disclosure relates to the field of nucleic acid sequencing. Specifically, the disclosed technology relates to determining the HBA1 / 2 copy number polymorphism genotypes in nucleic acid samples.

Background Art

[0004] Description of Related Art Mutations or copy number variations in HBA1 / 2 can cause α-thalassemia, one of the most common human single-gene diseases in the world, and its carrier frequency is 5% worldwide. Approximately 95% of α-thalassemia cases are caused by gene deletion(s) rather than non-deletion mutations, and the HBA1 gene and the HBA2 gene are at least 97% homologous. Common 1-copy and 2-copy HBA1 / 2 deletions in α-thalassemia include α3.7 deletion, α4.2 deletion, Southeast Asia (SEA) deletion, and Mediterranean (MED) deletion.

[0005] Approximately 95% of alpha thalassemia cases are caused by gene deletions (multiple), rather than non-deletion mutations. For example, FIG. 1A shows potential gene deletions (multiple), non-deletion mutations, and the final phenotypes. Detecting alpha thalassemia variants from standard whole-genome sequencing (WGS) data can be difficult because the high homology between the HBA1 gene region and the HBA2 gene region can sometimes lead to ambiguous read alignment. Determining the genotype of HBA1 / 2 copy number polymorphisms can be complicated by the high sequence similarity observed between the two genes. For example, in some cases, sequence reads of the HBA1 gene or the HBA2 gene may be misaligned to the wrong gene or mapped to both genes with the same confidence, resulting in a decrease in mapping quality. This can lead to inaccurate sequence assembly through the HBA1 gene and the HBA2 gene, and there is a risk that the determination of the copy number of HBA1 and / or HBA2 will be inaccurate.

[0006] Furthermore, in conventional variant detection methods, deletions such as the α3.7 deletion and the α4.2 deletion are difficult to accurately detect because there are boundaries where these deletions enter segmental duplication regions such as regions X, Y, and Z described herein. For example, deletions can occur in regions of segmental duplication. In conventional variant detection methods, sequence reads from segmental duplications may be discarded or not utilized, which can result in the failure to detect deletions such as the α3.7 deletion or the α4.2 deletion. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0007] In one aspect, the present specification discloses a computer-implemented method for determining the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample. In some embodiments, the method includes determining sequence reads from the nucleic acid sample, counting the sequence reads that align to the diploid regions of the human genome within the nucleic acid sample, counting the sequence reads that align to the target regions of one or more target regions that are proximal to the positions of the HBA1 gene and the HBA2 gene of the human genome, and determining the HBA1 / 2 copy number polymorphism genotype based on the count of the sequence reads that align to the target regions of one or more target regions compared to the count of the sequence reads that align to the diploid regions of the human genome.

[0008] In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype includes, for each of the one or more target regions, estimating an integer copy number. In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype includes normalizing the count of the sequence reads that align to each target region by the count of the sequence reads that align to the diploid regions of the human genome to determine a floating-point copy number for each of the one or more target regions.

[0009] In some embodiments, estimating an integer copy number for each of the one or more target regions further includes applying a mixture Gaussian model to the floating-point copy number of the sequence reads that align to each target region. In some embodiments, the mixture Gaussian model includes a predefined shift, a prior value, an average value, or a standard deviation as shown in Table 3.

[0010] In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include a first upstream region upstream of the HBA2 gene and the HBA1 gene. In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome further include a second upstream region upstream of the HBA2 gene and the HBA1 gene. In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include an intergenic region between the HBA2 gene and the HBA1 gene, or a downstream region downstream of the HBA2 gene and the HBA1 gene. In some embodiments, the one or more target regions include the first and second upstream regions upstream of the HBA2 gene and the HBA1 gene, the intergenic region between the HBA2 gene and the HBA1 gene, and the downstream region downstream of the HBA2 gene and the HBA1 gene.

[0011] In some embodiments, sequence reads align to each of the one or more target regions with at least a 30 alignment MAPQ score. In some embodiments, the first upstream region is adjacent to a segmental duplication region X upstream of the HBA2 gene. In some embodiments, the second upstream region corresponds to a region within the α4.2 deletion event. In some embodiments, the second upstream region is adjacent to a segmental duplication region Z upstream of the HBA2 gene. In some embodiments, the intergenic region corresponds to a region within the α3.7 deletion event. In some embodiments, the intergenic region is adjacent to a segmental duplication region Z upstream of the HBA1 gene. In some embodiments, the first upstream region, the second upstream region, the intergenic region, and the downstream region correspond to regions within deletion events in cis for both HBA1 and HBA2.

[0012] In some embodiments, the first upstream region has coordinates chr16:167503-169503 of the reference genome hg38, the second upstream region has coordinates chr16:170263-171875 of the reference genome hg38, the intergenic region has coordinates chr16:174519-175845 of the reference genome hg38, or the downstream region has coordinates chr16:178002-180501 of the reference genome hg38.

[0013] In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype is aaa 3.7 / aa genotype, aaa 4.2 / aa genotype, aa / aa genotype, -a 3.7 / aa genotype, -a 4.2 / aa genotype, -- / aaa 3.7 genotype, -- / aaa 4.2 genotype, -a 3.7 / -a 3.7 genotype, -a 4.2 / -a 4.2 genotype, -a 3.7 / -a 4.2 genotype, -- / aa genotype, -- / a 3.7 genotype, -- / a 4.2 genotype, or determining the -- / -- genotype, includes.

[0014] In other aspects, the present specification discloses a computer-implemented method for detecting one or more single nucleotide variants, or indels, in the HBA1 / 2 region in a nucleic acid sample. In some embodiments, the method includes determining sequence reads from the nucleic acid sample, obtaining sequence reads that align to a site of a single nucleotide variant, or indel, within the HBA1 gene, or the HBA2 gene of the human genome in the nucleic acid sample, counting the sequence reads that store bases corresponding to alternative alleles at the site of the single nucleotide variant, or indel, where counting the sequence reads includes counting sequence reads that align to HBA1 and sequence reads that align to the HBA2 gene, and creating a digital file that includes a variant call corresponding to the single nucleotide variant, or indel, where the variant call is not specific to the HBA1 gene, or the HBA2 gene.

[0015] In some embodiments, the single nucleotide variant, or indel, is HBA2_c.60del, HBA2_c.69C>T, HBA2_c.95+2_95+6delTGAGG, HBA2_c.95+1G>A, HBA1_c.179G>A, HBA2_c.377T>C, HBA2_c.427T>C, HBA2_c.427T>G, HBA2_c.429A>T, HBA2_c. * 92A>G, HBA2_c.428A>C, HBA2_c.314G+A, HBA2_c.379G>A, HBA2_c.179G>A, HBA2_c.75T>G, HBA1_c.96-1G>A, HBA1_c.358C>T, or HBA2_c. * 94A>G.

[0016] In other aspects, the present specification discloses an electronic system for determining the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample. In some embodiments, the electronic system includes a processor configured to execute a method that includes determining sequence reads from a nucleic acid sample, counting the sequence reads that align to a diploid region of the human genome within the nucleic acid sample, counting the sequence reads that align to a target region of one or more target regions that are proximal to the locations of the HBA1 gene and the HBA2 gene of the human genome, and determining the HBA1 / 2 copy number polymorphism genotype based on the count of the sequence reads that align to the target region of the one or more target regions as compared to the count of the sequence reads that align to the diploid region of the human genome.

[0017] In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype includes, for each of the one or more target regions, estimating an integer copy number. In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype includes normalizing the count of the sequence reads that align to each target region by the count of the sequence reads that align to the diploid region of the human genome to determine a floating point copy number for each of the one or more target regions. In some embodiments, estimating an integer copy number for each of the one or more target regions further includes applying a mixture of Gaussian models to the floating point copy number of the sequence reads that align to each target region.

[0018] In other aspects, described herein is an electronic system for detecting one or more single nucleotide variants, or indels, in the HBA1 / 2 region in a nucleic acid sample. In some embodiments, the electronic system includes a processor configured to execute a method that includes determining sequence reads from the nucleic acid sample, obtaining sequence reads that align to a location of a single nucleotide variant, or indel, within the HBA1 gene, or the HBA2 gene, of the human genome in the nucleic acid sample, counting the sequence reads to store bases corresponding to alternative alleles at the location of the single nucleotide variant, or indel, wherein counting the sequence reads includes counting sequence reads that align to the HBA1 gene and sequence reads that align to the HBA2 gene, and creating a digital file that includes a variant call corresponding to the single nucleotide variant, or indel, wherein the variant call is not specific to the HBA1 gene, or the HBA2 gene.

[0019] The features of the examples of this disclosure will become apparent by reference to the following detailed description and drawings. In the drawings, like reference numerals correspond to components that are similar, if not the same. For the sake of brevity, reference numerals or features having the aforementioned functions may or may not be described in connection with other drawings in which they appear.

Brief Description of the Drawings

[0020]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 3A

Figure 3B

Figure 4

[0021] All patents, patent applications, and other publications are hereby expressly incorporated by reference into this specification to the same extent as if each individual publication, patent, or patent application were specifically and individually indicated to be incorporated by reference. All cited references are incorporated by reference in their entirety into this specification for the purposes indicated by the context of the citation in this specification in the relevant portions. However, the citation of any reference should not be construed as an admission that it is prior art to the present disclosure.

[0022] One embodiment of the invention is a targeted gene calling approach for detecting deletion variants and / or non-deletion variants of the HBA1 / 2 gene from sequence reads such as standard whole genome sequencing data. In some embodiments, the use of one or more target regions further described herein provides for the detection of clinically relevant HBA1 / 2 copy number polymorphisms including single copy deletions such as α3.7 and α4.2, and two copy deletions in cis such as SEA. In embodiments of the present disclosure, the determination of multiple haplotypes of copy number genotypes in HBA1 / 2, such as -a 3.7 / aa, representing a heterozygous α3.7 deletion, is provided.

[0023] Overview This specification describes a method and a system for detecting the HBA1 / 2 copy number polymorphic genotype in a nucleic acid sample collected from a subject. In the disclosed system and method for detecting the HBA1 / 2 copy number polymorphic genotype in a nucleic acid sample, it has been found that the specificity and sensitivity regarding the determination of the HBA1 / 2 copy number polymorphic genotype and variant calls in the HBA1 region and / or HBA2 region in the nucleic acid sample are improved. In some embodiments, the disclosed system and method solve the technical problem of inaccurate HBA1 and HBA2 copy number determination due to ambiguous sequence read alignment to the HBA1 gene and HBA2 gene due to high sequence similarity.

[0024] In some embodiments, the disclosed system and method include determining sequence reads from a nucleic acid sample. Once the sequence reads are determined, they can be aligned to a reference genome. The method may further include counting sequence reads that align to the diploid region of the human genome within the nucleic acid sample. For example, the diploid region can be a region that is generally diploid in a nucleic acid sample from a human.

[0025] Next, the methods and systems of the present disclosure can count sequence reads that align to one or more target regions that are proximal to the locations of the HBA1 and HBA2 genes of the human genome. The target regions can include the HBA2 gene and the first and / or second upstream regions upstream of the HBA1 gene, the intergenic region between the HBA2 gene and the HBA1 gene, and / or the downstream region downstream of the HBA2 gene and the HBA1 gene. In some embodiments, as seen in FIG. 1B, the first upstream region 1012, the second upstream region 1027, the intergenic region 1032, and the downstream region 1042 can have substantially genetic positions. In FIG. 1C, in particular, the segmental duplication regions X110, Y111, and Z112 in the vicinity of the HBA2 locus 122 and the HBA1 locus 121 are depicted. Such regions X, Y, and Z are well known and studied by those skilled in the art and are described, for example, in "Molecular basis of α-thalassemia, Blood Cells, Molecules, and Diseases, 70:43-53 (2018)" by Farashi and Harteveld.

[0026] Next, the disclosed systems and methods can determine the HBA1 / 2 copy number polymorphism genotype based on the count of sequence reads that align to each of one or more target regions compared to the count of sequence reads that align to the diploid regions of the human genome. For example, the disclosed systems and methods can estimate an integer copy number for each of the one or more target regions. For example, the disclosed systems and methods can normalize the count of sequence reads that align to each target region by the count of sequence reads that align to the diploid regions of the human genome, such as non-repetitive regions where the diploid copy number is stable within a population, to determine a floating-point copy number for each of the one or more target regions. The disclosed systems and methods can apply a mixture Gaussian model to the floating-point copy numbers of the sequence reads that align to each target region to estimate an integer copy number for each of the one or more target regions.

[0027] The disclosed system and method can improve, for example, by increasing the true positive detection of variants, the specificity of single nucleotide polymorphisms (SNPs) and / or insertions / deletions (indels) associated with the HBA1 / 2 copy number polymorphic genotype, i.e., the proportion of true variants correctly detected, by 20%, 50%, 80%, 100%, or more, for example, due to the HBA1 / 2 copy number polymorphic genotype.

[0028] Definitions Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994), Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined below.

[0029] As used herein, "nucleotide" includes a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are the monomeric units of nucleic acid sequences. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is ribose, and in deoxyribonucleotides (DNA), the sugar is deoxyribose, i.e., a sugar without a hydroxyl group present at the 2'-position of ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Examples of purine bases include adenine (A) and guanine (G), and their modified derivatives or analogs. Examples of pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and their modified derivatives or analogs. The C-1 atom of deoxyribose is bonded to N-1 of pyrimidine or N-9 of purine. The phosphate group can be in mono-, di-, or triphosphate form. It should be further understood that these nucleotides may be natural nucleotides, but non-natural nucleotides, modified nucleotides, or analogs of the aforementioned nucleotides can also be used.

[0030] As used herein, "base" or "nucleic acid base" is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. The nucleic acid bases can be naturally occurring or synthetic. Non-limiting examples of nucleic acid bases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purine substituted with methyl or bromine at the 8-position, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)-alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-mudole, ethenoadenine, and non-naturally occurring nucleic acid bases described in U.S. Patent Nos. 5,432,272 and 6,150,510, and International Publications Nos. 92 / 002258, 93 / 10820, 94 / 22892, and 94 / 24144, and Fasman ("Practical Handbook of Biochemistry and Molecular Biology", pp. 385-394, 1989, CRC Press, Boca Raton, LO) (all of which are hereby incorporated by reference in their entirety).

[0031] The terms "nucleic acid" or "polynucleotide" refer to a deoxyribonucleotide or ribonucleotide polymer in single-stranded or double-stranded form, and include known analogs of natural nucleotides that hybridize to nucleic acids in a manner similar to natural nucleotides such as peptide nucleic acids (PNAs) and phosphorothioate DNA, unless otherwise specified. Unless otherwise stated, a particular nucleic acid sequence includes its complementary sequence. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha-thiotriphosphates for all of the above and 2'-O-methyl-ribonucleotide triphosphates for all of the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0032] As used herein, the term "chromosome" means a genetic carrier of the present invention in living cells, derived from chromatin strands containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.

[0033] "Genome" means the complete genetic information of an organism or virus, expressed as a nucleic acid sequence.

[0034] When used in the present invention, the term "reference genome" or "reference sequence" refers to any particular known genomic sequence, either partial or complete, of an organism or virus that can be used to reference a specified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (ncbi.nlm.nih.gov). In various embodiments, the reference sequence may be significantly larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 5 times larger, or at least about 10 6 times larger, or at least about 10 7 times larger. In one example, the reference sequence is that of a full-length genome. Such a sequence may also be referred to as a genomic reference sequence. For example, the reference sequence may be a reference human genomic sequence such as hg19 or hg38. In another example, the reference sequence may be limited to a particular human chromosome such as chromosome 13. In some embodiments, the reference Y chromosome is the Y chromosome sequence from human genome version hg19. Such a sequence may also be referred to as a chromosomal reference sequence. Other examples of reference sequences include genomes of other species, as well as chromosomes, partial chromosomal regions (such as strands), etc. of any species. In various embodiments, the reference sequence is a common base sequence or other combination derived from multiple individuals. However, for certain applications, the reference sequence may be taken from a particular individual.

[0035] As used herein, the term "nucleic acid sample" refers to a sample typically derived from a biological fluid, cell, tissue, organ, or organism, which contains a nucleic acid or a mixture of nucleic acids that includes at least one nucleic acid sequence to be screened for copy number variations. In certain embodiments, the nucleic acid sample contains at least one nucleic acid sequence suspected of having a copy number variation. Such samples can include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, or fine needle biopsy samples (such as surgical biopsies, fine needle biopsies), urine, ascites, pleural effusion, etc. Samples are often taken from human subjects (such as patients), but samples can also be taken from any mammal, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples can be used directly upon being obtained from a biological source, or they can be used after undergoing a pretreatment that modifies the characteristics of the sample. For example, such pretreatment can include preparing plasma from blood, diluting viscous fluids, etc. Further, pretreatment methods can include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc. When such pretreatment methods are employed on a sample, such pretreatment methods typically leave the nucleic acid(s) of interest remaining in the test sample, and in some cases, its concentration is proportional to the concentration in the untreated test sample (i.e., a sample not treated with such pretreatment method(s)). Such "treated" or "processed" samples are still considered biological "test" samples with respect to the methods described herein.

[0036] The term "read", or "array read" (or sequence determination read) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part, or the whole, of a nucleic acid molecule. Typically, but not necessarily, a read represents a short sequence of contiguous base pairs in a sample. A read may be symbolically represented by the base pair sequence (A, T, C, or G) of the sample portion. To determine whether a read aligns with a reference sequence or meets other criteria, the read may be stored in a memory device and appropriately processed. A read may be obtained directly from a sequencing device or indirectly from stored sequence information regarding the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, to specifically assign by alignment to a chromosome, or genomic region, or gene. For example, a sequence read can be a short string of nucleotides at one or both ends of a nucleic acid fragment (such as 20 - 150 bases) sequenced from the nucleic acid fragment, or the sequencing of the entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained by any method known in the art. For example, sequence reads can be obtained in various ways, such as by the use of sequencing techniques, or the use of probes such as hybridization arrays, or capture probes, or amplification techniques such as polymerase chain reaction (PCR), or linear amplification using a single primer, or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by binding. Sequence reads can be generated using devices such as the MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0037] As used herein, the term "sequencing depth" generally refers to the number of times a locus is covered by sequence reads that are aligned to that locus. The locus may be as small as a nucleotide, or as large as a chromosomal arm, or even as large as the entire genome. Sequencing depth can be expressed as 50×, 100×, etc., where "×" refers to the number of times the locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome, in which case x can refer to the average number of times each locus, or the haploid genome, or the entire genome is sequenced. When the average depth is cited, the actual depths of the different loci included in the dataset vary over a range of values. Ultra-deep sequencing can refer to at least 100× in sequencing depth.

[0038] As used herein, the terms "aligned", "alignment", or "aligning" refer to the process of comparing a read, or tag, to a reference sequence and thereby determining the likelihood that the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be located within the reference sequence, or in certain alternative embodiments, may be mapped to a specific location within the reference sequence. For example, from the alignment of a read to a reference sequence of human chromosome 13, the likelihood that the read is present in the reference sequence of chromosome 13 can be determined. In some cases, the alignment further indicates where the read or tag maps within the reference sequence. For example, if the reference sequence is the entire human genome sequence, the alignment may indicate that the read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A "site" can be a polynucleotide sequence, or a unique position (i.e., chromosome ID, chromosomal position, and orientation) on the reference genome. In some embodiments, the site may be the position of a residue, sequence tag, or segment on the sequence.

[0039] The aligned reads or tags are one or more sequences identified as matches with respect to the order of nucleic acid molecules from the reference genome to a known sequence. The alignment can be done manually, but is typically performed by a computer algorithm because it is not possible to align the reads in a reasonable time period to implement the methods disclosed herein. The matching of sequence reads during alignment can be 100% sequence identity or less than 100% (imperfect match).

[0040] The alignment can be performed by variants of methods, and / or combinations thereof, such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, Gensearch NGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTInvestigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, STORM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0041] As used herein, the term "mapping" refers to the specific assignment of sequence reads to a larger sequence, such as a reference genome, by alignment.

[0042] "Genetic variation" or "gene mutation" refers to a specific genotype present in a particular individual, and in many cases, genetic variations exist in a statistically significant subpopulation of individuals. The presence or absence of genetic variance can be determined using the methods or devices described herein. In certain embodiments, the presence or absence of one or more genetic variations is determined according to the results provided by the methods and devices described herein. In some embodiments, a genetic variation is a chromosomal abnormality (such as aneuploidy), a partial chromosomal abnormality, or a mosaic, etc., each of which will be further detailed herein. Non-limiting examples of genetic variations include one or more deletions (such as microdeletions), duplications (such as microduplications), insertions, mutations, polymorphisms (such as single nucleotide polymorphisms), fusions, repeats (such as short tandem repeats), different methylation sites, different methylation patterns, etc., and combinations thereof. An insertion, repeat, deletion, duplication, mutation, or polymorphism can be of any length, and in some embodiments, can be from about 1 base or base pair (bp) to about 250 megabase pairs (Mb) in length. In some embodiments, the length of an insertion, repeat, deletion, duplication, mutation, or polymorphism is from about 1 base, or base pair (bp) to about 1000 kilobases (kb) (e.g., about 10 bp, 50 bp, 100 bp, 500 bp, 1 kb, 5 kb, 10 kb, 50 kb, 100 kb, 500 kb, or 1000 kb).

[0043] In some cases, the genetic variation is a deletion. In certain embodiments, a deletion is a mutation (such as a genetic abnormality) in which a part of a chromosome or a DNA sequence is missing. A deletion is often a loss of genetic material. Any number of nucleotides may be deleted. A deletion can include the deletion of one or more entire chromosomes, segments of chromosomes, alleles, genes, introns, exons, any non-coding region, any coding region, segments thereof, or combinations thereof. A deletion can include a microdeletion. A deletion can include a single base deletion.

[0044] Genetic mutations are sometimes genetic duplications. In certain embodiments, a duplication is a mutation (such as a genetic abnormality) in which a portion of a chromosome or a DNA sequence is copied and reinserted into the genome. In certain embodiments, a genetic duplication (i.e., a duplication) is a duplication of a region of DNA. In some embodiments, the duplication is a nucleic acid sequence that is repeated, often tandemly, within the genome or chromosome. In some embodiments, the duplication can include a copy of one or more entire chromosomes, chromosomal segments, alleles, genes, introns, exons, any non-coding region, any coding region, segments thereof, or combinations thereof. The duplication can include microduplications. The duplication sometimes includes one or more copies of the duplicated nucleic acid. In some cases, the duplication is characterized in that the gene region is repeated one or more times (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times). The duplication can range, in some cases, from a small region (thousands of base pairs) to an entire chromosome. Duplications frequently occur as a result of errors in homologous recombination or due to retrotransposon events. Duplications are associated with certain types of proliferative diseases. Duplications can be characterized using genomic microarrays or comparative genetic hybridization (CGH).

[0045] Genetic mutations are sometimes insertions. An insertion is sometimes an addition of a nucleic acid sequence of one or more nucleotide base pairs. An insertion is sometimes a microinsertion. In certain embodiments, an insertion includes the addition of a chromosomal segment to the genome, a chromosome, or a segment thereof. In certain embodiments, an insertion includes the addition of an allele, gene, intron, exon, any non-coding region, any coding region, segments thereof, or combinations thereof to the genome or a segment thereof. In certain embodiments, an insertion includes the addition (i.e., insertion) of a nucleic acid of unknown origin to the genome, a chromosome, or a segment thereof. In certain embodiments, an insertion includes the addition of a single base (i.e., insertion).

[0046] Genetic variations may include copy number variations, i.e., variations in the copy number of a nucleic acid sequence present in a test sample as compared to the copy number of the nucleic acid sequence present in a reference sample. In certain embodiments, the nucleic acid sequence is 1 kb or more. In some cases, the nucleic acid sequence is an entire chromosome or a significant portion thereof. Copy number polymorphisms may also refer to nucleic acid sequences in which a difference in copy number has been found by comparison of a nucleic acid sequence of interest in a test sample with an expected level of the nucleic acid sequence of interest. For example, the level of a target nucleic acid sequence in a test sample is compared to that present in a qualified sample. Copy number polymorphisms / copy number variations may include deletions including microdeletions, insertions including microinsertions, duplications, amplifications, and translocations. CNV encompasses aneuploidy and segmental aneuploidy.

[0047] Embodiments of a method and system for determining an HBA1 / 2 copy number polymorphism genotype FIG. 2A is a block diagram schematically illustrating an exemplary method 200 for determining an HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample. In some embodiments, method 200 is executed on a computer. Method 200 may be embodied as a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives of a computing system. For example, server device 3102 shown in FIGS. 3A and 3B and described in more detail below can execute a set of executable program instructions for implementing method 200. When method 200 starts, the executable program instructions are loaded into a memory such as RAM and can be executed by one or more processors of server device 3102. Although method 200 is described with respect to server device 3102 shown in FIG. 3B, this description is illustrative only and not intended to be limiting. In some embodiments, method 200, or a portion thereof, can be executed continuously or in parallel by multiple computing systems.

[0048] As shown in FIG. 2A, an exemplary method 200 for determining the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample can start from start block 210. Method 200 can proceed to block 220, where sequence reads are determined from the nucleic acid sample. Next, the method can proceed to block 230, where the sequence reads are aligned to a reference genome. Next, method 200 can proceed to block 240, where sequence reads aligned to diploid regions of the human genome within the nucleic acid sample are counted. The diploid region can be a non-repetitive region having a stable diploid copy number in the population. Next, method 200 can proceed to decision state 250, where the system can determine whether there are further sequence reads to be counted that align to the diploid regions being counted. If there are additional sequence reads for counting, method 200 returns to block 240 and the method can proceed as described above. If there are no additional sequence reads for counting, method 200 can proceed to block 260, where sequence reads that align to target regions proximate to the positions of the HBA1 and HBA2 genes of the human genome are counted. The method can proceed to decision state 250, where the system can determine whether there are further target regions for sequence read counting. If there are additional target regions, method 200 returns to block 260 and the method can proceed as described above. If there are no additional target regions, method 200 can proceed to process block 280, where the HBA1 / 2 copy number genotype is determined. Process block 280 can be described in more detail with respect to FIG. 2B. Method 200 can end at end block 290.

[0049] FIG. 2B is a block diagram further illustrating process block 280 described above, where the HBA1 / 2 copy number genotype is determined.

[0050] As shown in FIG. 2B, the method of process block 280 for determining the HBA1 / 2 copy number polymorphism genotype can start from start block 2810. The method of process block 280 can proceed to block 2820, where the floating-point copy number of the target region is determined by normalizing the count of sequence reads aligned to the target region by the count of sequence reads aligned to the diploid region. The method of process block 280 can proceed to block 2830, where the estimated copy number of the target region is determined by applying a mixture Gaussian model to the floating-point copy number determined in block 2820. The method of process block 280 can proceed to decision state 2840, where the system can determine whether there are any additional target regions for integer copy number estimation. If there are additional target regions, the method of process block 280 can return to block 2820 and the method can proceed as described above. If there are no additional target regions, the method of process block 280 can proceed to block 2850, where the estimated integer copy numbers of one or more target regions are analyzed. The method of process block 280 can end at end block 2860.

[0051] Determination of sequence reads from a nucleic acid sample In some embodiments, the methods and systems disclosed herein include the step of determining sequence reads from a nucleic acid sample, e.g., block 220 of FIG. 2A. In some embodiments, the sequence reads are generated from a nucleic acid sample obtained from a subject.

[0052] Array reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by binding. Array reads can be generated using instruments such as the MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA). Array reads can be, for example, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000 or more base pairs (bps) in length each. For example, array reads are each about 100 base pairs to about 1000 base pairs in length. Array reads can include paired-end array reads. Array reads can include single-end array reads. Array reads can be generated by whole genome sequencing (WGS). WGS can be clinical WGS (cWGS). Samples can include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, blood samples, biopsy samples, or combinations thereof.

[0053] In some embodiments, array reads are aligned to a reference sequence, such as block 230 of FIG. 2A. For example, array reads obtained from a sample can be aligned to one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene of the reference sequence. As further described herein, array reads can also be aligned to diploid regions of the reference sequence. In some embodiments, a computing system stores a first plurality of array reads in memory. The computing system can load the first plurality of array reads into memory.

[0054] In some embodiments, the array reads are obtained from a digital file that stores the sequencing information. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotating magnetic disk drive, or a solid state drive). In some embodiments, the digital file is stored in the form of a BAM, SAM, CRAM, FASTQ, JSON, or VCF file.

[0055] Counting of Array Reads In some embodiments, the disclosed systems and methods include, for example, as in block 240 of FIG. 2A, counting array reads that align to a diploid region of the human genome within a nucleic acid sample. The diploid region can include a preselected region across the genome of the subject that has been measured to be consistently diploid across the population of nucleic acid samples. In some embodiments, the diploid region is non-repetitive. For example, in some embodiments, the alignment of the array reads to the diploid region is unambiguous. For example, in some embodiments, the array reads align to the diploid region with an alignment MAPQ score of at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90.

[0056] In some embodiments, the diploid region can include about 100, about 500, about 1000, about 2000, about 3000, about 4000, or more pre-selected regions across the subject's genome. In some embodiments, the length of the diploid region is about 100 bp, about 500 bp, about 1000 bp, about 2000 bp, about 3000 bp, about 5000 bp, or more, or a range composed of any of the foregoing values. In some embodiments, the length of the diploid region is about 2 kb. For example, the diploid region can be randomly selected from the genome to obtain stable coverage across the population sample, to infer sequencing depth, and to capture GC bias. The system can determine whether additional sequence reads that align to the diploid region should still be counted, as seen in decision state 250 of FIG. 2A.

[0057] In some embodiments, the disclosed systems and methods include, for example, counting sequence reads that align to a target region of one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome, such as block 260 of FIG. 2A. In some embodiments, counting sequence reads that uniquely align to the target region of one or more target regions. In some embodiments, the sequence reads align to the target region of one or more target regions with an alignment MAPQ score of at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, or at least 90. In some embodiments, the MAPQ median of each of the one or more target regions is about 60. In some embodiments, the system can count sequence reads that align to a first target region and then determine whether any further target regions (such as a second, third, and / or fourth target region) remain, as depicted in decision state 270 of FIG. 2A.

[0058] In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include the HBA2 gene, the first upstream region upstream of the HBA2 gene and the HBA1 gene, the second upstream region upstream of the HBA2 gene and the HBA1 gene, the intergenic region between the HBA2 gene and the HBA1 gene, and the downstream region downstream of the HBA2 gene and the HBA1 gene. In some embodiments, as seen in FIG. 1B, the first upstream region, the second upstream region, the intergenic region, and the downstream region have substantial positions.

[0059] For example, in some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include the first upstream region upstream of the HBA2 gene and the HBA1 gene. In some embodiments, the first upstream region is adjacent to the segmental duplication region X upstream of the HBA2 gene. In some embodiments, the first upstream region has coordinates of approximately chr16:167503-169503 of the reference genome hg38 (available, for example, at GenBank assembly accession GCA_000001405.15).

[0060] In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include the second upstream region upstream of the HBA2 gene and the HBA1 gene. In some embodiments, the second upstream region corresponds to the region within the α4.2 deletion event. In some embodiments, the second upstream region is adjacent to the segmental duplication region Z upstream of the HBA2 gene. In some embodiments, the second upstream region has coordinates of approximately chr16:170263-171875 of the reference genome hg38.

[0061] In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include the intergenic region between the HBA2 gene and the HBA1 gene. In some embodiments, the intergenic region corresponds to a region within the α3.7 deletion event. In some embodiments, the intergenic region is adjacent to the segmental duplication region Z upstream of the HBA1 gene. In some embodiments, the intergenic region has coordinates of approximately chr16:174519-175845 of the reference genome hg38.

[0062] In some embodiments, one or more target regions proximate to the positions of the HBA1 gene and the HBA2 gene in the human genome include the downstream region downstream of the HBA2 gene and the HBA1 gene. In some embodiments, the downstream region is adjacent to the downstream end of the HBA1 gene. In some embodiments, the downstream region has coordinates of approximately chr16:178002-180501 of the reference genome hg38.

[0063] In some embodiments, the first upstream region, the second upstream region, the intergenic region, and the downstream region correspond to regions within deletion events in cis for both HBA1 and HBA2. For example, the deletion event in cis for both HBA1 and HBA2 can be a two-gene deletion such as the Southeast Asian (SEA) deletion or the Mediterranean (MED) deletion.

[0064] Normalization and / or determination of GC-corrected copy number In some embodiments, for example, as in process block 280 of FIG. 2A, the disclosed systems and methods include determining the HBA1 / 2 copy number polymorphism genotype based on the count of sequence reads that align to one or more target regions compared to the count of sequence reads that align to the diploid regions of the human genome.

[0065] In some embodiments, determining the HBA1 / 2 copy number polymorphic genotype comprises determining the normalized count of sequence reads aligned to each of one or more target regions. For example, in block 2820 of FIG. 2B, the count of sequence reads aligned to the target region is normalized by the count of sequence reads aligned to the diploid region. In some embodiments, determining the HBA1 / 2 copy number polymorphic genotype comprises normalizing the sequence read counts (such as of the target region and / or diploid region) by the length of each region. In some embodiments, determining the normalized count of sequence reads aligned to each of one or more target regions comprises normalization using (1a) the depth of sequence reads aligned to each of one or more target regions, (1b) the length of each of one or more target regions, (2a) the depth of sequence reads aligned to the diploid region, and (2b) the length of each of the diploid regions.

[0066] In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype comprises determining, for each of one or more target regions, a floating point copy number by normalizing the count of sequence reads that align to each target region by the count of sequence reads that align to a diploid region of the human genome. For example, in some embodiments, the sequence read counts for each target region (e.g., sequence read counts normalized by the length of the region) are pooled with the sequence read counts for a diploid region (e.g., sequence read counts normalized by the length of the region) that includes approximately 3,000 separate 2 kb regions. In some embodiments, normalizing the count of sequence reads that align to a target region by the count of sequence reads that align to a diploid region can correct for biases in sequencing coverage caused by varying GC content between regions. For example, the count of sequence reads aligned to each of one or more target regions can be corrected for GC content using a sequence that uses (1) the GC content of each of the one or more target regions and (2) the GC content of each of the diploid regions. In some embodiments, for each of one or more target regions, a normalized and / or GC-corrected copy number is determined. In some embodiments, the normalized and / or GC-corrected copy number is a floating point copy number that includes non-integers, such as 1.2, 2.4, etc.

[0067] Determination of the putative integer copy number In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype comprises, for each of one or more target regions, estimating an integer copy number. In some embodiments, estimating an integer copy number for each of one or more target regions further comprises applying a mixture Gaussian model to the floating point copy number of sequence reads that align to each target region. For example, in block 2830 of FIG. 2B, a mixture Gaussian model is applied to the normalized counts of sequence reads aligned to the target region.

[0068] In some embodiments, after normalizing and / or determining the GC-corrected depth, a mixture of Gaussian models (GMM) is used to determine an estimated integer copy number (CN) for each of one or more target regions. In some embodiments, the GMM includes pre-defined parameters such as a shift, a prior value, a mean value, and a standard deviation (sd). In some embodiments, the normalized GC-corrected depth is scaled by a shift value that first corrects for bias in the alignment between the target region and the diploid region. In some embodiments, then, based on the pre-trained mean value, standard deviation, and prior value from the mixture of Gaussian models, the posterior probability of CN = i is calculated for i = 0-6 given the scaled depth. In some embodiments, then, the integer copy number with the highest posterior probability is selected as a candidate for the final integer copy number estimate.

[0069] In some embodiments, estimating the integer copy number includes binning the normalized counts of the array reads using a mixture of Gaussian models. For example, a mixture of Gaussian models can be used to infer the most likely copy number of a target region based on the observed normalized depth signal.

[0070] The estimated integer copy number can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more copies. The mixture Gaussian model may include a one-dimensional mixture Gaussian model. The plurality of Gaussians of the mixture Gaussian model can represent integer copy numbers, for example, 0 to 5, 0 to 6, 0 to 7, 0 to 8, 0 to 9, 0 to 10, 0 to 11, 0 to 12, 0 to 13, 0 to 14, or 0 to 15. For example, the plurality of Gaussians of the mixture Gaussian model can represent integer copy numbers from 0 to 10. The mean of each of the plurality of Gaussians can be the integer copy number represented by the Gaussian. The mean of each of the plurality of Gaussians can be an integer copy number (such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more copies) represented by the Gaussian. The standard deviation of the Gaussian can be, for example, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 or more, or can be about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 or more. The plurality of Gaussians of the mixture Gaussian model can include, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more Gaussians. For example, the plurality of Gaussians of the mixture Gaussian model can include 5 Gaussians.

[0071] To estimate the integer copy number, when the computing system is given the normalized number of array reads aligned to the target region, it can use a mixture Gaussian model and a predetermined posterior probability threshold to determine the copy number. The predetermined posterior probability threshold can be, for example, 0.7, 0.75, 0.8, 0.85, 0.95, or more. In some embodiments, the predetermined posterior probability threshold is 0.95.

[0072] In some embodiments, the mixture Gaussian model (GMM) includes an optimized mixture Gaussian model. In some embodiments, the GMM parameters are trained based on the expectation maximization method. For example, starting with three randomly placed Gaussians (parameters are randomly initialized), the optimized parameters can be trained. The floating-point copy numbers obtained from many nucleic acid samples as described herein can be used as training data for the mixture Gaussian model. For example, for each floating-point copy number x of a given sample, P(x / CN = 1), P(x[CN = 2], and P(x|CN = 3) can be calculated. Next, the sample integer copy number can be re-assigned to CN = k, where the posterior P(CN = kx) is the highest. Then, the parameters of the GMM can be adjusted to fit the points assigned to them. This process can be repeated until the parameters reach convergence. In some embodiments, the converged parameters can be used in the mixture Gaussian model described herein.

[0073] In some embodiments, the mixture Gaussian model includes parameters optimized for each of one or more target regions. In some embodiments, for the first upstream target region, the mixture Gaussian model has a shift of about 1.029, an average value of about 2:1.0 and about 3:1.5 (2, 3), prior values of about 0:0.001, about 1:0.01, about 2:0.987, about 3:0.0005, and about 4:0.0005 (0 - 4), and / or a standard deviation of about 0.062 (2). In some embodiments, for the second upstream target region, the mixture Gaussian model has a shift of about 1.02, an average value of about 2:1.0 and about 3:1.5 (2, 3), prior values of about 0:0.001, about 1:0.015, about 2:0.987, about 3:0.005, and about 4:0.0005 (0 - 4), and / or a standard deviation of about 0.0073 (2). In some embodiments, for the intergenic target region, the mixture Gaussian model has a shift of about 0.966, an average value of about 2:1.0 and about 3:1.476 (2, 3), prior values of about 0:0.012, about 1:0.13, about 2:0.834, about 3:0.023, and about 4:0.0005 (0 - 4), and / or a standard deviation of about 0.0077 (2). In some embodiments, for the downstream target region, the mixture Gaussian model has a shift of about 1.071, an average value of about 2:1.0 and about 3:1.5 (2, 3), prior values of about 0:0.001, about 1:0.01, about 2:0.987, about 3:0.001, and about 4:0.0005 (0 - 4), and / or a standard deviation of about 0.06 (2).

[0074] In some embodiments, the probability of the estimated integer copy number was calculated, for example, as a quality check of the estimated integer copy number. In some embodiments, the estimated integer copy number was determined only if the posterior probability was greater than 0.95 and the p - value of the scaled depth in the Gaussian distribution of the candidate copy number was greater than 0.001. In some embodiments, if any of the one or more target regions does not have an estimated integer copy number passing the quality check, the HBA1 / 2 copy number genotype is not determined.

[0075] In some embodiments, for each of one or more target regions, the estimation of the integer copy number is repeated. For example, in determination state 2840, as described above, the system can determine whether further target regions still need to be analyzed. For example, as described herein, based on determining the normalized GC-corrected floating-point copy number and as described herein, based on applying a mixture Gaussian model, for each of one or more target regions, an estimated integer copy number can be determined.

[0076] Determination of HBA1 / 2 copy number polymorphism genotype In some embodiments, for each of one or more target regions, the estimated integer copy numbers are accumulated and compared to determine the HBA1 / 2 copy number polymorphism genotype. For example, as depicted in block 2850 of FIG. 2B, this system and method can analyze the estimated integer copy number of a target region. For example, in some embodiments, based on the estimated values of the estimated integer copy number for each of four target regions, the HBA1 / 2 copy number genotype is deterministically generated.

[0077] In some embodiments, determining the HBA1 / 2 copy number polymorphism genotype is aaa 3.7 / aa genotype, aaa 4.2 / aa genotype, aa / aa genotype, -a 3.7 / aa genotype, -a 4.2 / aa genotype, -- / aaa 3.7 genotype, -- / aaa 4.2 genotype, -a 3.7 / -a 3.7 genotype, -a 4.2 / -a 4,2 genotype, -a 3.7 / -a 4,2 genotype, -- / aa genotype, -- / a 3.7 genotype, -- / a 4,2 including determining the genotype, or the -- / -- genotype.

[0078] For example, the following table represents the HBA1 / 2 copy number genotype that can be determined based on the estimated integer copy number of each of four target regions (first and second upstream regions, intergenic region, downstream region). In the table below, the interpretation is for research use only (RUO).

[0079]

Table 1

[0080] In some embodiments, the methods and systems disclosed herein further include performing a variant call of the HBA1 / 2 copy number polymorphism. In some embodiments, the variant call includes a copy number genotype that includes two or more copy number alleles.

[0081] In some embodiments, the methods and systems disclosed herein further include creating a digital file that includes the variant call. In some embodiments, the file includes the estimated integer copy number for each of one or more target regions, the floating point copy number for each of one or more target regions, and the copy number genotype. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, for example, a rotating magnetic disk drive or a solid state drive). In some embodiments, the digital file is stored in the form of a BAM, SAM, CRAM, FASTQ, JSON, or VCF file. In some embodiments, the digital file is a VCF file or a JSON file.

[0082] Method for Detecting Variants in the HBA1 / 2 Region In other aspects, methods and systems for detecting one or more single nucleotide variants, or indels, in the HBA1 / 2 region in a nucleic acid sample are disclosed herein. In some embodiments, the method and system determine sequence reads from the nucleic acid sample. For example, the sequence reads can be determined as described above herein with reference to methods and systems for determining HBA1 / 2 copy number polymorphism genotypes.

[0083] In some embodiments, the method and system obtain sequence reads that align to sites of single nucleotide variants, or indels, within the HBA1 gene or the HBA2 gene of the human genome in the nucleic acid sample. For example, the sequence reads can be aligned to a reference genome as described above herein with reference to methods and systems for determining HBA1 / 2 copy number polymorphism genotypes. In some embodiments, the sequence reads are derived from next generation sequencing. In some embodiments, the sequence reads are about 75 bp to about 500 bp in length. In other embodiments, the sequence reads are about 200 bp to about 400 bp in length.

[0084] In some embodiments, the method and system count sequence reads that store bases corresponding to alternative alleles at sites of single nucleotide variants, or indels. In some embodiments, counting sequence reads includes counting both sequence reads (including sites of single nucleotide variants, or indels) that align to the HBA1 gene and sequence reads (including sites of single nucleotide variants, or indels) that align to the HBA2 gene. In some embodiments, sequence read counting can be normalized and GC-corrected as described above herein with reference to methods and systems for determining HBA1 / 2 copy number polymorphism genotypes.

[0085] In some embodiments, the method and system create a digital file that includes variant calls corresponding to single nucleotide variants, or indels (collectively, "small variants"). In some embodiments, small variants are reported when the majority of the sequence reads support the alternative allele. For example, a small variant may be reported when about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, about 60% or more, about 70% or more, about 80% or more, or about 90% or more of the sequence reads covering the small variant store base calls corresponding to the alternative allele at that site compared to the reference allele at that site of the small variant. In some embodiments, a small variant may be reported when one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more sequence reads store the alternative allele at the variant site.

[0086] In some embodiments, sequence reads that contain the alternative allele and sequence reads that store the reference allele are counted. In some embodiments, an integer copy number is estimated for the alternative allele, or variant allele, based on: a) the combined count of sequence reads covering the corresponding positions of small variants in HBA1 and HBA2, b) the count of reads supporting the reference allele, and c) the count of reads supporting the alternative allele.

[0087] In some embodiments, the variant call is not specific to the HBA1 gene, or the HBA2 gene. For example, in some embodiments, the variant call is not assigned to HBA1, or HBA2, or is not phased to one of the candidate haplotypes further described herein. In some embodiments, the small variant can be further away from one sequence read length from one or more of the target regions described herein (such as 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, or more). In some embodiments, making the variant call ambiguous with respect to HBA1, or HBA2, eliminates the need to phase the detected small variant to a candidate haplotype, and in this method and system, there is no need to further analyze the sequence reads to determine whether the small variant is assigned to HBA1, or HBA2, thus enabling the user to advantageously detect one or more single nucleotide variants, or indels, in the HBA1 / 2 region of the nucleic acid sample while using computational power and memory more efficiently. In some embodiments, detecting small variants in a region in an ambiguous manner requires a more complex process, is much less computationally efficient, and in some cases provides lower accuracy, or recall, for the variants of interest, compared to de novo small variant calling, or calling of small variants, and phasing of small variants to a region, or haplotype, which improves computational resource efficiency and enables high accuracy and recall in discovering mutant alleles.

[0088] In some embodiments, variant calls that are ambiguous with respect to HBA1 or HBA2 advantageously enable a user to detect small variants using short-read sequencing. Without being bound by theory, in some embodiments, short-read sequencing reads (such as sequence reads including about 75 to 500 bp) on the HBA1 gene or the HBA2 gene do not store sufficient information to uniquely map small variants, and the user does not necessarily need to know the unique mapping of the variants. In some embodiments, the advantage of making calls for regions that are ambiguous is that the user can avoid the need to perform a more extensive sequencing assay, such as a long-read sequencing assay. The information needed can be obtained from the same whole-genome sequencing (WGS) assay that was used for variant calls in the rest of the genome.

[0089] In some embodiments, after an ambiguous variant call is made for HBA1 or HBA2, the placement of single nucleotide variants or indels in the HBA1 gene or the HBA2 gene can be confirmed by orthogonal (long-read) sequencing methods known to those of skill in the art. For example, if a single nucleotide variant or indel is detected by a method that is not specific to the HBA1 gene or the HBA2 gene, additional sequencing such as an orthogonal method is used to confirm the variant call and / or to phase the variant to a region.

[0090] In some embodiments, the single nucleotide variants or indels include the variants listed in the following table.

[0091] [Table 2]

[0092] In some embodiments, the methods and systems disclosed herein further include creating a digital file that includes variant calls. In some embodiments, the file includes, for each single nucleotide variant, or indel, a reference to the small variant, a count of the sequence reads that support the alternative allele, and a count of the sequence reads that support the reference allele. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotational magnetic disk drive, or a solid state drive). In some embodiments, the digital file is stored in the form of a BAM, SAM, CRAM, FASTQ, JSON, or VCF file. In some embodiments, the digital file is a VCF file or a JSON file.

[0093] Embodiments of the sequencing system FIG. 3A illustrates a diagram of an environment in which an HBA1 / 2 copy number detection system can operate according to one or more implementations. In the following paragraphs, the HBA1 / 2 copy number detection system will be described with respect to an exemplary implementation and a diagram of an example depicting the implementation. For example, FIG. 3A illustrates a schematic diagram of a computing system 3000 in which an HBA1 / 2 copy number detection system 3106 operates according to one or more implementations. As illustrated, the computing system 3000 includes one or more server devices (plural) 3102 connected to a user client device 3108, a local device 3118, and a sequencing device 3114 via a network 3112. The network 3112 can include any suitable network over which computing devices can communicate.

[0094] As shown in FIG. 3A, computing system 3000 includes server device(s) 3102. In various implementations, server device(s) 3102 can generate, receive, analyze, store, and transmit digital data, such as nucleic acid base calls or data of sequenced nucleic acid polymers. In some implementations, server device(s) 3102 receive various data, such as data from a sample genome and / or sequence reads, from sequencing device 3114. Further, server device(s) 3102 can communicate with user client device 3108. In particular, server device(s) 3102 can transmit data regarding sequence reads, direct nucleic acid base calls, nucleic acid base calls, and / or sequencing metrics to user client device 3108.

[0095] As shown, server device(s) 3102 includes sequencing application 3110. Generally, sequencing application 3110 analyzes data (such as call data) received from sequencing device 3114 or other locations to determine the nucleic acid base sequence of a nucleic acid polymer. For example, sequencing application 3110 can receive raw data from sequencing device 3114 and determine the nucleic acid base sequence for a sample genome or nucleic acid segment. In some implementations, sequencing application 3110 determines the sequence of nucleic acid bases in DNA and / or RNA segments, or oligonucleotides.

[0096] As further illustrated, the array determination application 3110 includes an HBA1 / 2 copy number detection system 3106. As described below, the HBA1 / 2 copy number detection system 3106 can determine the HBA1 / 2 copy number polymorphic genotype in a nucleic acid sample. For example, in some embodiments, the HBA1 / 2 copy number detection system 3106 receives sequence reads obtained from a nucleic acid sample. The HBA1 / 2 copy number detection system 3106 further counts the sequence reads that align to the diploid regions of the human genome within the nucleic acid sample. The HBA1 / 2 copy number detection system 3106 further counts the sequence reads that align to the target regions of one or more target regions that are proximal to the positions of the HBA1 gene and the HBA2 gene of the human genome. The HBA1 / 2 copy number detection system 3106 can determine the HBA1 / 2 copy number polymorphic genotype based on the count of the sequence reads that align to the target regions of one or more target regions compared to the count of the sequence reads that align to the diploid regions of the human genome.

[0097] Furthermore, although the HBA1 / 2 copy number detection system 31...

[0098] As further seen in FIG. 3A, computing system 3000 includes user client device 3108. In various implementations, user client device 3108 is capable of generating, storing, receiving, and transmitting digital data. In particular, user client device 3108 can receive data from sequencing device 3114. As a further illustration, user client device 3108 includes sequencing application 3110. Sequencing application 3110 can be a web application or a native application (e.g., a mobile application, a desktop application, or a web application) stored and executed on user client device 3108. Sequencing application 3110 can receive data from sequencing application 3110 and / or HBA1 / 2 copy number detection system 3106. For example, user client device 3108 can receive a variant call file and / or an alignment file from sequencing application 3110.

[0099] Furthermore, sequencing application 3110 can include instructions that cause (at runtime) user client device 3108 to receive data from HBA1 / 2 copy number detection system 3106 and present data from sequencing device 3114 and / or server device(s) 3102. Additionally, sequencing application 3110 can instruct user client device 3108 to display data of variant calls, such as nucleobase calls or indications of HBA1 / 2 copy number polymorphisms. Indeed, user client device 3108 can display nucleobase call results of a genomic sample and / or indications of predicted HBA1 / 2 copy number polymorphisms.

[0100] As further seen in FIG. 3A, computing system 3000 includes a sequencing device 3114. In various implementations, sequencing device 3114 can sequence a genomic sample or other nucleic acid polymers. For example, sequencing device 3114 can analyze nucleic acid segments or oligonucleotides extracted from a genomic sample and generate data either directly or indirectly on sequencing device 3114. More specifically, sequencing device 3114 receives and analyzes nucleic acid sequences extracted from a genomic sample within a nucleotide sample slide (such as a flow cell). In one or more implementations, sequencing device 3114 can sequence a genomic sample or other nucleic acid polymers using SBS. In addition to or as an alternative to communication via network 3112, in some implementations, sequencing device 3114 bypasses network 3112 and communicates directly with user client device 3108.

[0101] As further depicted in FIG. 3A, in some implementations, server device(s) 3102 includes a collection of distributed servers, and server device(s) 3102 is distributed throughout network 3112 and includes several server devices located at the same physical location or different physical locations. For example, server device(s) 3102 can be implemented in whole or in part on local device 3118. By way of example, local device 3118 can implement sequencing application 3110 and / or HBA1 / 2 copy number detection system 3106. Further, server device(s) 3102 and / or local device 3118 can include a content server, an application server, a communication server, a web hosting server, or other types of servers.

[0102] The user client device 3108 illustrated in FIG. 3A can include various types of client devices. For example, in some implementations, the user client device 3108 includes a non-mobile device such as a desktop computer or a server, or other types of client devices. In various implementations, the user client device 3108 includes a mobile device such as a laptop, a tablet, a mobile phone, or a smart phone.

[0103] FIG. 3A illustrates the components of a computing system 3000 that communicate via a network 3112. In certain implementations, however, the components of the computing system 3000 can also communicate directly with each other, bypassing the network 3112. For example, in some implementations, the user client device 3108 communicates directly with the array determination device 3114. Further, in some implementations, the user client device 3108 communicates directly with the HBA1 / 2 copy number detection system 3106 and / or the server device(s) 3102. In some implementations, the user client device 3108 communicates directly with the local device 3118. Further, the HBA1 / 2 copy number detection system 3106 can access one or more databases that are housed on the server device(s) 3102 or elsewhere within the computing system 3000, or that are accessible by such server devices.

[0104] Figure 3B is a block diagram of an exemplary server device 3102 that can be used in conjunction with the exemplary array determination system 3000 of Figure 3A. The server device 3102 can be configured to determine the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample. The general architecture of the server device 3102 depicted in Figure 3B includes the arrangement of computer hardware and software components. The server device 3102 can include more (or fewer) elements than those shown in Figure 3B. However, not all of these general conventional elements need to be shown to provide an effective disclosure. As illustrated, the server device 3102 includes a processing unit 310, a network interface 320, a computer-readable media drive 330, an input / output device interface 340, a display 350, and an input device 360, all of which can communicate with each other via a communication bus. The network interface 320 can provide connectivity to one or more networks or computing systems. Thus, the processing unit 310 can receive information and instructions from other computing systems or services via the network. Further, the processing unit 310 can communicate with the memory 370 and can also provide output information for any display 350 via the input / output device interface 340. Additionally, the input / output device interface 340 can receive input from any input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, game pad, accelerometer, gyroscope, or other input device.

[0105] Memory 370 may store computer program instructions (grouped as modules or components in some embodiments) that are executed by processing unit 310 to implement one or more embodiments. Generally, memory 370 includes RAM, ROM, and / or other persistent, auxiliary, or non-transitory computer-readable media. Memory 370 can store operating system 372 that provides computer program instructions used by processing unit 310 during general management and operation of server device 3102. Memory 370 can store reference genome 373, such as used by array determination application 3110. Memory 370 may further include computer program instructions for implementing aspects of the present disclosure and other information.

[0106] For example, in one embodiment, memory 370 may include array determination application 3110, which may include HBA1 / 2 copy number detection system 3106. HBA1 / 2 copy number detection system 3106 can execute the methods disclosed herein. Additionally, memory 370 may include or communicate with data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) related to determining the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample of the present disclosure, such as array determination reads, estimated copy number(s), and determined variant calls (e.g., detection of HBA1 / 2 copy number polymorphism).

[0107] In some embodiments, the disclosed systems and methods may involve an approach for shifting or distributing certain sequence data, analysis functions, and sequence data, storage devices to a cloud computing environment or cloud-based network. User interactions with sequencing data, genomic data, or other types of biological data may be mediated through a central hub that stores the data and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide for sharing of protocols, analysis methods, libraries, sequence data, and distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by a user. In some embodiments, the systems and methods may be implemented on a computer browser, on-demand, or online.

[0108] In some embodiments, software written to execute the methods described herein is stored on some form of computer-readable medium such as memory, CD-ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system, and the like.

[0109] In some embodiments, the method may be written in any of several compiled languages such as various suitable programming languages, for example, C, C#, C++, Fortran, and Java. Other programming languages may include scripting languages such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the method is written in C, C#, C++, Fortran, Java, Perl, R, Java, or Python. In some embodiments, the method may be an independent application having data input and data display modules. Alternatively, the method may be a computer software product, and distributed objects may include classes that include an application that includes the computational methods described herein.

[0110] In some embodiments, the method may be incorporated into existing data analysis software, such as that found in a sequencing instrument. Software that includes the computer-implemented methods described herein may be installed directly on a computer system or indirectly held on a computer-readable medium and loaded onto the computer system as needed. Further, the method may be located on a computer remote from where the data is generated, such as software found on a server that maintains the data in a location separate from where it is generated, such as that provided by a third-party service provider.

[0111] An assay instrument, desktop computer, laptop computer, or server may include a processor that operably communicates with an accessible memory that includes instructions for implementation of the system and method. In some embodiments, a desktop computer or laptop computer operably communicates with one or more computer-readable storage media or devices and / or an output device. The assay instrument, desktop computer, and laptop computer may operate under many different computer-based operating languages, such as those utilized by an Apple-based computer system or a PC-based computer system. The assay instrument, desktop, and / or laptop computer and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results, and monitoring experimental progress. In some embodiments, the output device may be a computer monitor, or a graphic user interface such as a computer screen, a printer, a handheld device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone, etc.), a tablet computer (i.e., iPad, etc.), a hard drive, a server, a memory stick, a flash drive, etc.

[0112] A computer-readable memory device or medium can be any device such as a server, mainframe, supercomputer, magnetic tape system, etc. In some embodiments, the memory device can be located in a location proximate to the assay device, e.g., adjacent to or in a close proximity to the assay device. For example, the memory device can be located in the same room, in the same building, in an adjacent building, on the same floor within a building, on different floors within a building, etc., in relation to the assay device. In some embodiments, the memory device can be located outside or distal to the assay device. For example, the memory device can be located in a different location in the city, in a different city, in a different state, in a different country, relative to the assay device. In embodiments where the memory device is located distal to the assay device, communication between the assay device and one or more of a desktop, laptop, or server typically occurs via an Internet connection either wirelessly through an access point or by a network cable. In some embodiments, the memory device can be maintained and managed by an individual or entity directly associated with the assay device, while in other embodiments, the memory device can typically be maintained and managed by a third party in a location distal to the individual or entity associated with the assay device. In the embodiments described herein, the output device can be any device for visualizing data.

[0113] An assay device, desktop, laptop, and / or server system may be used to store and / or retrieve a computer-implemented software program incorporating computer code for performing and implementing the calculation methods described herein, data for use in the implementation of the calculation methods, and the like. One or more of the assay device, desktop, laptop, and / or server may include a software program incorporating computer code for performing and implementing the calculation methods described herein, and one or more computer-readable storage media for storing and / or retrieving data for use in the implementation of the calculation methods. The computer-readable storage media may include, but is not limited to, one or more of a hard drive, SSD hard drive, CD-ROM drive, DVD-ROM drive, floppy disk, tape, flash memory stick or card. Further, a network including the Internet may be a computer-readable storage media. In some embodiments, the computer-readable storage media refers to computational resource storage accessible via a computer network or service provider provided company network over the Internet, rather than, for example, from a local desktop or laptop computer at a distal location to the assay device.

[0114] In some embodiments, a computer-implemented software program incorporating computer code for performing and implementing the calculation methods described herein, and a computer-readable storage media for storing and / or retrieving data used in the implementation of the calculation methods, are operated and maintained by a service provider that operably communicates with the assay device, desktop, laptop, and / or server system via an Internet connection or network connection.

[0115] In some embodiments, the hardware platform for providing a computing environment includes a processor (i.e., CPU) where memory layout such as processor time and random access memory (i.e., RAM) is a system consideration. For example, smaller computer systems provide inexpensive, high-speed processors and large memory and storage capabilities. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, the hardware platform for executing the computing methods described herein includes one or more computer systems having one or more processors. In some embodiments, smaller computers are clustered together to form a supercomputer network.

[0116] In some embodiments, the computing methods described herein are executed in an assembly of interconnected or intra-connected computer systems (i.e., grid technology) that can cooperatively execute various operating systems. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available from United Devices are examples of the cooperation of multiple independent computer systems for the purpose of handling large amounts of data. These systems may provide a Perl interface for submitting, monitoring, and managing large array analysis jobs on clusters in a serial or parallel configuration.

Examples

[0117] Some aspects of the above-described embodiments are disclosed in more detail in the following examples, which are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that many other embodiments are also within the scope of the present disclosure, as described above and in the claims.

[0118] Example 1 In the following examples, the GC-corrected and normalized depths for four target regions (upstream 1, upstream 2, intergenic, downstream) were incorporated as input from 2,407 unrelated samples of the 1000 Genomes Project.

[0119] The optimized Gaussian mixture model (GMM) parameters were determined using the expectation maximization process. Three Gaussians (with randomly initialized parameters) were randomly placed. For example, for each floating-point copy number x of a given sample (and for each sample), P(x|CN = 1), P(x|CN = 2), and P(x|CN = 3) were calculated to obtain the sample integer copy number. This sample integer copy number was then re-assigned to CN = k for which the posterior P(CN = k|x) was highest. And the parameters were adjusted to fit the points assigned to them. This process was repeated until the parameters converged. The resulting parameters are shown in the following table.

[0120] [Table 3]

[0121] The above table does not cover parameters for all assumed copy numbers (CN). The following strategy was used to input parameters for copy numbers not covered in the above table. The average value of CN0 was set to 0, and the average value of CN1 was set to 0.5. Based on the interval between CN3 and CN2, the average values of CN of 3 or more were input. For example, for the intergenic region, the average value of CN0 was 0, the average value of CN1 was 0.5, the average value of CN4 was 1.952, the average value of CN6 was 2.428, and so on. The prior values of copy numbers not covered in the above table were also input. The prior values of copy numbers not covered in the above table were uniformly distributed. A digital file of the gram parameter storing the standard deviation of CN2 was created. The standard deviation values of the states of other CNs were derived from the standard deviation of CN2. The standard deviation of CN = 0 was arbitrarily set to 0.032. The standard deviation of CN = x was set to the value obtained by multiplying the square root of x / 2 by CN2. Based on the low likelihood of samples with copy numbers exceeding 10, values were input as described for CN = 0 - 10 (11 states).

[0122] Example 2 In the following example, the method and system for determining the HBA1 / 2 copy number polymorphism genotype described in this specification were tested on 3,201 samples of the 1000 Genome Project. Using Illumina® short read technology, sequence reads were determined from nucleic acid samples. The sequence reads aligned to approximately 3,000 predetermined 2 kb diploid regions in the genome were counted.

[0123] Furthermore, sequence reads that aligned to four target regions proximal to the positions of the HBA1 gene and the HBA2 gene in the human genome were also counted. The median of the alignment MAPQ scores for each of the four target regions was 60. Furthermore, the four target regions included the first upstream region upstream of the HBA2 gene and the HBA1 gene, adjacent to the segmental duplication region X upstream of the HBA2 gene, at the coordinates of chr16:167503-169503 of the reference genome hg38. Furthermore, the four target regions included the second upstream region upstream of the HBA2 gene and the HBA1 gene, adjacent to the segmental duplication region Z upstream of the HBA2 gene, at the coordinates of chr16:170263-171875 of the reference genome hg38. Furthermore, the four target regions included the intergenic region between the HBA2 gene and the HBA1 gene, adjacent to the segmental duplication region Z upstream of the HBA1 gene, at the coordinates of chr16:174519-175845 of the reference genome hg38. Finally, the four target regions included the downstream region downstream of the HBA2 gene and the HBA1 gene, at the coordinates of chr16:178002-180501 of the reference genome hg38.

[0124] The array read counts for each of the target regions were normalized by the region length, GC corrected using the counts of array reads aligned to approximately 3,000 2-kb diploid regions, and for each of the four target regions, a floating-point copy number was obtained. After determining the normalized GC-corrected depth, the final copy number (CN) of the four target regions was determined using a Gaussian mixture model (GMM) with parameters (shift, prior value, mean value, standard deviation) defined in Example 1. The normalized GC-corrected depth was scaled by a shift value that first corrects for the bias in the alignment between the target region and the 3,000 normalized regions. Next, for i = 0-6, the posterior probability of CN = i given the scaled depth was calculated based on the pre-trained mean value, standard deviation, and prior value from the Gaussian mixture model listed in Example 1. Next, the CN with the highest posterior probability was selected as a candidate for the final copy number estimate. The estimated copy number was determined only if the posterior probability was greater than 0.95 and the p-value of the scaled depth in the Gaussian distribution of the candidate CN was greater than 0.001.

[0125] Next, based on the estimated integer copy number for each of the four regions, the HBA1 / 2 copy number polymorphism genotype was determined for each sample according to Table 1. If none of the four target regions had a copy number value above the quality inspection cut-off (NA in the table below), the copy number genotype was not determined.

[0126] The methods and systems described herein were able to determine the HBA1 / 2 copy number polymorphism genotype for 3,154 / 3,201 samples (98.5% of the samples). The proportion of genotypes determined among the samples is shown in the table below. In the table below, the interpretation is for research use only (RUO).

[0127]

Table 4

[0128] Example 3 In the following examples, concordance analysis was performed between samples sequenced by both Illumina® short reads and PacBio® orthogonal long reads.

[0129] Two hundred and forty-six cell line samples from the 1000 Genome Project were sequenced on both Illumina® and PacBio® sequencing systems to generate whole-genome sequencing (WGS) data. As described in Example 2, Illumina® sequence reads were used in the HBA-targeted k-mer method to determine the HBA1 / 2 copy number polymorphism genotype.

[0130] The following table shows the concordance analysis between the HBA-targeted k-mer and the orthogonal long-read technology. In the following table, "negative" refers to the aa / aa genotype, and "positive" refers to a call of deletion or duplication. For concordance, both the genotype and a specific deletion and / or duplication had to match.

[0131]

Table 5

[0132] In the above table, "PPV" refers to the positive predictive value, "NPV" refers to the negative predictive value, "PPA" refers to the positive agreement rate, and "NPA" refers to the negative agreement rate. As seen in Table 4, the positive predictive value of the HBA-targeted k-mer is 100%.

[0133] The following table is a concordance matrix between the HBA-targeted k-mer and the orthogonalization results separated by copy number genotype.

[0134]

Table 6

[0135] Example 4 In the following examples, concordance analysis was performed between samples sequenced with both Illumina® short reads and PacBio® orthogonal long reads.

[0136] Two hundred and forty-six cell line samples from the 1000 Genome Project were sequenced with both Illumina® and PacBio® sequencing systems to generate whole-genome sequencing (WGS) data. Illumina® sequence reads were used with another small variant that does not target HBA, and a copy number variant calling method.

[0137] The following table describes the concordance analysis between the untargeted caller and the orthogonal long-read technology. In the following table, "negative" refers to the aa / aa genotype, and "positive" refers to a call of deletion or duplication. For concordance, both the genotype and a specific deletion and / or duplication had to match.

[0138]

Table 7

[0139] In the above table, "PPV" refers to the positive predictive value, "NPV" refers to the negative predictive value, "PPA" refers to the positive agreement rate, and "NPA" refers to the negative agreement rate. As seen in Table 6, the positive predictive value of the untargeted caller is 9%.

[0140] The following table is a concordance matrix between the untargeted caller and the orthogonalization results separated by copy number genotype.

[0141]

Table 8

[0142] Example 5 In the following examples, the method for determining the HBA1 / 2 copy number polymorphism genotype described in Example 2, and the system and method of the system were tested in 575 trios from the 1000 Genomes Project without missing calls. In 575 / 575 trios, the genotype calls of the offspring were consistent with the genotype calls of the parents. All of the trio genotypes had copy numbers consistent with Mendelian inheritance.

[0143] For example, in the trio shown in Figure 4, the father sample HG00536 was determined to have an a-a3.7 / aa genotype, while the mother sample HG00537 was determined to have an aa / aa genotype. The offspring sample HG00538 was determined to have a -a3.7 / aa genotype, clearly inheriting the -a3.7 copy from the father and the aa copy from the mother. Thus, the offspring genotype was consistent with the Mendelian inheritance pattern in the trio shown in Figure 4 and in the other trios tested.

[0144] Other considerations The embodiments described herein are exemplary. Modifications, rearrangements, alternative processes, etc. may be made to these embodiments and still be encompassed within the teachings described herein. One or more of the steps, processes, or methods described herein may be preferably performed by one or more appropriately programmed processes and / or digital devices.

[0145] The various illustrative imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality may be implemented in various manners for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present disclosure.

[0146] The various illustrative detection systems described in connection with the embodiments disclosed herein can be implemented or executed by a machine device such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gates or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively, the processor may be a controller, a microcontroller, or a state machine, or combinations thereof. Also, the processor may be implemented as a combination of computing devices, such as, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with DSP cores, or other such configurations. For example, the systems described herein may be implemented using discrete memory chips, portions of memory within a microprocessor, flash, EPROM, or other types of memory.

[0147] The elements of a method, process, or algorithm described in connection with the embodiments disclosed in this specification can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. A software module can include computer-executable instructions that cause a hardware processor to execute the computer-executable instructions.

[0148] In particular, conditional language used in this specification such as "can", "might", "may", "for example", etc., generally conveys that a particular embodiment includes a particular feature, element, and / or state, while other embodiments do not include a particular feature, element, and / or state, unless otherwise specified or understood differently within the context in which it is used. Thus, such conditional language generally does not mean that a feature, element, and / or state is necessary in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or states are included or implemented in any particular embodiment, with or without author input or prompting. Terms such as "comprise", "including", "have", "involving", etc. are synonymous and are used in an inclusive open-ended manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or", when used, for example, to connect a list of elements, is used in its inclusive sense (not its exclusive sense) such that the term "or" means one, some, or all of the elements in the list.

[0149] Unless otherwise stated, disjunctive language such as "at least one of X, Y, or Z" is generally understood within the context to present that an item, term, etc. may be any of X, Y, or Z, or a combination thereof, as would be commonly used to suggest such. Thus, such disjunctive language generally does not and is not intended to imply that a particular embodiment requires the presence of at least one of X, at least one of Y, or at least one of Z, respectively.

[0150] Terms such as "about" or "approximately" are synonymous and are used to indicate that the value modified by such term has an understood range associated therewith, which can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term "substantially" is used to indicate that the result (such as a measured value) is close to the target value, where close can mean, for example, that the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.

[0151] Unless otherwise specified, articles such as "a" or "an" should generally be construed to include one or more of the items described. Thus, phrases such as "a device configured to" or "a device for" are intended to include one or more of the recited devices. Such one or more recited devices may also be collectively configured to perform the recited detailed description. For example, "a processor that executes processes A, B, and C" can include a first processor configured to execute process A in cooperation with a second processor configured to execute processes B and C.

[0152] The above detailed description has shown, described, and pointed out novel features applicable to exemplary embodiments, but it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms shown can be made without departing from the spirit of the present disclosure. As will be recognized, certain embodiments described herein may be embodied in forms that do not provide all of the features and advantages described herein, as some features may be used or implemented separately from others. All changes within the meaning and scope of the claims and equivalents thereto are included within their scope.

[0153] It should be understood that all combinations of the foregoing concepts (subject to the condition that such concepts are not mutually inconsistent) are intended to be part of the subject matter of the invention disclosed herein. Specifically, all combinations of the claimed subject matter that appear at the end of this disclosure are intended to be part of the subject matter of the invention disclosed herein.

[0154] The scope of this disclosure is not intended to be limited by the specific disclosure of examples in this section or elsewhere in this specification, but may be defined by the claims presented in this section or elsewhere in this specification, or claims presented in the future. The language of the claims should be broadly construed based on the language used in the claims and not limited to the examples described herein or during the prosecution of this application, and the examples should be construed as non-exclusive.

Explanation of Reference Signs

[0155] 110 Segment Duplication Region X 111 Segment Duplication Region Y 112 Segment Duplication Region Z 121 Locus 122 Locus 200 Method 210 Start Block 250 Judgment State 270 Judgment State 280 Process Block 290 End Block 310 Processing Unit 320 Network Interface 330 Computer Readable Media Drive 340 Output Device Interface 350 Display 360 Input Device 370 Memory 372 Operating System 373 Reference Genome 390 Data Store 1012 First Upstream Region 1027 Second upstream region 1032 Intergenic region 1042 Downstream region 2810 Start block 2840 Judgment state 2860 End block 3000 Computing system 3000 Array determination system 3102 Server device 3106 Number detection system 3108 User client device 3110 Array determination application 3112 Network 3114 Array determination device 3118 Local device

Claims

1. A computer-implemented method for determining the HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample, the method comprising: determining sequence reads from the nucleic acid sample; counting sequence reads that align to diploid regions of the human genome within the nucleic acid sample; counting sequence reads that align to target regions of one or more target regions proximal to the positions of the HBA1 gene and the HBA2 gene of the human genome; determining the HBA1 / 2 copy number polymorphism genotype based on the count of the sequence reads that align to the target regions of the one or more target regions compared to the count of the sequence reads that align to the diploid regions of the human genome.

2. Determining the HBA1 / 2 copy number polymorphism genotype includes, for each of the one or more target regions, estimating an integer copy number, the method according to claim 1.

3. Determining the HBA1 / 2 copy number polymorphism genotype includes normalizing the count of the sequence reads that align to each target region by the count of the sequence reads that align to the diploid regions of the human genome to determine a floating-point copy number for each of the one or more target regions, the method according to claim 2.

4. For each of the one or more target regions, estimating an integer copy number further includes applying a mixture Gaussian model to the floating-point copy number of the sequence reads that align to each target region, the method according to claim 3.

5. The mixture Gaussian model includes a predefined shift, a prior value, an average value, or a standard deviation as shown in Table 3, the method according to claim 4.

6. The one or more target regions proximal to the positions of the HBA1 gene and the HBA2 gene in the human genome include a first upstream region upstream of the HBA2 gene and the HBA1 gene, the method according to any one of claims 1 to 5.

7. The one or more target regions proximal to the positions of the HBA1 gene and the HBA2 gene in the human genome further include a second upstream region upstream of the HBA2 gene and the HBA1 gene, the method according to claim 6.

8. The method according to any one of claims 1 to 7, wherein the one or more target regions adjacent to the positions of the HBA1 gene and the HBA2 gene in the human genome include an intergenic region between the HBA2 gene and the HBA1 gene, or a downstream region downstream of the HBA2 gene and the HBA1 gene.

9. The method according to any one of claims 1 to 8, wherein the one or more target regions include first and second upstream regions upstream of the HBA2 gene and the HBA1 gene, an intergenic region between the HBA2 gene and the HBA1 gene, and a downstream region downstream of the HBA2 gene and the HBA1 gene.

10. The method according to any one of claims 1 to 9, wherein the sequence reads align to each of the one or more target regions with at least 30 alignment MAPQ scores.

11. The method according to any one of claims 6 to 9, wherein the first upstream region is adjacent to a segmental duplication region X upstream of the HBA2 gene.

12. The method according to any one of claims 7 or 9, wherein the second upstream region corresponds to a region within an α4.2 deletion event.

13. The method according to any one of claims 7, 9 or 12, wherein the second upstream region is adjacent to a segmental duplication region Z upstream of the HBA2 gene.

14. The method according to any one of claims 8 or 9, wherein the intergenic region corresponds to a region within an α3.7 deletion event.

15. The method according to any one of claims 8 to 9, or 14, wherein the intergenic region is adjacent to a segmental duplication region Z upstream of the HBA1 gene.

16. The method according to claim 9, wherein the first upstream region, the second upstream region, the intergenic region, and the downstream region correspond to regions within deletion events in cis for both HBA1 and HBA2.

17. The method according to any one of claims 6 to 16, wherein the first upstream region has coordinates chr16:167503 - 169503 of the reference genome hg38, the second upstream region has coordinates chr16:170263 - 171875 of the reference genome hg38, the intergenic region has coordinates chr16:174519 - 175845 of the reference genome hg38, or the downstream region has coordinates chr16:178002 - 180501 of the reference genome hg38.

18. Determining the HBA1 / 2 copy number polymorphism genotype is aaa 3.7 / aa genotype, aaa 4.2 / aa genotype, aa / aa genotype, -a 3.7 / aa genotype, -a 4.2 / aa genotype, -- / aaa 3.7 genotype, -- / aaa 4.2 genotype, -a 3.7 / -a 3.7 genotype, -a 4.2 / -a 4.2 genotype, -a 3.7 / -a 4.2 genotype, -- / aa genotype, -- / a 3.7 genotype, -- / a 4.2 genotype, or the method according to any one of claims 1 to 17, including determining the -- / -- genotype.

19. A computer-implemented method for detecting one or more single nucleotide variants, or indels, in the HBA1 / 2 region in a nucleic acid sample, the method comprising: Determining sequence reads from the nucleic acid sample; Obtaining sequence reads that align to sites of single nucleotide variants, or indels, within the HBA1 gene, or the HBA2 gene, of the human genome in the nucleic acid sample; Counting sequence reads that store bases corresponding to alternative alleles at the sites of the single nucleotide variants, or indels, wherein counting the sequence reads includes counting sequence reads that align to the HBA1 gene and sequence reads that align to the HBA2 gene; Creating a digital file that includes a variant call corresponding to the single nucleotide variant, or indel, wherein the variant call is not specific to the HBA1 gene or the HBA2 gene.

20. The single nucleotide variant or indel is HBA2_c.60del, HBA2_c.69C>T, HBA2_c.95+2_95+6delTGAGG, HBA2_c.95+1G>A, HBA1_c.179G>A, HBA2_c.377T>C, HBA2_c.427T>C, HBA2_c.427T>G, HBA2_c.429A>T, HBA2_c. * 92A>G, HBA2_c.428A>C, HBA2_c.314G>A, HBA2_c.379G>A, HBA2_c.179G>A, HBA2_c.75T>G, HBA1__c.96-1G>A, HBA1_c.358C>T, or HBA2_c. * The method according to claim 19, comprising 94A>G.

21. An electronic system for determining an HBA1 / 2 copy number polymorphism genotype in a nucleic acid sample, comprising a processor configured to execute a method, the method comprising: Determining sequence reads from the nucleic acid sample; Counting sequence reads that align to a diploid region of the human genome within the nucleic acid sample; Counting sequence reads that align to target regions of one or more target regions that are proximal to the locations of the HBA1 gene and the HBA2 gene of the human genome; Determining an HBA1 / 2 copy number polymorphism genotype based on the count of the sequence reads that align to the one or more target regions compared to the count of the sequence reads that align to the diploid region of the human genome.

22. The electronic system according to claim 21, wherein determining the HBA1 / 2 copy number polymorphism genotype includes estimating an integer copy number for each of the one or more target regions.

23. The electronic system according to claim 22, wherein determining the HBA1 / 2 copy number polymorphism genotype includes normalizing the count of the sequence reads that align to each target region by the count of the sequence reads that align to the diploid region of the human genome to determine a floating-point copy number for each of the one or more target regions.

24. For each of the one or more target regions, estimating the integer copy number further includes applying a mixture of Gaussian models to the floating point copy numbers of the array reads that align to each target region, the electronic system according to claim 23.