Methods and systems for detecting recombination events
The method and system enhance the detection of recombination events and variants in the RCCX region by estimating copy number and constructing haplotypes, addressing the challenges of sequence homology in the CYP21A2 and CYP21A1P genes, achieving up to 100% improved accuracy.
Patent Information
- Application Number
- JP2024576523
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-07
- Filing Date
- 2023-07-05
- Publication Date
- 2025-08-13
AI Technical Summary
Existing nucleic acid sequencing technologies struggle to accurately detect recombination events and small variants in the highly homologous RCCX region, particularly between the CYP21A2 and CYP21A1P genes, due to sequence homology, leading to ambiguous variant calls and missed detections.
A computer-implemented method and system that estimates the RCCX region copy number and constructs candidate haplotypes by phasing sequence reads at predetermined differentiation sites, enabling precise detection of recombination events and single-nucleotide variants.
Improves the specificity and sensitivity of detecting recombination events and variant calling by up to 100%, enhancing the accuracy of genetic variant detection in the RCCX region.
Smart Images

Figure 2025526252000001_ABST
Abstract
Description
[Technical Field]
[0001] Incorporation by reference of any priority application All applications for which a foreign or domestic priority claim is identified in the Application Data Sheet filed with this application are incorporated herein by reference under 37 CFR 1.57.
[0002] This application claims priority to U.S. Provisional Application No. 63 / 367,896, filed July 7, 2022, the entire contents of which are incorporated herein by reference.
[0003] The disclosed technology relates to the field of nucleic acid sequencing. More particularly, the disclosed technology relates to detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample. [Background technology]
[0004] 2. Description of Related Art CYP21A2 encodes 21-hydroxylase, a cytochrome P450 enzyme that helps regulate the adrenal hormones cortisol and aldosterone. These hormones play several roles, including regulating salt retention in the kidney. CYP21A2 inactivation accounts for 95% of 21-hydroxylase CAH cases, which can take one of three forms. The first form, salt-wasting CAH, is the most severe, in which complete deficiency of CYP21A2 results in very low levels of aldosterone synthesis and therefore reduced sodium retention. Symptoms include dehydration, diarrhea, vomiting, and adrenal crisis, and can be very severe and even fatal. Low cortisol levels also play a developmental role, leading to virilization. The second form, simple virilizing CAH, is a milder form caused by reduced CYP21A2 activity without a complete genetic defect. This form generally avoids the most severe and life-threatening symptoms, but typically still presents with virilization and developmental challenges. The third form, non-classic CAH, has symptoms similar to those of simple virilizing CAH. Non-classic CAH is characterized by higher aldosterone and cortisol hormone levels, resulting in less severe symptoms. Because of the lesser phenotypic impact, diagnosing non-classic CAH is more challenging.
[0005] CYP21A2 resides within a 30-kilobase segmental duplication within the major histocompatibility complex (MHC) class III region. The repeat, commonly referred to as RCCX, contains part or all of four genes: STK19, C4A / C4B, CYP21A2, and TNXB. RCCX repeats typically exist as two modules with nearly identical sequences. The first module contains the end of the STK19 gene, the active C4A gene, and two inactive pseudogenes, CYP21A1P and TNXA. The second module contains the ends of C4B, CYP21A2, and TNXB, all active genes that play important roles in human health.
[0006] The high sequence homology of the RCCX region drives a high rate of non-allelic homologous recombination. These recombination events can occur at any point within the repeat. If the breakpoint of a recombination event occurs within the CYP21A2 region, a chimeric gene fusion is created with part of the pseudogene sequence and part of the gene sequence. Despite approximately 98% sequence similarity between the gene and pseudogene, these chimeric fusion genes can be partially or completely inactivated by introducing several small mutations from the pseudogene into the gene. These can be considered partial gene conversions. CYP21A2 is also subject to more standard gene conversion mutations of the partial gene sequence, likely due to template switching during synthesis excision repair.
[0007] If the recombination breakpoint for the deletion occurs outside the gene, it can be completely deleted from the resulting chimeric RCCX module, leaving only CYP21A1P. This heterozygous CYP21A2 deletion creates a carrier state and subsequently leads to phenotypic effects when co-inherited with another defective allele. Summary of the Invention [Means for solving the problem]
[0008] In one aspect, disclosed herein is a computer-implemented method for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample. In some embodiments, the method includes receiving sequence reads and aligning them to the RCCX region of the human genome in the nucleic acid sample, estimating the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads, constructing one or more candidate haplotypes by phasing a plurality of sequence reads that are aligned to the CYP21A2 gene or the CYP21A1P gene in the human genome and include at least two predetermined differentiation sites of the CYP21A2 gene and the CYP21A1P gene, and detecting recombination events between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on the one or more candidate haplotypes.
[0009] In some embodiments, the one or more candidate haplotypes cover one or more breakpoints of the recombination events. In some embodiments, constructing the one or more candidate haplotypes comprises identifying at least one seed sequence read from the plurality of sequence reads. In some embodiments, the seed sequence read is selected from a 5' seed sequence read, a center sequence read, and a 3' seed sequence read. In some embodiments, constructing the one or more candidate haplotypes comprises iteratively extending the at least one seed sequence read in either the 5' or 3' direction by aligning the sequence reads using predetermined differentiation sites.
[0010] In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the RCCX region of the human genome. In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to a C4A gene, a CYP21A1P gene, a TNXA gene, a C4B gene, a CYP21A2 gene, or a TNXB gene in the human genome. In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to a region corresponding to positions chr6:32024461 to chr6:32043719 of reference genome hg38, chr6:31991723 to chr6:32010985 of reference genome hg38, chr6:31992238 to chr6:32011496 of reference genome hg19, or chr6:31959500 to chr6:31978762 of reference genome hg19.
[0011] In some embodiments, estimating the copy number comprises normalizing the counts of sequence reads that align to the RCCX region of the human genome. In some embodiments, estimating the copy number comprises binning the normalized counts of sequence reads that align to the RCCX region of the human genome using a Gaussian mixture model.
[0012] In some embodiments, the disclosed methods and systems further comprise making variant calls at predetermined differentiation sites in a plurality of predetermined differentiation sites. In some embodiments, the disclosed methods and systems further comprise making variant calls for recombination events. In some embodiments, the disclosed methods and systems further comprise creating a digital file comprising the variant calls. In some embodiments, the disclosed methods and systems further comprise creating a digital file comprising one or more candidate haplotypes.
[0013] In some embodiments, in the reference genome hg38, the plurality of predetermined differentiation sites are chr6:32038514, chr6:32038844, chr6:32039015, chr6:32039081, chr6:32039128, chr6:32039132, chr6:32039143, chr6:32039426, chr6:3203954 of the CYP21A2 gene. 8, chr6:32039802, chr6:32039807, chr6:32039810, chr6:32039816, chr6:32040110, chr6:32040182, chr6:32040216, chr6:32040421, or chr6:32040535, or the corresponding position in the pseudogene CYP21A1P. In some embodiments, in the reference genome hg19, the plurality of predetermined differentiation sites are chr6:32006291, chr6:32006621, chr6:32006792, chr6:32006858, chr6:32006905, chr6:32006909, chr6:32006920, chr6:32007203, chr6:3200732 of the CYP21A2 gene. 5, chr6:32007579, chr6:32007584, chr6:32007587, chr6:32007593, chr6:32007887, chr6:32007959, chr6:32007993, chr6:32008198, or chr6:32008312, or the corresponding position in the pseudogene CYP21A1P.
[0014] In another aspect, a computer-implemented method for detecting one or more single-nucleotide variants or indels in an RCCX region in a nucleic acid sample is disclosed herein. In some embodiments, the method includes determining sequence reads from the nucleic acid sample, obtaining sequence reads that align to the site of the single-nucleotide variant or indel in a CYP21A2 gene or a CYP21A1P gene of the human genome of the nucleic acid sample, counting sequence reads that contain a base corresponding to an alternative allele at the site of the single-nucleotide variant or indel, including counting sequence reads that align to the CYP21A2 gene and sequence reads that align to the CYP21A1P gene, and creating a digital file containing variant calls corresponding to the single-nucleotide variants or indels, where the variant calls are not specific to the CYP21A2 gene or the CYP21A1P gene.
[0015] The specifications for this are NM_000500.9:c.60G >A、NM_000500.9:c.92C>A、NM_000500.9:c.111del、NM_0 00500.9:c.159_160del、NM_000500.9:c.169G>A、NM_000500.9:c.274A>G、NM_000500.9:c.332_339del、NM_000500 .9:c.418G>A、NM_000500.9:c.421G>A、NM_000500.9:c.515T>A、NM_000500.9:c.710_719delinsACGAGGAGA、NM_00 0500.9:c.850A>G、NM_000500.9:c.874G>A、NM_000500.9:c.922T>G、NM_000500.9:c.923_924dup、NM_000500.9:c. 952C>T=、NM_000500.9:c.955C>G、NM_000500.9:c.1042G>A、NM_000500.9:c M_000500.9:c.1070G>A、NM_000500.9:c.1096C>T、NM_000500.9:c.1118G>A.9:c.1136T>A、NM_000500. 9:c.1226G>A、NM_000500.9:c.1273G>A、NM_000500.9:c.1274G>T、NM_000500.9:c.1279C>T、NM_000500.9:c.1357C >T=、NM_000500.9:c.1360C>T、NM_000500.9:c.1444C>T、NM_000500.9:c.
[0016] In another aspect, disclosed herein is an electronic system for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample. In some embodiments, the system includes a processor configured to perform a method including receiving sequence reads that align to the RCCX region of the human genome in the nucleic acid sample, estimating the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads, constructing one or more candidate haplotypes by phasing a plurality of sequence reads that align to the CYP21A2 gene or the CYP21A1P gene in the human genome and that include at least two predetermined differentiation sites in the CYP21A2 gene and the CYP21A1P gene, and detecting recombination events between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on the one or more candidate haplotypes.
[0017] In some embodiments, the processor is configured to perform a method comprising detecting a recombination event between the CYP21A2 gene and the CYP21A1P gene based on an estimated copy number of the RCCX region of the human genome and based on one or more candidate haplotypes.
[0018] In some embodiments, the one or more candidate haplotypes cover one or more breakpoints of the recombination events. In some embodiments, constructing the one or more candidate haplotypes comprises identifying at least one seed sequence read from the plurality of sequence reads. In some embodiments, the seed sequence read is selected from a 5' seed sequence read, a center sequence read, and a 3' seed sequence read. In some embodiments, constructing the one or more candidate haplotypes comprises iteratively extending the at least one seed sequence read in either the 5' or 3' direction by aligning the sequence reads using predetermined differentiation sites.
[0019] In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the RCCX region of the human genome.
[0020] In a further aspect, disclosed herein is an electronic system for detecting one or more single nucleotide variants or indels in an RCCX region in a nucleic acid sample. In some embodiments, the system comprises a processor configured to perform a method including determining sequence reads from the nucleic acid sample, obtaining sequence reads that align to the site of the single nucleotide variant or indel in a CYP21A2 gene or a CYP21A1P gene of the human genome from the nucleic acid sample, counting sequence reads that contain a base corresponding to an alternative allele at the site of the single nucleotide variant or indel, including counting sequence reads that align to the CYP21A2 gene and sequence reads that align to the CYP21A1P gene, and creating a digital file containing variant calls corresponding to the single nucleotide variants or indels, where the variant calls are not specific to the CYP21A2 gene or the CYP21A1P gene. [Brief explanation of the drawings]
[0021] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numbers correspond to similar, if not identical, components, and for the sake of brevity, reference numbers or features having a previously mentioned function may or may not be described with reference to other drawings in which they appear.
[0022] [Figure 1A] The RCCX domain and RCCX module are shown schematically. [Figure 1B] Schematic representation of recombination events within the RCCX region. [Figure 2A] FIG. 1 is a block diagram illustrating a method for detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample. [Figure 2B]FIG. 1 is a block diagram further outlining the process of constructing one or more candidate haplotypes. [Figure 3] 1 illustrates a schematic diagram of one embodiment of a candidate haplotype construction process. [Figure 4A] FIG. 1 is a block diagram of an exemplary sequencing system that can be used to perform the disclosed methods. [Figure 4B] FIG. 4B is a block diagram of an exemplary computing device that may be used in conjunction with the exemplary sequencing system of FIG. 4A. [Figure 5] Schematic representation of recombinant haplotypes constructed in a trio of congenital adrenal hyperplasia (CAH) cases. [Figure 6] A graphical comparison of RCCX module copy number estimates with copy number calls from Bionano optical mapping is shown. DETAILED DESCRIPTION OF THE INVENTION
[0023] All patents, patent applications, and other publications, including all sequences disclosed therein and referred to herein, are expressly incorporated by reference herein to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. All cited documents are, in relevant part, incorporated herein by reference in their entirety, for the purposes indicated by the context of the citation herein. However, the citation of any document should not be construed as an admission that it is prior art to the present disclosure.
[0024] overview CYP21A2 is located within the RCCX region, as shown schematically in Figure 1A. Recombination events occur at a high rate within the RCCX region due to the high sequence homology. For example, deletion and duplication events occurring between CYP21A2 and CYP21A1P are shown schematically in Figure 1B. Recombination events in the RCCX region, such as those between CYP21A2 and CYP21A1P, can be difficult to detect due to the high sequence homology between the CYP21A2 and CYP21A1P genes. For example, detecting gene conversion variants can be difficult because sequence reads across a gene conversion boundary may contain alleles from alternative RCCX modules at the gene conversion site and may preferentially map to the wrong gene.
[0025] Other small variants (single-base and insertion / deletion events) can also result in reduced CYP21A2 activity. These variants may reside in regions of the CYP21A2 gene whose nucleotide sequence is identical to that of the CYP21A1P pseudogene, making variant detection extremely difficult. This is because reads sequenced from either the gene or the pseudogene may lack identifying markers, meaning they may be randomly assigned to the wrong gene during the post-sequencing assembly process. This can result in weak and ambiguous evidence for variants at either position, potentially resulting in missed or low-confidence variant calls.
[0026] This combination of factors has made it difficult to accurately determine the sequence information of the CYP21A2 or CYP21A1P gene using whole genome sequencing (WGS) data. The disclosed method and system overcome the challenge of sequence homology between the CYP21A2 gene and the CYP21A1P pseudogene to detect multiple types of genetic variants in this genomic region. Such genetic variants can include small mutations, gene conversions, and deletions or duplications of the entire gene resulting from recombination.
[0027] Described herein are methods and systems for detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample obtained from one or more subjects. The disclosed systems and methods for detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample have been found to improve the specificity and sensitivity of detecting recombination events and variant calling between the CYP21A2 and CYP21A1P genes in the RCCX region of a nucleic acid sample.
[0028] In some embodiments, the disclosed systems and methods include receiving sequence reads that align to an RCCX region found in a biological sample taken from a subject. Once the sequence reads are received, the copy number of the RCCX region can be estimated. Estimating the RCCX copy number can include counting sequence reads that align to the RCCX region of a reference genome.
[0029] The disclosed systems and methods can then construct one or more candidate haplotypes by phasing a plurality of sequence reads that are aligned to the CYP21A2 or CYP21A1P gene of the human genome and include at least two predetermined differentiation sites in the CYP21A2 and CYP21A1P genes. These predetermined differentiation sites can include positions in the nucleic acid sequence of the CYP21A2 gene or a corresponding position in the CYP21A1P gene that includes at least one base that differs between the CYP21A2 and CYP21A1P genes, and that the difference is predetermined to be fixed in the population. Thus, these predetermined differentiation sites can be used to determine whether a particular sequence read corresponds to the CYP21A2 or CYP21A1P gene, which includes detecting recombination events between the CYP21A2 and CYP21A1P genes.
[0030] In some embodiments, the disclosed systems and methods detect recombination events between the CYP21A2 and CYP21A1P genes based on the estimated copy number of the RCCX region of the human genome and based on one or more candidate haplotypes. For example, the disclosed methods and systems can detect recombination events, such as gene conversions, duplications, or deletions, based on the RCCX copy number along a predetermined differentiation site in one or more candidate haplotypes and / or based on detecting a transition from a CYP21A2 specific base to a CYP21A1P specific base (or vice versa).
[0031] The disclosed systems and methods can improve the recall (also known as sensitivity, which is the proportion of true variants correctly detected) of single nucleotide polymorphisms (SNPs) generated by recombination events between the CYP21A2 gene and CYP21A1P by 20%, 50%, 80%, 100%, or more.
[0032] definition Unless otherwise defined, technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of this disclosure, the following terms are defined below.
[0033] As used herein, a "nucleotide" comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are the monomeric units of nucleic acid sequences. Examples of nucleotides include ribonucleotides and deoxyribonucleotides. In ribonucleotides (RNA), the sugar is ribose, while in deoxyribonucleotides (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group at the 2' position of the ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of deoxyribose is linked to the N-1 atom of a pyrimidine or the N-9 atom of a purine. The phosphate group can be mono-, di-, or triphosphate. These nucleotides may be naturally occurring nucleotides, although it should be further understood that non-naturally occurring nucleotides, modified nucleotides or analogs of the aforementioned nucleotides may also be used.
[0034] As used herein, a "base" or "nucleobase" is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. Nucleobases can be naturally occurring or synthetic. Non-limiting examples of nucleobases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purine substituted at the 8-position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)-alkynylcytosine, 5-fluorouracil, and 5-bromouracil. , thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine, and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510, and WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (Practical Handbook of Biochemistry and Molecular Biology, pp. 485-494, 1989, CRC Press, Boca Raton, LO), all of which are incorporated herein by reference in their entireties.
[0035] The terms "nucleic acid" or "polynucleotide" refer to deoxyribonucleotide or ribonucleotide polymers in single- or double-stranded form and, unless otherwise limited, encompass known analogs of natural nucleotides that hybridize to nucleic acids in a manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise specified, a particular nucleic acid sequence includes its complementary sequence. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha-thiotriphosphate for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all of the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl-dCTP, and 5-propynyl-dUTP.
[0036] As used herein, the term "chromosome" refers to a gene carrier of the present invention in a living cell, derived from a chromatin strand containing DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is used herein.
[0037] "Genome" means the complete genetic information of an organism or virus expressed in nucleic acid sequences.
[0038] As used herein, the term "reference genome" or "reference sequence" refers to a specific known genomic sequence, either partial or complete, of any organism or virus that can be used to reference a sequence identified from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information (ncbi.nlm.nih.gov). In various embodiments, the reference sequence may be significantly larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 10 times larger. 5 times larger, or at least about 10 6 times larger, or at least about 10 7 The reference sequence may be 1-fold larger. In one example, the reference sequence is of a full-length genome. Such a sequence may be referred to as a genome reference sequence. For example, the reference sequence may be a reference human genome sequence such as hg19 (e.g., available at GenBank assembly accession number GCA_000001405.1) or hg38 (e.g., available at GenBank assembly accession number GCA_000001405.15). In another example, the reference sequence is limited to a particular human chromosome, such as chromosome 13. In some embodiments, the reference Y chromosome is the Y chromosome sequence from human genome version hg19. Such a sequence may be referred to as a chromosome reference sequence. Other examples of reference sequences include genomes of other species, as well as chromosomes, partial chromosomal regions (such as strands) of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a specific individual.
[0039] As used herein, the term "nucleic acid sample" refers to a sample, typically derived from a biological fluid, cell, tissue, organ, or organism, that contains a nucleic acid or mixture of nucleic acids comprising at least one nucleic acid sequence to be screened for copy number variations. In certain embodiments, the nucleic acid sample contains at least one nucleic acid sequence suspected of being copy number varied. Such samples may include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, fine-needle biopsy samples (e.g., surgical biopsy, fine-needle biopsy), urine, peritoneal fluid, pleural effusion, etc. Samples are often collected from human subjects (e.g., patients), but samples may also be collected from any mammal, including, but not limited to, dogs, cats, horses, goats, sheep, cows, pigs, etc. Samples may be used directly upon acquisition from a biological source or may be used after pretreatment to alter the sample's characteristics. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Additionally, pretreatment methods can include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, lysis, etc. When such pretreatment methods are employed on a sample, such pretreatment methods typically result in the nucleic acid(s) of interest remaining in the test sample, in some cases at a concentration proportional to that in the untreated test sample (i.e., a sample not treated with such pretreatment method(s)). Such "treated" or "processed" samples are still considered to be biological "test" samples with respect to the methods described herein.
[0040] The term "read" or "sequence read" (or sequencing read) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any portion or the entire nucleic acid molecule. Typically, but not necessarily, a read represents a short sequence of consecutive base pairs in a sample. A read may be symbolically represented by the base pair sequence (A, T, C, or G) of the sample portion. A read may be stored in a memory device and appropriately processed to determine whether the read matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing device or indirectly from stored sequence information about the sample. In some cases, a read refers to a DNA sequence of sufficient length (e.g., at least about 25 bp) that it can be used to identify a larger sequence or region, e.g., aligned and specifically assigned to a chromosome, genomic region, or gene. For example, a sequence read can be a short string of nucleotides (e.g., 20-150 bases) sequenced from a nucleic acid fragment, at one or both ends of a nucleic acid fragment, or the entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained by any method known in the art. For example, sequence reads can be obtained by a variety of methods, including the use of sequencing techniques, probes such as hybridization arrays or capture probes, or amplification techniques such as polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing-by-synthesis, sequencing-by-ligation, or sequencing-by-ligation. Sequence reads can be generated using instruments such as the MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0041] As used herein, the term "sequencing depth" generally refers to the number of times a locus is covered by sequence reads aligned to that locus. A locus can be as small as a nucleotide, as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as 50x, 100x, etc., where "x" refers to the number of times a locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome, in which case x can refer to the average number of times a locus, a haploid genome, or the entire genome is sequenced, respectively. When average depth is cited, the actual depth of different loci included in a dataset spans a range of values. Ultra-deep sequencing can refer to a sequencing depth of at least 100x.
[0042] As used herein, the terms "aligned," "alignment," or "aligning" refer to the process of comparing a read or tag to a reference sequence, thereby determining the likelihood that the reference sequence contains the read sequence. When the reference sequence includes a read, the read may be mapped to the reference sequence, or in certain other embodiments, may be mapped to a specific location within the reference sequence. For example, alignment of a read to a reference sequence for human chromosome 13 indicates the likelihood that the read is present in the reference sequence for chromosome 13. In some cases, alignment also indicates where the read or tag maps within the reference sequence. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13 and may further indicate that the read is located on a specific strand and / or site of chromosome 13. A "site" can refer to a polynucleotide sequence or a unique location on the reference genome (i.e., chromosome ID, chromosomal location, and orientation). In some embodiments, a site can be a residue, a sequence tag, or the location of a segment on a sequence.
[0043] Aligned read or tag is one or more sequences that are identified as matching in the order of nucleic acid molecules from reference genome to known sequence.Alignment can be performed manually, but is typically performed by computer algorithm, because it is impossible to align reads in a reasonable time period for carrying out the method disclosed herein.The matching of sequence reads during alignment can be 100% sequence match or less than 100% (non-perfect match).
[0044] Alignment is Burrows-Wheeler Aligner(BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, EL AND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, GeneiousAssembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT It may be performed by variations and / or combinations of methods such as Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.
[0045] As used herein, the term "mapping" refers to the specific assignment of sequence reads to a larger sequence, such as a reference genome, by alignment.
[0046] A "genetic variation" or "genetic mutation" refers to a specific genotype present in a particular individual; often, genetic variations exist in a statistically significant subpopulation of individuals. The presence or absence of genetic variance can be determined using methods or devices described herein. In certain embodiments, the presence or absence of one or more genetic variations is determined according to the results provided by the methods and devices described herein. In some embodiments, the genetic variation is a chromosomal abnormality (e.g., aneuploidy), a partial chromosomal abnormality, or mosaicism, each of which is further described herein. Non-limiting examples of genetic variations include one or more deletions (e.g., microdeletions), duplications (e.g., microduplications), insertions, mutations, polymorphisms (e.g., single nucleotide polymorphisms), fusions, repeats (e.g., short tandem repeats), differential methylation sites, differential methylation patterns, and the like, and combinations thereof. An insertion, repeat, deletion, duplication, mutation, or polymorphism can be of any length and, in some embodiments, can be from about 1 base or base pair (bp) to about 250 megabase pairs (Mb) in length. In some embodiments, the length of the insertion, repeat, deletion, duplication, mutation, or polymorphism is from about 1 base or base pair (bp) to about 1000 kilobases (kb) (e.g., about 10 bp, 50 bp, 100 bp, 500 bp, 1 kb, 5 kb, 10 kb, 50 kb, 100 kb, 500 kb, or 1000 kb).
[0047] In some cases, the genetic variation is a deletion. In certain embodiments, a deletion refers to a mutation (such as a genetic abnormality) in which a portion of a chromosome or a sequence of DNA is missing. A deletion is often a loss of genetic material. Any number of nucleotides may be deleted. A deletion may include the deletion of one or more entire chromosomes, chromosomal segments, alleles, genes, introns, exons, any non-coding regions, any coding regions, segments thereof, or combinations thereof. A deletion may include a microdeletion. A deletion may include the deletion of a single base.
[0048] Genetic variations are sometimes genetic duplications. In certain embodiments, a duplication is a mutation (e.g., a genetic abnormality) in which a portion of a chromosome or a sequence of DNA is copied and reinserted into the genome. In certain embodiments, a genetic duplication (i.e., a duplication) is a duplication of a region of DNA. In some embodiments, a duplication is a nucleic acid sequence that is repeated, often in tandem, within a genome or chromosome. In some embodiments, a duplication can include copies of one or more entire chromosomes, chromosomal segments, alleles, genes, introns, exons, any non-coding regions, any coding regions, segments thereof, or combinations thereof. Duplications can include microduplications. Duplications sometimes include one or more copies of the duplicated nucleic acid. In some cases, duplications are characterized as a genetic region that is repeated one or more times (e.g., repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times). Duplications can range from small regions (thousands of base pairs) to entire chromosomes. Duplications frequently occur as a result of errors in homologous recombination or due to retrotransposon events. Duplications are associated with certain types of proliferative disorders. Duplications can be characterized using genomic microarrays or comparative genetic hybridization (CGH).
[0049] A genetic variation is sometimes an insertion. An insertion is sometimes the addition of one or more nucleotide base pairs to a nucleic acid sequence. An insertion is sometimes a microinsertion. In certain embodiments, an insertion comprises the addition of a chromosomal segment to a genome, chromosome, or segment thereof. In certain embodiments, an insertion comprises the addition of an allele, gene, intron, exon, any non-coding region, any coding region, segment thereof, or combination thereof to a genome or segment thereof. In certain embodiments, an insertion comprises the addition (i.e., insertion) of a nucleic acid of unknown origin to a genome, chromosome, or segment thereof. In certain embodiments, an insertion comprises the addition (i.e., insertion) of a single base.
[0050] Genetic variations may include copy number variations, i.e., variations in the copy number of a nucleic acid sequence present in a test sample compared to the copy number of a nucleic acid sequence present in a reference sample. In certain embodiments, the nucleic acid sequence is 1 kb or more. In some cases, the nucleic acid sequence is an entire chromosome or a significant portion thereof. Copy number variations may also refer to nucleic acid sequences in which copy number differences are found by comparing the nucleic acid sequence of interest in a test sample with the expected level of the nucleic acid sequence of interest. For example, the level of the nucleic acid sequence of interest in a test sample is compared with that present in a qualified sample. Copy number variations / copy number variations may include deletions, including microdeletions, insertions, including microinsertions, duplications, multiplications, and translocations. CNVs include chromosomal aneuploidy and partial aneuploidy.
[0051] "Phasing" refers to analyzing linkage information between sequence reads of a nucleic acid to determine whether two subsequences of a nucleic acid (such as alleles or variants) are located on a single chromosome or two separate chromosomes (e.g., maternally and paternally inherited chromosomes).
[0052] Embodiments of a method and system for detecting recombination events between the CYP21A2 and CYP21A1P genes FIG. 2 is a block diagram that schematically illustrates an exemplary method 200 for detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample. In some embodiments, method 200 is performed on a computer. Method 200 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives of a computing system. For example, server device 4102, shown in FIGS. 4A and 3B and described in more detail below, may execute a set of executable program instructions to implement method 200. When method 200 is initiated, the executable program instructions may be loaded into memory, such as RAM, and executed by one or more processors of server device 4102. Although method 200 is described with respect to server device 4102 shown in FIG. 4B, this description is by way of example only and not by way of limitation. In some embodiments, method 200, or portions thereof, may be performed serially or in parallel by multiple computing systems.
[0053] As shown in FIG. 2A, a method 200 for detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample may begin at start block 210. Method 200 may proceed to block 220, where sequence reads aligning to the RCCX region of the human genome in the nucleic acid sample are received. The method may then proceed to block 230, where the sequence reads are aligned across a reference genome, e.g., the RCCX region. Method 200 may then proceed to block 240, where the copy number of the RCCX region of the human genome in the nucleic acid sample is estimated from the aligned sequence reads. Method 200 may then proceed to process block 250, where one or more candidate haplotypes are constructed. The method performed within process block 250 is described in further detail with respect to FIG. 2B. Method 200 may then proceed to decision state 260, where the system may determine whether there are additional candidate haplotypes to be constructed. If there are additional candidate haplotypes to be constructed, method 200 may return to block 250, and the method may proceed as described above. If there are no additional candidate haplotypes to be constructed, method 200 can proceed to block 270 where recombination events between the CYP21A2 and CYP21A1P genes are detected. Method 200 can end at end block 280.
[0054] FIG. 2B is a block diagram further illustrating the method performed within process block 250 described above, in which one or more candidate haplotypes are constructed. As shown in FIG. 2B, the method of process block 250 may begin at start block 2510. The method of process block 250 may proceed to block 2520, where a 5′ seed sequence read, a center sequence read, or a 3′ seed sequence read is identified. The method of process block 250 may proceed to block 2530, where the seed sequence read is extended by alignment along a predetermined differentiation site. The method of process block 250 may proceed to decision state 2540, where the system may determine whether there are additional seed sequence reads to be extended. If there are additional seed sequence reads to be extended, the workflow may return to block 2520, and the workflow may proceed as described above. If there are no additional seed sequence reads to be extended, the workflow may proceed to block 2550, where the partial candidate haplotypes are assembled into a complete candidate haplotype. The method of process block 250 may end at end block 2560.
[0055] Receiving sequence reads that align to the RCCX region In some embodiments, the methods and systems disclosed herein include receiving a plurality of sequence reads that align to the RCCX region of the human genome in a nucleic acid sample, e.g., as shown in block 220 of Figure 2A. In some embodiments, the sequence reads are generated from a sample obtained from a subject.
[0056] In some embodiments, the RCCX region comprises two RCCX modules. For example, the RCCX modules have nearly identical sequences. In some embodiments, each RCCX module is about 10 kb, about 15 kb, about 20 kb, about 25 kb, or about 30 kb in length (or a range constructed from any of these values). In some embodiments, each RCCX module is about 20 kb in length. In some embodiments, each RCCX module is separated by about 5 kb, about 6 kb, about 7 kb, about 8 kb, about 9 kb, about 10 kb, about 11 kb, about 12 kb, about 13 kb, about 14 kb, about 15 kb, about 16 kb, about 17 kb, about 18 kb, about 19 kb, about 20 kb, about 25 kb, or about 30 kb (or a range constructed from any of these values). In some embodiments, each RCCX module is separated by about 13 kb.
[0057] In some embodiments, the first RCCX module comprises the end of the STK19 gene, the C4A gene, the CYP21A1P gene, and the TNXA gene. In some embodiments, the second RCCX module comprises the end of the C4B gene, the CYP21A2 gene, and the TNXB gene. In some embodiments, the first RCCX module comprises a HERV-K retrotransposon insertion within the C4A gene. In some embodiments, the HERV-K retrotransposon insertion is about 6.4 kb in length. In some embodiments, the second RCCX module comprises a HERV-K retrotransposon insertion within the C4B gene. In some embodiments, the HERV-K retrotransposon insertion is about 6.4 kb in length. In some embodiments, the first RCCX module covers a 120 bp deletion site in the TNXA gene relative to the TNXB gene.
[0058] In some embodiments, the RCCX region includes a region corresponding to positions chr6:32024461 to chr6:32043719 of reference genome hg38, chr6:31991723 to chr6:32010985 of reference genome hg38, chr6:31992238 to chr6:32011496 of reference genome hg19, and / or chr6:31959500 to chr6:31978762 of reference genome hg19. For example, in some embodiments, the RCCX region includes a region corresponding to positions chr6:32024461 to chr6:32043719 and chr6:31991723 to chr6:32010985 of reference genome hg38. In some embodiments, the RCCX region comprises the region corresponding to positions chr6:31992238 to chr6:32011496 and chr6:31959500 to chr6:31978762 of the reference genome hg19.
[0059] Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by ligation. Sequence reads can be generated using instruments such as the MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA). Sequence reads can be, for example, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000, or more base pairs (bps) in length, respectively. For example, sequence reads can be about 100 base pairs to about 1000 base pairs in length, respectively. Sequence reads can include paired-end sequence reads. The sequence reads may include single-end sequence reads. The sequence reads may be generated by whole genome sequencing (WGS). The WGS may be clinical WGS (cWGS). The sample may include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof.
[0060] In some embodiments, the sequence reads are obtained by aligning the reads to the RCCX region of a reference sequence. For example, the sequence reads may be aligned to a reference genome, as shown in block 230 of FIG. 2A. In some embodiments, the sequence reads are obtained by aligning a first plurality of sequence reads generated from a sample to a reference genome sequence to obtain a second plurality of sequence reads that align to the RCCX region in the reference genome sequence. In some embodiments, the computing system stores the first plurality of sequence reads in memory. The computing system can load the first plurality of sequence reads into memory. The sequence reads may be aligned to the RCCX region in the reference sequence with an alignment quality score of 0 or greater. The sequence reads may be aligned to any copy of the RCCX module in the reference sequence with an alignment quality score of about zero (e.g., when sequences are aligned to regions where genes and gene paralogs are highly homologous) or greater.
[0061] In some embodiments, the sequence reads are obtained from a digital file containing sequencing information. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotating magnetic disk drive, or a solid-state drive). In some embodiments, the digital file is stored in the form of a BAM, FASTQ, SAM, CRAM, or VCF file.
[0062] Estimation of copy number of RCCX region In some embodiments, the methods and systems disclosed herein include estimating the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads, for example, as shown in block 240 of FIG. 2A.
[0063] In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the RCCX region of the human genome. For example, the sequence reads may have previously been aligned to a reference sequence as described. In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the C4A gene, the CYP21A1P gene, the TNXA gene, the C4B gene, the CYP21A2 gene, and / or the TNXB gene in the human genome. In some embodiments, sequence reads that align to any copy of the RCCX module (such as any copy of a gene within the RCCX region) are counted.
[0064] In some embodiments, estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to regions corresponding to positions chr6:32024461 to chr6:32043719 of reference genome hg38, chr6:31991723 to chr6:32010985 of reference genome hg38, chr6:31992238 to chr6:32011496 of reference genome hg19, and / or chr6:31959500 to chr6:31978762 of reference genome hg19. In some embodiments, sequence reads are counted if they align to one or more sites within the aforementioned positions.
[0065] In some embodiments, estimating copy number includes normalizing counts of sequence reads that align to RCCX regions of the human genome. In some embodiments, sequence read counts are normalized by the length of the RCCX region. In some embodiments, read counts may be further normalized by the length of the region to a set of 3000 genomic regions of 2000 bp that are consistently predicted to be diploid across populations. In some embodiments, determining the normalized counts of sequence reads aligned to RCCX regions includes normalizing using (1a) the depth of sequence reads aligned to the RCCX region, (1b) the respective lengths of the RCCX regions (e.g., the length of each RCCX module), (2a) the depth of sequence reads aligned to the diploid regions, and (2b) the respective lengths of the diploid regions.
[0066] In some embodiments, estimating copy number includes GC-correcting sequence read counts. For example, in some embodiments, sequence read counts for the RCCX region (e.g., sequence read counts normalized by region length) are pooled with sequence read counts for a diploid region comprising approximately 3,000 distinct 2-kb regions (e.g., sequence read counts normalized by region length). Normalizing the counts of sequence reads aligning to the RCCX region by the counts of sequence reads aligning to the diploid region can, in some embodiments, correct for bias in sequencing coverage due to variable GC content between different regions. For example, the counts of sequence reads aligned to each of one or more target regions can be corrected for GC content using (1) the GC content of each of the RCCX region and (2) the GC content of each of the diploid region. In some embodiments, a normalized and / or GC-corrected copy number is determined for the RCCX region.
[0067] In some embodiments, estimating the copy number comprises using a Gaussian mixture model to bin the normalized counts of sequence reads that align to the RCCX region of the human genome. For example, a Gaussian mixture model can be used to infer the most likely copy number of the RCCX region based on the observed normalized depth signal.
[0068] The total copy number can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The Gaussian mixture model can include a one-dimensional Gaussian mixture model. The multiple Gaussians of the Gaussian mixture model can represent integer copy numbers, for example, 0 to 5, 0 to 6, 0 to 7, 0 to 8, 0 to 9, 0 to 10, 0 to 11, 0 to 12, 0 to 13, 0 to 14, or 0 to 15. For example, the multiple Gaussians of the Gaussian mixture model can represent integer copy numbers from 0 to 10. The average of each of the multiple Gaussians can be the integer copy number represented by the Gaussian. The average of each of the multiple Gaussians can be the integer copy number represented by the Gaussian (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more copies). The standard deviation of the Gaussians can be, for example, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 or more, or can be about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1 or more. The plurality of Gaussians of the Gaussian mixture model can include, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more Gaussians. For example, the plurality of Gaussians of the Gaussian mixture model can include 5 Gaussians.
[0069] To estimate the copy number of the RCCX region, a computing system can determine the copy number using a Gaussian mixture model and a predetermined posterior probability threshold given the normalized number of sequence reads aligned to the RCCX region. The posterior probability threshold can be, for example, 0.7, 0.75, 0.8, 0.85, 0.95, or higher. In some embodiments, the predetermined posterior probability threshold is 0.95.
[0070] Predetermined differentiation site In some embodiments, the methods and systems disclosed herein use a predetermined differentiation site. In some embodiments, the predetermined differentiation site comprises a site where the sequences of the CYP21A2 gene and the CYP21A1P pseudogene differ. In some embodiments, the sequences of the CYP21A2 gene and the CYP21A1P pseudogene differ at the predetermined differentiation site in at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% of a population of nucleic acid samples.
[0071] In some embodiments, in the reference genome hg38, the plurality of predetermined differentiation sites are chr6:32038514, chr6:32038844, chr6:32039015, chr6:32039081, chr6:32039128, chr6:32039132, chr6:32039143, chr6:32039426, chr6:3203954 of the CYP21A2 gene. 8, chr6:32039802, chr6:32039807, chr6:32039810, chr6:32039816, chr6:32040110, chr6:32040182, chr6:32040216, chr6:32040421, or chr6:32040535, or the corresponding position in the pseudogene CYP21A1P.
[0072] In some embodiments, in the reference genome hg19, the plurality of predetermined differentiation sites are chr6:32006291, chr6:32006621, chr6:32006792, chr6:32006858, chr6:32006905, chr6:32006909, chr6:32006920, chr6:32007203, chr6:3200732 of the CYP21A2 gene. 5, chr6:32007579, chr6:32007584, chr6:32007587, chr6:32007593, chr6:32007887, chr6:32007959, chr6:32007993, chr6:32008198, or chr6:32008312, or the corresponding position in the pseudogene CYP21A1P.
[0073] The table below lists 18 defined differentiation sites, 11 of which are pathogenic gene conversion variants. The locations listed in the table below are derived from chromosome 6 of the reference genome hg38.
[0074] [Table 1]
[0075] In another aspect, a method and system for identifying multiple predetermined differentiation sites are disclosed. In some embodiments, the method includes identifying single-base differences between the sequences of the CYP21A2 and CYP21A1P genes in a reference sequence. For example, the reference sequence of the CYP21A2 gene and the reference sequence of the CYP21A1P gene can be aligned and compared to each other to identify all sites containing single-base differences between the two gene sequences. The locations of these differentiation sites in both the CYP21A2 and CYP21A1P genes can then be stored in an electronic storage device. For example, a digital file containing a list of single-base differences can be created.
[0076] In some embodiments, the method includes selecting a single base difference that is fixed across the entire population as the differentiation site. For example, the method may include receiving a plurality of sequence reads that align to the CYP21A2 gene and the CYP21A1P gene for a plurality of nucleic acid samples (e.g., a plurality of nucleic acid samples from a population of individuals). In some embodiments, the plurality of nucleic acid samples are derived from individuals from a population of more than 100, more than 500, more than 1,000, more than 5,000, or more than 10,000 individuals. In some embodiments, the plurality of samples is obtained from the 1000 Genomes Project. In some embodiments, the population is a diverse population, such as a genetically diverse population including individuals from multiple ethnic groups, which takes into account differences between population types and increases the likelihood that the single base difference does not involve differences between population types. The method may further include estimating the gene-specific copy number of the CYP21A2 gene and the copy number of the CYP21A1P gene for each of the plurality of nucleic acid samples. The method may further include selecting a subset of nucleic acid samples from the plurality of nucleic acid samples, the subset of nucleic acid samples including nucleic acid samples predicted to be diploid for the CYP21A2 gene and predicted to be diploid for the CYP21A1P gene (e.g., using only data from samples predicted to be free of recombination events between the CYP21A2 gene and the CYP21A1P gene). The method may further include selecting single base differences having copy numbers consistent with diploidy for the CYP21A2 gene and the CYP21A1P gene in at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% of the nucleic acid samples of the subset of nucleic acid samples.
[0077] The method may further include generating a digital file containing a plurality of predetermined differentiation sites by creating a digital file listing the positions of the selected single-base differences. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotary magnetic disk drive, or a solid-state drive). In some embodiments, the digital file is saved in the format of a BAM, SAM, FASTQ, CRAM, JSON, or VCF file. The digital file may include information about the predetermined differentiation sites, such as the chromosome name on which the predetermined differentiation site is located, the 1-base global start position of CYP21A1P, the predicted base sequence for CYP21A1P reads mapped to the start position of CYP21A1P, the 1-base global start position of CYP21A2, the predicted base sequence for CYP21A2 reads mapped to the start position of CYP21A2, the CYP21A1P region corresponding to the CYP21A2 start position, a unique name for the predetermined differentiation site, and / or the orientation of the predetermined differentiation site given by the orientation of the gene.
[0078] Constructing one or more candidate haplotypes In some embodiments, methods and systems construct one or more candidate haplotypes, as shown in process block 250 of FIG. 2A . In some embodiments, methods and systems phase a plurality of sequence reads that align to the CYP21A2 gene or the CYP21A1P gene of the human genome. In some embodiments, the sequence reads include at least two predetermined sites of differentiation in the CYP21A2 gene and the CYP21A1P gene. In some embodiments, phasing the predetermined sites of differentiation includes constructing one or more candidate haplotypes based on all sequenced bases of a first predetermined site of differentiation and extending the one or more candidate haplotypes to a second predetermined site of differentiation by aligning sequence reads of the CYP21A2 gene or the CYP21A1P gene.
[0079] In some embodiments, constructing one or more candidate haplotypes includes identifying at least one seed sequence read from the plurality of sequence reads. In some embodiments, the seed sequence read aligns to the CYP21A2 gene or the CYP21A1P gene and includes at least two predetermined differentiation sites of the CYP21A2 gene and the CYP21A1P gene. In some embodiments, the seed sequence read is selected from a 5' seed sequence read, a center sequence read, and a 3' seed sequence read. For example, in block 2520 of FIG. 2B, a 5' seed sequence read, a center seed sequence read, or a 3' seed sequence read is identified. In some embodiments, constructing one or more candidate haplotypes includes iteratively extending at least one seed sequence read in either the 5' or 3' direction by aligning the sequence read using the predetermined differentiation sites. For example, in block 2530 of FIG. 2B, the seed sequence read is extended by alignment along the predetermined differentiation sites.
[0080] For example, a candidate haplotype can be formed from all sequenced bases at a first predetermined differentiation site. For example, two candidate haplotypes can be formed if two bases are possible at the first predetermined differentiation site based on base calls from sequencing reads covering the first predetermined differentiation site. In some embodiments, the haplotypes are then extended to the next predetermined differentiation site by considering all sequencing reads that can be uniquely assigned to a single candidate haplotype. In some embodiments, if such sequencing reads support only a single base at the next differentiation site for a given candidate haplotype, the haplotype is extended by that base. In some embodiments, if a candidate haplotype can be extended by two possible bases at a second predetermined differentiation site, both extended possible haplotypes are included in the set of candidate haplotypes, and the set is increased by one. In some embodiments, a next extension step is performed at a third predetermined differentiation site, and the steps can be repeated until all sites have been processed. In some embodiments, this process results in a set of candidate haplotypes based on bases observed at multiple predetermined differentiation sites.
[0081] In some embodiments, this process can be run more than once using alternative differentiation sites as starting points, with extensions occurring in either the 3' or 5' direction along the haplotype. For example, at decision state 2540 of FIG. 2B, the system can determine whether to perform an extension step on any additional seed sequence reads. In some embodiments, if the process is run multiple times using alternative starting differentiation sites and / or extension directions, the final run of the process can be performed to merge partial candidate haplotypes formed from previous runs of the process. For example, instead of using the original sequencing reads as input during this final run of the process, the partial candidate haplotypes from previous runs of the process are used as if they were input sequencing reads. For example, at block 2550 of FIG. 2B, the partial candidate haplotypes are assembled into complete candidate haplotypes.
[0082] For example, the embodiment of Figure 3 schematically shows a 5' seed sequence read 310, a central sequence read 320, and a 3' seed sequence read 330. Each seed sequence read aligns to the CYP21A2 or CYP21A1P gene and contains at least two predetermined differentiation sites. In the depiction of Figure 3, each site can include a "1" allele or a "2" allele. In the embodiment of Figure 3, each seed sequence read is extended in the 3' and / or 5' direction using other sequence reads that align to the CYP21A2 or CYP21A1P gene and contain at least two predetermined differentiation sites. In the embodiment of Figure 3, partial haplotypes 340 are constructed, which are then extended with other partial haplotypes 340 using alleles at the predetermined differentiation sites to generate final candidate haplotypes 350.
[0083] In some embodiments, a computing system uses sequence reads aligned to the CYP21A2 gene or the CYP21A1P gene, which includes multiple predetermined sites of differentiation, to construct one or more candidate haplotypes derived from the CYP21A2 gene or the CYP21A1P gene, which includes multiple predetermined sites of differentiation. For example, sequence reads can be aligned to a reference sequence so that the sequence reads overlap with the predetermined sites of differentiation. Sequence reads can be aligned to regions of the CYP21A2 gene or corresponding regions of the CYP21A1P gene that include multiple predetermined sites of differentiation with an alignment quality score of zero or greater.
[0084] To phase one or more haplotypes derived from the CYP21A2 gene or the CYP21A1P gene, the computing system can analyze linkage information between predetermined sites of differentiation in the plurality of predetermined sites of differentiation using sequence reads aligned to the CYP21A2 region or the CYP21A1P region, which includes a plurality of predetermined sites of differentiation. To phase one or more haplotypes derived from the CYP21A2 gene or the CYP21A1P gene, the computing system can phase one or more haplotypes derived from the CYP21A2 gene or the CYP21A1P gene using sequence reads aligned to two or more of the plurality of predetermined sites of differentiation.
[0085] In some embodiments, the one or more candidate haplotypes cover one or more breakpoints of a recombination event, for example, the one or more candidate haplotypes can cover one breakpoint, two breakpoints, three breakpoints, four breakpoints, five breakpoints, or more of a recombination event.
[0086] Detection of recombination events In some embodiments, the methods and systems detect recombination events between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on one or more candidate haplotypes. For example, the methods and systems may detect the probability of a recombination event between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on one or more candidate haplotypes.
[0087] In some embodiments, a recombination event can be detected based on a deviation from an estimated RCCX copy number of 4 and / or if at least one candidate haplotype contains both a CYP21A2-specific base and a CYP21A1P-specific base across a predetermined differentiation site. For example, in some embodiments, a deletion recombination event is detected if the estimated copy number of the RCCX region is 3 or less and / or if a deletion recombination variant is detected among one or more candidate haplotypes.
[0088] For example, in the embodiment of Figure 3, the candidate haplotype "22211111111121" may represent a breakpoint between the first three predetermined differentiation sites (including the pseudogene CYP21A1P alleles at these sites) and the fourth predetermined differentiation site (starting with the "1" column, representing the CYP21A2 gene alleles at these sites). Thus, recombination variants may be detected based on the candidate haplotype.
[0089] In some embodiments, if the estimated RCCX copy number is 4 and one or more candidate haplotypes do not exhibit a recombination event (e.g., each candidate haplotype includes only all CYP21A2-specific bases or all CYP21A1P-specific bases across a given differentiation site), no recombination event is detected.
[0090] In some embodiments, the methods and systems further comprise making variant calls for the recombination events. In some embodiments, the methods and systems disclosed herein further comprise creating a digital file comprising the variant calls. In some embodiments, the file comprises an estimated integer copy number for each of the one or more target regions, a floating copy number for each of the one or more target regions, and a copy number genotype. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotating magnetic disk drive, or a solid-state drive). In some embodiments, the digital file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, or VCF file. In some embodiments, the digital file is a VCF file or a JSON file.
[0091] In some embodiments, the digital file comprises one or more candidate haplotypes. In some embodiments, the digital file comprises RCCX copy number. In some embodiments, the digital file comprises information on whether a breakpoint has been detected. In some embodiments, the breakpoint is detected based on the RCCX copy number and one or more candidate haplotypes.
[0092] Methods for detecting mutations within the RCCX region In another aspect, the present specification discloses a method and a system for detecting one or more single base variants or indels in the RCCX region in a nucleic acid sample.In some embodiments, the method and the system determine sequence reads from the nucleic acid sample.For example, the sequence reads can be determined as described previously herein with reference to the method and the system for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene.
[0093] In some embodiments, methods and systems obtain sequence reads that align to the site of a single nucleotide variant or indel in the CYP21A2 or CYP21A1P gene of the human genome of a nucleic acid sample. For example, the sequence reads can be aligned to a reference genome as previously described herein with reference to methods and systems for detecting recombination events between the CYP21A2 and CYP21A1P genes. Sequence reads that align to either the CYP21A2 or CYP21A1P gene in the reference sequence can be collected, including sequence reads with low or zero mapping quality. In some embodiments, the sequence reads are derived from a short-read sequencing system or process. In some embodiments, the short-read sequence reads are about 75 bp to about 500 bp in length. In other embodiments, the short-read sequence reads are 200 bp to about 400 bp in length.
[0094] In some embodiments, the methods and systems count sequence reads that contain a base corresponding to an alternative allele at the site of a single-nucleotide variant or indel. In some embodiments, counting sequence reads includes counting both sequence reads that align to the CYP21A2 gene (and that contain the site of the single-nucleotide variant or indel) and sequence reads that align to the CYP21A1P gene (and that contain the site of the single-nucleotide variant or indel). In some embodiments, the sequence read counts may be normalized and GC-corrected as described above in this specification with reference to methods and systems for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene.
[0095] In some embodiments, methods and systems generate digital files containing variant calls corresponding to single-nucleotide variations or indels (collectively "minor variants"). In some embodiments, a minor variant is reported if a significant portion of sequence reads support an alternative allele. For example, a minor variant may be reported if about 10% or more, about 20% or more, about 30% or more, about 40% or more, about 50% or more, about 60% or more, about 70% or more, or about 80% or more, or about 90% or more of the sequence reads covering the minor variant contain base calls corresponding to the alternative allele at the site of the minor variant, compared to a reference allele at the site. In some embodiments, a minor variant may be reported if one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more sequence reads contain the alternative allele at the site of the variant.
[0096] In some embodiments, sequence reads containing the alternative allele and sequence reads containing the reference allele are counted. In some embodiments, integer copy numbers are estimated for alternative or variant alleles based on a) the total count of sequence reads covering corresponding positions of minor variants in CYP21A2 and CYP21A1P, b) the count of reads supporting the reference allele, and c) the count of reads supporting the alternative allele.
[0097] In some embodiments, the variant call is not specific to the CYP21A2 or CYP21A1P gene. For example, in some embodiments, the variant call is not phased to CYP21A2 or CYP21A1P or to one of the candidate haplotypes further described herein. In some embodiments, minor variants may be located farther than one sequence read length (e.g., farther than 100 bp, 150 bp, 200 bp, 250 bp, 400 bp, 450 bp, or more) from one or more target regions described herein. In some embodiments, by ambiguating variant calls to CYP21A2 or CYP21A1P, users can detect one or more single-nucleotide variants or indels in the RCCX region in a nucleic acid sample while using computational power and memory more efficiently, because the detected minor variants do not need to be phased to candidate haplotypes, and the methods and systems do not require further analysis of sequence reads to determine whether the minor variant is assigned to CYP21A2 or CYP21A1P. In some embodiments, detecting minor variants in a region ambiguous manner improves computational resource efficiency and allows for high precision and reproducibility in variant allele discovery compared to de novo minor variant calling or minor variant calling and phasing minor variants to regions or haplotypes, which require a much more complex process, are much less computationally efficient, and may result in lower precision or reproducibility for variants of interest.
[0098] In some embodiments, ambiguous variant calling for the CYP21A2 or CYP21A1P gene allows users to advantageously detect minor variants using short-read sequencing. Without being bound by theory, in some embodiments, short-read sequencing reads spanning the CYP21A2 or CYP21A1P gene (e.g., sequence reads comprising approximately 75-500 bp) do not contain sufficient information to uniquely locate minor variants, and users do not necessarily need to know the unique location of the variants. In some embodiments, an advantage of making ambiguous calls for a region is that users avoid the need to perform a more extensive sequencing assay, such as a long-read sequencing assay. The required information can be obtained from the same whole genome sequencing (WGS) assay used to perform variant calling on the rest of the genome.
[0099] In some embodiments, once an ambiguous variant call is made for CYP21A2 or CYP21A1P, the location of the single nucleotide variant or indel in the CYP21A2 or CYP21A1P gene can be confirmed using orthogonal (long-read) sequencing methods known to those skilled in the art. For example, the single nucleotide variant or indel is detected in a manner that is not specific for the CYP21A2 or CYP21A1P gene, and then additional sequencing, such as orthogonal techniques, is used to confirm the variant call and / or phase the variant into regions.
[0100] In some embodiments, the single nucleotide variants or indels are NM_000500.9:c.60G>A, NM_000500.9:c.92C>A, NM_000500.9:c.111del, NM_000500.9:c.159_160del, NM_000500.9:c.169G>A, NM_000500.9:c.274A>G, NM_000500.9:c.432_339del, NM_000500.9:c. .418G>A, NM_000500.9:c.421G>A, NM_000500.9:c.515T>A, NM_000500.9:c.710_719delinsACGAGGAGAA,NM_00050 0.9:c.850A>G, NM_000500.9:c.874G>A, NM_000500.9:c.922T>G, NM_000500.9:c.923_924dup,NM_000500.9:c.952 C>T=, NM_000500.9:c.955C>G, NM_000500.9:c.1042G>A, NM_000500.9:c.1051G>A, NM_000500.9:c.1066C>T=, NM_ 000500.9:c.1070G>A,NM_000500.9:c.1096C>T, NM_000500.9:c.1118G>A,NM_000500.9:c.1136T>A,NM_000500.9 :c.1226G>A,NM_000500.9:c.1273G>A,NM_000500.9:c.1274G>T,NM_000500.9:c.1279C>T,NM_000500.9:c.1357C >T=,NM_000500.9:c.1360C>T, NM_000500.9:c.1444C>T, NM_000500.9:c.1450dup, or NM_000500.9:c.1451G>A.
[0101] The table below lists the small variant sites, including 8 indels and 25 single nucleotide variants. The positions listed in the table below are from chromosome 6 of the reference genome hg38.
[0102] [Table 2-1]
[0103] [Table 2-2]
[0104] In some embodiments, the methods and systems disclosed herein further comprise creating a digital file comprising the variant calls. In some embodiments, the file comprises, for each single nucleotide variant or indel, a reference to the minor variant, a count of sequence reads supporting the alternative allele, and a count of sequence reads supporting the reference allele. In some embodiments, the digital file is on a computer storage medium (such as a computer hard drive, e.g., a rotating magnetic disk drive, or a solid-state drive). In some embodiments, the digital file is stored in the form of a BAM, SAM, CRAM, FASTQ, JSON, or VCF file. In some embodiments, the digital file is a VCF file or a JSON file. In some embodiments, the digital file also comprises recombination event detection information as described herein.
[0105] Sequencing System Embodiments 4A illustrates a diagram of an environment in which a recombination event detection system may operate according to one or more implementations. The following paragraphs describe the recombination event detection system with respect to example implementations and illustrative diagrams illustrating embodiments. For example, FIG. 4A illustrates a schematic diagram of a computing system 4106 in which a recombination event detection system 4000 operates according to one or more implementations. As illustrated, the computing system 4000 includes one or more server device(s) 4102 connected to a user client device 4108, a local device 4118, and a sequencing device 4114 via a network 4112. The network 4112 may include any suitable network over which computing devices can communicate.
[0106] 4A , the computing system 4000 includes server device(s) 4102. In various implementations, the server device(s) 4102 can generate, receive, analyze, store, and transmit digital data, such as nucleobase calls or data of sequenced nucleic acid polymers. In some implementations, the server device(s) 4102 receive various data, such as data from sample genomes and / or sequence reads, from a sequencing device 4114. Additionally, the server device(s) 4102 can communicate with a user client device 4108. Among other things, the server device(s) 4102 can transmit data related to sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 4108.
[0107] As shown, server device(s) 4102 include a sequencing application 4110. Generally, the sequencing application 4110 analyzes data (e.g., call data) received from a sequencing device 4114 or elsewhere to determine the nucleobase sequence of a nucleic acid polymer. For example, the sequencing application 4110 can receive raw data from the sequencing device 4114 and determine the nucleobase sequence of a sample genome or nucleic acid segment. In some implementations, the sequencing application 4110 determines the sequence of nucleobases in DNA and / or RNA segments or oligonucleotides.
[0108] As further shown, the sequencing application 4110 includes a recombination event detection system 4106. As described below, the recombination event detection system 4106 can detect recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample. For example, in some embodiments, the recombination event detection system 4106 receives sequence reads that align to the RCCX region of the human genome in the nucleic acid sample. The recombination event detection system 4106 further estimates the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads. The recombination event detection system 4106 can further construct one or more candidate haplotypes by phasing a plurality of sequence reads that align to the CYP21A2 gene or the CYP21A1P gene in the human genome and include at least two predetermined differentiation sites in the CYP21A2 gene and the CYP21A1P gene. The recombination event detection system 4106 further detects recombination events between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on one or more candidate haplotypes.
[0109] Additionally, while the recombination event detection system 4106 has been described as being implemented on the server device(s) 4102 as part of the sequencing application 4110, in some implementations the recombination event detection system 4106 is implemented (e.g., located entirely or partially) on the user client device 4108, the sequencing device 4114, and / or the local device 4118. As mentioned above, in some implementations the recombination event detection system 4106 is implemented by one or more other components of the computing system 4000, such as the sequencing device 4114. Notably, the recombination event detection system 4106 may be implemented in a variety of different ways across the server device(s) 4102, the network 4112, the user client device 4108, the local device 4118, and the sequencing device 4114.
[0110] 4A , the computing system 4000 includes a user client device 4108. In various implementations, the user client device 4108 is capable of generating, storing, receiving, and transmitting digital data. Among other things, the user client device 4108 can receive data from a sequencing device 4114. As further illustrated, the user client device 4108 includes a sequencing application 4110. The sequencing application 4110 can be a web application or a native application (e.g., a mobile application, a desktop application, or a web application) stored and executed on the user client device 4108. The sequencing application 4110 can receive data from the sequencing application 4110 and / or the recombination event detection system 4106. For example, the user client device 4108 can receive a variant call file and / or an alignment file from the sequencing application 4110.
[0111] Additionally, the sequencing application 4110 may include instructions that (when executed) cause the user client device 4108 to receive data from the recombination event detection system 4106 and present data from the sequencing device 4114 and / or server device(s) 4102. Additionally, the sequencing application 4110 may instruct the user client device 4108 to display data for the nucleobase calls and / or variant calls, such as one or more candidate haplotypes. Indeed, the user client device 4108 may display the nucleobase call results for the genomic sample and / or an indication of the detected recombination events between the CYP21A2 and CYP21A1P genes.
[0112] As further seen in FIG. 4A , the computing system 4000 includes a sequencing device 4114. In various implementations, the sequencing device 4114 can sequence a genomic sample or other nucleic acid polymer. For example, the sequencing device 4114 analyzes nucleic acid segments, or oligonucleotides, extracted from the genomic sample and generates data either directly or indirectly on the sequencing device 4114. More specifically, the sequencing device 4114 receives and analyzes nucleic acid sequences extracted from the genomic sample in a nucleotide sample slide (such as a flow cell). In one or more implementations, the sequencing device 4114 can utilize SBS to sequence the genomic sample or other nucleic acid polymer. In addition to or as an alternative to communicating via the network 4112, in some implementations, the sequencing device 4114 bypasses the network 4112 and communicates directly with the user client device 4108.
[0113] 4A , in some implementations, the server device(s) 4102 include a distributed collection of servers, where the server device(s) 4102 are distributed throughout the network 4112 and include several server devices located in the same physical location or different physical locations. For example, the server device(s) 4102 can be implemented in whole or in part on a local device 4118. By way of example, the local device 4118 can implement the sequencing application 4110 and / or the recombination event detection system 4106. Furthermore, the server device(s) 4102 and / or the local device 4118 can include a content server, an application server, a communication server, a web hosting server, or other types of servers.
[0114] 4A can include various types of client devices. For example, in some implementations, the user client device 4108 includes a non-mobile device such as a desktop computer or a server, or other types of client devices. In various implementations, the user client device 4108 includes a mobile device such as a laptop, tablet, mobile phone, or smartphone.
[0115] 4A illustrates components of computing system 4000 communicating over network 4112, in certain implementations, components of computing system 4000 may also communicate directly with each other, bypassing network 4112. For example, in some implementations, user client device 4108 communicates directly with sequencing device 4114. Further, in some implementations, user client device 4108 communicates directly with recombination event detection system 4106 and / or server device(s) 4102. In some implementations, user client device 4108 communicates directly with local device 4118. Further, recombination event detection system 4106 may have access to one or more databases housed on or accessed by server device(s) 4102 or elsewhere within computing system 4000.
[0116] FIG. 4B is a block diagram of an exemplary server device 4102 that can be used in conjunction with the exemplary computing system 4000 of FIG. 4A. The server device 4102 can be configured to detect recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample. The general architecture of the server device 4102 depicted in FIG. 4B includes an arrangement of computer hardware and software components. The server device 4102 can include more (or fewer) elements than those shown in FIG. 4B. However, not all of these generally conventional elements need be shown to provide a useful disclosure. As illustrated, the server device 4102 includes a processing unit 410, a network interface 420, a computer-readable medium drive 430, an input / output device interface 440, a display 450, and an input device 460, all of which can communicate with each other via a communication bus. The network interface 420 can provide connectivity to one or more networks or computing systems. Thus, the processing unit 410 can receive information and instructions from other computing systems or services via a network. Additionally, the processing unit 410 communicates with memory 470 and may also provide output information for an optional display 450 via an input / output device interface 440. Additionally, the input / output device interface 440 may accept input from any input device 460, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0117] The memory 470 may store computer program instructions (which in some embodiments are grouped as modules or components) that the processing unit 410 executes to implement one or more embodiments. Generally, the memory 470 includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. The memory 470 may store an operating system 472 that provides computer program instructions used by the processing unit 410 in the general management and operation of the server device 4102. The memory 470 may store a reference genome 473, such as used by the sequencing application 4110. The memory 470 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0118] For example, in one embodiment, memory 470 includes a sequencing application 4110, which may include a recombination event detection system 4106. Recombination event detection system 4106 is capable of performing the methods disclosed herein. Additionally, memory 470 may include or be in communication with a data store 490 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) related to detecting recombination events between the CYP21A2 and CYP21A1P genes in a nucleic acid sample of the present disclosure, such as sequencing reads, estimated copy number(s), one or more candidate haplotypes, and determined variant calls (e.g., detection of recombination events).
[0119] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis functions and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genomic data, or other types of biological data may be mediated through a central hub that stores the data and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide shared protocols, analytical methods, libraries, sequence data, and distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates user correction or annotation of sequence data. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand, or online.
[0120] In some embodiments, software written to perform the methods described herein is stored on some form of computer readable medium, such as memory, CD-ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system, etc.
[0121] In some embodiments, the method may be written in any of a variety of suitable programming languages, for example, compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages may include scripting languages such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the method is written in C, C#, C++, Fortran, Java, Perl, R, Java, or Python. In some embodiments, the method may be a stand-alone application having data entry and data display modules. Alternatively, the method may be a computer software product, and distributed objects may include classes that include applications that include the computational methods described herein.
[0122] In some embodiments, the methods can be incorporated into existing data analysis software, such as that found on sequencing instruments. Software comprising the computer-implemented methods described herein can be installed directly on a computer system or indirectly stored on a computer-readable medium and loaded onto the computer system as needed. Additionally, the methods can be located on a computer that is remote from where the data is generated, such as software found on a server maintained at a separate location from where the data is generated, such as that provided by a third-party service provider.
[0123] The assay instrument, desktop computer, laptop computer, or server may include a processor in operative communication with accessible memory containing instructions for implementing the systems and methods. In some embodiments, the desktop computer or laptop computer is in operative communication with one or more computer-readable storage media or devices and / or output devices. The assay instrument, desktop computer, and laptop computer may operate under many different computer-based operating languages, such as those utilized by Apple-based or PC-based computer systems. The assay instrument, desktop computer, and / or laptop computer and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results, and monitoring experimental progress. In some embodiments, the output device may be a graphic user interface, such as a computer monitor or computer screen, a printer, a handheld device, such as a personal digital assistant (i.e., PDA, Blackberry, iPhone, etc.), a tablet computer (iPAD, etc.), a hard drive, a server, a memory stick, a flash drive, etc.
[0124] The computer-readable storage device or medium may be any device, such as a server, mainframe, supercomputer, magnetic tape system, etc. In some embodiments, the storage device may be located in close proximity to the assay device, e.g., adjacent to or in close proximity to the assay device. For example, the storage device may be located relative to the assay device in the same room, in the same building, in an adjacent building, on the same floor within a building, on a different floor within a building, etc. In some embodiments, the storage device may be located external to or distal to the assay device. For example, the storage device may be located in a different part of a city, in a different city, in a different state, or in a different country relative to the assay device. In embodiments in which the storage device is located distal to the assay device, communication between the assay device and one or more of a desktop, laptop, or server is typically via an internet connection, either wirelessly or via a network cable via an access point. In some embodiments, the storage device may be maintained and managed by an individual or entity directly associated with the assay device, while in other embodiments, the storage device may be maintained and managed by a third party, typically in a location distal to the individual or entity associated with the assay device. In the embodiments described herein, the output device can be any device for visualizing data.
[0125] The assay instrument, desktop, laptop, and / or server system may be used to store and / or retrieve computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, etc. One or more of the assay instrument, desktop, laptop, and / or server may include one or more computer-readable storage media for storing and / or retrieving software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, etc. Computer-readable storage media may include, but are not limited to, one or more of a hard drive, SSD hard drive, CD-ROM drive, DVD-ROM drive, floppy disk, tape, flash memory stick, or card. Furthermore, a network, including the Internet, may be a computer-readable storage medium. In some embodiments, a computer-readable storage medium refers to computational resource storage accessible by a computer network via the Internet or a corporate network provided by a service provider, rather than, for example, from a local desktop or laptop computer at a remote location to the assay instrument.
[0126] In some embodiments, computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, computer-readable storage media for storing and / or retrieving data used in implementing the computational methods, etc., are operated and maintained by a service provider in operative communication with the assay instrument, desktop, laptop, and / or server system via an internet or network connection.
[0127] In some embodiments, a hardware platform for providing a computing environment includes a processor (i.e., CPU) where processor time and memory layout, such as random access memory (i.e., RAM), are system considerations. For example, smaller computer systems offer cheaper, faster processors and greater memory and storage capabilities. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, a hardware platform for performing the computational methods described herein includes one or more computer systems with one or more processors. In some embodiments, smaller computers are clustered together to create a supercomputer network.
[0128] In some embodiments, the computational methods described herein are implemented on a collection of inter- or intra-connected computer systems (i.e., grid technology) that may cooperatively run various operating systems. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available from United Devices are illustrative of the cooperation of multiple independent computer systems for the purpose of handling large amounts of data. These systems may provide a Perl interface for submitting, monitoring, and managing large sequence analysis jobs on a cluster in a serial or parallel configuration.
[0129] Example Some aspects of the above-described embodiments are disclosed in further detail in the following examples, which are not intended to limit the scope of the disclosure. Those skilled in the art will recognize that many other embodiments are within the scope of the disclosure, as described above and in the claims.
[0130] Example 1 In the following examples, recombination events between the CYP21A2 and CYP21A1P genes were detected in multiple nucleic acid samples. The copy number of the entire RCCX region was identified, and any recombined CYP21A1P-CYP21A2 gene fusions were reported. Additionally, 33 small variants (single-nucleotide variants or indels) were detected in either the gene or pseudogene. These variant calls included all CYP21A2 variants with multiple subtypes annotated as pathogenic or likely pathogenic in ClinVar17.
[0131] Total RCCX copies The copy number of the entire RCCX region was calculated by counting reads belonging to both copies of the segmental duplication. In most cases, reads could not be unambiguously mapped to either copy of the repeat due to high sequence similarity; however, counting all reads that mapped to either copy provided a more accurate measure of the total copy number of both regions. The region used for copy number calling covered most, but not all, of the RCCX region. The region used for copy number calling began after the polymorphic 6.4 kb HERV-K retrotransposon in both the C4A and C4B introns and extended 20 kb downstream to the 120 bp deletion in TNXA, encompassing the entire CYP21A2 gene. This region corresponded to positions chr6:32024461-chr6:32043719 and chr6:31991723-chr6:32010985 in the reference genome hg38. Therefore, the copy number calling subregion of RCCX was large enough to reach any non-allelic homologous recombination events (30 kb long, affecting all copies of RCCX). Read coverage was corrected for GC content by normalization to a panel of 3000 preselected 2 kb genomic sites with highly consistent diploid copy numbers. The normalized sequence read counts were then binned using a Gaussian mixture model. This estimated copy number was an accurate estimate of the total copy number of the RCCX segment duplication.
[0132] Detection of recombinant mutants A panel of 18 predefined regions spanning CYP21A2 was used to detect gene fusions between CYP21A1P and active CYP21A2. The gene and pseudogene sequences differed at these predefined regions. The 18 predefined regions were located at chr6:32038514, chr6:32038844, chr6:32039015, chr6:32039081, chr6:32039128, chr6:32039132, chr6:32039143, chr6:32039426, and chr6:32039 in the CYP21A2 gene of the reference genome hg38. 548, chr6:32039802, chr6:32039807, chr6:32039810, chr6:32039816, chr6:32040110, chr6:32040182, chr6:32040216, chr6:32040421, and chr6:32040535, or the corresponding positions in the pseudogene CYP21A1P.
[0133] To identify recombinant variants, haplotypes present within the genome were detected. Reads spanning a set of 18 predefined differentiation sites were collected. Reads spanning multiple (i.e., two or more) predefined differentiation sites were used to construct concatenated haplotypes across the entire region. Reads containing predefined differentiation sites were collected and assembled into partial haplotypes from the 5' end, center, and 3' end of the gene. The partial haplotypes were then assembled into a final complete haplotype spanning the entire gene region. Transitions from gene-allele to pseudogene-allele sequences within the resulting haplotypes indicated either complete chimeric gene fusions or smaller gene conversion events.
[0134] Targeted detection of small mutations For a set of 33 known sites where the gene and pseudogene sequences are identical, other deleterious mutations were detected. Reads aligning to the gene or pseudogene for each of these sites were collected. The number of reads containing the reference allele and any support for deleterious alternative alleles were counted and reported. The reads were then used to provide evidence for the presence or absence of pathogenic alleles in either the gene or pseudogene. The 33 sites used were NM_000500.9:c.60G>A, NM_000500.9:c.92C>A, NM_000500.9:c.111del, NM_000500.9:c.159_160del, NM_000500.9:c.169G>A, NM_000500.9:c.274A>G, NM_000500.9:c.332_339del, NM_000500.9:c.418G>A, NM_ 000500.9:c.421G>A, NM_000500.9:c.515T>A, NM_000500.9:c.710_719delinsACGAGGAGAA, NM_000500.9:c.850 A>G, NM_000500.9:c.874G>A, NM_000500.9:c.922T>G, NM_000500.9:c.923_924dup, NM_000500.9:c.952C>T=, NM _000500.9:c.955C>G, NM_000500.9:c.1042G>A, NM_000500.9:c.1051G>A, NM_000500.9:c.1066C>T=, NM_00050 0.9:c.1070G>A, NM_000500.9:c.1096C>T, NM_000500.9:c.1118G>A, NM_000500.9:c.1136T>A, NM_000500.9:c.1 226G>A, NM_000500.9:c.1273G>A, NM_000500.9:c.1274G>T, NM_000500.9:c.1279C>T, NM_000500.9:c.1357C>T= , NM_000500.9:c.1360C>T, NM_000500.9:c.1444C>T, NM_000500.9:c.1450dup, and NM_000500.9:c.1451G>A.
[0135] Results: CAH cases from Radboud UMC (N=16, cases) WGS data were obtained for 16 cases of congenital adrenal hyperplasia (CAH), with validation from Sanger sequencing or multiplex ligation-dependent probe amplification (MLPA). For each sample, recombination events between the CYP21A2 and CYP21A1P genes were determined, and small variants were identified using the targeted method described above and validated using Sanger sequencing or MLPA. In each of these case genomes, the targeted method accurately detected the total RCCX copy number and pathogenic variants (including small variants, complete gene deletions, and inactivating gene conversions).
[0136] The validation results for each of the 16 samples are summarized in the table below, where the causative allele and total RCCX copy number are reported in each genome. The targeted method was consistent with the MLPA / Sanger results in each case. All variant IDs correspond to the NM_000500.9 transcript, respectively.
[0137] [Table 3]
[0138] Example 2 In the following example, the method described in Example 1 was used to detect recombination events between the CYP21A2 and CYP21A1P genes, along with minor mutations, in four nucleic acid samples.
[0139] The method described in Example 1 was further validated by MLPA or long-range PCR confirmation of CYP21A2 variants using four sequenced cell lines. These included a trio in which the proband, NA14734, was affected by a severe salt-wasting form of CAH. This was caused by a complete deletion of two copies of the RCCX segment duplication and a complete loss of CYP21A2, as evidenced by MLPA validation. MLPA also revealed that both parents were carriers of the CYP21A2 deletion, demonstrating inheritance of the deleterious genotype in the proband.
[0140] Each of these genotypes was identified in the trio using the methods described in Example 1. Haplotypes generated by the deletion of the RCCX module were constructed to estimate the total RCCX copy number in each family member. Detailed information obtained from the candidate haplotypes also provided insight into the inherited alleles of the CYP21A1P pseudogene.
[0141] Figure 5 shows a schematic of the recombinant haplotypes constructed in the CAH case trio. In Figure 5, each haplotype is simplified to a series of one or two identifiers, indicating the gene (1) or pseudogene (2) allele at each site. The haplotype of CAH-affected proband NA14734 contained a copy of an RCCX segmental duplication with an inactive pseudogene CYP21A1P allele at most sites and no copies of the wild-type CYP21A2 gene. This identified the most likely parental origin of the two RCCX copies in the proband. A copy number call of 3 in each parent also indicates a risk of wild-type gene deletion. Each parent was identified as a possible CAH carrier due to their reduced RCCX copy number. The proband lacking both copies of the active gene was identified as a likely CAH case.
[0142] A fourth CAH cell line (NA12217) also represented a CAH case, but was affected by a milder, simpler, and virilizing form of the disorder. In this genome, MLPA and long-range PCR validation identified a single deletion in one copy of RCCX and the exonic single-nucleotide variant NM_000500.9:c.518T>A, which carries a known CAH risk. RCCX copy number was estimated and candidate haplotypes were constructed using the methods described in Example 1. The results are shown in the table below.
[0143] [Table 4]
[0144] As can be seen in the table above, the NM_000500.9:c.518T>A variant was identified in one haplotype, leading to an estimated total RCCX copy number of 4. A chimeric pseudogene-gene fusion likely resulting from a recombination-mediated deletion was also identified. This deletion event, combined with the total RCCX copy number of 4, indicated that this genome represents the result of both deletions and duplications in the RCCX module.
[0145] The "1211111211111111111" haplotype would be difficult to assemble using conventional methods but was unambiguous using the methods described in Example 1. The chimeric fusion haplotype structure is represented as "2222222111111111111," where "1" indicates the target gene allele and "2" indicates the pseudogene allele. The haplotype showed clear delineation between consistent pseudogene alleles at the first seven differentiation sites, followed by conversion to consistent target alleles at the last 11 sites, a refined representation of the fusion gene structure and deletion breakpoints.
[0146] Example 3 In the following example, the total RCCX copy number was estimated for several nucleic acid samples (N=204) using the method described in Example 1. The estimated RCCX copy numbers were compared with orthogonal sequencing results.
[0147] The RCCX total copy number calling results from the method described in Example 1 were compared to the RCCX copy number calling from the orthogonal Bionano Genomics optical mapping technology in 204 genomes from the 1000 Genomes Project cohort. The results are shown in Figure 6. The Pearson correlation coefficient and P value are annotated in the lower right corner of Figure 6.
[0148] Although optical mapping lacks the resolution to identify gene fusions or small mutations, these call comparisons demonstrated the overall copy number calling accuracy of the method described in Example 1 ("CYP21A2 targeted caller"). In 201 of 204 genomes, the copy number calls were concordant, but in three genomes there was a discrepancy for one RCCX copy. This concordance demonstrates the high accuracy of the method described in Example 1 in recovering the correct copy number of the RCCX region.
[0149] Example 4 In the following examples, 33 small variants (single-nucleotide variations or indels) were tested in either the CYP21A2 gene or the CYP21A1P pseudogene using the method described in Example 1. 3195 samples from the 1000 Genomes Project cohort were tested for the 33 small variants, and the results were examined. 11 of the 3195 (0.3%) contained strong evidence for the targeted variant (at least two supporting sequence reads from either the gene or the pseudogene). While these variant calls were highly reliable, they were not uniquely assigned to the gene or pseudogene, but rather ambiguously assigned to either the gene or the pseudogene.
[0150] Other Considerations The embodiments described herein are exemplary. Modifications, rearrangements, alternative processes, etc. may be made to these embodiments and still fall within the teachings described herein. One or more of the steps, processes, or methods described herein may be performed by one or more suitably programmed processing and / or digital devices.
[0151] The various illustrative imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. The described functionality may be implemented in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0152] The various exemplary detection systems described in connection with the embodiments disclosed herein may be implemented or performed by a mechanical apparatus such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively, the processor may be a controller, microcontroller, or state machine, combinations thereof, or the like. The processor may also be implemented with a combination of computing devices, such as, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or other such configurations. For example, the systems described herein may be implemented using discrete memory chips, a portion of memory within a microprocessor, flash, EPROM, or other types of memory.
[0153] Elements of a method, process, or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The software module may include computer-executable instructions that cause a hardware processor to execute the computer-executable instructions.
[0154] In particular, conditional language used herein, such as "can," "might," "may," "for example," and the like, is intended to generally convey that certain embodiments include certain features, elements, and / or conditions, while other embodiments do not include certain features, elements, and / or conditions, unless otherwise specified or understood differently within the context in which it is used. Thus, such conditional language generally does not imply that features, elements, and / or conditions are required in any manner for one or more embodiments, or that one or more embodiments necessarily include logic for determining, with or without author input or prompting, whether these features, elements, and / or conditions are included or performed in any particular embodiment. Terms such as "comprise," "including," "having," and "involving" are synonymous and used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or," when used to connect a list of elements, for example, is used in its inclusive sense (rather than its exclusive sense), so as to mean one, some, or all of the elements in the list.
[0155] Unless otherwise indicated, disjunctive language such as "at least one of X, Y, or Z" should be understood in conjunction with the context as it is commonly used to indicate that an item, term, etc. may be either X, Y, or Z, or any combination thereof. Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires the presence of at least one of X, at least one of Y, or at least one of Z, respectively.
[0156] Terms such as "about" or "approximately" are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, which may be ±20%, ±15%, ±10%, ±5%, or ±1%. The term "substantially" is used to indicate that a result (such as a measurement) is close to a target value, where close may mean, for example, that the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.
[0157] Unless otherwise specified, articles such as "a" or "an" should generally be construed to include one or more listed items. Thus, phrases such as "a device configured to" or "a device for" are intended to include one or more listed devices. Such one or more listed devices may also be collectively configured to perform the described detailed descriptions. For example, "a processor configured to perform detailed descriptions A, B, and C" may include a first processor to perform operations in conjunction with a second processor configured to perform detailed description A and to perform operations in conjunction with a second processor configured to perform detailed descriptions B and C.
[0158] While the foregoing detailed description has illustrated, described, and pointed out novel features applied to the exemplary embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms shown may be made without departing from the spirit of the present disclosure. It will be recognized that certain embodiments described herein may be embodied in forms that do not provide all of the features and advantages described herein, since some features may be used or practiced separately from others. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope.
[0159] It is to be understood that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are intended to be part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated to be part of the inventive subject matter disclosed herein.
[0160] The scope of the present disclosure is not intended to be limited by the specific disclosure of examples in this section or elsewhere herein, but rather may be defined by the claims presented in this section or elsewhere herein or presented in the future. Claim language is to be interpreted broadly based on the language used in the claims and is not limited to the examples described herein or during the prosecution of this application, and examples are to be construed as non-exclusive. [Explanation of symbols]
[0161] 310 seed sequence reads 320 central sequence reads 330 seed sequence reads 340 partial haplotypes 350 final candidate haplotypes 4000 Computing System 4102 Server Device 4106 Recombination Event Detection System 4108 User Client Device 4110 Sequencing Applications 4112 Network 4114 Sequencing Device 4118 Local Device 410 Processing Unit 420 network interface 430 Computer Readable Media Drive 440 Input / Output Device Interface 450 display 460 Input Devices 470 memory 472 Operating Systems 473 reference genomes 490 Datastore
Claims
1. 1. A computer-implemented method for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample, comprising: receiving sequence reads that align to a RCCX region of the human genome in the nucleic acid sample; estimating the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads; constructing one or more candidate haplotypes by phasing a plurality of sequence reads that are aligned to the CYP21A2 gene or the CYP21A1P gene of the human genome and that include at least two predetermined differentiation regions of the CYP21A2 gene and the CYP21A1P gene; and detecting a recombination event between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on the one or more candidate haplotypes. A method comprising:
2. The method of claim 1 , wherein the one or more candidate haplotypes cover one or more breakpoints of the recombination event.
3. 2. The method of claim 1, wherein constructing the one or more candidate haplotypes comprises identifying at least one seed sequence read from the plurality of sequence reads.
4. 4. The method of claim 3, wherein the seed sequence read is selected from a 5' seed sequence read, a center sequence read, and a 3' seed sequence read.
5. 4. The method of claim 3, wherein constructing the one or more candidate haplotypes comprises iteratively extending at least one seed sequence read in either the 5' or 3' direction by aligning the sequence read using the predetermined differentiation site.
6. 6. The method of claim 1, wherein estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the RCCX region of the human genome.
7. 7. The method of claim 6, wherein estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to a C4A gene, a CYP21A1P gene, a TNXA gene, a C4B gene, a CYP21A2 gene, or a TNXB gene in the human genome.
8. 8. The method of claim 7, wherein estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to a region corresponding to positions chr6:32024461 to chr6:32043719 of reference genome hg38, chr6:31991723 to chr6:32010985 of reference genome hg38, chr6:31992238 to chr6:32011496 of reference genome hgl 9, or chr6:31959500 to chr6:31978762 of reference genome hgl 9.
9. 7. The method of claim 6, wherein estimating the copy number comprises normalizing the count of the sequence reads that align to the RCCX region of the human genome.
10. 10. The method of claim 9, wherein estimating the copy number comprises binning the normalized counts of the sequence reads that align to the RCCX region of the human genome using a Gaussian mixture model.
11. The method according to any one of claims 1 to 10, further comprising performing variant calling at a predetermined differentiation site among the plurality of predetermined differentiation sites.
12. The method of any one of claims 1 to 11, further comprising performing variant calling of the recombination events.
13. The method of any one of claims 1 to 12, further comprising generating a digital file comprising the variant calls.
14. The method of any one of claims 1 to 13, further comprising creating a digital file containing one or more candidate haplotypes.
15. In the reference genome hg38, the plurality of predetermined differentiation sites are chr6:32038514, chr6:32038844, chr6:32039015, chr6:32039081, chr6:32039128, chr6:32039132, chr6:32039143, chr6:32039426, chr6:32039548, chr6:32039802 of the CYP21A2 gene. , chr6:32039807, chr6:32039810, chr6:32039816, chr6:32040110, chr6:32040182, chr6:32040216, chr6:32040421, or chr6:32040535, or the corresponding position of the pseudogene CYP21A1P.
16. In the reference genome hg19, the plurality of predetermined differentiation sites are chr6:32006291, chr6:32006621, chr6:32006792, chr6:32006858, chr6:32006905, chr6:32006909, chr6:32006920, chr6:32007203, chr6:32007325, chr6:32007579 of the CYP21A2 gene, 16. The method of any one of claims 1 to 15, comprising a site corresponding to a position selected from chr6:32007584, chr6:32007587, chr6:32007593, chr6:32007887, chr6:32007959, chr6:32007993, chr6:32008198, or chr6:32008312, or the corresponding position of the pseudogene CYP21A1P.
17. 1. A computer-implemented method for detecting one or more single nucleotide variants or indels in an RCCX region in a nucleic acid sample, comprising: determining sequence reads from the nucleic acid sample; obtaining sequence reads that align to the site of the single base variant or indel in the CYP21A2 gene or the CYP21A1P gene of the nucleic acid sample in the human genome; counting sequence reads that align to the CYP21A2 gene and that align to the CYP21A1P gene; and counting sequence reads that contain a base corresponding to an alternative allele at the site of a single base variant or indel, including counting sequence reads that align to the CYP21A2 gene and that align to the CYP21A1P gene. creating a digital file containing variant calls corresponding to the single nucleotide variants or indels, wherein the variant calls are not specific to the CYP21A2 gene or the CYP21A1P gene; A method comprising:
18. The one or more single nucleotide mutations or indels are selected from the group consisting of NM_000500.9:c. 60G>A, NM_000500.9:c. 92C>A, NM_000500.9:c. 111del, NM_000500.9:c. 159_160del, NM_000500.9:c. 169G>A, NM_000500.9:c. 274A>G, NM_000500.9:c. 332_339del, NM_000500.9:c. 418G>A, NM_000500.9:c. 421G>A, NM_000500.9:c. 515T>A, NM_000500.9:c. 710_719delinsACGAGGAGAA, NM_000500.9:c. 850A>G, NM_000500.9:c. 874G>A, NM_000500.9:c. 922T>G, NM_000500.9:c. 923_924dup, NM_000500.9:c. 952C>T=, NM_000500.9:c. 955C>G, NM_000500.9:c. 1042G>A, NM_000500.9:c. 1051G>A, NM_000500.9:c. 1066C>T=, NM_000500.9:c. 1070G>A, NM_000500.9:c. 1096C>T, NM_000500.9:c. 1118G>A, NM_000500.9:c. 1136T>A, NM_000500.9:c. 1226G>A, NM_000500.9:c. 1273G>A, NM_000500.9:c. 1274G>T, NM_000500.9:c. 1279C>T, NM_000500.9:c. 1357C>T=, NM_000500.9:c.
18. The method of claim 17, comprising NM_000500.9:c. 1360C>T, NM_000500.9:c. 1444C>T, NM_000500.9:c. 1450dup, or NM_000500.9:c. 1451G>A.
19. 1. An electronic system for detecting recombination events between the CYP21A2 gene and the CYP21A1P gene in a nucleic acid sample, comprising: receiving sequence reads that align to a RCCX region of the human genome in the nucleic acid sample; estimating the copy number of the RCCX region of the human genome in the nucleic acid sample from the aligned sequence reads; constructing one or more candidate haplotypes by phasing a plurality of sequence reads that are aligned to the CYP21A2 gene or the CYP21A1P gene of the human genome and that include at least two predetermined differentiation regions of the CYP21A2 gene and the CYP21A1P gene; and detecting a recombination event between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on the one or more candidate haplotypes.
1. An electronic system comprising: a processor configured to perform a method comprising:
20. 20. The electronic system of claim 19, wherein the processor is configured to execute a method comprising detecting a recombination event between the CYP21A2 gene and the CYP21A1P gene based on the estimated copy number of the RCCX region of the human genome and based on the one or more candidate haplotypes.
21. 20. The electronic system of claim 19, wherein the one or more candidate haplotypes cover one or more breakpoints of the recombination events.
22. 20. The electronic system of claim 19, wherein constructing the one or more candidate haplotypes comprises identifying at least one seed sequence read from the plurality of sequence reads.
23. 23. The electronic system of claim 22, wherein the seed sequence read is selected from a 5' seed sequence read, a central sequence read, and a 3' seed sequence read.
24. 23. The electronic system of claim 22, wherein constructing the one or more candidate haplotypes comprises iteratively extending at least one seed sequence read in either the 5' or 3' direction by aligning the sequence read using the predetermined differentiation site.
25. 25. The electronic system of claim 19, wherein estimating the copy number of the RCCX region of the human genome comprises counting sequence reads that align to the RCCX region of the human genome.