A method for detecting fetal genetic variation using a polymorphic site and a target site sequencing
Patent Information
- Application Number
- CN202180080432.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-21
- Filing Date
- 2021-10-21
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-10-21
AI Technical Summary
然而对于亚染色体水平的微缺失/微重复变异,其无创检测方法的灵敏性和特异性还不是很高,尤其是对于小片段的微缺失/微重复变异(Advani,Barrett et al.2017,Prenat Diagn 37:1067-1075;Hu,Wang et al.2019,HumanGenomics 13:14;Srebniak,Knapen et al.2020,Mol Genet Genomic Med 8:e1062)
Smart Images

Figure HPA0000340465310000011 
Figure HPA0000340465310000012 
Figure HPA0000340465310000021
Abstract
Description
Technical Field
[0001] This invention relates to the field of genetic variation detection, particularly to aneuploidy at the chromosomal level, microdeletion / microduplication at the subchromosomal level, or insertion / deletion and single nucleotide site variation at the short sequence level. Background Technology
[0002] In 1997, cell-free DNA derived from the fetus was discovered in the plasma of pregnant women (Lo, Corbetta et al. 1997, Lancet 350: 485-487). Based on this discovery and massively parallel sequencing, multiple research groups have developed methods based on sequencing analysis of pregnant women's plasma DNA (cfDNA) to detect chromosomal aneuploidy, microdeletion / microduplication at the subchromosomal level, or short sequence insertions / deletions and single nucleotide polymorphisms at the single gene level (Advani, Barrett et al. 2017, Prenat Diagn 37: 1067-1075; Breveglieri, D'Aversa et al. 2019, Mol Diagn Ther 23: 291-299; Andari, Bussamra et al. 2020, Ceska Gynekol 85: 41-48; Guseh 2020, Hum Genet 139: 1141-1148).
[0003] Currently, the use of next-generation sequencing for the detection of chromosomal aneuploidy is recognized and industrialized in many countries around the world due to its high sensitivity and specificity (Chiu, Chan et al. 2008, Proc Natl Acad Sci USA 105: 20458-20463; Fan, Blumenfeld et al. 2008, Proc Natl Acad Sci US A105: 16266-16271; Liao, Chan et al. 2012, PLoS One 7: e38154; Zimmermann, Hill et al. 2012, Prenat Diagn 32: 1233-1241). However, the sensitivity and specificity of non-invasive detection methods for microdeletion / microduplication variants at the subchromosomal level are not very high, especially for small fragment microdeletion / microduplication variants (Advani, Barrett et al. 2017, Prenat Diagn 37: 1067-1075; Hu, Wang et al. 2019, HumanGenomics 13: 14; Srebniak, Knapen et al. 2020, Mol Genet Genomic Med 8: e1062). Although various non-invasive detection methods for single-gene genetic diseases based on next-generation sequencing have been developed (Lun, Tsui et al. 2008, Proc Natl Acad Sci USA 105: 19920-19925; Lo, Chan et al. 2010, SciTransl Med 2: 61ra91; Lv, Wei et al. 2015, Clinical Chemistry 61: 172-181; Vermeulen, Geeven et al. 2017, Am J Hum Genet 101: 326-339; Allen, Young et al. 2018, Noninvasive Prenatal Testing (NIPT) 157-177; Yin, Du et al. 2018, J HumGenet 63: 1129-1137; Cutts, Vavolis et al. 2019, Blood 134:1190-1193; Zhang, Li et al. 2019, Nat Med 25:439-447), but these methods have not been widely used in clinical practice, mainly because they use different methods than those for detecting chromosomal or subchromosomal level variations, and therefore cannot be used to detect chromosomal and subchromosomal level variations at the same time.At the same time, these methods are very expensive to test for each single-gene genetic disease, making them not cost-effective for screening for single-gene genetic diseases with low incidence rates.
[0004] Therefore, a universal method that can simultaneously detect fetal chromosomal, subchromosomal, and single-gene short sequence variations using maternal plasma DNA would greatly benefit non-invasive detection of fetal genetic variations. Summary of the Invention
[0005] The purpose of this invention is to provide a method for simultaneously detecting chromosomal aneuploidy, subchromosomal microdeletion / microduplication, and single-gene genetic diseases caused by short sequence variations.
[0006] To achieve the above objectives, this invention designs a method for screening genetic variations based on high-throughput sequencing technology, including obtaining test samples and extracting DNA, selectively amplifying target sites, performing high-throughput sequencing on the target sites, and analyzing the sequencing data to obtain detection results.
[0007] This invention provides a method for detecting genetic variations, which includes the following steps:
[0008] (1) Receive biological samples to be tested and prepare nucleic acids;
[0009] (2) Enrich or amplify target DNA sites, wherein at least one target DNA site has more than one allele in the sample;
[0010] (3) Sequencing the amplified target DNA sites;
[0011] (4) For each target DNA site, count the number of each allele.
[0012] (5) Use the goodness-of-fit test of allele counts at target DNA sites and / or the relative distribution map of allele counts to determine the karyotype, genotype or wild-type mutant of the target to be detected in the sample.
[0013] This invention provides detection of chromosomal aneuploidy, subchromosomal microdeletion / microduplication, and short sequence fragment variation in mixed samples through amplification and sequencing technology of specific target DNA sites, wherein at least one of the specific target DNA sites has more than one allele in the sample.
[0014] The target DNA site described in this invention refers to a specific DNA sequence in which the bases may vary between different individuals, and this DNA sequence can be amplified by techniques such as PCR, multiplex PCR, or enriched by techniques such as nucleic acid hybridization. In this invention, the terms "target DNA sequence" and "target DNA site" are used interchangeably, and the term "site" does not limit the length of the target when referring to the target; that is, the length of the target can range from a single nucleotide to the length of an entire chromosome.
[0015] In another aspect, the present invention provides a method for detecting aneuploidy at the chromosome level and microdeletion / microduplication at the subchromosomal level in a single genome sample by amplification and sequencing of specific DNA sites (target sites), wherein at least one of the specific target DNA sites has more than one allele in the sample.
[0016] The biological samples described in this invention include fetal and maternal nucleic acids (such as cell-free DNA in maternal plasma) from the biological samples of the pregnant women or from single-genome samples (such as embryonic nucleic acids derived from preimplantation diagnostics).
[0017] The enrichment or amplification of target DNA sites described in this invention can be performed using any method known in the art, including but not limited to PCR, multiplex PCR, whole-genome amplification (WGA), multiple substitution amplification (MDA), rolling circle amplification (RCA), circular amplification (RCR), and hybridization capture. Among the enriched or amplified target DNA sites, a portion originates from regions of one or more chromosomes presumed to be euploid, and a portion originates from regions of one or more chromosomes suspected of having chromosomal, subchromosomal, or short-sequence-level variations to be tested. Chromosomes, regions, or sites presumed to be euploid are hereby designated as “reference chromosomes, reference regions, reference sequences, or reference sites,” and chromosomes, regions, or sites presumed to be indicative of genetic variation states to be detected are hereby designated as “target chromosomes, target regions, target sequences, or target sites.” In this invention, a collection consisting of at least one or more reference chromosomes, reference regions, reference sequences, or reference sites is referred to as a reference group. In this invention, a collection consisting of at least one or more target chromosomes, target regions, target sequences, or target sites is referred to as a target group.
[0018] The counting of alleles at each target DNA site, as described in this invention, refers to first mapping each amplified sequence to a chromosomal or genomic location, and then counting the number of mapped sequences in each chromosomal or genomic region. If a chromosomal or genomic region has different alleles, the number of sequences mapped to each allele in that region is also counted. Various computer methods can be used to map sequence readings to chromosomal or genomic locations / regions. Non-limiting examples of computer algorithms that can be used for sequence mapping include, but are not limited to, specific sequence lookup, BLAST, BLITZ, FASTA, BOWTIE, BOWTIE 2, BWA, NOVOALIGN, GEM, ZOOM, ELAN, MAQ, MATCH, SOAP, STAR, SEGEMEHL, MOSAIK, or SEQMAP, or variations thereof, or combinations thereof.
[0019] In this invention, for ease of understanding, a subchromosomal microdeletion is considered as one chromosome, and a subchromosomal microduplication is considered as two chromosomes. Therefore, for single-genome samples, chromosomes with heterozygous microdeletions at the subchromosomal level are marked as monosomy, chromosomes with homozygous microdeletions are marked as absent chromosomes, chromosomes with heterozygous microduplications are marked as trisomy, and chromosomes with homozygous microduplications are marked as tetrasomy. Correspondingly, in mixed samples, such as maternal plasma samples, chromosomes with normal maternal and fetal chromosomes are marked as disomy-disomy, chromosomes with a microdeletion in a normal maternal fetus are marked as disomy-monosomy, and chromosomes with a microduplication in a normal maternal fetus are marked as disomy-trisomy. In this invention, chromosomes and / or chromosome segments involving variations at the chromosomal or subchromosomal level are marked according to similar principles.
[0020] In this invention, subchromosomal microdeletions / microduplications refer to chromosomal aberrations where the deleted or added segments on a chromosome are not very long and are difficult to detect through traditional cytogenetic analysis. Chromosomal microdeletion / microduplication syndromes are another major category of neonatal birth defects besides chromosomal aneuploidy. In this invention, copy number variations of chromosomal segments are also used in some parts to refer to chromosomal microdeletion / microduplication variations.
[0021] In this invention, karyotype refers to variations at the chromosomal or subchromosomal level, while genotype refers to variations at the short sequence level. For example, for a maternal plasma sample, if the mother's chromosome 21 is normally disomic and the fetus's chromosome is trisomic, then this invention will label the karyotype of chromosome 21 in the sample as a disomic-trisomic karyotype. Similarly, for a maternal plasma sample, if one of the mother's chromosome 22 contains a 22q11 microdeletion while the other chromosome 22 does not, and one of the fetus's chromosome 22 contains a 22q11 microdeletion while the other chromosome 22 does not, then this invention will label the karyotype of the 22q11 chromosome segment in the sample as a monosomic-monosomal karyotype. For example, in a maternal plasma sample, if one chromosome 22 of the mother contains a 22q11 microduplication while the other chromosome 22 does not, and one chromosome 22 of the fetus contains a 22q11 microduplication while the other chromosome 22 does not, then this invention will label the karyotype of the 22q11 chromosome segment in the sample as a trisomy-trisomy karyotype. For example, in a maternal plasma sample, if the alleles for the 6th amino acid of the β subunit of hemoglobin in the mother are A and S, and the alleles for the 6th amino acid of the β subunit of hemoglobin in the fetus are S and C, this invention will label the genotype of the 6th amino acid of the β subunit of hemoglobin in the sample as AS|SC, where the part before the vertical line represents the mother's genotype and the part after the vertical line represents the fetal genotype. In this invention, wild type refers to the genotype with the highest frequency observed at the target site in a normal, disease-free population. On the other hand, wild type refers to a genotype that does not contain pathogenic or potentially pathogenic variants at the target site. In this invention, mutant type is used to refer to genotypes that are different from wild type at the target site.
[0022] In this invention, the concentration of the minimum component DNA in a portion of the sample to be tested is estimated using allele counts at each target locus in a reference group. The concentration of the minimum component DNA in the sample to be tested can be estimated using any currently reported method. Preferably, the concentration of the minimum component DNA in the sample to be tested is estimated using a relative proportion method of allele counts at each target locus in the reference group; preferably, the concentration of the minimum component DNA in the sample is estimated using an iterative fitting genotype method of allele counts at each target locus in the reference group; preferably, the concentration of the minimum component DNA in the sample is calculated using the mean and / or median of FC and TC.
[0023] In this invention, the concentration of the least significant DNA component in a sample is calculated using the relative proportion method of allele counts. For example, in a pregnant woman's plasma DNA sample, the least significant DNA component is fetal DNA, while the largest significant DNA component is maternal DNA. In a normal pregnant woman's plasma DNA sample, the fetus inherits one chromosome from the mother; therefore, the genotype of each target locus can only be one of the following five possible genotypes: AA|AA, AA|AB, AB|AA, AB|AB, or AB|AC, where A, B, and C represent the various alleles at the target locus. Among these five genotypes, if the target locus is the AA|AA or AB|AB genotype, the fetal DNA concentration does not affect the relative count of each allele; however, if the target locus is the AA|AB, AB|AA, or AB|AC genotype, the count of each allele is affected by the fetal DNA concentration. Therefore, the genotypes AA|AB, AB|AA, and AB|AC can be used to estimate the count (FC) of fetal DNA at each target DNA locus.
[0024] This invention provides a method for calculating the concentration of the least significant component DNA in a sample using the relative proportions of allele counts at each target locus in a reference group, the method comprising:
[0025] (a1) Set the noise threshold α for the sample;
[0026] (a2) For each target DNA site, first estimate its genotype using the allele counts of each site, and then estimate the count of DNA from the least component (FC) and the total count (TC) based on the estimated genotype.
[0027] (a3) Estimate the concentration of minimum component DNA using the minimum component DNA count (FC) and total count (TC) at each target site.
[0028] Further, in step (a1) above, setting the noise threshold α for the sample is a threshold used to distinguish between the count signal of real alleles and the count signal of non-real alleles; preferably, the noise threshold α is any value not greater than 0.05; preferably, the noise threshold α is 0.05, 0.04, 0.03, 0.02, 0.01, 0.0075, 0.005, 0.0025 or 0.001.
[0029] Furthermore, in step (a2) above, for each target DNA site, its genotype is first estimated using the allele counts, and then the counts of DNA originating from the least significant component (FC) and the total count (TC) are estimated based on the estimated genotypes, including the following steps:
[0030] (a2-i) Sort the allele counts of the target DNA site from largest to smallest, and label the three largest allele counts as R1, R2 and R3 respectively;
[0031] (a2-ii) Estimate the genotype of the target DNA locus by counting the alleles of each target DNA locus;
[0032] (a2-iii) Estimate the count of DNA originating from the least component (FC) and the total count (TC) based on the estimated genotype of the target DNA site and the count of each allele at the target DNA site.
[0033] Further, in step (a2-ii) above, the genotype of the target DNA locus is estimated by counting the alleles at each allele, with the three largest allele counts labeled R1, R2, and R3, respectively. This includes the following steps:
[0034] (a2-ii-1) Count the number of alleles at each target DNA site to determine the number of alleles detected at the target DNA site that are above the noise threshold; if the result is 1, proceed to step (a2-ii-2); if the result is 2, proceed to step (a2-ii-3); if the result is greater than 2, proceed to step (a2-ii-4).
[0035] (a2-ii-2) Estimate the genotype of the target DNA site to be AA|AA, and then perform the following steps (a2-ii-5);
[0036] (a2-ii-3) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold being 2 and the two largest allele counts at the target DNA site, and then perform the following steps (a2-ii-5).
[0037] (a2-ii-4) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold greater than 2 and the largest two allele counts of the target DNA site, and then perform the following steps (a2-ii-5).
[0038] (a2-ii-5) outputs the estimated genotype of the target site.
[0039] Furthermore, the above step (a2-ii-1) utilizes the counting of each allele at the target DNA site to determine the number of alleles detected at the target DNA site that exceed the noise threshold, and includes the following steps in sequence:
[0040] (a2-ii-1-1) Calculate the relative count of each allele at the target DNA site;
[0041] (a2-ii-1-2) Determine whether the relative count of each allele is higher than the set noise threshold, and then count the number of alleles that are higher than the set noise threshold.
[0042] The relative count of an allele is the quotient of the count of that allele and the count of all alleles at the target site. Preferably, the noise threshold α is any value not greater than 0.05; preferably, the predetermined noise threshold is 0.05, 0.04, 0.02, 0.01, 0.0075, 0.005, 0.0025, or 0.001.
[0043] Further, in the above step (a2-ii-3), based on the number of alleles detected above the noise threshold being 2 and the two largest allele counts at the target DNA site, the genotype of the target DNA site is estimated, where the two largest allele counts are labeled R1 and R2, respectively. This includes the following steps:
[0044] (a2-ii-3-1) Determine whether the value of R1 / (R1+R2) is less than 0.5+α. If the result is yes, estimate the genotype of the target DNA site as AB|AB, and then proceed to the following step (a2-ii-3-3); if the result is no, proceed to the following step (a2-ii-3-2).
[0045] (a2-ii-3-2) Determine if the value of R1 / (R1+R2) is less than 0.75. If the result is yes, estimate the genotype of the target DNA site as AB|AA, and then proceed to the following step (a2-ii-3-3); if the result is no, estimate the genotype of the target DNA site as AA|AB, and then proceed to the following step (a2-ii-3-3).
[0046] (a2-ii-3-3) outputs the estimated genotype of the target site.
[0047] Further, in step (a2-ii-4) above, the genotype of the target DNA site is estimated based on the number of alleles detected above the noise threshold being greater than 2 and the largest two allele counts of the target DNA site, wherein the two largest allele counts are labeled as R1 and R2, respectively. This includes the following steps:
[0048] (a2-ii-4-1) Determine whether R2 / R1 is greater than or equal to 0.5 and / or whether R1 / (R1+R2) is greater than or equal to 1 / 2 and less than or equal to 2 / 3 and / or whether R2 / (R1+R2) is greater than or equal to 1 / 3 and less than or equal to 1 / 2. If the determination result is yes, estimate the genotype of the target DNA site as AB|AC, and then perform the following step (a2-ii-4-3); if the determination result is no, perform the following step (a2-ii-4-2).
[0049] (a2-ii-4-2) Mark the allele count at this site as abnormal, and then either estimate the genotype of the target site as NA and perform the following step (a2-ii-4-3); or set the number of alleles detected at the target DNA site that are above the noise threshold to 2, and then estimate the genotype of the target site as described in step (a2-ii-3) and perform the following step (a2-ii-4-3); (a2-ii-4-3) Output the estimated genotype of the target site.
[0050] Among them, genotype NA represents a genotype for which the target site cannot be estimated.
[0051] Further, in steps (a2-iii) above, based on the estimated genotype of the target DNA site and the allele counts of each target DNA site, the counts of DNA originating from the least significant component (FC) and the total count (TC) are estimated, with the three largest allele counts labeled R1, R2, and R3, respectively. This includes the following steps:
[0052] (a2-iii-1) If the estimated genotype of the target site is AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0053] (a2-iii-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, the estimated total count (TC) is R1+R2 or R1+R2+R3, and then perform the following steps (a2-iii-7).
[0054] (a2-iii-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0055] (a2-iii-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0056] (a2-iii-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3, and then perform the following steps (a2-iii-7).
[0057] (a2-iii-6) If the genotype estimated at the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3, and then the following steps are performed (a2-iii-7).
[0058] (a2-iii-7) Outputs estimated counts of DNA derived from the least component (FC) and total count (TC).
[0059] In this context, the estimated count of DNA derived from the least component (FC) is NA, which represents the count of DNA derived from the least component (FC) that cannot be estimated.
[0060] Furthermore, in step (a3) above, the concentration of minimum component DNA is estimated using the minimum component DNA count (FC) and total count (TC) of each target site in the reference group, wherein the concentration of minimum component DNA in the sample is calculated using linear regression or robust linear regression, and / or the concentration of minimum component DNA in the sample is calculated using the mean or median of FC and TC.
[0061] Furthermore, in step (a3) above, the concentration of minimum component DNA is estimated using the minimum component DNA count (FC) and total count (TC) of each target site in the reference group, wherein the concentration of minimum component DNA is estimated by fitting a regression model.
[0062] Furthermore, in the above steps, the concentration of the least significant component DNA is estimated by fitting a regression model, wherein the regression model is selected from: linear regression model, robust linear regression model, simple regression model, ordinary least squares regression model, multiple regression model, general multiple regression model, polynomial regression model, general linear model, generalized linear model, discrete choice regression model, logistic regression model, polynomial logarithmic model, mixed logarithmic model, probabilistic unit model, polynomial probabilistic unit model, ordered logarithmic model, ordered probabilistic unit model, Poisson model, multivariate response regression model, multilevel model, fixed effects model, random effects model, mixed model, nonlinear regression model, nonparametric model, semiparametric model, robust model, quantile model, isotonic model, principal component model, least angle model, local model, piecewise model, and variable error model.
[0063] Furthermore, in the above steps, the concentration of the minimum component DNA is estimated by fitting a regression model, wherein in the fitted model, the total count (TC) of each target site in the reference group is the independent variable, and the count (FC) of the minimum component DNA at each target site is the dependent variable.
[0064] Furthermore, in the above steps, the concentration of the least component DNA is estimated by fitting a regression model, wherein the concentration of the least component DNA is estimated as the regression coefficient of the total count (TC) of the model parameters.
[0065] Preferably, the fitted regression model is a linear regression model; preferably, the fitted regression model is a robust linear regression model; preferably, the fitted regression model is a general linear model.
[0066] This invention provides a method for calculating the concentration of the least significant component DNA in a sample using an iterative fitting genotype method with allele counting at each target locus in a reference group, the method comprising:
[0067] (b1) Set the noise threshold α, the initial concentration estimate f0, and the iteration error accuracy ε for the sample;
[0068] (b2) For each target DNA site, estimate its genotype using its allele counts and the concentration value f0 of the least component DNA in the sample;
[0069] (b3) For each target DNA site, estimate the count of DNA derived from the least component (FC) and the total count (TC) based on its estimated genotype.
[0070] (b4) Estimate the concentration f of the minimum component DNA using the count (FC) and total count (TC) of the minimum component DNA;
[0071] (b5) Determine whether the absolute value of f-f0 is less than ε. If the result is no, set f0 = f and then execute step (b2). If the result is yes, estimate the minimum concentration of DNA component in the sample as f.
[0072] Further, in step (b1) above, setting the noise threshold α for the sample is a threshold used to distinguish between the count signal of real alleles and the count signal of non-real alleles; preferably, the noise threshold α is any value not greater than 0.05; preferably, the noise threshold α is 0.05, 0.04, 0.03, 0.02, 0.01, 0.0075, 0.005, 0.0025 or 0.001.
[0073] Further, in step (b1) above, the initial concentration estimate f0 is set to be any possible minimum component DNA concentration; preferably, the initial concentration estimate f0 is less than 0.5; preferably, the initial concentration estimate f0 is less than 0.5 and greater than the set noise threshold α; preferably, the initial concentration estimate f0 is any value that is not only less than 0.5 but also greater than the set noise threshold α; preferably, the initial concentration estimate f0 is 0.50, 0.45, 0.40, 0.35, 0.30, 0.25, 0.20, 0.15, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01 or 0.005.
[0074] Further, in step (b1) above, setting the iteration error precision value ε is to set ε as a very small cutoff threshold for iteration calculation; preferably, the set ε value is less than 0.01; preferably, the set ε value is any value less than 0.01; preferably, the set ε value is less than 0.001; preferably, the set ε value is less than 0.0001; preferably, the set ε value is 0.01, 0.001, 0.0001 or 0.00001.
[0075] Furthermore, in step (b2) above, the genotype of each target DNA site is estimated using the allele counts of its individual sites and the concentration value f0 of the least significant component DNA in the sample, including the following steps:
[0076] (b2-i) List all possible genotypes of the target DNA site based on the sample source;
[0077] (b2-ii) For each possible genotype of the target DNA site, the theoretical count of each allele is calculated using the concentration value f0 of the least component DNA in the sample and the total count (TC) of each allele at the target DNA site.
[0078] (b2-iii) For each possible genotype of the target DNA site, perform a goodness-of-fit test using the allele counts of each target DNA site and the theoretical allele counts of each allele.
[0079] (b2-iv) Analyze the goodness-of-fit test results of the target DNA site to all possible genotypes, and select the genotype that has the best fit to the allele count of each target DNA site as the estimated genotype of the target DNA site.
[0080] In this invention, goodness-of-fit test refers to one or more statistical test methods that can be used to test the consistency between observed numbers and theoretical numbers; preferably, the goodness-of-fit test is the chi-square test; preferably, the goodness-of-fit test is the G test; preferably, the goodness-of-fit test is the Fisher exact test; preferably, the goodness-of-fit test is the binomial distribution test; preferably, the goodness-of-fit test is the chi-square test and / or the G test and / or the Fisher exact test and / or the binomial distribution test and / or its variants and / or combinations thereof; preferably, the goodness-of-fit test is performed using the calculated G value and / or the AIC value and / or the corrected G value and / or the corrected AIC value and / or the G value or a variant of the AIC value and / or combinations thereof.
[0081] Further, in step (b3) above, for each target DNA site, the count (FC) and total count (TC) derived from the least component DNA are estimated based on its estimated genotype, with the four largest allele counts labeled R1, R2, R3, and R4, respectively, including the following steps:
[0082] (b3-1) If the genotype of the target site is estimated to be AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0083] (b3-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0084] (b3-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0085] (b3-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0086] (b3-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0087] (b3-6) If the estimated genotype of the target site is AA|BB, then the estimated count (FC) of the least component DNA is R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0088] (b3-7) If the genotype of the target site is estimated to be AA|BC, then estimate the count (FC) of the least component DNA as R2+R3 or twice R2 or twice R3, estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0089] (b3-8) If the estimated genotype of the target site is AB|CC, determine whether the current estimated value f0 is greater than or equal to 1 / 3. If the result is yes, estimate the count (FC) of DNA from the least component as R1 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11); if the result is no, estimate the count (FC) of DNA from the least component as R3 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11).
[0090] (b3-9) If the estimated genotype of the target site is AB|CD, then estimate the count (FC) of the least component DNA as R3+R4 or twice R3 or twice R4, estimate the total count (TC) as R1+R2+R3+R4, and then perform the following steps (b3-11).
[0091] (b3-10) If the genotype estimated for the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, and the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then the following steps (b3-11) are performed.
[0092] (b3-11) Output estimated counts of DNA derived from the least component (FC) and total count (TC).
[0093] Furthermore, in step (b4) above, the concentration f of the minimum component DNA is estimated using the count (FC) and total count (TC) of the minimum component DNA, which is done by using the method described in step (a3).
[0094] In this invention, the concentration of the minimum component DNA in a sample is calculated using an iterative fitting genotype method based on allele counting at each target locus in a reference group. This method can be used not only to estimate the concentration of the minimum component DNA in biologically related mixed samples, but also in mixed samples without biological relationships. Furthermore, this method is applicable not only to calculating the concentration of fetal DNA in plasma DNA samples from mothers who are biologically related, but also to calculating the concentration of fetal DNA in plasma DNA samples from mothers who have legally received donated eggs. Furthermore, this method can be used to estimate the concentration of the minimum component DNA in two independent mixed DNA samples. Furthermore, the method described above can be used to estimate the concentration of several components in a mixture of more than two samples. For example, for multiple pregnancies, a fetal DNA concentration value to be iterated can be set for each fetus; for example, for twin pregnancies, the fetal DNA concentration values to be iterated can be set as f1 and f2; for triplets, the fetal DNA concentration values to be iterated can be set as f1, f2, and f3, and so on. To estimate the concentration of multiple sample components, an initial value can be set for the concentration of each sample. Then, the estimated count of the target site in each sample component can be estimated using the allele counts of each target DNA site and all possible genotypes of that site. The concentration of each sample component can then be iteratively calculated using a goodness-of-fit test until the calculated change in the concentration of each sample component is less than the set precision value.
[0095] In this invention, the target samples to be detected include a single target DNA site, an entire chromosome containing one or more target DNA sites, and a subchromosomal segment containing one or more target DNA sites.
[0096] This invention provides a method for determining the karyotype, genotype, or wild-type mutant of a target DNA site in a sample using a goodness-of-fit test of allele counting at the target DNA site. The method includes:
[0097] (c1) Each target DNA site is divided into a reference site or a target site according to its location on the chromosome, wherein each reference site forms a reference group and each target site forms a target group.
[0098] (c2) Calculate the concentration of the least significant component DNA in the sample by counting alleles at each target DNA site in the reference group;
[0099] (c3) Using the allele counts of each target DNA site in the target group and the concentration of the least component DNA in the sample, the goodness-of-fit test method is adopted to estimate the karyotype, genotype or wild-type of the target to be detected in the sample.
[0100] Further, in step (c3) above, the method of estimating the genotype of the target DNA in the sample by using allele counts at each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, and employing a goodness-of-fit test, includes:
[0101] (c3-a1) For each target DNA site in the target group, list all possible genotypes;
[0102] (c3-a2) For each target DNA site in the target group, for each possible genotype, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site.
[0103] (c3-a3) For each target DNA site in the target group, for each possible genotype, the goodness of fit is tested using the allele counts and theoretical counts of each target DNA site.
[0104] (c3-a4) For each target DNA site in the target group, the best-fit genotype is selected as the genotype of the target DNA site based on the goodness-of-fit test results of all possible genotypes.
[0105] Further, in step (c3) above, the method of estimating the karyotype of the target in the sample by using allele counts at each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, and employing a goodness-of-fit test, includes:
[0106] (c3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested;
[0107] (c3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group;
[0108] (c3-b3) For each target DNA site in the target group, firstly, the goodness of fit of all possible genotypes is tested by counting each allele. Then, for each possible karyotype, a genotype that has the best fit for that karyotype is selected.
[0109] (c3-b4) The goodness-of-fit test results of all target DNA sites for each karyotype are comprehensively analyzed, and the karyotype that best fits all target DNA sites is selected as the karyotype of the target chromosome or subchromosome segment to be detected.
[0110] Further, in step (c3) above, the method of estimating the wild-type mutant of the target to be detected in the sample by using the allele count of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, and employing a goodness-of-fit test, includes:
[0111] (c3-c1a) For each target DNA site in the target group, list all possible wild-type mutant genotypes;
[0112] (c3-c2a) For each target DNA site in the target group, for each possible wild-type mutant, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site.
[0113] (c3-c3a) For each target DNA site in the target group, for each possible wild-type mutant, the goodness of fit is tested by using the allele counts of each target DNA site and its theoretical count.
[0114] (c3-c4a) Comprehensive analysis of all target DNA sites in the target group, and selection of the wild-type mutant genotype that has the best fit to all target sites as the wild-type mutant genotype of the target to be tested.
[0115] Further, in step (c3) above, the method of estimating the wild-type mutant of the target to be detected in the sample by using the allele count of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, and employing a goodness-of-fit test, includes:
[0116] (c3-c1b) For each target DNA site in the target group, the genotype is estimated by using a goodness-of-fit test based on the allele counts and the concentration of the least significant component DNA in the sample.
[0117] (c3-c2b) Based on the genotype of each target DNA site in the target group and the sequence of each allele, determine the wild-type mutant of each allele of the target to be tested in each component of the sample.
[0118] Further, in step (c3) above, the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample are used to estimate the genotype or wild-type mutant of the target to be detected in the sample using a goodness-of-fit test. The target group can be a single target locus or multiple independent repeats of a single target locus. Preferably, the independent repeats of the target locus are obtained using the same primers and independent PCR and / or multiplex PCR amplification reactions; preferably, the independent repeats of the target locus are obtained using different primers and independent PCR and / or multiplex PCR amplification reactions.
[0119] Further, in step (c3) above, the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample are used to estimate the genotype or wild-type mutant of the target to be detected in the sample using a goodness-of-fit test method. The goodness-of-fit test method employs one or more statistical tests that can be used to test the consistency between observed and theoretical numbers. Preferably, the goodness-of-fit test is a chi-square test; preferably, the goodness-of-fit test is a G-test; preferably, the goodness-of-fit test is a Fisher exact test; preferably, the goodness-of-fit test is a binomial distribution test; preferably, the goodness-of-fit test is a chi-square test and / or a G-test and / or a Fisher exact test and / or a binomial distribution test and / or its variants and / or combinations thereof; preferably, the goodness-of-fit test uses the calculated G-value and / or the AIC value and / or the corrected G-value and / or the corrected AIC value and / or the variants and / or combinations thereof of the G-value or AIC value to perform the goodness-of-fit test.
[0120] Furthermore, in step (c3) above, the allele counts of each target DNA site in the target group and the concentration of the least component DNA in the sample are used to estimate the genotype or wild-type mutant of the target to be detected in the sample by adopting a goodness-of-fit test method. The goodness-of-fit test method is performed by the method described in steps (b2-i) to (b2-iv).
[0121] In this invention, chromosome-level karyotype refers to the euploidy or aneuploidy status of a certain chromosome in the various components of the mixture. For example, in a maternal plasma sample, the karyotype of chromosomes in a normal mother and a monosomic fetus is called disoso-monosomic, the karyotype of chromosomes in a normal mother and a trisomic fetus is called disoso-trisomic, and the karyotype of chromosomes in both the mother and fetus being normal is called disoso-disomic.
[0122] In this invention, each subchromosome-level segment is considered a chromosome. Therefore, in maternal plasma samples, the subchromosome karyotype with homozygous microdeletion in both mother and fetus is vacuolation-vacuolation; the subchromosome karyotype with homozygous microdeletion in mother and heterozygous microdeletion in fetus is vacuolation-monosome; the subchromosome karyotype with heterozygous microdeletion in mother and normal fetus is monosomy-disomy; the subchromosome karyotype with heterozygous microdeletion in both mother and fetus is monosomy-monosome; the subchromosome karyotype with homozygous microdeletion in mother and homozygous microdeletion in fetus is monosomy-vacuolation; and the subchromosome karyotype with normal mother and heterozygous microdeletion in fetus is monosomy-vacuolation. The subchromosomal karyotypes are as follows: Disoso-Monososomal, normal in both mother and fetus; Disoso-Disosomal, homozygous microduplication in both mother and fetus; Tetrasoso-Tetrasomal, homozygous microduplication in mother and heterozygous microduplication in fetus; Tetrasoso-Trisosomal, heterozygous microduplication in mother and normal in fetus; Trisososo-Disosomal, heterozygous microduplication in both mother and fetus; Trisososo-Trisosomal, heterozygous microduplication in mother and homozygous microduplication in fetus; and Disososo-Trisosomal, normal in mother and heterozygous microduplication in fetus.
[0123] In this invention, genotype refers to the combination of genotypes of a target DNA locus in the various components of the mixture sample, wherein 0 or 1 alleles may be detected at this locus on each chromosome. For example, in a pregnant woman's plasma sample, a locus with a karyotype of dimer-monomer has four possible genotypes (excluding genotypes where the mother and / or fetus are chimeric), namely... and The possible genotypes for disomy and trisomy (excluding genotypes where the mother and / or fetus are chimeras and / or genotypes where the fetus has not inherited at least one allele from the mother due to de novo mutations, etc.) are AA|AAA, AA|AAB, AB|AAA, AB|AAB, AB|AAC, AB|ABC, AA|ABB, AA|ABC, AB|ACC, and AB|ACD, where A, B, C, and D represent alleles at different target DNA sites. This represents a deletion. In general, the genotype of a locus in a pooled sample represents all possible combinations of all alleles at that locus on each chromosome in each sample. Similarly, for subchromosomal variation, each chromosome may detect 0 (microdeletion), 1 (normal), or 2 (microduplication) alleles at that locus. Therefore, all possible genotypes corresponding to the subchromosomal karyotype of a pooled sample represent all possible combinations of all alleles at each locus on each chromosome in the pooled sample. For example, in a pregnant woman's plasma sample, there are 22 possible genotypes for the subchromosomal trisomy-trisomy locus (excluding genotypes where the mother and / or fetus are mosaics and / or genotypes where the fetus has not inherited at least one allele from the mother due to de novo mutations, etc.). These are AAA|AAA, AAA|AAB, AAA|ABB, AAA|ABC, AAB|AAA, AAB|AAB, AAB|AAC, AAB|ABB, AAB|ABC, AAB|ACC, AAB|ACD, AAB|BBB, AAB|BBC, AAB|BCC, AAB|BCD, ABC|AAA, ABC|AAB, ABC|AAD, ABC|ABC, ABC|ABD, ABC|ADD, and ABC|ADE, where A, B, C, D, and E represent different alleles at the target DNA site.
[0124] This invention provides a method for determining the karyotype, genotype, or wild-type of a target in a sample using a relative distribution map of allele counts at various target loci, the method comprising:
[0125] (d1) Each target DNA site is divided into a reference site or a target site according to its location on the chromosome. The reference sites form a reference group, and the target sites form a target group.
[0126] (d2) Calculate the concentration of the least significant component DNA in the sample by counting alleles at each target DNA site in the reference group;
[0127] (d3) Using the allele counts of each target DNA site in the target group and the concentration of the least component DNA in the sample, the relative distribution map method of allele counts is adopted to estimate the karyotype, genotype or wild-type of the target to be detected in the sample.
[0128] Furthermore, in step (d3) above, the method of estimating the genotype of the target DNA in the sample by using the allele count of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele count, includes:
[0129] (d3-a1) For each target DNA site in the target group, list all possible genotypes;
[0130] (d3-a2) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum theoretical relative count of the allele to plot against the maximum theoretical relative count of the allele to mark the theoretical position of the genotype.
[0131] (d3-a3) For each target DNA site in the target group, first calculate the relative count of each allele, and then select at least one non-maximum allele relative count to plot against the maximum allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0132] (d3-a4) Infer the genotype of the target group based on the theoretical and actual positional distribution of each target DNA locus in the relative allele count map.
[0133] Furthermore, in step (d3) above, the method of estimating the karyotype of the target to be detected in the sample by using the allele count of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele count, includes:
[0134] (d3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested;
[0135] (d3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group;
[0136] (d3-b3) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum allele relative count theoretical value to plot against the maximum allele relative count theoretical value to mark the theoretical position of the genotype.
[0137] (d3-b4) For each target DNA site in the target group, first calculate the relative count of each allele, and then select at least one non-maximum allele relative count to plot against the maximum allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0138] (d3-b5) The karyotype of the target to be tested is inferred based on the theoretical and actual positional distribution of each target DNA site in each karyotype in the relative allele count diagram.
[0139] Furthermore, in step (d3) above, the method of estimating the wild-type mutant of the target to be detected in the sample by using the allele count of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele count, includes:
[0140] (d3-c1) For each target DNA site in the target group, list its wild-type sequence and all possible wild-type mutant genotypes;
[0141] (d3-c2) For each possible wild-type mutant genotype, calculate the theoretical relative counts of its wild-type alleles and each of the other non-wild-type alleles, and select at least one non-wild-type allele relative count to plot the theoretical relative count of the wild-type alleles to mark the theoretical position of its wild-type mutant genotype.
[0142] (d3-c3) For each target DNA site in the target group, calculate the relative count values of its wild-type alleles and other non-wild-type alleles, and select at least one non-wild-type allele relative count to plot the wild-type allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0143] (d3-c4) The wild-type mutants are inferred based on the theoretical and actual positional distributions of all target DNA sites in the relative allele count map of the target group.
[0144] This invention provides a method for determining the karyotype, genotype, or wild-type mutant of a target DNA in a sample using a goodness-of-fit test of allele counts at target DNA loci and / or a relative distribution map of allele counts. The method is characterized in that, in step (c2) or step (d2), the concentration of the least significant DNA component in the sample is calculated using allele counts at each target DNA locus in a reference group, employing the method described in steps (a1)-(a3) and / or steps (b1)-(b5).
[0145] This invention provides a method for determining the karyotype of a target gene in a single genome sample using a relative distribution map of allele counts, the method comprising:
[0146] (e1) Calculate the relative count of each allele at each target DNA site;
[0147] (e2) For each target DNA site, plot the distribution map A of the relative allele count of its second largest allele to the relative allele count of its largest allele, or plot the distribution map B of the relative position of the target DNA site on the chromosome or subchromosome of its largest allele count.
[0148] (e3) Using the relative distribution map A and / or distribution map B of allele counts at each target DNA site, estimate the karyotype of the target to be detected in a single genome sample.
[0149] This invention can not only detect genetic alterations in individual components of a mixed genome—for example, by counting alleles at polymorphic sites in a pregnant woman's plasma DNA sample to detect genetic alterations at a single site in the mother and / or fetus, or variations at the chromosomal and subchromosomal levels—but also can be applied to karyotype or genotype detection of single-genome samples, such as for preimplantation genetic diagnosis of embryonic genetic diseases. This method can simultaneously detect genetic alterations at both the nucleotide and chromosomal or subchromosomal levels, showing promising development and application prospects for fetal genetic disease screening.
[0150] This invention relates to the detection of genetic abnormalities in a target using a mixture of maternal and fetal genetic material. Therefore, in one aspect, the invention provides a method for determining the presence or absence of fetal aneuploidy in a biological sample comprising fetal and maternal nucleic acids in the form of free-floating DNA from a biological sample of the mother, amplifying target DNA sites in a PCR or multiplex PCR reaction (i.e., amplifying template DNA such that the amplified DNA reproduces the original template DNA to a ratio), and then determining the presence or absence of fetal aneuploidy based on the relative count distribution of each allele at each target DNA site of the amplified target.
[0151] In another aspect, the present invention provides a method for determining the presence or absence of copy number variations of fetal chromosomal segments in a biological sample comprising fetal and maternal nucleic acids in the form of free-floating DNA from a biological sample of the mother, amplifying target DNA sites in a PCR or multiplex PCR reaction (i.e., amplifying template DNA such that the amplified DNA reproduces the original template DNA to a ratio), and then determining the presence or absence of copy number variations of the fetal chromosomal segments based on the relative count distribution of each allele at each target DNA site of the amplified target.
[0152] On the other hand, the present invention provides a method for determining the presence or absence of a variation at a fetal single-gene genetic disease pathogenic gene locus in a biological sample, the biological sample comprising fetal and maternal nucleic acids in the form of free-floating DNA from a biological sample of the mother, amplifying the target DNA locus in a PCR or multiplex PCR reaction (i.e., amplifying the template DNA such that the amplified DNA reproduces the original template DNA to a ratio), and then determining the presence or absence of a variation at the fetal single-gene genetic disease pathogenic gene locus based on the relative count distribution of each allele at the amplified target DNA locus (the single-gene genetic disease pathogenic gene locus).
[0153] In another aspect, the present invention provides a diagnostic kit for carrying out the method of the present invention, comprising at least one set of primers for amplifying target DNA sites. The at least one set of primers amplifies at least one reference group target DNA site and / or at least one target group target DNA site. The target group target DNA site is selected from chromosomes with possible chromosomal aneuploidy and / or chromosomal segments with possible copy number variations and / or sites that may be pathogenic variants of single-gene genetic diseases. The nucleic acid sequence of the target group target DNA site is generally polymorphic in the population being tested and / or the target group target DNA site is a pathogenic variant site of a possible single-gene genetic disease. The reference group target DNA site is selected from chromosomes that generally do not have chromosomal aneuploidy and / or chromosomal segments that generally do not have copy number variations. The nucleic acid sequence of the reference group target DNA site is generally polymorphic in the population being tested.
[0154] In another aspect, the present invention provides a diagnostic kit for carrying out the methods of the present invention. This diagnostic kit includes primers for performing steps (2) and / or (3). Optional reagents that may be included in the diagnostic kit are instructions for use, polymerases and buffers for performing PCR and / or multiplex PCR reactions, and reagents required for constructing high-throughput sequencing libraries from the amplified fragments.
[0155] In another aspect, the present invention provides a system for carrying out the method of the present invention. This system is used to carry out one or more steps in a method for predicting the karyotype or genotype or wild-type mutant of a target to be detected from a biological test sample, such as one or more of steps (4) to (5). In another aspect, the present invention provides an apparatus and / or computer program product and / or system and / or module for carrying out the method of the present invention, the apparatus and / or computer program product and / or system and / or module comprising any of the steps (1)-step (5), (a1)-step (a3), (b1)-step (b5), (c1)-step (c3), (d1)-step (d3), and / or (e1)-step (e3) described above.
[0156] In some embodiments, the method of the present invention is performed in vitro or ex vivo. In some embodiments, the sample of the present invention is an in vitro or ex vivo sample.
[0157] In one aspect, the present invention relates to apparatus for performing the method of the invention. For example, in some embodiments, the present invention relates to an apparatus for detecting genetic variations in a sample, characterized by comprising:
[0158] (1) Configure a module for receiving biological samples to be tested and preparing nucleic acids;
[0159] (2) Configure a module for enriching or amplifying target DNA sites, wherein at least one target DNA site has more than one allele in the sample;
[0160] (3) Configure the module for sequencing the amplified target DNA sites;
[0161] (4) Statistical module, which is configured to count the number of alleles for each target DNA site;
[0162] (5) Determine the module, which is configured to determine the karyotype or genotype or wild-type mutant of the target in the sample by using the goodness-of-fit test of allele counts at the target DNA site and / or the relative distribution map of allele counts.
[0163] In some embodiments, a statistics module is configured to count the number of alleles at each target DNA site, the statistics comprising the following steps: (4-1) mapping each amplified sequence to a chromosomal or genomic location; (4-2) counting the number of mapped sequences in each chromosomal or genomic region; wherein if a chromosomal or genomic region has different alleles, the number of sequences mapped to each allele in that region is also counted. In some embodiments, any computer method is used to map the sequence readings to chromosomal or genomic locations / regions. In some embodiments, the computer algorithms used for mapping sequences in step (4-1) include, but are not limited to, specific sequence lookup, BLAST, BLITZ, FASTA, BOWTIE, BOWTIE 2, BWA, NOVOALIGN, GEM, ZOOM, ELAN, MAQ, MATCH, SOAP, STAR, SEGEMEHL, MOSAIK, or SEQMAP or variants thereof or combinations thereof. In some embodiments, specific sequences (unique mapping sequences) are extracted from the chromosomal or genomic sequences corresponding to each target DNA site, and then the readings are mapped to chromosomal or genomic locations / regions using the specific sequences. In some embodiments, sequence reads may be aligned to sequences at chromosomal or genomic locations / regions. In some embodiments, sequence reads may be aligned to sequences at chromosomal or genomic locations. In some embodiments, sequence reads may be obtained from and / or aligned to sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (Japan DNA Database). BLAST or similar tools may be used to search for identical sequences against the sequence database. Then, for example, search hits may be used to sort identical sequences into appropriate chromosomal or genomic locations / regions. In some embodiments, reads may be uniquely or non-uniquely mapped to portions of a reference genome. If a read aligns to a single sequence in the genome, it is called a “unique mapping.” If a read aligns to two or more sequences in the genome, it is called a “non-unique mapping.” In some embodiments, non-uniquely mapped reads are removed from further analysis (e.g., quantification).
[0164] In some implementations, the determination module is configured to determine the karyotype, genotype, or wild-type mutant of the target in the sample using a goodness-of-fit test of allele counts at the target DNA site. This determination includes the following steps:
[0165] (c1) Each target DNA site is classified into a reference site or a target site based on its location on the chromosome;
[0166] (c2) Calculate the concentration of the least significant component DNA in the sample by counting alleles at each target DNA site in the reference group;
[0167] (c3) Using the allele counts of each target DNA site in the target group and the concentration of the least component DNA in the sample, the goodness-of-fit test method is adopted to estimate the karyotype, genotype or wild-type of the target to be detected in the sample.
[0168] In some implementations, the determination module is configured to determine the karyotype, genotype, or wild-type of the target to be detected in a sample using a relative distribution map of allele counts at the target DNA site. This determination includes the following steps:
[0169] (d1) Each target DNA site is classified into a reference site or a target site based on its location on the chromosome;
[0170] (d2) Calculate the concentration of the least significant component DNA in the sample by counting alleles at each target DNA site in the reference group;
[0171] (d3) Using the allele counts of each target DNA site in the target group and the concentration of the least component DNA in the sample, the relative distribution map method of allele counts is adopted to estimate the karyotype, genotype or wild-type of the target to be detected in the sample.
[0172] In some implementations, one or more goodness-of-fit statistical tests are used to test the agreement between observed and expected numbers. In some implementations, the goodness-of-fit test is a chi-square test. In some implementations, the goodness-of-fit test is a G-test. In some implementations, the goodness-of-fit test is a Fisher exact test. In some implementations, the goodness-of-fit test is a binomial distribution test. In some implementations, the goodness-of-fit test is a chi-square test, a G-test, a Fisher exact test, a binomial distribution test, a variant thereof, or a combination thereof. In some implementations, the goodness-of-fit test is performed using the calculated G-value, the AIC value, a corrected G-value, a corrected AIC value, a G-value or a variant of the AIC value, or a combination thereof.
[0173] In some implementations, the determination module is configured to determine the karyotype of the target in a sample using a relative distribution map of allele counts at the target DNA site, wherein the sample to be tested is a single-genome sample, and the determination includes the following steps in sequence:
[0174] (e1) Calculate the relative count of each allele at each target DNA site in the target group;
[0175] (e2) For each target DNA site in the target group, plot the distribution map A of the second largest relative allele count against the largest relative allele count, or plot the distribution map B of the relative position of the target DNA site on the chromosome or subchromosome of the largest relative allele count.
[0176] (e3) Using the relative distribution map A and / or distribution map B of allele counts at each target DNA site in the target group, estimate the karyotype of the target to be detected in a single genome sample.
[0177] In some implementations, the concentration of the least significant component DNA in the sample is calculated in step (c2) or step (d2) using the allele counting relative proportion method, and the calculation includes the following steps in sequence:
[0178] (a1) Set the noise threshold α for the sample;
[0179] (a2) For each target DNA site, first estimate its genotype using the allele counts of each target site, and then estimate the count of DNA from the least component (FC) and the total count (TC) based on the estimated genotype.
[0180] (a3) The concentration of minimum component DNA was estimated using the minimum component DNA count (FC) and total count (TC) at each target site in the reference group.
[0181] In some implementations, in step (c2) or step (d2), an allele counting iterative fitting genotype method is used to calculate the concentration of the least significant component DNA in the sample. This calculation includes the following steps in sequence:
[0182] (b1) Set the noise threshold α, the initial concentration estimate f0, and the iteration error accuracy ε for the sample;
[0183] (b2) Estimate the genotype of each target DNA site using the allele counts of each site and the concentration value f0 of the least component DNA in the sample;
[0184] (b3) For each target DNA site, estimate the count of DNA derived from the least component (FC) and the total count (TC) based on its estimated genotype.
[0185] (b4) Estimate the concentration f of the minimum component DNA using the count (FC) and total count (TC) of the minimum component DNA;
[0186] (b5) Determine whether the absolute value of f-f0 is less than ε. If the result is no, set f0 = f and then execute step (b2). If the result is yes, estimate the minimum concentration of DNA component in the sample as f.
[0187] In some implementations, in step (c3), the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample are used to estimate the genotype of the target to be detected in the sample using a goodness-of-fit test. The estimation includes the following steps in sequence:
[0188] (c3-a1) For each target DNA site in the target group, list all possible genotypes;
[0189] (c3-a2) For each target DNA site in the target group, for each possible genotype, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site.
[0190] (c3-a3) For each target DNA site in the target group, for each possible genotype, the goodness of fit is tested using the allele counts and theoretical counts of each target DNA site.
[0191] (c3-a4) For each target DNA site in the target group, the best-fit genotype is selected as the genotype of the target DNA site based on the goodness-of-fit test results of all possible genotypes.
[0192] In some implementations, in step (c3), the karyotype of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing a goodness-of-fit test method. This estimation includes the following steps:
[0193] (c3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested;
[0194] (c3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group;
[0195] (c3-b3) For each target DNA site in the target group, firstly, the goodness of fit of all possible genotypes is tested by counting each allele. Then, for each possible karyotype, a genotype that has the best fit for that karyotype is selected.
[0196] (c3-b4) The goodness-of-fit test results of all target DNA sites for each karyotype are comprehensively analyzed, and the karyotype that best fits all target DNA sites is selected as the karyotype of the target chromosome or subchromosome segment to be detected.
[0197] In some implementations, in step (c3), the wild-type mutant of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing a goodness-of-fit test. This estimation includes the following steps:
[0198] (c3-c1a) For each target DNA site in the target group, list all possible wild-type mutant genotypes;
[0199] (c3-c2a) For each target DNA site in the target group, for each possible wild-type mutant, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site.
[0200] (c3-c3a) For each target DNA site in the target group, for each possible wild-type mutant, the goodness of fit is tested by using the allele counts of each target DNA site and its theoretical count.
[0201] (c3-c4a) Comprehensive analysis of all target DNA sites in the target group, and selection of the wild-type mutant genotype that has the best fit to all target sites as the wild-type mutant genotype of the target to be tested.
[0202] In some implementations, in step (c3), the wild-type mutant of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing a goodness-of-fit test. This estimation includes the following steps:
[0203] (c3-c1b) For each target DNA site in the target group, the genotype is estimated by using a goodness-of-fit test based on the allele counts and the concentration of the least significant component DNA in the sample.
[0204] (c3-c2b) Based on the genotype of each target DNA site in the target group and the sequence of each allele, determine the wild-type mutant of each allele of the target to be tested in each component of the sample.
[0205] In some implementations, a goodness-of-fit test is performed using one or more statistical tests that can be used to examine the consistency between observed and theoretical numbers. In some implementations, the goodness-of-fit test is a chi-square test. In some implementations, the goodness-of-fit test is a G-test. In some implementations, the goodness-of-fit test is a binomial distribution test. In some implementations, the goodness-of-fit test is a chi-square test and / or a G-test and / or a Fisher exact test and / or a binomial distribution test. In some implementations, the goodness-of-fit test is performed using the calculated G-value and / or the AIC value and / or the corrected G-value and / or the corrected AIC value and / or values derived from the G-value or AIC value.
[0206] In some implementations, in step (d3), the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample are used to estimate the genotype of the target to be detected in the sample using the relative distribution map method of allele counts. The estimation includes the following steps in sequence:
[0207] (d3-a1) For each target DNA site in the target group, list all possible genotypes;
[0208] (d3-a2) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum theoretical relative count of the allele to plot against the maximum theoretical relative count of the allele to mark the theoretical position of the genotype.
[0209] (d3-a3) For each target DNA site in the target group, first calculate the relative count of each allele, and then select at least one non-maximum allele relative count to plot against the maximum allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0210] (d3-a4) Infer the genotype of the target group based on the theoretical and actual positional distribution of each target DNA locus in the relative allele count map.
[0211] In some implementations, in step (d3), the karyotype of the target DNA in the sample is estimated using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing a relative allele count distribution map method. This estimation includes the following steps:
[0212] (d3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested;
[0213] (d3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group;
[0214] (d3-b3) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum allele relative count theoretical value to plot against the maximum allele relative count theoretical value to mark the theoretical position of the genotype.
[0215] (d3-b4) For each target DNA site in the target group, first calculate the relative count of each allele, and then select at least one non-maximum allele relative count to plot against the maximum allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0216] (d3-b5) The karyotype of the target to be tested is inferred based on the theoretical and actual positional distribution of each target DNA site in each karyotype in the relative allele count diagram.
[0217] In some implementations, in step (d3), the wild-type mutant of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the allele count relative distribution map method. The estimation includes the following steps in sequence:
[0218] (d3-c1) For each target DNA site in the target group, list its wild-type sequence and all possible wild-type mutant genotypes;
[0219] (d3-c2) For each possible wild-type mutant genotype, calculate the theoretical relative counts of its wild-type alleles and each of the other non-wild-type alleles, and select at least one non-wild-type allele relative count to plot the theoretical relative count of the wild-type alleles to mark the theoretical position of its wild-type mutant genotype.
[0220] (d3-c3) For each target DNA site in the target group, calculate the relative count values of its wild-type alleles and other non-wild-type alleles, and select at least one non-wild-type allele relative count to plot the wild-type allele relative count to mark the actual position of the target DNA site on the allele relative count map.
[0221] (d3-c4) The wild-type mutants are inferred based on the theoretical and actual positional distributions of all target DNA sites in the relative allele count map of the target group.
[0222] In some implementations, in step (a2), for each target DNA site, the genotype is first estimated using the allele counts of each target site, and then the counts of DNA derived from the least component (FC) and the total count (TC) are estimated based on the estimated genotypes. The estimation includes the following steps in sequence:
[0223] (a2-i) Sort the allele counts of the target DNA site from largest to smallest, and label the three largest allele counts as R1, R2 and R3 respectively;
[0224] (a2-ii) Estimate the genotype of the target DNA locus by counting the alleles of each target DNA locus;
[0225] (a2-iii) Estimate the count of DNA originating from the least component (FC) and the total count (TC) based on the estimated genotype of the target DNA site and the count of each allele at the target DNA site.
[0226] In some implementations, the genotype of the target DNA site is estimated by counting alleles at each target DNA site in step (a2-ii), and the estimation includes the following steps in sequence:
[0227] (a2-ii-1) Count the number of alleles at each target DNA site to determine the number of alleles detected at the target DNA site that are above the noise threshold; if the result is 1, proceed to step (a2-ii-2); if the result is 2, proceed to step (a2-ii-3); if the result is greater than 2, proceed to step (a2-ii-4).
[0228] (a2-ii-2) Estimate the genotype of the target DNA site to be AA|AA, and then perform the following steps (a2-ii-5);
[0229] (a2-ii-3) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold being 2 and the two largest allele counts at the target DNA site, and then perform the following steps (a2-ii-5).
[0230] (a2-ii-4) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold greater than 2 and the largest two allele counts of the target DNA site, and then perform the following steps (a2-ii-5).
[0231] (a2-ii-5) outputs the estimated genotype of the target site.
[0232] In some implementations, the estimation of the genotype of the target DNA site in step (a2-ii-3) based on the number of detected alleles above the noise threshold being 2 and the count of the two largest alleles at the target DNA site, the estimation comprising the following steps in sequence:
[0233] (a2-ii-3-1) Determine whether the value of R1 / (R1+R2) is less than 0.5+α. If the result is yes, estimate the genotype of the target DNA site as AB|AB, and then proceed to the following step (a2-ii-3-3); if the result is no, proceed to the following step (a2-ii-3-2).
[0234] (a2-ii-3-2) Determine if the value of R1 / (R1+R2) is less than 0.75. If the result is yes, estimate the genotype of the target DNA site as AB|AA, and then proceed to the following step (a2-ii-3-3); if the result is no, estimate the genotype of the target DNA site as AA|AB, and then proceed to the following step (a2-ii-3-3).
[0235] (a2-ii-3-3) outputs the estimated genotype of the target site.
[0236] In some implementations, the estimation of the genotype of the target DNA site in step (a2-ii-4) based on the number of detected alleles greater than 2 above a noise threshold and the largest two allele counts of the target DNA site, the estimation comprising the following steps in sequence:
[0237] (a2-ii-4-1) Determine whether R2 / R1 is greater than or equal to 0.5 and / or whether R1 / (R1+R2) is greater than or equal to 1 / 2 and less than or equal to 2 / 3 and / or whether R2 / (R1+R2) is greater than or equal to 1 / 3 and less than or equal to 1 / 2. If the determination result is yes, estimate the genotype of the target DNA site as AB|AC, and then perform the following step (a2-ii-4-3); if the determination result is no, perform the following step (a2-ii-4-2).
[0238] (a2-ii-4-2) Mark the allele count at this site as abnormal, and then either estimate the genotype of the target site as NA and perform the following steps (a2-ii-4-3); or set the number of alleles detected at the target DNA site that are above the noise threshold to 2, and then estimate the genotype of the target site as described in step (a2-ii-3) and perform the following steps (a2-ii-4-3).
[0239] (a2-ii-4-3) outputs the estimated genotype of the target site.
[0240] In some embodiments, the estimation of the count of DNA originating from the least significant component (FC) and the total count (TC) in steps (a2-iii) based on the estimated genotype of the target DNA site and the allele counts of each target DNA site, wherein the three largest allele counts are labeled R1, R2, and R3 respectively, includes the following steps in sequence:
[0241] (a2-iii-1) If the estimated genotype of the target site is AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0242] (a2-iii-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, the estimated total count (TC) is R1+R2 or R1+R2+R3, and then perform the following steps (a2-iii-7).
[0243] (a2-iii-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0244] (a2-iii-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0245] (a2-iii-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3, and then perform the following steps (a2-iii-7).
[0246] (a2-iii-6) If the genotype estimated at the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3, and then the following steps are performed (a2-iii-7).
[0247] (a2-iii-7) Outputs estimated counts of DNA derived from the least component (FC) and total count (TC).
[0248] In some implementations, the genotype estimation for each target DNA site in step (b2) using its individual allele counts and the concentration value f0 of the least significant component DNA in the sample, the estimation comprising the following steps in sequence:
[0249] (b2-i) List all possible genotypes of the target DNA site based on the sample source;
[0250] (b2-ii) For each possible genotype of the target DNA site, the theoretical count of each allele is calculated using the concentration value f0 of the least component DNA in the sample and the total count (TC) of each allele at the target DNA site.
[0251] (b2-iii) For each possible genotype of the target DNA site, perform a goodness-of-fit test using the allele counts of each target DNA site and the theoretical allele counts of each allele.
[0252] (b2-iv) Analyze the goodness-of-fit test results of the target DNA site to all possible genotypes, and select the genotype that has the best fit to the allele count of each target DNA site as the estimated genotype of the target DNA site.
[0253] In some implementations, step (b3) involves estimating the count of DNA derived from the least component (FC) and the total count (TC) for each target DNA site based on its estimated genotype, wherein the four largest allele counts are labeled R1, R2, R3, and R4, respectively. This estimation process includes the following steps:
[0254] (b3-1) If the genotype of the target site is estimated to be AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0255] (b3-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0256] (b3-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0257] (b3-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0258] (b3-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0259] (b3-6) If the estimated genotype of the target site is AA|BB, then the estimated count (FC) of the least component DNA is R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0260] (b3-7) If the genotype of the target site is estimated to be AA|BC, then estimate the count (FC) of the least component DNA as R2+R3 or twice R2 or twice R3, estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0261] (b3-8) If the estimated genotype of the target site is AB|CC, determine whether the current estimated value f0 is greater than or equal to 1 / 3. If the result is yes, estimate the count (FC) of DNA from the least component as R1 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11); if the result is no, estimate the count (FC) of DNA from the least component as R3 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11).
[0262] (b3-9) If the estimated genotype of the target site is AB|CD, then estimate the count (FC) of the least component DNA as R3+R4 or twice R3 or twice R4, estimate the total count (TC) as R1+R2+R3+R4, and then perform the following steps (b3-11).
[0263] (b3-10) If the genotype estimated for the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, and the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then the following steps (b3-11) are performed.
[0264] (b3-11) Output estimated counts of DNA derived from the least component (FC) and total count (TC).
[0265] In some embodiments, the present invention relates to an apparatus for calculating the concentration of a minimum component DNA in a sample, the apparatus comprising:
[0266] (a1) Module used to set the noise threshold α for the sample;
[0267] (a2) is a module used for each target DNA site to first estimate its genotype by counting the alleles of each target site, and then to estimate the count (FC) and total count (TC) of DNA from the least component based on the estimated genotype.
[0268] (a3) The calculation module estimates the concentration of minimum component DNA using the minimum component DNA count (FC) and total count (TC) at each target site in the reference group.
[0269] In some embodiments, the step (a2) of the present invention, which involves first estimating the genotype of each target DNA site using allele counts at that target site, and then estimating the count of DNA originating from the least significant component (FC) and the total count (TC) based on the estimated genotype, includes the following steps:
[0270] (a2-i) Sort the allele counts of the target DNA site from largest to smallest, and label the three largest allele counts as R1, R2 and R3 respectively;
[0271] (a2-ii) Estimate the genotype of the target DNA locus by counting the alleles of each target DNA locus;
[0272] (a2-iii) Estimate the count of DNA originating from the least component (FC) and the total count (TC) based on the estimated genotype of the target DNA site and the count of each allele at the target DNA site.
[0273] In some implementations, the genotype of the target DNA locus is estimated by counting the individual alleles at the target DNA locus, with the three largest allele counts labeled R1, R2, and R3, respectively. This includes the following steps:
[0274] (a2-ii-1) Count the number of alleles at each target DNA site to determine the number of alleles detected at the target DNA site that are above the noise threshold; if the result is 1, proceed to step (a2-ii-2); if the result is 2, proceed to step (a2-ii-3); if the result is greater than 2, proceed to step (a2-ii-4).
[0275] (a2-ii-2) Estimate the genotype of the target DNA site to be AA|AA, and then perform the following steps (a2-ii-5);
[0276] (a2-ii-3) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold being 2 and the two largest allele counts at the target DNA site, and then perform the following steps (a2-ii-5).
[0277] (a2-ii-4) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold greater than 2 and the largest two allele counts of the target DNA site, and then perform the following steps (a2-ii-5).
[0278] (a2-ii-5) outputs the estimated genotype of the target site.
[0279] In some implementations, the genotype of the target DNA site is estimated based on the number of alleles detected above a noise threshold being 2 and the two largest allele counts at the target DNA site, wherein the two largest allele counts are labeled R1 and R2, respectively. This includes the following steps:
[0280] (a2-ii-3-1) Determine whether the value of R1 / (R1+R2) is less than 0.5+α. If the result is yes, estimate the genotype of the target DNA site as AB|AB, and then proceed to the following step (a2-ii-3-3); if the result is no, proceed to the following step (a2-ii-3-2).
[0281] (a2-ii-3-2) Determine if the value of R1 / (R1+R2) is less than 0.75. If the result is yes, estimate the genotype of the target DNA site as AB|AA, and then proceed to the following step (a2-ii-3-3); if the result is no, estimate the genotype of the target DNA site as AA|AB, and then proceed to the following step (a2-ii-3-3).
[0282] (a2-ii-3-3) outputs the estimated genotype of the target site.
[0283] In some implementations, the genotype of the target DNA site is estimated based on the number of alleles detected above a noise threshold greater than 2 and the largest two allele counts of the target DNA site, wherein the two largest allele counts are labeled R1 and R2, respectively. This includes the following steps:
[0284] (a2-ii-4-1) Determine whether R2 / R1 is greater than or equal to 0.5 and / or whether R1 / (R1+R2) is greater than or equal to 1 / 2 and less than or equal to 2 / 3 and / or whether R2 / (R1+R2) is greater than or equal to 1 / 3 and less than or equal to 1 / 2. If the determination result is yes, estimate the genotype of the target DNA site as AB|AC, and then perform the following step (a2-ii-4-3); if the determination result is no, perform the following step (a2-ii-4-2).
[0285] (a2-ii-4-2) Mark the allele count at this site as abnormal, and then either estimate the genotype of the target site as NA and perform the following steps (a2-ii-4-3); or set the number of alleles detected at the target DNA site that are above the noise threshold to 2, and then estimate the genotype of the target site as described in step (a2-ii-3) and perform the following steps (a2-ii-4-3).
[0286] (a2-ii-4-3) outputs the estimated genotype of the target site.
[0287] In some implementations, the count of DNA derived from the least significant component (FC) and the total count (TC) are estimated based on the estimated genotype of the target DNA site and the allele counts of each target DNA site, wherein the three largest allele counts are labeled R1, R2, and R3, respectively. This includes the following steps:
[0288] (a2-iii-1) If the estimated genotype of the target site is AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0289] (a2-iii-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, the estimated total count (TC) is R1+R2 or R1+R2+R3, and then perform the following steps (a2-iii-7).
[0290] (a2-iii-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0291] (a2-iii-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7).
[0292] (a2-iii-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3, and then perform the following steps (a2-iii-7).
[0293] (a2-iii-6) If the genotype estimated at the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3, and then the following steps are performed (a2-iii-7).
[0294] (a2-iii-7) Outputs estimated counts of DNA derived from the least component (FC) and total count (TC).
[0295] In some implementations, the calculation module of step (a3) calculates the concentration of the least component DNA in the sample using linear regression or robust linear regression based on FC and TC counts, or calculates the concentration of the least component DNA in the sample using the mean or median of FC and TC.
[0296] In some embodiments, the present invention relates to an apparatus for calculating the concentration of a minimum component DNA in a sample, the apparatus being:
[0297] (b1) Setting module, which sets the noise threshold α, the initial concentration estimate f0, and the iteration error accuracy ε of the sample;
[0298] (b2) A module for estimating the genotype of each target DNA site using the allele counts of its individual sites and the concentration value f0 of the least component DNA in the sample.
[0299] (b3) Estimation module, which estimates the count (FC) and total count (TC) of DNA originating from the least component for each target DNA site based on its estimated genotype.
[0300] (b4) A module for estimating the concentration f of the least component DNA using the count (FC) and total count (TC) of the least component DNA;
[0301] (b5) The judgment module determines whether the absolute value of f-f0 is less than ε. If the judgment result is no, then f0 = f and then step (b2) is executed; if the judgment result is yes, then the minimum component DNA concentration in the sample is estimated as f.
[0302] In some implementations, estimating the genotype of each target DNA site using its allele count and the concentration value f0 of the least significant component DNA in the sample includes the following steps:
[0303] (b2-i) List all possible genotypes of the target DNA site based on the sample source;
[0304] (b2-ii) For each possible genotype of the target DNA site, the theoretical count of each allele is calculated using the concentration value f0 of the least component DNA in the sample and the total count (TC) of each allele at the target DNA site.
[0305] (b2-iii) For each possible genotype of the target DNA site, perform a goodness-of-fit test using the allele counts of each target DNA site and the theoretical allele counts of each allele.
[0306] (b2-iv) Analyze the goodness-of-fit test results of the target DNA site to all possible genotypes, and select the genotype that has the best fit to the allele count of each target DNA site as the estimated genotype of the target DNA site.
[0307] In some implementations, for each target DNA site, the count of DNA derived from the least component (FC) and the total count (TC) are estimated based on its estimated genotype, with the four largest allele counts labeled R1, R2, R3, and R4, respectively. This includes the following steps:
[0308] (b3-1) If the genotype of the target site is estimated to be AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0309] (b3-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0310] (b3-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0311] (b3-4) If the estimated genotype of the target site is AA|AB, then the estimated count (FC) of the least component DNA is twice that of R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0312] (b3-5) If the estimated genotype of the target site is AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0313] (b3-6) If the estimated genotype of the target site is AA|BB, then the estimated count (FC) of the least component DNA is R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11).
[0314] (b3-7) If the genotype of the target site is estimated to be AA|BC, then estimate the count (FC) of the least component DNA as R2+R3 or twice R2 or twice R3, estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11).
[0315] (b3-8) If the estimated genotype of the target site is AB|CC, determine whether the current estimated value f0 is greater than or equal to 1 / 3. If the result is yes, estimate the count (FC) of DNA from the least component as R1 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11); if the result is no, estimate the count (FC) of DNA from the least component as R3 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11).
[0316] (b3-9) If the estimated genotype of the target site is AB|CD, then estimate the count (FC) of the least component DNA as R3+R4 or twice R3 or twice R4, estimate the total count (TC) as R1+R2+R3+R4, and then perform the following steps (b3-11).
[0317] (b3-10) If the genotype estimated for the target site is not one of the genotypes described above, the count (FC) derived from the least component DNA is estimated to be NA, and the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then the following steps (b3-11) are performed.
[0318] (b3-11) Output estimated counts of DNA derived from the least component (FC) and total count (TC).
[0319] In some embodiments, the sample is a maternal plasma sample, and the minimum component DNA is fetal DNA. In some embodiments, the sample is embryonic nucleic acid derived from preimplantation genetic diagnosis.
[0320] In some embodiments, the present invention provides a diagnostic kit for carrying out the methods of the present invention. This diagnostic kit includes at least one set of primers to amplify a reference group target DNA site and / or a target group target DNA site. The target group target DNA site is selected from chromosomes with possible aneuploidy and / or chromosomal segments with possible copy number variations and / or sites that may be pathogenic variants of single-gene genetic diseases. The nucleic acid sequence of the target group target DNA site is generally polymorphic in the population being tested and / or is a site that may be a pathogenic variant of a single-gene genetic disease. The reference group target DNA site is selected from chromosomes that generally do not have chromosomal aneuploidy and / or chromosomal segments that generally do not have copy number variations. The nucleic acid sequence of the reference group target DNA site is generally polymorphic in the population being tested. In some embodiments, the reference group target DNA site is selected from chromosomal regions in the sample that are considered to be free of chromosomal aneuploidy or chromosomal segment copy number variations. In some embodiments, the reference chromosome or reference chromosome region is selected from chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, and Y, and sometimes, the reference chromosome or reference chromosome region is selected from autosomes (i.e., not X and Y). In some embodiments, the target DNA site is selected from chromosomal regions in the sample that are considered to be likely to contain chromosomal aneuploidy or chromosomal segment copy number variations. In some embodiments, the target DNA site is selected from nucleic acid regions in the sample that are considered to contain and / or may contain pathogenic variants of single-gene hereditary diseases. In some embodiments, the target chromosome region is selected from chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, and Y. Preferably, the target DNA sites are selected from chromosomes 13 and / or 18 and / or 21 and / or X and / or Y. Preferably, the kit includes primers for amplifying target nucleic acids derived from chromosomes 13, 18, 21, X, and / or Y. Preferably, the target DNA sites are selected from chromosomal regions of 1p36 deletion syndrome, Cri-du-chat syndrome, peroneal muscular dystrophy, Digeorge syndrome, Duchenne muscular dystrophy, Williams-Beuren syndrome, Wolf-Hirschhorn syndrome, 15q13.3 microdeletion syndrome, Miller-Dieker syndrome, Smith-Magenis syndrome, Angelman syndrome, and Langer-Giedion syndrome.Preferably, the kit includes primers for amplifying target nucleic acids derived from chromosomal regions of 1p36 deletion syndrome, Cri-du-chat syndrome, peroneal muscular dystrophy, Digeorge syndrome, Duchenne muscular dystrophy, Williams-Beuren syndrome, Wolf-Hirschhorn syndrome, 15q13.3 microdeletion syndrome, Miller-Dieker syndrome, Smith-Magenis syndrome, Angelman syndrome, and Langer-Giedion syndrome. It should be understood that the reference chromosome or a portion thereof containing the target site region is an euploid chromosome. Euploidy refers to a normal number of chromosomes. Optional reagents that may be included in the diagnostic kit are instructions for use, polymerases and buffers for performing PCR and / or multiplex PCR reactions, and reagents required for constructing high-throughput sequencing libraries of the amplified fragments.
[0321] In some embodiments, the present invention provides a diagnostic kit for carrying out the methods of the present invention. This diagnostic kit includes primers for performing steps (2) and / or (3). Optional reagents that may be included in the diagnostic kit are instructions for use, polymerases and buffers for performing PCR and / or multiplex PCR reactions, and reagents required for constructing high-throughput sequencing libraries from the amplified fragments.
[0322] In some embodiments, the present invention provides a system for carrying out the method of the present invention, which is used to perform one or more steps in a method for predicting the karyotype or genotype or wild-type of a target to be detected from a biological test sample, such as one or more of steps (4) to (5). In some embodiments, the present invention provides an apparatus and / or computer program product and / or system and / or module for carrying out the method of the present invention, the apparatus and / or computer program product and / or system and / or module being used to perform any of the steps (1)-step (5), steps (a1)-step (a3), steps (b1)-step (b5), steps (c1)-step (c3), steps (d1)-step (d3) and / or steps (e1)-step (e3) above.
[0323] In one aspect, the present invention relates to the following embodiments:
[0324] 1. A method for detecting genetic variations in a sample, characterized by comprising the following steps in sequence:
[0325] (1) Receive biological samples to be tested and prepare nucleic acids;
[0326] (2) Enrich or amplify target DNA sequences, wherein at least one target DNA sequence has more than one allele in the sample;
[0327] (3) Sequencing the amplified target DNA;
[0328] (4) For each target DNA sequence, count the number of each allele;
[0329] (5) Use the goodness-of-fit test of allele counts and / or the relative distribution map of allele counts to determine the karyotype, genotype or wild mutant of the target locus to be detected in the sample.
[0330] 2. The method as described in Scheme 1, characterized in that in step (5), the karyotype, genotype, or wild-type of the target locus to be detected in the sample is determined by using a goodness-of-fit test of allele counting, wherein the determination includes the following steps in sequence:
[0331] (A1) The target DNA sequence is divided into reference group sequence and target group sequence according to its location on the chromosome;
[0332] (A2) Using the allele count of each target DNA sequence in the reference group, the concentration of the least component DNA in the sample is calculated by the relative proportion method of allele count or the iterative fitting method of allele count to fit the genotype.
[0333] (A3) Use the goodness-of-fit test of allele counts of each target DNA sequence in the target group to estimate the karyotype, genotype or wild-type of the target site to be detected in the sample.
[0334] 3. The method as described in Scheme 1, characterized in that in step (5), the karyotype, genotype, or wild-type of the target locus to be detected in the sample is determined using a relative distribution map of allele counts, wherein the determination includes the following steps in sequence:
[0335] (B1) The target DNA sequence is divided into reference group sequence and target group sequence according to its location on the chromosome;
[0336] (B2) Using the allele count of each target DNA sequence in the reference group, the concentration of the least component DNA in the sample is calculated by the relative proportion method of allele count or the iterative fitting method of allele count to fit the genotype.
[0337] (B3) Using the relative distribution map of allele counts of each target DNA sequence in the target group, estimate the karyotype, genotype or wild-type of the target site to be detected in the sample.
[0338] 4. The method as described in Implementation Scheme 1, characterized in that in step (5), the karyotype, genotype, or wild-type of the target locus to be detected in the sample is determined using a relative distribution map of allele counts, wherein the sample to be detected is a single genome sample, and the determination includes the following steps in sequence:
[0339] (C1) Calculate the relative count of each allele for each target DNA sequence in the target group;
[0340] (C2) For each target DNA sequence, plot a distribution map A of the second largest relative allele count against the largest relative allele count, or plot a distribution map B of the relative position of the largest relative allele count against the target DNA sequence on a chromosome or subchromosome.
[0341] (C3) Using the relative distribution map A and / or distribution map B of allele counts of each target DNA sequence in the target group, estimate the karyotype of the target region to be detected in a single genome sample.
[0342] 5. The method as described in embodiment 2 or 3, characterized in that the concentration of the least significant component DNA in the sample is calculated in step (A2) or step (B2) using the allele counting relative proportion method, the calculation comprising the following steps in sequence:
[0343] (a1) Set the noise background value α for the sample;
[0344] (a2) Estimate the count (FC) and total count (TC) of DNA originating from the least component for each target DNA sequence;
[0345] (a3) Estimate the concentration of minimal component DNA using the count of minimal component DNA (FC) and the total count (TC).
[0346] 6. The method as described in embodiment 2 or 3, characterized in that in step (A2) or step (B2), the concentration of the least significant component DNA in the sample is calculated using an allele counting iterative fitting genotype method, the calculation comprising the following steps in sequence:
[0347] (b1) Set the noise background value α, the initial concentration estimate f0, and the iteration error accuracy value ε for the sample;
[0348] (b2) Estimate the genotype of each target DNA sequence using its allele count and f0;
[0349] (b3) For each target DNA sequence, estimate the count (FC) and total count (TC) of DNA derived from the least component based on its estimated genotype.
[0350] (b4) Estimate the concentration f of the minimum component DNA using the count (FC) and total count (TC) of the minimum component DNA;
[0351] (b5) Determine whether the absolute value of f-f0 is less than ε. If the result is no, set f0 = f and then execute step (b2). If the result is yes, estimate the minimum concentration of DNA component in the sample as f.
[0352] 7. The method as described in Scheme 2, characterized in that in step (A3), the genotype of the target locus to be detected in the sample is estimated by using a goodness-of-fit test of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0353] (A3.a1) Analyze the target DNA sequence sites and list all possible genotypes;
[0354] (A3.a2) For each possible genotype, calculate the theoretical count of each allele based on the estimated minimum component DNA concentration and the total count of the target DNA sequence, and then perform a goodness-of-fit test on the allele counts of the target DNA sequence and their theoretical counts.
[0355] (A3.a3) Based on the goodness-of-fit test results of all possible genotypes of the target DNA sequence, the genotype with the best fit is selected as the genotype of the target DNA sequence site.
[0356] 8. The method as described in Scheme 2, characterized in that in step (A3), the karyotype of the target site to be detected in the sample is estimated by using a goodness-of-fit test of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0357] (A3.b1) Analyze the sample to be tested and list all possible karyotypes of the sample on the target chromosome or subchromosome segment;
[0358] (A3.b2) For each possible karyotype, list all possible genotypes of the target DNA sequence on the chromosome or subchromosome of that karyotype in the sample;
[0359] (A3.b3) For each target DNA sequence, first use the allele count to perform a goodness-of-fit test on all possible genotypes, and then select the genotype that best fits the karyotype for each karyotype.
[0360] (A3.b4) Analyze the goodness-of-fit test results of all target DNA sequences to all karyotypes, and select the karyotype that best fits all target DNA sequences as the karyotype of the target chromosome or subchromosome fragment to be detected.
[0361] 9. The method as described in Scheme 2, characterized in that in step (A3), the wild-type mutant of the target site to be detected in the sample is estimated by using a goodness-of-fit test of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0362] (A3.c1) Estimate the genotype of the target DNA sequence by counting alleles of the target DNA and goodness-of-fit test;
[0363] (A3.c2) Based on the genotype of the target DNA sequence and the wild-type mutant sequences of its various alleles, determine the wild-type mutants of each allele of the target DNA sequence in each component of the sample.
[0364] 10. The method as described in embodiment 3, characterized in that in step (B3), the genotype of the target locus to be detected in the sample is estimated using the relative distribution map of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0365] (B3.a1) Analyze the sample to be tested and list all possible genotypes of the target sequence site;
[0366] (B3.a2) Calculate the theoretical relative count of each allele in each possible genotype, and plot the theoretical relative count of at least one non-maximum allele relative count against the maximum allele relative count for each genotype to mark the theoretical positions of all possible genotypes.
[0367] (B3.a3) Calculate the relative counts of each allele in the target DNA sequence, and plot at least one non-maximum allele relative count against the maximum allele relative count to mark the actual position of the target DNA sequence's allele relative count; (B3.a4) Infer the genotype of the target DNA sequence based on the theoretical and actual position distributions of the target DNA sequence in the allele relative count graph.
[0368] 11. The method as described in embodiment 3, characterized in that in step (B3), the karyotype of the target site to be detected in the sample is estimated using the relative distribution map of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0369] (B3.b1) Analyze the sample to be tested and list all possible karyotypes of the sample on the target chromosome or subchromosome segment;
[0370] (B3.b2) For each possible karyotype, list all possible genotypes of the target DNA sequence of the target group on the chromosome or subchromosome of that karyotype in the sample. Then, for each genotype, select at least one non-maximum theoretical allele relative count value and plot it against the maximum theoretical allele relative count value to mark the theoretical position of the genotype.
[0371] (B3.b3) For each target DNA sequence in the target group, calculate the relative count of each allele and select at least one non-maximum relative allele count to plot against the maximum relative allele count to mark the actual location of the site.
[0372] (B3.b4) The karyotype of the target chromosome or subchromosome segment to be detected is inferred based on the theoretical and actual positional distribution of all target DNA sequences in the relative allele count graph.
[0373] 12. The method as described in embodiment 3, characterized in that in step (B3), the wild-type mutant of the target site to be detected in the sample is estimated using the relative distribution map of allele counts of each target DNA sequence in the target group, wherein the estimation includes the following steps in sequence:
[0374] (B3.c1) Analyze the sample to be tested and list the wild-type sequence of the target sequence site and all its possible genotypes;
[0375] (B3.c2) Calculate the theoretical relative counts of wild-type alleles and other non-wild-type alleles in each possible genotype, and plot the theoretical relative counts of at least one non-wild-type allele against the theoretical relative counts of wild-type alleles for each genotype to mark the theoretical positions of all possible genotypes.
[0376] (B3.c3) Calculate the relative counts of wild-type alleles and other non-wild-type alleles of the target DNA sequence, and plot the relative counts of at least one non-wild-type allele against the relative counts of wild-type alleles to mark the actual position of the relative counts of alleles in the target DNA sequence.
[0377] (B3.c4) Infer wild-type mutants based on the theoretical and actual positional distribution of target DNA sequences in the allele relative counting graph.
[0378] 13. The method as described in embodiment 5, characterized in that the estimation performed in step (a2) sequentially includes the following steps:
[0379] (i) Sort the allele counts of the target DNA sequence from largest to smallest, and label the three largest allele counts as R1, R2 and R3 respectively;
[0380] (ii) Determine the number of alleles detected in the target DNA sequence that are above the noise threshold; if the determination result is 1, estimate the genotype of the target DNA sequence as AA|AA, and then perform the following step (vi); if the determination result is 2, perform the following step (iii); if the determination result is greater than 2, perform the following step (v).
[0381] (iii) Determine whether the value of R1 / (R1+R2) is less than 0.5+α. If the result is yes, estimate the genotype of the target DNA sequence as AB|AB, and then proceed to step (vi); if the result is no, proceed to step (iv).
[0382] (iv) Determine whether the value of R1 / (R1+R2) is less than 0.75. If the result is yes, estimate the genotype of the target DNA sequence as AB|AA, and then proceed to step (vi) below. If the result is no, estimate the genotype of the target DNA sequence as AA|AB, and then proceed to step (vi) below.
[0383] (v) Determine whether the value of R2 / R1 is less than 0.5. If the result is no, estimate the genotype of the target DNA sequence as AB|AC, and then perform the following step (vi); if the result is yes, mark the target DNA sequence as an outlier, and then either estimate the genotype of the target DNA sequence as NA, and then perform the following step (vi), or perform the above step (iii).
[0384] (vi) Estimate the count (FC) and total count (TC) of DNA derived from the least component based on the estimated target DNA sequence genotype.
[0385] 14. The method as described in embodiment 6, characterized in that the estimation performed in step (b2) sequentially includes the following steps:
[0386] (i) List all possible genotypes of the target DNA sequence based on the sample source;
[0387] (ii) For all possible genotypes of the target DNA sequence, calculate the theoretical count of each allele using f0 and the total count (TC) of each allele in the target DNA sequence.
[0388] (iii) Use the allele counts of each target DNA sequence and the theoretical counts of each allele to perform a goodness-of-fit test.
[0389] (iv) Analyze the goodness-of-fit test results of the target DNA sequence to all possible genotypes, and select the genotype that has the best fit to the allele count of each target DNA sequence as the estimated genotype of the target DNA sequence. Attached Figure Description
[0390] Figure 1 This is a flowchart illustrating the process of estimating fetal DNA concentration by counting alleles at multiple polymorphic sites in a pregnant woman's plasma cfDNA sample.
[0391] Figure 2 This is a schematic diagram of a process for estimating the concentration of the least significant component DNA by counting alleles at multiple polymorphic sites in a mixed sample of two components.
[0392] Figure 3This method estimates fetal DNA concentration by sequencing polymorphic sites in maternal plasma cfDNA samples. First, the fetal DNA count (FC) and total maternal and fetal DNA count (TC) are estimated using the allele counts of each polymorphic site. Then, robust RLM regression fitting is performed on the FC and TC counts of all polymorphic sites, and the fetal DNA concentration is estimated as the slope (model coefficient) of the fitted line.
[0393] Figure 4 It estimates the minimum component DNA concentration by sequencing polymorphic sites in a mixed component DNA sample. The minimum component DNA count (FC) and the total component DNA count (TC) of each polymorphic site are estimated using the allele counts of each polymorphic site. Figure 4 In step a, robust RLM regression through the origin is performed using the FC and TC values of each polymorphic site, while the minimum component DNA concentration is estimated as the slope of the line (model coefficient). Figure 4 b represents the results of RLM robust regression estimation of the minimum component DNA concentration for multiple different samples or different biological replicates. Four pooled samples were replicated multiple times at the library preparation or sequencing level, with expected minimum component DNA concentrations of 0.01, 0.02, 0.10, or 0.20 (x-axis), while the estimated minimum component DNA concentration for each sample is represented by the y-axis. The dashed line in the figure indicates the position of the line y = x.
[0394] Figure 5 It uses the counting of alleles at polymorphic sites to detect monosomy of fetal chromosomes. Figure 5 a) uses the comprehensive goodness-of-fit test results to detect whether the dimer-diso karyotype chromosome in the simulated maternal plasma cfDNA sample is a fetal monosomic abnormality. Figure 5 b) utilizes the comprehensive goodness-of-fit test results to detect whether the dimer-monomer karyotype chromosome in simulated maternal plasma cfDNA samples represents a fetal monosomy abnormality. The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value of the G test at that locus by the fetal concentration and then by the total count of all alleles at that locus.
[0395] Figure 6 It uses the counting of alleles at polymorphic sites to detect trisomy variations in fetal chromosomes. Figure 6 a) uses the comprehensive goodness-of-fit test results to detect whether the dimer-diso karyotype chromosome in the simulated maternal plasma cfDNA sample is a fetal trisomy abnormality. Figure 6 b uses the comprehensive goodness-of-fit test results to detect whether the disomy-trisomy karyotype chromosome in the simulated maternal plasma cfDNA sample is a fetal trisomy abnormality.
[0396] Figure 7It estimates microdeletion variations at the subchromosome level in the fetus being tested by counting alleles at polymorphic sites. Figure 7 a) uses the results of a comprehensive goodness-of-fit test to detect whether monosomy-disomy karyotype chromosomes in simulated maternal plasma cfDNA samples are microdeletion abnormalities of fetal chromosomes. Figure 7 b is Figure 7 A local magnification of a. Figure 7 c uses the comprehensive goodness-of-fit test results to detect whether the monomer-monomer karyotype chromosome in the simulated maternal plasma cfDNA sample is a fetal chromosomal microdeletion abnormality. Figure 7 d is Figure 7 A local magnification of c.
[0397] Figure 8 It estimates the microrepetitive variations at the subchromosomal level of the fetus under test by counting alleles at polymorphic sites. Figure 8 a) uses the results of the comprehensive goodness-of-fit test to detect whether trisomy-disomy karyotype chromosomes in simulated maternal plasma cfDNA samples are microduplications of fetal chromosomes. Figure 8 b is Figure 8 A local magnification of a. Figure 8 c uses the comprehensive goodness-of-fit test results to detect whether trisomy-trisomy karyotype chromosomes in simulated maternal plasma cfDNA samples are microduplications of fetal chromosomes. Figure 8 d is Figure 8 A local magnification of c.
[0398] Figure 9 It utilizes the counting of alleles at polymorphic sites to detect wild-type mutants in fetuses at the short sequence level. Figure 9 a) uses the goodness-of-fit test results to detect the genotype of a short sequence at a simulated site where the mother has a heterozygous mutation but the fetus is normal. Figure 9 b is Figure 9 A local magnification of 'a' shows that the estimated genotype of this genetic locus is AB|AA, meaning the mother is heterozygous and the fetus is homozygous. Further analysis of the allele sequences revealed that allele A is wild-type while allele B is mutant. Therefore, the wild-type at this locus is determined to be Aa|AA, indicating a heterozygous mutant mother and a normal fetus. Figure 9 c is used to detect the genotype of the short sequence at the site of the simulated heterozygous mutation in both the mother and fetus, based on the goodness-of-fit test results. Figure 9 d is Figure 9A magnified view of c. The results indicate that the estimated genotype of this genetic locus is AB|AC, meaning both the mother and fetus are heterozygous. Further analysis of the allele sequences revealed that allele A is wild-type while alleles B and C are mutant. Therefore, the wild-type mutant at this locus is determined to be a heterozygous mutation in both the mother and fetus (Aa|Ab), and the fetus either developed a de novo mutation or inherited a paternal allele mutation.
[0399] Figure 10 The diagram shows how to estimate the genotype of a target locus using a relative distribution map of allele counts. Figure 10 'a' represents the theoretical distribution of the relative allele counts of a polymorphic site on a normal dimo-diso karyotype chromosome. Figure 10 b is the distribution of the second largest relative allele count relative to the largest relative allele count at a polymorphic site on a normal dimo-diso karyotype chromosome.
[0400] Figure 11 The figure shows the theoretical distribution of the relative allele counts at each polymorphic site on the normal chromosome of the mother in the maternal plasma cfDNA sample. Figure 11 'a' represents the theoretical relative count of all possible genotypes and their alleles at each polymorphic locus on a hemi-hemi-karyotype or hemi-monotype chromosome. Figure 11 b is the theoretical distribution of the second largest relative allele count relative to the largest relative allele count at each polymorphic site on the dimer-disomorphic and dimer-monomorphic chromosomes. Figure 11 c represents the theoretical relative count of all possible genotypes and all alleles at each polymorphic locus on the chromosome of the dimo-diso or dimo-triso karyotype. Figure 11 d is the theoretical distribution of the relative count of the second or fourth largest allele relative to the relative count of the largest allele at each polymorphic site on the chromosomes of the dimo-diso and dimo-triso karyotypes.
[0401] Figure 12 The figure shows the theoretical distribution of the relative counts of alleles at each polymorphic site at the subchromosome level in the target group of the pregnant woman's plasma cfDNA sample. Figure 12 'a' represents the theoretical relative count of all possible genotypes and their alleles at each polymorphic site on the chromosome with or without microdeletion karyotypes in the mother or fetus. Figure 12 b is the theoretical distribution of the second largest relative allele count relative to the largest relative allele count at each polymorphic site on the chromosome with or without microdeletion karyotypes in the mother or fetus. Figure 12 c represents the theoretical relative count of all possible genotypes and all alleles at each polymorphic locus on the subchromosome where the mother has or does not have microduplications and the fetus is normal. Figure 12 d is the theoretical distribution of the relative counts of the second or third largest alleles relative to the relative counts of the largest alleles at each polymorphic site on the subchromosome with a normal fetal karyotype, whether the mother has microduplications or not.
[0402] Figure 13 The figure shows the theoretical distribution of all possible genotypes and the relative counts of each allele at the test site on a normal dimo-diso karyotype chromosome in a pregnant woman's plasma cfDNA sample. Figure 13 'a' is the theoretical value of the total number of all possible genotypes and their relative counts of all alleles at the test site on a normal dimo-diso karyotype chromosome. Figure 13 b is a theoretical distribution diagram of the maximum relative count of non-wild-type alleles relative to the relative count of wild-type alleles for each possible genotype of the test site on a normal dimo-diso karyotype chromosome.
[0403] Figure 14 It uses a relative distribution map of allele counts at polymorphic sites to detect monosomy variations in fetal chromosomes. Figure 14 a) uses the relative distribution map of allele counts to estimate the karyotype of normal dimo-diso chromosomes in simulated maternal plasma cfDNA samples. Figure 14 b is the estimation of the karyotype of dimelo-monomer chromosomes in simulated maternal plasma cfDNA samples using a relative distribution map of allele counts.
[0404] Figure 15 It uses a relative distribution map of allele counts at polymorphic sites to detect trisomy variations in fetal chromosomes. Figure 15 a) uses the relative distribution map of allele counts to estimate the karyotype of normal dimo-diso chromosomes in simulated maternal plasma cfDNA samples. Figure 15 b is the estimation of the karyotype of diploid-trisomy chromosomes in simulated maternal plasma cfDNA samples using a relative distribution map of allele counts.
[0405] Figure 16 It utilizes the relative distribution map of allele counts at polymorphic sites to detect microdeletion variations at the fetal subchromosomal level. Figure 16 a) uses the relative distribution map of allele counts to estimate the microdeletion karyotype of monosomy-disomy subchromosomes in simulated maternal plasma cfDNA samples. Figure 16 b is the estimation of the microdeletion karyotype of monosomy-monosomy subchromosomes in simulated maternal plasma cfDNA samples using the relative distribution map of allele counts.
[0406] Figure 17 It uses the relative distribution map of allele counts at polymorphic sites to detect microrepetitive variations at the fetal subchromosomal level. Figure 17a) uses the relative distribution map of allele counts to estimate the microduplication karyotype of trisomy-disomy subchromosomes in simulated maternal plasma cfDNA samples. Figure 17 b is the estimation of the microduplication karyotype of trisomy-trisomy subchromosomes in simulated maternal plasma cfDNA samples using the relative distribution map of allele counts.
[0407] Figure 18 It uses the relative distribution map of allele counts at polymorphic sites to detect wild-type mutants in fetuses at the short sequence level. Figure 18 'a' represents the wild-type mutant of the ab|Aa genotype locus in a simulated pregnant woman's plasma cfDNA sample, estimated using a relative distribution map of allele counts. Figure 18 b represents the wild-type mutant of the Aa|ab genotype locus in a simulated pregnant woman's plasma cfDNA sample, estimated using a relative distribution map of allele counts.
[0408] Figure 19 The diagram illustrates the use of relative allele counts at polymorphic loci to detect karyotypes of target chromosomes or subchromosomal segments in a single-genome sample. For each polymorphic locus in the target group, the relative allele count of its second-largest allele is plotted against its largest allele count (relative count map), or the relative allele count of its largest allele is plotted against the locus's relative position on a simulated chromosome (relative count position map). The karyotype of the target locus can be estimated based on the distribution characteristics of each polymorphic locus on the relative count map or relative count position map. Detailed Implementation
[0409] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of protection claimed by the present invention. Any modifications or substitutions made to the methods, steps, or conditions of the present invention without departing from the spirit and substance of the invention are within the scope of the present invention.
[0410] Example 1: Analysis and calculation of allele counts at each polymorphic site in a pregnant woman's plasma DNA sample.
[0411] In this embodiment, the sequencing result file (Barrett, Xiong et al. 2017, PLoS One 12: e0186771) was obtained from the NIH SRA database (BioProject ID: PRJNA387652).
[0412] 1. Sample Collection: In this embodiment, Barrett et al. collected 10-20 ml of peripheral blood from each pregnant woman and then extracted plasma DNA (cfDNA) using the QiaAmp Circulating Nucleic Acid Kit (Qiagen) according to the manufacturer's protocol. In this embodiment, we analyzed 157 plasma cfDNA samples collected using the above method.
[0413] 2. Polymorphic Site Amplification and Sequencing: Barrett et al. selected 44 polymorphic sites with high minimum allele frequencies (MAF>0.25) in the population and designed 45 pairs of amplification primers, including 44 pairs of sequence-specific polymorphic site amplification primers and one pair of ZFX / ZFY site amplification primers. Each sample was amplified using the 45 primer pairs and multiplex PCR. The amplification products were used to prepare sequencing libraries using TruSeq Nano DNA Sample Preparation kits (Illumina) according to the manufacturer's instructions. Sequencing was then performed using a MiSeq sequencer according to the manufacturer's instructions.
[0414] 3. Prepare the data analysis index:
[0415] (3.1) Prepare reference sequences for polymorphic sites: Using the forward and reverse primers of the 45 amplification sites reported by Barrett et al., as well as the specific location of the polymorphic sites on the chromosome, extract the reference sequence of each amplification product from the human genome sequence database.
[0416] (3.2) Preparation of localization indexes for polymorphic sites: For each amplification product reference sequence, manually divide it into three regions: the 5' region, the variant region, and the 3' region. The variant region is the nucleic acid sequence region in the amplification product reference sequence affected by any allele of the polymorphic site. The 5' region is the nucleic acid sequence from the beginning of the reference sequence to the beginning of the variant region in the 5' to 3' direction, and the 3' region is the nucleic acid sequence from the end of the variant region to the end of the reference sequence in the 5' to 3' direction. Then, for each polymorphic site, select at least one unique sequence from each of the 5' and 3' regions as the localization index for that polymorphic site. The unique sequence is one that is unique among all the amplification product reference sequences of the polymorphic sites, and can be used to uniquely locate the amplification product to the specific polymorphic site.
[0417] (3.3) Prepare allele counting index for polymorphic sites: For each polymorphic site, first download all its allele sequences from the NCBI dbSNP database. Then, for each allele sequence, select a unique nucleic acid sequence as the allele counting index for that polymorphic site. The unique nucleic acid sequence means that the unique nucleic acid sequence is unique among all possible amplification product reference sequences of that polymorphic site. For the same allele at that site, the unique nucleic acid sequence is the same, while for different alleles at that site, the unique nucleic acid sequence is different.
[0418] 4. Sequencing Data Analysis: For each sequencing sequence, low-quality sequences are first filtered out. Then, the location index of each polymorphic site is searched from beginning to end in the filtered sequences. If at least one location index of a polymorphic site is found, the sequence is located at that specific polymorphic site; otherwise, the sequence is discarded. Finally, for each sequence located at a specific polymorphic site, the allele count index of that polymorphic site is searched from beginning to end. If at least one allele count index is found, one of the allele count indices is selected, and the sequence is marked as the allele of that polymorphic site; otherwise, the sequence is discarded.
[0419] 5. Count the number of alleles at each polymorphic site: For each sample, count the number of sequencing sequences of each allele at each polymorphic site, which is the count of each allele at each polymorphic site.
[0420] Example 2: Analysis and calculation of alleles at each polymorphic site in two independent genomic pool samples. count
[0421] In this embodiment, the sequencing result file (Kim, Kim et al. 2019, Nat Commun 10:1047) was obtained from the NIH SRA database (BioProject ID: PRJNA517742).
[0422] 1. Sample Collection: In this embodiment, Kim et al. extracted genomic DNA from two independent blood samples, one as the major component and the other as the minor component. The genomic DNA from the two samples were mixed in a certain proportion to obtain mixed samples with minor component proportions of 0.01, 0.02, 0.10 and 0.20, respectively.
[0423] 2. Polymorphic Site Amplification and Sequencing: Kim et al. selected 645 polymorphic sites in two genomic samples and designed amplification primers. Each pooled sample was amplified using amplification primers and multiplex PCR. Sequencing libraries were prepared from the amplification products according to the manufacturer's instructions, and then sequenced using either an Ion Torrent or Illumina sequencer. In this example, we analyzed the Illumina sequencing dataset (ILA dataset) from the sequencing results.
[0424] 3. Preparation of data analysis index: We used the information on the specific location of the 645 amplification sites on chromosomes reported by Kim et al. to extract the reference sequence of each amplification product from the human genome sequence database. Then, we used the methods described in steps (3.2) and (3.3) of step 3 in Example 1 to prepare the location index of each polymorphic site and the counting index of each allele.
[0425] 4. Sequencing Data Analysis: For each sequencing sequence, low-quality sequences are first filtered out. Then, the location index of each polymorphic site is searched from beginning to end in the filtered sequences. If at least one location index of a polymorphic site is found, the sequence is located at that specific polymorphic site; otherwise, the sequence is discarded. Finally, for each sequence located at a specific polymorphic site, the allele count index of that polymorphic site is searched from beginning to end. If at least one allele count index is found, one of the allele count indices is selected, and the sequence is marked as the allele of that polymorphic site; otherwise, the sequence is discarded.
[0426] 5. Count the number of alleles at each polymorphic site: For each sample, count the number of sequencing sequences of each allele at each polymorphic site, which is the count of each allele at each polymorphic site.
[0427] Example 3: Computer simulation of each polymorphic site in a mixed sample and counting of each allele.
[0428] In this embodiment, we generate the allele sequences of the simulated polymorphic sites according to the following steps.
[0429] 1. Simulating Polymorphic Sites: First, a unique 70bp sequence is randomly generated and divided into three regions: a 5' region (30bp), a variant region (10bp), and a 3' region (30bp). Then, mutations (including insertions, deletions, point mutations, and multiple site variations) are randomly generated in the 10bp variant region to obtain at least six different nucleic acid sequences of length ≥60bp containing the 5', variant, and 3' regions, which are then labeled as different alleles of the polymorphic site. Finally, at least one unique 12bp sequence is selected as the location index of the polymorphic site as described in step (3.2) of Example 1; and at least one unique 12bp sequence containing the variant region is selected as the allele counting index of the polymorphic site as described in step (3.3) of Example 1.
[0430] 2. Simulate specific chromosomes or chromosome segments in the sample: For each chromosome, simulate at least 100 polymorphic sites according to step 1 above, and determine the number of simulated alleles and the count of each allele for each site based on the simulated genotype.
[0431] For example, simulating a polymorphic site on a chromosome in the plasma cfDNA of a pregnant woman with a karyotype of disoproximity-disoproximity, the genotype could be AA|AA, AA|AB, AB|AA, AB|AB, and AB|AC. Assuming the concentration of fetal DNA in the sample is 10%, and the simulated genome copy number is 200, then the fetal genome has 20 copies, while the maternal genome has 180 copies. First, select a polymorphic site, list its various allele sequences, and label them A, B, C, D, E, F, etc. Then, for genotype AA|AA, 200 copies of allele A are simulated; for genotype AA|AB, 180 copies of maternal allele A, 10 copies of fetal allele A, and 10 copies of fetal allele B are simulated, i.e., 190 copies of allele A and 10 copies of allele B are simulated; for genotype AB|AA, 110 copies of allele A and 90 copies of allele B are simulated; for genotype AB|AB, 100 copies of allele A and 100 copies of allele B are simulated; for genotype AB|AC, 100 copies of allele A, 90 copies of allele B, and 10 copies of allele C are simulated.
[0432] For example, a polymorphic site on a certain chromosome in the plasma cfDNA of a pregnant woman with a simulated karyotype of dimo-monotype, or a polymorphic site on the plasma cfDNA of a pregnant woman with a simulated karyotype of dimo-monotype, may have the following genotype: or Assuming the concentration of fetal DNA in the sample is 10%, and the simulated normal genome copy number is 200, then the fetal genome has 20 copies and the mother's genome has 180 copies. First, select a polymorphic site, list its allele sequences, and label them A, B, C, D, E, F, etc. Then, for the genotype... Simulate 190 copies of allele A; for genotype Simulate 100 copies of allele A and 90 copies of allele B; for genotype Simulate 180 copies of allele A and 10 copies of allele B; for genotype Simulate 90 copies of allele A, 90 copies of allele B, and 10 copies of allele C.
[0433] The number of alleles at polymorphic sites on other karyotype chromosomes or chromosome segments, and the genomic copy number of each allele, can be simulated using a similar method.
[0434] 3. Simulate specific samples: Each sample simulates different chromosomes according to the experimental purpose, and each chromosome or chromosome segment simulates at least 100 polymorphic sites according to step 2 above. Each site simulates the genomic copy number of different alleles according to different genotypes, where the total copy number of all alleles of each polymorphic site corresponds to 200 genomic copies under the normal dimo-diso karyotype.
[0435] 4. Simulate high-throughput sequencing results: Using the genome copy sequences of different polymorphic sites simulated for each sample as input files, the high-throughput sequencing results were simulated using ART simulation software (Huang, Li et al. 2012, Bioinformatics 28: 593-594).
[0436] 5. Sequencing Data Analysis: For each sequencing sequence, low-quality sequences are first filtered out. Then, the location index of each polymorphic site is searched from beginning to end in the filtered sequences. If at least one location index of a polymorphic site is found, the sequence is located at that specific polymorphic site; otherwise, the sequence is discarded. Finally, for each sequence located at a specific polymorphic site, the allele count index of that polymorphic site is searched from beginning to end. If at least one allele count index is found, one of the allele count indices is selected, and the sequence is marked as the allele of that polymorphic site; otherwise, the sequence is discarded.
[0437] 6. Count the number of alleles at each polymorphic site: For each sample, count the number of sequencing sequences of each allele at each polymorphic site, which is the count of each allele at each polymorphic site.
[0438] Example 4: Estimating the number of alleles above the noise threshold detected at a polymorphic site.
[0439] Select a polymorphic locus, and arrange the allele counts in descending order, labeling them as R1, R2, R3, ..., Rn or R1, R2, R3, ..., Rn. n The total count of each allele is the sum of the counts of each allele, denoted as TC.
[0440] Assuming the noise threshold for the sample is α, for a polymorphic locus, if the count of any allele is less than TC × α, then that allele count is marked as noise. The number of alleles at that polymorphic locus that are not marked as noise is the number of alleles at that locus that are above the noise threshold. For example, if the four allele counts at a polymorphic locus are 27, 3552, 5809, and 11, then TC = 27 + 3552 + 5809 + 11 = 9399, R1 = 5809, R2 = 3552, R3 = 27, and R4 = 11. If the noise threshold α = 0.01, then the cutoff threshold (Th) = TC × α = 93.99. Since R1 and R2 are both greater than 93.99 while R3 and R4 are both less than 93.99, the alleles at that locus that are above the noise threshold are R1 and R2, and the number of alleles at that locus that are above the noise threshold is 2.
[0441] Preferably, after sorting the allele counts of a polymorphic site in descending order and labeling them as R1, R2, ..., Rn, the number of alleles detected at the polymorphic site that are above the noise threshold is estimated according to the following steps:
[0442] (1) Set the noise threshold for sequencing to α;
[0443] (2) Calculation
[0444] (3) If C i-1 ≥α and C i If <α, then it is estimated that there are i-1 alleles at this polymorphic site.
[0445] For example, for a polymorphic locus, if i = 3, C2 = R2 / (R1+R2) ≥ α and C3 = R3 / (R1+R2+R3) < α, then it is estimated that there are i-1 = 2 detected alleles above the noise threshold at this locus. For example, if the allele counts of the four polymorphic loci are 27, 3552, 5809, and 11, then TC = 27 + 3552 + 5809 + 11 = 9399, R1 = 5809, R2 = 3552, R3 = 27, and R4 = 11. If the noise threshold α = 0.01 is set, then the cutoff threshold α = 0.01. Since C1 = R1 / R1 = 1.0, C2 = R2 / (R1+R2) = 0.38, C3 = R3 / (R1+R2+R3) = 0.003, and C4 = R4 / (R1+R2+R3+R4) = 0.001, and C2 is greater than or equal to 0.01 while C3 is less than 0.01, the alleles above the noise threshold at this locus are R1 and R2, and the total number of alleles above the noise threshold at this locus is 2.
[0446] Example 5: Estimating the total allele count (TC) at a polymorphic locus.
[0447] The total count (TC) of all alleles at a polymorphic site can be calculated using any of the following methods:
[0448] (1) For a polymorphic site, sum the counts of each allele to obtain the total count of each allele in the polymorphic site.
[0449] (2) For a polymorphic site, first calculate the number of alleles detected above the noise threshold according to the method described in Example 4. Then the total count of each allele in the polymorphic site is the sum of the counts of each allele above the noise threshold.
[0450] (3) Based on the sample characteristics, consider the maximum number of alleles (let's say k) that a polymorphic site might have in the sample. Then, the total count of all alleles at the polymorphic site is the sum of the counts of its k largest alleles, i.e.
[0451] Example 6: Estimating the probable genotype of a polymorphic locus in a plasma cfDNA sample using allele counting.
[0452] For pregnant women who are the biological mothers of the fetus (biological pregnant women), the genotype of each polymorphic locus on chromosomes with normal disomy in both mother and fetus in their plasma cfDNA can only be one of five genotypes (not considering cases where the mother and / or fetus have mosaic genotypes and / or the fetus did not inherit the mother's genotype for various reasons). For each polymorphic locus, firstly, the number of alleles detected above the noise threshold is calculated according to the method described in Example 4, and then the possible genotypes of the polymorphic locus can be estimated according to the following steps:
[0453] (1) Set the noise threshold for sequencing to α;
[0454] (2) Determine the number of alleles above the noise threshold. If the result is 1, proceed to step (3); if the result is 2, proceed to step (4); if the result is greater than 2, proceed to step (8).
[0455] (3) Estimate the genotype of the polymorphic site to be AA|AA, and then perform the following steps (11);
[0456] (4) Determine the value of R1 / (R1+R2). If the result is less than 0.5+α, proceed to step (5); if the result is greater than or equal to 0.5+α and less than 0.75, proceed to step (6); if the result is greater than or equal to 0.75, proceed to step (7).
[0457] (5) Estimate the genotype of the polymorphic site as AB|AB, and then perform the following steps (11);
[0458] (6) Estimate the genotype of the polymorphic site to be AB|AA, and then perform the following steps (11);
[0459] (7) Estimate the genotype of the polymorphic site to be AA|AB, and then perform the following steps (11);
[0460] (8) Determine the value of R2 / R1. If the result is less than 0.5, proceed to step (9); if the result is greater than or equal to 0.5, proceed to step (10).
[0461] (9) Mark the polymorphic site as an outlier, and then either estimate the genotype of the polymorphic site as NA and perform the following step (11); or perform the above step (4);
[0462] (10) Estimate the genotype of the polymorphic site to be AB|AC, and then perform the following steps (11);
[0463] (11) Output the estimated genotype of the polymorphic site.
[0464] Example 7: Estimation of the number of polymorphic sites derived from fetal DNA in the plasma cfDNA sample of the biological mother (FC)
[0465] Select a polymorphic locus, first estimate the total count (TC) of DNA derived from the pregnant woman and fetus according to the method described in Example 5, then estimate the possible genotype of the polymorphic locus according to the method described in Example 6, and estimate the count (FC) of fetal DNA derived from the polymorphic locus according to the following steps:
[0466] (1) If the genotype of the polymorphic site is AA|AA, then FC is estimated to be NA;
[0467] (2) If the genotype of the polymorphic site is AA|AB, then the FC estimate is 2.0×R2;
[0468] (3) If the genotype of the polymorphic site is AB|AA, then the FC estimate is R1-R2;
[0469] (4) If the genotype of the polymorphic site is AB|AB, then FC is estimated to be NA;
[0470] (5) If the genotype of the polymorphic site is AB|AC, then the FC estimate is R1-R2+R3 or 2.0×R3;
[0471] (6) If the genotype of the polymorphic site is not any of the above genotypes, then FC is estimated to be NA.
[0472] Example 8: Estimating the concentration (f) of the least significant component DNA in a mixed sample.
[0473] Select multiple polymorphic sites, and then estimate the concentration of the least significant component DNA in the sample according to the following steps:
[0474] (1) Estimate the total count (TC) of DNA from all samples for each polymorphic site according to the method described in Example 5;
[0475] (2) Estimate the number of DNA components (FC) originating from each polymorphic site according to the methods described in Examples 6 and 7;
[0476] (3) Based on the FC and TC counts for each polymorphic site, calculate the concentration of the least component DNA in the sample using linear regression or robust linear regression, or calculate the concentration of the least component DNA in the sample using the mean or median of FC and TC.
[0477] Figure 1 This is a flowchart illustrating the estimation of fetal DNA concentration in plasma cfDNA samples from biological pregnant women, as described in Example 8.
[0478] Example 9: Estimating the individual polymorphic sites based on the sample concentration f of the least significant component in a mixture of two samples. Expected count of alleles
[0479] For biological pregnant women's plasma cfDNA samples, the two samples refer to maternal cfDNA and fetal cfDNA respectively, with the least significant component being fetal cfDNA and the largest component being maternal cfDNA; for a mixture of two independent genomic samples, the least significant component refers to the DNA component of the sample with a smaller proportion and the largest component is the DNA component with a larger proportion; for pregnant women who have legally accepted donated eggs, the least significant component is fetal cfDNA and the largest component is maternal cfDNA.
[0480] Select a polymorphic locus and first estimate the total count (TC) of the polymorphic locus from the two sample DNAs, following the method described in Example 5. If the concentration of the least abundant component is f, then the concentration of the other, most abundant component sample is 1-f. For any polymorphic locus genotype, estimate the theoretical expected count of each allele at that polymorphic locus according to the following steps:
[0481] (1) For each allele from the largest component sample, label its relative value as 1-f;
[0482] (2) For each allele from the minimum component sample, label its relative value as f;
[0483] (3) Calculate the relative total value of each allele and the relative total value of all alleles in the polymorphic site;
[0484] (4) For each allele, calculate the ratio of its relative total value to the relative total value of all alleles, and then multiply the ratio by TC to obtain the theoretical expected count of the allele.
[0485] For example, for a biological pregnant woman's plasma DNA sample, assuming the fetal DNA concentration is f, and for any polymorphic locus, the total count of all alleles is TC. Then, for the polymorphic locus of genotype AA|AA, there are two chromosomal locations originating from the largest component sample (mother's DNA), namely A and A (marked before the vertical line), and two chromosomal locations originating from the smallest component sample (fetal DNA), namely A and A (marked after the vertical line). Therefore, the relative total value of allele A is (1-f) + (1-f) + f + f = 2, and the relative total value of all alleles is (1-f) + (1-f) + f + f = 2; the ratio is 2 / 2 = 1. Therefore, the theoretical expected value of allele A is TC*1 = TC. For genotypes AB|AC, the relative total value of all alleles is (1-f)+(1-f)+f+f=2; the relative total value of allele A is (1-f)+f=1, and the proportion is 1 / 2, so its theoretical expected value is 1 / 2×TC=TC / 2; the relative total value of allele B is 1-f, and the proportion is (1-f) / 2, so its theoretical expected value is (1-f) / 2×TC; the relative total value of allele C is f, and the proportion is f / 2, so its theoretical expected value is f / 2×TC. For the genotype AB|AAB at a polymorphic locus on a disomic-trisomic chromosome, the relative total of all alleles is (1-f) + (1-f) + f + f + f = 2 + f; the relative total of allele A is (1-f) + f + f = 1 + f, with a ratio of (1 + f) / (2 + f), so its theoretical expected value is (1 + f) / (2 + f) × TC; the relative total of allele B is 1-f + f = 1, with a ratio of 1 / (2 + f), so its theoretical expected value is 1 / (2 + f) × TC. The theoretical expected counts for other genotypes can be obtained using a similar method.
[0486] For example, for a pregnant woman's plasma DNA sample who has legally accepted donated eggs, assuming the fetal DNA concentration is f, the total count of all alleles at any polymorphic site is TC. For genotype AB|AC, the relative total value of all alleles is (1-f)+(1-f)+f+f=2; the relative total value of allele A is (1-f)+f=1, with a proportion of 1 / 2, so its theoretical expected value is 1 / 2×TC=TC / 2; the relative total value of allele B is 1-f, with a proportion of (1-f) / 2, so its theoretical expected value is (1-f) / 2×TC; the relative total value of allele C is f, with a proportion of f / 2, so its theoretical expected value is f / 2×TC. For genotype AA|BC, the relative total of all alleles is (1-f) + (1-f) + f + f = 2; the relative total of allele A is (1-f) + (1-f) = 2 - 2f, and the proportion is (2-2f) / 2 = 1-f, so its theoretical expected value is (1-f) × TC; the relative total of allele B is f, and the proportion is f / 2, so its theoretical expected value is f / 2 × TC; the relative total of allele C is f, and the proportion is f / 2, so its theoretical expected value is f / 2 × TC. The theoretical expected counts for other genotypes can be obtained using a similar method.
[0487] Example 10: Goodness-of-fit test for allele counts at polymorphic sites.
[0488] Select a polymorphic locus and perform a goodness-of-fit test on the possible genotypes of that locus according to the following steps:
[0489] (1) Calculate the count of each allele at the polymorphic locus and label them in descending order as observation counts O1, O2, ..., O m ;
[0490] (2) Calculate the expected count of each allele according to the method described in Example 9, and label them as E1, E2, ..., E in descending order. n ;
[0491] (3) Use the observed counts and expected counts of each allele to perform a goodness-of-fit test.
[0492] The goodness-of-fit test in step (3) above can be performed using, but is not limited to, Fisher's exact test, binomial distribution test, chi-square test, or G test.
[0493] For example, for a given genotype, if the observed allele counts are O1, O2, and O3, and the expected counts are E1, E2, and E3, then the goodness of fit of the G-test can be calculated as follows:
[0494]
[0495] or
[0496] Where df represents the degrees of freedom.
[0497] Preferably, if the observed allele count is less than the expected allele count, the missing observed allele count is set to a very small value, such as 0.1; if the expected allele count is less than the observed allele count, the expected value of the missing position is set to a very small value or background noise value, such as 5 or TC×α.
[0498] For example, if the two allele counts at a certain polymorphic locus are observed to be 4105 and 577 respectively, the fetal DNA concentration f = 0.25, and the noise threshold is set to α = 0.01, then O1 = 4105, O2 = 577, and TC = 4105 + 577 = 4682. To determine which genotype the allele counts at this polymorphic locus best fit, a goodness-of-fit test is performed on the observed allele counts against the theoretical counts of all possible genotypes at this polymorphic locus. Examples of the goodness-of-fit test results for this polymorphic locus to genotypes AA|AA, AA|AB, and AB|AC are shown below:
[0499] Genotype AA|AA: Degrees of freedom df = 1, expected allele counts are E1 = TC × (1 - α) = 4682 × (1 - 0.01) = 4635.18, E2 = TC × α = 46.82; then G = 1901.045, AIC = G - 2 × df = 1899.045. Alternatively, degrees of freedom df = 0, expected allele counts are E1 = TC = 4682, E2 = 0 (discarded); then G = 0.0, AIC = G - 2 × df = 0.0.
[0500] Genotype AA|AB: Degrees of freedom df=1, expected allele counts are E1=TC×(2-f) / 2=4682×(2-0.25) / 2=4096.75, E2=TC×f / 2=4682×0.25 / 2=585.25; then G=0.1334, AIC=G-2×df=-1.8666.
[0501] Genotype AB|AC: Degrees of freedom df = 2. Since we expect three alleles but only observe two alleles, O3 is set to a very small value, such as O3 = 0.1. The expected allele counts are E1 = TC × 1 / 2 = 4682 × 1 / 2 = 2341, E2 = TC × (1-f) / 2 = 4682 × (1-0.25) / 2 = 1755.75, E3 = TC × f / 2 = 4682 × 0.25 / 2 = 585.25. Then G = 3325.046, AIC = G - 2 × df = 3321.046.
[0502] Alternatively, the same number of allele counts can be used for all goodness-of-fit tests. Since this polymorphic locus may have a maximum of three alleles, the three largest values are retained for both the observed and expected allele counts. The observed allele counts can be padded with smaller values, while the expected allele counts can be padded with a threshold. For example, to fit the genotype AA|AB to the two observed allele values above, we set O3 = 0.1, E3 = TC × α = 46.82, df = 2, therefore E1 = TC × (1-α) × (2-f) / 2 = 4055.783, E2 = TC × (1-α) × f / 2 = 579.398; G = 94.24, AIC = G - 2 × df = 90.24.
[0503] Example 11: Using the sample concentration f of the least component in a mixture of samples and the allelic concentration of a polymorphic site. Gene counting is used to estimate the possible genotypes of this polymorphic site.
[0504] For a polymorphic locus, estimate the genotype of that polymorphic locus according to the following steps:
[0505] (1) According to the method described in Example 10, the observed allele counts were used to perform a goodness-of-fit test on the theoretical counts of each allele for each possible genotype.
[0506] (2) Select the genotype that has the best fit test for each observed allele count as the genotype of the polymorphic site.
[0507] Example 12: Using the sample concentration f of the least significant component in a mixture of two independent samples and a polymorphic site Allelic counts and genotype estimations at each site represent the minimum number of alleles derived from the least subset of samples at that polymorphic locus (FC).
[0508] In a mixture of two independent samples, if the concentration of the least significant component is f, then the concentration of the largest component is 1-f. Allele counts, arranged from largest to smallest, are labeled R1, R2, R3, and R4. The count of the least significant component (FC) from which the polymorphic site originates is then estimated according to the following steps:
[0509] (1) If the genotype of the polymorphic site is AA|AA, then FC is estimated to be NA;
[0510] (2) If the genotype of the polymorphic site is AA|AB, then the FC estimate is 2.0×R2;
[0511] (3) If the genotype of the polymorphic site is AB|AA, then the FC estimate is R1-R2;
[0512] (4) If the genotype of the polymorphic site is AB|AB, then FC is estimated to be NA;
[0513] (5) If the genotype of the polymorphic site is AB|AC, then the FC estimate is R1-R2+R3 or 2.0×R3;
[0514] (6) If the genotype of the polymorphic site is AA|BB, then the FC estimate is R2;
[0515] (7) If the genotype of the polymorphic site is AA|BC, then the FC is estimated to be R2+R3 or 2.0×R2 or 2.0×R3;
[0516] (8) If the genotype of the polymorphic site is AB|CC, then determine whether f>1 / 3. If the result is yes, then FC is estimated as R1; if the result is no, then FC is estimated as R3.
[0517] (9) If the genotype of the polymorphic site is AB|CD, then the FC estimate is R3+R4 or 2.0×R3 or 2.0×R4;
[0518] (10) If the genotype of the polymorphic site is not any of the above genotypes, then FC is estimated to be NA.
[0519] Example 13: Estimating the minimum component in a mixture of two samples using allele counts at polymorphic sites. Sample concentration
[0520] Select multiple polymorphic sites, and then estimate the sample concentration of the least abundant component in a mixture of two independent samples according to the following steps:
[0521] (1) Set the background noise value α, the concentration accuracy value ε, and the initial concentration value f0;
[0522] (2) Estimate the genotype of each polymorphic site according to the method described in Example 11;
[0523] (3) Estimate the total count (FC) of DNA from the least component sample for each polymorphic site in the mixture according to the method described in Example 7 or Example 12;
[0524] (4) Estimate the total count (TC) of each polymorphic site from the two samples according to the method described in Example 5;
[0525] (5) Estimate the concentration of the least component DNA in the mixed sample based on the FC and TC counts for each polymorphic site, as described in Example 8;
[0526] (6) Determine whether |f-f0| is less than ε. If the result is yes, the concentration of the minimum component in the mixture is f. If the result is no, set f0 = f and then execute step (2).
[0527] For plasma DNA samples from pregnant women who have legally accepted donated eggs, the minimum component is fetal DNA, and the maximum component is maternal DNA. Since the fetus does not inherit genetic material from the chromosomes of the pregnant woman who has legally accepted donated eggs, each polymorphic site in the plasma DNA of this pregnant woman may be one of nine genotypes (not considering cases where the mother and / or fetus have chromosomal aneuploidy or chromosomal copy number variations, and / or the mother and / or fetus are chimeric, and / or the fetus has other non-diploid karyotypes for various reasons). The concentration of fetal DNA can be estimated iteratively using the steps described above.
[0528] For a biological pregnant woman's plasma DNA sample, the smallest component is fetal DNA, and the largest component is maternal DNA. Since the fetus inherits genetic material from the biological mother's chromosomes, each polymorphic site in the biological pregnant woman's plasma DNA may be one of five genotypes (not considering cases where the mother and / or fetus have chromosomal aneuploidy or chromosomal copy number variations and / or the mother and / or fetus are mosaic genotypes and / or the fetus did not inherit the mother's genotype for various reasons). The concentration of fetal DNA can be estimated iteratively using the steps described above.
[0529] Figure 2 This is a flowchart for estimating the concentration of fetal DNA in plasma DNA samples from pregnant women who have legally accepted donated eggs, as described in Example 13.
[0530] Example 14: Estimating fetal DNA concentration using simulated sequencing of polymorphic sites in maternal plasma DNA samples.
[0531] The following section uses the allele counts of five hypothetical polymorphic sites in simulated pregnant woman plasma cfDNA as an example to briefly explain the method and steps for estimating the fetal DNA concentration in a sample using the relative proportion of allele counts.
[0532] (1) Simulate sequencing results of multiple polymorphic sites on the reference genome
[0533] Polymorphic sites on the reference genome were selected and labeled Id001-Id005. The allele counts for each of the five simulated polymorphic sites as described in Example 3 are shown in Table 1. In the hypothetical maternal plasma cfDNA, the reference genome is considered to be a chromosomal region with a normal disomy karyotype in both the mother and fetus; therefore, each polymorphic site theoretically contains a maximum of three alleles. Here, each site shows a maximum of five allele counts (some of which represent system noise from sample processing, sequencing, etc.). It should be understood that each polymorphic site may be detected to contain multiple alleles, and each allele should be counted.
[0534] Table 1: Allele counts for each of the five hypothetical polymorphic loci
[0535]
[0536] (2) Estimate the count of fetal DNA in each polymorphic site count according to the methods described in Examples 6 and 7.
[0537] Since each polymorphic site in pregnant women's plasma cfDNA theoretically has a maximum of three alleles, the allele counts for each polymorphic site were sorted in descending order, and the three largest alleles were labeled as R1, R2, and R3. The results are shown in Table 2.
[0538] Table 2: Allele counts for each of the five hypothetical, sorted polymorphic loci.
[0539] Id001 14127 35 0 Id002 4105 577 13 Id003 3148 3101 54 Id004 5809 3552 27 Id005 4007 3028 1011
[0540] The sequencing noise threshold was set to α = 0.01. The amplification count (FC) theoretically derived from fetal DNA and the total count (TC) derived from maternal and fetal DNA were calculated for each polymorphic site.
[0541] For locus Id001, R2 / (R1+R2)=35 / (14127+35)=0.002<0.01, the number of alleles is estimated to be one, the genotype is estimated to be AA|AA, FC=NA, TC=R1=14127.
[0542] For locus Id002, R2 / (R1+R2)=577 / (4105+577)=0.123≥0.01, R3 / (R1+R2+R3)=13 / (4105+577+13)=0.003<0.01, so the number of alleles is estimated to be two. Since R1 / (R1+R2)=0.877≥0.75, the genotype is estimated to be AA|AB, FC=2×R2=1154, TC=R1+R2=4682.
[0543] For locus Id003, R2 / (R1+R2)=0.496≥0.01, R3 / (R1+R2+R3)=0.009<0.01, the number of alleles is estimated to be two, because R1 / (R1+R2)=0.504<0.5+α, the genotype is estimated to be AB|AB, FC=NA, TC=R1+R2=6249.
[0544] For locus Id004, R2 / (R1+R2)=0.379≥0.01, R3 / (R1+R2+R3)=0.003<0.01, the number of alleles is estimated to be two, because 0.5+α≤R1 / (R1+R2)=0.621<0.75, the genotype is estimated to be AB|AA, FC=R1-R2=2257, TC=R1+R2=9361.
[0545] For locus Id005, R2 / (R1+R2)=0.430≥0.01, R3 / (R1+R2+R3)=0.126≥0.01, the number of alleles is estimated to be two, because R2 / R1=0.756≥0.5, the genotype is estimated to be AB|AC, FC=R1-R2+R3=1990, TC=R1+R2+R3=8046.
[0546] (3) Estimating the concentration of fetal DNA
[0547] The concentration of fetal DNA in the sample was calculated using R software and linear or robust linear regression, or by using the mean or median of FC and TC. The results are shown in Table 3.
[0548] (a) Input the values of FC and TC
[0549] FC=c(NA,1154,NA,2257,1990)
[0550] TC=c(14127,4682,6249,9361,8046)
[0551] (b) Calculation of fetal DNA concentration using linear regression
[0552] lmfit = lm(FC ~ TC + 0)
[0553] f = lmfit$coefficients["TC"]
[0554] (c) Calculation of fetal DNA concentration using robust regression
[0555] library(MASS)
[0556] rlmfit=rlm(FC~TC+0, maxit=1000)
[0557] f = rlmfit$coefficients["TC"]
[0558] (d) Calculate the concentration of fetal DNA in the sample using the mean or median of FC and TC.
[0559] (d1)f=median(FC / TC, na.rm=T)
[0560] (d2)f=median(FC[c(2,4,5)]) / median(TC[c(2,4,5)])
[0561] (d3)f=mean(FC / TC, na.rm=T)
[0562] (d4)f=mean(FC[c(2,4,5)]) / mean(TC[c(2,4,5)])
[0563] (e) The results of the fetal DNA concentration calculation are shown in Table 3.
[0564] Table 3: Estimation of fetal DNA concentration in samples using different methods.
[0565] Linear regression (b) 0.2441 Steady recovery (c) 0.2441 Median of the ratio (d1) 0.2465 The average of the ratios (d3) 0.2450 The ratio of the median (d2) 0.2473 The ratio of the means (d4) 0.2445
[0566] Example 15: Utilizing simulated polymorphic sites in plasma cfDNA samples from pregnant women who have legally accepted donor eggs. Sequencing to estimate fetal DNA concentration
[0567] The following example uses the allele counts of nine hypothetical polymorphic sites in the plasma cfDNA of a pregnant woman who has legally accepted donated eggs to briefly illustrate the method and steps for estimating the fetal DNA concentration in the sample using the allele count iterative fitting genotype method.
[0568] (1) The sequencing results of plasma cfDNA samples from pregnant women who have legally accepted donated eggs were used to count alleles at multiple polymorphic sites in the genome.
[0569] Polymorphic sites on the reference genome were selected and labeled Id001-Id009. The allele counts for each of the nine polymorphic sites simulated as described in Example 3 are shown in Table 4. In the hypothetical plasma cfDNA of a legally permitted donor mother, the reference genome is considered to be a chromosomal region with a normal disomy karyotype in both the mother and fetus; therefore, each polymorphic site theoretically contains a maximum of four alleles. Here, the counts for each site are shown as a maximum of five alleles. It should be understood that each polymorphic site may be detected to contain multiple alleles, and each allele should be counted.
[0570] Table 4: Allele counts at nine hypothetical polymorphic loci.
[0571]
[0572]
[0573] (2) The concentration of fetal DNA was iteratively estimated according to the method described in Example 13.
[0574] Since each polymorphic locus in the plasma of pregnant women who have legally accepted donated eggs can theoretically have a maximum of four alleles, the allele counts for each polymorphic locus are sorted in descending order and the four largest numbers are labeled as R1, R2, R3, and R4. Then, the sequencing noise threshold is set to α = 0.01, the iteration precision value is ε = 0.001, and the initial fetal concentration estimate is f0 = 0.10. Finally, the fetal DNA concentration is calculated according to the following steps.
[0575] Step (a) For each polymorphic locus, the genotype of the locus is estimated according to the methods described in Examples 11 and 12, based on the allele counts and f0, as well as the amplification count (FC) theoretically derived from fetal DNA and the total count (TC) derived from maternal and fetal DNA.
[0576] For example, for locus number Id0006, R1 to R4 are 3322, 936, 36, and 28 respectively, then O1 = 3322, O2 = 936, O3 = 36, and O4 = 28. Since R2 / (R1+R2)≥0.01 and R3 / (R1+R2+R3)<0.01, this locus has two detected allele counts above the noise threshold.
[0577] The goodness-of-fit test for all possible genotypes at this locus is as follows:
[0578] AA|AA: TC=R1+R2=4258, E1=(1-α)×TC=4215.42, E2=α×TC=42.58,
[0579] AA|AB: TC=R1+R2=4258, E1=(1-f0 / 2)×TC=4045.10, E2=f0 / 2×TC=212.90,
[0580] AB|AA: TC=R1+R2=4258, E1=(1+f0) / 2×TC=2341.90, E2=(1-f0) / 2×TC=1916.10,
[0581] AB|AB: TC=R1+R2=4258, E1=1 / 2×TC=2129.00, E2=1 / 2×TC=2129.00,
[0582] AB|AC: TC=R1+R2+R3=4294, E1=1 / 2×TC=2147.00, E2=(1-f0) / 2×TC=1932.30, E3=f0 / 2×TC=214.70,
[0583] AA|BB:TC=R1+R2=4258,E1=(1-f0)×TC=3832.20,E2=f0×TC=425.80,
[0584] AA|BC: TC=R1+R2+R3=4294, E1=(1-f0)×TC=3864.60, E2=f0 / 2×TC=214.70, E3=f0 / 2×TC=214.70,
[0585] AB|CC: TC=R1+R2+R3=4294, E1=(1-f0) / 2×TC=1932.30, E2=(1-f0) / 2×TC=1932.30, E3=f0×TC=429.40,
[0586] AB|CD:TC=R1+R2+R3+R4=4322,E1=(1-f0) / 2×TC=1944.90,E2=(1 -f0) / 2×TC=1944.90,E3=f0 / 2×TC=216.10,E4=f0 / 2×TC=216.10,G AB|CD =1944.34。
[0587] Due to G AA|BB <G AB|AA <G AB|AC <G AB|AB <G AA|AB <G AA|BC <G AB|CD <G AB|CC <G AA|AA Therefore, the genotype of locus Id006 is estimated to be AA|BB. Then, according to the method described in Example 12, FC = R2 = 936 and TC = R1 + R2 = 4258 are estimated.
[0588] Following the same rules, the FC and TC values were estimated for the above nine sites respectively.
[0589] Step (b) uses the FC and TC values of each polymorphic site to estimate the fetal DNA concentration f according to the method described in Example 8.
[0590] Step (c) determines whether the absolute value of f-f0 is less than ε. If the result is yes, the fetal DNA concentration is output as f, and the calculation ends. If the result is no, f0 is set to f, and then step (a) above is executed.
[0591] The results of iteratively executing the above example are shown in Table 5 below.
[0592] Table 5: Iterative parameter estimates of fetal DNA concentration.
[0593] 1 0.1 0.2385 0.1385 2 0.2385 0.2436 0.0051 3 0.2436 0.2436 0
[0594] Therefore, the concentration of fetal DNA in this example is estimated to be f = 0.2436.
[0595] Example 16: Estimation of fetal DNA concentration and allele count of the analyte in a pregnant woman's plasma DNA sample Calculate the genotype of this locus
[0596] According to Example 3, a set of reference group polymorphic sites and two target group polymorphic sites were obtained from a simulated pregnant woman's plasma cfDNA sample. Assuming the fetal DNA concentration f = 0.20 was estimated using the set of reference genomic polymorphic sites as described in Example 14, and the allele counts for the two target group polymorphic sites were A: 16994, 1896, 23; B: 9146, 7355, 1892, 58, respectively. If the chromosomes containing sites A and B in both the mother and fetus are normal dimers and there are no large insertion or deletion variations affecting sites A and B, then sites A and B can only be one of the following five genotypes: AA|AA, AA|AB, AB|AA, AB|AB, and AB|AC. The following uses the above allele count results for sites A and B as examples to estimate their most likely genotypes according to the method described in Example 11.
[0597] The goodness-of-fit of all possible genotypes at loci A and loci B was tested using the G-test, and the results are shown in Table 6 below.
[0598] Table 6: Estimation of target site genotypes using goodness-of-fit test.
[0599]
[0600] As can be seen from the results in Table 6, locus A has the best goodness-of-fit test result for genotype AA|AB, while locus B has the best goodness-of-fit test result for genotype AB|AC. Therefore, the estimated genotype of locus A is AA|AB, and the genotype of locus B is AB|AC.
[0601] Example 17: Utilizing the sample concentration f of the least component in the sample mixture and a set of polymorphic sites within the target region. Allele counting at points to estimate the karyotype of the target gene
[0602] The main steps for estimating the karyotype of the target chromosome or subchromosome segment using multiple polymorphic sites within the target region and a comprehensive goodness-of-fit test are as follows:
[0603] (1) Analyze the sample to be tested and list all possible karyotypes of the sample in the target area;
[0604] (2) For each possible karyotype, list all possible genotypes corresponding to that karyotype at each polymorphic site in the target region;
[0605] (3) For each polymorphic site in the target region, select a genotype with the best fit for each karyotype according to the method described in Example 11;
[0606] (4) Analyze the goodness-of-fit test results of all polymorphic sites in the target region to all karyotypes, and take the karyotype with the best comprehensive fit of polymorphic sites as the karyotype of the target (chromosome or subchromosome segment) to be detected.
[0607] Example 18: Using the fetal DNA concentration f in a pregnant woman's plasma DNA sample and the chromosome or subchromosome to be analyzed. Allele counts of a set of polymorphic loci within a horizontal region are used to estimate chromosome-level aneuploidy within the region to be analyzed. XOR subchromosome level deletion duplication variation
[0608] Two maternal plasma cfDNA samples were simulated as described in Example 3, with each sample simulating a set of reference genomic polymorphic sites and a set of polymorphic sites in a target region derived from a specific chromosomal or subchromosomal segment. Assuming that the concentration of fetal DNA in both samples was estimated to be f = 0.20 using a set of reference genomic polymorphic sites as described in Example 14, the allele counts of each polymorphic site in the target region of Sample 1 and Sample 2 are shown in Table 7 below.
[0609] Table 7: Allele counts of a set of polymorphic sites on the target chromosome in two hypothetical samples.
[0610]
[0611] Assume that a set of polymorphic loci in the target region of Sample 1 and Sample 2 originate from chromosome 21, and our goal is to detect whether the fetuses in Sample 1 and Sample 2 have trisomy 21, i.e., whether the karyotype of chromosome 21 in these two samples is disomy-disomy (both the mother and fetus have normal disomy of chromosome 21) or disomy-trisomy (a pregnant woman with normal disomy of chromosome 21 carries a fetus with trisomy 21). For disomy-disomy, all polymorphic loci can only be one of the following 5 genotypes: AA|AA, AA|AB, AB|AA, AB|AB, or AB|AC. For disomy-trisomy, all polymorphic loci can only be one of the following 10 genotypes: AA|AAA, AA|AAB, AA|ABB, AA|ABC, AB|AAA, AB|AAB, AB|AAC, AB|ABC, AB|ACC, or AB|ACD. For each set of polymorphic sites in the target region of chromosome 21 in Sample 1 and Sample 2, the goodness-of-fit test was performed according to the method described in Example 17, based on the karyotypes of disomic-disomic and disomic-trisomic, and the results are shown in Table 8 below.
[0612] Table 8: Results of karyotype goodness-of-fit test for allele counts at each polymorphic site in the target region.
[0613]
[0614]
[0615] For Sample 1, the allele counts of most polymorphic sites fit the genotypes in the dimer-disomy better than those in the dimer-trisomy, therefore the karyotype of Sample 1 is estimated to be dimer-disomy, meaning that both the mother and the fetus are normal dimers.
[0616] For Sample 2, the allele counts of all polymorphic sites fit the genotypes in trisomy-disomy better than those in disomy-disomy. Therefore, the karyotype of Sample 2 is estimated to be disomy-trisomy, i.e., the mother is normal disomy and the fetus is abnormal trisomy 21.
[0617] When considering the fitting results of multiple polymorphic sites, one can consider the karyotype that best fits most samples, or use G value, AIC value, modified G value and / or modified AIC value for judgment.
[0618] For example, if we fit a two-body-two-body karyotype to sample 1, then:
[0619] The overall value of G is ΣG i =0.0 + 0.039 + 0.025 + 2.138 + 0.054 = 2.256
[0620] The overall AIC value is ΣAIC i =0.0 + (-1.961) + (-1.975) + 0.138 + (-3.946) = -7.744
[0621] The sum of AIC and the total value is Σ(AIC) i / TC i = 0.0 / 9565 + (-1.961 / 6472) + (-1.975 / 11183) + 0.138 / 15494 + (-3.946 / 18915) = -0.00068
[0622] The sum of AIC / total count / f is Σ(AIC) i / TC i / f)=0.0 / 9565 / 0.2+(-1.961 / 6472 / 0.2)+(-1.975 / 11183 / 0.2)+0.138 / 15494 / 0.2+(-3.946 / 18915 / 0.2)=-0.0034.
[0623] For sample 1, fitted with a two-trisomal karyotype, then:
[0624] The overall value of G is ΣG i =319.73
[0625] The overall AIC value is ΣAIC i =309.73
[0626] The sum of AIC and total count is Σ(AIC) i / TC i ) = 0.02017
[0627] The sum of AIC / total count / f is Σ(AIC) i / TC i / f)=0.10087.
[0628] For sample 1, the combined G value, combined AIC value, combined AIC / total value, and combined AIC / total count / f value all showed a better fit for the dimelo-diso genotype than the corresponding fit for the dimelo-triso genotype. Therefore, these values or values derived from them can also be used to determine the fit of each allele at multiple polymorphic loci to different karyotypes.
[0629] When detecting microdeletion and microduplication variations at the subchromosomal level, the possibility that the mother may carry homozygous or heterozygous subchromosomal microdeletions or microduplications should be considered. Therefore, all possible genotypes for each affected polymorphic locus need to be taken into account and detected using a goodness-of-fit test. For example, detecting subchromosomal microdeletion mutations requires detecting all possible genotype combinations of the mother and fetus, assuming the mother is homozygous, heterozygous, or normal, and the fetus is also homozygous, heterozygous, or normal. Similarly, detecting subchromosomal microduplication mutations requires detecting all possible genotype combinations of the mother and fetus, assuming the mother is homozygous, heterozygous, or normal, and the fetus is also homozygous, heterozygous, or normal.
[0630] Example 19: Estimating the sample using high-throughput sequencing results of a set of polymorphic sites in a pregnant woman's plasma DNA sample. Mid-fetal DNA concentration
[0631] For each sample in the amplicon sequencing dataset of pregnant women's plasma cfDNA insertion and deletion markers (Barrett, Xiong et al. 2017, PLoS One 12: e0186771) as described in Example 1, the count of each allele in each insertion and deletion marker (polymorphic site) was counted. Then, for each polymorphic site in each sample, the count of fetal DNA (FC) and the total count of maternal and fetal DNA (TC) were estimated. Using the FC and TC of each polymorphic site in each sample, the concentration of fetal DNA in each sample was estimated.
[0632] Figure 3 This is the analysis result of a pregnant woman's plasma cfDNA sample from this dataset. The count of fetal DNA (FC) and the total count of maternal and fetal DNA (TC) at each insertion / deletion polymorphism site in the sample are represented by a point in the graph. Robust regression fitting (fitting model: FC ~ TC + 0) was performed using the FC and TC values of each polymorphic site in the sample and the rlm function from the MASS library in the R software package to estimate the concentration of fetal DNA. The result of the rlm robust regression fitting is the straight line in the graph, while the fetal DNA concentration is estimated as the slope of this line (the model coefficient of TC).
[0633] Example 20: Using high-throughput sequencing results of a set of polymorphic sites in a mixed DNA sample to estimate the most common polymorphic sites in the sample. DNA concentration of fewer components
[0634] For each sample in the mixed sample amplicon sequencing dataset (Kim, Kim et al. 2019, Nat Commun 10:1047) as described in Example 2, the count of each allele at each polymorphic site is counted. Then, for each polymorphic site in each sample as described in Example 8, the count of the least component DNA (FC) and the total count of all DNA (TC) are estimated. Using the FC and TC of each polymorphic site in each sample, the concentration of the least component DNA in each sample is estimated.
[0635] Figure 4 'a' represents the results of analyzing a mixed DNA sample from this dataset. The count of the least abundant DNA component (FC) and the total count (TC) of all DNA components at each polymorphic site in the sample are represented by a point in the graph. Robust RLM regression (model: FC ~ TC + 0) was performed using the FC and TC values for each polymorphic site to estimate the concentration of the least abundant DNA component in the sample. The RLM robust regression result is the fitted straight line in the graph, while the concentration of the least abundant DNA component is estimated as the slope of this line (the model coefficient of TC). Figure 4 b shows the analysis results for all mixed DNA samples in this dataset. The four mixed samples were replicated multiple times at the library preparation or sequencing level, with expected minimum component DNA concentrations of 0.01, 0.02, 0.10, or 0.20 (x-axis), while the estimated minimum component DNA concentration for each sample is shown on the y-axis. The dashed line in the figure represents the position of the line y = x.
[0636] Example 21: Computer simulation of chromosome level, subchromosome level and short sequence water content in pregnant women's plasma DNA samples Variations on flat surfaces
[0637] To detect genetic variation at the chromosomal, subchromosomal, or short-sequence levels, we simulated variations with karyotypes of disomy-monomy and disomy-trisomy at the chromosomal level; at the subchromosomal level, we simulated variations of absent-absent, absent-monomy, monosomy-absent, monosomy-monomy, monosomy-disomy, disomy-monomy, disomy-trisomy, trisomy-disomy, trisomy-trisomy, trisomy-tetrasomy, tetrasomy-trisomy, and tetrasomy-tetrasomy; and at the short-sequence level, we simulated all possible genotypes of any polymorphic locus within the normal disomy-disomy karyotype. The specific simulation process for different polymorphic loci in each sample is briefly described below:
[0638] 1. A simulated plasma DNA sample from a pregnant woman containing chromosomal monosomy.
[0639] To detect chromosomal aneuploidy at the chromosome level, we simulated maternal plasma DNA samples containing chromosomal monomorphisms. In each sample, three pairs of chromosomes were simulated for both the mother and fetus, numbered 1 (Chr01), 2 (Chr02), and 3 (Chr03). In each sample, 100 polymorphic sites were simulated on chromosomes 1, 2, and 3 as described in Example 3. A concentration of the simulated fetal DNA was randomly selected from the following concentrations (0.02, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45) for each sample.
[0640] The simulated chromosome 1 serves as the reference chromosome in the sample, where the genotype of each polymorphic locus is simulated as one of the normal dimo-diso genotypes, and the total count of alleles at each polymorphic locus is 200.
[0641] The simulated chromosome 2 in the sample is a dimeric-disomy chromosome, where the genotype of each polymorphic locus is simulated as one of the normal dimeric-disomy genotypes, and the total count of alleles at each polymorphic locus is 200.
[0642] The simulated chromosome 3 in the sample was a disomy-monomysome chromosome, with the genotype of each polymorphic locus simulated as one of the disomy-monomysome genotypes. Due to the lack of one fetal chromosome, the total count of alleles at each polymorphic locus was 200-100f.
[0643] Using the simulated allele sequence of each sample as the input file, the high-throughput sequencing results were simulated using ART simulation software (Huang, Li et al. 2012, Bioinformatics 28:593-594), where the fold parameter of the ART simulation software was set to 50 or 100.
[0644] 2. Simulate plasma DNA samples from pregnant women with trisomy 11.
[0645] To detect trisomy aneuploidy at the chromosomal level, we simulated plasma DNA samples from pregnant women with trisomy. In each sample, both the mother and fetus were simulated with three pairs of chromosomes, numbered 1 (Chr01), 2 (Chr02), and 3 (Chr03). In each sample, 100 polymorphic sites were simulated on chromosomes 1, 2, and 3 as described in Example 3. A concentration of the simulated fetal DNA was randomly selected from the following concentrations (0.02, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45) for each sample.
[0646] The simulated chromosome 1 serves as the reference chromosome in the sample, where the genotype of each polymorphic locus is simulated as one of the normal dimo-diso genotypes, and the total count of alleles at each polymorphic locus is 200.
[0647] The simulated chromosome 2 in the sample is a dimeric-disomy chromosome, where the genotype of each polymorphic locus is simulated as one of the normal dimeric-disomy genotypes, and the total count of alleles at each polymorphic locus is 200.
[0648] The simulated chromosome 3 in the sample is a disoproximity-trisomy chromosome, with the genotype of each polymorphic locus simulated as one of the disoproximity-trisomy genotypes. Due to the extra fetal chromosome, the total count of alleles at each polymorphic locus is 200 + 100f.
[0649] Using the simulated allele sequence of each sample as the input file, the ART simulation software was used to simulate high-throughput sequencing results, where the fold parameter of the ART simulation software was set to 50 or 100.
[0650] 3. Simulate plasma DNA samples from pregnant women containing subchromosomal microdeletions.
[0651] To detect microdeletion variations at the subchromosomal level, we simulated maternal plasma DNA samples containing chromosomal microdeletions. In each sample, both the mother and fetus were simulated with seven pairs of chromosomes, numbered 1 (Chr01), 2 (Chr02), 3 (Chr03), 4 (Chr04), 5 (Chr05), 6 (Chr06), and 7 (Chr07). In each sample, 100 polymorphic sites were simulated on chromosomes 1-7 as described in Example 3. For each sample, a concentration of the simulated fetal DNA was randomly selected from the following concentrations (0.02, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45). Here, each microdeletion region is treated as a whole chromosome, and polymorphic sites are selected from that microdeletion region. In a single genome, a chromosome with one normal chromosome and one chromosome with a microdeletion is marked as a monosomy, while two chromosomes with microdeletions are marked as a null chromosome.
[0652] The simulated chromosome 1 serves as the reference chromosome in the sample, where the genotype of each polymorphic locus is simulated as one of the normal dimo-diso genotypes, and the total count of alleles at each polymorphic locus is 200.
[0653] The simulated chromosome 2 in the sample is a dimeric-disomy chromosome, where the genotype of each polymorphic locus is simulated as one of the normal dimeric-disomy genotypes, and the total count of alleles at each polymorphic locus is 200.
[0654] The simulated chromosome 3 in the sample was a disomy-monomysome chromosome, with the genotype of each polymorphic locus simulated as one of the disomy-monomysome genotypes. Since one fetal chromosome contains microdeletions, the total count of alleles at each polymorphic locus was 200-100f.
[0655] The simulated chromosome 4 in the sample is a monosomy-disomy chromosome, with the genotype of each polymorphic locus simulated as one of the monosomy-disomy genotypes. Since one maternal chromosome contains a microdeletion, the total count of alleles at each polymorphic locus is 100+100f.
[0656] The simulated chromosome 5 in the sample is a monosomy-monosomy chromosome, with the genotype of each polymorphic locus simulated as one of the monosomy-monosomy genotypes. Since both the maternal chromosome and the fetal chromosome contain microdeletions, the total count of alleles at each polymorphic locus is 100.
[0657] The simulated chromosome 6 in the sample was a monosomy-deletion chromosome, with the genotype of each polymorphic locus simulated as one of the monosomy-deletion genotypes. Since both the maternal chromosome and the two fetal chromosomes contained microdeletions, the total count of alleles at each polymorphic locus was 100-100f.
[0658] The simulated chromosome 7 in the sample is a null-null chromosome, with the genotype of each polymorphic locus simulated as one of the null-null genotypes. Since both maternal chromosomes and both fetal chromosomes contain microdeletions, the total count of alleles at each polymorphic locus is 0, meaning that specific amplified sequences are not simulated, or some random sequences are simulated that cannot be located on any chromosome.
[0659] Using the simulated allele sequence of each sample as the input file, the ART simulation software was used to simulate high-throughput sequencing results, where the fold parameter of the ART simulation software was set to 50 or 100.
[0660] 4. Simulate plasma DNA samples from pregnant women containing subchromosomal microduplications.
[0661] To detect subchromosomal microduplications, we simulated maternal plasma DNA samples containing subchromosomal microduplications. In each sample, both the mother and fetus were simulated with seven pairs of chromosomes, numbered 1 (Chr01), 2 (Chr02), 3 (Chr03), 4 (Chr04), 5 (Chr05), 6 (Chr06), and 7 (Chr07). In each sample, 100 polymorphic sites were simulated on chromosomes 1-7 as described in Example 3. For each sample, a concentration was randomly selected from the following concentrations (0.02, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45) as the simulated fetal DNA concentration. Here, each microduplication region is treated as two chromosomes, and the polymorphic site is selected from the microduplication region. Therefore, in a single genome, a normal chromosome and a chromosome containing a microduplication are marked as trisomy, while two chromosomes containing microduplications are marked as tetrasomy.
[0662] The simulated chromosome 1 serves as the reference chromosome in the sample, where the genotype of each polymorphic locus is simulated as one of the normal dimo-diso genotypes, and the total count of alleles at each polymorphic locus is 200.
[0663] The simulated chromosome 2 in the sample is a dimeric-disomy chromosome, where the genotype of each polymorphic locus is simulated as one of the normal dimeric-disomy genotypes, and the total count of alleles at each polymorphic locus is 200.
[0664] The simulated chromosome 3 in the sample is a disoproximity-trisomy chromosome, with the genotype of each polymorphic locus simulated as one of the disoproximity-trisomy genotypes. Since a fetal chromosome contains microduplications, the total count of alleles at each polymorphic locus is 200 + 100f.
[0665] The simulated chromosome 4 in the sample is a trisomy-disomy chromosome, with the genotype of each polymorphic locus simulated as one of the trisomy-disomy genotypes. Since a maternal chromosome contains microduplications, the total count of alleles at each polymorphic locus is 300-100f.
[0666] The simulated chromosome 5 in the sample is a trisomy-trisomy chromosome, with the genotype of each polymorphic locus simulated as one of the trisomy-trisomy genotypes. Since both the maternal chromosome and the fetal chromosome contain microduplications, the total count of alleles at each polymorphic locus is 300.
[0667] The simulated chromosome 6 in the sample is a trisomy-tetrasomy chromosome, with the genotype of each polymorphic locus simulated as one of the trisomy-tetrasomy genotypes. Since both the maternal chromosome and the two fetal chromosomes contain microduplications, the total count of alleles at each polymorphic locus is 300 + 100f.
[0668] The simulated chromosome 7 in the sample was a tetrasomic-tetrasomic chromosome, with the genotype of each polymorphic locus simulated as one of the tetrasomic-tetrasomic genotypes. Since both maternal chromosomes and both fetal chromosomes contain microduplications, the total count of alleles at each polymorphic locus was 400.
[0669] Using the simulated allele sequence of each sample as the input file, the ART simulation software was used to simulate high-throughput sequencing results, where the fold parameter of the ART simulation software was set to 50 or 100.
[0670] 5. Simulate plasma DNA samples from pregnant women containing short-sequence level variations.
[0671] To detect short-sequence level variations, we simulated maternal plasma DNA samples containing short-sequence level variation sites. In each sample, two pairs of chromosomes were simulated for both the mother and fetus, designated as Chr01 and Chr02, respectively. In each sample, 100 polymorphic sites were simulated on chromosomes 1 and 2 as described in Example 3. A concentration of the simulated fetal DNA was randomly selected from the following concentrations (0.02, 0.05, 0.10, 0.15, 0.20, 0.25, 0.30, 0.35, 0.40, 0.45) for each sample.
[0672] The simulated chromosome 1 serves as the reference chromosome in the sample, where the genotype of each polymorphic locus is simulated as one of the normal dimo-diso genotypes, and the total count of alleles at each polymorphic locus is 200.
[0673] The simulated chromosome 2 in the sample was a dimeric-disomy chromosome, with a total count of 200 alleles simulated at each locus. For any simulated locus, one allele was selected as wild-type (normal, represented by uppercase A), and the remaining alleles were selected as mutants (represented by lowercase a, b, c, or d, respectively). Each simulated locus could only be one of the following 14 genotypes: AA|AA, AA|Aa, Aa|AA, Aa|Aa, Aa|Ab, Aa|aa, Aa|ab, aa|Aa, aa|aa, aa|ab, ab|Aa, ab|aa, ab|ab, or ab|ac. One hundred test loci on chromosome 2 were randomly simulated, with each locus randomly selected from the 14 genotypes. The sequences of each allele were then simulated proportionally according to the set fetal DNA concentration and the method described in Example 3.
[0674] Using the simulated allele sequence of each sample as the input file, the ART simulation software was used to simulate high-throughput sequencing results, where the fold parameter of the ART simulation software was set to 50 or 100.
[0675] 6. Simulate a single genome sample.
[0676] To detect variations in a single genome at the chromosomal or subchromosomal level, we simulated genomic DNA samples from non-pregnant women (e.g., preimplantation embryonic genomic DNA samples), with each sample simulating chromosome 5, numbered 1 through 5 (Chr01). In each sample, chromosomes 1 through 5 were simulated with 100 polymorphic sites as described in Example 3. Here, normal chromosomes were labeled as disomy, each microdeletion region was treated as a single chromosome, and each microduplication region was treated as two chromosomes, with polymorphic sites selected from these microdeletion / microduplication regions. Specifically, a single normal chromosome and a chromosome with a microdeletion were labeled as monosomy, two chromosomes with microdeletions were labeled as nullisomes, a single normal chromosome and a chromosome with a microduplication were labeled as trisomy, and two chromosomes with microduplications were labeled as tetrasomy.
[0677] The simulated chromosome 1 in the sample is a normal disomy chromosome, where the genotype of each polymorphic locus is simulated as one of the normal disomy genotypes (AA or AB), and the total count of each allele at each polymorphic locus is 200.
[0678] The simulated chromosome 2 in the sample is either a deletion or a homozygous microdeletion chromosome, with the genotype of each polymorphic locus simulated as a normal deletion or homozygous microdeletion genotype. Since the total count of all alleles at each polymorphic site is 0, it does not simulate the generation of specific amplified sequences or simulates some random sequences that cannot be located on any chromosome.
[0679] The simulated chromosome 3 in the sample is either a monosomy or a heterozygous microdeletion chromosome, with the genotype of each polymorphic locus simulated as either a monosomy or a heterozygous microdeletion genotype. The total count of all alleles at each polymorphic site is 100.
[0680] The simulated chromosome 4 in the sample is a trisomic or heterozygous microduplicated chromosome, in which the genotype of each polymorphic locus is simulated as one of the trisomic or heterozygous microduplicated genotypes (AAA, AAB, or ABC), and the total count of each allele at each polymorphic locus is 300.
[0681] The simulated chromosome 5 in the sample is a tetrasomic or homozygous microduplicated chromosome, in which the genotype of each polymorphic locus is simulated as one of the tetrasomic or homozygous microduplicated genotypes (AAAA, AAAB, AABB, AABC, or ABCD), and the total count of each allele at each polymorphic locus is 400.
[0682] Using the simulated allele sequence of each sample as the input file, the ART simulation software was used to simulate high-throughput sequencing results, where the fold parameter of the ART simulation software was set to 50 or 100.
[0683] Example 22: Detection of fetal DNA concentration and allele counting of the analyte in maternal plasma DNA samples. Testing for fetal chromosomal monosomy
[0684] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing a single chromosome, wherein chromosomes 1, 2, and 3 are the reference chromosome, the chromosome with a normal dimocytic-disomic karyotype, and the chromosome with an abnormal dimocytic-monocytic karyotype, respectively.
[0685] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosome 2 or 3, the karyotype of chromosome 2 or 3 was estimated respectively, following the method described in Example 17. To detect whether there is a monosomy abnormality on fetal chromosome 2 or 3, we need to consider whether the allele counts of each polymorphic site on chromosome 2 or 3 provide a better overall goodness-of-fit test result for the genotype with a karyotype of disomy-disomy or the genotype with a karyotype of disomy-monomy.
[0686] Figure 5The image shows the detection of chromosomal abnormalities in a simulated fetal sample using a goodness-of-fit test. Figure 5 'a' utilizes the comprehensive goodness-of-fit test results to detect fetal monosomy abnormalities in normal disomy-disomy karyotype chromosomes in simulated samples. The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value of the G test at that locus by the fetal concentration and then by the total allele count at that locus. Figure 5 Method b utilizes the comprehensive goodness-of-fit test results to detect fetal monosomy abnormalities in the simulated sample of disomy-monomyocytic chromosomes. For the normal chromosome (disomy-monomyocytic chromosome 2), almost all polymorphic sites fit the disomy-monomyocytic genotype well, but not well for the disomy-monomyocytic genotype. For the abnormal chromosome (disomy-monomyocytic chromosome 3), almost all polymorphic sites fit the disomy-monomyocytic genotype well, but not well for the disomy-monomyocytic genotype. Therefore, the test results show no monosomy abnormality on fetal chromosome 2, but a monosomy abnormality on fetal chromosome 3.
[0687] Example 23: Detection of fetal DNA concentration and allele counting of the analyte in maternal plasma DNA samples. Testing for fetal trisomy
[0688] The plasma DNA sample of a pregnant woman with trisomy was simulated according to the method described in Example 21, wherein chromosomes 1, 2 and 3 are the reference chromosome, the chromosome with a normal dimo-disomy karyotype and the chromosome with an abnormal dimo-trisomy karyotype, respectively.
[0689] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosome 2 or 3, the karyotype of chromosome 2 or 3 was estimated respectively, following the method described in Example 17. To detect whether fetal chromosome 2 or 3 has trisomy, we need to consider whether the allele counts of each polymorphic site on chromosome 2 or 3 provide a better overall goodness-of-fit test result for genotypes with a karyotype of disomy-disomy or genotypes with a karyotype of disomy-trisomy.
[0690] Figure 6 The image shows the detection of trisomy chromosomal abnormalities in a simulated fetus using a goodness-of-fit test. Figure 6 'a' utilizes the comprehensive goodness-of-fit test results to detect fetal trisomy abnormalities in simulated samples with normal disomy-disomy karyotype chromosomes. The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value of the G test at that locus by the fetal concentration and then by the total allele count at that locus. Figure 6Method b utilizes the comprehensive goodness-of-fit test results to detect fetal trisomy abnormalities in the simulated sample with a disomy-trisomy karyotype. For the normal chromosome (disomy-trisomy chromosome 2), almost all polymorphic sites fit the disomy-trisomy karyotype genotype well, but not well with the disomy-trisomy genotype. For the abnormal chromosome (disomy-trisomy chromosome 3), almost all polymorphic sites fit the disomy-trisomy karyotype genotype well, but not well with the disomy-trisomy karyotype. Therefore, the test results show no trisomy abnormality on fetal chromosome 2, but trisomy abnormality on fetal chromosome 3.
[0691] Example 24: Detection of fetal DNA concentration and allele counting of the analyte in maternal plasma DNA samples. Detection of fetal chromosomal microdeletion abnormalities
[0692] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing chromosomal microdeletions, wherein chromosomes 1 to 7 are respectively the reference chromosome, a chromosome in which both the mother and fetus are normal (a normal disodium-disodium karyotype chromosome), a chromosome in which the mother is normal and the fetus has one chromosome with a microdeletion (a disodium-monocytoplasmic karyotype chromosome), a chromosome in which the mother has one chromosome with a microdeletion and the fetus is normal (a monosodium-disodium karyotype chromosome), a chromosome in which both the mother and fetus have one chromosome with a microdeletion (a monosodium-monocytoplasmic karyotype chromosome), a chromosome in which the mother has one chromosome with a microdeletion and both of the fetus's chromosomes have microdeletions (a monosodium-delete karyotype chromosome), and a chromosome in which both the mother and fetus have two chromosomes with microdeletions (a delete karyotype chromosome).
[0693] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosomes 2 to 7, the karyotypes of chromosomes 2 to 7 were estimated respectively, following the method described in Example 17. To detect whether a fetal chromosome has a microdeletion abnormality, a goodness-of-fit test was performed on each possible maternal-fetal microdeletion karyotype using the allele counts of each polymorphic site on that chromosome. The karyotype that best fits the allele counts of all polymorphic sites was then used to determine whether the fetal chromosome has a microdeletion abnormality.
[0694] Figure 7 The image shows the detection of microdeletion anomalies in fetal chromosomes in a simulated sample using a goodness-of-fit test. Figure 7'a' utilizes the comprehensive goodness-of-fit test results to detect fetal chromosomal microdeletion abnormalities in simulated samples with monosomy-disomy karyotypes (heterozygous microdeletion in the mother is considered normal in the fetus). The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value from the G-test at that locus by the fetal concentration and then by the total allele count at that locus. Figure 7 b is Figure 7 A local magnification of a. Figure 7 c uses the results of the comprehensive goodness-of-fit test to detect fetal chromosomal microdeletion abnormalities in monosomy-monosomy karyotype chromosomes (both the mother and fetus are heterozygous microdeletions) in the simulated sample. Figure 7 d is Figure 7 A magnified view of c. For chromosomes with a normal monosomy-disomy karyotype in the fetus, almost all polymorphic sites fit the monosomy-disomy genotype well, but poorly for other possible karyotypes. For chromosomes with a monosomy-monomyctic karyotype containing microdeletions in the fetus, almost all polymorphic sites fit the monosomy-monomyctic genotype well, but poorly for other possible karyotypes. Therefore, Figure 7 a and Figure 7 The test results for chromosome b showed no microdeletion abnormalities on that chromosome of the fetus. Figure 7 c and Figure 7 The test results for d showed a microdeletion abnormality on that chromosome number in the fetus.
[0695] Example 25: Detection of fetal DNA concentration and allele counting of the analyte in maternal plasma DNA samples. Detection of fetal chromosomal microduplications
[0696] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing chromosomal microduplications, wherein chromosomes 1 to 7 are respectively the reference chromosome, chromosomes in which both the mother and fetus are normal (normal disodium-disodium karyotype chromosome), chromosomes in which the mother is normal and the fetus has one chromosome with microduplication (disodium-trisodium karyotype chromosome), chromosomes in which the mother has one chromosome with microduplication and the fetus is normal (trisodium-disodium karyotype chromosome), chromosomes in which both the mother and fetus have one chromosome with microduplication (trisodium-trisodium karyotype chromosome), chromosomes in which the mother has one chromosome with microduplication and both of the fetus's chromosomes have microduplications (trisodium-tetrasodium karyotype chromosome), and chromosomes in which both the mother and fetus have two chromosomes with microduplications (tetrasodium-tetrasodium karyotype chromosome).
[0697] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosomes 2 to 7, the karyotypes of chromosomes 2 to 7 were estimated respectively, following the method described in Example 17. To detect whether a fetal chromosome has a microduplication abnormality, a goodness-of-fit test was performed on each possible maternal-fetal microduplication karyotype using the allele counts of each polymorphic site on that chromosome. The karyotype that best fits the allele counts of all polymorphic sites was then used to determine whether the fetal chromosome has a microduplication abnormality.
[0698] Figure 8 The image shows the detection of microduplications in fetal chromosomes in a simulated sample using a goodness-of-fit test. Figure 8 'a' utilizes the comprehensive goodness-of-fit test results to detect fetal chromosomal microduplication abnormalities in a simulated sample with trisomy-disomy karyotype chromosomes (the fetus is considered normal if the mother is heterozygous for microduplication). The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value from the G-test at that locus by the fetal concentration and then by the total allele count at that locus. Figure 8 b is Figure 8 A local magnification of a. Figure 8 c uses the results of the comprehensive goodness-of-fit test to detect fetal chromosomal microduplication abnormalities in trisomy-trisomy karyotype chromosomes (both the mother and fetus are heterozygous microduplications) in the simulated sample. Figure 8 d is Figure 8 A local magnification of c. For chromosomes with a normal trisomy-disomy karyotype in the fetus, almost all polymorphic sites fit the trisomy-disomy karyotype well, but poorly for other possible karyotypes. For chromosomes with a trisomy-trisomy karyotype containing microduplications in the fetus, almost all polymorphic sites fit the trisomy-trisomy karyotype well, but poorly for other possible karyotypes. Therefore, Figure 8 a and Figure 8 The test results for chromosome b showed no microduplication abnormalities on that chromosome in the fetus. Figure 8 c and Figure 8 The test results for d showed that a microduplication was found on chromosome number d in the fetus.
[0699] Example 26: Detection of fetal DNA concentration and allele counting of the analyte in maternal plasma DNA samples. Wild-type mutants of the sites to be analyzed
[0700] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing a specific short sequence site variation, wherein chromosomes 1 and 2 are the reference chromosome and the chromosome containing the specific short sequence site variation, respectively. Specifically, the polymorphic sites on chromosome 1 are selected from different chromosomal regions, while the multiple polymorphic sites on chromosome 2 are selected from the same specific site but are the results of independent amplification using the same and / or different primers. In other words, the simulated polymorphic sites on chromosome 2 represent different independent repeats of a specific site.
[0701] To detect wild-type mutants at specific loci, we adopted two approaches: (1) directly performing goodness-of-fit tests on all possible wild-type mutant genotypes and comprehensively analyzing the results of the goodness-of-fit tests; (2) first estimating the genotype of the locus to be tested without distinguishing wild-type mutant alleles, and then determining the wild-type mutant of each allele in the estimated genotype, thereby determining the wild-type mutant of each allele of the mother and / or fetus.
[0702] (1) Perform a goodness-of-fit test on all possible wild-type mutant genotypes. (a) Estimate the concentration f of fetal DNA in the sample using the allele counts of each polymorphic locus on reference chromosome 1, as described in Example 8. (b) List all possible wild-type mutant genotypes at this specific locus on chromosome 2, namely AA|AA, AA|Aa, Aa|AA, Aa|Aa, Aa|Ab, Aa|aa, Aa|ab, aa|Aa, aa|aa, aa|ab, ab|Aa, ab|aa, ab|ab, and ab|ac, where A represents the wild-type allele and a, b, and c represent the mutant alleles. (c) For each wild-type mutant genotype, estimate the theoretical count of its alleles based on the concentration f of fetal DNA in the sample estimated in step (a) above. (d) For each wild-type mutant genotype, determine its actual count based on the nucleic acid sequence of its alleles. (e) Perform a goodness-of-fit test on each wild-type mutant genotype for each locus's independent replicate. (f) Analyze the results of the goodness-of-fit test and select the wild-type mutant that has the best overall fit for all repeat sites as the estimated genotype for that specific site. (g) Determine the wild-type mutants of each allele in the mother and / or fetus based on the estimated wild-type mutant genotype.
[0703] (2) First estimate the genotype without distinguishing wild mutant alleles, and then determine the wild mutant type of each allele of the mother and / or fetus based on the wild mutant nucleic acid sequence of each allele.
[0704] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic locus on chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each specific short sequence locus to be tested on chromosome 2, the genotype of each specific short sequence locus on chromosome 2 was estimated according to the method described in Example 11. To detect whether the fetus has short-sequence level genetic variations, such as point mutations or short insertion / deletion mutations leading to certain single-gene genetic diseases, the genotype of each repeat locus to be tested was first estimated according to the method described in Example 11, without considering whether the allele sequences belong to the wild-type sequence. Then, the presence of variation at that locus in the mother and fetus was determined based on whether the sequences of each allele were normal wild-type sequences.
[0705] Figure 9 The image shows the wild-type mutation of a short fetal sequence site in a simulated sample, detected using a goodness-of-fit test. Figure 9 'a' uses the goodness-of-fit test results to detect the genotype of a short sequence locus simulating a maternal heterozygous mutation but with a normal fetus (different points represent different independent repeats of the target locus to be tested). The y-axis AIC value is a corrected AIC value, obtained by dividing the AIC value of the locus obtained from the G-test by the fetal concentration and then by the total allele count at that locus. Figure 9 b is Figure 9 A magnified view of the locus 'a'. The results showed that the mother was heterozygous while the fetus had a homozygous genotype (AB|AA). Further analysis revealed that allele A was wild-type while allele B was mutant, thus confirming that the mother had a heterozygous mutation at this locus while the fetus was normal. Figure 9 c is used to detect the genotype of short sequence sites with heterozygous mutations in both the simulated mother and fetus, based on the goodness-of-fit test results. Figure 9 d is Figure 9 A magnified view of c. The results showed that both the mother and fetus were heterozygous (AB|AC). Further analysis revealed that allele A was wild-type while alleles B and C were mutant. Therefore, it was determined that both the mother and fetus had a heterozygous mutation at this locus, and the fetus either developed a de novo mutation or inherited a paternal allele mutation.
[0706] Example 27: Using the sample concentration f of the least component in the sample mixture and the allele count of the locus to be analyzed. Relative distribution plots are used to estimate the genotype at this locus.
[0707] For a locus to be detected, estimate the genotype of that locus according to the following steps:
[0708] (1) Analyze the sample to be tested and list all possible genotypes of the target site;
[0709] (2) Calculate the theoretical relative count of each allele in each possible genotype, and plot the theoretical relative count of at least one non-maximum allele relative count against the maximum allele relative count for each genotype to mark the theoretical position of all possible genotypes.
[0710] (3) Calculate the relative counts of each allele at the target DNA site, and select at least one non-maximum relative allele count to plot against the maximum relative allele count to mark the actual position of the relative allele count at the target DNA site.
[0711] (4) Infer the genotype based on the theoretical and actual position distribution of the target DNA site in the relative allele count graph.
[0712] Figure 10 The diagram shows the theoretical distribution of polymorphic sites derived from normal karyotype chromosomes in a pregnant woman's plasma DNA sample, as shown in the relative allele distribution map. Figure 10 'a' represents the theoretical value of the total number of all possible genotypes and the relative counts of each allele at the polymorphic site on a normal dimo-diso karyotype chromosome. Figure 10 b shows the distribution of the second-largest relative allele count (RR2) relative to the largest relative allele count (RR1) at each polymorphic locus on a normal dimelo-diso karyotype chromosome. The results indicate that each polymorphic locus is located in a different position on the relative allele count distribution map due to different genotypes, and its genotype can be inferred from its specific distribution position.
[0713] Example 28: Utilizing the sample concentration f of the least component in the sample mixture and a set of polymorphic sites within the target region. Allele count relative distribution map of points to estimate the karyotype of the target.
[0714] We utilize the relative distribution map of allele counts at various polymorphic loci within the target region to detect aneuploidy at the chromosomal level or deletion / duplication variations at the subchromosomal level. The main steps are as follows:
[0715] (1) Analyze the sample to be tested and list all possible karyotypes of the sample on the target chromosome or subchromosome segment;
[0716] (2) For each possible karyotype, list all possible genotypes of the target DNA site on the chromosome or subchromosome of the karyotype in the sample. Then, for each genotype, select at least one non-maximum theoretical allele relative count value and plot it against the maximum theoretical allele relative count value to mark the theoretical position of the genotype.
[0717] (3) For each target DNA site in the target group, calculate the relative count of each allele and select at least one non-maximum relative allele count to plot against the maximum relative allele count to mark the actual location of the site.
[0718] (4) The karyotype of the target chromosome or subchromosome segment to be detected is inferred based on the theoretical and actual position distribution of all target DNA sites in the relative allele count graph.
[0719] Figure 11 The diagram shows the theoretical distribution of various polymorphic sites on chromosomes with normal maternal function but aneuploidy in the fetus, as shown in the relative allele distribution map. Figure 11 'a' represents the theoretical value of the total number of all possible genotypes and their relative counts at polymorphic sites on the hemi-hemi-karyotype and hemi-monotype chromosomes. Figure 11 b is the theoretical distribution of the second largest relative allele count (RR2) relative to the largest relative allele count (RR1) at each polymorphic site on the dimer-diso and dimer-monosomal chromosomes. Figure 11 c represents the theoretical value of the total number of all possible genotypes and the relative counts of all alleles at each polymorphic locus on the chromosomes of the dimo-diso and dimo-triso karyotypes. Figure 11 d is the theoretical distribution of the second or fourth largest relative allele count (RR2 or RR4) relative to the largest relative allele count (RR1) at each polymorphic locus on the dimer-diso and dimer-triso karyotype chromosomes.
[0720] Figure 12 The diagram shows the theoretical distribution of various polymorphic sites on subchromosomes with microdeletions or microduplications in maternal or fetal plasma DNA samples, in terms of the relative distribution of alleles. Figure 12 'a' represents the theoretical value of the total number of all possible genotypes and their relative counts of all alleles at the polymorphic site on the chromosome with a microdeletion karyotype in the mother or fetus. Figure 12 b is the theoretical distribution of the second largest relative allele count (RR2) relative to the largest relative allele count (RR1) at each polymorphic site on the chromosome with a microdeletion karyotype in the mother or fetus. Figure 12 c represents the theoretical value of the total number of all possible genotypes and their relative allele counts at each polymorphic locus on a normal subchromosome in a fetus with a mother having microduplication. Figure 12 d is the theoretical distribution of the relative counts of the second or third largest alleles (RR2 or RR3) relative to the largest relative allele (RR1) at each polymorphic locus on a normal subchromosome in a fetus with microduplication in the mother.
[0721] Example 29: Utilizing the sample concentration f of the least component in the sample mixture and the wild-type of the site to be analyzed, and The relative counts of each non-wildtype allele are used to estimate the wild-type mutant at that locus.
[0722] We use the wild-type allele count and the count of each non-wild-type allele at the locus to be analyzed to detect the wild-type mutant at that locus. The main steps are as follows:
[0723] (1) Analyze the sample to be tested and list the wild-type sequence of the target site and all possible genotypes of the target site;
[0724] (2) Calculate the theoretical relative counts of wild-type alleles and other non-wild-type alleles in each possible genotype, and plot the theoretical relative counts of at least one non-wild-type allele against the theoretical relative counts of wild-type alleles for each genotype to mark the theoretical positions of all possible genotypes.
[0725] (3) Calculate the relative counts of wild-type alleles and other non-wild-type alleles at the target DNA site in the sample, and select at least one non-wild-type allele relative count to plot the relative count of wild-type alleles to mark the actual location of the genotype of the target DNA site.
[0726] (4) Infer wild mutants based on the theoretical and actual position distribution of target DNA sites in the relative allele count graph.
[0727] Figure 13 The image shows the relative distribution of allele counts for each possible genotype on the test site at a normal dimo-diso chromosome in a pregnant woman's plasma DNA sample. Figure 13 'a' represents the theoretical value of all possible genotypes and the relative counts of each allele at the test site on a normal dimotic-disosome chromosome. Figure 13 b is the theoretical distribution of the largest relative count of non-wild-type alleles (RR2) relative to the relative count of wild-type alleles (RR1) at the test site on a normal dimelo-diso chromosome. Here, A represents the wild-type allele, and a, b, or c represent the non-wild-type (mutant) allele.
[0728] Example 30: Using fetal DNA concentration and allele counting of the analyte in a pregnant woman's plasma DNA sample Detection of fetal chromosomal monosomy using distribution maps
[0729] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing a single chromosome, wherein chromosomes 1, 2, and 3 are the reference chromosome, the chromosome with a normal dimocytic-disomic karyotype, and the chromosome with an abnormal dimocytic-monocytic karyotype, respectively.
[0730] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosome 2 or 3, the karyotype of chromosome 2 or 3 was estimated respectively, following the method described in Example 28. To detect whether there is a monosomy abnormality on fetal chromosome 2 or 3, we need to determine whether chromosome 2 or 3 has a normal disoproximity-disoproximity karyotype (both mother and fetus are disoproximities) or an abnormal disoproximity-monomoproximity karyotype (the mother has normal disoproximity while the fetus has abnormal monosomy). Therefore, we first marked the theoretical positions of all disoproximity-disoproximity and disoproximity-monomoproximity genotypes on the relative distribution map of allele counts, and then determined the karyotype of the chromosome based on the distribution of each polymorphic site on the chromosome to be analyzed on the relative distribution map of allele counts.
[0731] Figure 14 It uses a relative distribution map of allele counts at polymorphic sites to detect monosomy variations in fetal chromosomes. Figure 14 'a' plots the relative allele counts at all polymorphic sites on a simulated normal dimo-diso chromosome. Figure 14 b is a plot of the relative allele counts for all polymorphic sites on a simulated disomy-monomysome chromosome. The results show that... Figure 14 In genotype a, the relative counts of almost all polymorphic sites are distributed around the corresponding disomic-disomic genotype clusters, while they are almost entirely distributed around the corresponding disomic-monomorphic genotype clusters. Figure 14 In b, almost all polymorphic sites were relatively counted around the corresponding disomy-monad genotype clusters, while almost none were found around the corresponding disomy-disomy genotype clusters. Therefore, Figure 14 The karyotype of chromosome a to be analyzed is disoproximity-disoproximity, meaning that the fetus has a normal chromosome for that chromosome; while Figure 14 The karyotype of chromosome b to be analyzed is disomy-monomymy, meaning that the fetal chromosome is an abnormal monosomy.
[0732] Example 31: Using fetal DNA concentration and allele counting of the analyte in a pregnant woman's plasma DNA sample Detection of fetal trisomy chromosomal abnormalities using distribution maps
[0733] The plasma DNA sample of a pregnant woman with trisomy was simulated according to the method described in Example 21, wherein chromosomes 1, 2 and 3 are the reference chromosome, the chromosome with a normal dimo-disomy karyotype and the chromosome with an abnormal dimo-trisomy karyotype, respectively.
[0734] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosome 2 or 3, the karyotype of chromosome 2 or 3 was estimated respectively, following the method described in Example 28. To detect whether fetal chromosome 2 or 3 has trisomy, we need to determine whether chromosome 2 or 3 has a normal disoprosomy-disomy karyotype (both mother and fetus are disoprosomy) or an abnormal disoprosomy-trisomy karyotype (the mother is normal disoprosomy while the fetus is abnormal trisomy). Therefore, we first marked the theoretical positions of all disoprosomy-disomy and disoprosomy-trisomy genotypes on the relative distribution map of allele counts, and then determined the karyotype of the chromosome based on the distribution of each polymorphic site on the chromosome to be analyzed on the relative distribution map of allele counts.
[0735] Figure 15 It uses a relative distribution map of allele counts at polymorphic sites to detect trisomy variations in fetal chromosomes. Figure 15 'a' plots the relative allele counts at all polymorphic sites on a simulated normal dimo-diso chromosome. Figure 15 b) plots the relative allele counts at all polymorphic sites on simulated disomic-trisomic chromosomes. The results show that... Figure 15 In genotype a, the relative counts of almost all polymorphic sites are distributed around the corresponding dimer-diso genotype clusters, while they are almost entirely absent around the corresponding dimer-triso genotype clusters. Figure 15 In b, almost all polymorphic sites were relatively counted around the corresponding disomic-trisomic genotype clusters, while almost none were found around the corresponding disomic-monomic genotype clusters. Therefore, Figure 15 The karyotype of chromosome a to be analyzed is disoproximity-disoproximity, meaning that the fetus has a normal chromosome for that chromosome; while Figure 15 The karyotype of chromosome b to be analyzed is disomy / trisomy, meaning that the fetus has an abnormal trisomy of this chromosome.
[0736] Example 32: Using fetal DNA concentration and allele counting of the analyte in a pregnant woman's plasma DNA sample Detection of fetal chromosomal microdeletion abnormalities using distribution maps
[0737] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing chromosomal microdeletions, wherein chromosomes 1 to 7 are respectively the reference chromosome, a chromosome in which both the mother and fetus are normal (a normal disodium-disodium karyotype chromosome), a chromosome in which the mother is normal and the fetus has one chromosome with a microdeletion (a disodium-monocytoplasmic karyotype chromosome), a chromosome in which the mother has one chromosome with a microdeletion and the fetus is normal (a monosodium-disodium karyotype chromosome), a chromosome in which both the mother and fetus have one chromosome with a microdeletion (a monosodium-monocytoplasmic karyotype chromosome), a chromosome in which the mother has one chromosome with a microdeletion and both of the fetus's chromosomes have microdeletions (a monosodium-delete karyotype chromosome), and a chromosome in which both the mother and fetus have two chromosomes with microdeletions (a delete karyotype chromosome).
[0738] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosomes 2 to 7, the karyotypes of chromosomes 2 to 7 were estimated respectively, following the method described in Example 28. To detect whether a fetal chromosome has a microdeletion abnormality, we need to determine whether the chromosome has a normal disoproximity-disoproximity karyotype (both mother and fetus are disoproximities) or an abnormal karyotype containing microdeletions (the mother and / or fetus have microdeletions on this chromosome). Therefore, we first marked the positions of all possible genotypes in cases where the mother and fetus's chromosomes might contain microdeletions on the relative allele count distribution map, and then determined the karyotype of the chromosome based on the distribution of each polymorphic site on the chromosome to be analyzed on the relative allele count distribution map.
[0739] Figure 16 It uses a relative distribution map of allele counts at polymorphic sites to detect microdeletion variations in fetal chromosomes. Figure 16 'a' plots the relative allele counts for all polymorphic sites on a simulated monosomy-disomy chromosome. Figure 16 b is a plot of the relative allele counts for all polymorphic sites on a simulated monosomy-monomer chromosome. The results show that... Figure 16 In karyotype a, the relative counts of almost all polymorphic sites are distributed around the corresponding monosomy-disomy genotype clusters, while they are almost not distributed around the genotype clusters of other karyotypes. Figure 16 In karyotype b, almost all polymorphic sites were relatively counted around the corresponding monomer-monomer genotype clusters, while almost none were distributed around genotype clusters of other karyotypes. Therefore, Figure 16 The karyotype of chromosome a to be analyzed is monosomy-disomy, meaning that the fetus has a normal karyotype without microdeletions; while Figure 16The karyotype of chromosome b to be analyzed is monosomy-monosomy, meaning that one of the chromosomes in the fetus contains a microdeletion variant.
[0740] Example 33: Using fetal DNA concentration and allele counting of the analyte in a pregnant woman's plasma DNA sample Detection of fetal chromosomal microduplications using distribution maps
[0741] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing chromosomal microduplications, wherein chromosomes 1 to 7 are respectively the reference chromosome, chromosomes in which both the mother and fetus are normal (normal disodium-disodium karyotype chromosome), chromosomes in which the mother is normal and the fetus has one chromosome with microduplication (disodium-trisodium karyotype chromosome), chromosomes in which the mother has one chromosome with microduplication and the fetus is normal (trisodium-disodium karyotype chromosome), chromosomes in which both the mother and fetus have one chromosome with microduplication (trisodium-trisodium karyotype chromosome), chromosomes in which the mother has one chromosome with microduplication and both of the fetus's chromosomes have microduplications (trisodium-tetrasodium karyotype chromosome), and chromosomes in which both the mother and fetus have two chromosomes with microduplications (tetrasodium-tetrasodium karyotype chromosome).
[0742] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on reference chromosome 1, following the method described in Example 8. Then, based on the fetal DNA concentration f and the allele counts of each polymorphic site on chromosomes 2 to 7, the karyotypes of chromosomes 2 to 7 were estimated respectively, following the method described in Example 28. To detect whether a fetal chromosome has a microduplication abnormality, we need to determine whether the chromosome has a normal disomy-disomy karyotype (both mother and fetus are disomy) or an abnormal karyotype containing microduplications (the mother and / or fetus have microduplications on this chromosome). Therefore, we first marked the positions of all possible genotypes where the mother and fetus's chromosomes might contain microduplications on the relative distribution map of allele counts, and then determined the karyotype of the chromosome based on the distribution of each polymorphic site on the chromosome to be analyzed on the relative distribution map of allele counts. Since the total number of genotypes containing microduplications on maternal and fetal chromosomes can reach dozens or even hundreds, marking all these genotypes on an allele count relative distribution map would be very difficult for the classification and analysis of the relative allele counts at each polymorphic locus. Therefore, here we only mark the distribution of genotypes in the fetus that are normal and do not contain microduplications. If the relative allele counts at each polymorphic locus on the tested chromosome do not show clustering at the positions relative to the normal fetal genotypes, but clustering is observed at other positions, it means that the chromosome in the sample contains fetal microduplication variants or other types of variants.
[0743] Figure 17 It uses a relative distribution map of allele counts at polymorphic sites to detect microduplications in fetal chromosomes. Figure 17 'a' plots the relative allele counts at all polymorphic sites on a simulated trisomic-disomic chromosome. Figure 17 b is a plot of the relative allele counts for all polymorphic sites on the simulated trisomy-trisomy chromosome. The results show that... Figure 17 In genotype a, the relative counts of almost all polymorphic sites are distributed around the corresponding normal fetal genotype clusters. However... Figure 17 In case b, the relative counts of all polymorphic sites clearly show that they are clustered into several groups, but they do not cluster around the normal fetal genotype clusters. Therefore, Figure 17 In chromosome a to be analyzed, the fetal chromosome is normal and does not contain microduplications; while Figure 17 The chromosome to be analyzed in b, or at least one of the chromosomes in the fetus, contains a microduplication, or the chromosome has other types of variation.
[0744] Example 34: Using fetal DNA concentration and allele counting of the analyte in a pregnant woman's plasma DNA sample Detect wild-type mutants of the sites to be analyzed using the distribution map.
[0745] The method described in Example 21 simulates a pregnant woman's plasma DNA sample containing a specific short sequence site variation, wherein chromosomes 1 and 2 are the reference chromosome and the chromosome containing the specific short sequence site variation, respectively. Specifically, the polymorphic sites on chromosome 1 are selected from different chromosomal regions, while the multiple polymorphic sites on chromosome 2 are selected from the same specific site but are the results of independent amplification using the same and / or different primers. In other words, the simulated polymorphic sites on chromosome 2 represent different independent repeats of a specific site.
[0746] To analyze the sequencing data of the simulated sample, the concentration f of fetal DNA in the sample was first estimated using the allele counts of each polymorphic site on chromosome 1, according to the method described in Example 8. Then, based on the concentration f of fetal DNA in the sample and the allele counts of each specific short sequence site to be tested on chromosome 2, the wild-type of each specific short sequence site on chromosome 2 was estimated according to the method described in Example 29. To detect short genetic variations in the fetus, such as point mutations or short insertion / deletion mutations that cause certain single-gene genetic diseases, all possible genotypes of the fetus and mother need to be considered for each locus (wild-type alleles are marked with the uppercase letter A, and variants are marked with the lowercase letter ac according to the allele count from largest to smallest). This includes variants where all four copies of the mother and fetus are non-wild-type (aa|aa, aa|ab, ab|aa, ab|ab, or ab|ac), variants where the mother has two copies of the non-wild-type variant and the fetus is a wild-type heterozygous variant (aa|Aa or ab|Aa), variants where the mother is a wild-type heterozygous variant and the fetus is normal (Aa|AA), variants where both the mother and fetus are wild-type heterozygous (Aa|Aa or Aa|Ab), variants where the mother is a wild-type heterozygous variant and the fetus is a non-wild-type variant (Aa|aa or Aa|ab), variants where the mother is normal and the fetus is a wild-type heterozygous variant (AA|Aa), and variants where both the mother and fetus are normal wild-type (AA|AA). Each simulated site to be measured was biologically replicated at the sequencing level 20-fold.
[0747] Figure 18 It uses the relative distribution map of allele counts at polymorphic sites to detect fetal variation at the short sequence level. Figure 18 'a' plots the relative allele counts of polymorphic sites in the simulated ab|Aa genotype. Based on the clustered distribution of biologically repetitive polymorphic sites at the sequencing level on the relative count distribution map, the genotype of the polymorphic site is estimated to be ab|Aa, meaning the mother has a double-mutant heterozygous variant and the fetus has a wild-type heterozygous variant. Figure 18 b is a plot of the relative allele counts of polymorphic sites in the simulated Aa|ab genotype. Based on the clustered distribution of biologically repetitive polymorphic sites at the sequencing level on the relative count distribution map, the genotype of the polymorphic site is estimated to be Aa|ab, meaning the mother has a wild-type heterozygous mutant variant and the fetus has a double-mutant heterozygous variant.
[0748] Example 35: Detection of a single genome sample using allele counting and relative distribution mapping of the locus to be analyzed. Genetic variation
[0749] We utilize the relative distribution map of allele counts at various polymorphic sites in the target region to detect aneuploidy at the chromosomal level or deletion / duplication variations at the subchromosomal level in a single genome sample. The main steps are as follows:
[0750] (1) Calculate the relative count of each allele at each target DNA site in the target group;
[0751] (2) For each target DNA site, plot the distribution map A of the relative count of its second largest allele to the relative count of its largest allele, or plot the distribution map B of the relative position of the target DNA site on the chromosome or subchromosome based on the relative count of its largest allele.
[0752] (3) Using the relative distribution map A and / or distribution map B of allele counts at each target DNA site in the target group, estimate the karyotype of the target to be detected in a single genome sample.
[0753] The method described in Example 21 simulates a single genome sample, wherein chromosomes 1 to 5 are respectively a dimer, a deletion (or homozygous microdeletion), a monosomy (or heterozygous microdeletion), a trisomy (or heterozygous microduplication), and a tetrasomy (or homozygous microduplication).
[0754] To detect whether a single genome sample has chromosomal or subchromosomal variations, the following five scenarios need to be considered: (1) both chromosomes are missing (neutropenia) or both chromosomes have microdeletions in the same region (homozygous microdeletion); (2) one chromosome is normal while the other is missing (monotropenia) or the other chromosome has microdeletions (heterozygous microdeletion); (3) both chromosomes are normal; (4) three chromosomes (trisomy) or one chromosome is normal while the other chromosome has microduplications (heterozygous microduplications); (5) four chromosomes (tetrasomy) or both chromosomes have microduplications in the same region (homozygous microduplications).
[0755] Figure 19 The diagram illustrates the use of relative allele counts at polymorphic loci to detect the karyotype of a target chromosome or subchromosome in a single genome sample. For each polymorphic locus in the target region (chromosome or subchromosome region), the relative allele count of its second largest allele is plotted against the relative allele count of its largest allele (relative count map A), or the relative allele count of its largest allele is plotted against the relative position of the locus on a simulated chromosome (relative count position map B). The results show that genotypes of different karyotype chromosomes exhibit different characteristic distributions in relative count map A or relative count position map B. Based on these characteristic distributions, the karyotype (variation type) of the target chromosome or subchromosome can be detected.
[0756] Furthermore, unless otherwise indicated herein or otherwise clearly contradicted by the context, all methods described herein can be performed in any suitable order. The use of any and / or all instances and / or exemplary language provided in certain embodiments herein is intended only to better illustrate the invention and not to limit the scope of the otherwise claimed invention. The language in this specification should not be construed as indicating that any unclaimed element is necessary for practicing the invention.
[0757] The alternative elements or sets of embodiments of the invention disclosed herein should not be construed as limiting. Each member of a group may be mentioned or claimed individually or in any combination with other members of the group or other elements discovered herein. For convenience and / or patentability reasons, one or more members of a group may be included in or removed from the group.
[0758] Although the present invention has been described in full detail with reference to one or more specific embodiments, those skilled in the art will recognize that changes may be made to the embodiments specifically disclosed herein, and such modifications and alterations are within the scope and spirit of the present invention. Therefore, the subject matter of the invention is not limited except as defined in the appended claims. Furthermore, in interpreting the specification and claims, all terms should be interpreted in the broadest possible manner, consistent with the context.
Claims
1. A method for calculating the concentration of the least significant component DNA in a sample, characterized in that... The method includes the following steps: (a1) Set the noise threshold α for the sample; (a2) For each target DNA site, first estimate its genotype using its individual allele counts, then estimate the count of DNA derived from the least significant component (FC) and the total count (TC) based on the estimated genotype; and (a3) Estimate the concentration of minimum component DNA using the minimum component DNA count (FC) and total count (TC) at each target DNA site; Step (a2) includes the following steps: (a2-i) Sort the allele counts of the target DNA site from largest to smallest, and label the three largest allele counts as R1, R2 and R3 respectively; (a2-ii) Estimate the genotype of the target DNA locus by counting the alleles at each locus; and (a2-iii) Estimate the count of DNA derived from the least component (FC) and the total count (TC) based on the estimated genotype of the target DNA site and the allele counts of each target DNA site; and Step (a2-ii) includes the following steps: (a2-ii-1) Count the number of alleles at each target DNA site to determine the number of alleles detected at the target DNA site that are above the noise threshold; if the result is 1, proceed to step (a2-ii-2); if the result is 2, proceed to step (a2-ii-3); if the result is greater than 2, proceed to step (a2-ii-4). (a2-ii-2) Estimate the genotype of the target DNA site to be AA|AA, and then perform the following steps (a2-ii-5); (a2-ii-3) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold being 2 and the two largest allele counts at the target DNA site, and then perform the following steps (a2-ii-5). (a2-ii-4) Estimate the genotype of the target DNA site based on the number of alleles detected above the noise threshold greater than 2 and the largest two allele counts of the target DNA site, and then perform the following steps (a2-ii-5); and (a2-ii-5) outputs the estimated genotype for this target site. The sample is a pregnant woman's plasma sample, and the minimum component DNA is fetal DNA. The method described herein is for non-diagnostic purposes.
2. The method as described in claim 1, characterized in that... Step (a2-ii-3) includes the following steps: (a2-ii-3-1) Determine whether the value of R1 / (R1+R2) is less than 0.5+α. If the result is yes, estimate the genotype of the target DNA site as AB|AB, and then proceed to the following step (a2-ii-3-3); if the result is no, proceed to the following step (a2-ii-3-2). (a2-ii-3-2) Determine if the value of R1 / (R1+R2) is less than 0.
75. If the result is yes, estimate the genotype of the target DNA site as AB|AA, and then proceed to step (a2-ii-3-3); if the result is no, estimate the genotype of the target DNA site as AA|AB, and then proceed to step (a2-ii-3-3); and (a2-ii-3-3) outputs the estimated genotype of the target site.
3. The method as described in claim 1, characterized in that... Step (a2-ii-4) includes the following steps: (a2-ii-4-1) Determine whether R2 / R1 is greater than or equal to 0.5 and / or whether R1 / (R1+R2) is greater than or equal to 1 / 2 and less than or equal to 2 / 3 and / or whether R2 / (R1+R2) is greater than or equal to 1 / 3 and less than or equal to 1 / 2. If the determination result is yes, estimate the genotype of the target DNA site as AB|AC, and then perform the following step (a2-ii-4-3); if the determination result is no, perform the following step (a2-ii-4-2); (a2-ii-4-2) Mark the allele count at this locus as abnormal, and then either estimate the genotype of the target locus as NA and perform the following steps (a2-ii-4-3); or set the number of alleles detected at the target DNA locus above the noise threshold to 2, and then estimate the genotype of the target locus as described in step (a2-ii-3) and perform the following steps (a2-ii-4-3); and (a2-ii-4-3) outputs the estimated genotype of the target site.
4. The method as described in claim 1, characterized in that... Step (a2-iii) includes the following steps: (a2-iii-1) If the genotype estimated for the target site is AA|AA, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1 or R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7). (a2-iii-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, the estimated total count (TC) is R1+R2 or R1+R2+R3, and then perform the following steps (a2-iii-7). (a2-iii-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3. Then perform the following steps (a2-iii-7). (a2-iii-4) If the genotype of the target site is estimated to be AA|AB, the count (FC) derived from the least component DNA is estimated to be twice R2, the total count (TC) is estimated to be R1+R2 or R1+R2+R3, and then the following steps (a2-iii-7) are performed. (a2-iii-5) If the genotype of the target site is estimated to be AB|AC, then estimate the count (FC) of the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3, and then perform the following steps (a2-iii-7). (a2-iii-6) If the estimated genotype at the target site is not one of the genotypes described above, then estimate the count (FC) derived from the least component DNA as NA, estimate the total count (TC) as R1 or R1+R2 or R1+R2+R3, and then perform the following steps (a2-iii-7); and (a2-iii-7) Outputs estimated counts of DNA derived from the least component (FC) and total count (TC).
5. A method for calculating the concentration of the least significant component DNA in a sample, characterized in that... The method includes the following steps: (b1) Set the noise threshold α, the initial concentration estimate f0, and the iteration error accuracy ε for the sample; (b2) For each target DNA site, estimate its genotype using its allele counts and the concentration value f0 of the least component DNA in the sample; (b3) For each target DNA site, estimate the count of DNA derived from the least component (FC) and the total count (TC) based on its estimated genotype. (b4) Estimate the concentration f of the minimum component DNA using the counts (FC) and total counts (TC) of the minimum component DNA at each target site; and (b5) Determine whether the absolute value of f-f0 is less than ε. If the result is no, set f0=f and then execute step (b2). If the result is yes, estimate the minimum concentration of DNA component in the sample as f.
6. The method as described in claim 5, characterized in that... Step (b2) includes the following steps: (b2-i) List all possible genotypes of the target DNA site based on the sample source; (b2-ii) For each possible genotype of the target DNA site, the theoretical count of each allele is calculated using the concentration value f0 of the least component DNA in the sample and the total count (TC) of each allele at the target DNA site. (b2-iii) For each possible genotype of the target DNA site, perform a goodness-of-fit test using the allele counts of each target DNA site and the theoretical allele counts of each allele. and (b2-iv) Analyze the goodness-of-fit test results of the target DNA site to all possible genotypes, and select the genotype that has the best fit to the allele count of each target DNA site as the estimated genotype of the target DNA site.
7. The method as described in claim 5, characterized in that... In step (b3), for each target DNA site, the count of DNA originating from the least component (FC) and the total count (TC) are estimated based on its estimated genotype. The four largest allele counts are labeled R1, R2, R3, and R4 in descending order, and the steps include the following: (b3-1) If the genotype of the target site is estimated to be AA|AA, then the count (FC) derived from the least component DNA is estimated to be NA, and the total count (TC) is estimated to be R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11). (b3-2) If the estimated genotype of the target site is AB|AB, then the estimated count (FC) of the least component DNA is NA, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11). (b3-3) If the genotype of the target site is estimated to be AB|AA, then the count (FC) derived from the least component DNA is estimated to be R1-R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then the following steps (b3-11) are performed. (b3-4) If the genotype of the target site is estimated to be AA|AB, then the count (FC) derived from the least component DNA is estimated to be twice R2, and the total count (TC) is estimated to be R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11). (b3-5) If the genotype estimated for the target site is AB|AC, then estimate the count (FC) derived from the least component DNA as R1-R2+R3 or twice R3 or twice (R1-R2), estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11). (b3-6) If the genotype of the target site is estimated to be AA|BB, then the estimated count (FC) of the least component DNA is R2, and the estimated total count (TC) is R1+R2 or R1+R2+R3 or R1+R2+R3+R4. Then perform the following steps (b3-11). (b3-7) If the genotype of the target site is estimated to be AA|BC, then estimate the count (FC) of the least component DNA as R2+R3 or twice R2 or twice R3, estimate the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11). (b3-8) If the estimated genotype of the target site is AB|CC, determine whether the current estimated value f0 is greater than or equal to 1 / 3. If the result is yes, estimate the count (FC) from the least component DNA as R1 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11); if the result is no, estimate the count (FC) from the least component DNA as R3 and the total count (TC) as R1+R2+R3 or R1+R2+R3+R4, and then perform the following step (b3-11). (b3-9) If the genotype estimated for the target site is AB|CD, then estimate the count (FC) derived from the least component DNA as R3+R4 or twice R3 or twice R4, estimate the total count (TC) as R1+R2+R3+R4, and then perform the following steps (b3-11). (b3-10) If the estimated genotype at the target site is not one of the genotypes described above, then estimate the count (FC) derived from the least component DNA as NA, estimate the total count (TC) as R1 or R1+R2 or R1+R2+R3 or R1+R2+R3+R4, and then perform the following steps (b3-11); and (b3-11) Outputs estimated counts of DNA derived from the least component (FC) and total count (TC).
8. The method as described in claim 1 or claim 5, characterized in that... In step (a3) or step (b4), the concentration of the minimum component DNA is estimated by fitting a regression model.
9. The method as described in claim 1 or claim 5, characterized in that... In step (a3) or step (b4), based on the FC and TC counts, the concentration of the least component DNA in the sample is calculated using linear regression and / or robust linear regression and / or the mean and / or median of FC and TC.
10. A method for detecting genetic variations in a sample, characterized in that... The steps are as follows: (1) Receive the biological sample to be tested and prepare nucleic acid; (2) Enrich or amplify target DNA sites, wherein at least one target DNA site has more than one allele in the sample; (3) The target DNA sites amplified by sequencing; (4) For each target DNA site, count the number of alleles; and (5) Use the goodness-of-fit test of allele counts at target DNA sites and / or the relative distribution map of allele counts to determine the karyotype, genotype or wild-type mutant of the target to be detected in the sample; in: I: In step (5), the goodness-of-fit test of allele counts at the target DNA site is used to determine the karyotype, genotype, or wild-type mutant of the target to be detected in the sample. The determination includes the following steps in sequence: (c1) Each target DNA site is divided into a reference site or a target site according to its location on the chromosome, wherein each reference site forms a reference group and each target site forms a target group. (c2) Calculate the concentration of the least significant component DNA in the sample using allele counts at each target DNA locus in the reference group; and (c3) Using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, a goodness-of-fit test is employed to estimate the karyotype, genotype, or wild-type mutant of the target DNA in the sample; and / or II: In step (5), the relative distribution map of allele counts at the target DNA site is used to determine the karyotype, genotype, or wild-type mutant of the target to be detected in the sample. The determination includes the following steps in sequence: (d1) Each target DNA site is divided into a reference site or a target site according to its location on the chromosome. The reference sites form a reference group, and the target sites form a target group. (d2) Calculate the concentration of the least significant component DNA in the sample by counting alleles at each target DNA site in the reference group; (d3) Using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, the relative distribution map method of allele counts is employed to estimate the karyotype, genotype, or wild-type of the target to be detected in the sample; and In step (c2) or step (d2), the concentration of the least significant component DNA in the sample is calculated using the method described in any one of claims 1-9, and The method described herein is for non-diagnostic purposes.
11. The method as described in claim 10, characterized in that... In step (c3), using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, a goodness-of-fit test is employed to estimate the genotype of the target to be detected in the sample. The estimation includes the following steps in sequence: (c3-a1) For each target DNA site in the target group, list all possible genotypes; (c3-a2) For each target DNA site in the target group, for each possible genotype, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site. (c3-a3) For each target DNA site in the target group, for each possible genotype, the goodness of fit is tested using the allele counts and theoretical counts of each target DNA site. and (c3-a4) For each target DNA site in the target group, the best-fit genotype is selected as the genotype of the target DNA site based on the goodness-of-fit test results of all possible genotypes.
12. The method as described in claim 10, characterized in that... In step (c3), using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, a goodness-of-fit test is employed to estimate the karyotype of the target to be detected in the sample. The estimation includes the following steps in sequence: (c3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested; (c3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group; (c3-b3) For each target DNA locus in the target group, firstly, the goodness-of-fit of all possible genotypes is tested using allele counts; then, for each possible karyotype, the genotype that best fits that karyotype is selected; and (c3-b4) Comprehensively analyze the goodness-of-fit test results of all target DNA sites for each karyotype, and select the karyotype that best fits all target DNA sites as the karyotype of the target chromosome or subchromosome segment to be detected.
13. The method as described in claim 10, characterized in that... In step (c3), using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, a goodness-of-fit test is employed to estimate the wild-type mutant of the target to be detected in the sample. The estimation includes the following steps in sequence: (c3-c1a) For each target DNA site in the target group, list all possible wild-type mutant genotypes; (c3-c2a) For each target DNA site in the target group, for each possible wild-type mutant, calculate the theoretical count of each allele based on the minimum component DNA concentration in the sample and the total count of each allele at that site. (c3-c3a) For each target DNA site in the target group, for each possible wild-type mutant, the goodness of fit is tested by using the allele counts of each target DNA site and its theoretical count. and (c3-c4a) Comprehensive analysis of all target DNA sites in the target group, and selection of the wild-type mutant genotype that has the best fit to all target sites as the wild-type mutant genotype of the target to be tested.
14. The method as described in claim 10, characterized in that... In step (c3), using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, a goodness-of-fit test is employed to estimate the wild-type mutant of the target to be detected in the sample. The estimation includes the following steps in sequence: (c3-c1b) For each target DNA site in the target group, its genotype is estimated by using a goodness-of-fit test based on the allele counts and the concentration of the least significant component DNA in the sample. and (c3-c2b) Based on the genotype of each target DNA site and the sequence of each allele in the target group, determine the wild-type mutant of each allele of the target to be tested in each component of the sample.
15. The method as described in claim 10, characterized in that... The goodness-of-fit test method described in step (c3) is performed using the chi-square test, G test, Fisher exact test, binomial distribution test, or variations thereof.
16. The method as described in claim 10, characterized in that... The goodness-fit test method described in step (c3) uses the calculated G-test value, AIC value, corrected G-test value, corrected AIC value, G-test value or variant of AIC value, or combination thereof to perform the goodness-fit test.
17. The method as described in claim 10, characterized in that... The goodness-of-fit test method in step (c3) is performed using the method described in claim 6.
18. The method as described in claim 10, characterized in that... In step (d3), the genotype of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele counts. The estimation includes the following steps in sequence: (d3-a1) For each target DNA site in the target group, list all possible genotypes; (d3-a2) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum theoretical relative count of alleles to plot against the maximum theoretical relative count of alleles to mark the theoretical position of the genotype. (d3-a3) For each target DNA locus in the target group, first calculate the relative counts of each allele, then select at least one non-maximum allele relative count and plot it against the maximum allele relative count to mark the actual position of the target DNA locus on the allele relative count map; and (d3-a4) Based on the theoretical and actual positional distribution of each target DNA locus in the relative allele count map of the target group, the genotype of the target to be tested is inferred.
19. The method as described in claim 10, characterized in that... In step (d3), the karyotype of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele counts. The estimation includes the following steps in sequence: (d3-b1) Analyze the sample to be tested and list all possible karyotypes of the target chromosome or subchromosome segment to be tested; (d3-b2) For each possible karyotype, list all possible genotypes for each target DNA site in the target group; (d3-b3) For each possible genotype of the target DNA site in the target group, first calculate the theoretical relative count of each allele based on the concentration of the least fractional DNA in the sample, and then select at least one non-maximum theoretical relative count of the allele to plot against the maximum theoretical relative count of the allele to mark the theoretical position of the genotype. (d3-b4) For each target DNA locus in the target group, first calculate the relative count of each allele, then select at least one non-maximum allele relative count and plot it against the maximum allele relative count to mark the actual position of the target DNA locus on the allele relative count map; and (d3-b5) Based on the theoretical and actual location distribution of each target DNA site in each karyotype in the relative allele count diagram, the karyotype of the target to be tested is inferred.
20. The method as described in claim 10, characterized in that... In step (d3), the wild-type mutant of the target DNA in the sample is estimated by using the allele counts of each target DNA locus in the target group and the concentration of the least significant component DNA in the sample, employing the relative distribution map method of allele counts. The estimation includes the following steps in sequence: (d3-c1) For each target DNA site in the target group, list its wild-type sequence and all possible wild-type mutant genotypes; (d3-c2) For each possible wild-type mutant genotype, calculate the theoretical relative counts of its wild-type alleles and each of the other non-wild-type alleles, and select at least one non-wild-type allele relative count theoretical value to plot the relative count theoretical value of the wild-type allele to mark the theoretical position of its wild-type mutant genotype. (d3-c3) For each target DNA site in the target group, calculate the relative count values of its wild-type alleles and other non-wild-type alleles, and select at least one non-wild-type allele relative count to plot the wild-type allele relative count to mark the actual position of the target DNA site on the allele relative count map. (d3-c4) Based on the theoretical and actual positional distribution of all target DNA sites in the relative allele count map, the wild-type mutant can be inferred.
21. A system for detecting genetic variation in a sample, comprising means and / or computer program products and / or modules for performing any step of the method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Methods for non-invasive prenatal ploidy calling
US20130274116A1
Methods for simultaneous amplification of target loci
US20190256919A1
Diagnosing fetal chromosomal aneuploidy using genomic sequencing
WO2009013496A1