Method for determining fetal nucleic acid concentration and method for fetal genotyping

By analyzing the sequencing data and reference genomic sequences of pregnant women, determining the fetal free nucleic acid concentration and inferring the genotype, the problems of high invasiveness, high cost and low detection specificity of existing non-invasive prenatal diagnostic methods are solved, and accurate fetal genotype analysis is achieved at low sequencing depth and low fetal concentration.

CN114945685BActive Publication Date: 2025-07-04SHENZHEN HUADA GENE INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080093338.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-17
Publication Date
2025-07-04
Estimated Expiration
2040-01-17

AI Technical Summary

Technical Problem

The existing non-invasive prenatal diagnostic methods have problems such as high invasiveness, high cost, low detection specificity and high requirements for sequencing depth, especially in the case of low sequencing depth and low fetal concentration, which is difficult to accurately determine the fetal genotype.

Method used

By obtaining sequencing data and reference genomic sequences of pregnant women's nucleic acid samples, selecting predetermined regions, analyzing mutation information, determining fetal free nucleic acid concentration, and inferring the fetal genotype at the pre-local site based on Bayesian model, it is suitable for lower-deep sequencing and low-fetal concentration samples.

Benefits of technology

Accurate fetal genotype analysis was achieved at lower sequencing depths and fetal concentrations, expanding the detection range to the whole genome of single-base mutations, reducing detection costs and experimental requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005351319560000021
    Figure FDA0005351319560000021
  • Figure FDA0005351319560000022
    Figure FDA0005351319560000022
  • Figure FDA0005351319560000031
    Figure FDA0005351319560000031
Patent Text Reader

Abstract

The present invention provides a method for determining the concentration of fetal nucleic acid and a method for fetal genotyping. According to an embodiment of the present invention, the method for determining the concentration of fetal free nucleic acid includes: (1) obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genomic sequence, wherein the first nucleic acid sample of the pregnant woman contains fetal free nucleic acid, and the sequencing data is composed of a plurality of sequencing reads; (2) selecting a predetermined region on the reference genomic sequence, and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and (3) determining the concentration of fetal free nucleic acid corresponding to the predetermined region based on the mutation information within the predetermined region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, particularly non-invasive prenatal genetic testing, and specifically to a method for determining the nucleic acid concentration of a fetus and a method for fetal genotyping. Background Art

[0002] Prenatal diagnostic methods are generally divided into two categories according to different sampling materials and examination means, namely invasive methods and non-invasive methods. The former mainly includes amniocentesis, chorionic villus sampling, cordocentesis, etc.; the latter includes ultrasonic examination, determination of maternal peripheral serum markers, fetal cell detection, etc. By using invasive methods such as chorionic villus sampling (CVS) or amniocentesis, cells isolated from the fetus can be obtained, and these cells can be used for routine prenatal diagnosis. Although the accuracy of diagnosing fetal aneuploidy by this method is relatively high, however, these conventional methods are invasive and pose certain risks to the mother and the fetus.

[0003] Currently, the method for inferring the fetal genome using maternal plasma cfDNA and parental WGS data requires the simultaneous collection of blood samples from both the father and the mother, and the performance of two WGS sequencings on both maternal plasma cfDNA and parental blood cells. The sampling difficulty is relatively high, and both the sequencing cost and the calculation cost are relatively high. At the same time, if this method is used to complete the detection of fetal de novo mutations, it is necessary to further increase the cost to achieve extreme sequencing of maternal plasma, and the specificity of the detection is very low. The method for inferring the fetal genome using maternal plasma cfDNA data and single-cell sequencing data, although only requires the maternal whole blood for sampling, when inferring the fetal genotype in the whole genome range, it requires the simultaneous performance of two methods, namely maternal plasma cfDNA sequencing and lymphocyte single-cell sequencing, which has too high requirements for experiments, and the cost requirements for inferring the haplotypes of the father and the mother are also very high. Moreover, this method cannot perform fetal de novo mutation detection. In addition, the method for inferring the fetal genome only using ultra-high-depth sequencing data of maternal plasma cfDNA, since only a fixed threshold is used for each locus to judge the fetal genotype, thus requires an extremely high sequencing depth (above 300x), and when the sequencing depth is relatively low (within 100x) or the fetal concentration is relatively low, the fetal genotype cannot be accurately inferred.

[0004] However, the current non-invasive diagnostic methods for fetuses still need to be improved. Summary of the Invention

[0005] The present invention aims to at least solve one of the technical problems existing in the prior art. To this end, an object of the present invention is to propose a method capable of accurately and efficiently determining chromosomal aneuploidy.

[0006] In the first aspect of the present invention, the present invention provides a method for determining the concentration of fetal free nucleic acid. According to an embodiment of the present invention, the method includes: (1) obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genome sequence, wherein the first nucleic acid sample of the pregnant woman contains fetal free nucleic acid, and the sequencing data is composed of a plurality of sequencing reads; (2) selecting a predetermined region on the reference genome sequence, and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and (3) determining the concentration of fetal free nucleic acid corresponding to the predetermined region based on the mutation information within the predetermined region. According to an embodiment of the present invention, by determining the distribution of mutations in the predetermined region, the concentration of fetal nucleic acid corresponding to this region can be effectively determined, so that it can be effectively applied to in-depth analysis of fetal free nucleic acid. In addition, according to an embodiment of the present invention, it is found that this method can be applied to sequencing results with a relatively low depth, for example, the sequencing depth does not exceed 100X, such as not exceeding 60X, and can also be applied to samples with a relatively low fetal concentration, for example, the fetal concentration does not exceed 10%, so that it can be effectively applied to early prenatal diagnosis of pregnant women, such as detection at 15 weeks of gestation.

[0007] In the second aspect of the present invention, the present invention provides a method for determining the fetal genotype at a predetermined locus. According to an embodiment of the present invention, the method includes: (a) determining the concentration of fetal free nucleic acid in the predetermined region according to the method described above, wherein the predetermined region contains the predetermined locus; (b) respectively determining the number of sequencing reads A j , where j is A, T, G, or C, that support base A, base T, base G, or base C for the predetermined locus; (c) constructing a genotype set {M 1i M 2i F 1i F 2i} for the predetermined locus, where i represents the genotype number, M 1i represents the base type on the first chromosome of the mother for the i-th genotype, M 2i represents the base type on the second chromosome of the mother for the i-th genotype, F 1i represents the base type on the first chromosome of the fetus for the i-th genotype, F 2i represents the base type on the second chromosome of the fetus for the i-th genotype, wherein the first chromosome of the mother and the second chromosome of the mother belong to a pair of homologous chromosomes, the first chromosome of the fetus and the second chromosome of the fetus belong to a pair of homologous chromosomes, and M 1i , M 2i , F 1i and F 2i are each independently base A, base T, base G, or base C; (d) for the genotype set {M1i M 2i F 1i F 2i For each of {}, based on the fetal nucleic acid concentration, determine the occurrence probability P of each base j , where j is A, T, G, or C; (e) for each of the genotype sets {M 1i M 2i F 1i F 2i}, based on the occurrence probability P of each base j and the number of sequencing reads A j , determine the cumulative probability P(M 1i M 2i F 1i F 2i ); (f) based on the cumulative probability P(M 1i M 2i F 1i F 2i ) of each genotype, determine the combination of the pregnant woman's genotype and the fetal genotype at the predetermined locus, and then obtain the fetal genotype at the predetermined locus. By this method, it is possible to effectively determine the combination of the fetal and pregnant woman's genotypes with the highest probability based on the fetal nucleic acid concentration within the predetermined region and the number of sequencing reads of the sequencing. In other words, it is possible to determine the haplotype at a specific locus within this region.

[0008] In the third aspect of the present invention, the present invention proposes a device for determining the concentration of fetal free nucleic acid. In an embodiment of the present invention, the device includes: a reading module for obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genomic sequence, the first nucleic acid sample of the pregnant woman containing fetal free nucleic acid, and the sequencing data consisting of a plurality of sequencing reads; a mutation determination module for selecting a predetermined region on the reference genomic sequence and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and a fetal free nucleic acid concentration determination module for determining the concentration of fetal free nucleic acid corresponding to the predetermined region based on the mutation information within the predetermined region. According to an embodiment of the present invention, the device can effectively implement the method for determining the concentration of fetal free nucleic acid described above. The advantages and features described for the above method are all applicable to this device and will not be elaborated further.

[0009] In the fourth aspect of the present invention, the present invention proposes a device for determining the fetal genotype at a predetermined locus, characterized in that it includes: the device for determining the concentration of fetal free nucleic acid described above for determining the concentration of fetal free nucleic acid in the predetermined region, the predetermined region including the predetermined locus; a sequencing read number determination module for respectively determining the number of sequencing reads A that support base A, base T, base G, or base C for the predetermined locusj , where j is A, T, G or C; a genotype set construction module for constructing a genotype set {M 1i M 2i F 1i F 2i} for the predetermined locus, where i represents the genotype number, M 1i represents the base type on the first chromosome of the mother for the i-th genotype, M 2i represents the base type on the second chromosome of the mother for the i-th genotype, F 1i represents the base type on the first chromosome of the fetus for the i-th genotype, F 2i represents the base type on the second chromosome of the fetus for the i-th genotype, M 1i 、M 2i 、F 1i and F 2i are each independently base A, base T, base G or base C, where the first chromosome of the mother and the second chromosome of the mother belong to a pair of homologous chromosomes, and the first chromosome of the fetus and the second chromosome of the fetus belong to a pair of homologous chromosomes; an occurrence probability determination module for determining the occurrence probability P 1i M 2i F 1i F 2i} of each of the bases based on the nucleic acid concentration of the fetus, where j is A, T, G or C; a cumulative probability determination module for determining the cumulative probability P(M j for each of the genotype sets {M 1i M 2i F 1i F 2i} based on the occurrence probability P j of each of the bases and the number A j of the sequencing reads, and determining the cumulative probability P(M 1i M 2i F 1i F 2i ) of the genotype; a genotype combination determination module for determining the genotype combination of the pregnant woman and the fetus at the predetermined locus based on the cumulative probability P(M 1i M 2i F 1i F 2i ) of each of the genotypes, and further obtaining the fetal genotype at the predetermined locus.

[0010] In the fifth aspect of the present invention, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, the steps of the method described above are implemented.

[0011] The advantages of the present invention are as follows:

[0012] (1) Currently, most fetal genome analyses based on whole-genome sequencing data of cell-free DNA in pregnant women's plasma focus on the detection of chromosomal structural variations, large-fragment copy number mutations, and paternally specific variations. The present invention can expand the region of fetal genome analysis to the whole genome, reduce the detection accuracy to single-base mutations, and expand the application scope of fetal genetic disease detection.

[0013] (2) Most of the existing fetal whole-genome analysis methods currently require the assistance of whole-genome sequencing (WGS) data of the father's and mother's blood cells. The present invention provides multiple analysis schemes that rely on parental WGS data and do not rely on parental WGS data, providing diverse methodological support for different types of samples and data.

[0014] (3) Currently, fetal genome analysis methods that only use cell-free DNA (cfDNA) in pregnant women's plasma have extremely high requirements for sequencing depth (above 300x). The present invention still achieves high accuracy when the sequencing depth is relatively low (about 60x - 100x).

[0015] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, wherein:

[0017] Figure 1 is a schematic flowchart of a method for determining the nucleic acid concentration of a fetus according to an embodiment of the present invention;

[0018] Figure 2 is a schematic flowchart of a method for determining the fetal genotype at a predetermined locus according to an embodiment of the present invention;

[0019] Figure 3 is a schematic flowchart of a method for determining the fetal genotype at a predetermined locus according to another embodiment of the present invention;

[0020] Figure 4 is a schematic flowchart of a method for determining the fetal genotype at a predetermined locus according to an embodiment of the present invention;

[0021] Figure 5 is a schematic flowchart of a method for determining the fetal genotype at a predetermined locus according to another embodiment of the present invention;

[0022] Figure 6Schematic structural diagram of a device for determining the concentration of fetal cell-free nucleic acid according to an embodiment of the present invention;

[0023] Figure 7 Schematic structural diagram of a device for determining the fetal genotype at a predetermined locus according to an embodiment of the present invention;

[0024] Figure 8 Schematic structural diagram of a device for determining the fetal genotype at a predetermined locus according to another embodiment of the present invention;

[0025] Figure 9 Schematic structural diagram of a device for determining the fetal genotype at a predetermined locus according to another embodiment of the present invention;

[0026] Figure 10 Schematic structural diagram of a device for determining the fetal genotype at a predetermined locus according to an embodiment of the present invention;

[0027] Figure 11 Schematic structural diagram of a device for determining the fetal genotype at a predetermined locus according to another embodiment of the present invention;

[0028] Figure 12 Schematic diagram showing the relationship between the frequency of sub-bases and the site proportion distribution;

[0029] Figure 13 Schematic diagram showing Schemes 1-4 according to an embodiment of the present invention; and

[0030] Figure 14 Shows the comparison of the accuracy after correcting the mother heterozygous and fetus heterozygous loci in Pedigree 1 by setting different numbers of flanking loci in Scheme 4 with Scheme 3 in Embodiment 1 of the present invention. Detailed implementation manners

[0031] Embodiments of the present invention will be described in detail below. The embodiments described below are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. Embodiments of the present invention will be described in detail below. The embodiments described below are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. It should be noted that the present application can be used in many general or special computing device environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or equipment, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0032] Method for determining the concentration of fetal free nucleic acid

[0033] In a first aspect of the present invention, the present invention provides a method for determining the concentration of fetal free nucleic acid. According to an embodiment of the present invention, the method includes:

[0034] S100 Obtain sequencing data and a reference genome sequence

[0035] According to an embodiment of the present invention, first, sequencing data and a reference genome sequence from a pregnant woman containing fetal free nucleic acid are obtained.

[0036] According to an embodiment of the present invention, the pregnant woman samples that can be used include but are not limited to the peripheral blood of the pregnant woman. Dennis Lo et al. found cell-free fetal DNA in maternal plasma and serum, providing a new idea for non-invasive prenatal diagnosis (NIPT). By using the peripheral blood of the pregnant woman, it will not cause trauma to the pregnant woman and avoid the risk of miscarriage caused by sampling. According to an embodiment of the present invention, after obtaining the pregnant woman samples, such as the peripheral blood of the pregnant woman, these samples can be subjected to nucleic acid sequencing to obtain the nucleic acid sequencing data of the pregnant woman samples. Generally, the nucleic acid sequencing data is composed of multiple or a large number of sequencing reads. According to an embodiment of the present invention, the method for sequencing the nucleic acid molecules of the pregnant woman samples is not particularly limited. Specifically, any sequencing method known to those skilled in the art can be used, for example, including but not limited to sequencing the nucleic acid molecules of the pregnant woman samples by paired-end sequencing, single-end sequencing, or single-molecule sequencing.

[0037] In addition, according to the embodiments of the present invention, there is no special requirement for the sequencing depth of sequencing the pregnant woman's sample. According to the embodiments of the present invention, the inventors of the present invention found that by analyzing the distribution of mutation types using the method of the present invention, even for the sequencing results with low sequencing depth, high accuracy can be achieved. For example, for the sequencing data with a sequencing depth lower than 100X, specifically for the sequencing data of 60X to 100X, high accuracy can also be achieved. Those skilled in the art can understand that after obtaining the nucleic acid sequencing data, according to the quality control standard, the sequencing data composed of a large number of sequencing reads obtained can be filtered and screened to remove the sequencing reads with sequencing quality problems, thereby improving the accuracy of subsequent data analysis.

[0038] According to the embodiments of the present invention, the reference genome sequence that can be adopted is not particularly limited and can be a part of any known human genome sequence. For example, it can be obtained through the sequence published by NCBI (https: / / www.ncbi.nlm.nih.gov / grc / human). Of course, those skilled in the art can understand that the reference genome sequence adopted here does not need to be the full length of the human genome sequence, and even does not need to cover the entire predetermined region, as long as it covers a sufficient number of mutation sites that may appear in the predetermined region. In order to determine the fetal nucleic acid concentration of the predetermined sequence, according to the embodiments of the present invention, it is necessary to determine the base type distribution of approximately 50 to 200 mutation sites. Therefore, as long as the obtained reference genome sequence can cover these numbers of mutation sites.

[0039] S200 Select a predetermined region and determine the mutation information within the predetermined region

[0040] After obtaining the sequencing data and the reference genome sequence, a predetermined region can be first selected on the reference genome sequence for subsequent analysis. The existing methods for determining the fetal free nucleic acid concentration usually target the total content of all free nucleic acids, and the obtained fetal free nucleic acid concentration cannot be effectively applied to genotyping. After in-depth analysis and combining with practical experience, the inventors of the present invention proposed to divide the reference genome into multiple regions and analyze the fetal free nucleic acid concentration of these regions respectively. Thus, subsequently, according to the obtained data of the free nucleic acid concentration and combined with the sequencing alignment results, the fetal genotypes within these regions can be typed.

[0041] According to an embodiment of the present invention, the length of the predetermined region is not less than 50 kb. According to an embodiment of the present invention, the length of the predetermined region is 50 - 200 kb. According to an embodiment of the present invention, the length of the predetermined region is 100 kb. The inventors found that the number of mutations within the predetermined region of this length can meet the requirement for measuring the concentration of fetal cell-free nucleic acids (herein, sometimes the "concentration of fetal cell-free nucleic acids" is abbreviated as "fetal concentration", which can be used interchangeably). Thus, the efficiency of determining the concentration of fetal nucleic acids can be improved. According to an embodiment of the present invention, by analyzing the base type distribution of mutation sites, the peaks or valleys of the base type distribution can be determined. Usually, these peaks and valleys are determined by the fetal concentration. Therefore, the concentration of fetal cell-free nucleic acids in this region can be determined by analyzing the distribution of these base types. For this purpose, when determining the concentration of fetal cell-free nucleic acids, a certain number of mutation sites are required as the analysis basis. According to the research of the inventors, 50 mutation sites can effectively analyze the concentration of fetal cell-free nucleic acids in this region. Considering the mutation probability, a predetermined nucleic acid region with a length of not less than 50 kb is adopted. Additionally, according to an embodiment of the present invention, considering that the obtained fetal cell-free nucleic acid concentration will be used for fetal genotyping later, if the length of the predetermined region is too large, the true fetal nucleic acid ratio of each mutation site cannot be represented. Thus, according to an embodiment of the present invention, the length of the predetermined region is 50 - 200 kb. According to an embodiment of the present invention, the length of the predetermined region is 100 kb. The inventors found that by adopting the predetermined region of this length, not only can the concentration of fetal nucleic acids in this region be effectively determined, but also the fetal nucleic acid concentration can truly reflect the fetal nucleic acid ratio of the mutation sites, and thus can be effectively applied to fetal genotyping within this region.

[0042] The term "concentration of fetal cell-free nucleic acids" used herein is sometimes also referred to as "fetal concentration" or "fetal nucleic acid concentration", and these terms can be used interchangeably. Its meaning is that for a specified genomic region, among the cell-free nucleic acids in the peripheral blood of a pregnant woman, the proportion of fetal cell-free nucleic acids corresponding to this region in the total amount of cell-free nucleic acids of the pregnant woman and the fetus corresponding to this region. For example, for a certain specified genomic region, there are 100 cell-free nucleic acid molecules corresponding to this region. These 100 cell-free nucleic acid molecules include both cell-free nucleic acid molecules from the pregnant woman and cell-free nucleic acid molecules from the fetus. If the number of cell-free nucleic acid molecules from the fetus is 30, then for this specified genomic region, its fetal cell-free nucleic acid concentration is 30%.

[0043] After obtaining the sequencing data and the reference genome sequence, mutation information within a predetermined region can be determined based on the sequencing data of the pregnant woman. According to an embodiment of the present invention, the first nucleic acid sample of the pregnant woman is derived from the peripheral blood of the pregnant woman, and the mutation information includes at least one of SNP, Indel, and SV, preferably SNP. According to an embodiment of the present invention, this information can be easily obtained through sequence alignment, and those skilled in the art can easily obtain relevant locus information, such as obtaining loci known to have corresponding mutations through a database, so that the corresponding mutation information can be determined at a fixed point to improve the efficiency of determining the fetal nucleic acid concentration. Additionally, according to an embodiment of the present invention, by using the peripheral blood of the pregnant woman, no trauma is caused to the pregnant woman, avoiding the risk of miscarriage caused by sampling.

[0044] For a specific mutation locus, such as SNP, the base types it can select are usually A, T, G, or C. Each different base type can have a different proportion. The inventors of the present invention propose to divide them into a main base type and a secondary base type according to the occurrence proportion of each base type. Among them, the meaning of "main base type" refers to the base type with the highest occurrence proportion for a specific locus, such as an SNP locus. The meaning of "secondary base type" refers to the base type with the second highest occurrence proportion for a specific locus, such as an SNP locus. Obviously, when the pregnant woman and the fetus have different base types at this locus, the proportions of the main base type and the secondary base type are both affected by the concentration of fetal free nucleic acid, so it can be used to infer the concentration of fetal free nucleic acid inversely.

[0045] According to an embodiment of the present invention, the occurrence proportion of each base type can be obtained by determining the number of sequencing reads supporting each base type, or the occurrence frequency of each base type at a specific mutation locus. The mutation information that needs to be determined includes the mutation locus and the main base type or secondary base type among the base types at this mutation locus, as well as the frequency (occurrence proportion) of the corresponding main base type or secondary base type. Additionally, after determining the specific base type, i.e., the main base type or secondary base type, at each mutation locus, the number of mutation loci corresponding to each frequency or the proportion of the corresponding mutation loci among all mutation loci in this predetermined region can be further determined. According to an embodiment of the present invention, the main base type or secondary base type and the corresponding frequency are determined through the following steps:

[0046] S210: Align the sequencing data with the reference genome sequence to determine the mutation loci within the predetermined region and the base types of each of the mutation loci.

[0047] Those skilled in the art can align the sequencing data with the reference genome by conventional means. For example, known software such as SOAP can be used to align the sequencing data with the reference genome sequence. Through alignment, the mutation sites in the predetermined region and the possible base types of each mutation site can be determined.

[0048] S220: Determine the number of sequencing reads corresponding to each of the base types.

[0049] In determining the base types of the mutation sites, by counting the number of sequencing reads supporting a specific base type, the number of sequencing reads corresponding to each base type is determined.

[0050] S230: Determine the total number of sequencing reads corresponding to the mutation site.

[0051] As described above, in determining the base types of the mutation sites, by counting the number of sequencing reads supporting a specific mutation site, the total number of sequencing reads corresponding to the mutation site can be determined.

[0052] S240: Determine the proportion of sequencing reads of the base type.

[0053] For a specific mutation site, after determining the sequencing reads corresponding to each base type and the number of sequencing reads corresponding to the mutation site, the proportion of sequencing reads of each base type can be determined. For example, for a certain mutation site, there are 100 sequencing reads supporting the mutation site, and for base A, there are 50 sequencing reads supporting it. Then the proportion of sequencing reads of base A is 50 / 100 = 50%.

[0054] S250: Determine the specific base type of the mutation site

[0055] According to an embodiment of the present invention, the specific base type refers to at least one of the main base type and the secondary base type. For a certain mutation site, after determining the proportion of sequencing reads of each base type, the base type with the highest proportion of sequencing reads can be selected as the main base type, and the base type with the second highest proportion of sequencing reads can be selected as the secondary base type.

[0056] S260: Determine the frequency of the main base type or the secondary base type at the corresponding mutation site.

[0057] After determining the specific base type, that is, the main base type or the secondary base type, the proportion of sequencing reads corresponding to the specific base type is used as the frequency of the specific base type at the corresponding mutation site.

[0058] According to an embodiment of the present invention, for the convenience of calculation, it may be assumed that there are only a main base type and a secondary base type. Thus, the sum of the frequencies of the main base type and the secondary base type is 100%.

[0059] Specifically, according to an embodiment of the present invention, determining the mutation information within the predetermined region includes: (2-1) aligning the sequencing data with the reference genome sequence to determine the mutation sites within the predetermined region and the base types of each of the mutation sites; (2-2) for each of the mutation sites, determining a specific base type based on the number of sequencing reads corresponding to each base type and determining the frequency of the specific base type to obtain a plurality of the frequencies; (2-3) based on each of the plurality of frequencies, determining a mutation site ratio corresponding to the frequency, where the mutation site ratio represents the number of mutation sites corresponding to the frequency. It should be noted that the mutation site ratio used in step (2-3) can be directly represented by the number of mutation sites corresponding to the frequency, or can be represented by any mathematical operation result based on the number of mutation sites and the total number of mutation sites, as long as it can represent the number of mutation sites corresponding to the same frequency, thereby reflecting the site distribution of the specific base type in this region.

[0060] According to an embodiment of the present invention, after determining the main base type or the secondary base type of each mutation site, as well as the corresponding occurrence frequencies, and obtaining the number or ratio of mutation sites at each same frequency, the concentration of fetal free nucleic acids in the predetermined region can be further determined through the following:

[0061] (3-1) Determine the distribution of the mutation site ratio relative to the frequency;

[0062] (3-2) Based on the distribution, determine the concentration of the fetal free nucleic acids corresponding to the predetermined region.

[0063] More specifically, in step (3-1), within a predetermined frequency range, make a two-dimensional plot of the mutation site ratio relative to the plurality of frequencies; and in step (3-2), based on the frequencies corresponding to the peaks and valleys of the two-dimensional plot, determine the concentration of the fetal free nucleic acids.

[0064] For ease of understanding, taking the minor allele type as an example below, in order to calculate the regional fetal concentration using pregnant women's plasma cfDNA data, the whole genome is divided into different sliding windows, and the window size can be a 100 kb window. Within each sliding window, the minor allele frequency (MAF) of each mutation site is calculated respectively, and the distribution of each minor allele frequency is calculated, that is, the number of mutation sites corresponding to the same minor allele frequency. The inventor has concluded through research that for the distribution of minor allele frequencies, there can theoretically be four peaks, corresponding to four types of sites: mother homozygous and fetus homozygous (mother homozygous and fetus homozygous), mother homozygous and fetus heterozygous (mother homozygous and fetus heterozygous), mother heterozygous and fetus homozygous (mother heterozygous and fetus homozygous), and mother heterozygous and fetus heterozygous (mother heterozygous and fetus heterozygous). Since the MAF distribution is composed of sequencing reads from both the pregnant woman and the fetus, for each site, when the fetal concentration is f, the MAF corresponding to the peak of the mother homozygous and fetus heterozygous site is f / 2. Therefore, for MAF ∈ [0, 0.25], its distribution can be simulated as a mixture model of two normal distributions, and it can be inferred that the MAF corresponding to the peak of the fetal-derived reads distribution is f / 2, that is, the fetal concentration within this window is obtained. Refer to Figure 12 . Of course, those skilled in the art can understand that similar analysis can also be adopted to analyze the fetal concentration using the major allele type.

[0065] Thus, according to the embodiments of the present invention, the predetermined frequency range is selected from 0 to 0.5 or 0.5 to 1.

[0066] According to the embodiments of the present invention, the specific base type is the minor allele type, and the predetermined frequency range is 0 to 0.25 or its subset or 0.25 to 0.5 or its subset. Among them, when the predetermined frequency range is 0 to 0.25 or its subset, along the direction from small to large, the base frequency a corresponding to the second peak is selected, and the fetal free nucleic acid concentration is 2a; when the predetermined frequency range is 0.25 to 0.5 or its subset, along the direction from small to large, the base frequency b corresponding to the first peak is selected, and the fetal free nucleic acid concentration is 1 - 2b.

[0067] According to the embodiments of the present invention, the specific base type is the major allele type, and the predetermined frequency range is 0.5 to 0.75 or its subset or 0.75 to 1 or its subset. Among them, when the predetermined frequency range is 0.5 to 0.75 or its subset, along the direction from small to large, the base frequency c corresponding to the second trough is selected, and the fetal free nucleic acid concentration is 2c; when the predetermined frequency range is 0.75 to 1 or its subset, along the direction from small to large, the base frequency d corresponding to the first trough is selected, and the fetal free nucleic acid concentration is 1 - 2d.

[0068] Thus, by the method according to the embodiments of the present invention, it is possible to effectively determine the concentration of fetal free nucleic acids in a specific region based on the base type distribution of mutation sites, and to truly reflect the ratio of fetal free nucleic acids to maternal free nucleic acids in this region, so that it can be more effectively used to guide fetal gene analysis.

[0069] In view of this, in the second aspect of the present invention, the present invention proposes a method for determining the fetal genotype at a predetermined site.

[0070] Reference Figure 2 , the embodiments of the present invention, the method includes:

[0071] S1000: In this step, according to the method described above, determine the concentration of fetal free nucleic acids in the predetermined region, and the predetermined region contains the predetermined site. How to determine the concentration of fetal free nucleic acids in the predetermined region has been described in detail above and will not be repeated here. Regarding the predetermined site, it can be any site where mutations may exist that are of interest to the analyst, such as sites that may cause diseases or can be used to distinguish the identity of the fetus, such as for paternity testing.

[0072] S2000: For the predetermined site, respectively determine the number of sequencing reads A that support base A, base T, base G, or base C j , where j is A, T, G, or C.

[0073] S3000: For the predetermined site, construct a genotype set {M 1i M 2i F 1i F 2i}, where i represents the genotype number, M 1i represents the base type on the mother's first chromosome for the i-th genotype, M 2i represents the base type on the mother's second chromosome for the i-th genotype, F 1i represents the base type on the fetus's first chromosome for the i-th genotype, F 2i represents the base type on the fetus's second chromosome for the i-th genotype, M 1i , M 2i , F 1i and F 2i are independently base A, base T, base G, or base C respectively, where the mother's first chromosome and the mother's second chromosome belong to a pair of homologous chromosomes, and the fetus's first chromosome and the fetus's second chromosome belong to a pair of homologous chromosomes.

[0074] S4000: For the genotype set {M 1i M 2i F1i F 2i For each of {}, based on the fetal nucleic acid concentration, determine the occurrence probability P of each base j , where j is A, T, G, or C.

[0075] S5000: For each of the genotype sets {M 1i M 2i F 1i F 2i}, based on the occurrence probability P of each base j and the number A of sequencing reads j , determine the cumulative probability P(M 1i M 2i F 1i F 2i ) of the genotype.

[0076] S6000: Based on the cumulative probability P(M 1i M 2i F 1i F 2i ) of each genotype, determine the combination of the pregnant woman's genotype and the fetal genotype at the predetermined locus, and then obtain the fetal genotype at the predetermined locus. Through this method, it is possible to effectively determine the combination of the fetal genotype and the pregnant woman's genotype based on the fetal nucleic acid concentration in the predetermined region and the number of sequencing reads of the sequencing. In other words, it is possible to determine the haplotype at a specific locus in this region, that is, for a specific locus, it is possible to determine the base types carried on each of the four chromosomes (two homologous chromosomes of the mother and the corresponding two homologous chromosomes of the fetus).

[0077] For ease of understanding, the principle of the above analysis method is analyzed below:

[0078] According to an embodiment of the present invention, after obtaining the fetal concentration, a Bayesian model is established based on this to infer the fetal genotype. For each locus, when the read depths of different alleles in the pregnant woman's plasma and the fetal concentration are known, the probability calculation of the combination i of the genotype of the pregnant woman and the fetus at the specific locus is:

[0079]

[0080] where P(A i ) is the prior probability of the occurrence of combination i in the n combinations of the genotype of the pregnant woman and the fetus obtained based on the fetal concentration; P(B|A i ) is the cumulative probability that the combination of the genotype of the pregnant woman and the fetus is i calculated based on the read depth at this locus and the prior probability; the number n of combinations of the genotype of the pregnant woman and the fetus = (10 genotypes of the pregnant woman × 10 genotypes of the fetus).

[0081] According to an embodiment of the present invention, the occurrence probability P j is determined based on the following formula:

[0082]

[0083] where B jF is an integer from 0 to 2, representing the occurrence times of base j in F 1i F 2i , C represents the fetal nucleic acid concentration in the predetermined region,

[0084] B jM is an integer from 0 to 2, representing the occurrence times of base j in M 1i M 2i .

[0085] According to an embodiment of the present invention, the cumulative probability P(M 1i M 2i F 1i F 2i ) is determined based on the following formula:

[0086]

[0087] where j represents the base that appears in M 1i M 2i F 1i F 2i .

[0088] According to an embodiment of the present invention, in step (f), it further includes determining the final cumulative probability P final (M 1i M 2i F 1i F 2i ) of each genotype, and selecting the genotype with the highest final probability as the combination of the pregnant woman's genotype and the fetal genotype at the predetermined locus, and further obtaining the fetal genotype at the predetermined locus, where the final cumulative probability P final (M 1i M 2i F 1i F 2i ) is determined by the following formula:

[0089]

[0090] According to an embodiment of the present invention, referring to Figure 3 after step S3000, it may include:

[0091] S3100: obtaining at least one of the maternal genotype information and the paternal genotype information of the fetus; and

[0092] S3200: Optimize the genotype set {M 1i M 2i F 1i F 2i} based on at least one of the maternal genotype information and the paternal genotype information.

[0093] According to an embodiment of the present invention, the maternal genotype information and the paternal genotype information are generated by gene sequencing of nucleated cells, and the optimization process includes removing genotypes that do not conform to Mendel's genetic laws from the genotype set {M 1i M 2i F 1i F 2i}. Thus, by determining the maternal genotype information and the paternal genotype information, the possible genotype set can be optimized, thereby improving the accuracy of determining the fetal genotype. In other words, when the genotypes of the father and mother are known, for the above Bayesian model, the combination of prior probabilities can be modified, removing the combinations of genotypes of the pregnant woman and the fetus that do not conform to Mendel's genetic laws, and recalculating the probability values of each combination. Finally, the combination with the highest probability is selected, that is, the fetal genotype is obtained.

[0094] According to an embodiment of the present invention, for the genotype information of the father or mother, it can be obtained by any known means, such as by sequencing the nucleated cells of the father or mother, and can also be obtained by long-read sequencing. Thus, the haplotype of the father or mother in this region can be effectively obtained, and the obtained results can be corrected more effectively. According to an embodiment of the present invention, there are multiple predetermined loci. Thus, after obtaining the fetal genotype of one predetermined locus, the genotypes of more loci can be analyzed, and the fetal genotype results of more loci can be obtained. Those skilled in the art can understand that when determining the fetal genotype of multiple predetermined loci, each locus needs to be analyzed independently, that is, for each predetermined locus, the above method steps are executed.

[0095] According to an embodiment of the present invention, after obtaining the reference genome sequence, the reference genome sequence is divided into multiple regions, the multiple regions include the predetermined region, and the predetermined region includes the predetermined locus.

[0096] According to an embodiment of the present invention, the reference genome sequence can be divided into multiple regions with a length of 50 to 200 kb, which can be multiple consecutive regions or discrete regions. Thus, a sufficient number of mutations are included in these regions, which can be used to analyze the concentration of fetal free nucleic acids in the regions. At the same time, it can ensure that the obtained concentration of fetal free nucleic acids can truly reflect the distribution or proportion of fetal free nucleic acids in the regions, and ensure the accuracy of genotype analysis.

[0097] Reference Figure 5 , according to an embodiment of the present invention, after S6000, the method for determining gene analysis of a pregnant woman and a fetus at a predetermined locus may further include:

[0098] S333: Obtain at least one of the maternal haploid information and paternal haploid information of the fetus, and based on at least one of the maternal haploid information and the paternal haploid information, perform correction processing on the obtained results of the combination of the genotypes of the pregnant woman and the fetus.

[0099] According to an embodiment of the present invention, the correction processing is based on linkage disequilibrium. According to an embodiment of the present invention, the correction processing is performed using the linkage disequilibrium of 200 to 400 loci (preferably 300 loci) flanking the predetermined locus. According to an embodiment of the present invention, when the number of selected loci is too small, problems may occur in the determination of paternal and maternal sources due to insufficient locus information or incorrect inference of the original loci themselves. On the other hand, when too many flanking loci are used, additional errors may be introduced due to the increase in the chromosomal recombination rate (the longer the selected genomic region, the greater the possibility of recombination). In addition, the calculation efficiency problem also needs to be considered. The more flanking loci are selected, the longer the calculation time.

[0100] According to an embodiment of the present invention, at least one of the maternal haploid information and the paternal haploid information is generated by performing long-read sequencing on nucleated cells.

[0101] Specifically, according to the embodiments of the present invention, long fragment (single-tube Long fragment reads, stLFR) sequencing can be performed on the nucleated cells in the blood of pregnant women and their husbands, and for the obtained sequencing data, tools such as Hapcut2 or Longhap can be used to complete the assembly of long fragments and the haplotype analysis of samples. Of course, if conventional sequencing data is used, tools such as SHAPEIT or BEAGLE can be used to complete haplotype analysis only using short sequencing reads. When the haplotype results of the father and mother are known, based on the above single-site fetal genotypes, the LD relationship (linkage disequilibrium) of reliable sites flanking the locus can be used to correct the sites with lower accuracy. In short, for the fetal loci that need to be corrected, the two haplotypes on which the two alleles (base types) inherited from the father and mother at this locus are respectively speculated are extracted, and the alleles (base types) actually at this locus in the two haplotypes are used as the new fetal genotype. For example, when speculating the paternal haplotype, the paternal heterozygous and maternal homozygous loci are used: when the father is AC, the mother is AA, and the fetus is AC, it is determined that the haplotype where the father's C allele (base type) is located is inherited to the fetus; when the father is AC and the fetus is AA, it is determined that the haplotype where the father's A allele (base type) is located is inherited to the fetus. When speculating the maternal haplotype, the maternal heterozygous and fetal homozygous loci are used: when the mother is AC and the fetus is CC, it is determined that the mother's C allele is inherited to the fetus. Then, with the analyzed locus as the center, a certain number of available loci are extended to both sides, and the haplotype with the highest speculated proportion among these loci is calculated, which is the inherited haplotype. Returning to the haplotype results of the father or mother, the alleles (base types) of these two haplotypes inherited to the fetus at this locus are found, which are combined as the new genotype of the fetus.

[0102] At present, for fetal genome analysis based on maternal plasma cell-free DNA whole-genome sequencing data, most of the focus is on chromosomal structural variations, large fragment copy number mutations, and the detection of paternally specific variations. According to the embodiments of the present invention, the scope of fetal genome inference using maternal plasma cell-free DNA whole-genome sequencing data can be extended to single-site variations across the entire genome, the detection accuracy can be reduced to single-base mutations, and the application scope of fetal genetic disease detection can be expanded. Additionally, according to the embodiments of the present invention, the present invention can also provide multiple analysis methods (using only maternal plasma data, using different types of data such as maternal plasma and nucleated cells in the blood of the pregnant woman and her husband), which have application value for different projects and different sample types. According to the embodiments of the present invention, the present invention further improves the accuracy of fetal genotype inference using only maternal plasma by combining local fetal concentration with a Bayesian model. Additionally, the current fetal genome analysis method using only maternal plasma cfDNA has extremely high requirements for sequencing depth (above 300x). According to the embodiments of the present invention, the inventors unexpectedly found that the present invention still achieves high accuracy when the sequencing depth is relatively low (about 60x - 100x).

[0103] Corresponding to the above method, an embodiment of the present application also provides a corresponding device for implementing the above method.

[0104] Specifically, in the third aspect of the present invention, the present invention proposes a device for determining the concentration of fetal free nucleic acids. Referring to Figure 6 , according to the embodiments of the present invention, the device includes:

[0105] A reading module 100 for obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genome sequence, where the first nucleic acid sample of the pregnant woman contains fetal free nucleic acids, and the sequencing data is composed of multiple sequencing reads;

[0106] A mutation determination module 200 for selecting a predetermined region on the reference genome sequence and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and

[0107] A fetal free nucleic acid concentration determination module 300 for determining the concentration of fetal free nucleic acids corresponding to the predetermined region based on the mutation information within the predetermined region. According to the embodiments of the present invention, the device can effectively implement the method for determining the concentration of fetal free nucleic acids. The advantages and features described for the above method are all applicable to the device and will not be elaborated further.

[0108] According to the embodiments of the present invention, the length of the predetermined region is 50 - 200 kb.

[0109] According to an embodiment of the present invention, the first nucleic acid sample of the pregnant woman is from the peripheral blood of the pregnant woman, and the mutation information includes at least one of SNP, Indel, and SV.

[0110] According to an embodiment of the present invention, the mutation determination module 200 is adapted to:

[0111] Align the sequencing data with the reference genome sequence to determine the mutation sites within the predetermined region and the base types of each of the mutation sites;

[0112] For each of the mutation sites, determine a specific base type based on the number of sequencing reads corresponding to each base type, and determine the frequency of the specific base type, so as to obtain a plurality of the frequencies;

[0113] Based on each of the plurality of frequencies, determine the mutation site proportion corresponding to the frequency, and the mutation site proportion characterizes the number of mutation sites corresponding to the frequency.

[0114] According to an embodiment of the present invention, the fetal free nucleic acid concentration determination module 300 is adapted to:

[0115] Determine the distribution of the mutation site proportion relative to the frequency;

[0116] Based on the distribution, determine the fetal free nucleic acid concentration corresponding to the predetermined region.

[0117] According to an embodiment of the present invention, the fetal free nucleic acid concentration determination module 300 is adapted to:

[0118] Within a predetermined frequency range, perform a two-dimensional mapping of the mutation site proportion relative to the plurality of frequencies; and

[0119] Based on the frequencies corresponding to the peaks and valleys of the two-dimensional mapping, determine the fetal free nucleic acid concentration.

[0120] According to an embodiment of the present invention, the predetermined frequency range is selected from 0 to 0.5 or 0.5 to 1.

[0121] According to an embodiment of the present invention, the specific base type is a minor base type, and the predetermined frequency range is 0 to 0.25 or a subset thereof, or 0.25 to 0.5 or a subset thereof. Wherein, when the predetermined frequency range is 0 to 0.25 or a subset thereof, along the direction from small to large, select the base frequency a corresponding to the second peak, and the fetal free nucleic acid concentration is 2a; when the predetermined frequency range is 0.25 to 0.5 or a subset thereof, along the direction from small to large, select the base frequency b corresponding to the first peak, and the fetal free nucleic acid concentration is 1 - 2b.

[0122] According to an embodiment of the present invention, the specific base type is the main base type, and the predetermined frequency range is 0.5 to 0.75 or a subset thereof, or 0.75 to 1 or a subset thereof. Wherein, when the predetermined frequency range is 0.5 to 0.75 or a subset thereof, along the direction from small to large, the base frequency c corresponding to the second trough is selected, and the fetal free nucleic acid concentration is 2c; when the predetermined frequency range is 0.75 to 1 or a subset thereof, along the direction from small to large, the base frequency d corresponding to the first trough is selected, and the fetal free nucleic acid concentration is 1 - 2d.

[0123] In a fourth aspect of the present invention, the present invention provides an apparatus for determining the fetal genotype at a predetermined locus. Referring to Figure 7 , the apparatus includes:

[0124] The aforementioned device 1000 for determining the concentration of fetal free nucleic acids, which is used to determine the concentration of fetal free nucleic acids in the predetermined region, and the predetermined region includes the predetermined locus;

[0125] The sequencing read count determination module 2000 is configured to respectively determine the number of sequencing reads A that support base A, base T, base G, or base C for the predetermined locus j , where j is A, T, G, or C;

[0126] The genotype set construction module 3000 is configured to construct a genotype set {M 1i M 2i F 1i F 2i} for the predetermined locus, where i represents the genotype number, M 1i represents the base type on the mother's first chromosome for the i-th genotype, M 2i represents the base type on the mother's second chromosome for the i-th genotype, F 1i represents the base type on the fetus's first chromosome for the i-th genotype, F 2i represents the base type on the fetus's second chromosome for the i-th genotype, M 1i , M 2i , F 1i and F 2i are each independently base A, base T, base G, or base C. Wherein, the mother's first chromosome and the mother's second chromosome belong to a pair of homologous chromosomes, and the fetus's first chromosome and the fetus's second chromosome belong to a pair of homologous chromosomes;

[0127] The occurrence probability determination module 4000 is configured to, for the genotype set {M 1i M 2i F 1i F2i For each of {}, based on the fetal nucleic acid concentration, determine the occurrence probability P of each base j , where j is A, T, G, or C;

[0128] Cumulative probability determination module 5000, for the genotype set {M 1i M 2i F 1i F 2i}, for each of them, based on the occurrence probability P of each base j and the number of sequencing reads A j , determine the cumulative probability P(M 1i M 2i F 1i F 2i ); and

[0129] Genotype combination determination module 6000, for based on the cumulative probability P(M 1i M 2i F 1i F 2i ), determine the pregnant woman genotype and fetal genotype combination at the predetermined locus, and then obtain the fetal genotype at the predetermined locus.

[0130] According to an embodiment of the present invention, the occurrence probability P j is determined based on the following formula:

[0131]

[0132] where B jF is an integer from 0 to 2, representing the number of occurrences of base j in F 1i F 2i , C represents the fetal nucleic acid concentration in the predetermined region,

[0133] B jM is an integer from 0 to 2, representing the number of occurrences of base j in M 1i M 2i .

[0134] According to an embodiment of the present invention, the cumulative probability P(M 1i M 2i F 1i F 2i ) is determined based on the following formula:

[0135]

[0136] where j represents M 1i M 2i F 1i F 2iThe bases that appear in

[0137] Reference Figure 8 , according to an embodiment of the present invention, the genotype combination determination module includes 6000: a final cumulative probability determination unit 6100 for determining the final cumulative probability P of each of the genotypes final (M 1i M 2i F 1i F 2i ); a genotype combination selection unit 6200 for selecting the genotype with the highest final cumulative probability as the genotype combination of the pregnant woman and the fetus at the predetermined locus, and further obtaining the fetal genotype at the predetermined locus, wherein the final cumulative probability P final (M 1i M 2i F 1i F 2i ) is determined by the following formula:

[0138]

[0139] Reference Figure 9 As shown, according to an embodiment of the present invention, it further includes a genotype set optimization module 3200 for: obtaining at least one of the maternal genotype information and the paternal genotype information of the fetus; and based on at least one of the maternal genotype information and the paternal genotype information, optimizing the genotype set {M 1i M 2i F 1i F 2i}.

[0140] According to an embodiment of the present invention, the maternal genotype information and the paternal genotype information are generated by gene sequencing of nucleated cells, and the optimization processing module includes a filtering unit for removing genotypes that do not conform to Mendelian inheritance laws from the genotype set {M 1i M 2i F 1i F 2i}.

[0141] Reference Figure 10 As shown, according to an embodiment of the present invention, it further includes a region division module Mi for dividing the reference genome sequence into multiple regions after obtaining the reference genome sequence, the multiple regions including the predetermined region, and the predetermined region including the predetermined locus.

[0142] According to an embodiment of the present invention, the predetermined locus includes multiple mutation sites.

[0143] Reference Figure 11, according to an embodiment of the present invention, the device for determining the fetal gene at the predetermined locus further includes:

[0144] A correction module Miii, configured to obtain at least one of the maternal haploid information and the paternal haploid information of the fetus, and perform a correction process on the result of the combination of the pregnant woman's genotype and the fetus's genotype based on at least one of the maternal haploid information and the paternal haploid information.

[0145] According to an embodiment of the present invention, the correction process is based on linkage disequilibrium. According to an embodiment of the present invention, the correction process is performed using the linkage disequilibrium of 200 - 400 loci flanking the predetermined locus. Preferably, the linkage disequilibrium of 300 loci flanking the predetermined locus is used.

[0146] According to an embodiment of the present invention, at least one of the maternal haploid information and the paternal haploid information is generated by performing long - fragment sequencing on nucleated cells.

[0147] In a fifth aspect of the present invention, a computer - readable storage medium is proposed, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method described above are implemented.

[0148] It should be noted that those skilled in the art can understand that the features and advantages described in each aspect are equally applicable to other aspects, and will not be elaborated here.

[0149] The solution of the present invention will be explained below in conjunction with embodiments. Those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention. For those not specified in the embodiments regarding specific techniques or conditions, they are carried out according to the techniques or conditions described in the literature in the art or according to the product specifications. For those not specified in the embodiments regarding specific conditions, they are carried out according to the conventional conditions or the conditions recommended by the manufacturer. For the reagents or instruments not specified regarding the manufacturer, they are all conventional products that can be obtained in the market.

[0150] Example 1

[0151] Experimental method

[0152] Three families were selected, and the sequencing data of maternal plasma cfDNA (cell-free nucleic acid) sequencing (whole-genome sequencing WGS, paired-end 100bp sequencing, with an average depth of about 100x), and stLFR sequencing of nucleated blood cells from the father and mother (whole-genome sequencing WGS, paired-end 100bp sequencing, with an average depth of about 40-60x) were obtained respectively. The whole-genome sequencing data of fetal umbilical cord blood (paired-end 100bp, with an average depth of about 40x) was used as the true set to calculate the accuracy of fetal genotype inference. The overall accuracy of all loci in the 3 families reached over 95% (that is, regardless of the heterozygous and homozygous situations, the subsequent various methods can achieve an overall accuracy of over 95%).

[0153] Reference Figure 13 , for the above three families, the following schemes can be used for analysis respectively:

[0154] Scheme 1: Only use the sequencing data of maternal plasma cfDNA (cell-free nucleic acid). After determining the fetal concentration in the region, Bayesian inference is performed to obtain the fetal genotype.

[0155] Scheme 2 (this scheme is not directly applied): On the basis of Scheme 1, the genotype data from the stLFR sequencing of maternal nucleated blood cells is used to optimize the fetal genotype in Scheme 1.

[0156] Scheme 3: On the basis of Scheme 1, the genotype data from the stLFR sequencing of maternal and paternal nucleated blood cells is used to optimize the fetal genotype in Scheme 1.

[0157] Scheme 4: On the basis of Scheme 1, the haplotype data from the stLFR sequencing of maternal and paternal nucleated blood cells is used to optimize the fetal genotype in Scheme 1.

[0158] Experimental results

[0159] Based on the parental stLFR data of each family, the haplotype results of the paternal and maternal are obtained respectively, and the switch error rate is calculated to evaluate the accuracy of the stLFR data.

[0160] Based on the data of the above 3 families, the tests of the aforementioned Scheme 1 and Scheme 3 were completed respectively, and the accuracies of different types of loci were obtained and summarized in the following table.

[0161] Table 1. Summary of the accuracy of fetal genome inference by different schemes at different types of loci

[0162]

[0163]

[0164] Note: Accuracy refers to the accuracy of fetal genotype inference. The accuracy is calculated by directly comparing the inferred fetal genotype at each locus with the true fetal genotype obtained from the fetal umbilical cord blood sequencing data, and it is the proportion of the number of loci with consistent genotypes to the total number of analyzed loci.

[0165] Among them, for Family 1, the fetal genotype correction scheme of Scheme 4 is adopted. For the maternal heterozygous and fetal heterozygous loci with relatively low accuracy in Schemes 1 and 3, the accuracy can be further increased by about 15%. As Figure 14 shown, after setting different numbers of flanking loci in Scheme 4 to complete the correction of the maternal heterozygous and fetal heterozygous loci in Family 1, the comparison of the accuracy with Scheme 3 shows that the number of flanking loci has a relatively obvious impact on the correction effect. Therefore, considering comprehensively, 200 - 400 flanking loci can be adopted, and preferably 300 flanking loci are adopted.

[0166] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0167] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for determining the concentration of fetal free nucleic acid, characterized in that, Comprising: (1) Obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genome sequence, wherein the first nucleic acid sample of the pregnant woman contains cell-free fetal nucleic acid, and the sequencing data is composed of a plurality of sequencing reads; (2) Selecting a predetermined region on the reference genome sequence, and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and (3) Determining the concentration of cell-free fetal nucleic acid corresponding to the predetermined region based on the mutation information within the predetermined region; In step (2), determining the mutation information within the predetermined region includes: (2-1) Aligning the sequencing data with the reference genome sequence to determine mutation sites within the predetermined region and the base types of each of the mutation sites; (2-2) For each of the mutation sites, determining a specific base type based on the number of sequencing reads corresponding to each base type, and determining the frequency of the specific base type, so as to obtain a plurality of the frequencies; the specific base type is the major base type or the minor base type; (2-3) Based on each of the plurality of frequencies, determining the proportion of mutation sites corresponding to the frequency, and the proportion of mutation sites represents the number of mutation sites corresponding to the frequency; Step (3) further includes: (3-1) Determining the distribution of the proportion of mutation sites relative to the frequency; (3-2) Based on the distribution, determining the concentration of cell-free fetal nucleic acid corresponding to the predetermined region; In step (3-1), within a predetermined frequency range, making a two-dimensional plot of the proportion of mutation sites relative to the frequency; and In step (3-2), based on the frequencies corresponding to the peaks and valleys of the two-dimensional plot, determining the concentration of cell-free fetal nucleic acid; The length of the predetermined region is 50-200 kb; The specific base type is the minor base type, wherein when the predetermined frequency range is 0-0.25, along the direction from small to large, selecting the base frequency a corresponding to the second peak, and the concentration of cell-free fetal nucleic acid is 2a; The specific base type is the minor base type, wherein when the predetermined frequency range is 0.25-0.5, along the direction from small to large, selecting the base frequency b corresponding to the first peak, and the concentration of cell-free fetal nucleic acid is 1-2b; The specific base type is the major base type, wherein when the predetermined frequency range is 0.5-0.75, along the direction from small to large, selecting the base frequency c corresponding to the second peak, and the concentration of cell-free fetal nucleic acid is 2c.

2. The method according to claim 1, wherein The length of the predetermined region is 100 kb.

3. The method according to claim 1, wherein The sequencing depth of the sequencing data does not exceed 100X.

4. The method according to claim 1, characterized in that The sequencing depth of the sequencing data is 60X-100X.

5. The method according to claim 1, wherein The first nucleic acid sample of the pregnant woman is from the peripheral blood of the pregnant woman, and the mutation information includes at least one of SNP, Indel and SV.

6. A method for determining the fetal genotype at a predetermined locus, characterized in that, Comprising: (a) According to the method according to any one of claims 1-5, determining the concentration of cell-free fetal nucleic acid in the predetermined region, and the predetermined region contains a predetermined site; (b) For the pre-determined loci, determine the number of sequencing reads A that support base A, base T, base G, or base C, respectively, where j is A, T, G, or C. j , where j is A, T, G, or C. (c) For the said pre - positioned locus, construct a genotype set {M 1i M 2i F 1i F 2i}, where i represents the genotype number, M 1i represents the base type on the first chromosome of the mother for the i - th genotype, M 2i represents the base type on the second chromosome of the mother for the i - th genotype, F 1i represents the base type on the first chromosome of the fetus for the i - th genotype, F 2i represents the base type on the second chromosome of the fetus for the i - th genotype, M 1i 、M 2i 、F 1i and F 2i are each independently base A, base T, base G or base C. Among them, the first chromosome of the mother and the second chromosome of the mother belong to a pair of homologous chromosomes, and the first chromosome of the fetus and the second chromosome of the fetus belong to a pair of homologous chromosomes; (d) For each of the genotype sets {M 1i M 2i F 1i F 2i}, based on the fetal free nucleic acid concentration, determine the occurrence probability P of each base j , where j is A, T, G, or C; (e) For each of the genotype sets {M 1i M 2i F 1i F 2i}, based on the occurrence probability P j of each of the bases and the number A j of the sequencing reads, determine the cumulative probability P(M 1i M 2i F 1i F 2i ); (f) Based on the cumulative probability P(M 1i M 2i F 1i F 2i ), determine the combination of the pregnant woman's genotype and the fetus's genotype at the predetermined locus, and further obtain the fetus's genotype at the predetermined locus; The occurrence probability P j is determined based on the following formula: where B jF is an integer from 0 to 2, representing the number of occurrences of base j in F 1i F 2i and C represents the concentration of the fetal free nucleic acid in the predetermined region B jM is an integer from 0 to 2, representing the occurrence times of base j in M 1i M 2i ; the number of occurrences of base j in M The cumulative probability P(M 1i M 2i F 1i F 2i ) is determined based on the following formula:

7. The method according to claim 6, wherein In step (f), it further includes determining the final cumulative probability P final (M 1i M 2i F 1i F 2i ) of each of the genotypes, and selecting the genotype with the highest final cumulative probability as the combination of the pregnant woman's genotype and the fetal genotype at the predetermined locus, thereby obtaining the fetal genotype at the predetermined locus, where the final cumulative probability P final (M 1i M 2i F 1i F 2i ) is determined by the following formula:

8. The method according to claim 6, characterized in that, After step (c), including: (c-1) Obtain at least one of the maternal genotype information and the paternal genotype information of the fetus; and (c-2) Optimize the genotype set {M 1i M 2i F 1i F 2i} based on at least one of the female parent genotype information and the male parent genotype information.

9. The method according to claim 8, wherein The maternal genotype information and the paternal genotype information are generated by gene sequencing of nucleated cells, The optimization process includes removing genotypes that do not conform to Mendelian inheritance laws from the genotype set {M 1i M 2i F 1i F 2i}.

10. The method according to claim 6, wherein including: After obtaining the reference genome sequence, divide the reference genome sequence into multiple regions, the multiple regions include the predetermined region, and the predetermined region includes the predetermined locus.

11. The method according to claim 6, wherein There are multiple predetermined loci.

12. The method according to claim 6, wherein Further include: (g) Obtain at least one of the maternal haploid information and the paternal haploid information of the fetus, and based on at least one of the maternal haploid information and the paternal haploid information, perform correction processing on the pregnant woman genotype and fetus genotype results obtained in step (f).

13. The method according to claim 12, characterized in that, The correction processing is based on linkage disequilibrium.

14. The method according to claim 13, characterized in that, The correction processing is performed using the linkage disequilibrium of 200-400 loci flanking the predetermined locus.

15. The method according to claim 12, wherein At least one of the maternal haploid information and the paternal haploid information is generated by long-read sequencing of nucleated cells.

16. An apparatus for determining the concentration of fetal free nucleic acid, characterized in that, including: A reading module for obtaining sequencing data of a first nucleic acid sample of a pregnant woman and a reference genome sequence, the first nucleic acid sample of the pregnant woman containing cell-free fetal nucleic acids, and the sequencing data consisting of multiple sequencing reads; A mutation determination module for selecting a predetermined region on the reference genome sequence and determining mutation information within the predetermined region based on the sequencing data of the first nucleic acid sample of the pregnant woman; and A cell-free fetal nucleic acid concentration determination module for determining the cell-free fetal nucleic acid concentration corresponding to the predetermined region based on the mutation information within the predetermined region; The mutation determination module is adapted to: Align the sequencing data with the reference genome sequence to determine mutation sites within the predetermined region and the base types of each mutation site; For each mutation site, determine a specific base type based on the number of sequencing reads corresponding to each base type and determine the frequency of the specific base type to obtain multiple frequencies; the specific base type is the major base type or the minor base type; Based on each of the multiple frequencies, determine the mutation site proportion corresponding to the frequency, and the mutation site proportion represents the number of mutation sites corresponding to the frequency; The cell-free fetal nucleic acid concentration determination module is adapted to: Determine the distribution of the mutation site proportion relative to the frequency; Based on the distribution, determine the cell-free fetal nucleic acid concentration corresponding to the predetermined region; The cell-free fetal nucleic acid concentration determination module is adapted to: Within a predetermined frequency range, perform two-dimensional mapping of the mutation site proportion relative to the frequency; and Based on the frequencies corresponding to the peaks and valleys of the two-dimensional mapping, determine the cell-free fetal nucleic acid concentration; The length of the predetermined region is 50-200 kb; The specific base type is the minor base type, wherein when the predetermined frequency range is 0-0.25, along the direction from small to large, select the base frequency a corresponding to the second peak, and the cell-free fetal nucleic acid concentration is 2a; The specific base type is a secondary base type. When the predetermined frequency range is 0.25 to 0.5, along the direction from small to large, the base frequency b corresponding to the first peak is selected, and the concentration of fetal free nucleic acid is 1 - 2b; The specific base type is a primary base type. When the predetermined frequency range is 0.5 to 0.75, along the direction from small to large, the base frequency c corresponding to the second peak is selected, and the concentration of fetal free nucleic acid is 2c.

17. The device according to claim 16, characterized in that, The length of the predetermined region is 100 kb.

18. The device according to claim 16, characterized in that, The sequencing depth of the sequencing data does not exceed 100X.

19. The device according to claim 16, characterized in that, The sequencing depth of the sequencing data is 60X to 100X.

20. The device according to claim 16, characterized in that, The first nucleic acid sample of the pregnant woman is from the peripheral blood of the pregnant woman, and the mutation information includes at least one of SNP, Indel, and SV.

21. A device for determining the fetal genotype of a predetermined locus, characterized in that, Including: The device for determining the concentration of fetal free nucleic acid according to any one of claims 16 to 20, which is used to determine the concentration of fetal free nucleic acid in the predetermined region, and the predetermined region contains a predetermined locus; A sequencing read count determination module, which is used to respectively determine the number of sequencing reads A that support base A, base T, base G, or base C for the pre-positioned site j , where j is A, T, G, or C; Genotype set construction module, which is used to construct a genotype set {M 1i M 2i F 1i F 2i} for the predetermined locus, where i represents the genotype number, M 1i represents the base type on the first chromosome of the mother for the i-th genotype, M 2i represents the base type on the second chromosome of the mother for the i-th genotype, F 1i represents the base type on the first chromosome of the fetus for the i-th genotype, F 2i represents the base type on the second chromosome of the fetus for the i-th genotype, M 1i 、M 2i 、F 1i and F 2i are each independently base A, base T, base G or base C. Among them, the first chromosome of the mother and the second chromosome of the mother belong to a pair of homologous chromosomes, and the first chromosome of the fetus and the second chromosome of the fetus belong to a pair of homologous chromosomes; An occurrence probability determination module, which is used to determine the occurrence probability P of each base based on the fetal free nucleic acid concentration for each of the genotype sets {M 1i M 2i F 1i F 2i}, where j is A, T, G or C; j ,j is A, T, G or C; Cumulative probability determination module, which is used to, for each of the genotype sets {M 1i M 2i F 1i F 2i}, based on the occurrence probability P j of each of the bases and the number A j of the sequencing reads, determine the cumulative probability P(M 1i M 2i F 1i F 2i ); Genotype combination determination module, which is used to determine the genotype combination of the pregnant woman and the fetus at the predetermined locus based on the cumulative probability P(M 1i M 2i F 1i F 2i ), and further obtain the fetal genotype at the predetermined locus; The occurrence probability P j is determined based on the following formula: Among them, B jF is an integer from 0 to 2, representing the number of occurrences of base j in F 1i F 2i and C represents the concentration of the fetal free nucleic acid in the predetermined region B jM is an integer from 0 to 2, representing the number of occurrences of base j in M 1i M 2i ; The cumulative probability P(M 1i M 2i F 1i F 2i ) is determined based on the following formula:

22. The device according to claim 21, wherein The genotype combination determination module includes: Final cumulative probability determination unit, configured to determine the final cumulative probability P of each of the genotypes final (M 1i M 2i F 1i F 2i ); A genotype combination selection unit, which is used to select the genotype with the highest final cumulative probability as the genotype combination of the pregnant woman and the fetus at the predetermined locus, and further obtain the fetal genotype at the predetermined locus. Among them, The final cumulative probability P final (M 1i M 2i F 1i F 2i ) is determined by the following formula:

23. The device according to claim 21, characterized in that, It further includes a genotype set optimization module, which is used for: Obtaining at least one of the maternal genotype information and paternal genotype information of the fetus; and Based on at least one of the maternal genotype information and the paternal genotype information, optimize the genotype set {M 1i M 2i F 1i F 2i}.

24. The device according to claim 23, wherein, The maternal genotype information and the paternal genotype information are generated by gene sequencing of nucleated cells. The optimization processing module includes a filtering unit for removing genotypes that do not conform to Mendelian inheritance laws from the genotype set {M 1i M 2i F 1i F 2i}.

25. The device according to claim 21, wherein It further includes: A region division module, which is used to divide the reference genome sequence into multiple regions after obtaining the reference genome sequence. The multiple regions include the predetermined region, and the predetermined region contains the predetermined locus.

26. The device according to claim 25, wherein There are multiple predetermined loci.

27. The device according to claim 21, characterized in that, It further includes: A correction module, which is used to obtain at least one of the maternal haploid information and paternal haploid information of the fetus, and perform correction processing on the genotype combination result of the pregnant woman and the fetus based on at least one of the maternal haploid information and the paternal haploid information.

28. The device according to claim 27, wherein The correction processing is based on linkage disequilibrium.

29. The device according to claim 28, characterized in that, The correction processing is performed using the linkage disequilibrium of 200 to 400 loci flanking the predetermined locus.

30. The device according to claim 27, characterized in that, At least one of the maternal haploid information and the paternal haploid information is generated by long - fragment sequencing of nucleated cells.

31. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and device for simultaneously determining fetal nucleic acid content and aneuploidy of chromosome

    CN104232777A