Noninvasive prenatal parent-child detection and analysis method and device and storage medium
By acquiring SNP locus information and calculating genetic laws, the problem of identifying complex kinship in non-invasive prenatal paternity testing has been solved, providing an accurate method and device for determining parentage, meeting the needs of the judiciary and hospitals.
Patent Information
- Application Number
- CN202410847837.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-12-30
AI Technical Summary
Existing non-invasive prenatal paternity testing technology cannot effectively handle complex kinship relationships, especially cases where the biological father suspects the mother or the biological mother suspects the father, resulting in a lack of rigorous theoretical basis in the fields of law and hospitals.
By employing methods such as SNP locus information acquisition, exclusion probability calculation, and paternity index calculation, and by deriving the exclusion probability and paternity index of various complex kinship relationships through Mendelian inheritance laws, a non-invasive prenatal paternity testing and analysis method and device are provided.
It enables accurate identification of complex kinship relationships, provides scientific evidence for the judiciary, public security and hospitals, and improves the accuracy and reliability of non-invasive prenatal paternity testing.
Smart Images

Figure CN121237202A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of genetic testing, and in particular to a method, apparatus and storage medium for non-invasive prenatal paternity testing and analysis. Background Technology
[0002] Prenatal paternity testing is the identification of parentage before birth. Traditional prenatal paternity testing techniques include chorionic villus sampling, amniocentesis, and cordocentesis. These traditional techniques are all invasive and can easily cause fetal malformations or miscarriages, and sampling can generally only be performed after 11 weeks of gestation. With the development of high-throughput sequencing technology, non-invasive prenatal paternity testing (NIPPT) has emerged. In 2016, Jiang et al. published "Noninvasive prenatal paternity testing (NIPAT) through maternal plasma DNA sequencing: A pilot study" [J], which explored the impact of low allele frequency, number of SNP sites, proportion of fetal cell-free DNA (cfDNA), and effective sequencing depth design on non-invasive prenatal paternity testing by screening a series of effective SNP sites through whole-genome sequencing. In their 2012 paper, *Evaluation of 96 SNPs in 14 Populations for Worldwide Individual Identification*, Zhaoshu Zeng et al. used Illumina Global Gate technology to genotype 96 SNP loci in 192 samples, demonstrating that SNPs are effective molecular genetic markers in the field of forensic identification and are applicable to highly degraded samples. In their 2019 paper, *Development and comprehensive evaluation of a noninvasive prenatal paternity testing method through a scaled trial*, Liao Chang et al. enriched over 5000 SNP loci using a probe capture method, demonstrating its applicability for noninvasive prenatal paternity testing in 36 pregnant women. In their 2023 paper, *A theoretical base for non-invasive prenatal paternity testing*, Shengjie Gao et al. established a theoretical model for prenatal paternity testing using high-throughput sequencing technology. This technology calculates fetal concentration based on the Poisson distribution model, and can accurately analyze fetal concentration using only the sequencing depth of maternal blood cfDNA. Based on the fetal concentration, the sequencing samples are divided into regular samples and low fetal concentration samples. This statistical model is applicable to triplet samples of regular samples and low fetal concentration samples.
[0003] With the development of science and technology, the relationship between the fetus and the pregnant mother and the alleged father has become increasingly complex, such as the biological father suspecting the mother, or the biological mother suspecting the father. Current non-invasive prenatal paternity testing cannot accurately and effectively identify these complex kinship relationships, which makes it impossible for judicial, public security or hospital work that requires strict theoretical basis to be carried out effectively. Summary of the Invention
[0004] The purpose of this application is to provide a new method, apparatus, and storage medium for non-invasive prenatal paternity testing and analysis.
[0005] The following technical solution is adopted in this application:
[0006] One aspect of this application discloses a data analysis method for non-invasive prenatal paternity testing, including steps for obtaining SNP locus information, calculating exclusion probability, calculating paternity index, and making a judgment.
[0007] The steps for obtaining SNP locus information include obtaining SNP locus information from the pregnant mother, fetus, and the sample to be tested. The SNP locus information includes whether the SNP locus is homozygous or heterozygous, the frequency of each allele at the SNP locus, and the number of alleles at the SNP locus. The sample to be tested is either a suspected paternal sample or a suspected maternal sample. If the sample to be tested is a suspected maternal sample, and the biological father's information is available, the steps also include obtaining the biological father's SNP locus information.
[0008] The exclusion probability calculation steps include screening SNP loci based on the identification content, and calculating the exclusion probability PE using the corresponding formula based on the screened SNP loci; the identification content includes prenatal maternal-suspected father identification, prenatal low-degree maternal-suspected father identification, prenatal father-child identification, prenatal mother-child identification, and prenatal father-suspected mother identification.
[0009] Prenatal paternity testing involves identifying the suspected father of a fetus from a known pregnant mother. The test sample (the suspected father) is used to determine whether the fetus's biological father is the biological father. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with homozygous AA genotypes in the pregnant mother and then calculating the probability of exclusion (PE) for the suspected father sample according to Formula 1. PAF ;
[0010] Formula 1:
[0011] Prenatal low-level maternal paternity identification refers to the identification of a sample with a known maternal mother who is the biological mother of the fetus, and a low fetal concentration. The test sample (i.e., the suspected paternal sample) is used to determine whether it is the biological father of the fetus. The exclusion probability of the test sample not being the biological father is calculated by screening for SNP loci with the maternal × fetal genotype of AA × AB, and then calculating the exclusion probability PE of the suspected paternal sample according to Formula 2. PLAF ;
[0012] Formula 2:
[0013] Prenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the suspected father sample is the biological father of the fetus. The probability of exclusion for a non-biological father sample is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC ;
[0014] Formula 3:
[0015] Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother. The identification of the test sample (suspected mother) determines whether it is the biological mother of the fetus. The exclusion probability of the test sample being a non-biological mother is calculated by screening for SNP loci with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMC ;
[0016] Formula 4:
[0017] Prenatal paternity testing is used when the suspected mother is a known biological father of the fetus, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM ;
[0018] Formula 5:
[0019] In formulas 1 to 5, P i denoted as the frequency of the i-th allele at the SNP locus, and n is the number of alleles at the SNP locus.
[0020] The steps for calculating the parentage index include: deriving the probability X of the test sample being the biological father or mother of the fetus based on the SNP locus and Mendel's laws of inheritance; deriving the probability Y of the random sample being the biological father or mother of the fetus; and using the ratio of X to Y, i.e., X / Y, as the parentage index PI.
[0021] The determination process includes calculating the cumulative exclusion probability (CPE) based on the exclusion probability (PE) of each SNP locus, calculating the cumulative paternity index (CPI) based on the paternity index (PI) of each SNP locus, and determining the kinship relationship based on the CPE and CPI.
[0022] The non-invasive prenatal paternity testing method of this application derives and verifies the formulas for Power of Exclusion (PE) and Parentage Index (PI) for different relationships between the fetus and the pregnant mother, the alleged father, etc., thereby enabling accurate and effective paternity testing for complex kinship relationships and providing a scientific basis for fields such as the judiciary, public security, or hospitals that require strict theoretical basis.
[0023] In this application, "biological father" refers to the biological father, "biological mother" refers to the biological mother, and "pregnant mother" refers to the mother who carries the fetus. Exclusion probability is the probability of excluding the test sample from being the biological father or biological mother, and the parentage index is a parameter supporting the view that the test sample is the biological father or biological mother.
[0024] Preferably, the paternity index calculation step also includes calculating the paternity index (PI) in different ways for five cases: prenatal maternal-suspected father identification, prenatal low-degree maternal-suspected father identification, prenatal paternal-child identification, prenatal mother-child identification, and prenatal paternal-suspected mother identification, based on the exclusion probability calculation step. For prenatal maternal-suspected father identification, the paternity index (PI) is calculated according to Table 1. PAF Prenatal low-grade maternal-paternity identification was performed according to the paternity index (PI) calculated in Table 2. PLAF Prenatal paternity testing was performed to calculate the parentage index (PI) according to Table 3. PFC Prenatal maternal-fetal paternity index (PI) was calculated according to Table 4. PMC Prenatal paternity testing for suspected maternity is performed according to the paternity index (PI) calculated in Table 5. PAM ;
[0025] Table 1
[0026]
[0027]
[0028] Table 2
[0029] Maternal genotype fetal genotype Genotype of suspected father sample X Y <![CDATA[PI PLAF ]]> AA AB BB 1 / 2 b / 2 1 / b AA AB AB 1 / 4 b / 2 1 / 2b AA AB AA μ b / 2 2μ / b
[0030] Table 3
[0031] Maternal genotype fetal genotype Genotype of suspected father sample X Y <![CDATA[PI PFC ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a 1 / 2a AA AA BB μ a μ / a
[0032] Table 4
[0033] Maternal genotype fetal genotype suspected maternal sample genotype X Y <![CDATA[PI PMC ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a 1 / 2a AA AA BB μ a μ / a
[0034] Table 5
[0035] Maternal genotype fetal genotype Father's genotype X Y <![CDATA[PI PAM ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a / 2 1 / a AA AA BB μ μa 1 / a AA AB / BB AA <![CDATA[(2μa+μ 2 b) / (2a+b)]]> <![CDATA[(ab+μb 2 ) / (2a+b)]]> <![CDATA[(2μa+μ 2 b) / (ab+μb 2 )]]>
[0036] In Tables 1 to 5, X represents the probability that the sample being tested is the biological father or mother of the fetus, Y represents the probability that a random sample is the biological father or mother of the fetus, μ represents the probability of a single gene mutation, a represents the frequency of allele A, and b represents the frequency of allele B. The specific calculation formulas for X and Y in Tables 1 to 5 of this application were derived by this application based on different parent-child relationships and specific genotypes, according to Mendel's laws.
[0037] Preferably, μ is 0.00001.
[0038] Preferably, the random sample is a sample randomly obtained from a database of 1,000 people.
[0039] Preferably, in the judgment step, In the formula, PE i Let PI be the exclusion probability of the i-th SNP site. i denoted as the parentage index of the i-th SNP locus, and n is the number of SNP loci.
[0040] Preferably, the determination step involves determining kinship based on CPE and CPI. Specifically, if CPI > 10000 and CPE > 0.9999, the child is determined to be biologically related; if CPI < 0.0001 and CPE > 0.9999, the child is determined to be non-biologically related; otherwise, the relationship cannot be determined.
[0041] Preferably, the sample with low fetal concentration is the test sample with a fetal cfDNA concentration of less than 4%, where fetal cfDNA concentration is the proportion of fetal cfDNA in the pregnant woman's peripheral blood to the total cfDNA in the pregnant woman.
[0042] Preferably, in the determination step, the number of SNP sites used for detection is at least 100.
[0043] The second aspect of this application discloses a non-invasive prenatal paternity testing and analysis device, including an SNP locus information acquisition module, an exclusion probability calculation module, a paternity index calculation module, and a judgment module;
[0044] The SNP locus information acquisition module includes SNP locus information for the pregnant mother, fetus, and sample to be tested. The SNP locus information includes whether the SNP locus is homozygous or heterozygous, the frequency of each allele at the SNP locus, and the number of alleles at the SNP locus. The sample to be tested is either a suspected paternal sample or a suspected maternal sample.
[0045] The exclusion probability calculation module includes a method for screening SNP loci based on the identification content and calculating the exclusion probability PE using the corresponding formula based on the screened SNP loci. The identification content includes prenatal maternal-suspected father identification, prenatal low-degree maternal-suspected father identification, prenatal father-child identification, prenatal mother-child identification, and prenatal father-suspected mother identification.
[0046] Prenatal paternity testing involves identifying the suspected father of a fetus from a known pregnant mother. The test sample (the suspected father) is used to determine whether the fetus's biological father is the biological father. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with homozygous AA genotypes in the pregnant mother and then calculating the probability of exclusion (PE) for the suspected father sample according to Formula 1. PAF ;
[0047] Formula 1:
[0048] Prenatal low-level maternal paternity identification refers to the identification of a sample with a known maternal mother who is the biological mother of the fetus, and a low fetal concentration. The test sample (i.e., the suspected paternal sample) is used to determine whether it is the biological father of the fetus. The exclusion probability of the test sample not being the biological father is calculated by screening for SNP loci with the maternal × fetal genotype of AA × AB, and then calculating the exclusion probability PE of the suspected paternal sample according to Formula 2. PLAF ;
[0049] Formula 2:
[0050] Prenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the suspected father sample is the biological father of the fetus. The probability of exclusion for a non-biological father sample is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC ;
[0051] Formula 3:
[0052] Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother. The identification of the test sample (suspected mother) determines whether it is the biological mother of the fetus. The exclusion probability of the test sample being a non-biological mother is calculated by screening for SNP loci with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMC ;
[0053] Formula 4:
[0054] Prenatal paternity testing is used when the suspected mother is a known biological father of the fetus, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM ;
[0055] Formula 5:
[0056] In formulas 1 to 5, P i denoted as the frequency of the i-th allele at the SNP locus, and n is the number of alleles at the SNP locus.
[0057] The parentage index calculation module includes a module for deriving the probability X of a test sample being the biological father or mother of the fetus based on SNP loci and Mendel's laws of inheritance, and deriving the probability Y of a random sample being the biological father or mother of the fetus. The ratio of X to Y, i.e., X / Y, is used as the parentage index PI.
[0058] The judgment module includes calculating the cumulative non-father or non-mother exclusion probability (CPE) based on the exclusion probability (PE) of each SNP locus, calculating the cumulative paternity index (CPI) based on the paternity index (PI) of each SNP locus, and determining the kinship relationship based on the CPE and CPI.
[0059] It should be noted that the device for non-invasive prenatal paternity testing and analysis in this application actually implements each step of the method for non-invasive prenatal paternity testing and analysis in this application through each module; therefore, the specific limitations of each module can be referred to the method of this application, such as calculating the parentage index PI for five situations according to Tables 1 to 5, etc., which will not be elaborated here.
[0060] A third aspect of this application discloses an apparatus for non-invasive prenatal paternity testing and analysis, the apparatus comprising a memory and a processor; the memory including a program for storing programs; and the processor including a method for implementing the non-invasive prenatal paternity testing and analysis method of this application by executing the program stored in the memory.
[0061] The fourth aspect of this application discloses a computer-readable storage medium storing a program that can be executed by a processor to implement the non-invasive prenatal paternity testing and analysis method of this application.
[0062] The beneficial effects of this application are as follows:
[0063] The non-invasive prenatal paternity testing and analysis method proposed in this application derives formulas for exclusion probability and paternity index for different relationships between the fetus and the pregnant mother, the alleged father, etc. It can accurately and effectively identify paternity in various complex kinship relationships, providing accurate and effective scientific evidence for fields such as the judiciary, public security, or hospitals that require strict theoretical basis. Attached Figure Description
[0064] Figure 1 PE in the embodiments of this application PAF The fitting curve between the formula value and the simulated value;
[0065] Figure 2 PE in the embodiments of this application PLAF The fitting curve between the formula value and the simulated value;
[0066] Figure 3 PE in the embodiments of this application PFC The fitting curve between the formula value and the simulated value;
[0067] Figure 4 PE in the embodiments of this application PAM The fitted curve of the formula value and the simulated value. Detailed Implementation
[0068] Existing non-invasive prenatal paternity testing techniques lack standardization, theoretical derivation, and verification, severely hindering their application in the industry. This application, for the first time, derives and verifies the core formulas and statistical models for all types of non-invasive prenatal paternity testing, providing a rigorous theoretical basis for the technology in judicial, public security, and hospital fields.
[0069] This application provides a comprehensive definition and mathematical model for various prenatal paternity tests. Based on these definitions, it independently derives formulas for the Power of Exclusion (PE) and Parentage Index (PI) for each test, including PI formulas for cases conforming to genetic laws and those involving mutations. Finally, the PE formula is validated, and the effectiveness of the PI formula is statistically analyzed.
[0070] Based on the above research and understanding, this application creatively develops a non-invasive prenatal paternity testing and analysis method, including a SNP locus information acquisition step, an exclusion probability calculation step, a paternity index calculation step, and a judgment step. The SNP locus information acquisition step includes acquiring SNP locus information from the pregnant mother, fetus, and the sample to be tested. The SNP locus information includes whether the SNP locus is homozygous or heterozygous, the frequency of each allele at the SNP locus, and the number of alleles at the SNP locus. The sample to be tested is either the suspected father's sample or the suspected mother's sample. The exclusion probability calculation step includes... The content involves screening SNP loci and calculating the exclusion probability (PE) based on the selected SNP loci using the corresponding formula. The identification content includes prenatal maternal-paternity identification, prenatal low-degree maternal-paternity identification, prenatal paternity-child identification, prenatal mother-child identification, and prenatal paternity-maternity identification. Prenatal maternal-paternity identification involves identifying whether the tested sample (the suspected father) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother. The exclusion probability for a sample that is not the biological father is calculated by screening SNP loci for which the pregnant mother's genotype is homozygous (AA) and calculating the exclusion probability (PE) of the suspected father sample according to Formula 1. PAFPrenatal low-level maternal-paternity identification refers to the identification of whether the test sample (i.e., the suspected father sample) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother and the fetal concentration is low. The exclusion probability of the test sample being the non-biological father is calculated by screening for SNP loci with genotypes of AA×AB in the pregnant mother × fetus, and then calculating the exclusion probability PE of the suspected father sample according to Formula 2. PLAF Prenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the sample to be tested (i.e., the suspected father) is the biological father of the fetus. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother, to determine whether the sample to be tested (i.e., the suspected mother) is the biological mother of the fetus. The exclusion probability of the sample being a non-biological mother is calculated by screening for SNP loci with the genotype of AA×AA between the pregnant mother and the fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMC Prenatal paternity testing is used when the biological father of the fetus is known, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM In formulas 1 to 5, P i Let n be the frequency of the i-th allele at the SNP locus, and n be the number of alleles at the SNP locus. The paternity index calculation steps include: deriving the probability X of the test sample being the biological father or mother of the fetus based on Mendel's laws of inheritance at the SNP locus; deriving the probability Y of the random sample being the biological father or mother of the fetus; and using the ratio of X to Y, i.e., X / Y, as the paternity index PI. The judgment steps include: calculating the cumulative exclusion probability CPE based on the exclusion probability PE of each SNP locus; calculating the cumulative paternity index CPI based on the paternity index PI of each SNP locus; and determining the kinship based on CPE and CPI.
[0071] Those skilled in the art will understand that all or part of the functions of the methods described above can be implemented in hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. Alternatively, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or portable hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0072] Therefore, based on the non-invasive prenatal paternity testing analysis method of this application, this application proposes a non-invasive prenatal paternity testing analysis device, including an SNP locus information acquisition module, an exclusion probability calculation module, a paternity index calculation module, and a judgment module; the SNP locus information acquisition module includes a module for acquiring SNP locus information of the pregnant mother, fetus, and the sample to be tested, the SNP locus information including whether the SNP locus is homozygous or heterozygous, the frequency of each allele of the SNP locus, and the number of alleles of the SNP locus; the sample to be tested is a sample of the suspected father or the suspected mother; the exclusion probability calculation module includes a module for acquiring SNP locus information of the pregnant mother, fetus, and the sample to be tested; the exclusion probability calculation module includes a module for acquiring SNP locus information of the pregnant mother, fetus, and the sample to be tested; the exclusion probability calculation module includes a module for acquiring SNP locus information of the pregnant mother, fetus, and sample to be tested, ... SNP loci are screened based on the identification criteria, and the exclusion probability (PE) is calculated using the corresponding formula based on the screened SNP loci. The identification criteria include prenatal maternal-suspected father identification, prenatal low-degree maternal-suspected father identification, prenatal father-child identification, prenatal mother-child identification, and prenatal father-suspected mother identification. Prenatal maternal-suspected father identification is used when the pregnant mother is known to be the biological mother of the fetus, and the test sample (i.e., the suspected father sample) is used to identify whether the test sample is not the biological father. The exclusion probability for the test sample is calculated by screening SNP loci with homozygous AA genotype of the pregnant mother and calculating the exclusion probability (PE) of the suspected father sample according to Formula 1. PAF Prenatal low-level maternal-paternity identification refers to the identification of whether the test sample (i.e., the suspected father sample) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother and the fetal concentration is low. The exclusion probability of the test sample being the non-biological father is calculated by screening for SNP loci with genotypes of AA×AB in the pregnant mother × fetus, and then calculating the exclusion probability PE of the suspected father sample according to Formula 2. PLAFPrenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the sample to be tested (i.e., the suspected father) is the biological father of the fetus. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother, to determine whether the sample to be tested (i.e., the suspected mother) is the biological mother of the fetus. The exclusion probability of the sample being a non-biological mother is calculated by screening for SNP loci with the genotype of AA×AA between the pregnant mother and the fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMC Prenatal paternity testing is used when the biological father of the fetus is known, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM In formulas 1 to 5, P i Let n be the frequency of the i-th allele at the SNP locus, and n be the number of alleles at the SNP locus. The paternity index calculation module includes a module for deriving the probability X of a test sample being the biological father or mother of the fetus based on the SNP locus and Mendel's laws of inheritance, and deriving the probability Y of a random sample being the biological father or mother of the fetus. The ratio of X to Y, i.e., X / Y, is used as the paternity index PI. The judgment module includes a module for calculating the cumulative exclusion probability CPE based on the exclusion probability PE of each SNP locus, calculating the cumulative paternity index CPI based on the paternity index PI of each SNP locus, and determining the kinship based on CPE and CPI.
[0073] Furthermore, this application also provides an apparatus comprising a memory and a processor; the memory includes a program for storing data; the processor includes a method for executing the program stored in the memory to implement the following steps: an SNP locus information acquisition step, comprising acquiring SNP locus information of the pregnant mother, fetus, and sample to be tested, wherein the SNP locus information includes whether the SNP locus is homozygous or heterozygous, the frequency of each allele at the SNP locus, and the number of alleles at the SNP locus; wherein the sample to be tested is a suspected paternal or suspected maternal sample; and an exclusion probability calculation step, comprising calculating the exclusion probability based on the identification content. SNP loci are screened, and the exclusion probability (PE) is calculated using the corresponding formula based on the screened SNP loci. The identification process includes prenatal maternal-paternity identification, prenatal low-degree maternal-paternity identification, prenatal paternity-child identification, prenatal mother-child identification, and prenatal paternity-maternity identification. Prenatal maternal-paternity identification involves identifying whether the tested sample (the suspected father) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother. The exclusion probability for a sample that is not the biological father is calculated by screening SNP loci with homozygous AA genotypes of the pregnant mother and calculating the exclusion probability (PE) of the suspected father sample according to Formula 1. PAF Prenatal low-level maternal-paternity identification refers to the identification of whether the test sample (i.e., the suspected father sample) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother and the fetal concentration is low. The exclusion probability of the test sample being the non-biological father is calculated by screening for SNP loci with genotypes of AA×AB in the pregnant mother × fetus, and then calculating the exclusion probability PE of the suspected father sample according to Formula 2. PLAF Prenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the sample to be tested (i.e., the suspected father) is the biological father of the fetus. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother, to determine whether the sample to be tested (i.e., the suspected mother) is the biological mother of the fetus. The exclusion probability of the sample being a non-biological mother is calculated by screening for SNP loci with the genotype of AA×AA between the pregnant mother and the fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMC Prenatal paternity testing is used when the biological father of the fetus is known, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM In formulas 1 to 5, P iLet n be the frequency of the i-th allele at the SNP locus, and n be the number of alleles at the SNP locus. The paternity index calculation steps include: deriving the probability X of the test sample being the biological father or mother of the fetus based on Mendel's laws of inheritance at the SNP locus; deriving the probability Y of the random sample being the biological father or mother of the fetus; and using the ratio of X to Y, i.e., X / Y, as the paternity index PI. The judgment steps include: calculating the cumulative exclusion probability CPE based on the exclusion probability PE of each SNP locus; calculating the cumulative paternity index CPI based on the paternity index PI of each SNP locus; and determining the kinship based on CPE and CPI.
[0074] This application also provides a computer-readable storage medium including a program that can be executed by a processor to implement the following method: a SNP locus information acquisition step, including acquiring SNP locus information of the pregnant mother, fetus, and sample to be tested, wherein the SNP locus information includes whether the SNP locus is homozygous or heterozygous, the frequency of each allele of the SNP locus, and the number of alleles of the SNP locus; wherein the sample to be tested is a suspected father sample or a suspected mother sample; an exclusion probability calculation step, including screening SNP loci according to the identification content, and calculating the exclusion probability PE according to the screened SNP loci using the corresponding formula; the identification content includes prenatal maternal suspected father identification, prenatal low-degree maternal suspected father identification, prenatal father-child identification, prenatal mother-child identification, and prenatal paternal suspected mother identification; prenatal maternal suspected father identification is an identification of whether the sample to be tested, i.e., the suspected father sample, is the biological father of the fetus when the pregnant mother is known to be the biological mother; the exclusion probability calculation method for the sample to be tested being a non-biological father is to screen SNP loci with the pregnant mother's genotype being homozygous AA, and calculate the exclusion probability PE of the suspected father sample according to Formula 1. PAF Prenatal low-level maternal-paternity identification refers to the identification of whether the test sample (i.e., the suspected father sample) is the biological father of the fetus, given that the pregnant mother is known to be the biological mother and the fetal concentration is low. The exclusion probability of the test sample being the non-biological father is calculated by screening for SNP loci with genotypes of AA×AB in the pregnant mother × fetus, and then calculating the exclusion probability PE of the suspected father sample according to Formula 2. PLAF Prenatal paternity testing is used when the relationship between the pregnant mother and fetus is unknown. It involves determining whether the sample to be tested (i.e., the suspected father) is the biological father of the fetus. The probability of exclusion for a sample that is not the biological father is calculated by screening for SNPs with genotypes of AA×AA between the pregnant mother and fetus, and then calculating the probability of exclusion (PE) of the suspected father sample according to Formula 3. PFC Prenatal maternal-fetal identification is used when the known pregnant mother is not the biological mother, to determine whether the sample to be tested (i.e., the suspected mother) is the biological mother of the fetus. The exclusion probability of the sample being a non-biological mother is calculated by screening for SNP loci with the genotype of AA×AA between the pregnant mother and the fetus, and then calculating the exclusion probability PE of the suspected mother sample according to Formula 4. PMCPrenatal paternity testing is used when the biological father of the fetus is known, but the relationship between the pregnant mother and the fetus is unknown. The test determines whether the pregnant mother is the biological mother of the fetus. The probability of exclusion (if the pregnant mother is not the biological mother) is calculated by screening for SNP loci with genotypes of AA×AA in the sample of the pregnant mother × fetus, or AA×AB (or BB)×AA in the sample of the pregnant mother × fetus × biological father. The probability of exclusion (PE) for the pregnant mother is then calculated according to Formula 5. PAM In formulas 1 to 5, P i Let n be the frequency of the i-th allele at the SNP locus, and n be the number of alleles at the SNP locus. The paternity index calculation steps include: deriving the probability X of the test sample being the biological father or mother of the fetus based on Mendel's laws of inheritance at the SNP locus; deriving the probability Y of the random sample being the biological father or mother of the fetus; and using the ratio of X to Y, i.e., X / Y, as the paternity index PI. The judgment steps include: calculating the cumulative exclusion probability CPE based on the exclusion probability PE of each SNP locus; calculating the cumulative paternity index CPI based on the paternity index PI of each SNP locus; and determining the kinship based on CPE and CPI.
[0075] The basic principle of the data analysis method for non-invasive prenatal paternity testing in this application is as follows:
[0076] NIPPT performs prenatal paternity testing by analyzing cfDNA in the blood of pregnant women. First, cfDNA is extracted from the pregnant woman's peripheral blood. Fragments of cfDNA originate from the placenta and fetus and carry fetal genetic information. Next, a cfDNA library is constructed, and the DNA fragments are appropriately processed and amplified. High-throughput sequencing is then used to obtain a large amount of DNA sequence data. Finally, the sequencing data undergoes bioinformatics analysis and alignment. The sequences are compared with a reference genome to identify fetal SNP information, which is then compared with the SNP genotypes of the mother and father to determine paternity.
[0077] The basic framework of the data analysis method for non-invasive prenatal paternity testing in this application is as follows:
[0078] 1) Cell-free DNA (cfDNA): DNA molecules released from cells and existing in the extracellular environment.
[0079] 2) Cell-free fetal DNA (cffDNA): Fetal DNA molecules present in the peripheral blood of pregnant women.
[0080] 3) Fetal concentration: The proportion of fetal cell-free DNA in the peripheral blood of pregnant women to the total cell-free DNA of pregnant women.
[0081] 4) Routine samples: Test samples with a fetal cfDNA concentration of not less than 4%.
[0082] 5) Low fetal concentration samples: Test samples with a fetal cfDNA concentration of less than 4%.
[0083] 6) Pregnant mother: A pregnant woman whose relationship with the fetus is unknown. She may be the biological mother or a non-biological mother.
[0084] Let the allele frequencies of genes A and B at SNP loci be a and b, respectively. Let n be the number of alleles at a given SNP locus. Let i and j be any values from 1 to 2, and i and j are mutually exclusive. Let P... i The frequency of the i-th allele at this locus.
[0085] The mutation setting of the data analysis method for non-invasive prenatal paternity testing in this application:
[0086] Let P(W→Z) be the probability that gene W at a certain SNP locus mutates into gene Z in a single mutation. SNP mutations have the following characteristics: First, SNP mutations are single-point mutations, unlike STR mutations which have upward and downward mutations; second, because the probability of a true SNP mutation is less than that from PCR and sequencing errors, there is no distinction between paternal and maternal mutations as with STR mutations; third, SNP mutations are extremely low-probability events, so homologous mutations only consider mutations in one gene. Therefore, P(W→Z) can be directly expressed as the SNP mutation rate μ. In practical applications, the standard "SF / Z JD0105003—2015 Forensic SNP Typing and Application Specification" is usually referenced, and μ is set to 0.00001.
[0087] Definitions of each identification in the data analysis method for non-invasive prenatal paternity testing in this application:
[0088] 1) Prenatal parentage testing of alleged father with biological mother (PAF): This test determines whether the alleged father is the biological father of the fetus, given that the pregnant mother is the biological mother of the fetus.
[0089] 2) Prenatal parentage testing of low fetal concentration alleged father (PLAF): This test is conducted in cases where the pregnant mother is known to be the biological mother of the fetus and the fetal concentration is low, to determine whether the alleged father is the biological father of the fetus.
[0090] 3) Prenatal parentage testing of Father-Child (PFC): This test determines whether the alleged father is the biological father of the fetus when the relationship between the pregnant mother and the fetus is unknown.
[0091] 4) Prenatal parentage testing of mother-child (PMC): This is a process of determining whether the suspected mother is the biological mother of the fetus when the pregnant mother is known to be a non-biological mother.
[0092] 5) Prenatal parentage testing of alleged mother with biological father (PAM): The biological father of the fetus is known, but the relationship between the pregnant mother and the fetus is unknown. This test determines whether the pregnant mother is the biological mother of the fetus.
[0093] The present application will be further described in detail below through specific embodiments. The following embodiments are only for further illustration of the present application and should not be construed as limiting the present application.
[0094] Example
[0095] 1. Sample Source
[0096] High-throughput sequencing data were obtained from the 1000-person database. Let the allele frequencies of genes A and B at SNP loci be a and b, respectively. Let n be the number of alleles at a given SNP locus. Let i and j be any values from 1 to 2, and i and j are mutually exclusive. Let P... i Let μ be the frequency of the i-th allele at this locus. μ = 0.00001.
[0097] 2. Derivation of the probability formula and parentage index.
[0098] 2.1 Derivation of the Formula for Prenatal Paternity Testing (PAF)
[0099] Judgment rule: Only SNP loci where the mother's genotype is homozygous (AA) will be analyzed. Since the extracted cfDNA is a mixture of the mother's cfDNA and the fetal cffDNA, the mother's SNPs are heterozygous loci, and the true fetal genotype cannot be determined at these loci. In this case, only SNP loci where the mother's genotype is homozygous will be analyzed.
[0100] (1)PE formula
[0101] Prenatal paternity testing to rule out non-paternity (PE) rate PAFThe probability that the genotypes of the alleged father, biological mother (i.e., the surrogate mother), and fetus do not show a genetic relationship at a single SNP locus that meets the criteria. The combined probability = maternal probability × child probability × paternal probability, with 6 possible scenarios:
[0102] 1) Biological mother A i A i ×Fetal A i A i ×Suspicious Father A i A i Mother probability = Child probability = Parent probability = P i 2 The probability of merging is 1 = P i 6 .
[0103] 2) Biological mother A i A i ×Fetal A i A i ×Suspicious Father A i A j Mother probability = Child probability = P i 2 The probability of the parent is 2P. i (1-P i The probability of merging is 2 = 2P. i 5 (1-P i ).
[0104] 3) Biological mother A i A i ×Fetal A i A i ×Suspicious Father A j A j Mother probability = Child probability = P i 2 Parent probability = (1-P) i ) 2 The probability of merging is 3 = P i 4 (1-P i ) 2 .
[0105] 4) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A j A j Mother probability = P i 2 Subprobability = 2P i (1-P i ), Parent probability = (1-P i ) 2The probability of merging is 4 = 2P i 3 (1-P i ) 3 .
[0106] 5) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A i A j Mother probability = P i 2 Child probability = Parent probability = 2P i (1-P i The probability of merging is 5 = 4P. i 4 (1-P i ) 2 .
[0107] 6) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A i A i Mother probability = Father probability = P i 2 Subprobability = 2P i (1-P i The probability of merging is 6 = 2P. i 5 (1-P i ).
[0108] Merging probability X = Merging probability 3 + Merging probability 6 = P i 4 (1-P i 2 )
[0109] Merging probability Y = Merging probability 1 + Merging probability 2 + Merging probability 3 + Merging probability 4 + Merging probability 5 + Merging probability 6 = 2P i 3 -P i 4
[0110] Therefore PE PAF The formula is Formula 1:
[0111] Formula 1:
[0112] The merging probability X represents all possibilities that can be used for analysis, the merging probability Y represents the possibility of excluding biological children, and PE PAFThe formula is the ratio of the sum of the merging probabilities X of all alleles at the SNP locus to the sum of the merging probabilities Y, and the same applies below.
[0113] (2) PI formula
[0114] Prenatal paternity index (PI) for suspected paternity testing of the mother. PAF PI: The ratio of the probability (X) that the suspected father is the biological father of the fetus to the probability (Y) that the random father is the biological father of the fetus at a single SNP locus that meets the determination rules. PIs conforming to both genetic laws and mutation scenarios. PAF The formula is shown in Table 1.
[0115] Table 1
[0116] Maternal genotype fetal genotype Genotype of suspected father sample X Y <![CDATA[PI PAF ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a 1 / 2a AA AA BB μ a μ / a AA AB BB 1 / 2 b / 2 1 / b AA AB AB 1 / 4 b / 2 1 / 2b AA AB AA μ b / 2 2μ / b
[0117] 2.2 Derivation of the Formula for Prenatal Low-Grade Plausible Father Identification (PLAF)
[0118] Judgment rule: Only SNP sites with a maternal × fetal genotype of AA × AB are analyzed. This is because when fetal genotype concentrations are too low, it is possible that non-maternal alleles from the fetus may not be detected, leading to analytical errors. Therefore, in this case, SNP sites with homozygous fetal genotypes are not analyzed.
[0119] (1)PE formula
[0120] Prenatal low fetal concentration of maternal suspected father identification non-paternity exclusion rate (PE) PLAF The probability that the genotypes of the alleged father, biological mother, and fetus do not show a genetic relationship at a single SNP locus that meets the criteria. The combined probability = maternal probability × child probability × paternal probability, with three possible scenarios:
[0121] 1) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A j A j Mother probability = P i 2 Subprobability = 2P i (1-P i ), Parent probability = (1-P i ) 2 The probability of merging is 1 = 2P i 3 (1-P i ) 3 .
[0122] 2) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A iA j Mother probability = P i 2 Child probability = Parent probability = 2P i (1-P i The probability of merging is 2 = 4P. i 4 (1-P i ) 2 .
[0123] 3) Biological mother A i A i ×Fetal A i A j ×Suspicious Father A i A i Mother probability = Father probability = P i 2 Subprobability = 2P i (1-P i The probability of merging is 3 = 2P. i 5 (1-P i ).
[0124] Merging probability X = Merging probability 3 = 2P i 5 (1-P i )
[0125] Merging probability Y = Merging probability 1 + Merging probability 2 + Merging probability 3 = 2P i 3 (1-P i )
[0126] Therefore PE PLAF Formula 2:
[0127] Formula 2:
[0128] (2) PI formula
[0129] Low prenatal fetal concentration of the paternity index (PI) for suspected paternity testing PLAF PI: The ratio of the probability (X) that the suspected father is the biological father of the fetus to the probability (Y) that the random father is the biological father of the fetus at a single SNP locus that meets the determination rules. PIs conforming to both genetic laws and mutation scenarios. PLAF The formula is shown in Table 2.
[0130] Table 2
[0131] Maternal genotype fetal genotype Genotype of suspected father sample X Y <![CDATA[PI PLAF ]]> AA AB BB 1 / 2 b / 2 1 / b AA AB AB 1 / 4 b / 2 1 / 2b AA AB AA μ b / 2 2μ / b
[0132] 2.3 Derivation of the Formula for Prenatal Paternity Testing (PFC)
[0133] Judgment rule: Only SNP loci with genotypes of AA×AA between the pregnant mother and fetus are considered. High fetal concentrations are required in this case, and testing after 10 weeks of gestation is recommended.
[0134] (1)PE formula
[0135] Prenatal paternity testing non-father exclusion rate (PE) PFC The probability that the genotypes of the alleged father and fetus do not show a genetic relationship at a single SNP locus that meets the criteria. The probability of merging is calculated as: probabilities of offspring × probabilities of father, with three possible scenarios:
[0136] 1) Fetal A i A i ×Suspicious Father A i A i Child probability = Parent probability = P i 2 The probability of merging is 1 = P i 2 ×P i 2 =P i 4 .
[0137] 2) Fetal A i A i ×Suspicious Father A i A j Subprobability = P i 2 The probability of the parent is 2P. i (1-P i The probability of merging is 2 = 2P. i 3 (1-P i ).
[0138] 3) Fetal A i A i ×Suspicious Father A j A j Subprobability = P i 2 Parent probability = (1-P) i ) 2 The probability of merging is 3 = P i 2 (1-P i ) 2 .
[0139] Merging probability X = Merging probability 3 = P i 2 (1-P i ) 2
[0140] Merging probability Y = Merging probability 1 + Merging probability 2 + Merging probability 3 = Pi 2
[0141] Therefore PE PFC Formula 3:
[0142] Formula 3:
[0143] (2) PI formula
[0144] Prenatal paternity index (PI) PFC PI: The ratio of the probability (X) that the suspected father is the biological father of the fetus to the probability (Y) that the random father is the biological father of the fetus. PI is calculated for cases conforming to genetic laws and mutation scenarios. PFC The formula is shown in Table 3.
[0145] Table 3
[0146] Maternal genotype fetal genotype Genotype of suspected father sample X Y <![CDATA[PI PFC ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a 1 / 2a AA AA BB μ a μ / a
[0147] 2.4 Derivation of the Formula for Prenatal Maternal-Fetal Identification (PMC)
[0148] Judgment rules: Only SNP sites with genotype AA×AA between non-biological mothers and fetuses are judged, requiring high fetal concentration, similar to PFC identification.
[0149] (1)PE formula
[0150] Prenatal maternal-fetal fertilization exclusion rate (PE) PMC The probability that the genotypes of the suspected mother and the fetus do not show a genetic relationship at a single SNP locus that meets the criteria. The probability of merging is calculated as: Probability of suspected mother × Probability of offspring, with three possible scenarios:
[0151] 1) Fetal A i A i ×suspicious mother A i A i Probability of the child = Probability of the doubtful mother = P i 2 The probability of merging is 1 = P i 2 ×P i 2 =P i 4 .
[0152] 2) Fetal A i A i ×suspicious mother A i A j Subprobability = P i 2 The probability of doubting the mother is 2P. i (1-P i The probability of merging is 2 = 2P. i3 (1-P i ).
[0153] 3) Fetal A i A i ×suspicious mother A j A j Subprobability = P i 2 The probability of doubting the mother is (1-P). i ) 2 The probability of merging is 3 = P i 2 (1-P i ) 2 .
[0154] Merging probability X = Merging probability 3 = P i 2 (1-P i ) 2
[0155] Merging probability Y = Merging probability 1 + Merging probability 2 + Merging probability 3 = P i 2
[0156] Therefore PE PMC Formula 4 is as follows:
[0157] Formula 4:
[0158] As can be seen, Formula 3 and Formula 4 are exactly the same, that is, PE PMC =PE PFC .
[0159] (2) PI formula
[0160] Prenatal maternal paternity index (PI) PMC PI: The ratio of the probability (X) that the suspected mother is the biological mother of the fetus to the probability (Y) that the random mother is the biological mother of the fetus at a single SNP locus that meets the determination rules. PI conforms to both genetic and mutation scenarios. PMC The formula is shown in Table 4.
[0161] Table 4
[0162] Maternal genotype fetal genotype suspected maternal sample genotype X Y <![CDATA[PI PMC ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a 1 / 2a AA AA BB μ a μ / a
[0163] 2.5 Derivation of the Formula for Prenatal Paternity Testing (PAM)
[0164] Judgment rules: The SNP locus is selected based on the genotype of the mother × fetus being AA×AA, or the genotype of the mother × fetus × father being AA×AB (or BB)×AA. When the biological mother cannot be determined, only SNPs where both the fetus and mother are homozygous can confirm the genotype. However, when the fetus is not the mother's biological child, detecting an AB genotype does not definitively determine whether the fetal genotype is AB or BB; only the probability of either AB or BB can be given. Therefore, SNP selection can also target SNP loci with a genotype of AA×AB (or BB)×AA for the mother × fetus × father. In this case, a high fetal concentration is required, and testing after 10 weeks of gestation is recommended.
[0165] (1)PE formula
[0166] Prenatal paternity testing exclusion rate (PE) PAM The probability that the genotypes of the biological father, pregnant mother, and fetus do not show a genetic relationship at a single SNP locus that meets the criteria. Combined probability = maternal probability × child probability × paternal probability. Five scenarios are considered:
[0167] 1) Pregnant mother A i A i ×Fetal A i A i ×Biological Father A i A i Mother probability = Child probability = Parent probability = P i 2 The probability of merging is 1 = P i 2 ×P i 2 ×P i 2 =P i 6 .
[0168] 2) Pregnant mother A i A i ×Fetal A i A i ×Biological Father A i A j Mother probability = Child probability = P i 2 The probability of the parent is 2P. i (1-P i The probability of merging is 2 = P. i 2 ×P i 2 ×2P i (1-P i ) = 2P i 5 (1-P i ).
[0169] 3) Pregnant mother A i A i ×Fetal A i A i ×Biological Father A j A j Mother probability = Child probability = P i 2, Parent probability = (1-P) i ) 2 The probability of merging is 3 = P i 2 ×P i 2 ×(1-P i ) 2 =P i 4 (1-P i ) 2 .
[0170] 4) Pregnant mother A i A i ×Fetal A i A j ×Biological Father A i A i Mother probability = Father probability = P i 2, Subprobability = 2P i (1-P i The probability of merging is 4 = 2P. i 5 (1-P i ).
[0171] 5) Pregnant mother A i A i ×Fetal A j A j ×Biological Father A i A i Mother probability = Father probability = P i 2 Subprobability = (1-P) i ) 2 The probability of merging is 5 = P i 4 (1-P i ) 2 .
[0172] Merging probability X = Merging probability 3 + Merging probability 4 + Merging probability 5 = P i 4 (1-P i ) 2 +2P i 5 (1-P i )+P i4 (1-P i ) 2 =2P i 4 (1-P i )
[0173] Merging probability Y = Merging probability 1 + Merging probability 2 + Merging probability 3 + Merging probability 4 + Merging probability 5 = P i 6 +2P i 5 (1-P i )+P i 4 (1-P i ) 2 +2P i 5 (1-P i )+P i 4 (1-P i ) 2 =P i 4 (2-P i 2 )
[0174] Therefore PE PAM Formula 5:
[0175] Formula 5:
[0176] (2) PI formula
[0177] Prenatal paternity index (PI) for suspected maternity testing PAM PI: The ratio of the probability (X) that the suspected mother is the biological mother of the fetus to the probability (Y) that the random mother is the biological mother of the fetus at a single SNP locus that meets the determination rules. PI conforms to both genetic and mutation scenarios. PAM The formula is shown in Table 5.
[0178] Table 5
[0179] Maternal genotype fetal genotype Father's genotype X Y <![CDATA[PI PAM ]]> AA AA AA 1 a 1 / a AA AA AB 1 / 2 a / 2 1 / a AA AA BB μ μa 1 / a AA AB / BB AA <![CDATA[(2μa+μ 2 b) / (2a+b)]]> <![CDATA[(ab+μb 2 ) / (2a+b)]]> <![CDATA[(2μa+μ 2 b) / (ab+μb 2 )]]>
[0180] In Tables 1 to 5, X represents the probability that the sample to be tested is the biological father or mother of the fetus, Y represents the probability that the random sample is the biological father or mother of the fetus, μ represents the probability of a single gene mutation with a value of μ = 0.00001, a represents the frequency of allele A, and b represents the frequency of allele B.
[0181] In formulas 1 to 5, P i denoted as , where is the frequency of the i-th allele at the SNP locus, and n is the number of alleles at the SNP locus.
[0182] The exclusion probability (PE) and parentage index (PI) derived above can be used to select the appropriate formula based on the specific identification content, such as prenatal maternal-suspected-paternity identification, prenatal low-degree maternal-suspected-paternity identification, prenatal paternal-child identification, prenatal mother-child identification, or prenatal paternal-suspected-mother identification.
[0183] The results were interpreted by calculating the exclusion probability (PE) and paternity index (PI) for all screened SNP loci. The cumulative exclusion probability (CPE) was calculated based on the PE of each SNP locus, and the cumulative paternity index (CPI) was calculated based on the PI of each SNP locus. Phylogenetic relationships were then determined based on the CPE and CPI. The specific formulas for calculating CPE and CPI are as follows:
[0184]
[0185] In the above formulas for calculating CPE and CPI, PE i Let PI be the exclusion probability of the i-th SNP site. i denoted as the parentage index of the i-th SNP locus, and n is the number of SNP loci.
[0186] If CPI > 10000 and CPE > 0.9999, the child is determined to be biologically related; if CPI < 0.0001 and CPE > 0.9999, the child is determined to be non-biologically related; otherwise, it cannot be determined.
[0187] 3. PE Formula Verification
[0188] Based on the SNP gene frequencies in the Chinese population, 10 SNP loci were selected from the 1000-person database for comparative analysis. Specifically, the P values of these 10 SNP loci were analyzed. i The frequency information is shown in Table 7, and the comparison between the various PE formula values and the simulated values is also shown in Table 7.
[0189] Scatter plots were created using the formula and simulated values of PE, and a linear fit was performed between the two. Since PE... PMC =PE PFC Therefore, this example only applies to PE. PAF PE PLAF PE PFC and PE PAM Perform a linear fit. The results are as follows: Figures 1 to 4 As shown.
[0190] For PE PAF Plot a scatter plot of the formula values and simulated values, see... Figure 1 Linear regression analysis was performed using SPSS 19.0, and the R-value was obtained. 2 =1, P<0.001. The results show that the PE of each SNP is... PAFThe formula value and the simulation value are consistent, and Formula 1 is consistent with the simulation experiment verification.
[0191] For PE PLAF Plot a scatter plot of the formula values and simulated values, see... Figure 2 Linear regression analysis yields R0 2 =1, P<0.001. The results show that the PE of each SNP is... PLAF The formula value and the simulation value are consistent, and Formula 2 is consistent with the simulation experiment verification.
[0192] For PE PFC Plot a scatter plot of the formula values and simulated values, see... Figure 3 Linear regression analysis yields R0 2 =1, P<0.001. The results show that the PE of each SNP is... PFC The formula values and simulation results are consistent, and Formula 3 conforms to the simulation experiment verification. Formula 3 and Formula 4 are exactly the same, and similarly, Formula 4 also conforms to the simulation experiment verification.
[0193] For PE PAM Plot a scatter plot of the formula values and simulated values, see... Figure 4 Linear regression analysis yields R0 2 =1, P<0.001. The results show that the PE of each SNP is... PAM The formula value and the simulation value are consistent, and Formula 5 is consistent with the verification of the simulation experiment.
[0194] Table 7. Comparison of PE at 10 SNP sites
[0195]
[0196] 4. CPI Calculation and Statistics
[0197] Based on the 10 SNP loci identified in “3.PE Formula Verification”, the statistical data of the CPI for each identified locus are shown in Table 8, where the 95% confidence interval was calculated using SPSS 19.0.
[0198] Table 8. CPI statistics in the simulation experiment
[0199]
[0200]
[0201] In Table 8, m refers to the number of SNPs that can be used for analysis, the CPI of the related group is the CPI of the biological group, the 95% CI refers to the 95% confidence interval, and the irrelevant group is the non-biological group.
[0202] The results in Table 8 show that the CPI calculated from all identifications can be used to determine kinship according to the "Technical Specifications for Paternity Testing". When m = 100, all identifications can provide either a supporting or excluding conclusion. For PLAFs with m ≥ 50, definite conclusions can be given for more than 75.90% of related groups and 100% of unrelated individuals, while more than 24.10% of related groups require additional testing of more SNP loci.
[0203] 5. Clinical sample testing
[0204] Following the above methods, this case involved testing 313 samples, including 89 samples for prenatal mother-to-father identification, 87 samples for low-degree mother-to-father identification, 91 samples for prenatal father-to-child identification, 29 samples for prenatal mother-to-child identification, and 17 samples for prenatal father-to-mother identification. The collection and testing time for prenatal mother-to-father identification samples was 6-10 weeks, for low-degree mother-to-father identification samples it was 4-6 weeks, and for prenatal father-to-child, mother-to-child, and father-to-mother identification samples it was 10-12 weeks. The tested samples were followed up, and after the baby's birth, samples were collected for routine paternity testing to verify the accuracy of the method used in this case.
[0205] The results showed that the test results of 313 samples using the method in this example were consistent with the actual situation, indicating that the test method in this example can accurately identify the kinship in five situations.
[0206] The application of NIPPT in public security and forensic identification fields needs to address several key questions: Who is the father of the fetus? Is he / she a non-biological mother? If not, who is the biological mother? This case study combines first-generation sequencing paternity testing algorithms with the practical application of second-generation sequencing in NIPPT to establish a complete NIPPT algorithm system based on SNP genetic markers. This system can specifically answer five types of identification: First, when the pregnant mother is known to be the biological mother, identify the father of the fetus, including routine samples of the pregnant mother (PAF) and low fetal concentration samples (PLAF); second, when the relationship between the pregnant mother and the fetus is unknown, identify the father of the fetus (PFC); third, when the pregnant mother is known to be a non-biological mother, identify the biological mother of the fetus (PMC); fourth, when the biological father of the fetus is known, identify whether the pregnant mother is the biological mother of the fetus (PAM); fifth, when the relationship between the pregnant mother, the alleged father, and the fetus is unknown, PFC and PAM can be applied sequentially to identify whether the pregnant mother is the biological mother of the fetus. Based on Mendelian inheritance laws, algorithms for STR paternity testing that conform to inheritance laws and mutation scenarios, and the special patterns of SNP mutations, this example derives the complete set of PE and PI for the above five types of NIPPT identification. The CPI value is used to determine kinship, which has high forensic application value.
[0207] STR paternity testing requires calculating all genetic markers, while NIPPT only calculates some SNP genotypes; this is the biggest difference between the two algorithms. This is because prenatal cell-free DNA is a mixture of maternal and fetal DNA. In this case, heterozygous SNPs from the mother cannot determine the fetal genotype. In this example, SNP loci were selected based on an allele frequency greater than 0.3 in the subpopulation. According to the characteristics of SNP polymorphism, the algorithm model in this example only considers SNP loci with biselequities. If ternary SNP loci are used, or rare SNP genes with extremely low probability of occurrence in this case, they are not calculated. Normally, the calculable SNP genotypes are sufficient for identification requirements.
[0208] Of the five types of paternity testing, PAF, PLAF, PFC, and PMC correspond to the triad (i.e., mother-almost-father) and diad (almost-father) tests in STR paternity testing. The corresponding formula derivations are relatively simple, and most PI formulas can be directly applied to the STR paternity testing PI formulas. However, the genotyping combination for PAM (paternity-almost-father) is very special, especially the "pregnant mother AA - fetus AB / BB - alleged father AA" type. This is because when the fetus is not born to the pregnant mother, and the fetal genotype is detected as AB, it cannot be completely determined whether the fetal genotype is AB or BB; only the probability of either AB or BB can be given. This is because the genotype frequency of AB is 2ab, and the genotype frequency of BB is b. 2 Therefore, the frequency ratio of the two is 2a:b, and the PI formula is derived based on this ratio. Furthermore, directly deriving the PE formula using conventional methods is also difficult. This example starts from the definition of PE, deriving the merging probability X for those not related and the merging probability Y for all those involved in the calculation; the quotient of the two is the PE. Linear regression analysis of the formula verification experiment shows that the PE formulas for each identification completely conform to the simulation experiment verification.
[0209] According to the "Technical Specifications for Paternity Testing," the criteria for determining the identification opinion are as follows: when CPI < 0.0001, an exclusion opinion is given; when CPI > 10000, a supporting opinion is given; when CPI is between 0.0001 and 10000, further testing should be conducted. Regardless of whether the individuals being tested are related, it is necessary to determine whether the genotype of each genetic marker conforms to Mendelian inheritance laws, calculate PI separately using both the conventional formula and the mutation formula, and then combine and accumulate the two to calculate the CPI to give the identification opinion. NIPPT also follows the same criteria as STR testing for determining the identification opinion. Considering that NIPPT SNP mutations do not distinguish between upstream and downstream mutations, nor do they consider paternal and maternal lineage, it differs from the empirical decreasing mutation model of STR paternity testing. According to the statistical results of CPI calculation, when 100 SNP loci (m) are involved in the calculation, a definite identification conclusion can be given for all identified related groups and unrelated individuals. When the designed capture SNP loci are 5000, except for PLAF, other identification methods typically involve more than 500 SNP loci in the calculation, inevitably leading to a definitive conclusion. Even with PLAF, as long as 50 SNP loci are involved in the calculation, a relatively high proportion of cases can still yield a definitive conclusion. In our real-world data studies, we found that as long as the fetal concentration is greater than 2%, there will still be more than 50 loci available for analysis and judgment. When the fetal concentration is less than 2%, the number of detected loci is likely to be less than 50, making it impossible to draw a conclusion. Therefore, when identifying fetal concentrations not higher than 2%, UMI technology can be used. UMI tags can significantly reduce the sequencing error rate, thereby improving the detection sensitivity in early gestational weeks. However, it cannot be applied to samples with low fetal concentrations when encountering suspected mothers or non-biological mothers; instead, samples from over 10 weeks of gestation with fetal concentrations greater than 5% are required.
[0210] In summary, the non-invasive prenatal paternity testing method presented in this example can accurately and effectively identify paternity in various complex kinship relationships, providing accurate and effective scientific evidence for fields requiring strict theoretical basis, such as the judiciary, public security, and hospitals.
[0211] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.
Claims
1. A method of non-invasive prenatal paternity testing analysis, the method comprising: The method comprises a SNP site information obtaining step, an exclusion probability calculating step, a paternity index calculating step and a judging step. The SNP site information obtaining step comprises obtaining SNP site information of a pregnant mother, a fetus and a sample to be tested, wherein the SNP site information comprises whether a SNP site is homozygous or heterozygous, frequencies of each allele of the SNP site, and the number of alleles of the SNP site; and the sample to be tested is a putative father sample or a putative mother sample. The exclusion probability calculating step comprises screening SNP sites according to identification content, and calculating an exclusion probability PE according to the screened SNP sites by using a corresponding formula; and the identification content comprises prenatal paternity identification of a biological father, prenatal low-degree paternity identification of a biological father, prenatal paternity identification of a biological son, prenatal paternity identification of a biological daughter and prenatal paternity identification of a biological mother. The prenatal paternity identification is to identify whether the known pregnant mother is the mother of the fetus, and whether the sample to be tested, i.e., the suspected father sample, is the biological father of the fetus. The calculation method of the exclusion probability of the sample to be tested, which is a non-biological father, is as follows: screening SNP sites with homozygous AA genotype of the pregnant mother, and calculating the exclusion probability PE of the suspected father sample according to formula 1 PAF ; Formula 1: The prenatal low-degree paternal identification is to identify whether the sample to be tested, i.e., the suspected father sample, is the biological father of the fetus, and the sample is of a known pregnant mother and a fetus with low concentration. The calculation method of the exclusion probability of the sample to be tested, i.e., the suspected father sample, is that the SNP site of which the genotype of the pregnant mother and the fetus is AA and AB is screened, and the exclusion probability PE of the suspected father sample is calculated according to formula 2. PLAF ; Equation 2: The prenatal father-son identification is the identification of the relationship between an unknown pregnant mother and a fetus, and the identification of whether a test sample, i.e., a suspected father sample, is the biological father of the fetus. The calculation method of the exclusion probability of the test sample being a non-biological father is as follows: screening SNP sites with genotypes AA×AA of the pregnant mother and the fetus, and calculating the exclusion probability PE of the suspected father sample according to Formula 3 PFC ; Formula 3: The prenatal mother-child identification is to identify whether the sample to be tested, i.e., the suspected mother sample, is the biological mother of the fetus for the known pregnant mother who is a non-biological mother. The calculation method of the exclusion probability of the sample to be tested being the non-biological mother is as follows: screening SNP sites with genotypes AA x AA of the pregnant mother x fetus, and calculating the exclusion probability PE of the suspected mother sample according to formula 4 PMC ; Formula 4: The prenatal father doubtful mother identification is known as the father of the fetus, the relationship between the unknown pregnant mother and the fetus, the identification of whether the pregnant mother is the biological mother of the fetus, the calculation method of the exclusion probability of the pregnant mother not being the biological mother of the fetus is that the genotype of the pregnant mother and the fetus is AA x AA, or the genotype of the pregnant mother, the fetus and the father sample is AA x AB (or BB) x AA SNP site, and the exclusion probability PE of the pregnant mother is calculated according to formula 5 PAM ; Equation 5: In Formulas 1 to 5, P i is the frequency of the ith allele of the SNP locus, and n is the number of alleles of the SNP locus. The paternity index calculating step comprises deriving a probability X that the sample to be tested is a biological father or mother of the fetus according to the SNP sites by using Mendelian inheritance law, deriving a probability Y that a random sample is the biological father or mother of the fetus, and taking a ratio X / Y of X and Y as a paternity index PI. The judging step comprises calculating a cumulative exclusion probability CPE according to the exclusion probability PE of each SNP site, calculating a cumulative paternity index CPI according to the paternity index PI of each SNP site, and judging a kinship relationship according to the CPE and the CPI.
2. The method of claim 1, wherein: The paternity index calculating step further comprises calculating the paternity index PI in different ways according to the prenatal paternity identification of a biological father, the prenatal low-degree paternity identification of a biological father, the prenatal paternity identification of a biological son, the prenatal paternity identification of a biological daughter and the prenatal paternity identification of a biological mother in the exclusion probability calculating step. The prenatal paternity test is calculated according to Table 1 to obtain the paternity index PI PAF ; The prenatal low-paternity identification is calculated according to the paternity index PI in Table 2 PLAF ; The prenatal paternity is identified according to the paternity index PI calculated in table 3 PFC ; The prenatal parent-offspring identification calculates a paternity index PI according to Table 4 PMC ; The prenatal paternity test is calculated according to Table 5 to determine the paternity index PI PAM ; Table 1 Table 2 Table 3 Table 4 Table 5 In Tables 1 to 5, X is a probability that the sample to be tested is a biological father or mother of the fetus, Y is a probability that a random sample is the biological father or mother of the fetus, μ is a probability of a single mutation of a gene, a is a frequency of an allele A, and b is a frequency of an allele B. Preferably, μ is 0.00001. Preferably, the random sample is a sample randomly obtained from a thousand-person database. Preferably, in the judging step, wherein PE i is the probability of exclusion calculated for the ith SNP site, PI i is the probability of inclusion calculated for the ith SNP site, n is the number of SNP sites; Preferably, in the judging step, the kinship relationship is judged according to the CPE and the CPI, and specifically, when CPI>10000 and CPE>0.9999, it is determined to be a biological relationship; when CPI<0.0001 and CPE>0.9999, it is determined to be a non-biological relationship; and in other cases, it is determined to be indeterminable.
3. The method of claim 1, wherein: The sample with a low concentration of the fetus is a detection sample with a fetal cfDNA concentration lower than 4%, and the fetal cfDNA concentration is a proportion of fetal cfDNA in total cfDNA in peripheral blood of a pregnant woman.
4. The method according to any one of claims 1 to 3, characterized in that: In the judging step, the number of SNP sites used for detection is at least 100.
5. A device for non-invasive prenatal paternity testing analysis, characterized by: The method comprises a SNP site information obtaining module, an exclusion probability calculating module, a paternity index calculating module and a judging module. The SNP site information obtaining module comprises a module for obtaining SNP site information of a pregnant mother, a fetus and a sample to be tested, wherein the SNP site information comprises whether a SNP site is homozygous or heterozygous, frequencies of each allele of the SNP site, and the number of alleles of the SNP site; and the sample to be tested is a putative father sample or a putative mother sample. The exclusion probability calculation module comprises a module for screening SNP sites according to identification content, and a module for calculating the exclusion probability PE according to the screened SNP sites using a corresponding formula; the identification content comprises prenatal paternity identification, prenatal low-degree paternity identification, prenatal father-child identification, prenatal mother-child identification, and prenatal maternity identification; The prenatal paternity identification is to identify whether the known pregnant mother is the mother of the fetus, and whether the sample to be tested, i.e., the suspected father sample, is the biological father of the fetus. The calculation method of the exclusion probability of the sample to be tested, which is a non-biological father, is as follows: screening SNP sites with homozygous AA genotype of the pregnant mother, and calculating the exclusion probability PE of the suspected father sample according to formula 1 PAF ; Formula 1: The prenatal low-degree paternal identification is to identify whether the sample to be tested, i.e., the suspected father sample, is the biological father of the fetus, and the sample is of a known pregnant mother and a fetus with low concentration. The calculation method of the exclusion probability of the sample to be tested, i.e., the suspected father sample, is that the SNP site of which the genotype of the pregnant mother and the fetus is AA and AB is screened, and the exclusion probability PE of the suspected father sample is calculated according to formula 2. PLAF ; Equation 2: The prenatal father-son identification is the identification of the relationship between an unknown pregnant mother and a fetus, and the identification of whether a test sample, i.e., a suspected father sample, is the biological father of the fetus. The calculation method of the exclusion probability of the test sample being a non-biological father is as follows: screening SNP sites with genotypes AA×AA of the pregnant mother and the fetus, and calculating the exclusion probability PE of the suspected father sample according to Formula 3 PFC ; Equation 3: The prenatal mother-child identification is to identify whether the sample to be tested, i.e., the suspected mother sample, is the biological mother of the fetus for the known pregnant mother who is a non-biological mother. The calculation method of the exclusion probability of the sample to be tested being the non-biological mother is as follows: screening SNP sites with genotypes AA x AA of the pregnant mother x fetus, and calculating the exclusion probability PE of the suspected mother sample according to formula 4 PMC ; Equation 4: The prenatal paternity identification is known as the father of the fetus, the relationship between the unknown pregnant mother and the fetus, the identification of whether the pregnant mother is the biological mother of the fetus, the calculation method of the exclusion probability of the pregnant mother who is not the biological mother of the fetus is that the genotype of the pregnant mother and the fetus is AA x AA, or the genotype of the pregnant mother, the fetus and the father sample is AA x AB (or BB) x AA SNP site, and the exclusion probability PE of the pregnant mother is calculated according to formula 5 PAM ; Equation 5: In Formulas 1 to 5, P i is the frequency of the ith allele of the SNP locus, and n is the number of alleles of the SNP locus. The parentage index calculation module comprises a module for deriving the probability X that the sample to be tested is the biological father or mother of the fetus according to the SNP sites using Mendelian inheritance law, deriving the probability Y that a random sample is the biological father or mother of the fetus, and taking the ratio X / Y of X and Y as the parentage index PI; The judgment module comprises a module for calculating the cumulative exclusion probability CPE according to the exclusion probability PE of each SNP site, a module for calculating the cumulative parentage index CPI according to the parentage index PI of each SNP site, and a module for judging the kinship according to CPE and CPI.
6. The apparatus of claim 5, wherein: The parentage index calculation module further comprises a module for calculating the parentage index PI in five different ways according to prenatal paternity identification, prenatal low-degree paternity identification, prenatal father-child identification, prenatal mother-child identification, and prenatal maternity identification in the exclusion probability calculation module; The prenatal paternity test is calculated according to Table 1 to obtain the paternity index PI PAF ; The prenatal low-paternity identification is calculated according to the paternity index PI in Table 2 PLAF ; The prenatal paternity is identified according to the paternity index PI calculated in table 3 PFC ; The prenatal parentage is calculated according to Table 4 a paternity index PI PMC ; The prenatal paternity test is calculated according to Table 5 to determine the paternity index PI PAM ; Table 1 Table 2 Table 3 Table 4 Table 5 In Tables 1 to 5, X is the probability that the sample to be tested is the biological father or mother of the fetus, Y is the probability that a random sample is the biological father or mother of the fetus, μ is the probability of a single mutation of a gene, a is the frequency of allele A, and b is the frequency of allele B; Preferably, μ is 0.00001. Preferably, the random sample is a sample randomly obtained from a thousand-person database. Preferably, in the judging module, wherein PE i is the exclusion probability of the ith SNP site, PI i is the kinship index of the ith SNP site, and n is the number of SNP sites. Preferably, in the judgment module, the kinship is judged according to CPE and CPI, specifically comprising: CPI>10000 and CPE>0.9999, which is determined as biological; CPI<0.0001 and CPE>0.9999, which is determined as non-biological; and other cases, which are determined as indeterminate.
7. The apparatus of claim 5, wherein: The sample with low fetal concentration is a detection sample with fetal cfDNA concentration lower than 4%, and the fetal cfDNA concentration is the proportion of fetal cfDNA in the total cfDNA in the peripheral blood of the pregnant woman.
8. The apparatus of any one of claims 5-7, wherein: In the judgment module, the number of SNP sites detected is at least 100.
9. A device for non-invasive prenatal paternity testing analysis, characterized by: The device comprises a memory and a processor; The memory comprises a program for storing; The processor comprises a program for executing the program stored in the memory to realize the method of non-invasive prenatal parentage detection analysis according to any one of claims 1-4.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, which can be executed by the processor to realize the method of non-invasive prenatal parentage detection analysis according to any one of claims 1-4.