A method for parentage testing

By analyzing cell-free DNA samples from two pregnancies, calculating the proportion P of fetal signal sites, and comparing it with clustering models M(PT) and M(NPT), the problem of accurately determining the father-child relationship when samples of suspected fathers cannot be obtained is solved, and paternity testing without suspected fathers is achieved.

CN116312767BActive Publication Date: 2026-02-24SUZHOU SUYIN ZHIQI BIOTECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211404459.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-10
Publication Date
2026-02-24
Estimated Expiration
2042-11-10

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately determine the father-son relationship without obtaining a sample of the suspected father, especially when the child is working away from home or privacy issues prevent obtaining a sample of the suspected father.

Method used

By analyzing cell-free DNA samples from two pregnancies, the proportion of fetal signal sites P was calculated and compared with clustering models M(PT) and M(NPT) to determine whether the fetuses share the same father.

Benefits of technology

It can accurately determine whether a fetus belongs to the same father without requiring a sample from the alleged father, thus improving the feasibility and accuracy of paternity testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312767B_ABST
    Figure CN116312767B_ABST
Patent Text Reader

Abstract

The application discloses a method for parentage identification, and belongs to the technical field of parentage identification. S101: sequencing of the biallelic loci of the two free DNA samples of the pregnant woman during pregnancy respectively obtains DNA data S1 and S2; S102: respectively obtaining the typing locus set X(S1) and X(S2) containing the fetal signal in S1 and S2; S103: respectively calculating the DNA concentration p1 and p2 of the fetus according to S1 and S2; S104: calculating the common rate K according to formula II, K=[X(S1)∩X(S2)] / [X(S1)∪X(S2)](II); S105: judging the genetic relationship of the fetus during the two pregnancies according to the clustering model of K, p1 and p2. Without a doubtful father sample, whether the fetus is the same father can be judged by comparing the DNA data between the two independent pregnancies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of paternity testing technology, and specifically relates to a method for paternity testing. Background Technology

[0002] DNA paternity testing utilizes theories and techniques from forensic medicine, biology, and genetics to analyze genetic characteristics based on morphological or physiological similarities between parents and offspring, determining whether a parent and child are biologically related. The most common type of paternity testing is determining parentage based on the father's relationship. This usually requires samples from both the alleged father and the child or fetus. However, due to factors such as working away from home or privacy concerns, directly obtaining samples from the alleged father is often inconvenient. Summary of the Invention

[0003] Cell-free DNA in maternal peripheral blood is a mixed signal of maternal and fetal DNA. By calculation, the proportion P (K) of fetal signal sites shared between two independent pregnancies of the same pregnant woman can be determined. When the fetuses in the two independent pregnancies share the same father (i.e., they are full mothers and fathers), this proportion P can cluster and fit a model M (PT). When the fetuses in the two independent pregnancies do not share the same father (i.e., they are half-fathers), P can cluster and fit another model M (NPT). When there is a significant difference between the two models, this difference can be used to identify and determine the relationship between the fetuses in the two independent pregnancies. If the P values ​​obtained from samples S1 and S2 clearly cluster with the P values ​​of fetuses in the two independent pregnancies sharing the same father (i.e., they are full mothers and fathers), we consider the fetuses in samples S1 and S2 to be full mothers and fathers, meaning the fetuses share the same father. Conversely, if the fetuses in the two independent pregnancies are half-fathers, the fetuses do not share the same father. Therefore, embodiments of the present invention provide a method for paternity testing, the method comprising the following steps:

[0004] S101: DNA data S1 and S2 were obtained by sequencing the dimorphic sites of cell-free DNA samples from two pregnancies of a pregnant woman.

[0005] S102: Obtain the genotyping site sets X(S1) and X(S2) containing fetal signals in S1 and S2, respectively;

[0006] S103: The fetal DNA concentrations p1 and p2 are calculated based on S1 and S2, respectively;

[0007] S104: Calculate the common ownership rate K according to Formula II.

[0008] K=[X(S1)∩X(S2)] / [X(S1)∪X(S2)](Ⅱ);

[0009] S105: Determine whether the fetuses in the two pregnancies are of the same mother and father or the same mother but different fathers based on the correspondence between K, p1, and p2 and the clustering model of the simulated samples. The clustering model is obtained by clustering fetal concentration and K in steps S102 and S104 based on the fetal relationship between the two pregnancies of the pregnant woman and the simulated samples with clear fetal concentration.

[0010] Specifically, step S105 includes: obtaining clustering models M(PT) for fetal concentration and K under same-parent relationships and M(NPT) for fetal concentration and K under different-parent relationships; determining whether K, p1, and p2 conform to model M(PT) or M(NPT); if K, p1, and p2 conform to model M(PT), then the two fetuses are determined to be same-parent; if K, p1, and p2 conform to model M(NPT), then the two fetuses are determined to be different-parent; simulating cell-free DNA samples from a first pregnancy with different fetal concentrations, and simulating cell-free DNA samples from a second pregnancy with different fetal concentrations; some samples between the cell-free DNA samples from the first pregnancy and the cell-free DNA samples from the second pregnancy conform to the same-parent relationship, and the remaining samples conform to the different-parent relationship; clustering different fetal concentrations and corresponding K values ​​under the same-parent relationship to obtain clustering model M(PT), and clustering different fetal concentrations and corresponding K values ​​under the different-parent relationship to obtain clustering model M(NPT); p is 0-25%.

[0011] Specifically, M(NPT) and M(PT) are obtained through the following methods:

[0012] S201: Generate paternal DNA sample F1, maternal DNA sample M, and unrelated male DNA sample F2 randomly based on the frequency of the Chinese population; according to Mendel's laws of inheritance, offspring Z1 and Z2 are generated from F1 and M, and offspring Z3 is generated from F2 and M.

[0013] S202: Fetal concentrations of 0-25% were mixed in an isogradient manner to obtain simulated cell-free DNA samples of pregnant women, such as S1' (mixed with Z1 and M), S2' (mixed with Z2 and M), and S3' (mixed with Z3 and M). Multiple samples were generated for each fetal concentration, with a simulated sequencing depth of 100X-150X. Autosomal SNP loci with a mutation frequency between [0.05-0.95] on the sample genome were selected as genetic markers, and SNP genotyping was performed based on the simulated sequencing depth.

[0014] S203: Calculate the K value for each S1' and each S2' and the K value for each S1' and each S3' according to the method of steps S102 and S104;

[0015] S204: Under different simulated concentrations, the fetal concentrations of cluster S1' and S2' and their corresponding K are used to obtain model M(PT), and the fetal concentrations of cluster S1' and S3' and their corresponding K are used to obtain model M(NPT).

[0016] Further, for each simulated concentration of a simulated sample, the simulated concentration of another simulated sample is used as the x-axis and K as the y-axis to obtain the cluster correspondence map of p and K at each simulated concentration; the cluster correspondence maps of all simulated concentrations are drawn in the same coordinate system to form a clustering model, and each cluster correspondence map on it can be hidden; in step S105, only the cluster correspondence map corresponding to the simulated concentration of p1 or p2 is selected not to be hidden; in the unhidden cluster correspondence map, it is determined whether the simulated concentration corresponding to p2 or p1 and K conform to the M(NPT) or M(PT) clustering model.

[0017] Specifically, the gradient difference for equal gradients is 0.5% or 1%, 50 samples are generated for each fetal concentration, and more than 2,000 SNP sites are selected from the sample genome.

[0018] In step S105, p1 and p2 are rounded or rounded down to correspond to the simulated concentration in step S202.

[0019] Furthermore, in the unhidden cluster correspondence map, under the same simulated concentration, the K values ​​of multiple samples form vertical line regions; under the correspondence between p1 and p2 and p, it is determined whether K is located in the vertical line region.

[0020] The dimorphic loci are selected from SNP loci, and the population frequency of the dimorphic loci is 0.05-0.95.

[0021] In step S102, the set of genotyping sites containing fetal signals is obtained according to formula I.

[0022] X(S) = {X i |0 <na i (S) / n i (S)<0.2∪0 <nA i (S) / n i (S)<0.2}(I),

[0023] Where nA and na represent the observed values ​​of dimorphic sites A and a, respectively, and n = nA + na.

[0024] This invention provides a method for paternity testing that does not require a sample from the alleged father and can determine whether a fetus shares the same father by comparing DNA data from two separate pregnancies. Attached Figure Description

[0025] Figure 1This is a flowchart of the paternity testing method provided in an embodiment of the present invention;

[0026] Figure 2 This is a clustering correspondence diagram under a simulated concentration.

[0027] Figure 3 This is the clustering correspondence diagram at a concentration of 1.5% and under the same parent-child relationship;

[0028] Figure 4 yes Figure 3 A magnified view of a portion of the image;

[0029] Figure 5 This is the clustering correspondence diagram at a concentration of 4% and under the same parent-different parent relationship;

[0030] Figure 6 yes Figure 5 A magnified view of a portion of the image.

[0031] in, Figure 2 The lower part represents random relationships, the middle part represents half-father relationships, and the upper part represents full-father relationships. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0033] Example 1

[0034] See Figure 1 Example 1 provides a method for paternity testing, which includes the following steps:

[0035] S101: DNA data S1 and S2 were obtained by sequencing the dimorphic sites of cell-free DNA samples from two pregnancies of a pregnant woman.

[0036] S102: Obtain the genotyping site sets X(S1) and X(S2) containing fetal signals in S1 and S2 respectively according to Formula I;

[0037] X(S) = {X i |0 <na i (S) / n i (S)<0.2∪0 <nA i (S) / n i (S)<0.2}(I).

[0038] Specifically, X(S1) = {X i |0 <na i (S1) / n i (S1)<0.2∪0 <nA i (S1) / n i(S1)<0.2},X(S2)={X i |0 <na i (S2) / n i (S2)<0.2∪0 <nA i (S2) / n i (S2)<0.2}.

[0039] Where nA and na represent the observed values ​​of dimorphic sites A and a, respectively, and n = nA + na.

[0040] S103: The fetal DNA concentrations p1 and p2 are calculated based on S1 and S2, respectively. The calculation of fetal DNA concentration can be found using methods from application numbers CN2021111978446, CN2021111978145, CN2021111978357, CN2021111978465, etc., and can be calculated using only the pregnant woman's DNA data.

[0041] S104: Calculate the common ownership rate K according to Formula II.

[0042] K=[X(S1)∩X(S2)] / [X(S1)∪X(S2)](Ⅱ).

[0043] K represents the proportion of the number of fetal signal sites to the total number of fetal signal sites.

[0044] S105: Determine whether the fetuses in the two pregnancies are related as mother and father or mother and father but different fathers based on the correspondence between K, p1, and p2 and the clustering model of the simulated samples. The clustering model is obtained by clustering fetal concentration and K in steps S102 and S104, based on the fetal relationship between the two pregnancies and the simulated samples with clear fetal concentration (or by calculating a large number of actual samples, although obtaining such samples is more difficult).

[0045] Specifically, step S105 includes: obtaining a clustering model M(PT) of fetal concentration and K under same mother and father relationship and a clustering model M(NPT) of fetal concentration and K under same mother and different father relationship; determining whether K, p1, and p2 conform to model M(PT) or M(NPT); if K, p1, and p2 conform to model M(PT), then the two fetuses are determined to be same mother and father; if K, p1, and p2 conform to model M(NPT), then the two fetuses are determined to be different mother and father; simulating cell-free DNA simulation samples of first pregnancy with different fetal concentrations and simulating cell-free DNA simulation samples of second pregnancy with different fetal concentrations; some samples between the cell-free DNA simulation samples of first pregnancy and second pregnancy conform to same mother and father relationship, and the remaining samples conform to different mother and father relationship with the previously matched samples; clustering different fetal concentrations and corresponding K values ​​under same mother and father relationship to obtain clustering model M(PT), and clustering different fetal concentrations and corresponding K values ​​under different mother and father relationship to obtain clustering model M(NPT). A p-value of 0-25% is advantageous for calculating fetal concentration with high accuracy.

[0046] Among them, the dimorphic loci were selected from SNP loci, and the population frequency of the dimorphic loci was 0.05-0.95.

[0047] SNP genotyping method: After sequencing and analysis, each SNP locus in each sample will have a total sequencing depth, as well as the depth of "wild-type" and "mutant" loci determined based on the human genome reference sequence. Taking a certain SNP locus as an example, let A represent the wild-type locus and a represent the mutant locus. If the total depth of the locus in the sequencing results is 100X, where A is 100X and a is 0X, then the locus is a homozygous wild-type locus, denoted as AA; if A is 0X and a is 100X, then it is a homozygous mutant locus, denoted as aa; if the sequencing depth of A and a is close to 1:1, then the locus is heterozygous, denoted as Aa. The SNP genotyping results of the sample can be obtained accordingly. For cell-free DNA samples from pregnant women, if the sequencing depth ratio of A and a at this locus is A / (A+a) < 0.2, then the pregnant woman's genotype is denoted as aa, and the fetal genotype contains one A. Similarly, if the sequencing depth ratio of A and a at this locus is a / (A+a) < 0.2, then the fetal genotype is found to contain one a, and such locus is recorded as a locus containing fetal signal.

[0048] Example 2

[0049] Example 2 provides the process for obtaining the clustering model, as follows:

[0050] S201: Generate paternal DNA sample F1, maternal DNA sample M, and unrelated male DNA sample F2 randomly based on the frequency of the Chinese population; and generate offspring Z1 and Z2 from F1 and M, and offspring Z3 from F2 and M, according to Mendel's laws of inheritance.

[0051] S202: Immunografting is performed at fetal concentrations of 0-25%, with a gradient of 1%; specifically 0%, 1%, 2%...25%. The samples of Z1 and M are mixed to obtain simulated cell-free DNA samples of the pregnant woman, S1', which serves as a simulated sample of cell-free DNA from the first pregnancy. The samples of Z2 and M are mixed to obtain simulated cell-free DNA samples of the pregnant woman, S2', and the samples of Z3 and M are mixed to obtain simulated cell-free DNA samples of the pregnant woman, S3'. S2' and S3' serve as simulated samples of cell-free DNA from the second pregnancy. The fetuses corresponding to S1' and S2' are full parents, while the fetuses corresponding to S1' and S3' are half parents. Multiple samples are generated for each fetal concentration. The number of samples for S1', S2', and S3' may be equal or unequal, preferably equal; specifically, the number of samples for each simulated concentration of S1', S2', and S3' is 50, resulting in a total of 1250 samples for each of S1', S2', and S3'. The simulated sequencing depth was 100X-150X; 2500 autosomal SNP loci with bimorphic mutation frequencies between [0.05-0.95] on the sample genome were selected as genetic markers. The SNP locus data were obtained from: ftp: / / ftp.ncbi.nlm.nih.gov / snp / .redesign / .archive / b155 / VCF / GCF_000001405.39.gz; SNP genotyping was performed based on the simulated sequencing depth.

[0052] S203: Following the methods in steps S102 and S104, calculate the K value for each S1' and each S2', and the K value for each S1' and each S3'. The total number of K values ​​for the S1' and S2' clustering model is 50 * 25 * 25, and the total number of K values ​​for the S1' and S3' clustering model is also 50 * 25 * 25.

[0053] S204: Under different simulated concentrations, the fetal concentrations of clusters S1' and S2', and the corresponding K, yield the model M(PT), as follows: Figure 2 The points distributed in the upper part (at a simulated concentration). Clustering the fetal concentrations of S1' and S3' and their corresponding K yields the model M(NPT) as follows: Figure 1 The points distributed in the middle (at a simulated concentration). Furthermore, this embodiment can also design clustering models corresponding to two unrelated fetuses, such as... Figure 2 The distribution points of the lower part.

[0054] Furthermore, for each simulated concentration of a simulated sample (such as 1.5% in Example 4 and 4% in Example 5), the simulated concentration of another simulated sample is used as the x-axis, specifically S2' or S3' (S2' and S3' can be in the same coordinate system), and K is used as the y-axis. This yields a clustering correspondence map of p and K for each simulated concentration. All clustering correspondence maps for the simulated concentrations are plotted in the same coordinate system to form a clustering model, and each clustering correspondence map can be hidden. With a total of 25 concentration gradients, there are 25 clustering correspondence maps for same-parent relationships and 25 clustering correspondence maps for same-parent relationships. Same-parent and same-parent relationships can be plotted on a single map, such as... Figure 2 As shown; they may also be on different parts of the same image, such as... Figure 3 and Figure 5 As shown. In step S105, only the cluster correspondence graphs corresponding to the simulated concentrations of p1 or p2 are not hidden. Taking Example 5 as an example, p1 and p2 are 4.1% and 3.6% respectively, and the cluster correspondence graph corresponding to 4.0% for p1=4.1% is not hidden. In the unhidden cluster correspondence graphs, it is determined whether the simulated concentrations corresponding to p2 or p1 and K conform to the M(NPT) or M(PT) clustering model. Again, taking Example 5 as an example, the range of K calculated in the cluster correspondence graph is queried for 4.0% corresponding to p2=3.6%, and it is checked whether the K calculated by actual sampling is within the aforementioned range.

[0055] Of course, clustering models can also be implemented in other ways, such as using a mapping table, like p1=4%, p2=4%, K=37-44.3%. A common approach is to query whether the actual K value falls within the range of the simulated K value. Alternatively, specialized clustering software can be used for analysis.

[0056] In step S105, p1 and p2 are rounded to the nearest integer or a value between the two simulated concentrations to correspond to the simulated concentration in step S202. For example, in Example 4, p1 and p2 are 1.7% and 1.3% respectively, both corresponding to 1.5%; in Example 5, p1 and p2 are 4.1% and 3.6% respectively, both corresponding to 4%. If p1 and p2 are two sides of the same simulated concentration and the difference between them is less than 1%, then, as in Example 5, p1 and p2 can both correspond to the simulated concentration between the two values.

[0057] Furthermore, in the unhidden cluster correspondence map, at the same simulated concentration, the K values ​​of multiple samples form vertical line regions, such as... Figure 2 , Figure 3 and Figure 5 In the context of points closely arranged vertically, with p1 and p2 corresponding to p, determine whether K lies within the vertical region (points that are clearly discrete at the top and bottom can be removed if necessary).

[0058] Example 3

[0059] Example 3 discloses a theoretical calculation method for the commonality rate K of two independent pregnancies of the same pregnant woman:

[0060] After sequencing and analysis, each SNP locus in each sample will have a total sequencing depth, as well as the depths of "wild-type" and "mutant" loci determined based on the human genome reference sequence. This invention only discusses dimorphic loci, where the dimorphisms are denoted by A and a, respectively. Taking a specific SNP locus as an example, let A represent the wild-type locus with a population frequency denoted as p, and a represent the mutant locus with a population frequency denoted as q. Then, the frequency of the AA genotype is p. 2 The frequency of the aa genotype is q. 2 The frequency of the Aa genotype is 2pq.

[0061] When the mother has the AA genotype, if the father can provide one 'a' gene, a fetal signal will appear at that locus with a probability of p. 2 q; When the mother has the aa genotype, if the father can provide one A gene, a fetal signal will appear at that locus with a probability of q. 2 Theoretically, when M SNP sites are selected, the total number of sites where fetal signals appear is M*(p). 2 q+q 2 p) = M*p*q, where the number of fetal signals containing 'a' is m and the number of signals containing 'A' is n, then M*p*q = m + n.

[0062] Fetal origination from the same father: When the pregnant woman's genotype is AA at this locus and fetal signal 'a' is present, the father's genotype could be Aa with probability P; the father's genotype could also be aa with probability q. If a second pregnancy occurs at this time, at the same locus (the pregnant woman's genotype is AA), if the father's genotype is aa, a fetal signal will definitely appear; if the father's genotype is Aa, there is a 1 / 2P probability that the signal locus will not appear. Similarly, if the pregnant woman's genotype is aa at this locus, there is a 1 / 2q probability that a fetal signal will not appear. Based on this, the number of loci containing fetal signals shared by the two pregnancies is M*p*q - (m*1 / 2p + n*1 / 2q); the total number of fetal signal loci in the two pregnancies is M*p*q + (m*1 / 2p + n*1 / 2q). The proportion of shared fetal signal sites to the total number of fetal signal sites is K=[M*p*q-(m*1 / 2p+n*1 / 2q)] / [M*p*q+(m*1 / 2p+n*1 / 2q)]; when p=q=0.5, K=0.6, that is, when the fetuses of two independent pregnancies originate from the same father (same father and same mother), the theoretical value of the commonality rate of fetal signals is 0.6.

[0063] Fetuses from different fathers:

[0064] In the first pregnancy, the pregnant woman's genotype at this locus is AA, resulting in fetal signal 'a'. If in the second pregnancy, the fetus's true father is not the same person, then at the same AA locus, the probability of the other father providing 'a' is 'q', and the probability of providing 'A' is 'p'. Therefore, the probability that no fetal signal will appear at this locus is 'p'. Similarly, if the pregnant woman's genotype at the locus is 'aa', the probability that no fetal signal will appear in the second pregnancy is 'q'. Based on this, the number of loci containing fetal signals shared by the two pregnancies is M*p*q - (m*p + n*q); the total number of fetal signal loci in the two pregnancies is M*p*q + (m*p + n*q). The proportion of shared fetal signal loci to the total number of fetal signal loci is K = [M*p*q - (m*p + n*q)] / [M*p*q + (m*p + n*q)]. When p = q = 0.5, K = 0.33, meaning that the theoretical commonality rate of fetal signals when the fetuses from two independent pregnancies originate from different fathers (same mother, different fathers) is 0.33.

[0065] Example 4:

[0066] Example 4 discloses a method for paternity testing, comprising: a family numbered QZ19648, submitting a pregnant woman's peripheral blood sample numbered QZ19648S1 for testing; and another family numbered QZ28943, submitting a pregnant woman's peripheral blood sample numbered QZ29843S1 and a toothbrush sample of the alleged father, numbered QZY28943F. Sequencing analysis revealed that the pregnant women in both families were the same person, and the fetuses in both families matched QZY28943F. Consultation confirmed that both families were the same woman with two independent pregnancies. Therefore, the fetuses in both families are identical (same mother, same father). The fetal concentration of QZ19648S1 is known to be 1.7%; the fetal concentration of QZ29843S1 is known to be 1.3%. The co-occurrence rate K between QZ19648S1 and QZ29843S1 was calculated according to the method described in Example 1. Detection conclusion: The co-occurrence rate (K) between the two samples is 42.7%. (See attached image) Figure 3 When the simulated samples are q=0.15 and 0.15 (with a concentration gradient of 0.5%), K is approximately 35-43%, with 42.7% falling within this range. This cluster clearly shows that the fetuses from two independent pregnancies share the same father (i.e., the fetuses are full siblings) and the K values ​​of the fetuses with concentrations of 1.5% and 1.5% are clustered together. Therefore, it can be determined that the fetuses of the pregnant women corresponding to samples QZ19648S1 and QZ29843S1 are full siblings, which is consistent with the actual situation.

[0067] Example 5:

[0068] Example 5 discloses a method for paternity testing, comprising: a family numbered QZ19899, ​​for which peripheral blood samples from the pregnant woman and bloodstain samples from the alleged father were submitted for testing, numbered QZ19899S1 and QZH19899F, respectively. Another family numbered QZ26979, for which peripheral blood samples from the pregnant woman and hair samples from the alleged father were submitted for testing, numbered QZ26979S1 and QZM26979F, respectively. Sequencing analysis revealed that the fetus in QZ19899S1 is parented to both QZH19899F and QZ26979S1; the fetus in QZ26979S1 is parented to both QZM26979F. The pregnant women in QZ19899S1 and QZ26979S1 are the same person, while those in QZH19899F and QZM26979F are not the same person. Consultation confirmed that the two families belong to the same pregnant woman and two independent pregnancies. Therefore, the fetuses from the two families are half-fathers, sharing the same mother. The fetal concentration of QZ19899S1 is known to be 4.1%; the fetal concentration of QZ26979S1 is known to be 3.6%. The co-occurrence rate k between QZ19899S1 and QZ26979S1 was calculated according to the method in Example 1. Detection conclusion: The co-occurrence rate k between the two samples is 30.48%. (See [link to relevant documentation]). Figure 3 When the simulated samples are q=0.4 and 0.4 (with a concentration gradient of 1%), k is approximately 28-33%, with 30.48% falling within this range. This cluster clearly shows that the fetuses from two independent pregnancies with different fathers (i.e., half-fathers) and fetal concentrations of 4% and 4% respectively have k values. Therefore, it can be determined that the pregnant women and fetuses corresponding to samples QZ19899S1 and QZ26979S1 are half-fathers, which is consistent with the actual situation.

[0069] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for paternity testing, characterized in that, The method includes the following steps: S101: DNA data S1 and S2 were obtained by sequencing the dimorphic sites of cell-free DNA samples from two pregnancies of a pregnant woman. S102: Obtain the genotyping site sets X(S1) and X(S2) containing fetal signals in S1 and S2, respectively; S103: The fetal DNA concentrations p1 and p2 are calculated based on S1 and S2, respectively; S104: Calculate the common ownership rate K according to Formula II. K=[X(S1)∩X(S2)] / [X(S1)∪X(S2)](Ⅱ); S105: Determine whether the fetuses in the two pregnancies are of the same mother and father or the same mother but different fathers based on the correspondence between K, p1, and p2 and the clustering model of the simulated samples. The clustering model is obtained by clustering fetal concentration and K in steps S102 and S104 based on the fetal relationship between the two pregnancies of the pregnant woman and the simulated samples with clear fetal concentration. Step S105 specifically includes: Obtain the clustering model M(PT) of fetal concentration and K for mother-son and father-son relationships and the clustering model M(NPT) of fetal concentration and K for mother-son and half-son relationships. Determine whether K, p1, and p2 conform to model M(PT) or M(NPT). If K, p1, and p2 conform to model M(PT), then the two fetuses are mother-son and father-son relationships. If K, p1, and p2 conform to model M(NPT), then the two fetuses are half-son relationships. Cell-free DNA samples from first pregnancy and second pregnancy with different fetal concentrations were simulated. Some samples from the first and second pregnancy cell-free DNA samples were found to be from the same mother and father, while the remaining samples were found to be from different mothers and fathers. Clustering was performed to obtain a clustering model M(PT) for different fetal concentrations and corresponding K values ​​under the same mother and father relationship, and a clustering model M(NPT) for different fetal concentrations and corresponding K values ​​under the same mother and father relationship. The M(NPT) and M(PT) are obtained by the following method: S201: Generate paternal DNA sample F1, maternal DNA sample M, and unrelated male DNA sample F2 randomly based on the frequency of the Chinese population; according to Mendel's laws of inheritance, offspring Z1 and Z2 are generated from F1 and M, and offspring Z3 is generated from F2 and M. S202: Fetal concentrations of 0-25% were mixed in an isogradient manner to obtain simulated cell-free DNA samples of pregnant women, such as S1' (mixed with Z1 and M), S2' (mixed with Z2 and M), and S3' (mixed with Z3 and M). Multiple samples were generated for each fetal concentration, with a simulated sequencing depth of 100X-150X. Autosomal SNP loci with a mutation frequency between [0.05-0.95] on the sample genome were selected as genetic markers, and SNP genotyping was performed based on the simulated sequencing depth. S203: Calculate the K value for each S1' and each S2' and the K value for each S1' and each S3' according to the method of steps S102 and S104; S204: Under different simulated concentrations, the fetal concentrations of cluster S1' and S2' and their corresponding K are used to obtain model M(PT), and the fetal concentrations of cluster S1' and S3' and their corresponding K are used to obtain model M(NPT).

2. The method for paternity testing according to claim 1, characterized in that, For each simulated concentration of a simulated sample, the simulated concentration of another simulated sample is plotted on the x-axis, and K is plotted on the y-axis. Obtain the cluster correspondence map of p and K for each simulated concentration; draw the cluster correspondence maps of all simulated concentrations in the same coordinate system to form a clustering model, and each cluster correspondence map on it can be hidden; in step S105, only select the cluster correspondence map corresponding to the simulated concentration of p1 or p2 to not hide; in the unhidden cluster correspondence map, determine whether the simulated concentration of p2 or p1 and K conform to the M(NPT) or M(PT) clustering model.

3. The method for paternity testing according to claim 2, characterized in that, The gradient difference for equal gradients is 0.5% or 1%, and 50 samples are generated for each fetal concentration. More than 2,000 SNP sites on the genome of each sample are selected.

4. The method for paternity testing according to claim 3, characterized in that, In step S105, p1 and p2 are rounded or rounded down to correspond to the simulated concentration in step S202.

5. The method for paternity testing according to claim 2, characterized in that, In the unhidden cluster correspondence map, under the same simulated concentration, the K values ​​of multiple samples form a vertical line region; under the correspondence between p1 and p2 and p, determine whether K is located in the vertical line region.

6. The method for paternity testing according to claim 1, characterized in that, The dimorphic loci are selected from SNP loci, and the population frequency of the dimorphic loci is 0.05-0.

95.

7. The method for paternity testing according to claim 1, characterized in that, In step S102, the set of genotyping sites containing fetal signals is obtained according to Formula I. X(S)={X i |0<na i (S) / n i (S)<0.2∪0<nA i (S) / n i (S)<0.2}(I), Where nA and na represent the observed values ​​of dimorphic sites A and a, respectively, and n = nA + na.

Citation Information

Patent Citations

  • Method for evaluating DNA concentration of fetus by using free DNA of pregnant woman and application

    CN113999900A

  • Method for judging parent-child relationship between pregnant woman and fetus by calculating fetus concentration

    CN114496078A