Biomarker for early prediction of preeclampsia based on free DNA of body fluid and application of biomarker
By analyzing biological samples from healthy and diseased pregnant subjects, extracting and comparing motif sets, and constructing a prediction model, we solved the problem of efficient and accurate prediction of preeclampsia and achieved low-cost disease detection during pregnancy.
Patent Information
- Application Number
- CN202410329658.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies make it difficult to predict and detect preeclampsia efficiently and accurately, and the lack of highly accurate and low-cost prediction methods leads to insufficient early prevention and treatment.
By analyzing biological samples from healthy and diseased pregnant subjects, motif sets are extracted and compared, differential motif sets are determined, and a prediction model is constructed. Specific motifs and motif enrichment rates are used as biomarkers to predict and diagnose diseases during pregnancy.
It achieves high-accuracy and low-cost prediction of pregnancy diseases, especially early detection of preeclampsia, reduces the depth of data analysis and the targeted area of the genome, and improves detection efficiency and accuracy.
Smart Images

Figure CN120683239A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of prenatal diagnosis, and in particular to a biomarker for early prediction of preeclampsia based on free DNA in body fluids and its application. Background Art
[0002] Preeclampsia (also known as preeclampsia, preeclampsia, PE) / eclampsia is one of the three major causes of maternal morbidity and mortality worldwide. Severe eclampsia is difficult to detect accurately and efficiently before onset. For the early prediction of preeclampsia, related technologies have developed from traditional mid-pregnancy diastolic blood pressure measurement, mean arterial pressure measurement, turning test, 24-hour dynamic blood pressure and hemorheology monitoring to uterine artery blood flow index, vascular endothelial growth factor, soluble vascular endothelial growth factor receptor-1 and placental growth factor combined detection and early pregnancy serum placental protein-13 measurement, and then to genomic research. However, there is still no efficient, accurate and low-cost method for predicting preeclampsia, and early detection and prevention of preeclampsia are extremely limited.
[0003] Extracellular free DNA (cfDNA) primarily originates from apoptosis, necrosis, and secretion of extracellular networks (netosis). This free DNA (including plasma) carries a rich array of physiological and pathological signals from tissues. Maternal plasma cfDNA data contains both maternal and fetal DNA information, providing a theoretical basis for using cfDNA data to conduct genetic analyses of maternal-fetal dual genome associations. cfDNA can effectively detect fetal RhD blood type, sex, common chromosomal trisomies, and single-gene genetic disorders. Developing highly accurate, low-cost cfDNA-based methods for predicting preeclampsia is crucial for the early prediction and diagnosis of eclampsia. Summary of the Invention
[0004] The present application solves at least one of the problems of the related art from the following aspects.
[0005] To this end, an embodiment of the present application provides a method for determining biomarkers for predicting or diagnosing pregnancy diseases, comprising: based on sequencing data from biological samples of healthy subjects and subjects with the pregnancy disease in the first trimester, the second trimester, and an optional third trimester, taking motifs in each of the biological samples, wherein each motif consists of N bases, to obtain initial motif sets H1 and P1 corresponding to healthy subjects and subjects with the pregnancy disease in the first trimester, initial motif sets H2 and P2 corresponding to healthy subjects and subjects with the pregnancy disease in the second trimester, and optional motif sets H3 and P4 corresponding to healthy subjects and subjects with the pregnancy disease in the third trimester. Initial motif sets H3 and P3 corresponding to healthy subjects in the third trimester and subjects suffering from the pregnancy disease; comparing the initial motif sets H1, P1, H2, P2 and optionally H3, P3 to determine the difference motif sets A, B and optionally C in P1, P2 and optionally P3, respectively; determining a candidate motif set X for the pregnancy disease based on the difference motif sets A, B and optionally C; and determining, based on the candidate motif set X, a biomarker for predicting or diagnosing the pregnancy disease based on a prediction model, wherein the N bases of the motif are continuous or discontinuous in the genomic sequence corresponding to the biological sample, and N is a positive integer greater than 0.
[0006] In some embodiments, comparing the initial motif sets H1, P1, H2, P2 and optionally H3, P3 to respectively determine the differential motif sets A, B and optionally C in P1, P2 and optionally P3 includes: comparing the frequency distribution of each of the motifs in the initial motif sets H1 and P1 to determine the motifs in P1 that are upregulated or downregulated relative to H1, and using them as the differential motif set A; comparing the frequency distribution of each of the motifs in the initial motif sets H2 and P2 to determine the motifs in P2 that are upregulated or downregulated relative to H2, and using them as the differential motif set B; and optionally comparing the frequency distribution of each of the motifs in the initial motif sets H3 and P3 to determine the motifs in P3 that are upregulated or downregulated relative to H3, and using them as the differential motif set C, optionally, determining the differential motif sets A, B and optionally C based on the fold change of the frequency distribution and the test probability value, preferably, the test is a rank sum test.
[0007] In some embodiments, based on the difference motif sets A, B and optionally C, determining the candidate motif set X for the pregnancy disease includes: taking the intersection motifs with consistent difference changes in the difference motif sets A, B and optionally C as the candidate motif set X for the pregnancy disease, wherein the consistent difference changes are all up-regulated or all down-regulated.
[0008] In some embodiments, according to the candidate motif set X, determining a biomarker for predicting or diagnosing a pregnancy disease based on a prediction model includes: constructing the prediction model, using the motifs in the candidate motif set X and the corresponding motif frequencies as initial features for model training and cross-validation; and determining the model features of the prediction model based on the model test results in which the AUC value in the cross-validation is higher than the AUC threshold, wherein the model features include the specific motifs of the pregnancy disease, and using the specific motifs of the pregnancy disease as the first biomarker for predicting or diagnosing the pregnancy disease. Optionally, the method further includes: analyzing the A, T, G, and C bases in the specific up-regulated motifs and the specific down-regulated motifs in the specific motifs of the pregnancy disease. and base enrichment rates of a combination thereof, based on the fact that the base enrichment rate Up-ER / Dw-ER of a certain base or a base combination containing it is significantly higher than the base enrichment rate Up-ER' / Dw-ER' of other bases or base combinations containing it, the base enrichment rate Up-ER and / or Dw-ER corresponding to the base or the base combination containing it is used as the second biomarker for predicting or diagnosing pregnancy diseases, wherein the AUC threshold is selected from any value in the range of (0.6-1], preferably any value in the range of (0.7-1]. Optionally, the prediction model is based on a linear model, and the linear model is selected from one or more of a linear regression model, a logistic regression model, a Lasso regression model, a ridge regression model, and a linear discriminant analysis model.
[0009] In some embodiments, the biological sample is cell-free DNA (cfDNA), and optionally, the motif composed of N bases is one or more of the following: a. N bases among the M bases located at the end of the sequence of the biological sample, the end being the 5' end and / or the 3' end, preferably the 5' end, and M is a positive integer greater than 0 and M≥N; b. N bases located in the genomic sequence corresponding to the biological sample, preferably N bases among the M' bases upstream and downstream of the end of the sequence of the biological sample on its corresponding genomic sequence, the end being the 5' end and / or the 3' end, preferably the 5' end, and M' is a positive integer greater than 0 and M'≥N.
[0010] In some embodiments, the first trimester, the second trimester, and the third trimester are respectively selected from any one of the early trimester, the mid-pregnancy, and the late trimester. Optionally, the early trimester is 1-11 weeks of pregnancy, the mid-pregnancy is 12-24 weeks of pregnancy, and the late trimester is 25-40 weeks of pregnancy. Optionally, the pregnancy diseases are selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction.
[0011] An embodiment of the present application also provides the use of a reagent for detecting biomarkers in the preparation of a product for predicting or diagnosing pregnancy diseases, wherein the biomarker is determined based on a biological sample according to the method for determining a biomarker for predicting or diagnosing a pregnancy disease as described in any of the above embodiments, and the biomarker includes a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction.
[0012] In some embodiments, the pregnancy disease is eclampsia, and the first biomarker of eclampsia is an eclampsia-specific motif, wherein each motif in the eclampsia-specific motif consists of N bases, N is a positive integer greater than 0, optionally, N is taken from [1, 30], preferably [1, 15], more preferably [1, 10], and most preferably N=4, wherein the first biomarker of eclampsia includes one or more of the following: TCAT, GCGA, GCCG, TCGT, TAGT, GGGT, GCAG, CTTC, CACA, TCAG, TT GC, TCTG, TCGC, AGTG, TCGA, GCGG, TCGG, TCAA, GCAA, CAAC, TTGA, GAGT, GCGC, GCGT, GCAC, GAGC, GAGA and CTTG, preferably, the first biomarker for eclampsia includes one or more of the following: GGGT, CTTC, GAGA, TTGC, CACA, CTTG, TCGA, TCGT, GCCG, GCGA, TCGC, GCGT, GCGC, CAAC, TCGG and TTGA.
[0013] In some embodiments, the pregnancy disease is eclampsia, and the second biomarker for eclampsia is the base enrichment rate Up-ER and / or Dw-ER of a certain base or a base combination containing the same calculated based on the first biomarker for eclampsia, wherein the first biomarker for eclampsia includes an up-regulated motif and a down-regulated motif, wherein the second biomarker for eclampsia is the enrichment rate Up-ERAC of the base combination of A and C in the up-regulated motif; and / or the enrichment rate Dw-ERTG of the base combination of T and G in the down-regulated motif, optionally, the Up-ERAC ≥ 70%, preferably Up-ERAC ≥ 80%, more preferably Up-ERAC ≥ 90%, most preferably Up-ERAC ≥ 95%; and the Dw-ERAC ≥ 70%, preferably Dw-ERAC ≥ 80%, more preferably Dw-ERAC ≥ 90%, most preferably Dw-ERAC ≥ 95%.
[0014] In some embodiments, the reagents for detecting biomarkers include primers and / or probes that can amplify the biomarkers, and / or sequencing reagents that can sequence the biomarkers.
[0015] An embodiment of the present application also proposes a method for predicting or diagnosing a pregnancy disease, comprising: using a biomarker as a predictive variable, calculating a prediction result based on the level of the biomarker in a biological sample from a subject; and indicating whether the subject is at risk of suffering from the pregnancy disease or suffers from the pregnancy disease based on the prediction result, wherein the biomarker includes a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction.
[0016] In some embodiments, the step of using a biomarker as a predictor variable and calculating a prediction result based on the level of the biomarker in a biological sample from a subject comprises:
[0017] (a) Based on the number of biological samples to be predicted being m and the number of specific motifs of the first biomarker being n:
[0018] X∈R m*(n+1) ;
[0019] where X i,j (j>1) represents the frequency of the jth specific motif in the i-th biological sample, X i,1 =1;
[0020] (b)θ∈R (n+1)*1 ;
[0021] where θ j,1 (j>1) represents the model weight of the j-th specific motif; and
[0022] (c) P = sigmoid(Xθ) = (1 + e -xθ ) -1
[0023] P∈R m*1
[0024] Each P value in the matrix P is the prediction result of each of the m biological samples.
[0025] In some embodiments, based on the pregnancy disease being eclampsia, the first biomarker for eclampsia includes one or more of the following: TCAT, GCGA, GCCG, TCGT, TAGT, GGGT, GCAG, CTTC, CACA, TCAG, TTGC, TCTG, TCGC, AGTG, TCGA, GCGG, TCGG, TCAA, GCAA, CAAC, TTGA, GAGT, GCGC, GCGT, GCAC, GAGC, GAGA and CTTG. Preferably, the first biomarker for eclampsia includes one or more of the following: GGGT, CTTC, GAGA, TTGC, CACA, CTTG, TCGA, TCGT, GCCG, GCGA, TCGC, GCGT, GCGC, CAAC, TCGG and TTGA.
[0026] In some embodiments, the model weight of the model for predicting eclampsia corresponding to the first biomarker of eclampsia is selected from one or more of the following:
[0027]
[0028]
[0029] In some embodiments, the indication of whether the subject is at risk of suffering from the pregnancy disease or suffering from the pregnancy disease based on the prediction result includes: based on the P value of the biological sample being greater than or equal to a threshold value, indicating that the subject is at risk of suffering from the pregnancy disease or suffering from the pregnancy disease; based on the P value of the biological sample being less than the threshold value, indicating that the subject is not at risk of suffering from the pregnancy disease or does not suffer from the pregnancy disease. In some embodiments, the threshold value can optionally be selected from [0.5-1), such as 0.5, 0.6, 0.7, 0.8 or any value therebetween. Taking a threshold value of 0.5 as an example, based on a p value greater than or equal to 0.5, it indicates that the subject is at risk of suffering from the pregnancy disease or suffers from the pregnancy disease (such as eclampsia); based on a p value less than 0.5, it indicates that the subject is not at risk of suffering from the pregnancy disease or does not suffer from the pregnancy disease.
[0030] An embodiment of the present application also proposes a system for predicting or diagnosing pregnancy diseases, the system comprising: a processor; an input module for inputting the levels of biomarkers in a biological sample from a subject, wherein the biomarkers include a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker; a computer-readable medium comprising instructions, which, when executed by the processor, implement the method for predicting or diagnosing pregnancy diseases as described in any of the above embodiments of the present application; and an output module for indicating whether the subject is at risk of gestational diabetes or suffers from the pregnancy disease.
[0031] An embodiment of the present application also proposes a model for predicting or diagnosing pregnancy diseases, comprising: a calculation module for using biomarkers as model features to calculate a prediction result based on the level of the biomarker in a biological sample from a subject; and an indication module for indicating whether the subject is at risk of suffering from the pregnancy disease or suffers from the pregnancy disease based on the prediction result, wherein the biomarker comprises a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction.
[0032] The embodiments of the present application achieve the following beneficial effects:
[0033] The methods for predicting or diagnosing biomarkers for pregnancy disorders and related applications based thereon (including applications, models, and systems for predicting or diagnosing pregnancy disorders based on the identified biomarkers) proposed in the embodiments of this application can identify specific motifs for the pregnancy disorders through machine learning models, and based on these specific motifs, can effectively predict / diagnose pregnancy disorders. The biomarkers and related applications proposed in the embodiments of this application greatly reduce the depth required for data analysis because they only utilize short sequence fragments in biological samples, and their genomic targeting regions are small, which has the advantages of high accuracy and low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 Schematic diagram of the process of constructing eclampsia-specific motifs and related models according to Example 1 of the present application;
[0036] Figure 2 This is a schematic diagram of a process for predicting pregnancy diseases based on eclampsia-specific motifs according to an embodiment of the present application;
[0037] Figure 3 This is a distribution diagram of differential motifs between eclampsia and healthy pregnant women according to Example 1 of the present application;
[0038] Figure 4 The importance of each model feature in the eclampsia prediction model according to Example 1 of the present application is shown;
[0039] Figure 5 is the ROC evaluation result of the prediction model constructed based on the eclampsia-specific motif according to Example 1 of the present application;
[0040] Figure 6 The AUC distribution for eclampsia prediction under different cfDNA quantity gradients in downsampling according to Example 2 of the present application;
[0041] Figure 7 The figure is a ROC diagram for eclampsia prediction in the NIPT dataset according to Example 3 of the present application. DETAILED DESCRIPTION
[0042] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.
[0043] This application is made by the inventor based on the following understanding:
[0044] Previous studies have shown that noninvasive early screening for pregnancy-related disorders can be achieved through early detection of specific signals in plasma related to pregnancy-related disorders. This includes testing for maternal serum levels of pregnancy-associated plasma protein A (PAPP-A), placental growth factor (PLGF), and its related receptor sFlt-1. However, due to its clinical manifestations (serum PAPP-A levels between pregnant women with mid-trimester eclampsia and normal pregnant women are not significantly different), testing is usually performed between 11 and 13 weeks of gestation. The area under the guidance of a single PAPP-A marker in early pregnancy is low, necessitating its combination with other markers. PLGF testing also requires combination with sFlt-1 testing, making the test more complex and preventing effective eclampsia diagnosis using a single marker.
[0045] Related technologies also involve predicting preeclampsia using deconvolution methods based on cfDNA methylation and transcriptome features. Based on the low coverage of WGBS data and the percentage of placental cfDNA during early pregnancy, a sensitive tissue deconvolution method was developed to obtain its tissue origin from cfDNA methylation data (Del Vecchio G, Li Q, Li W, et al. Cell-free DNA methylation and transcriptomic signature prediction of pregnancies with adverse outcomes [J]. Epigenetics, 2021, 16 (6): 642-661.). In addition, by detecting gene changes in the cfRNA transcriptome, normal and preeclamptic pregnancies can be distinguished throughout the pregnancy process (Moufarrej MN, Vorperian SK, Wong RJ, et al. Early prediction of preeclampsiain pregnancy with cell-free RNA [J]. Nature, 2022, 602 (7898): 689-694.). However, cfDNA methylation testing requires a high plasma volume, is complex to operate, and is costly. Furthermore, current cfRNA detection technology is not yet mature. Because RNA is unstable and easily degraded, the method places extremely high demands on sample quality and experimental operation, and its AUC (Area Under the Curve, ROC curve) for detecting eclampsia is relatively low.
[0046] In this regard, this application proposes a complete set of methods for efficiently predicting pre-eclampsia based on cfDNA. The specific motifs identified in the cfDNA of pregnant women with eclampsia can be effectively used for eclampsia screening in pregnant women in the second trimester (same stage as NIPT, 12-24 weeks of gestation). Only the terminal base information is used, the required sequencing depth is low, the genomic targeted area is small, the analysis efficiency and accuracy are high, and the cost is low.
[0047] In the embodiments of the present application, "motif" refers to any sequence pattern that may be related to molecular function, structural properties or family members. In some embodiments, "motif" refers to a short characteristic DNA fragment whose sequence length can be N (i.e., composed of N bases), where N is a positive integer greater than 0. In some embodiments, N is any positive integer greater than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10. In some embodiments, N is taken from [1, 30], preferably [1, 15], more preferably [1, 10], and most preferably N=4. It is understandable that since cfDNA is composed of 4 bases, A, T, G, and C, the number of motif types involved in the analysis is at most 4. n Species(s).
[0048] In some embodiments, the N bases of the motif are continuous or discontinuous in the genomic sequence corresponding to the biological sample, and the N bases can be a. N bases located at the end of the sequence of the biological sample, where M is a positive integer greater than 0 and M ≥ N; and / or b. N bases located in the genomic sequence corresponding to the biological sample, preferably N bases located at the end of the sequence of the biological sample in the upstream and downstream of the genomic sequence corresponding to the biological sample, wherein the end is the 5' end and / or the 3' end, preferably the 5' end, where M' is a positive integer greater than 0 and M ≥ N. It will be understood that when M or M'=N, the N bases are continuous in the genomic sequence, and when M or M'>N, the N bases may be discontinuous in the genomic sequence. For example, when the biological sample is cfDNA, when N is 3, in case a, the three bases of the motif can be the 1st, 3rd, and 4th bases from the end of the cfDNA sequence (i.e., discontinuous in the genome), or the 1st, 2nd, and 3rd bases from the end (i.e., continuous in the genome); in case b, the N bases of the motif can be continuous or discontinuous bases upstream and downstream of the cfDNA end at the corresponding genomic coordinates, with the end of the cfDNA sequence as the reference point. For example, the four bases of the motif can be -1 before the end, the 1st, 3rd, and 4th bases from the end (N=4, discontinuous in the genome), or -2 bases before the end, the 1st, 2nd, and 3rd bases from the end (N=5, continuous in the genome). The end is the 5' end and / or the 3' end, preferably the 5' end.
[0049] In the embodiments of the present application, the first trimester, the second trimester and the optional third trimester are any one of the early pregnancy, the mid-pregnancy and the late pregnancy, respectively, wherein the first trimester, the second trimester and the third trimester are different. It will be appreciated that the number of weeks corresponding to each stage of pregnancy can be divided according to different classification criteria. In some embodiments, the early pregnancy is 1-11 weeks of pregnancy, the mid-pregnancy is 12-24 weeks of pregnancy, and the late pregnancy is 25-40 weeks of pregnancy. In other embodiments, the early pregnancy is 1-13 weeks of pregnancy, the mid-pregnancy is 14-28 weeks of pregnancy, and the late pregnancy is 29-40 weeks of pregnancy. The embodiments of the present application can effectively identify biomarkers that can be used to predict / diagnose the disease during pregnancy by comparing sick pregnant women in different trimesters and / or healthy pregnant women and sick pregnant women in the same trimester.
[0050] In the embodiments of the present application, "pregnancy-related diseases" can be selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction. In some embodiments, the "pregnancy-related disease" is eclampsia / preeclampsia, particularly early-onset severe eclampsia.
[0051] In the embodiments of the present application, the biological sample can be cfDNA obtained from a subject. In some embodiments, the biological sample can be extracted from a pregnant woman's body fluid. In some embodiments, the biological sample can be extracted from the pregnant woman's peripheral blood, including blood, plasma, serum, etc.
[0052] In the embodiments of the present application, "biomarkers" refer to biochemical indicators that mark changes or possible changes in the structure or function of systems, organs, tissues, cells and subcellular structures, and can be used for the diagnosis of diseases, determination of disease stages, or evaluation of the safety and effectiveness of new drugs or new therapies in target populations. In the embodiments of the present application, "biomarkers" are based on the motifs of biological samples. In some embodiments, the biomarker can be a specific motif (i.e., a first biomarker) for predicting / diagnosing the pregnancy disease. In other embodiments, the biomarker can be a base enrichment rate (i.e., a second biomarker) of a base or a combination of bases containing the base that undergoes specific changes, obtained by comparing the motif distribution frequencies between sample sets of a pregnancy disease (e.g., sick pregnant women at different stages of pregnancy and / or healthy pregnant women and sick pregnant women at the same stage of pregnancy).
[0053] In some embodiments, the first biomarker can be a specific motif for a pregnancy disease, specifically a model feature corresponding to a model test result with a high AUC threshold (or ROC curve) in a model constructed by the method for determining a biomarker. In other words, it is a specific motif that can be used for effective disease typing, determined by a machine learning model constructed based on sample data. These motifs can be used for the prediction / diagnosis of corresponding pregnancy diseases. In some embodiments, all identified specific motifs can be used as model features (i.e., first biomarkers) for the prediction / diagnosis of late pregnancy diseases, or some identified specific motifs can be used as model features for the prediction / diagnosis of late pregnancy diseases. In other embodiments, based on all identified specific motifs, correlation values between motifs can be used as model features (i.e., first biomarkers) for the prediction / diagnosis of late pregnancy diseases. In some embodiments, the correlation value can be a difference, ratio, sum, and / or product value between motifs, preferably a difference and / or ratio.
[0054] It is understood that the first biomarker (i.e., the specific motif of the pregnancy disease) may include up-regulated motifs and down-regulated motifs. In some embodiments, when the pregnancy disease is eclampsia, based on N=4, the specific motif in the first biomarker may be selected from one or more of the following: TCAT, GCGA, GCCG, TCGT, TAGT, GGGT, GCAG, CTTC, CACA, TCAG, TTGC, TCTG, TCGC, AGTG, TCGA, GCGG, TCGG, TCAA, GCAA, CAAC, TTGA, GAGT, GCGC, GCGT, GCAC, GAGC, GAGA, and CTTG (a total of 28 species), preferably selected from one or more of the following: GGGT, CTTC, GAGA, TTGC, CACA, CTTG, TCGA, TCGT, GCCG, GCGA, TCGC, GCGT, GCGC, CAAC, TCGG, and TTGA (a total of 16 species). It can be understood that the method for determining biomarkers for predicting or diagnosing pregnancy diseases proposed in the embodiments of the present application can be effectively used to identify specific motifs for a certain pregnancy disease, and the identified specific motifs can be effectively used to construct a prediction model for accurate prediction / diagnosis of pregnancy diseases.
[0055] In the embodiments of the present application, in the machine learning model constructed based on the sample set data, the first biomarker also has a corresponding model weight. In some embodiments, when the pregnancy disorder is eclampsia, based on N=4, the corresponding weights of the 28 or 16 biomarkers listed above are shown in Table 1 below.
[0056] In the embodiments of the present application, a "machine learning model," "prediction model," "typing model," etc., is a mapping that can be viewed as a function f(x) that outputs a certain result given an input (x). In the embodiments of the present application, the model can be based on a linear model, which can be selected from one or more of a linear regression model, a logistic regression model, a lasso regression model, a ridge regression model, and a linear discriminant analysis model.
[0057] In an embodiment of the present application, the authenticity of the prediction model (detection method) is evaluated based on the AUC value, and the value range of AUC is between 0.5 and 1. The closer the AUC is to 1.0, the higher the authenticity of the detection method; when it is equal to 0.5, the authenticity is the lowest. In an embodiment of the present application, by setting an AUC threshold for the prediction model, a high-accuracy model covering better biomarker indicators can be screened out. In some embodiments, the AUC threshold is selected from any value in (0.6-1], preferably any value in (0.7-1].
[0058] In the embodiment of the present application, for the second biomarker, the base enrichment rates of A, T, G, C and their combinations in the specific up-regulated motif and the specific down-regulated motif in the specific motif of the pregnancy disease (which can be the first biomarker) are analyzed. The base enrichment rate Up-ER / Dw-ER of a certain base or a base combination containing it is significantly higher than the base enrichment rate Up-ER' / Dw-ER' of other bases or base combinations containing it, then the base enrichment rate Up-ER and / or Dw-ER corresponding to the base or the base combination containing it is used as the second biomarker for predicting or diagnosing pregnancy diseases. In other words, the second biomarker in the embodiment of the present application reflects the changing trend of the motif at a certain position of its cfDNA during the onset of a pregnancy disease. It is understandable that this trend (i.e., the second biomarker) is more flexible than the specific motif that has been determined to be used as the first biomarker. Therefore, in the embodiments of the present application, the first biomarker and the second biomarker can be used alone for the prediction of pregnancy diseases, or the combination of the two can be used for the prediction of pregnancy diseases, both of which have high accuracy.
[0059] In some embodiments, when the pregnancy disease is eclampsia, based on N=4, the second biomarker for eclampsia is the enrichment rate Up-ER of the base combination of A and C in the up-regulated motif. AC and / or the enrichment rate Dw-ER of the base combination of T and G in the down-regulated motif TG In some embodiments, the Up-ER AC ≥70%, preferably Up-ER AC ≥80%, more preferably Up-ER AC≥90%, most preferably Up-ER AC ≥95%; and the Dw-ER AC ≥70%, preferably Dw-ER AC ≥80%, more preferably Dw-ER AC ≥90%, most preferably Dw-ER AC In the examples of the present application, the second biomarker identified can be effectively used for early prediction / diagnosis of pregnancy diseases.
[0060] In some embodiments, the comparison between the sample sets of sick pregnant women in different trimesters and / or healthy pregnant women and sick pregnant women in the same trimester can correspond to the comparison between H1 (i.e., a healthy pregnant woman in the first trimester, wherein the first trimester is one of the early, middle and late trimesters), P1 (i.e., a sick pregnant woman in the first trimester, wherein the first trimester is one of the early, middle and late trimesters), H2 (i.e., a healthy pregnant woman in the second trimester, wherein the second trimester is one of the pregnancy stages other than the first trimester), P2 (i.e., a sick pregnant woman in the second trimester, wherein the second trimester is one of the pregnancy stages other than the first trimester) and optional H3 (i.e., a healthy pregnant woman in the third trimester, wherein the third trimester is a pregnancy stage other than the first and second trimesters), P3 (i.e., a sick pregnant woman in the third trimester, wherein the third trimester is a pregnancy stage other than the first and second trimesters), which can specifically include: the comparison between H1 and P1, the comparison between H2 and P2 and the optional comparison between H3 and P3. In other embodiments, the comparison may further include: comparing the 4 n The method comprises the following steps: analyzing the frequency distribution of motifs to determine the relatively up-regulated or down-regulated motifs in P1, P2 and optionally P3; and determining the differential motif sets A, B and optionally C in P1, P2 and optionally P3 based on the relatively up-regulated or down-regulated motifs in P1, P2 and optionally P3.
[0061] Figure 1 This is a schematic diagram of the process of constructing eclampsia-specific motifs and related models according to Example 1 of the present application. Figure 1 In some specific embodiments, a method for determining a biomarker for predicting or diagnosing a pregnancy disease may include: Step 1-Step 2.
[0062] Step 1. Identification of eclampsia-specific fragment end motifs in plasma cfDNA
[0063] (1) 5 ml of peripheral blood was collected from pregnant women with eclampsia and healthy pregnant women in the second and third trimesters (corresponding to the second and first trimesters, respectively). After plasma separation, plasma free DNA was extracted and high-throughput sequenced. The 4mer bases at the 5' end of the cfDNA fragments (i.e., 4mer motif, N = 4) were extracted, and the frequency distribution of all 256 motifs was calculated.
[0064] (2) The fold change of the frequencies of motifs at the end of cfDNA between the eclamptic and healthy pregnant women in the late pregnancy (first trimester) (P1 and H1, respectively) was calculated, and the P value was calculated by Wilcoxon rank-sum test. The motifs that were specifically increased or decreased in eclampsia were screened based on the fold change and P value, and were defined as motif set A (i.e., corresponding to differential motif set A).
[0065] (3) The fold difference of the motif frequencies at the end of cfDNA between the eclamptic and healthy pregnant women in the second trimester (second trimester) (P2 and H2, respectively) was calculated, and the P value was calculated using the Wilcoxon rank-sum test. The motifs that were specifically increased or decreased in eclampsia were screened based on the fold difference and P value, and were defined as motif set B (i.e., corresponding to differential motif set B).
[0066] (4) Screening motifs with significant consistency differences between motif set A and set B as eclampsia-specific candidate motifs (i.e., corresponding to candidate motif set X), wherein the consistency difference changes are both up-regulated or both down-regulated.
[0067] Step 2: Screening of eclampsia-specific motifs and construction of an eclampsia prediction model based on eclampsia-specific motif features
[0068] The frequencies of eclampsia-specific candidate motifs were used as initial features to construct a machine learning model for training and cross-validation. The optimal feature set (i.e., eclampsia-specific motif, also known as the first biomarker) and model parameters were determined based on the test results with higher or even highest AUC values in cross-validation.
[0069] The method for predicting or diagnosing biomarkers for diseases during pregnancy proposed in the embodiment of the present application is based on a machine learning model and can effectively screen biological samples for specific motif sequences (i.e., first biomarkers) for a disease during pregnancy. It has high accuracy and strong stability when used to predict or diagnose diseases during pregnancy. At the same time, since the screening method for biomarkers only involves the motif of biological samples, the required sequencing depth is extremely low, thereby effectively improving the efficiency of biomarker identification and greatly reducing costs. In addition, the biomarkers screened out by the biomarker screening method proposed in this application can be effectively used to predict or diagnose diseases during pregnancy, and can be compatible with existing pregnancy detection technologies and processes, which is of great significance for efficient early detection and early judgment of diseases during pregnancy.
[0070] The second aspect of the present application also proposes the use of a reagent for detecting a biomarker in the preparation of a product for predicting or diagnosing a disease during pregnancy. The biomarker may include a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the disease during pregnancy, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. In some embodiments, the biomarker is determined according to the method for determining a biomarker for predicting or diagnosing a disease during pregnancy proposed in the first aspect of the present application. It is understandable that the reagent for detecting a biomarker may include: a primer and / or probe that can amplify the biomarker, and / or a sequencing reagent that can sequence the biomarker.
[0071] Based on the biomarkers identified in the first aspect of the present application, the third aspect of the present application also proposes a method for predicting or diagnosing pregnancy diseases, the method comprising: using the biomarker as a predictive variable, and calculating a prediction result based on the level of the biomarker in a biological sample from a subject; and indicating whether the subject is at risk of suffering from the pregnancy disease or suffers from the pregnancy disease based on the prediction result, wherein the biomarker comprises a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same, calculated based on the first biomarker.
[0072] Figure 2 The figure is a flow chart of the prediction process of pregnancy diseases based on eclampsia-specific motifs according to an embodiment of the present application. Figure 2 In some specific embodiments, the method for predicting or diagnosing pregnancy diseases proposed in the embodiments of the present application may include:
[0073] (1) Analyze the terminal 4mer motif of plasma cfDNA of pregnant women to be tested and extract the frequency of eclampsia-specific motif.
[0074] (2) The motif features of the sample to be tested are input into the prediction model to obtain the predicted value, and the eclampsia classification and sample set prediction AUC are evaluated.
[0075] It will be understood that, according to the method for predicting or diagnosing a pregnancy disease, it can be indicated whether a subject is at risk of suffering from the pregnancy disease or whether the subject suffers from the pregnancy disease, and a person skilled in the art will be able to treat a subject at risk of suffering from the disease or suffering from the disease based on the prediction result.
[0076] The fourth aspect of the present application also proposes a system for predicting or diagnosing pregnancy diseases, the system comprising: a processor; an input module for inputting the levels of biomarkers in a biological sample from a subject, wherein the biomarkers include a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker; a computer-readable medium comprising instructions, which, when executed by the processor, implement the method for predicting or diagnosing pregnancy diseases as described in claim 12; and an output module for indicating whether the subject is at risk of gestational diabetes or suffers from the pregnancy disease.
[0077] The fifth aspect of the present application also proposes a model for predicting or diagnosing pregnancy diseases, including: a calculation module for using biomarkers as model features to calculate prediction results based on the levels of the biomarkers in biological samples from subjects; and an indication module for indicating whether the subject is at risk of suffering from the pregnancy disease or suffers from the pregnancy disease based on the prediction results, wherein the biomarkers include a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same, calculated based on the first biomarker.
[0078] In some embodiments, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth, and fetal growth restriction. In some embodiments, the biological sample is cfDNA.
[0079] In some embodiments, the biological sample is cfDNA.
[0080] It should be noted that the aforementioned explanations of the method embodiment for determining biomarkers for predicting or diagnosing pregnancy diseases also apply to the biomarkers and the methods, models, and systems for predicting diseases in the above embodiments, and will not be repeated here.
[0081] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.
[0082] Unless otherwise specified, the quantitative analysis experiments in the following examples were performed three times, and the results were averaged.
[0083] Example 1 Identification of specific biomarkers for pregnancy-related diseases and model construction
[0084] This example uses eclampsia as an example to identify eclampsia-specific biomarkers and construct a model.
[0085] Sample information:
[0086] The first batch of samples (T1) included plasma samples from 56 healthy pregnant women (mid-trimester: 12-24 weeks) (NT2), 71 pregnant women with epilepsy (mid-trimester: 12-24 weeks) (PT2), 8 healthy pregnant women (NB) in the third trimester (late trimester: 25-40 weeks), and 16 pregnant women with epilepsy (PB). cfDNA was extracted and sequenced from the first batch of samples at a sequencing depth of 7×.
[0087] The second batch of samples (test set T2) included plasma samples from 48 healthy pregnant women (mid-trimester: 12-24 weeks) (NT2) and 26 pregnant women with epilepsy (mid-trimester: 12-24 weeks) (PT2), with a sequencing depth of 16×. cfDNA was extracted and sequenced from the second batch of samples at a sequencing depth of 16×.
[0088] 1.1 For the first batch of samples (T1), the motifs of cfDNA 5' end 4mers (i.e., N = 4, a total of 256 types) between NT2 and NB, PT2 and PB, NT2 and PT2, and NB and PB were compared, and the log fold change (FC) of the motifs between samples was calculated. 10 Value (log 10 (FC)). Taking NT2 and PT2 as an example, At the same time, the 256 motifs of the two groups of samples were subjected to Wilcoxon rank-sum test to calculate the P value. 10 (FC) is sorted into the top 25% and bottom 25%, and -log 10Motifs with a p value of ≤ 2 were defined as significantly increased (up) or significantly decreased (down, Dw) in eclampsia. This method allowed us to compare and identify upregulated and downregulated motifs between pregnant women with eclampsia and healthy women at the same gestational stage, as well as upregulated and downregulated motifs between pregnant women with eclampsia at different gestational stages.
[0089] Comparative analyses of 1.2NT2 vs. PT2 and NB vs. PB revealed consistent patterns in eclampsia: a high enrichment of A and C bases in Up Motifs (approximately 95%) and an enrichment of T and G bases in Down Motifs (approximately 95%). Furthermore, comparative analyses of pregnancies with eclampsia during the second trimester (PT2) and third trimester (PB) revealed the same pattern of high enrichment of A and C bases in Up Motifs and T and G bases in Down Motifs in PB, indicating that the signal for this motif increases with gestational age, closer to the onset of eclampsia. In contrast, no significant enrichment of this motif was found in the comparison of pregnancies between the second and third trimesters of healthy women. These results suggest that this specific motif pattern of high enrichment of A and C bases in Up Motifs and T and G bases in Down Motifs is a specific disease signature associated with eclampsia and increases with the onset of eclampsia. Thus, the enrichment rate Up-ERAC (here 95%) of the base combination of A and C in the up-regulated motif and / or the enrichment rate Dw-ERTG (here 95%) of the base combination of T and G in the down-regulated motif can be determined as specific markers for eclampsia (corresponding to the second biomarker).
[0090] 1.3 The samples from different periods of the above analysis all verified that the high enrichment of A and C bases and the significant reduction of T and G base terminal motifs are specific disease signals for eclampsia (corresponding to the second biomarker). Based on this, this embodiment further constructed a model based on logistic regression for feature learning and classification training of samples. This method uses the significant Up Motifs and Down Motifs in the NB vs PB analysis as candidate motifs, and adds the motif features with the best single classification effect to the model one by one until the classification effect of the model tends to be stable. Table 1 shows the final selected motifs (a total of 28 types, and only the top 16 types can be selected) and their weights. Figure 4 The feature importance of the finally selected motifs can be used as specific biomarkers for eclampsia (corresponding to the first biomarker, i.e., the model feature described below) for the prediction and diagnosis of eclampsia based on cell-free DNA.
[0091] The specific steps for prediction and diagnosis using the final selected motif and its model weights are as follows:
[0092] (a) Based on the number of samples to be predicted being m, the number of specific motifs of the first biomarker being n:
[0093] X∈R m*(n+1) ;
[0094] where X i,j (j>1) represents the frequency (value, frequency) of the j-th specific motif in the i-th sample, X i,1 =1;
[0095] (b)θ∈R (n+1)*1 ;
[0096] where θ j,1 (j>1) represents the model weight of the j-th specific motif (as shown in Table 1, the first 16 can be selected or all 28 can be used); and
[0097] (c) P = sigmoid(Xθ) = (1 + e -xθ ) -1
[0098] P∈R m*1
[0099] Each P value in the matrix P is the predicted value of each of the m samples. The matrix P has a total of m elements, corresponding to the m samples. The P value of each sample is used to determine whether the sample is eclamptic or healthy. If the P value is greater than or equal to 0.5 or 0.6, the sample is determined to be an eclamptic sample; otherwise, it is a healthy sample.
[0100] Table 1 Eclampsia-specific motifs and model weights
[0101] Serial number Eclampsia-specific motif Model weights 1 GGGT 5111.4746 2 CTTC 3480.3069 3 GAGA -969.8014 4 TTGC 99.528967 5 CACA 1202.1121 6 CTTG -417.4065 7 TCGA -8739.084 8 TCGT 10650.596 9 GCCG 11268.117 10 GCGA 11697.553 11 TCGC -13858.21 12 GCGT -2479.137 13 GCGC -2492.197 14 CAAC -3281.586 15 TCGG -5091.481 16 TTGA -2977.342 17 TCTG 63.276857 18 GCGG -7568.539 19 AGTG -12656.95 20 TCAT 15640.368 21 TAGT 7042.7615 22 GAGC -1651.055 23 GAGT -2805.391 24 GCAA -4219.036 25 GCAC -1706.898 26 TCAA -4840.762 27 GCAG 4292.8747 28 TCAG 529.28185
[0102] 1.4 After the model characteristics are determined, this example uses cross-validation to find the optimal model parameters to build a highly accurate and efficient eclampsia prediction model based on the biomarkers identified above.
[0103] 1.4.1 Five-fold cross-validation was performed on the T1 sample set for pregnant women with eclampsia and healthy pregnant women. Model parameters were adjusted and screened based on the AUC of the five-fold cross-validation. The optimal parameters finally determined achieved an AUC of 0.90 in the cross-validation test set, CI: (0.79, 0.97) ( Figure 5 ).
[0104] 1.4.2 After determining the model characteristics and optimal parameters, this example uses the T2 sample set as an external independent validation set to perform model validation and classification effect evaluation. The evaluation results are as follows: Figure 5 As shown. Figure 5It can be seen that the AUC in this independent validation set can reach 0.82 ( Figure 5 ). This indicates that the pregnancy disease-specific motif identification method of this embodiment can effectively identify specific biomarkers that can be used for the pregnancy disease. At the same time, based on the identified biomarkers, a prediction model for the prediction / diagnosis of the disease can be constructed, and based on the prediction model, highly accurate and efficient disease prediction / diagnosis can be achieved.
[0105] Example 2 Exploring the available sequencing depth based on the biomarkers and prediction models of Example 1
[0106] To further evaluate the impact of different sequencing depths on the prediction and diagnosis of eclampsia in maternal samples, this example performed a downsampling analysis on the T2 sample set. This analysis randomly sampled cfDNA molecules at varying numbers, ranging from 100,000 to 100 million, from all samples. Ten random samplings were performed for each order of cfDNA molecule, and the predicted value was output based on the eclampsia prediction model described in this example.
[0107] Figure 6 The results of the entire downsampling analysis and the AUC distribution for eclampsia prediction at different cfDNA sequencing depths are presented. These results demonstrate that with 900,000 cfDNA sequences, or a sequencing depth of approximately 0.05×, an AUC of 0.75 can be achieved for eclampsia prediction, demonstrating that the eclampsia-specific motif features identified in this example and the prediction model constructed are applicable to very low sequencing depths.
[0108] Example 3: The biomarkers and prediction model based on Example 1 can stably and accurately predict / diagnose pregnancy diseases at ultra-low sequencing depth
[0109] Sample Information: The third batch of samples (NIPT test set) consisted of 78 second-trimester non-invasive prenatal testing (NIPT) sequencing samples collected from Suzhou Municipal Hospital. These samples included 42 healthy pregnant women (gestational age: 12-24 weeks) and 36 pregnant women with epilepsy (gestational age: 12-24 weeks). cfDNA was extracted from plasma samples and sequenced using the Illumina platform using single-end 35bp sequencing, with an average sequencing depth of approximately 0.1x.
[0110] To evaluate the clinical application potential of the biomarkers identified in the examples of this application and the models constructed based on them, this example conducted further testing on a NIPT dataset. Specifically, this example performed terminal motif analysis on cfDNA sequencing data from 36 pregnant women with eclampsia and 42 matched healthy pregnant women at the same gestational stage (gestational period: 12-24 weeks), and extracted the frequencies of the 28 eclampsia-specific motifs shown in Table 1 from their terminal motifs.
[0111] Based on the frequency of 28 eclampsia-specific motifs, a prediction model for eclampsia was trained and evaluated with five-fold cross-validation based on logistic regression. The final model had a prediction AUC of 0.84 ( Figure 7 ), further verifying the stability and accuracy of this early eclampsia screening biomarker under different sequencing platforms, different sequencing strategies and ultra-low sequencing depths.
[0112] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0113] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for determining biomarkers for predicting or diagnosing pregnancy diseases, characterized in that: include: Based on sequencing data from biological samples of healthy subjects and subjects with the pregnancy disease in a first trimester, a second trimester, and an optional third trimester, extracting motifs in each of the biological samples, wherein each motif consists of N bases, to obtain initial motif sets H1 and P1 corresponding to the healthy subjects and subjects with the pregnancy disease in the first trimester, initial motif sets H2 and P2 corresponding to the healthy subjects and subjects with the pregnancy disease in the second trimester, and optional initial motif sets H3 and P3 corresponding to the healthy subjects and subjects with the pregnancy disease in the third trimester; Comparing the initial motif sets H1, P1, H2, P2 and optionally H3, P3 to determine differential motif sets A, B and optionally C in P1, P2 and optionally P3, respectively; Determining a candidate motif set X for the pregnancy disease based on the difference motif sets A, B and optionally C; and According to the candidate motif set X, a biomarker for predicting or diagnosing a pregnancy disease is determined based on a prediction model, The N bases of the motif are continuous or discontinuous in the genome sequence corresponding to the biological sample, and N is a positive integer greater than 0.
2. Use of a reagent for detecting a biomarker in the preparation of a product for predicting or diagnosing a pregnancy disease, characterized in that: Based on the biological sample, the biomarker is determined according to the method for determining a biomarker for predicting or diagnosing a pregnancy disease according to claim 1, The biomarker includes a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif of the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth and fetal growth restriction.
3. The use according to claim 2, characterized in that The pregnancy disease is eclampsia, and the first biomarker of eclampsia is an eclampsia-specific motif, wherein each motif in the eclampsia-specific motif consists of N bases, N is a positive integer greater than 0, optionally, N is selected from [1, 30], preferably [1, 15], more preferably [1, 10], and most preferably N=4, wherein the first biomarker for eclampsia comprises one or more of the following: TCAT, GCGA, GCCG, TCGT, TAGT, GGGT, GCAG, CTTC, CACA, TCAG, TTGC, TCTG, TCGC, AGTG, TCGA, GCGG, TCGG, TCAA, GCAA, CAAC, TTGA, GAGT, GCGC, GCGT, GCAC, GAGC, GAGA, and CTTG, Preferably, the first biomarker for eclampsia comprises one or more of the following: GGGT, CTTC, GAGA, TTGC, CACA, CTTG, TCGA, TCGT, GCCG, GCGA, TCGC, GCGT, GCGC, CAAC, TCGG and TTGA.
4. A method for predicting or diagnosing a disease during pregnancy, characterized in that: The method comprises: Using the biomarker as a predictor variable, calculating a prediction result based on the level of the biomarker in a biological sample from the subject; and According to the prediction result, it is indicated whether the subject has a risk of suffering from the pregnancy disease or suffers from the pregnancy disease, The biomarker comprises a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif of the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth and fetal growth restriction.
5. The method according to claim 4, characterized in that The method of using the biomarkers as predictive variables and calculating the prediction results based on the levels of the biomarkers in the biological sample from the subject comprises: (a) Based on the number of biological samples to be predicted being m and the number of specific motifs of the first biomarker being n: X∈R m*(n+1) ; where X i,j (j>1) represents the frequency of the jth specific motif in the i-th biological sample, X i,1 =1; (b)θ∈R (n+1)*1 ; where θ j,1 (j>1) represents the model weight of the j-th specific motif; and (c)P=sigmoid(Xθ)=(1+e -xθ ) -1 P∈R m*1 Each P value in the matrix P is the prediction result of each of the m biological samples.
6. The method according to claim 5, characterized in that Based on the pregnancy disease being eclampsia, the first biomarker for eclampsia includes one or more of the following: TCAT, GCGA, GCCG, TCGT, TAGT, GGGT, GCAG, CTTC, CACA, TCAG, TTGC, TCTG, TCGC, AGTG, TCGA, GCGG, TCGG, TCAA, GCAA, CAAC, TTGA, GAGT, GCGC, GCGT, GCAC, GAGC, GAGA and CTTG, Preferably, the first biomarker for eclampsia comprises one or more of the following: GGGT, CTTC, GAGA, TTGC, CACA, CTTG, TCGA, TCGT, GCCG, GCGA, TCGC, GCGT, GCGC, CAAC, TCGG and TTGA.
7. The method according to claim 6, characterized in that The model weight of the model for predicting eclampsia corresponding to the first biomarker of eclampsia is selected from one or more of the following:
8. The method according to any one of claims 4 to 7, characterized in that Indicating, based on the prediction result, whether the subject is at risk of suffering from the pregnancy disease or suffering from the pregnancy disease comprises: A P value based on the biological sample being greater than or equal to a threshold value indicates that the subject is at risk of having the pregnancy disease or has the pregnancy disease; The P value based on the biological sample is less than the threshold value, indicating that the subject is not at risk of having the pregnancy disease or does not have the pregnancy disease, Optionally, the threshold is 0.5 or 0.
6.
9. A system for predicting or diagnosing diseases during pregnancy, characterized in that: The system comprises: processor; an input module for inputting the levels of biomarkers in a biological sample from a subject, wherein the biomarkers include a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif for the pregnancy disease, and the second biomarker is a base enrichment ratio (Up-ER) and / or a base enrichment ratio (Dw-ER) corresponding to a certain base or a base combination containing the same, calculated based on the first biomarker; A computer-readable medium comprising instructions that, when executed by the processor, implement the method for predicting or diagnosing a pregnancy disorder according to any one of claims 4 to 8; and An output module is used to indicate whether the subject has a risk of suffering from gestational diabetes or suffers from the pregnancy-related disease.
10. A model for predicting or diagnosing diseases during pregnancy, characterized in that: The model includes: a calculation module for calculating a prediction result based on the level of the biomarker in a biological sample from a subject using the biomarker as a model feature; and an indication module, configured to indicate whether the subject has a risk of suffering from the pregnancy disease or has suffered from the pregnancy disease according to the prediction result, The biomarker comprises a first biomarker and / or a second biomarker, wherein the first biomarker is a specific motif of the pregnancy disease, and the second biomarker is a base enrichment rate Up-ER and / or Dw-ER corresponding to a certain base or a base combination containing the same calculated based on the first biomarker. Optionally, the pregnancy disease is selected from one or more of the following: anemia during pregnancy, hypertension during pregnancy, gestational diabetes, eclampsia, premature birth and fetal growth restriction.
Citation Information
Cited By
Preeclampsia noninvasive screening method based on deep sequencing of 4-mer terminal motif spectrum characteristics
CN121380322A