Classifier for early identification of the risk of placenta accreta spectrum disorders in high-risk pregnant women and its training system
Through a machine learning strategy based on the coverage of cfDNA promoter in pregnant women, 23 gene combinations were identified and a classifier was constructed to conduct PAS risk assessment, which solved the problem of early identification of placental implantable diseases, realized early risk assessment and non-invasive detection in high-risk pregnant women, and reduced perinatal risk.
Patent Information
- Application Number
- CN202411383048.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-09-30
AI Technical Summary
There is a lack of effective methods for early identification of the risk of placental implant disease in the prior art, resulting in late prenatal diagnosis and missed diagnosis, increasing the risk of major bleeding during delivery and hysterectomy during delivery.
Using a machine learning strategy based on the coverage of cfDNA promoter in pregnant women, a classifier was constructed to conduct PAS risk assessment for high-risk pregnant women by identifying specific combinations and expression patterns of 23 genes such as ABHD1 and ALG1L2, and non-invasive detection was performed using NIPT data.
It realizes early identification of high-sensitivity and specificity of PAS high-risk pregnant women, reduces perinatal risk, provides a non-invasive early prediction tool, and has significant clinical application value.
Smart Images

Figure CN119230114B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of model prediction, and particularly relates to a classifier for early identifying the risk of placenta accreta spectrum (PAS) in high-risk pregnant women and its training system. Background Art
[0002] Placenta accreta spectrum (PAS) is one of the main causes of obstetric critical illnesses. However, its prenatal diagnosis has limitations such as late detection and many missed diagnoses, especially posing a huge challenge to primary hospitals. Research shows that clinically, 2 / 3 of PAS cases are still missed. Failure to identify PAS prenatally is a risk factor for postpartum hemorrhage, blood transfusion, emergency interventional procedures, and hysterectomy during and after delivery. Therefore, early and accurate prediction of PAS helps high-risk pregnant women make decisions about their childbearing intentions, undergo high-risk referrals, achieve multidisciplinary joint consultations, and reduce the risks to pregnant women and perinatal infants.
[0003] Cell-free DNA (cfDNA) in plasma is derived from apoptotic cells. CfDNA contains nucleosome footprints and can reflect the gene expression information of the tissue from which cfDNA originated. During pregnancy, about 10% of the cfDNA in circulating blood comes from the placenta. Therefore, plasma cfDNA in the early pregnancy contains the gene expression information of the placenta and decidua. The whole-genome promoter nucleosome coverage profile of cfDNA in pregnant women in the early and middle pregnancy can reflect the expression patterns of each source tissue and has extremely high predictive value for placenta-derived diseases, especially PAS. Non-invasive prenatal DNA testing (NIPT) is a common prenatal screening item clinically. Hospitals at home and abroad perform NIPT relying on the whole-genome low-coverage sequencing technology of different sequencing platforms, such as the Illumina, Life, and BGI platforms. In recent years, in addition to being applied to the screening of fetal chromosomal abnormalities, the extraction of the cfDNA promoter nucleosome coverage profile based on NIPT has also shown great value in the early prediction of pregnancy complications, such as fetal growth restriction, macrosomia, preeclampsia, etc. However, in terms of placenta accreta spectrum diseases, there is no effective early prediction model. Summary of the Invention
[0004] In view of the actual needs and the deficiencies of the existing technology, the present invention provides a classifier for early identifying the risk of PAS in high-risk pregnant women and its training system, aiming to solve the problem that there is currently no method to accurately predict the occurrence of PAS in the early and middle pregnancy.
[0005] Technical Solution:
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a classifier for early identification of the risk of PAS occurrence in high-risk pregnant women based on the promoter coverage rate of plasma cfDNA. The target gene combination includes ABHD1, ALG1L2, EYS, FAM157C, KDSR, KRT5, LANCL2, LINC00390, LINC00964, LOC105371998, LOC107987394, LOC644090, LYZL2, MIR184, MIR4802, MYT1L, NGDN, NSD2, PACRG.AS3, SAP30L.AS1, SLC16A12.AS1, TADA3, TMEM147.AS1.
[0008] The present invention utilizes NIPT data to discover the cfDNA whole-genome promoter coverage feature spectrum in the early and mid-term plasma of pregnant women with PAS. Through machine learning strategies (see the second aspect), the optimal gene combination and the optimal cut-off value of each gene are obtained, and the optimal classifier is trained. The area under the curve (AUC) in the receiver operating characteristic curve (ROC) of the prediction effect in an independent validation dataset reaches above 0.85, showing good potential as a screening method for high-risk pregnant women with PAS.
[0009] A method for evaluating the risk of placental implantation diseases in high-risk pregnant women includes:
[0010] Data collection and preprocessing: Collect low-coverage whole-genome sequencing data of high-risk pregnant women during NIPT and perform necessary preprocessing to ensure data quality.
[0011] Promoter region identification and coverage rate extraction: Use bwa-mem, SAMtools, BEDtools software and sequence alignment algorithms to align the NIPT data with the human reference genome hg19. Determine the promoter region pTSS of 1000 bp above and below the transcription start site (TSS) of 23 genes including ABHD1, ALGIL2, EYS, FAM157C, KDSR, KRT5, LANCL2, LINC00390, LINC00964, LOC105371998, LOC107987394, LOC644090, LYZL2, MIR184, MIR4802, MYT1L, NGDN, NSD2, PACRG.AS3, SAP30L.AS1, SLC16A12.AS1, TADA3, TMEM147.AS1, and obtain the original read coverage rate of the pTSS regions of the 23 genes.
[0012] Feature factor normalization: Normalize the original read coverage of the pTSS regions of 23 genes by the class TPM method to obtain the class TPM-normalized promoter coverage (TPM-like Normalized pTSS Coverage, NPC-TPM) of the 23 genes.
[0013]
[0014] Among them, NPC-TPM i represents the class TPM-normalized promoter coverage of the pTSS region of the i-th gene, q i represents the original read coverage of the pTSS region, l i represents the transcript length (both are 2000), ∑ j (q j / l j ) represents the sum of the read coverages of each pTSS region after transcript length normalization.
[0015] Feature factor discretization: Compare the NPC-TPM values of the 23 genes with the optimal cut-off values (obtained from the second aspect) of each gene respectively. When the feature factor NPC-TPM is greater than the corresponding optimal cut-off value, set it to 1; otherwise, set it to 0.
[0016] Risk assessment: Input the discretized NPC-TPM values of the 23 genes of the pregnant woman into the classifier to calculate the occurrence risk of PAS.
[0017] This method is applicable to singleton pregnant women with at least one of the following high-risk factors for PAS: ① Previous uterine surgery history, such as cesarean section, myomectomy, metroplasty, etc.; ② Previous intrauterine operation history, such as hysteroscopic surgery history, uterine curettage history, etc.; ③ In vitro fertilization-embryo transfer (IVF-ET).
[0018] This process is based on the extraction of pTSS coverage from NIPT data, providing a non-invasive assessment method. It indicates that the specific combination and expression pattern of the above 23 genes have great potential in the production of diagnostic kits. In the development of diagnostic kits, by detecting the expression of these genes through precise molecular diagnostic techniques, a highly sensitive and highly specific detection method can be designed to accurately predict PAS in high-risk pregnant women at an early stage, and is expected to become a routine clinical detection item to serve a wider patient population.
[0019] In the second aspect, the present invention provides a classifier training system for the occurrence risk of PAS in high-risk pregnant women based on the promoter sequencing of plasma cfDNA in the early and middle pregnancy stages;
[0020] A classifier training method for the risk of PAS in high-risk pregnant women based on cfDNA promoter sequencing in the first and second trimesters of pregnancy, including:
[0021] Dataset partitioning module: Extract the medical data of PAS high-risk pregnant women who have undergone NIPT on different sequencing platforms in advance. Match the PAS pregnant women and those without PAS according to age, gestational week at the time of NIPT, fetal sex, and distribution of high-risk factors. Randomly divide the samples on the main platform into a training dataset and an internal validation dataset, and use the samples on the remaining platforms as an independent external validation dataset.
[0022] Impact factor extraction module: Perform cfDNA promoter nucleosome coverage annotation and feature extraction on the NIPT data of PAS high-risk pregnant women collected in advance.
[0023] NIPT is to perform low-depth high-throughput sequencing on cell-free DNA in pregnant women's peripheral blood. Using an accurate sequence alignment algorithm, use bwa-mem, SAMtools, and BEDtools software to align the NIPT data with the human reference genome hg19, delete PCR duplicates, determine the region 1000bp upstream and downstream of the transcription start site (TSS) as the promoter region pTSS, and calculate the raw read coverage of the pTSS region.
[0024] The class TPM-like normalized promoter coverage (NPC-TPM) of each gene is calculated through class TPM to reduce the impact of sequencing depth on data extraction and analysis.
[0025]
[0026] Among them, NPC-TPM i represents the class TPM-like normalized promoter coverage of the pTSS region of the i-th gene, q i represents the raw read coverage of the pTSS region, l <000…… i represents the transcript length (all 2000), ∑ j (q j / l j ) represents the sum of the read coverages of the pTSS regions of all genes after transcript length normalization in a certain sample.
[0027] Input the NPC-TPM values corresponding to all genes into the system as impact factors.
[0028] Feature screening module: Use propensity score to perform 1:1 matching on pregnant women with PAS recombination and those without PAS in the training dataset, and then include them in the differential analysis. Perform DESeq2, limma-voom, and rank sum test analysis on each influencing factor respectively. Screen out the influencing factors with p-values calculated by all three differential analysis methods < 0.05 or at least two methods with p-values < 0.05 as feature factors.
[0029] Feature factor discretization module: To enhance the universality and clinical practicability of the classifier for different sequencing platforms, adopt a discretization strategy:
[0030] Set the optimal cut-off value of each feature factor as the NPC-TPM value with the maximum sum of sensitivity and specificity in the training dataset. When the NPC-TPM of the feature factor is greater than the corresponding optimal cut-off value, set it to 1; otherwise, set it to 0.
[0031] Model acquisition module: Input the screened feature factors into machine learning feature selection processes such as Recursive Feature Elimination (RFE), and gradually construct a PAS disease prediction classifier using various machine learning methods such as Support Vector Machine (SVM)-linear kernel and SVM-Gaussian kernel function (Radial Basis Function, RBF).
[0032] Set pregnant women with PAS as 1 and high-risk pregnant women without PAS as 0 and input them into the system for classifier training. To improve the sensitivity of the classifier, set the parameter value of class_weight as {[0:0.1,1:0.3]}.
[0033] Obtain the disease risk assessment results of the target to be predicted, and apply k-fold cross-validation (k = 10) to increase the robustness of the assessment.
[0034] Extract the optimal feature factor combination and output the optimal classifier and its assessment results.
[0035] Advantages of the present invention
[0036] Placenta accreta spectrum (PAS) is a general term for a group of diseases characterized by abnormal adhesion or invasion of the placenta into the myometrium. After the fetus is delivered in patients with PAS, the placenta cannot be normally exfoliated, leading to massive bleeding at the placental exfoliation surface. It is one of the main causes of obstetric critical illnesses such as emergency hysterectomy, multiple organ failure, disseminated intravascular coagulation, shock, and even perinatal death. In recent years, with the increase in cesarean section surgeries and the progress of intrauterine operations and assisted reproductive technologies, the incidence of PAS has been rising year by year. According to statistics, it occurs in 1 case per 300 to 400 pregnancies. Although PAS cases diagnosed prenatally based on existing clinical means often have a more severe degree of invasion into the myometrium, the emergency cesarean section rate, blood loss, and blood transfusion volume are all lower than those of PAS cases not diagnosed prenatally. Therefore, prenatal identification and perinatal management of PAS are crucial.
[0037] The classifier developed based on 23 genes NPC in this study, which is based on non-invasive blood tests of pregnant women in the early and middle stages, can effectively predict the risk of PAS in high-risk pregnant women and has important clinical application value. The introduction of this innovative method is expected to be directly applied to clinical practice, providing a scientific basis and practical guidance for the early diagnosis and treatment of PAS, and having important clinical significance for the early prediction of the clinical outcome of pregnant women with placenta accreta, referral of high-risk pregnant women, multidisciplinary joint consultation, and reduction of the risks of pregnant women and perinatal infants. Brief Description of the Drawings
[0038] Figure 1 Venn diagram for differential gene screening in Example 1
[0039] Figure 2 Venn diagram for differential gene screening in Example 2
[0040] Figure 3 Schematic diagram of the performance of the 25-gene LR classifier obtained by inputting the NPC-TPM values of 702 genes in Example 2
[0041] Figure 4 Schematic diagram of the performance of the 23-gene SVM-RBF kernel classifier obtained by inputting the NPC-TPM values of 702 genes in Example 2 Detailed Embodiments
[0042] The present invention will be further described below in conjunction with the embodiments, but the protection scope of the present invention is not limited thereto:
[0043] Example 1
[0044] Methods: A high-risk PAS cohort was established at the Maternal and Child Health Hospital of City A (BGI sequencing platform) and the Maternal and Child Health Hospital of City B (Illumina sequencing platform). NIPT data and clinical information of pregnant women who underwent non-invasive DNA screening and had PAS high-risk factors (at least one of the following: previous cesarean section, uterine surgery such as uterine fibroids, hysteroscopy or intrauterine operation, assisted reproductive technology) were prospectively collected. The patients were followed up until 28 days after delivery.
[0045] Exclusion criteria included: ① stillbirth, neonatal death, or other non-viable pregnancy outcomes; ② fetal chromosomal abnormalities or structural developmental abnormalities; and ③ multiple pregnancies. Pregnant women with PAS and those without PAS were matched based on age, gestational age at NIPT, fetal sex, and distribution of high-risk factors.
[0046] Among high-risk pregnant women undergoing NIPT using the BGI sequencing platform, 54 PAS and 157 non-PAS high-risk women were included in the training dataset, 26 PAS and 71 non-PAS high-risk women were included in the internal validation dataset, and 25 PAS and 77 non-PAS high-risk women were included in the time validation dataset for differential analysis and model training. In addition, 54 PAS and 162 non-PAS high-risk women undergoing NIPT using the Illumina sequencing platform were included in an independent external validation dataset.
[0047] The NIPT data of the enrolled pregnant women were annotated with cfDNA promoter nucleosome coverage profiles, and sequence alignment was performed with the human reference genome hg19. PCR duplicates were removed, and the raw read coverage of the pTSS region was calculated. The raw read coverage was normalized using two methods to obtain the NPC-RPKM value and NPC-TPM value of each gene, respectively.
[0048]
[0049] Among them, NPC-RPKM i represents the RPKM-normalized promoter coverage of the pTSS region of the i-th gene, NPC-TPM i represents the TPM-normalized promoter coverage of the pTSS region of the i-th gene, q i represents the raw read coverage of the pTSS region, l i represents the transcript length (all 2000), tmr represents the total length of all genes, ∑ j (q j / l j ) represents the sum of the read coverage of the pTSS regions of all genes in a sample after normalization of transcript length.
[0050] In the training dataset, PAS pregnant women were matched with high-risk pregnant women without PAS at a ratio of 1:1, with 54 samples in each group. DESeq2, limma-voom, and the Wilcoxon rank-sum test were used for analysis. Genes with p-values < 0.05 and log2FC > 0.5 calculated by at least two of the three differential analysis methods were screened, and 226 differentially pTSS-covered genes were obtained. As Figure 1 shown in the Venn diagram of differential gene screening.
[0051] To verify the effectiveness of the model after standardization by different means, in the embodiment, the optimal cut-off values of the 226 feature factors were set as the NPC-RPKM value and NPC-TPM value with the maximum sum of sensitivity and specificity in the training dataset. When the feature factor NPC (NPC-RPKM or NPC-TPM) is greater than the corresponding optimal cut-off value, it is set to 1; otherwise, it is set to 0.
[0052] The screened feature factors were input into the RFE process, and multiple machine learning algorithms were used to gradually construct a PAS disease prediction classifier. SVM-Linear kernel, SVM-RBF kernel, and LR were respectively selected for classifier training. PAS pregnant women were set to 1, and high-risk pregnant women without PAS were set to 0 and input into the system for classifier training. To improve the sensitivity of the classifier, the parameter values of class_weight were set as {0:0.1, 1:0.3}. The disease risk assessment results of the target to be predicted were obtained, and k-fold cross-validation (k = 10) was applied to increase the robustness of the assessment. The optimal feature factor combinations under different modeling methods were extracted, and the optimal classifier and its evaluation results were output.
[0053] For the NPC-RPKM value, the output result was an SVM-Linear classifier with 23 genes. The verification performance of this classifier was poor (AUC < 0.5) in the external validation sets of different platforms, as shown in Table 1.
[0054] Table 1
[0055]
[0056] For the NPC-TPM value, the output result was an SVM-Linear kernel classifier with 31 genes. The verification performance of this classifier increased significantly in both the temporal validation set and the external validation set. It can be seen that the NPC-TPM value is more suitable for data on different platforms, as shown in Table 2.
[0057] Table 2
[0058]
[0059]
[0060] For the NPC-TPM values, the output results are the LR classifier of 38 genes and the SVM-RBF kernel classifier of 38 genes. See Tables 3 and 4 for details. The validation performance of both classifiers in the external validation set has increased significantly (AUC > 0.7), and the SVM-RBF kernel classifier has the best performance. It can be seen that the NPC-TPM discrete data applicable to different platforms is more in line with the non-linear model of the SVM-RBF kernel, and the SVM-RBF kernel is more suitable for constructing classifiers for different platforms.
[0061] Table 3
[0062]
[0063] Table 4
[0064]
[0065] Example 2
[0066] On the basis of Example 1, the PAS high-risk cohort of the Maternal and Child Health Hospital of City C (Life sequencing platform) was added, and the inclusion and exclusion criteria and the method of prospectively collecting sample data were the same as those in Example 1.
[0067] Among the high-risk pregnant women undergoing NIPT using the BGI sequencing platform, 70 PAS pregnant women and 210 non-PAS high-risk pregnant women were included in the training data set, and 35 PAS pregnant women and 95 non-PAS high-risk pregnant women were included in the internal validation data set for differential analysis and model training. In addition, 51 PAS pregnant women and 163 non-PAS high-risk pregnant women who underwent NIPT PLUS testing (higher sequencing depth) using the BGI sequencing platform at the Maternal and Child Health Hospital of City A, 54 PAS pregnant women and 162 non-PAS high-risk pregnant women enrolled in the Maternal and Child Health Hospital of City B, and 55 PAS pregnant women and 165 non-PAS high-risk pregnant women enrolled in the Maternal and Child Health Hospital of City C were respectively included in three independent external validation data sets (NIPT PLUS data set, Illumina validation set, Life validation set).
[0068] Annotate the cfDNA promoter nucleosome coverage profile of the NIPT data of the enrolled pregnant women using the same method as in Example 1, and input the NPC-TPM values into the classifier training system.
[0069] In the training data set, PAS recombinant pregnant women and non-PAS high-risk pregnant women were matched 1:1, with 35 samples in each group. DESeq2, limma-voom, and the rank sum test were used for analysis to screen genes with p-values < 0.05 calculated by all three differential analysis methods, and 702 differentially expressed pTSS-covered genes were obtained. As Figure 2 shown in the Venn diagram of differential gene screening.
[0070] Set the optimal cut-off value of 702 feature factors as the NPC-TPM value with the maximum sum of sensitivity and specificity in the training dataset. When the feature factor NPC (i.e., NPC-TPM) is greater than the corresponding optimal cut-off value, set it to 1; otherwise, set it to 0.
[0071] Input the selected feature factors into the RFE process, and use multiple machine learning methods to gradually construct a PAS disease prediction classifier. Select the SVM-RBF kernel and LR method respectively for classifier training.
[0072] Set PAS pregnant women as 1 and non-PAS high-risk pregnant women as 0 and input them into the system for classifier training. To improve the sensitivity of the classifier, set the parameter values of class_weight as {0:0.1, 1:0.3}.
[0073] Obtain the disease risk assessment results of the target to be predicted, and apply k-fold cross-validation (k = 10) to increase the evaluation robustness.
[0074] Extract the optimal feature factor combinations under different modeling methods, and output the optimal classifier and its evaluation results.
[0075] Using the LR training method, the output result is an LR classifier of 25 genes. Calculate the AUC, precision, sensitivity, and specificity of this classifier in the training dataset, internal validation dataset, and external validation dataset, as shown in Table 5. Figure 3 。
[0076] Table 5
[0077]
[0078]
[0079] Using the SVM-RBF kernel training method, the output result is an SVM-RBF kernel classifier of 23 genes. The discretization thresholds of each gene are shown in Table 6.
[0080] Table 6
[0081]
[0082] Calculate the AUC, precision, sensitivity, and specificity of this classifier in the training dataset, internal validation dataset, and external validation dataset, as shown in Table 7. Figure 4It can be seen that the SVM classifier has fewer genes and better prediction performance than the LR classifier. The 702 feature factor set selected after 1:1 matching of 35 PAS recombinant pregnant women and 35 non-PAS high-risk pregnant women is better than the 226 feature factor set selected from 54 PAS pregnant women and 54 non-PAS high-risk pregnant women. This result suggests that by integrating the expression levels of these 23 genes, namely ABHD1, ALG1L2, EYS, FAM157C, KDSR, KRT5, LANCL2, LINC00390, LINC00964, LOC105371998, LOC107987394, LOC644090, LYZL2, MIR184, MIR4802, MYT1L, NGDN, NSD2, PACRG.AS3, SAP30L.AS1, SLC16A12.AS1, TADA3, TMEM147.AS1, precise early prediction of PAS can be achieved, a detection method with high sensitivity and high specificity can be designed, and it is expected to develop a diagnostic kit widely used in clinical practice, promoting the development of personalized medicine and precision treatment.
[0083] Table 7
[0084]
[0085] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A classifier training system for early identification of the risk of placenta accreta in high-risk pregnant women, characterized by It includes: Dataset partitioning module: extracts pre-collected medical data of pregnant women at high risk of placenta accreta (PAS), and divides the data into discovery dataset, training dataset, internal validation dataset, and external validation dataset; Impact factor extraction module: The non-invasive prenatal DNA test (NIPT) data of PAS high-risk pregnant women collected in advance are used to annotate the whole-genome promoter nucleosome coverage spectrum of free DNA and extract features. The raw read coverage of the pTSS region is calculated and the raw read coverage of the pTSS region is normalized to obtain the TPM-normalized promoter coverage (NPC-TPM). The NPC-TPM corresponding to all genes is used as the impact factor; Feature screening module: Screen the influencing factors of high-risk pregnant women with and without PAS in the discovery data set to determine the characteristic factors; Characteristic factor discretization module: determine the optimal cutoff value for each characteristic factor and discretize the characteristic factor based on the optimal cutoff value; Model acquisition module: The screened feature factors are input into the feature recursive elimination (RFE) process, and a PAS disease prediction classifier is gradually constructed using multiple machine learning algorithms; the disease risk assessment results of the target to be predicted are obtained, and k-fold cross-validation is applied to increase the robustness of the assessment; the best feature factor combination is extracted, and the optimal classifier and its evaluation results are output; in the model acquisition module: the best feature factor combination is the target genes: ABHD1, ALG1L2, EYS, FAM157C, KDSR, KRT5, LANCL2, LINC00390, LINC00964, LOC105371998, LOC107987394, LOC644090, LYZL2, MIR184, MIR4802, MYT1L, NGDN, NSD2, PACRG.AS3, SAP30L.AS1, SLC16A12.AS1, TADA3, TMEM147.AS1.
2. The system according to claim 1, characterized in that In the dataset division module, pregnant women with PAS and those without PAS were matched according to age, gestational age at NIPT, fetal sex, and distribution of high-risk factors. The samples from the main center were randomly divided into a training dataset and an internal validation dataset, and the samples from each branch center served as independent external validation datasets. The discovery dataset was obtained by performing a 1:1 matching of the training dataset based on gestational age at NIPT and fetal sex.
3. The system according to claim 1, characterized in that In the impact factor extraction module, the NIPT data were aligned with the human reference genome hg19 using a sequence alignment algorithm, PCR duplicates were removed, the region 1000 bp above and below the transcription start site TSS was determined as the promoter region pTSS, and the raw read coverage of the pTSS region was calculated.
4. The system according to claim 3, characterized in that In the impact factor extraction module, the raw read coverage of the pTSS region is normalized to TPM-like normalization, and NPC-TPM is obtained by the following formula: Among them, NPC-TPM i represents the TPM-normalized promoter coverage of the pTSS region of the i-th gene, q i represents the raw read coverage of the pTSS region, l i represents the transcript length, ∑ j (q j / l j ) represents the sum of the read coverage of the pTSS regions of all genes in a sample after normalization of transcript length.
5. The system according to claim 1, characterized in that feature screening In this module, the influencing factors of high-risk pregnant women with and without PAS in the discovery dataset were analyzed by DESeq2, limma-voom and rank sum test, and the influencing factors with p values of <0.05 calculated by the three difference analysis methods were selected as characteristic factors.
6. The system according to claim 1, characterized in that the characteristic factor In the discrete module, to enhance the universality and clinical applicability of the classifier to different sequencing platforms, the optimal cutoff value of each feature factor was set to the NPC-TPM value with the maximum sum of sensitivity and specificity in the training dataset; When the characteristic factor NPC-TPM is greater than the corresponding optimal cutoff value, it is set to 1; otherwise, it is set to 0.
7. The system according to claim 1, characterized in that In the model acquisition module, a variety of machine learning algorithms are used to gradually build a PAS disease prediction classifier, and logistic regression LR, support vector machine SVM linear kernel and RBF kernel function are selected for classifier training.
8. The system according to claim 1, characterized in that In the model acquisition module: the optimal classifier uses the support vector machine SVM-RBF kernel.
9. A classifier for early identification of the risk of placenta accreta in high-risk pregnant women, characterized by It is obtained by the training system according to any one of claims 1-8.
Citation Information
Patent Citations
Target gene combination related to premature severe preeclampsia and application thereof
CN114822682A
Target gene combination related to premature delivery and application thereof
CN117672350A