Early pregnancy and preeclampsia risk early warning and typing method based on free DNA oligonucleotide characteristics
By obtaining low-coverage whole-genome sequencing data from pregnant women's plasma samples, extracting terminal motif sequences from cell-free DNA fragments in plasma, and constructing a deep learning model, the accuracy issues of risk assessment and subtyping of preeclampsia in early pregnancy were resolved. This enabled early warning and subtype differentiation, reduced testing costs, and improved testing stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-17
- Publication Date
- 2026-05-01
AI Technical Summary
Current technologies struggle to accurately assess the risk of preeclampsia in early pregnancy and differentiate between early and late-onset preeclampsia, leading to missed opportunities for clinical intervention and insufficient accuracy in predictive methods.
By obtaining low-coverage whole-genome sequencing data from plasma samples of pregnant women in early pregnancy, the terminal motif sequences of cell-free DNA fragments in the plasma are extracted, feature vectors are constructed, and input into a deep learning classification model to achieve risk scoring and clinical subtyping.
It achieves highly sensitive risk warning and accurate subtype classification in early pregnancy, reduces testing costs, improves the stability and repeatability of testing, and provides a valuable window of opportunity for early intervention.
Smart Images

Figure CN121963870A_ABST
Abstract
Description
A method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides. Technical Field
[0001] This invention relates to the field of medical testing technology, and more specifically, to a method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of free DNA oligonucleotides. Background Technology
[0002] Preeclampsia is a serious complication specific to pregnancy and one of the leading causes of maternal and perinatal mortality worldwide. Currently, there is a lack of effective means for accurate risk prediction in early pregnancy, causing many patients to miss the optimal window for early intervention. Developing a technology capable of precise risk assessment and classification of preeclampsia in early pregnancy is of great clinical significance for early warning, targeted monitoring, and timely intervention, and is expected to significantly improve maternal and infant outcomes.
[0003] Current technologies primarily rely on maternal clinical characteristics, ultrasound indicators, and serum protein biomarkers for risk assessment, but these methods have significant limitations. Clinical indicators such as hypertension and proteinuria typically appear only in the middle to late stages of the disease, failing to provide early warning. While serum biomarkers such as placental growth factor (PlGF) and soluble FMS-like tyrosine kinase-1 (sFlt-1) have some predictive value, their concentrations do not change significantly in early pregnancy, resulting in limited sensitivity and specificity. Some methods based on cell-free DNA concentration fail to fully utilize the molecular characteristics of cell-free DNA. More importantly, current technologies struggle to differentiate between different clinical subtypes of preeclampsia, particularly the crucial distinction between early-onset and late-onset types, which differ fundamentally in their pathological mechanisms, severity, and clinical management strategies. These shortcomings limit the accuracy and clinical applicability of existing predictive methods.
[0004] Therefore, this paper proposes a risk warning and classification method for preeclampsia in early pregnancy based on the characteristics of cell-free DNA oligonucleotides. The core issue to be addressed is how to overcome the deficiency of insufficient sensitivity of existing prediction technologies in early pregnancy and how to accurately distinguish different clinical subtypes of preeclampsia, thereby providing clinicians with an earlier and more accurate risk assessment tool. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method for early warning and typing of preeclampsia risk in early pregnancy based on the characteristics of free DNA oligonucleotides, in order to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for risk warning and classification of preeclampsia in early pregnancy based on the characteristics of cell-free DNA oligonucleotides, the method comprising the following steps: S1: acquiring low-coverage whole-genome sequencing data of plasma samples from pregnant women in early pregnancy; S2: extracting terminal motif sequences of cell-free DNA fragments from the sequencing data, wherein the terminal motifs are nucleotide sequences of a predetermined length at the ends of the fragments; S3: converting the terminal motif sequences into numerical feature vectors, wherein the feature vectors characterize the distribution and combination patterns of the terminal motifs in the samples; S4: inputting the feature vectors into a trained deep learning classification model, wherein the model automatically learns and outputs an overall risk score for preeclampsia; S5: generating high-risk or low-risk warning signals based on a comparison of the overall risk score with a preset threshold; S6: when the warning signal is high-risk, further determining the clinical classification of preeclampsia based on the intermediate layer output or branch network output of the deep learning classification model, wherein the clinical classification includes early-onset preeclampsia and late-onset preeclampsia.
[0007] Preferably, step S3 specifically includes: constructing a complete motif space containing all possible terminal motif types, wherein the complete motif space consists of all possible sequence combinations of length 2 to 5 nucleotides; calculating the absolute abundance or relative frequency of each terminal motif type in the sample; and organizing the absolute abundance or relative frequency into a feature vector of fixed dimension, wherein the dimension of the feature vector ranges from 100 to 500, and each dimension corresponds to a value of a specific terminal motif type.
[0008] Preferably, in step S3, the construction of the feature vector further includes: normalizing the absolute abundance or relative frequency of the terminal motifs to eliminate the influence of sequencing depth differences between samples; and the normalization process adopts the minimum-maximum scaling or Z-score normalization method.
[0009] Preferably, in step S4, the trained deep learning classification model is a convolutional neural network or a recurrent neural network; the training data of the model comes from the terminal motif feature vectors of multiple pregnant women's plasma samples and the corresponding clinical outcome labels, the clinical outcome labels including healthy, early-onset preeclampsia and late-onset preeclampsia; and the ratio of preeclampsia samples to healthy samples in the training data is adjusted to a range of 1:1 to 1:3 by oversampling or undersampling techniques.
[0010] Preferably, in step S4, the deep learning classification model employs an attention mechanism during training to automatically weight the contribution of different terminal motif features to risk prediction; the output of the attention mechanism is used to generate feature importance scores and visualize key terminal motifs to aid in explaining the model's decision-making process.
[0011] Preferably, in step S6, determining the clinical subtype of preeclampsia specifically includes: the deep learning classification model includes a shared feature extraction layer and multiple subtype output branches; the shared feature extraction layer processes the input feature vector; the multiple subtype output branches correspond to early-onset preeclampsia risk and late-onset preeclampsia risk, respectively; by calculating the output probability of each branch and comparing the magnitude of the output probabilities, the final clinical subtype is determined, wherein the type corresponding to the branch with the highest output probability is taken as the subtype result.
[0012] Preferably, in step S1, the sequencing coverage in the acquisition of low-coverage whole-genome sequencing data is 0.05 to 0.5 times; and the plasma sample is obtained from peripheral blood of pregnant women with a gestational age of 8 to 16 weeks, and is obtained after centrifugation, wherein the centrifugation includes centrifugation at 1600×g to 2000×g for 10 to 15 minutes.
[0013] Preferably, in step S2, the extraction of the terminal motif sequence of the cell-free DNA fragment from the plasma includes obtaining sequences from the paired ends of the cell-free DNA fragment; the paired-end sequences include a 5' end sequence and a 3' end sequence, which are treated as independent features; and the predetermined length of the terminal motif is 2 to 4 nucleotides.
[0014] Preferably, the method further includes step S7: S7: Integrating the warning signal and clinical classification results with the pregnant woman's basic information to generate a structured risk assessment report, which is output in electronic or paper form; wherein, the pregnant woman's basic information includes age, gestational age and body mass index, and the report also includes clinical recommendations based on the risk score and classification.
[0015] The technical effects and advantages of this invention are as follows: Compared with existing technologies, this invention achieves a significant improvement in technical performance by constructing a complete terminal motif feature space and employing a deep learning model for end-to-end pattern recognition. This method systematically extracts all possible sequence combinations of the last 2-5 nucleotides of cell-free DNA fragments in plasma, forming a high-dimensional feature vector. A deep learning model is then used to automatically learn the complex nonlinear relationship between these motif combinations and the pathogenesis of preeclampsia. This approach avoids the subjectivity and information omissions inherent in manual feature selection, and can capture weak but synergistic molecular signals that traditional methods cannot identify. Therefore, it maintains high sensitivity even with low sequencing coverage (0.05x-0.5x), significantly reducing detection costs and technical barriers.
[0016] Compared to existing technologies, this invention achieves a balance between interpretability and accurate subtyping in the prediction process by introducing an attention mechanism and a multi-branch network architecture. The attention mechanism in the deep learning model automatically evaluates the contribution weights of different terminal motifs to the prediction results, generating feature importance scores and making the model's decision-making process transparent and interpretable. Meanwhile, the network design, featuring a shared feature extraction layer and multiple subtyping output branches, enables the model to simultaneously learn the overall risk of preeclampsia and its subtype-specific patterns. This integrated design ensures that risk warning and clinical subtyping are based on the same source of molecular evidence, avoiding the accumulation of errors that may result from staged predictions, and providing clinicians with complete decision support from risk assessment to subtype determination.
[0017] Compared to existing technologies, this invention improves the stability and reproducibility of detection by separately processing and normalizing the paired-end sequences for feature construction. This method extracts the 5' and 3' terminal motif sequences of cell-free DNA fragments as independent features, fully utilizing the complete information at the fragment ends; simultaneously, it normalizes the motif abundance, effectively eliminating technical biases caused by differences in sequencing depth between samples. This approach can accurately capture subtle molecular feature changes in early pregnancy, enabling early prediction within the critical time window of 8-16 weeks of gestation, providing valuable time margin for clinical intervention. Attached Figure Description
[0018] Figure 1 is a flowchart of the overall method of the present invention.
[0019] Figure 2 is a diagram of the feature extraction and vectorization process of the present invention.
[0020] Figure 3 is a diagram of the deep learning model and classification mechanism of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1, as shown in Figures 1 to 3, is a method for early warning and classification of preeclampsia risk in early pregnancy based on cell-free DNA oligonucleotide features. It mainly includes the following core steps: S1, obtaining low-coverage whole-genome sequencing data from plasma samples of pregnant women in early pregnancy; S2, extracting terminal motif sequences of cell-free DNA fragments from the sequencing data; S3, converting the terminal motif sequences into numerical feature vectors; S4, inputting the feature vectors into a trained deep learning classification model to output an overall risk score for preeclampsia; S5, generating a warning signal based on the score; and S6, determining the clinical classification of preeclampsia when the warning signal indicates a high risk.
[0023] By combining the above steps, accurate risk warning and classification of preeclampsia can be achieved in the early stages of pregnancy, providing an important basis for clinical intervention.
[0024] Furthermore, the process of converting terminal motif sequences into numerical feature vectors specifically includes constructing a complete motif space, statistically analyzing motif abundance, and organizing feature vectors. This step is achieved by systematically enumerating all possible terminal motif sequences and calculating their distribution frequencies. The complete motif space consists of all possible sequence combinations ranging from 2 to 5 nucleotides in length. For example, for a motif of 3 nucleotides in length, the number of possible types is... After determining the absolute abundance of each terminal motif type in the sample, its relative frequency with respect to the total sequencing fragments was calculated. These values were then organized into a feature vector with dimensions ranging from 100 to 500, where each dimension corresponds to a value for a specific terminal motif type. This comprehensive coverage of all possible sequence combinations ensures the representativeness and completeness of the feature vector, providing a reliable data foundation for subsequent model analysis.
[0025] Furthermore, the construction of the feature vector also includes normalization of the terminal motif abundance, which is achieved by eliminating technical bias between samples through standardization techniques. Specifically, a min-max scaling method is used to scale the relative frequency of each terminal motif to the range of 0 to 1, calculated as follows: ,in Represents the original feature values. This represents the set of values for the feature across all samples. Alternatively, the Z-score standardization method can be used to transform the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula is: ,in The mean, The standard deviation is given. These treatments effectively eliminate technical variations caused by differences in sequencing depth, improving data comparability and consistency.
[0026] Furthermore, the trained deep learning classification model employs a convolutional neural network (CNN) or recurrent neural network (RNN) architecture and is trained on a specific ratio of training data. In practical applications, CNNs contain multiple convolutional layers, pooling layers, and fully connected layers, effectively capturing the spatial local correlations of terminal motif features. The training data comes from the terminal motif feature vectors and corresponding clinical outcome labels of multiple pregnant women's plasma samples. To ensure balanced model learning, oversampling techniques are used to adjust the ratio of preeclampsia samples to healthy samples to a range of 1:1 to 1:3. Model training employs the backpropagation algorithm, optimizing network parameters by minimizing the cross-entropy loss function. This process ensures the model's balanced learning ability across different categories of samples.
[0027] Furthermore, deep learning classification models incorporate an attention mechanism during training, achieved through the construction of a weight calculation layer and a visualization module. The attention mechanism automatically calculates the contribution weight of each terminal motif feature to the prediction result, using the following formula: ,in The energy score of the i-th feature is calculated through a fully connected layer. The total number of features is represented by these weights. These weights are visualized in the form of a heatmap, visually showing the key terminal motifs that have the greatest impact on prediction. This design makes the model's decision-making process transparent, allowing clinicians to understand the basis of the model's judgments, while also providing clues for research into biological mechanisms.
[0028] Furthermore, a shared feature extraction layer and a multi-branch output network architecture were used to determine the clinical subtyping of preeclampsia. The shared feature extraction layer consists of three fully connected layers that learn a general representation of terminal motif features; the two subtyping output branches correspond to the risk prediction of early-onset and late-onset preeclampsia, respectively. The model training uses a joint loss function: ,in The binary cross-entropy loss represents the overall risk prediction. This represents the multi-class cross-entropy loss for fractal tasks. and To balance hyperparameters, the final clinical classification is determined by comparing the output probabilities of each branch, with the branch corresponding to the highest output probability being used as the classification result.
[0029] Furthermore, obtaining low-coverage whole-genome sequencing data involves specific sequencing parameters and sample processing standards. Sequencing coverage is controlled within the range of 0.05 to 0.5 times, with 0.1 times coverage being preferred to balance cost and data quality. Plasma samples are derived from peripheral blood of pregnant women at 8 to 16 weeks of gestation and undergo a two-step centrifugation process: first at 1600... g to 2000 Centrifuge at 16000 g for 10 to 15 minutes to separate the plasma, then centrifuge at 16000 g for 10 to 15 minutes to separate the plasma. Centrifuge at 1000°C for 10 minutes to remove cell debris. This standardized sample processing procedure ensures the integrity of free DNA and the reliability of sequencing data.
[0030] Furthermore, the terminal motif sequences of cell-free DNA fragments extracted from plasma are processed separately, including the paired-end sequences. Nucleotide sequences of predetermined lengths, ranging from 2 to 4 nucleotides, are extracted from the 5' and 3' ends of each cell-free DNA fragment. The paired-end sequences are treated as independent features, and their abundance and distribution patterns are statistically analyzed separately. Finally, the feature vectors from the two ends are concatenated to form a complete input feature. This processing method fully utilizes the complete sequence information at the ends of cell-free DNA fragments, improving the sensitivity and specificity of detection.
[0031] Furthermore, this method also includes generating structured risk assessment reports. It integrates early warning signals and clinical classification results with basic information such as the pregnant woman's age, gestational age, and body mass index to generate a standardized report containing four parts: core conclusions, technical parameters, clinical recommendations, and explanations. The report can be output in both electronic and printed formats, facilitating clinical use in different scenarios. This comprehensive output approach provides clinicians with full decision support, effectively linking test results with clinical management.
[0032] Furthermore, the software implementation of this method can be carried on a computer-readable storage medium, and the computer program stored thereon, when executed by a processor, can fully implement all the above steps. The program adopts a modular design, including a data preprocessing module, a feature extraction module, a model inference module, and a report generation module. Each module exchanges data through standardized interfaces, ensuring the system's scalability and maintainability.
[0033] To illustrate this technical solution more clearly, a specific application scenario is provided below. In this scenario, the aim is to accurately predict and classify the risk of preeclampsia in pregnant women in early pregnancy. Sample processing is completed in a standard molecular biology laboratory equipped with centrifuges, and the data analysis module is deployed on a local server equipped with a GPU accelerator for deep learning model inference.
[0034] The detailed working process is as follows: First, the sample processing module receives a peripheral blood sample from a pregnant woman at 12 weeks of gestation, collecting 5 mL of venous blood using an EDTA anticoagulant tube. The sample is immediately processed at 1600... Centrifuge for 10 minutes at a centrifugation radius of 15 cm under 16000 g conditions, generating a relative centrifugal force of approximately 2000 g. Carefully aspirate the supernatant plasma, avoiding contact with the blood cell layer. Centrifuge the obtained plasma at 16000 g / L. Centrifuge again for 10 minutes under g conditions to completely remove residual cell debris, and collect the clear plasma supernatant for subsequent analysis. The entire centrifugation process is performed at 4°C to ensure the stability of free DNA.
[0035] Next, cell-free DNA was extracted from the processed plasma samples using a commercially available cell-free DNA extraction kit, strictly following the instructions. One mL of plasma sample was digested with proteinase K, and DNA was bound to a silica gel column. Finally, the sample was eluted with 30 μL of elution buffer to obtain cell-free DNA. The extracted cell-free DNA underwent quality testing, with DNA concentration measured using a quantitative real-time analyzer to ensure a concentration above 0.5 ng / μL and a DNA integrity index greater than or equal to 7.0. Sequencing libraries were then constructed using transposase assays, including end repair, A addition, adapter ligation, and PCR amplification. Low-coverage whole-genome sequencing was performed, with a sequencing depth controlled at 0.1-fold, using the Illumina sequencing platform to generate paired-end 150 bp reads, averaging approximately 5 million reads per sample.
[0036] After obtaining the sequencing data, the data analysis module begins its work. First, the raw sequencing data undergoes quality control, using the FastQC tool to assess data quality and remove aptamer sequences and low-quality reads. Then, paired-terminal motif sequences are extracted from each free DNA fragment from the quality-controlled data, extracting 3-nucleotide sequences from both the 5' and 3' ends. The frequency of all possible 3-nucleotide combinations in the sample is calculated, constructing a 64-dimensional feature vector. The value of each dimension is calculated as follows: ,in This represents the number of occurrences of the i-th basis order. This represents the total number of segments. The feature vectors are standardized using Z-score, using the formula... Each feature value is converted into a standard score, where and They were calculated from the training set data respectively.
[0037] The standardized feature vectors are then input into the trained deep learning model. This model employs a convolutional neural network architecture, containing two convolutional layers, one max-pooling layer, and three fully connected layers. The convolutional layers use the ReLU activation function with a kernel size of 3 to extract local feature patterns. The pooling layer uses a 22... 2. Max pooling was used to reduce feature dimensionality. The number of neurons in the fully connected layers was 128, 64, and 32, respectively. Finally, a risk score was output through the Sigmoid activation function. The model output that the pregnant woman's preeclampsia risk score was 0.82, which is higher than the preset threshold of 0.5, thus generating a high-risk warning signal.
[0038] The model is further analyzed through a subtyping branch. The subtyping network, based on shared feature representations, calculates the probabilities of early-onset and late-onset preeclampsia through two independent fully connected layers. The early-onset subtyping branch outputs a probability of 0.75, and the late-onset subtyping branch outputs a probability of 0.25. Based on the probability comparison, the subtyping result is determined to be high-risk early-onset preeclampsia. Throughout the prediction process, an attention mechanism calculates the importance weights of each terminal motif feature using a formula... A feature importance heatmap was generated, showing that multiple terminal motifs, led by "ACG" and "CGT", contributed the most to the prediction results, and the attention weights of these motifs all exceeded 0.85.
[0039] Finally, the system integrates all analysis results to generate a structured risk assessment report. The report includes the pregnant woman's basic information (age 28 years, gestational age 12 weeks, BMI 23.5), testing parameters (sequencing depth 0.1x, coverage 98.5%), risk score (0.82), warning level (high risk), and clinical classification (early onset). The report also provides specific clinical recommendations, including increased blood pressure monitoring frequency (twice weekly), regular urine protein testing, and consideration of low-dose aspirin intervention. The report is output in PDF format, supporting both electronic viewing and printing, and also generates machine-readable JSON data for hospital information system integration.
[0040] This scenario fully demonstrates how terminal motif feature analysis, deep learning models, and clinical subtyping can be organically combined to achieve accurate risk warning and subtyping of preeclampsia in the early stages of pregnancy. Through standardized sample processing, comprehensive feature extraction, and an advanced model architecture, a reliable decision support tool is provided for clinical practice. The entire process, from sample collection to report generation, can be completed within 36 hours, meeting the timeliness requirements of clinical applications. Simultaneously, the automated analysis process reduces manual operation time to within 2 hours, significantly improving testing efficiency.
[0041] Finally, several points should be noted: First, in the description of this application, it should be noted that, unless otherwise specified and limited, the terms "installation," "connection," and "linkage" should be interpreted broadly, and can refer to mechanical or electrical connections, or internal connections between two components, or direct connections. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may change. Second, the accompanying drawings of the embodiments disclosed in this invention only involve structures relevant to the embodiments disclosed in this invention; other structures can refer to common designs. Where there is no conflict, the same embodiment and different embodiments of this invention can be combined with each other. Finally, the above descriptions are merely preferred embodiments of this invention and are not intended to limit this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides, characterized in that, The method includes the following steps: S1: Obtaining low-coverage whole-genome sequencing data of plasma samples from pregnant women in early pregnancy; S2: Extracting terminal motif sequences of cell-free DNA fragments from the sequencing data, wherein the terminal motif is a nucleotide sequence of a predetermined length at the end of the fragment; S3: Converting the terminal motif sequences into numerical feature vectors, wherein the feature vectors characterize the distribution and combination patterns of the terminal motifs in the sample; S4: Inputting the feature vectors into a trained deep learning classification model, wherein the model automatically learns and outputs an overall risk score for preeclampsia; S5: Generating high-risk or low-risk warning signals based on a comparison of the overall risk score with a preset threshold; S6: When the warning signal is high-risk, further determining the clinical classification of preeclampsia based on the intermediate layer output or branch network output of the deep learning classification model, wherein the clinical classification includes early-onset preeclampsia and late-onset preeclampsia.
2. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 1, characterized in that, In step S3, the specific steps include: constructing a complete motif space containing all possible terminal motif types, wherein the complete motif space consists of all possible sequence combinations of length 2 to 5 nucleotides; calculating the absolute abundance or relative frequency of each terminal motif type in the sample; and organizing the absolute abundance or relative frequency into a feature vector of fixed dimension, wherein the dimension of the feature vector is between 100 and 500, and each dimension corresponds to a value of a specific terminal motif type.
3. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 2, characterized in that, In step S3, the construction of the feature vector further includes: normalizing the absolute abundance or relative frequency of the terminal motifs to eliminate the influence of sequencing depth differences between samples; and the normalization process adopts the minimum-maximum scaling or Z-score normalization method.
4. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 1, characterized in that, In step S4, the trained deep learning classification model is a convolutional neural network or a recurrent neural network; the training data of the model comes from the terminal motif feature vectors of multiple pregnant women's plasma samples and the corresponding clinical outcome labels, the clinical outcome labels include healthy, early-onset preeclampsia and late-onset preeclampsia; and the ratio of preeclampsia samples to healthy samples in the training data is adjusted to a range of 1:1 to 1:3 by oversampling or undersampling techniques.
5. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 4, characterized in that, In step S4, the deep learning classification model employs an attention mechanism during training to automatically weight the contribution of different terminal motif features to risk prediction; the output of the attention mechanism is used to generate feature importance scores and visualize key terminal motifs to aid in explaining the model's decision-making process.
6. The method for early warning and classification of preeclampsia risk in early pregnancy based on cell-free DNA oligonucleotide characteristics according to claim 1, characterized in that, In step S6, determining the clinical classification of preeclampsia specifically includes: the deep learning classification model contains a shared feature extraction layer and multiple classification output branches; the shared feature extraction layer processes the input feature vector; the multiple classification output branches correspond to the risk of early-onset preeclampsia and the risk of late-onset preeclampsia, respectively; by calculating the output probability of each branch and comparing the magnitude of the output probabilities, the final clinical classification is determined, wherein the type corresponding to the branch with the highest output probability is taken as the classification result.
7. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 1, characterized in that, In step S1, the sequencing coverage in the acquisition of low-coverage whole-genome sequencing data is 0.05 to 0.5 times; and the plasma sample is obtained from peripheral blood of pregnant women with a gestational age of 8 to 16 weeks, and is obtained after centrifugation, which includes centrifugation at 1600×g to 2000×g for 10 to 15 minutes.
8. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to claim 1, characterized in that, In step S2, the extraction of terminal motif sequences from cell-free DNA fragments from plasma includes obtaining sequences from the paired ends of the cell-free DNA fragments; the paired-end sequences include 5' end sequences and 3' end sequences, which are treated as independent features; and the predetermined length of the terminal motif is 2 to 4 nucleotides.
9. The method for early warning and classification of preeclampsia risk in early pregnancy based on the characteristics of cell-free DNA oligonucleotides according to any one of claims 1 to 8, characterized in that, The method further includes step S7: S7: Integrate the warning signal and clinical classification results with the pregnant woman's basic information to generate a structured risk assessment report, which is output in electronic or paper form; wherein, the pregnant woman's basic information includes age, gestational age and body mass index, and the report also includes clinical recommendations based on the risk score and classification.