Prediction method, device and equipment for pregnancy risk, storage medium and program product

By acquiring and screening the nucleic acid and clinical indicator characteristics of pregnant subjects, a parallel neural network model was constructed, which solved the problem of time-consuming and labor-intensive traditional pregnancy risk assessment and achieved efficient and accurate pregnancy risk prediction.

CN120998476APending Publication Date: 2025-11-21SHENZHEN HUADA GENE INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410622954.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional pregnancy risk assessment relies on doctors' expertise and experience, which is time-consuming, energy-intensive, and costly, and cannot efficiently predict pregnancy risks.

Method used

By acquiring target features of pregnant subjects, including nucleic acid features and clinical indicator features, a pregnancy risk prediction model is constructed. A parallel neural network structure is used for feature selection and model optimization to obtain prediction results, reduce redundant feature interference, and improve prediction accuracy.

Benefits of technology

It can accurately predict pregnancy risks without requiring a lot of doctors' time and energy, reducing costs and improving prediction accuracy. It is suitable for predicting the risks of premature birth and full-term pregnancy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998476A_ABST
    Figure CN120998476A_ABST
Patent Text Reader

Abstract

The invention relates to a pregnancy risk prediction method, device and equipment, a storage medium and a program product. The method comprises the steps of obtaining target features of a sample pregnancy subject; the target feature refers to a feature related to a preset pregnancy risk; inputting the target features of the sample pregnancy subject into a pregnancy risk prediction model; performing pregnancy risk prediction on the target pregnancy subject based on the pregnancy risk prediction model to obtain a prediction result; the prediction result is used for representing the probability of the preset pregnancy risk of the target pregnancy subject. By adopting the method, the cost of pregnancy risk prediction can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biotechnology, and in particular to a method, apparatus, device, storage medium, and program product for predicting pregnancy risks. Background Technology

[0002] With rapid societal development and constantly changing living environments, pregnant women face an increasing number of risk factors during pregnancy. Advanced maternal age, unhealthy lifestyle habits, chronic diseases, and genetic factors can all increase pregnancy risks. Therefore, accurate prediction of pregnancy risks is necessary to reduce adverse pregnancy outcomes.

[0003] In traditional techniques, doctors assess pregnancy risks using their professional knowledge and clinical experience. However, this method requires a significant amount of time and effort from doctors, inevitably leading to high costs. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, device, equipment, storage medium, and program product for predicting pregnancy risks that can reduce costs, in response to the above-mentioned technical problems.

[0005] In a first aspect, this application provides a method for predicting pregnancy risk, including:

[0006] Obtain the target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk.

[0007] Input the target characteristics of the sample pregnant subjects into the pregnancy risk prediction model;

[0008] Based on the pregnancy risk prediction model, the pregnancy risk of the target pregnant subject is predicted, and the prediction result is obtained; the prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

[0009] Secondly, this application also provides a pregnancy risk prediction device, comprising:

[0010] The acquisition module is used to acquire target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk.

[0011] The input module is used to input the target features of the sample pregnant subjects into the pregnancy risk prediction model;

[0012] The prediction module is used to predict the pregnancy risk of the target pregnant subject based on the pregnancy risk prediction model and obtain the prediction result; the prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

[0013] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method.

[0014] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0015] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method.

[0016] The aforementioned pregnancy risk prediction method, device, computer equipment, storage medium, and computer program product acquire the target features of the sample pregnant subjects. The target features refer to features related to the preset pregnancy risk. By removing redundant and irrelevant features, the model is avoided from being interfered with by irrelevant features, ensuring the accuracy of pregnancy risk prediction. Then, the target features of the sample pregnant subjects are input into the pregnancy risk prediction model. Based on the pregnancy risk prediction model, the pregnancy risk of the target pregnant subjects is predicted, and the prediction result is obtained. The prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk. Pregnancy risk can be accurately predicted without doctors spending a lot of time and energy, effectively reducing the cost of pregnancy risk prediction. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating a pregnancy risk prediction method provided in an embodiment of this application.

[0019] Figure 2 This is a receiver operating characteristic curve provided for an embodiment of this application.

[0020] Figure 3 This is a schematic diagram of the structure of a pregnancy risk prediction model provided in an embodiment of this application.

[0021] Figure 4 This is a structural block diagram of a pregnancy risk prediction device provided in an embodiment of this application.

[0022] Figure 5 This is an internal structural diagram of a computer device provided in an embodiment of this application.

[0023] Figure 6 This is an internal structural diagram of another computer device provided in an embodiment of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0025] In one exemplary embodiment, such as Figure 1 As shown, a method for predicting pregnancy risk is provided. Taking the application of this method to a computer device as an example, the method includes the following steps 102 to 106. Wherein:

[0026] Step 102: Obtain the target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk.

[0027] Pregnancy is the process by which an embryo or fetus grows and develops within the mother's body. The mother is the subject of pregnancy.

[0028] In some embodiments, the acquisition of target features of the sample pregnancy subject includes: acquiring nucleic acid features and clinical indicator features of the sample pregnancy subject using a computer device; and performing feature screening on the nucleic acid features and clinical indicator features to obtain the target features of the sample pregnancy subject.

[0029] In some embodiments, nucleic acid features may include features of at least one of the following nucleic acids: messenger RNA (mRNA), long noncoding RNA (lncRNA), microRNA (miRNA), and transfer RNA (tRNA).

[0030] In some embodiments, nucleic acid features may be, but are not limited to, features of cell-free ribonucleic acid (cfRNA) in plasma. Computer equipment can acquire nucleic acid features collected from plasma samples of the sampled pregnant subject.

[0031] In some embodiments, the sample pregnancy subject may be, but is not limited to, a pregnant woman. The nucleic acid characterization process may include, but is not limited to, obtaining peripheral blood from pregnant women at 16 weeks of gestation and immediately storing it at 4 degrees Celsius, performing plasma separation on the peripheral blood within 8 hours; immediately storing the plasma at -80 degrees Celsius for further processing after separation; wherein, the gestational age for obtaining peripheral blood from pregnant women may be limited to 11 to 25 weeks of gestation; adding Trizol LS to the plasma at a ratio of 1:3 and immediately shaking to mix, and then extracting cell-free ribonucleic acid from the plasma; sequencing and quantifying the expression profile of the cell-free ribonucleic acid from the plasma to obtain nucleic acid characteristics.

[0032] In some embodiments, the presupposed pregnancy risk may include at least one of the pregnancy risks such as preterm birth or full-term delivery. cfRNA sequencing can be performed using whole transcriptome sequencing, employing next-generation sequencing to sequence cfRNA in plasma samples from peripheral blood of pregnant women with spontaneous preterm or full-term delivery. The sequencing methods described above can simultaneously sequence multiple cfRNAs in plasma.

[0033] In some embodiments, at least one of the clinical indicator features or nucleic acid features may be, but is not limited to, collected from bodily fluids such as urine or amniotic fluid.

[0034] In some embodiments, the quantitative process of plasma cell-free ribonucleic acid expression profile includes: quality control of the raw cfRNA sequencing data, including cutting adapters, removing low-quality reads, removing reads shorter than 17 bp, and removing rRNA, value RNA, and Y RNA sequences; aligning the remaining reads after removal to the human transcriptome, in the order of miRNA, tRNA, and piRNA, mRNA and lncRNA, and finally other RNAs besides the aforementioned RNAs; and correcting the expression level of the predicted sample based on the existing training set expression level. RNA alignment is performed using Bowtie software, mRNA and lncRNA quantification is performed using RSEM software, and expression level is expressed as TPM, where formula (1) is the TPM calculation formula; miRNA and tRNA quantification is directly extracted from the alignment result file, and expression level is expressed as RPM, where formula (2) is the RPM calculation formula. Bowtie is an open-source software for aligning high-throughput sequencing data, mainly used to align DNA or RNA sequences to a reference genome. RSEM (RNA-Seq by Expectation-Maximization) is a software tool used to estimate gene or transcript expression levels from RNA-Seq data. Total Mapping Reads is the sum of all aligned reads. Total Mapping Reads is an important metric in RNA-Seq data analysis, reflecting the degree of matching between sequencing data and reference sequences, and providing crucial foundational data for subsequent analysis.

[0035] TPM=(Ni / Li)*1000000 / (sum(N1 / L1+N2 / L2+N3 / L3+…+Nn / Ln))(1).

[0036] RPM=Ni*1000000 / sum(N1+N2+N3+…Nn)(2).

[0037] Where Ni is the number of reads aligned to the i-th gene. Li is the length of the i-th gene; sum(N1 / L1+N2 / L2+...+Nn / Ln) is the sum of the values ​​of all (n) genes after standardization by length.

[0038] In some embodiments, clinical indicator characteristics may include at least one indicator such as plasma indicator data, liver function indicator data, lipid detection indicator data, or thyroid function indicator data.

[0039] In some embodiments, the nucleic acid characteristics and clinical indicators of the sample pregnant subjects correspond to the same period. It can be understood that the nucleic acid characteristics and clinical indicators of the sample pregnant subjects are collected from the sample pregnant subjects at the same time.

[0040] In some embodiments, the computer device may include at least one of a terminal or a server. The terminal may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices may include smartwatches, smart bracelets, head-mounted devices, etc. The server may be implemented using a standalone server or a server cluster consisting of multiple servers.

[0041] In some embodiments, the computer device may select target features from nucleic acid features based on the correlation between preset pregnancy risk and nucleic acid features. Target features may also be selected from clinical indicator features based on the correlation between preset pregnancy risk and clinical indicator features. Target features may include at least one of nucleic acid features or clinical indicator features.

[0042] In some embodiments, the computer device can perform feature filtering based on the correlation between candidate features and preset pregnancy risks, as well as the correlation between candidate features, to obtain the target features of the sample pregnancy subjects.

[0043] Step 104: Input the target features of the sample pregnant subjects into the pregnancy risk prediction model.

[0044] In some embodiments, the computer device can perform feature screening on nucleic acid features and clinical indicator features to obtain target features of the sample pregnant subjects. Then, a pregnancy risk prediction model is constructed based on these target features. After the pregnancy risk prediction model is built, the target features of the sample pregnant subjects can be input into the model. It can be seen that using features from both nucleic acid and clinical indicator dimensions ensures that the model learns more features and patterns, guaranteeing the accuracy of pregnancy risk prediction. Simultaneously, target features refer to features related to the preset pregnancy risk. Feature screening removes redundant and irrelevant features, preventing the model from being interfered with by irrelevant features and ensuring the accuracy of pregnancy risk prediction.

[0045] For example, a computer device can construct a pregnancy risk prediction model based on the risk labels and target features of a sample pregnant subject. The risk labels of the sample pregnant subject indicate the pregnancy risk that the subject may experience. The target features of the sample pregnant subject are input into the pregnancy risk prediction model, which then predicts the pregnancy risk for the subject, yielding an inference result. The pregnancy risk prediction model is optimized based on the difference between the inference result and the risk label, resulting in an optimized pregnancy risk prediction model. The inference result characterizes the probability that the sample pregnant subject will experience a predetermined pregnancy risk.

[0046] In some embodiments, the computer device can construct an input layer in the pregnancy risk prediction model that matches the number of target features. The input layer includes an input channel corresponding to each target feature. It is understood that each target feature has its own dedicated input channel, and inputting the target features into the pregnancy risk prediction model through the input layer ensures the orderly input of each target feature.

[0047] In some embodiments, computer devices can construct pregnancy risk prediction models based on parallel neural network structures. By optimizing and adjusting the parallel neural network structure, better prediction results can be achieved.

[0048] In some embodiments, the computer device can construct an output layer in the pregnancy risk prediction model that matches the number of preset pregnancy risks. The output layer includes neurons corresponding to each preset pregnancy risk. It can be understood that the neurons in the output layer are actually the neurons in the last level of the pregnancy risk prediction model.

[0049] In some embodiments, the computer device can construct a Receiver Operating Characteristic Curve (ROC) based on the inference results and risk labels. The pregnancy risk prediction model is then optimized based on the Area Under the Curve (AUC) to obtain an optimized pregnancy risk prediction model.

[0050] In some embodiments, such as Figure 2 As shown, receiver operating characteristic curves are provided. Pregnancy risk prediction models, including but not limited to preterm birth prediction models, are used to predict the probability of preterm or full-term birth in a pregnancy and are a type of binary classification model. A true positive is the number of pregnant women whose inference results indicate preterm birth and whose risk label is preterm birth. A false positive is the number of pregnant women whose inference results indicate preterm birth but whose risk label is full-term. Based on... Figure 2 The receiver operating characteristic curve (AUC) in the model can be used to determine the predictive accuracy of the pregnancy risk prediction model, which is 0.854.

[0051] Step 106: Based on the pregnancy risk prediction model, the pregnancy risk of the target pregnant subject is predicted, and the prediction results are obtained; the prediction results are used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

[0052] For example, a computer device can fuse target features of a target pregnancy subject based on a pregnancy risk prediction model to obtain fused features. Pregnancy risk is then predicted based on these fused features to obtain the prediction result.

[0053] In some embodiments, the pregnancy risk prediction model may include an input layer, at least one hidden layer, and an output layer. The input layer includes input channels corresponding to each target feature. The hidden layers and output layer correspond to different levels. The first hidden layer after the input layer corresponds to the first level, the next hidden layer after the first hidden layer corresponds to the next level after the first level, and so on, while the output layer corresponds to the last level. Each hidden layer and output layer includes at least one neuron. Each neuron is used to perform weighted fusion of the input data for that neuron.

[0054] In some embodiments, in the pregnancy risk prediction model, the levels gradually increase from the first level to the last level, with the number of neurons in lower levels exceeding the number of neurons in higher levels.

[0055] In the above-mentioned pregnancy risk prediction method, the target features of the sample pregnant subjects are obtained. The target features refer to the features related to the preset pregnancy risk. By removing redundant and irrelevant features, the model is avoided from being interfered with by irrelevant features, thus ensuring the accuracy of pregnancy risk prediction. Then, the target features of the sample pregnant subjects are input into the pregnancy risk prediction model. Based on the pregnancy risk prediction model, the pregnancy risk of the target pregnant subjects is predicted, and the prediction result is obtained. The prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk. Pregnancy risk can be accurately predicted without doctors spending a lot of time and energy, effectively reducing the cost of pregnancy risk prediction.

[0056] In some embodiments, feature screening is performed on nucleic acid features and clinical indicator features to obtain target features of the sample pregnancy subjects, including: determining a candidate feature set based on nucleic acid features and clinical indicator features; updating the current evaluation model based on the candidate feature for each candidate feature in the candidate feature set to obtain an intermediate evaluation model corresponding to the candidate feature; if the intermediate evaluation model corresponding to the candidate feature satisfies the performance gain condition, determining the intermediate evaluation model corresponding to the candidate feature as the new current evaluation model, and determining the candidate feature as the target feature, and removing the candidate feature from the candidate feature set; returning to the step of updating the current evaluation model based on the candidate feature for each candidate feature in the candidate feature set to obtain an intermediate evaluation model corresponding to the candidate feature and continuing to execute until the evaluation stop condition is met.

[0057] The performance gain condition is used to determine whether the performance gain of the updated preset evaluation model is sufficient compared to the original preset evaluation model. The evaluation stopping condition is used to determine when feature selection should stop. It can be understood that the more correlated a candidate feature is with the preset pregnancy risk, the better the performance of the intermediate evaluation model corresponding to that candidate feature.

[0058] For example, the computer device can use nucleic acid features and clinical features as candidate features to obtain a candidate feature set. The candidate feature set can contain at least two candidate features. Each candidate feature in the candidate feature set is added to the current evaluation model to obtain an intermediate evaluation model corresponding to each candidate feature. If the intermediate evaluation model corresponding to a candidate feature does not meet the performance gain condition, the intermediate evaluation model corresponding to the candidate feature is not used as the current evaluation model, and the candidate feature is not used as the target feature. If the intermediate evaluation model corresponding to a candidate feature meets the performance gain condition, the intermediate evaluation model corresponding to the candidate feature is used as the current evaluation model, and the current feature is used as the target feature of the pregnant subject in the sample. Then, the process of adding each candidate feature in the candidate feature set to the current evaluation model and subsequent steps continues until the number of target features in the preset evaluation model is not less than a preset number or the intermediate evaluation models corresponding to each candidate feature do not meet the performance gain condition.

[0059] In some embodiments, the performance gain condition may be, but is not limited to, the optimal performance among the intermediate evaluation models corresponding to each candidate feature, and the performance gain compared to the current evaluation model being no less than a preset gain threshold. The computer device can determine the intermediate evaluation model that satisfies the performance gain condition from the intermediate evaluation models corresponding to each candidate feature, use the candidate feature corresponding to the intermediate evaluation model that satisfies the performance gain condition as the target feature, and use the intermediate evaluation model that satisfies the performance gain condition as the new current evaluation model.

[0060] In some embodiments, the computer device can add candidate features as independent variables to the current evaluation model and determine the weights corresponding to the candidate features based on the training set, thereby obtaining an intermediate evaluation model corresponding to the candidate features. It can be understood that in the model, features can be independent variables, while the preset pregnancy risk is the dependent variable.

[0061] In some embodiments, the computer device may employ a forward feature selection method to screen nucleic acid features and clinical indicator features to obtain the target features of the sample pregnant subjects.

[0062] For example, feature selection is a data analysis step performed on a computer device. Suppose that through preliminary feature acquisition, nucleic acid features of 19,943 mRNAs, 15,477 lncRNAs, 2,138 miRNAs, and 49 tRNAs were collected. Simultaneously, 56 clinical indicator features were also collected, resulting in a total of 39,894 candidate features. To simplify the model and select features that play a crucial role in predicting pregnancy risk, a forward feature selection method is used for feature optimization.

[0063] During the feature selection process, the computer device can construct an initial current evaluation model, which may be, but is not limited to, an empty model. Candidate features are then continuously added to the current evaluation model, updating it accordingly. The goal of this step is to find the feature that contributes the most to the model's performance at each step, as shown in Equation 3.

[0064]

[0065] Where β is the weight of each feature, f is the feature, and y' is the current evaluation model.

[0066] Initially, neither the β nor the f vector contains any elements. Starting with an empty model, a current feature, f, is progressively added from 39,894 candidate features. For each f vector, the computer adds a weight β, which can be automatically learned from the training data. Target features are selected by comparing the performance of this intermediate evaluation model with the previous current evaluation model. The computer uses a likelihood ratio chi-square test to evaluate the improvement in model performance; features with the smallest p-value (p < 0.05) in the chi-square test are selected and added to the model, resulting in a new current evaluation model. Then, the best-performing candidate feature is selected for the next step. This forward selection process continues until no suitable candidate feature can further improve the model's performance (i.e., the p-value in the chi-square test is less than 0.05). Finally, these selected features are used to build a pregnancy risk prediction model.

[0067] Specific process: Assume an empty model: Model empty y = β0. By substituting each of the 39894 features into the model, a candidate feature f1 was found, which has the smallest p-value (p < 0.05) in the chi-square test. The new model will be: Model new y = β0 + β1f1. Thus, the current evaluation model is updated by adding a feature f1 and its corresponding weight. This process continues until no more features can significantly improve model performance. The final selected target features can include 1 clinical indicator feature, 1 tRNA feature, 16 lncRNA features, 4 miRNAs, and 73 mRNA features, for a total of 95 features.

[0068] In this embodiment, a candidate feature set is determined based on nucleic acid characteristics and clinical indicator characteristics. For each candidate feature in the candidate feature set, the current evaluation model is updated based on the candidate feature to obtain an intermediate evaluation model corresponding to the candidate feature. If the intermediate evaluation model corresponding to the candidate feature meets the performance gain condition, the intermediate evaluation model corresponding to the candidate feature is determined as the new current evaluation model, and the candidate feature is determined as the target feature. Candidate features are also removed from the candidate feature set. The process of updating the current evaluation model based on the candidate feature for each candidate feature in the candidate feature set to obtain the intermediate evaluation model corresponding to the candidate feature continues until the evaluation stop condition is met. This process can continuously screen out target features related to the preset pregnancy risk, remove redundant and irrelevant features, avoid interference from irrelevant features, and ensure the accuracy of pregnancy risk prediction.

[0069] In some embodiments, the pregnancy risk of a target pregnant subject is predicted based on a pregnancy risk prediction model to obtain a prediction result, including: inputting the target features of the target pregnant subject into the pregnancy risk prediction model through the input layer of the pregnancy risk prediction model; the input layer includes an input channel corresponding to each target feature; and predicting the pregnancy risk based on the target features of the target pregnant subject through the pregnancy risk prediction model to obtain a prediction result.

[0070] It is understandable that the target features of the target pregnancy subject are consistent with the target features of the sample pregnancy subjects. For example, if the target features obtained after feature screening are 1 clinical indicator feature, 1 tRNA feature, 16 lncRNA features, 4 miRNA features, and 73 mRNA features, then the target features of the target pregnancy subject are the same type of 1 clinical indicator feature, 1 tRNA feature, 16 lncRNA features, 4 miRNA features, and 73 mRNA features collected from the target pregnancy subject.

[0071] Optionally, the features, feature types, p-values, and feature importance values ​​used for feature selection are shown in Table 1 below. The p-value (Probability value) refers to the probability value obtained when testing statistical hypotheses (such as gene expression differences) on a feature. The feature importance values ​​are calculated using the SHAP (Shapley Additive exPlanations) algorithm.

[0072]

[0073]

[0074]

[0075]

[0076] Table 1

[0077] In some embodiments, a specific feature may be a feature whose feature importance value meets a preset condition. Optionally, the preset condition may be that the feature importance value of the specific feature is among the top preset values ​​among all features available for feature filtering. Further optionally, the preset value may be 5, that is, the specific feature may be a feature whose feature importance value is among the top 5 features.

[0078] In some embodiments, specific features may be TUBGCP3, SYNPO, MYO9B, SELENBP1 and / or LINC01702.

[0079] In some embodiments, the pregnancy risk prediction model can be a model built on a multilayer perceptron or a specific neural network model. For example, a computer device can build a multilayer perceptron model to predict the probability of a pregnant subject experiencing a predetermined pregnancy risk.

[0080] In some embodiments, the pregnancy risk prediction model may include a fully connected layer. A computer device can input various target features into the pregnancy risk prediction model. Each target feature is transmitted to the fully connected layer in the pregnancy risk prediction model through a corresponding input channel in the input layer. The fully connected layer performs fusion processing on the target features to obtain a fusion result. Based on the fusion result, pregnancy risk is predicted to obtain the prediction result.

[0081] In this embodiment, the target features of the target pregnant subject are input into the pregnancy risk prediction model through the input layer. The input layer includes an input channel corresponding to each target feature, which can ensure the orderly input of the target features. Thus, the pregnancy risk prediction model can accurately predict the pregnancy risk based on the target features of the target pregnant subject and obtain the prediction result.

[0082] In some embodiments, pregnancy risk prediction is performed based on the target features of the target pregnant subject using a pregnancy risk prediction model to obtain prediction results. This includes: performing multi-level fusion processing on the target features of the target pregnant subject using the pregnancy risk prediction model to obtain multi-level fusion results; wherein the fusion result of the next level below the current level is obtained by fusing based on the fusion result of the current level; and performing pregnancy risk prediction based on the multi-level fusion results to obtain prediction results.

[0083] In some embodiments, the first-level fusion processing refers to a nonlinear mapping after weighted fusion of target features. The non-first-level fusion processing refers to a nonlinear mapping after weighted fusion of the fusion results from the previous level.

[0084] In some embodiments, the pregnancy risk prediction model may include multiple fully connected layers corresponding to different levels. A computer device can input various target features into the pregnancy risk prediction model and perform multi-level fusion processing on these features through multiple fully connected layers corresponding to different levels, obtaining a multi-level fusion result. The output of the fully connected layer corresponding to the current level is used as the input of the fully connected layer of the next level; that is, the fusion result of the current level is used as the input of the fully connected layer of the next level, and the fully connected layer of the next level is used to perform further fusion of the fusion result of the current level to obtain the fusion result of the next level.

[0085] In some embodiments, a fully connected layer includes at least one neuron. It can be understood that neurons at each level constitute the fully connected layer corresponding to that level.

[0086] In some embodiments, the fully connected layer is used to perform a weighted fusion of the received input data and then apply an activation function for nonlinear mapping. The activation function may be, but is not limited to, the ReLU (Rectified Linear Unit) function.

[0087] In this embodiment, the target features of the target pregnancy subject are fused in multiple levels through a pregnancy risk prediction model, which can capture deeper features, process complex patterns, and obtain richer multi-level fusion results. The fusion result of the next level is obtained by fusing the fusion result of the current level. Thus, accurate pregnancy risk prediction can be performed based on the multi-level fusion results to obtain the prediction result.

[0088] In some embodiments, the pregnancy risk prediction model includes multi-level neurons; the fusion result of each level includes the output of each neuron in each level; each neuron in the last level corresponds to a different preset pregnancy risk; pregnancy risk prediction is performed based on the multi-level fusion result to obtain the prediction result, including: for each neuron in the last level, normalization processing is performed according to the output of the neuron to obtain the prediction result corresponding to the neuron; the prediction result corresponding to the neuron is used to characterize the probability of the target pregnancy subject experiencing the preset pregnancy risk corresponding to the neuron.

[0089] In some embodiments, the pregnancy risk prediction model includes a normalization layer. The output of each neuron in the last layer is used as the input to the normalization layer, and the normalization process is performed by the normalization layer to obtain the prediction result corresponding to that neuron at the output of the normalization layer.

[0090] In some embodiments, the output of each neuron in each layer is used as the input of each neuron in the next layer. Each neuron in the next layer performs a weighted fusion and nonlinear mapping on the outputs of each neuron in each layer to obtain the output of the neuron in the next layer.

[0091] In some embodiments, both preterm birth and full term are considered pre-defined pregnancy risks. The neurons in the last level correspond to these two pre-defined pregnancy risks, namely preterm birth and full term; that is, the pregnancy risk prediction model can be a binary classification model. The output of the neuron corresponding to preterm birth is normalized to obtain the prediction result for preterm birth, i.e., the probability of preterm birth in the target pregnancy. The output of the neuron corresponding to full term birth is normalized to obtain the pre-defined result for full term birth, i.e., the probability of full term birth in the target pregnancy.

[0092] In some embodiments, such as Figure 3The diagram shows the structure of a pregnancy risk prediction model. The model includes an input layer, hidden layers corresponding to the first level, hidden layers in the next level after the first level, and an output layer corresponding to the last level. Both the hidden and output layers are fully connected layers. The fully connected layers receive input data and generate output using the ReLU function. The processing of the i-th fully connected layer can be described as: L i =[Relu(L i-1 ·W i-1 +b i-1 )], L i = (x1, x2, ... x n );W i = (w1, w2…w n Relu = max(0, L). Where L i Let L1 represent the i-th fully connected layer, where i > 1. L0 represents the input layer. W i Let b be the weight matrix of the i-th fully connected layer. i Let be the bias vector of the i-th fully connected layer.

[0093] In the pregnancy risk prediction model, the input layer consists of 93 target features selected through feature screening, including 1 clinical indicator feature, 73 mRNAs, 16 incRNAs, 4 miRNAs, and 1 tRNA. Following the input layer are three fully connected layers: a hidden layer corresponding to the first layer, a hidden layer corresponding to the middle layer, and an output layer corresponding to the final layer. The number of neurons in each fully connected layer decreases as the layer level increases. For example, the first fully connected layer consists of 728 neurons, the second consists of 128 neurons, and the third consists of 2 neurons. The output of the third fully connected layer represents the raw values ​​for preterm and full-term births. These raw values ​​are normalized and converted into probability values ​​to obtain the prediction result. This design of three fully connected layers allows the model to handle large amounts of input data and learn and extract complex patterns from this data. ReLU (Rectified Linear Unit) is used as the activation function between the fully connected layers, effectively solving the gradient vanishing problem without affecting forward propagation. This allows the pregnancy risk prediction model to converge faster during training and to better learn and understand the characteristics of the input data. At the end of the pregnancy risk prediction model, a normalization (softmax) layer is used to transform the output of the previous layer into probability values ​​between 0 and 1. The main function of the softmax function is to transform the input real-valued vector into a probability distribution, allowing the pregnancy risk prediction model to provide a probability prediction for each possible output category. This approach enables the model not only to predict whether preterm birth will occur but also to provide the confidence level of that prediction. The area under the ROC curve (AUC) can be chosen as the performance metric for the pregnancy risk prediction model.

[0094] The softmax process can be described as follows: Among them, z k The output value of the k-th neuron, where C is the number of neurons, i.e. the number of categories. For a binary classification problem of preterm and full-term births, C = 2.

[0095] In this embodiment, for each neuron in the last level, the output of the neuron is normalized to obtain the prediction result corresponding to the neuron. The prediction result corresponding to the neuron is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk corresponding to the neuron. It can accurately predict the probability of the preset pregnancy risk, which can be used as a reference for doctors without consuming a lot of doctors' energy and can reduce costs.

[0096] In some embodiments, obtaining the nucleic acid characteristics and clinical indicators of the sample pregnant subject includes: obtaining the nucleic acid characteristics and clinical indicators of the sample pregnant subject within a preset collection time window; the preset collection time window covers a preset non-invasive screening time window.

[0097] It is understood that the selection of target features required for the pregnancy risk prediction model provided in this application is based on the prenatal pregnancy subject, such as the general pregnant woman population, including primiparous and multiparous women, and does not require differentiation of different subtypes of preterm birth, which is more conducive to clinical application. In addition, the preset collection time window for target features covers the preset non-invasive prenatal screening time window; for example, the preset collection time window can be from 11 to 25 weeks of gestation, which is an earlier gestational age, providing a longer time window for preterm birth intervention. Moreover, the preset collection time window covering the preset non-invasive prenatal screening time window is more conducive to clinical promotion. In specific clinical applications, target features can be collected when the pregnant subject undergoes non-invasive prenatal screening, for example, by collecting target features from the remaining plasma sample from the non-invasive prenatal screening, as well as the clinical indicator characteristics of the pregnant subject at the time of non-invasive prenatal screening.

[0098] It should be noted that the pregnancy risk prediction model provided in this application is not sensitive to the actual collection time of features and has high adaptability. Target features outside the preset collection time window can still be used to predict pregnancy risk using the pregnancy risk prediction model.

[0099] In this embodiment, nucleic acid characteristics and clinical indicators of the pregnant subjects are acquired within a preset collection time window. This preset collection time window covers a preset non-invasive screening time window. Compared to traditional techniques that rely solely on a single type of RNA or solely on clinical information, the combined use of biomarkers (i.e., nucleic acid characteristics) and clinical data (i.e., clinical indicator characteristics) provides a more accurate and reliable method for predicting pregnancy risks. By screening and analyzing nucleic acid characteristics and clinical indicator characteristics, the accuracy of pregnancy risk prediction is improved, thereby aiding in the diagnosis and intervention of pregnancy risks, improving pregnancy outcomes, and preventing related complications.

[0100] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0101] Based on the same inventive concept, this application also provides a pregnancy risk prediction device for implementing the pregnancy risk prediction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more pregnancy risk prediction device embodiments provided below can be found in the limitations of the pregnancy risk prediction method described above, and will not be repeated here.

[0102] In one exemplary embodiment, such as Figure 4 As shown, a pregnancy risk prediction device 400 is provided, including: an acquisition module 402, an input module 404, and a prediction module 406, wherein:

[0103] The acquisition module 402 is used to acquire the target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk.

[0104] Input module 404 is used to input the target features of the sample pregnant subjects into the pregnancy risk prediction model.

[0105] The prediction module 406 is used to predict the pregnancy risk of the target pregnant subject based on the pregnancy risk prediction model and obtain the prediction result; the prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

[0106] In some embodiments, the acquisition module 402 is used to acquire the nucleic acid characteristics and clinical indicator characteristics of the sample pregnant subject; and to perform feature screening on the nucleic acid characteristics and clinical indicator characteristics to obtain the target characteristics of the sample pregnant subject.

[0107] In some embodiments, the apparatus further includes a screening module 408, configured to: determine a set of candidate features based on the nucleic acid features and the clinical indicator features; update the current evaluation model for each candidate feature in the set of candidate features to obtain an intermediate evaluation model corresponding to the candidate feature; if the intermediate evaluation model corresponding to the candidate feature satisfies the performance gain condition, determine the intermediate evaluation model corresponding to the candidate feature as the new current evaluation model, determine the candidate feature as the target feature, and remove the candidate feature from the set of candidate features; return to the step of updating the current evaluation model for each candidate feature in the set of candidate features to obtain an intermediate evaluation model corresponding to the candidate feature, and continue execution until the evaluation stop condition is met.

[0108] In some embodiments, the prediction module 406 is used to input the target features of the target pregnant subject into the pregnancy risk prediction model through the input layer of the pregnancy risk prediction model; the input layer includes an input channel corresponding to each target feature; and the pregnancy risk prediction model is used to predict the pregnancy risk based on the target features of the target pregnant subject to obtain a prediction result.

[0109] In some embodiments, the prediction module 406 is used to perform multi-level fusion processing on the target features of the target pregnancy subject through the pregnancy risk prediction model to obtain multi-level fusion results; wherein, the fusion result of the next level of the current level is obtained by fusion based on the fusion result of the current level; and pregnancy risk prediction is performed based on the multi-level fusion results to obtain prediction results.

[0110] In some embodiments, the prediction module 406 is used to perform normalization processing on the output of each neuron in the last layer to obtain the prediction result corresponding to the neuron; the prediction result corresponding to the neuron is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk corresponding to the neuron.

[0111] In some embodiments, the acquisition module 402 is used to acquire the nucleic acid characteristics and clinical indicator characteristics of the sample pregnant subject within a preset collection time window; the preset collection time window covers a preset non-invasive screening time window.

[0112] The modules in the aforementioned pregnancy risk prediction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0113] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores target characteristics. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a pregnancy risk prediction method.

[0114] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a pregnancy risk prediction method. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0115] Those skilled in the art will understand that Figure 5 or Figure 6The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0116] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0117] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0118] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0120] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0122] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for predicting pregnancy risk, characterized in that, The method includes: Obtain the target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk. Input the target characteristics of the sample pregnant subjects into the pregnancy risk prediction model; Based on the pregnancy risk prediction model, the pregnancy risk of the target pregnant subject is predicted, and the prediction result is obtained; the prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

2. The method according to claim 1, characterized in that, The target features for obtaining the sample pregnant subjects include: Obtain the nucleic acid characteristics and clinical indicators of the pregnant subjects in the sample; Feature screening is performed on the nucleic acid features and the clinical indicator features to obtain the target features of the pregnant subjects in the sample.

3. The method according to claim 2, characterized in that, The feature screening of the nucleic acid features and the clinical indicator features to obtain the target features of the sample pregnancy subjects includes: A candidate feature set is determined based on the nucleic acid characteristics and the clinical indicator characteristics; For each candidate feature in the candidate feature set, the current evaluation model is updated based on the candidate feature to obtain the intermediate evaluation model corresponding to the candidate feature; If the intermediate evaluation model corresponding to the candidate feature satisfies the performance gain condition, the intermediate evaluation model corresponding to the candidate feature is determined as the new current evaluation model, the candidate feature is determined as the target feature, and the candidate feature is removed from the candidate feature set. The process of returning to each candidate feature in the candidate feature set, updating the current evaluation model based on the candidate feature, and obtaining the intermediate evaluation model corresponding to the candidate feature continues until the evaluation stop condition is met.

4. The method according to claim 1, characterized in that, The specific features are TUBGCP3, SYNPO, MYO9B, SELENBP1 and / or LINC01702.

5. The method according to claim 1, characterized in that, The process of predicting pregnancy risk for a target pregnancy based on the pregnancy risk prediction model, and obtaining prediction results, includes: The target features of the target pregnancy subject are input into the pregnancy risk prediction model through the input layer; the input layer includes an input channel corresponding to each target feature; The pregnancy risk prediction model is used to predict pregnancy risk based on the target characteristics of the target pregnant subject, and the prediction results are obtained.

6. The method according to claim 5, characterized in that, The process of predicting pregnancy risk based on the target characteristics of the target pregnancy subject using the pregnancy risk prediction model, and obtaining the prediction results, includes: The pregnancy risk prediction model performs multi-level fusion processing on the target features of the target pregnancy subject to obtain multi-level fusion results; wherein, the fusion result of the next level is obtained by fusing the fusion result of the current level. Pregnancy risk is predicted based on the multi-level fusion results, and the prediction results are obtained.

7. The method according to claim 6, characterized in that, The pregnancy risk prediction model includes multiple levels of neurons; the fusion result of each level includes the output of each neuron in each level; each neuron in the last level corresponds to a different preset pregnancy risk. The pregnancy risk prediction based on the multi-level fusion results yields prediction results including: For each neuron in the last level, the output of the neuron is normalized to obtain the prediction result corresponding to the neuron. The prediction result corresponding to the neuron is used to characterize the probability that the target pregnant subject will experience the preset pregnancy risk corresponding to the neuron.

8. A pregnancy risk prediction device, characterized in that, The device includes: The acquisition module is used to acquire target features of the sample pregnant subjects; the target features refer to features related to the preset pregnancy risk. The input module is used to input the target features of the sample pregnant subjects into the pregnancy risk prediction model; The prediction module is used to predict the pregnancy risk of the target pregnant subject based on the pregnancy risk prediction model and obtain the prediction result; the prediction result is used to characterize the probability of the target pregnant subject experiencing the preset pregnancy risk.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.