Construction method, prediction method and prediction device of pre-eclampsia prediction model based on time sequence multiple modes
By constructing a missing value fitting model based on graph attention network and fully connected network, the problem of missing and irregularity in the preeclampsia detection data is solved by traditional methods, and a more accurate preeclampsia prediction is achieved.
Patent Information
- Application Number
- CN202510111179.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Traditional time series analysis methods are difficult to effectively deal with the missing and irregularities in preeclampsia detection data, resulting in inaccurate prediction of preeclampsia.
A missing value fit model based on graph attention network and fully connected network is constructed, and the potential correlation between pregnant women's health data is captured through graph attention network, and the timing characteristics are captured and predicted in combination with the gated loop unit.
Effectively dealing with missing and irregularities in the data improves the prediction accuracy of preeclampsia and reduces the negative impact of missing values on model prediction accuracy.
Smart Images

Figure CN120015336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of preeclampsia risk prediction, and in particular to a method for constructing a preeclampsia prediction model based on time series multimodality, a prediction method and a prediction device. Background Art
[0002] Preeclampsia (PE) is a multi-system disease unique to pregnancy, posing a major safety threat to pregnant women and perinatal infants. Currently, the diagnosis of preeclampsia mainly relies on indicators such as hypertension and urine protein, but these methods have obvious lags and cannot meet the needs of early prediction. In addition, other characteristic data that are highly correlated with preeclampsia are often ignored during clinical testing.
[0003] The pathogenesis of preeclampsia has not yet been fully elucidated, and because electronic health record (EHR) data are usually missing and irregular, traditional time series analysis methods are difficult to effectively process these data, resulting in the inability to fully explore the potential factors and characteristics that affect the occurrence of preeclampsia, and the inability to accurately predict preeclampsia.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] Based on the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide a method for constructing a preeclampsia prediction model based on time series multimodality, a prediction method and a prediction device, aiming to solve the problem that traditional time series analysis methods are difficult to effectively handle missing data, resulting in inaccurate preeclampsia prediction.
[0006] The technical solution of the present invention is as follows:
[0007] A first aspect of the present invention provides a method for constructing a preeclampsia prediction model based on time series multimodality, which comprises the following steps:
[0008] Obtain multimodal time series feature data of pregnant women with known pregnancy outcomes and construct a time series dataset; pregnant women with known pregnancy outcomes include pregnant women with preeclampsia and pregnant women without preeclampsia;
[0009] Using the time series dataset to train the graph attention network and the fully connected network to obtain a missing value fitting model;
[0010] The missing value fitting model is embedded into a gated recurrent unit as an embedding layer, and after training with the time series data set, a preeclampsia prediction model is obtained.
[0011] Optionally, the steps of obtaining multimodal time series feature data of pregnant women with known pregnancy outcomes and constructing a time series dataset specifically include:
[0012] The gestational period of pregnant women with known pregnancy outcomes is divided into multiple time nodes, and the multimodal feature data of pregnant women with known pregnancy outcomes at each time node are collected. The multimodal feature data of each time node are merged, and the feature data that does not appear at the current time node is set to a null value, and the feature data that appears repeatedly is taken as the feature data that appears the last time, so as to obtain the preprocessed multimodal time series feature data;
[0013] The number of times the feature data appears is greater than a preset value as a screening criterion, and the pre-processed multimodal time series feature data is screened to construct a time series data set.
[0014] Optionally, the step of constructing a time series data set after screening the preprocessed multimodal time series feature data is based on the number of times the feature data appears being greater than a preset value as a screening criterion, and specifically includes:
[0015] Taking the number of times the feature data appears greater than a preset value as a screening criterion, the preprocessed multimodal time series feature data is screened to obtain screened multimodal time series feature data; the screened multimodal time series feature data includes constant features and variable features;
[0016] The constant features are filled in the missing positions of the constant features in the filtered multimodal time series feature data to obtain the filled multimodal time series feature data and construct a time series data set.
[0017] Optionally, the step of using the time series data set to train the graph attention network and the fully connected network to obtain the missing value fitting model specifically includes:
[0018] Converting the time series data set into graph structure data, and masking the characteristic values of a part of the nodes in the graph structure data as vacant values to form missing nodes, thereby obtaining graph structure data with missing nodes;
[0019] After inputting the graph structure data with missing nodes into the graph attention network, outputting first data;
[0020] Fusing the first data with the graph structure data to obtain fused feature data;
[0021] Perform feature expansion on the fused feature data to obtain one-dimensional feature data;
[0022] Inputting the one-dimensional feature data into a three-layer stacked fully connected network to obtain a filling value for the vacancy value and a preeclampsia prediction result;
[0023] A first loss value is obtained according to the filling value and the masked true value at the node, and a second loss value is obtained according to the preeclampsia prediction result and the preeclampsia true result;
[0024] The parameters in the graph attention network and the fully connected network are adjusted according to the first loss value and the second loss value until the weighted sum of the first loss value and the second loss value converges. Then the trained graph attention network and the two-layer stacked fully connected networks in the three-layer stacked fully connected network connected to the graph attention network constitute a missing value fitting model.
[0025] Optionally, after setting an information transmission mechanism in the graph attention network so that missing nodes can only receive information from adjacent nodes but not transmit information to adjacent nodes, the graph structure data with missing nodes is input into the graph attention network and the first data is output.
[0026] Optionally, the missing value fitting model is embedded into a gated recurrent unit as an embedding layer, and after training with the time series data set, the step of obtaining a preeclampsia prediction model specifically includes:
[0027] The missing value fitting model is embedded as an embedding layer into the input end of the gated recurrent unit, and the time series data set is converted into graph structure data as input to obtain the probability of correctly predicting whether the pregnant woman suffers from preeclampsia;
[0028] According to the probability of correctly predicting whether a pregnant woman has preeclampsia and the actual situation of whether a pregnant woman has preeclampsia, a cross entropy loss function is obtained;
[0029] The gated recurrent unit embedded with the missing value fitting model is trained according to the cross entropy loss function until the cross entropy loss function converges to obtain the preeclampsia prediction model.
[0030] A second aspect of the present invention provides a method for predicting preeclampsia, wherein the multimodal time series feature data of the pregnant woman to be tested is input into a preeclampsia prediction model constructed using the construction method of the present invention as described above, to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
[0031] A third aspect of the present invention provides a device for predicting preeclampsia, comprising:
[0032] The prediction unit is used to input the multimodal time series feature data of the pregnant woman to be tested into the preeclampsia prediction model constructed by the construction method of the present invention as described above, so as to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
[0033] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the prediction method of the present invention as described above is implemented.
[0034] According to a fifth aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the computer program is executed by the processor, the prediction method of the present invention as described above is implemented.
[0035] Beneficial effects: The present invention obtains a missing value fitting model through graph attention network and fully connected grid construction, which can effectively handle the missing and irregularities in the test data of pregnant women during pregnancy (such as medical indicators, physical signs, and HER data, etc.), can extract effective features from highly missing data, and use graph attention network to capture the potential correlation between various health data of pregnant women, and minimize the negative impact of missing values on the prediction accuracy of the model. Through the special processing of missing data by the missing value fitting model, the missing information can be effectively restored to ensure that the model can process incomplete medical data more accurately. At the same time, combined with the gated recurrent unit to capture and predict the timing features, the prediction accuracy of preeclampsia is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Schematic diagram of the process of constructing a preeclampsia prediction model based on time series multimodality.
[0037] Figure 2 The following is a statistical chart of partial data and data gaps.
[0038] Figure 3 A schematic diagram of constructing graph structured data with features as nodes.
[0039] Figure 4 Figure 2 is a diagram of information transmission adjustment strategy, where (a) is a schematic diagram of node information transmission, (b) is a schematic diagram of replacing some nodes with vacant values, (c) is a schematic diagram of vacant nodes transmitting noise signals to adjacent nodes, and (d) is a schematic diagram of adjusting the information transmission strategy to reduce the impact of vacant values.
[0040] Figure 5 This is a framework diagram of the preeclampsia prediction model based on time series multimodality.
[0041] Figure 6 This is a diagram of the preeclampsia prediction model based on time series multimodality.
[0042] Figure 7 Confusion matrix for preeclampsia prediction.
[0043] Figure 8A diagram showing the influence of different features on the prediction effect of the preeclampsia prediction model. DETAILED DESCRIPTION
[0044] The present invention provides a method for constructing a preeclampsia prediction model based on time series multimodality, a prediction method and a prediction device. In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0045] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0046] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their collections. It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms of "one", "an" and "the" are intended to include plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the items listed in the association and all possible combinations, and includes these combinations. As used in the present specification and the appended claims, the term "if" can be interpreted as "when..." or "once" or "in response to determination" or "in response to detection" according to the context. Similarly, the phrase "if determined" or "if [described condition or event] is detected" can be interpreted to mean "once determined" or "in response to determining" or "once [described condition or event] is detected" or "in response to detecting [described condition or event]" according to the context. The technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0047] The embodiment of the present invention provides a method for constructing a preeclampsia prediction model based on time series multimodality, wherein: Figure 1 As shown, the construction method comprises the following steps:
[0048] S1. Obtain multimodal time series feature data of pregnant women with known pregnancy outcomes and construct a time series data set (containing multidimensional feature data); pregnant women with known pregnancy outcomes include pregnant women with preeclampsia and pregnant women without preeclampsia;
[0049] S2. Using the time series dataset, a graph attention network (i.e., an attention-driven graph convolutional network, referred to as GAT) and a fully connected network are trained to obtain a missing value fitting model;
[0050] S3 embeds the missing value fitting model as an embedding layer into a gated recurrent unit (i.e., a recurrent neural network unit with a gating mechanism, referred to as GRU), and obtains a preeclampsia prediction model after training using the time series data set.
[0051] In practice, Figure 2 As shown, due to the sampling process at fixed time intervals, some data of pregnant women will be missing in most cases, and the missing rate is as high as 90%, and it cannot be guaranteed that pregnant women can do all the tests at the time nodes of fixed time intervals. Therefore, the present invention obtains a missing value fitting model through GAT and a fully connected grid construction, which can effectively handle the missing and irregularities in the test data of pregnant women during pregnancy (such as medical indicators, physical signs, and HER data, etc.), can extract effective features from highly missing data, and use GAT to capture the potential correlation between various health data of pregnant women, and minimize the negative impact of missing values on the prediction accuracy of the model. Through the special treatment of missing data by the missing value fitting model, the missing information can be effectively restored to ensure that the model can more accurately process incomplete medical data, and at the same time, the capture and prediction of time series features in combination with GRU effectively improves the prediction accuracy of preeclampsia.
[0052] Therefore, the preeclampsia prediction model constructed by the present invention can effectively extract features and more effective features in the case of missing data by integrating the advantages of GAT and GRU, and fully tap the feature information that could not be used in the past due to missing data. When facing highly missing HER data, it can effectively reduce the impact of missing data on the prediction results, greatly enhancing the robustness of the model in practical applications. Through the fusion of multimodal data and feature association analysis, the preeclampsia prediction model significantly improves the prediction accuracy of preeclampsia. The preeclampsia prediction model constructed by the present invention can still stably predict preeclampsia in the case of severe data loss, and this ability is difficult to replace by other methods in actual clinical practice.
[0053] In this embodiment, the preeclampsia prediction model can identify potential preeclampsia risks in advance by analyzing multi-dimensional data of pregnant women during pregnancy; through time series data of multiple time steps, the GRU network can capture the changing trends of pregnant women's health from historical data and achieve early warning of preeclampsia.
[0054] In addition, the experimental results show that the preeclampsia prediction model based on time series multimodality, namely the GAT-GRU model, has an AUC (area under the ROC curve, i.e., the receiver operating characteristic curve) value of 0.87 for early prediction of preeclampsia, and has strong clinical application value.
[0055] In step S1, the data of each pregnant woman is sampled regularly to ensure the time series of the data, that is, all data records of the pregnant woman are arranged in chronological order from pregnancy to delivery, and sampled at fixed time intervals. If there is no data detected at some time nodes, the time node will be marked as a null value, and the position of the time node will still be retained to form a continuous time series set.
[0056] In some embodiments, the step of obtaining multimodal time series feature data of pregnant women with known pregnancy outcomes and constructing a time series data set specifically includes step S11 and step S12:
[0057] S11. Divide the pregnancy period of pregnant women with known pregnancy outcomes (including pregnant women with preeclampsia and pregnant women without preeclampsia) into multiple time nodes, collect multimodal feature data of pregnant women with known pregnancy outcomes at each time node, merge the multimodal time series data of each time node, set the feature data that does not appear in the current time node to a null value, take the last appearing feature data for repeated feature data, arrange the feature data obtained from multiple time nodes in chronological order from pregnancy to delivery, and obtain preprocessed multimodal time series feature data.
[0058] At each time point, the multimodal feature data of pregnant women are in the form of text, numerical values, categories, etc. For example, electronic health records come from various information systems of hospitals, which contain various types and modalities of data, including various examination data records of patients in hospitals, and are composed of structured data (such as numerical data such as blood pressure and blood sugar), semi-structured data (such as examination records, diagnosis conclusions, etc.) and text data (such as text data in medical records), that is, unstructured data. That is, feature data is mainly divided into structured data, semi-structured data and text data according to modality. Therefore, the multimodal feature data of each time node are merged. In the merging process, structured data is usually numerical data, such as blood pressure, blood sugar, etc., which are already quantitative numerical features and can be used directly. Semi-structured data contains discrete category information (such as examination records, diagnosis conclusions, etc.), which usually has a limited range of values. Specifically, the category data is converted into a sparse binary vector. For example, the diagnosis result "positive / negative" can be converted into [1,0]. Text data (such as medical records) is unstructured data, which is rich in information but needs to be converted into numerical data through language processing. Specifically, label encoding (Label Encoder) can be used to convert text data into numerical data.
[0059] Among them, multimodal time series feature data specifically include medical indicators, physical signs and other data. Medical indicators include outpatient visits, hospital admissions, diagnosis and test results of the International Statistical Classification of Diseases and Related Health Problems (ICD10) code, and physical signs include body temperature, pulse, blood pressure and heart rate. Some of the feature data are as follows: Figure 2 As shown in the figure (where pain represents pain, weight_x represents the weight at each sampling time node, nutrition represents nutrition, high risk score represents high risk score, high risk level represents high risk level, and fall represents whether the pregnant woman has a record of falling).
[0060] Specifically, each gestational week (i.e., gestational age) during the pregnancy of a pregnant woman with a known pregnancy outcome can be used as a time node, i.e., a data collection node, and multimodal feature data of the pregnant woman with a known pregnancy outcome at each gestational week is collected, the multimodal feature data of each gestational week is merged, and the feature data that does not appear in the current gestational week is set to a null value, and the repeated feature data takes the last appearance of the feature data, and the data are arranged in order of gestational age from small to large (for example, arranged in the order of first week, second week, third week, fourth week, ...), to obtain the pre-processed multimodal time series feature data.
[0061] More specifically, the 10th to 25th week of pregnancy for pregnant women with known pregnancy outcomes is divided into 16 time nodes, and the multimodal feature data of each gestational week from the 10th to the 25th week of pregnancy is used. The reason for this selection is that this stage is a critical period when the risk of preeclampsia gradually emerges, which is helpful for early prediction, and the data at this stage is generally more complete and standardized, which is helpful for feature extraction of time series models. In addition, this selection method meets the model input requirements while containing sufficient dynamic information and controlling the complexity of the time series length.
[0062] S12, using the number of times the feature data appears greater than a preset value as a screening criterion, screening the preprocessed multimodal time series feature data, and constructing a time series data set.
[0063] In some embodiments, the preprocessed multimodal time series feature data is filtered using the number of times the feature data appears greater than a preset value as a screening criterion to obtain filtered multimodal time series feature data; the filtered multimodal time series feature data includes constant features and variable features, and the constant features are filled in the missing positions of the constant features in the filtered multimodal time series feature data to obtain the filled multimodal time series feature data, and to construct a time series data set.
[0064] Because many features in the preprocessed multimodal time series feature data may be specific detections and are not suitable for the purpose of early screening of preeclampsia in regular pregnant women, the data is universally screened based on the number of occurrences and vacancies of the feature data. Specifically, the number of occurrences of the feature data is greater than the preset value as the screening criterion. After the preprocessed multimodal time series feature data is screened, those feature data with a number of occurrences greater than the preset value are selected to construct the time series data set. This screening strategy ensures that the data input to the model is more concise and effective, removes redundant and irrelevant features, and thus improves the accuracy and efficiency of the prediction.
[0065] At each time node, the multimodal feature data of each pregnant woman includes constant features and variable features. Constant features are features that do not change over time, and constant features remain unchanged throughout the pregnancy. Variable features are features that change over time, and variable features will change as the pregnancy progresses. As an example, constant features include: blood type, age, family history (such as diabetes, hypertension), medical history (such as personal medical history, allergy history, surgical history, etc.) and fetal sex, etc. Variable features include: blood pressure (such as systolic pressure, diastolic pressure, abnormal blood pressure), pulse, uterine height, fetal heart rate (such as abnormal fetal heart rate), blood parameters (such as white blood cells, red blood cells, platelets, mean platelet volume, hemoglobin concentration, etc.), biochemical indicators (such as triglycerides, urate, calcium, protein, etc.) and inflammatory markers (such as total number of white blood cell classifications, red blood cell distribution width, etc.). In the embodiment of the present invention, particular attention is paid to the changing trend of variable features, because they are of great significance in the prediction of preeclampsia and can reveal whether pregnant women have potential preeclampsia risks. For constant features, in an embodiment of the present invention, constant features are filled in the vacant positions of constant features in the screened multimodal time series feature data (the specific filling method can be: using the mean or mode to fill the constant data; using label encoder for text data to convert the text data into numerical data), and the filled multimodal time series feature data is obtained, and then the time series data set is constructed with the filled multimodal time series feature data. That is, the constant features are reused in each time node (for example, for the constant of blood type, when some time nodes are vacant, the blood type positions in these time nodes are filled with blood type, so the constant feature of blood type will be reused at each time node) to assist the model to better extract the changing trend of variable features. In this way, the model can identify the dynamic change pattern of variable features in the time series, thereby improving the early prediction ability of preeclampsia. This process ensures that the model not only focuses on the characteristics of the current node, but also effectively captures the long-term changing trend of the health status of pregnant women.
[0066] In this step, the constructed time series data set can be expressed as: T = {T 1 ,T 2 ,...,T m}; where m is the total number of time nodes, each of which T i Represents the set of all feature data at the i-th time point, that is, each time node T i (i is 1, 2, ..., m) contains the multidimensional feature data of the pregnant woman at this time point, which can be expressed as: T i ={f i1 ,f i2 ,...f il}, where 1≤l≤total number of features (e.g., 1≤l≤1048, 1≤l≤131, 1≤l≤95 or 1≤l≤31); f i1 , f i2 , f il The specific value or vectorized representation of a single feature at each time node.
[0067] In a specific embodiment, after step S11, 1048 features are obtained. Because of the uneven distribution of each feature data of pregnancy detection, most of the 1048 features are specific detections, which are not suitable for the purpose of early screening of preeclampsia in conventional pregnant women. Therefore, the data is universally screened according to the number of occurrences and vacancies of a certain feature data. Specifically, according to the standard that the number of occurrences of a certain feature in the overall data is more than 2000 times, 1500 times, 1200 times and 1000 times, respectively, 31 features, 79 features, 95 features and 131 features are retained on the 1048 features, and time series data sets are constructed with these four groups of features.
[0068] The 31 features include: systolic blood pressure, diastolic blood pressure, abnormal blood pressure, weight_x, abnormal weight, pulse, edema, fundal height, fetal heart rate, fetal heart rate abnormality, presenting part, abnormal fetal position, engagement, nutrition, fall, pain, functional demand, high risk score, high risk level, protein, urobilinogen, nitrite, occult blood, leukocytes, calcium oxalate, triphosphate, urate, hyaline casts, non-squamous epithelial cells, squamous epithelial cells, and mucus filaments.
[0069] The 79 features include: systolic blood pressure, diastolic blood pressure, abnormal blood pressure, weight_x, abnormal weight, pulse, edema, uterine height, fetal heart rate, abnormal fetal heart rate, presenting part, abnormal fetal position, engagement, nutrition, fall, pain, functional needs, high riskscore, high risk level, protein, urobilinogen, nitrite, occult blood, white blood cells, calcium oxalate, triphosphate, urate, hyaline casts, non-squamous epithelial cells, squamous epithelial cells, mucus filaments, and red blood cells, fine granular casts, coarse granular casts, red blood cell casts, white blood cell casts, vitamin C, platelet count, large platelet ratio, mean platelet volume, platelet distribution width, red blood cell distribution width-CV (CV means coefficient of variation), red blood cell distribution width-SD (SD means standard deviation), absolute value of basophils, absolute value of eosinophils, absolute value of neutrophils, absolute value of monocytes, and lymphocytes. Absolute value, percentage of basophils, percentage of eosinophils, percentage of neutrophils, percentage of monocytes, percentage of lymphocytes, platelets, mean hemoglobin concentration, mean hemoglobin content, mean corpuscular volume, hematocrit, hemoglobin, atypical lymphocytes, total number of white blood cell classifications, absolute value of nucleated red blood cells, percentage of nucleated red blood cells, glucose, pH, indirect bilirubin, total protein, albumin, total bilirubin, direct bilirubin, r-glutamyl transferase, total bile acid, creatinine, uric acid, A / G (serum albumin to globulin ratio), globulin, cystatin C, bilirubin and ketone bodies.
[0070] The 95 characteristics include: systolic blood pressure, diastolic blood pressure, abnormal blood pressure, weight_x, abnormal weight, pulse, edema, uterine height, fetal heart rate, abnormal fetal heart rate, presenting part, abnormal fetal position, engagement, nutrition, fall, pain, functional needs, high riskscore, high risk level, protein, urobilinogen, nitrite, occult blood, leukocytes, calcium oxalate, triphosphate, urate, hyaline casts, non-squamous epithelial cells, squamous epithelial cells, mucus filaments, and erythrocytes, fine granular casts, coarse granular casts, erythrocyte casts, leukocyte casts, vitamin C, platelet volume, large platelet ratio, mean platelet volume, platelet distribution width, erythrocyte distribution width-CV, erythrocyte distribution width-SD, absolute basophil count, absolute eosinophil count, absolute neutrophil count, absolute monocyte count, absolute lymphocyte count, basophil percentage, eosinophil percentage, neutrophil percentage, monocyte percentage, lymphocyte percentage, platelets, mean hemoglobin concentration, Mean hemoglobin content, mean red blood cell volume, hematocrit, hemoglobin, atypical lymphocytes, total number of classified white blood cells, absolute value of nucleated red blood cells, percentage of nucleated red blood cells, glucose, pH, indirect bilirubin, total protein, albumin, total bilirubin, direct bilirubin, r-glutamyl transferase, total bile acid, creatinine, uric acid, A / G, globulin, cystatin C, bilirubin and ketone bodies, as well as specific gravity, alanine aminotransferase, aspartate aminotransferase, spouse's health, height, weight_y (weight of the pregnant woman when the card is first created), bmi (body mass index), hypertension, pregnancy number, gestational age of fetal movement, early pregnancy reaction, early pregnancy reaction week, diabetes, no history of disease, no history of drug allergy, and last menstrual period.
[0071] The 131 characteristics include: systolic blood pressure, diastolic blood pressure, abnormal blood pressure, weight_x, abnormal weight, pulse, edema, uterine height, fetal heart rate, abnormal fetal heart rate, presenting part, abnormal fetal position, engagement, nutrition, fall, pain, functional needs, high risks core, high risk level, protein, urobilinogen, nitrite, occult blood, leukocytes, calcium oxalate, triphosphate, urate, hyaline casts, non-squamous epithelial cells, squamous epithelial cells, mucus filaments, and erythrocytes, fine granular casts, coarse granular casts, erythrocyte casts, leukocyte casts, vitamin C, platelet volume, large platelet ratio, mean platelet volume, platelet distribution width, erythrocyte distribution width-CV, erythrocyte distribution width-SD, absolute basophils, absolute eosinophils, absolute neutrophils, absolute monocytes, absolute lymphocytes, percentage of basophils, percentage of eosinophils, percentage of neutrophils, percentage of monocytes, percentage of lymphocytes, platelets, mean hemoglobin concentration, mean hemoglobin content, mean corpuscular volume, hematocrit, hemoglobin, atypical lymphocytes, total number of leukocyte classification, absolute nucleated erythrocytes, percentage of nucleated erythrocytes, glucose , pH, indirect bilirubin, total protein, albumin, total bilirubin, direct bilirubin, r-glutamyl transferase, total bile acid, creatinine, uric acid, A / G, globulin, cystatin C, bilirubin and ketone bodies, as well as specific gravity, alanine aminotransferase, aspartate aminotransferase, spouse's health, height (height), weight_y, bmi, hypertension, pregnancy number, fetal movement gestational age, early pregnancy reaction, early pregnancy reaction week, diabetes, no medical history, no medication Drug allergy history, last menstrual period, as well as calcium, alkaline phosphatase, potassium, sodium, magnesium, atypical lymphocytes, chloride, glycosylated hemoglobin (HPLC), lactate dehydrogenase, creatine kinase, α-hydroxybutyrate dehydrogenase, high-sensitivity CRP, estradiol, white blood cell morphology, platelet morphology, α-amylase, red blood cell morphology, lipase, cholinesterase, creatine kinase isoenzyme, risk value, free hCGβ, gestational age at sampling, number of fetuses, expected age at delivery, gestational age at B-ultrasound, PAPP-A MoM value (PAPP-A refers to pregnancy-associated plasma protein A, and PAPP-A MoM value is an important value in prenatal screening for Down syndrome), nuchal translucency thickness (direct measurement value), NT MoM value (NT value refers to the thickness of the fetal nuchal translucency detected under B-ultrasound, and NT MoM value is the standardized multiple of the nuchal translucency thickness), free T4, thyrotropin, rapid C-reactive protein (instrument), reticulocyte percentage, immature reticulocyte fraction, neutrophil band nuclei and triglycerides.
[0072] In step S2, Figure 5As shown, in some embodiments, the step of training the graph attention network and the fully connected network using the time series dataset specifically includes steps S21 to S27:
[0073] S21. Convert the time series data set into graph structure data, and respectively mask the characteristic values of a part of the nodes in the graph structure data as missing values to form missing nodes, thereby obtaining graph structure data with missing nodes.
[0074] Specifically, at the input end, graph structure data is constructed with features as nodes (in some specific examples, such as Figure 3 As shown in the figure, the graph structure data is constructed with the feature of a gestational week as a node). The time series data set is converted into graph structure data. In the graph structure data, each node represents a feature, and each feature value is called the attribute of the node. Generally speaking, the expression of a graph is:
[0075] G=(V,E)
[0076] Where V is the total set of node features in the graph, and E is the edge set. Node feature V = {v 1 ,v 2 ,v 3 ,...,v m}; Each node v i ∈V has multiple feature attributes, which are used as node v i The attribute value of v i The i in the range is 1, 2, 3, ..., m. Every two nodes (i.e., v i and v j ) through the edge (v i ,v j ), the edge features together form the edge set E, and the edge weight represents the correlation between nodes. In this way, time series data can be converted into a graph structure, so that the correlation of each feature and the context information of the time series can be effectively represented. This graph structure provides a suitable input form for subsequent prediction models and helps capture the complex relationship between features.
[0077] S22, such as Figure 5 As shown, after the graph structure data with missing nodes is input into GAT, the first data is output.
[0078] In some specific embodiments, Figure 4 and Figure 5 As shown, after setting an information transmission mechanism in the graph attention network so that the missing nodes can only receive information from adjacent nodes but not transmit information to adjacent nodes, the graph structure data with missing nodes is input into the graph attention network and the first data is output.
[0079] More specifically, a propagation threshold is set in the GAT network structure. For example, nodes with vacant node values less than 0 are set not to propagate feature data to adjacent nodes. This prevents these nodes from transmitting features to adjacent nodes and allows adjacent nodes to transmit feature information to vacant nodes as much as possible. The results show that setting a higher threshold makes it easier for vacant nodes to obtain information from adjacent nodes, thereby enhancing the model's ability to fit and fill in missing data and increasing the prediction effect of subsequent time series models.
[0080] In some specific embodiments, the GAT has a three-layer structure.
[0081] After GAT processing, the first data output (corresponding to Figure 5 The output on the left (01) and the original input graph structure data (corresponding to Figure 5 The input on the left side) has the same length after one-dimensional expansion, so the first data can be expressed as Output characteristics (i is 1, 2, ..., N) is the final node representation of the Uth layer. At this point, the relationship between nodes has been optimized. Through the graph attention mechanism, the model can more accurately capture the potential connections between nodes, thereby improving the accuracy of feature extraction and the prediction effect of the model. This step provides high-quality feature input for subsequent preeclampsia prediction.
[0082] In step S21 and step S22, in order to simulate the common missing situations in actual medical data, some nodes are intentionally masked in the graph structure data, and the feature values of some nodes are replaced with vacant values. At the same time, for these missing nodes, a special information transmission machine is designed so that these missing nodes can only receive information from adjacent nodes, but not transmit information to adjacent nodes. Figure 4As shown in (d) in . In this way, the interference of missing nodes on the information transmission of adjacent nodes is reduced, ensuring that the model will not be negatively affected by missing data during the training process. Specifically, by adjusting the propagation threshold of the vacant nodes (for example, when the characteristic value of the vacant node is set to be less than 0, no feature transmission is performed. When training is performed, the data set is randomly divided into a training set and a test set in a ratio of 8:2. The training set data is trained 96 times in a loop, and each time a node value (a whole column) is set to -9, and a restriction condition is added to restrict the nodes less than -9 from transmitting information to the outside), the model can effectively control the noise caused by missing data and maximize the transmission of effective information between adjacent nodes, further improving the prediction effect of the time series model. The present invention artificially sets the characteristic values of some nodes to vacant and designs a specific propagation mechanism so that the vacant nodes only receive feature information from adjacent nodes, and do not transmit feature information to the outside. This design effectively reduces the interference of vacant nodes on information propagation during model training and improves the model's ability to fit and fill in missing data.
[0083] The graph structure data with missing value nodes is input into a three-layer graph neural network based on graph attention mechanism, namely three-layer GAT. The graph structure data input into GAT includes node features and graph structure information. i The characteristic is expressed as where x ij is represented as the jth feature of the i-th node. The graph structure information is represented by the edge set E, which describes the connection relationship between nodes. In GAT, node features are updated through the attention mechanism. i The feature update formula can be expressed as:
[0084] 1≤u≤3;
[0085] in, Represents node v i The feature vector at the u+1th layer, Represents node v j The feature vector at the uth layer, represents the attention coefficient of the u-th layer, that is, node v i Its adjacent node v j The association strength represents the weighted coefficient in the graph attention mechanism, W (u) is the weight matrix of the u-th layer, b (u) is the bias term, σ is the activation function, N(v i ) refers to the node v i The set of directly connected nodes.
[0086] The network can adaptively assign different weights according to the relationship between nodes, thereby effectively extracting the potential correlation between nodes in the graph. When dealing with missing nodes, it can automatically adjust the propagation mode of node information to minimize the information interference caused by missing nodes. This mechanism can improve the learning ability of the model under the condition of missing data.
[0087] S23, such as Figure 5 As shown, the first data (ie, H (U) ,correspond Figure 5 The output 01 on the left side and the graph structure data (denoted as G, corresponding to Figure 5 The input on the left side) is fused to obtain the fused feature data (T', corresponding to Figure 5 Input 02 on the left); then T'=G+H (U) .
[0088] In some embodiments, the fusion method is residual connection. Residual connection is a structure used in deep neural networks, which aims to improve the training effect of deep networks. The core idea is to introduce a "shortcut" or "skip connection" in the network so that the input can be directly passed to the subsequent layer and added to the output after one or more layers of processing. This structure allows the network to learn identity mapping of the input, thereby alleviating the problems of gradient disappearance and gradient explosion in deep network training. Therefore, the use of residual connection as a fusion method in this embodiment helps to retain key information in the original data and combine the features extracted by the graph neural network with the original features.
[0089] S24, performing feature expansion on the fused feature data to obtain one-dimensional feature data (i.e., T flat ), that is, T flat =Flatten(T′); Flatten means one-dimensional expansion processing.
[0090] S25, such as Figure 5 As shown, the one-dimensional feature data is input into a three-layer stacked fully connected network to obtain the filling value of the vacancy value and the preeclampsia prediction result;
[0091] S26, obtaining a first loss value according to the filling value and the masked true value at the node, and obtaining a second loss value according to the preeclampsia prediction result and the preeclampsia true result;
[0092] S27. Adjust the parameters in the graph attention network and the fully connected network according to the first loss value and the second loss value until the weighted sum of the first loss value and the second loss value converges. Then the trained graph attention network and the two-layer stacked fully connected networks in the three-layer stacked fully connected network connected to the graph attention network constitute a missing value fitting model.
[0093] In this step, if Figure 5 As shown, the three-layer stacked fully connected network includes the first layer of fully connected network (corresponding to Figure 5 The fourth layer in the second layer of the fully connected network (corresponding to Figure 5 The fifth layer in the network) and the third layer of the fully connected network (corresponding to Figure 5 The one-dimensional feature data is input into the first-layer fully connected network, the output of the first-layer fully connected network is used as the input of the second-layer fully connected network, and the output of the second-layer fully connected network is used as the input of the third-layer fully connected network. The trained three-layer GAT and the first-layer fully connected network and the second-layer fully connected network sequentially connected to the three-layer GAT constitute a missing value fitting model.
[0094] In a specific embodiment, when masking is performed, the characteristic values of a part of the nodes in the graph structure data are masked (referred to as mask01, which is used to set the mask of the vacant value, and is used to realize the loss calculation of the vacant value, i.e., the missing value filling task) as vacant values to form missing nodes, and obtain graph structure data 01 with missing nodes; the characteristic values of another part of the nodes in the graph structure data are respectively masked (referred to as mask02, which is a mask for the preeclampsia classification task, and is used to realize the loss calculation of the preeclampsia classification task) as vacant values to form missing nodes, and obtain graph structure data 02 with missing nodes.
[0095] After inputting the graph structure data 01 with missing nodes into the graph attention network (setting an information transmission mechanism in which missing nodes can only receive information from adjacent nodes but not transmit information to adjacent nodes), output the first data; fuse the first data with the graph structure data to obtain fused feature data; perform feature expansion on the fused feature data to obtain one-dimensional feature data; input the one-dimensional feature data into a three-layer stacked fully connected network to obtain a filling value for the missing value (the model will automatically fill a filling value at the feature position of the node with the mask set); obtain a first loss value based on the filling value and the masked true value at the node.
[0096] Next, the graph structure data 02 with missing nodes is input into the graph attention network (an information transmission mechanism is set in which missing nodes can only accept information from adjacent nodes but not transmit information to adjacent nodes), and the first data is output; the first data is fused with the graph structure data to obtain fused feature data; the fused feature data is feature expanded to obtain one-dimensional feature data; the one-dimensional feature data is input into a three-layer stacked fully connected network to obtain a preeclampsia prediction result (i.e., the model's predicted output for preeclampsia classification); and a second loss value is obtained based on the preeclampsia prediction result and the true preeclampsia result (i.e., the true label of preeclampsia).
[0097] The parameters in the graph attention network and the fully connected network are adjusted according to the first loss value and the second loss value until the weighted sum of the first loss value and the second loss value converges.
[0098] The weighted sum of the first loss value and the second loss value is the loss value loss GAT , loss GAT =α×loss01+β×loss02, where α=0.8, β=0.2; the first loss value, loss01, is used to measure the model's ability to fill in missing values (that is, its ability to fit missing values), and the second loss value, loss02, is used to measure the model's ability to extract features related to preeclampsia, because the main function of the missing value fitting model is the ability to fill in missing values (or the ability to fit), so the weight of loss01 is relatively large. GAT Training until the loss value converges allows the model to effectively handle missing data while ensuring its accuracy in predicting the risk of preeclampsia.
[0099] In this step, the one-dimensional feature data T flat Input into the three-layer stacked fully connected network, the final output is the binary classification result y (then the mask02 and loss02 in the above text are related to the target output y), then:
[0100] y=σ(W 3 ·σ(W 2 ·σ(W 1 ·T flat +b 1 )+b 2 )+b 3 )
[0101] Among them, W 1 , W 2 , W 3 is the weight matrix of each layer, b 1 , b 2 , b 3is the bias of each layer and σ is the activation function.
[0102] The fully connected network extracts more abstract features through nonlinear mapping of data. The design of the fully connected layer can improve the model's ability to express features and enhance the accuracy of prediction.
[0103] In step S3, Figure 5 As shown, in some embodiments, the missing value fitting model is embedded into a gated recurrent unit as an embedding layer, and after training with the time series data set, the step of obtaining a preeclampsia prediction model specifically includes:
[0104] S31, embedding the missing value fitting model as an embedding layer into the input end of the gated recurrent unit, converting the time series data set into graph structure data as input, and obtaining the probability of correctly predicting whether the pregnant woman suffers from preeclampsia;
[0105] S32, obtaining a cross entropy loss function according to the probability of correctly predicting whether the pregnant woman suffers from preeclampsia and the actual situation of whether the pregnant woman suffers from preeclampsia;
[0106] Specifically, the cross entropy loss function is:
[0107]
[0108] in, is the total loss function, which represents the prediction error on the current model data set; N is the total number of samples; y n Indicates whether the nth sample suffers from preeclampsia, and its value is 0 or 1, 0 represents no preeclampsia, and 1 represents preeclampsia; It represents the predicted probability of the nth sample (i.e., the probability of correctly predicting whether the pregnant woman has preeclampsia), and it represents the model's prediction confidence that the sample belongs to category 1 (i.e., the category of preeclampsia in the target classification task), ranging from 0 to 1.
[0109] S33. According to the cross entropy loss function, the gated recurrent unit embedded with the missing value fitting model is trained (the data set can be divided into 80% training set and 20% test set during training) until the cross entropy loss function converges to obtain the preeclampsia prediction model.
[0110] Through the above construction, the obtained preeclampsia prediction model has the following confusion matrix for preeclampsia prediction: Figure 7As shown, the performance indicators of the preeclampsia prediction model are AUC ≥ 0.87, sensitivity ≥ 76%, specificity ≥ 94%, and F1 score (representing the harmonic mean of precision and recall) is 0.85. AUC ≥ 0.87 indicates that the model has a good overall classification ability for the prediction of preeclampsia; sensitivity ≥ 76% indicates that the model has excellent sensitivity in identifying preeclampsia, and specificity ≥ 94% indicates that the model has high specificity in excluding non-preeclampsia patients. The F1 score of 0.85 indicates that the model performs well in balancing accuracy and sensitivity. Since the data set is an unbalanced data set, although the accuracy rate is high, the present invention mainly uses the AUC value as the main evaluation indicator.
[0111] In this step, the principle of the gated recurrent unit is: if the input data is The output of each layer is used as the input of the next layer to form a recurrent neural network for processing time series data. The update formula of the network is:
[0112]
[0113] Among them, h m is the hidden state at the mth time step, z m It is the update gate. is a candidate hidden state. Each layer of the network contains 256 units, which effectively captures the contextual information in the time series data.
[0114] In the present invention, the trained missing value fitting model is used as an embedding layer to perform feature extraction and missing value fitting on the input data. After the preeclampsia prediction model is trained, it is assumed that the time series data of the pregnant woman to be tested is T seq ={T″ 1 ,T″ 2 ,...,T″ m}, where T″ i represents the feature vector of time step i (i is 1, 2, ..., m), then the data is at the input end (such as Figure 5 The embedding layer uses GAT to extract the features of each time series node with the help of graph structure data, captures the correlation between nodes, and provides more accurate input features for subsequent prediction models, namely:
[0115]
[0116] Among them, F(T″ i ) represents the time step T i The result after embedding the features is recorded as The graph structure features of the nodes are extracted through GAT. It is the mapping function of the embedding layer. i represents the index of the time step. Since the entire time series contains m time steps, i = 1, 2, ..., m, indicating that the i-th time node in the sequence is processed.
[0117] In this embodiment, during the training process, the Adam optimizer can be used to optimize the model, and the initial learning rate is set to 0.001. Through these settings, overfitting can be effectively avoided and the convergence of the model can be accelerated, so that the model can learn effective feature representations in a shorter time and improve the prediction accuracy.
[0118] like Figure 6 As shown in the figure, the multimodal time series feature data of the pregnant woman to be tested is converted into graph structure data at the input end, and then input into the trained preeclampsia prediction model. The input data is subjected to feature extraction and missing value fitting processing through the embedding layer, and then the time series feature steps and predictions are performed through the GRU. Finally, the output of each time step is obtained through the linear layer, and the output of the last time step is selected as the final binary classification result.
[0119] In the above, 31, 79, 95 and 131 features were selected to construct the time series dataset. Figure 8 As shown, when the number of selected features is 79, the model classification prediction effect is the best. That is, when 79 features are selected, the relationship between the features can be obtained and retained, and the missing value fitting model can identify more effective features and is less affected by the missing values. The data selected by the present invention include test results of pregnant women during pregnancy, family medical history, disease history, physical signs, health factors and other related data as features for processing, which are mainly divided into numerical and text data according to the modality. The above features are mainly based on the feature after time series sampling, according to the vacancy of the feature for screening, and the features with low missing rate after sampling are selected as the selected features. Each feature is regarded as a node, and the correlation coefficient between the features is used as an edge to construct the graph structure data. The constructed graph structure data regards each node as the same dimension.
[0120] In addition, the present invention also uses the GCN (Graph Convolutional Network) model to process the vacancies. The overall prediction effect of the model is poor, the comparison results are not obvious, and different filling methods can be barely derived. There is a certain impact on the GCN fitting ability. Whether the feature data is normalized has the opposite effect on the same filling method.
[0121] Table 1. Impact of different processing methods on vacant nodes
[0122]
[0123] An embodiment of the present invention also provides a method for predicting preeclampsia, wherein the multimodal time series feature data of the pregnant woman to be tested is input into a preeclampsia prediction model constructed using the construction method described above in the embodiment of the present invention to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
[0124] The embodiment of the present invention further provides a device for predicting preeclampsia, which includes:
[0125] The prediction unit is used to input the multimodal time series feature data of the pregnant woman to be tested into the preeclampsia prediction model constructed by the construction method described above in the embodiment of the present invention, so as to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
[0126] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the protein structure model training method described above in the embodiment of the present invention or implements the protein structure prediction method described above in the embodiment of the present invention.
[0127] The computer-readable medium described in this embodiment can be a computer-readable storage medium or a computer-readable signal medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. In some embodiments, a computer-readable signal medium can include a data signal propagated in a baseband or as a part of a carrier wave, wherein a computer-readable program code is carried. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0128] The present invention also provides an electronic device, which includes a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the computer program is executed by the processor, the protein structure model training method described above in the embodiment of the present invention or the protein structure prediction method described above in the embodiment of the present invention is implemented.
[0129] In this embodiment, the memory may be a volatile memory, such as a random access memory; the memory may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk, etc. The processor may be a central processing unit, a controller, a microcontroller, a microprocessor or other data processing chip.
[0130] In summary, the present invention provides a method for constructing a preeclampsia prediction model based on time series multimodality, a prediction method and a prediction device, and the preeclampsia prediction model can effectively solve the problems of insufficient multimodal data processing and high missing data in the prior art. Unlike the traditional preeclampsia prediction method, the present invention uses GAT to extract features from data, especially for the case of high data missing rate, and can extract effective features from highly missing data. Combined with the gated recurrent unit GRU to capture and predict time series features, the prediction accuracy of preeclampsia is effectively improved. Traditional preeclampsia prediction methods usually rely on a single clinical indicator and are mostly used in the later stages of disease occurrence, while the present invention can identify potential preeclampsia risks in advance by analyzing multi-dimensional data of pregnant women during pregnancy. Through the input of multi-time step time series data, the GRU network can capture the changing trend of pregnant women's health from historical data and realize early warning of preeclampsia. In summary, the present invention combines GAT feature extraction technology with GRU time series prediction to propose a new strategy for processing high missing data. In EHR data, missing value problems are common, and traditional methods cannot effectively cope with this challenge. The prediction model of the present invention can effectively extract features even when there are many missing values, greatly enhancing the robustness of the prediction model in practical applications. The prediction model provided by the present invention not only improves the accuracy of preeclampsia prediction, but also provides a more efficient data processing and prediction tool for clinical practice.
[0131] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A method for constructing a preeclampsia prediction model based on time series multimodality, characterized in that: The steps include: Obtain multimodal time series feature data of pregnant women with known pregnancy outcomes and construct a time series dataset; pregnant women with known pregnancy outcomes include pregnant women with preeclampsia and pregnant women without preeclampsia; Using the time series dataset to train the graph attention network and the fully connected network to obtain a missing value fitting model; The missing value fitting model is embedded into a gated recurrent unit as an embedding layer, and after training with the time series data set, a preeclampsia prediction model is obtained.
2. The construction method according to claim 1, characterized in that: The steps to obtain multimodal time series feature data of pregnant women with known pregnancy outcomes and construct a time series dataset include: The pregnancy period of pregnant women with known pregnancy outcomes is divided into multiple time nodes, and the multimodal feature data of pregnant women with known pregnancy outcomes at each time node are collected. The multimodal feature data of each time node are merged, and the feature data that does not appear at the current time node is set to a null value, and the feature data that appears repeatedly is taken as the feature data that appears the last time, so as to obtain the preprocessed multimodal time series feature data; The number of times the feature data appears is greater than a preset value as a screening criterion, and the pre-processed multimodal time series feature data is screened to construct a time series data set.
3. The construction method according to claim 2, characterized in that: Taking the number of occurrences of feature data greater than a preset value as a screening criterion, after screening the preprocessed multimodal time series feature data, the step of constructing a time series data set specifically includes: Taking the number of times the feature data appears greater than a preset value as a screening criterion, the preprocessed multimodal time series feature data is screened to obtain screened multimodal time series feature data; the screened multimodal time series feature data includes constant features and variable features; The constant features are filled in the missing positions of the constant features in the filtered multimodal time series feature data to obtain the filled multimodal time series feature data and construct a time series data set.
4. The construction method according to any one of claims 1 to 3, characterized in that: The steps of training the graph attention network and the fully connected network using the time series data set to obtain the missing value fitting model specifically include: Converting the time series data set into graph structure data, and masking the characteristic values of a part of the nodes in the graph structure data as vacant values to form missing nodes, thereby obtaining graph structure data with missing nodes; After inputting the graph structure data with missing nodes into the graph attention network, outputting first data; Fusing the first data with the graph structure data to obtain fused feature data; Perform feature expansion on the fused feature data to obtain one-dimensional feature data; Inputting the one-dimensional feature data into a three-layer stacked fully connected network to obtain a filling value for the vacancy value and a preeclampsia prediction result; A first loss value is obtained according to the filling value and the masked true value at the node, and a second loss value is obtained according to the preeclampsia prediction result and the preeclampsia true result; The parameters in the graph attention network and the fully connected network are adjusted according to the first loss value and the second loss value until the weighted sum of the first loss value and the second loss value converges. Then the trained graph attention network and the two-layer stacked fully connected networks in the three-layer stacked fully connected network connected to the graph attention network constitute a missing value fitting model.
5. The construction method according to claim 4, characterized in that: After setting an information transmission mechanism in the graph attention network so that missing nodes can only receive information from adjacent nodes but not transmit information to adjacent nodes, the graph structure data with missing nodes is input into the graph attention network, and the first data is output.
6. The construction method according to claim 5, characterized in that: The steps of embedding the missing value fitting model as an embedding layer into a gated recurrent unit and training the model using the time series data set to obtain a preeclampsia prediction model specifically include: The missing value fitting model is embedded as an embedding layer into the input end of the gated recurrent unit, and the time series data set is converted into graph structure data as input to obtain the probability of correctly predicting whether the pregnant woman suffers from preeclampsia; According to the probability of correctly predicting whether a pregnant woman has preeclampsia and the actual situation of whether a pregnant woman has preeclampsia, a cross entropy loss function is obtained; The gated recurrent unit embedded with the missing value fitting model is trained according to the cross entropy loss function until the cross entropy loss function converges to obtain the preeclampsia prediction model.
7. A method for predicting preeclampsia, characterized in that: The multimodal time series feature data of the pregnant woman to be tested is input into the preeclampsia prediction model constructed by the construction method described in any one of claims 1 to 6 to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
8. A device for predicting preeclampsia, characterized in that: include: A prediction unit is used to input the multimodal time series feature data of the pregnant woman to be tested into the preeclampsia prediction model constructed by the construction method described in any one of claims 1 to 6, so as to predict whether the pregnant woman to be tested will develop preeclampsia in the future.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the prediction method of claim 7 is implemented.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the prediction method according to claim 7 is implemented.
Citation Information
Patent Citations
Prediction model for early stage and middle stage of pregnancy eclampsia
CN115938575A
Preeclampsia poor pregnancy outcome prediction method based on COX proportional risk model
CN116705314A
Prediction method and system for preeclampsia risk of early pregnancy
CN117672514A