Machine learning-based audit abortion related risk factor analysis method, system and product
By using modular data processing and feature selection algorithms, combined with a multilayer perceptron model, the problems of data missingness and feature redundancy in the analysis of risk factors for missed miscarriages were solved, and a more accurate and interpretable risk assessment was achieved.
Patent Information
- Application Number
- CN202510993791.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-14
Smart Images

Figure CN120954744A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical artificial intelligence technology, specifically relating to a method, system, and product for analyzing risk factors related to missed miscarriage based on machine learning. Background Technology
[0002] The risk factor analysis method and apparatus provided by this invention are only used to extract characteristic relationships and risk factors related to missed abortion from multidimensional clinical data. The generated quantitative results are only used for clinical research or doctors' reference and are not used for direct judgment or diagnosis of diseases, nor do they constitute any medical diagnostic conclusions.
[0003] Missed abortion, also known as delayed miscarriage or fetal death, refers to a special type of spontaneous abortion in which the embryo or fetus dies in the uterus but fails to be expelled naturally and remains in the mother's uterine cavity. Its characteristic feature is that after the embryo or fetus stops developing, it does not exit the body in a timely manner but remains in the uterus for a period of time. Its pathogenesis is highly complex, involving multiple factors such as genetic traits, endocrine imbalances, immune system disorders, environmental factors, and maternal systemic diseases.
[0004] Currently, the analysis of missed abortions primarily relies on clinicians' experience, traditional statistical methods, patient symptom descriptions (such as postmenopausal vaginal bleeding and abdominal pain), and auxiliary methods such as ultrasound examinations. However, these traditional analytical methods have some limitations. Relying on clinicians' experience can lead to subjectivity in the analysis results due to differences in the experience and skill levels of different doctors, affecting the consistency and accuracy of the analysis. Traditional statistical methods may not be able to fully capture complex and multi-factor interaction patterns in clinical data, limiting their effectiveness in risk factor identification and assessment. Furthermore, patient symptom descriptions may be vague or incomplete, especially in cases of early fetal death, where symptoms may be subtle, leading to delays in risk assessment. While ultrasound examination is an important auxiliary tool, differences in operator skill levels and equipment limitations may sometimes cause the optimal time for observing risk factors to be missed.
[0005] In this context, the advantages of using artificial intelligence (AI) for data analysis to assist in risk factor assessment become particularly important. AI technology, especially machine learning algorithms, can process and analyze large-scale medical datasets, identifying complex patterns and correlations that are difficult to capture with traditional methods for understanding and assessing risk factors.
[0006] The closest existing similar technology (e.g., patent publication number: CN111653354A) proposes a risk modeling method based on the levels of specific biomarkers in the serum of pregnant women in early pregnancy to assess the risk of spontaneous abortion, emphasizing the importance of the combined use of these biomarkers in risk assessment. However, this method is not specifically designed for detailed risk factor analysis of missed abortion, as it uses serum prenatal screening biomarkers from four subtypes of spontaneous abortion: threatened abortion, inevitable abortion, incomplete abortion, and missed abortion for modeling. It concludes that the combined application of three biomarkers has a relatively high overall risk value for assessing spontaneous abortion, and its ultimate goal is only to assess the overall risk of spontaneous abortion as a broad category, failing to deeply analyze the specific risk factors of missed abortion. Furthermore, the specificity and generalization ability of this method indicative of risk are limited. In addition, existing technologies have bottlenecks in medical data processing. Medical data generally has missing values, high-dimensional feature redundancy, and potential correlations between features. Missed abortion-related data is no exception. Traditional imputation and dimensionality reduction methods are prone to introducing noise or losing information, resulting in inaccurate risk factor analysis.
[0007] To address the above deficiencies and shortcomings, there is an urgent need for an artificial intelligence method that integrates standardized processing, innovative missing value imputation, and interpretable feature screening. This method would address challenges such as high-dimensional feature redundancy, insufficient nonlinear correlation modeling, and insufficient interpretability of risk factors in data related to missed miscarriages, and provide a new technical approach for the quantitative analysis and assessment of risk factors in missed miscarriages. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the existing technology and to propose a method, system and product for analyzing risk factors related to missed miscarriage based on machine learning.
[0009] In view of this, the present invention proposes a machine learning-based method for analyzing risk factors associated with missed miscarriage, including:
[0010] Step 1: Standardize the obtained physical examination data to be analyzed to obtain integrated normalized data;
[0011] Step 2: For missing values in the normalized data, use a missing value imputation method based on feature statistical distribution and dynamic feature correlation to imputation and obtain complete data;
[0012] Step 3: Use an interpretable feature filtering algorithm to extract feature associations from the complete data, and filter to construct key data containing important risk indicator features;
[0013] Step 4: Input the key data into the pre-established and trained risk factor analysis model to obtain the quantitative feature vector combination of the risk factors for missed abortion; the risk factor analysis model is a multilayer perceptron model.
[0014] Preferably, step 1 includes:
[0015] The process involves de-identifying physical examination data and constructing well-organized data with aligned features.
[0016] Based on the degree of data loss in the neat data, remove noisy data with a missing rate greater than a set value;
[0017] The Min-Max normalization method was used to eliminate dimensional differences, resulting in normalized data X for each data dimension after integration. ′ :
[0018]
[0019] Among them, X min and X max These are the minimum and maximum values of a certain data dimension after noise removal; X represents the data corresponding to that data dimension.
[0020] When X max =X min In this case, Z-score standardization is used instead.
[0021] Preferably, step 2 includes:
[0022] Calculate the feature mean vector μ and covariance matrix Σ, and determine the target missing feature f and its related observed feature subset F based on the missing data and the correlation between features;
[0023] For the missing feature f, the imputation value Z is generated according to the conditions of the observed feature subset F by the following formula. f :
[0024]
[0025] in, Z is the covariance matrix of the eigenset. oF The observed eigenvalue matrix is denoted by α, which is a regularization parameter used to avoid matrix singularity issues.
[0026] Preferably, step 3 includes:
[0027] Generate a decision tree, process the features in the complete data, and calculate the global importance score I of each feature j. j :
[0028]
[0029] Where t represents the node number with feature j as the splitting condition, GiniGain(j,t) represents the Gini index improvement brought about by splitting node t with feature j, ∑ represents the summation of all nodes that meet the condition, and T represents the total number of nodes.
[0030] Sort all features by their global importance scores and select the top Q most important features.
[0031] A weighted Lasso optimization algorithm is used to jointly optimize and screen the top Q important features to construct key data containing important risk indicator features.
[0032] The method also includes a training step for a feature selection algorithm, the optimization objective of which is:
[0033]
[0034] Where N is the number of training samples, y is the training label vector, A is the risk assessment matrix constructed from the training data, w is the feature weight vector to be solved, D is the diagonal weighted matrix in the dimension of feature quantity, used to apply differential sparsity constraints, and α is a hyperparameter controlling the regularization strength.
[0035] Preferably, the risk factor analysis model adopts a multilayer perceptron model, including an input layer, several hidden layers, and an output layer; wherein,
[0036] The input layer receives the processed key feature vector;
[0037] The hidden layer uses a non-linear activation function to perform feature transformation and combination learning, which is used to extract deep-level feature association patterns;
[0038] The output layer extracts the activation result of the previous hidden layer as a risk indicator feature vector related to missed miscarriage.
[0039] Preferably, the method further includes a training step for the risk factor analysis model, which employs a supervised learning strategy, introduces historical label data on whether missed abortions have occurred as the learning target, and sets the loss function to binary cross-entropy loss to optimize the risk factor analysis model's ability to capture risk-related structures.
[0040] Secondly, the present invention provides a machine learning-based system for analyzing risk factors associated with missed miscarriage, comprising:
[0041] The standardization processing module is used to standardize the acquired physical examination data to be analyzed, and obtain the integrated normalized data.
[0042] The missing value imputation module is used to impute missing values in normalized data using a missing value imputation method based on feature statistical distribution and dynamic feature correlation to obtain complete data;
[0043] The feature filtering module is used to extract feature associations from complete data using an interpretable feature filtering algorithm, and to filter and construct key data containing important risk indicator features.
[0044] The risk factor analysis module is used to input key data into a pre-established and trained risk factor analysis model to obtain a combination of quantitative feature vectors of risk factors for missed miscarriages; the risk factor analysis model is a multilayer perceptron model.
[0045] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0046] Fourthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0047] Compared with the prior art, the advantages of the present invention are:
[0048] This invention achieves standardized processing of clinical data through modular design, innovative missing value imputation based on feature association, interpretable risk feature screening, and modeling for risk quantification feature vector generation. It is particularly suitable for processing high-dimensional, high-noise medical data.
[0049] Modular architecture enhances flexibility: Due to the adoption of a modular architecture, each processing step operates independently, supporting system expansion and function upgrades, and improving adaptability to different medical data sources and analysis needs.
[0050] Precise missing value imputation based on correlation, improving data quality: The dynamic missing value imputation method based on covariance matrix and Gaussian conditional distribution proposed in this invention can better preserve the complex relationships between data structure and features, reduce noise introduction, and significantly improve data integrity and the quality of subsequent analysis compared with traditional mean / median imputation or simple regression imputation methods.
[0051] Efficient and interpretable feature selection: The method designed in this invention, which combines weighted Lasso and tree models, can automatically and efficiently select key risk indicator features, reduce data dimensionality and model complexity, and enhance the interpretability of the model selection results through specific weight constraints (such as non-negativity) and optimization strategies (such as dynamic weight adjustment), which helps relevant personnel understand the relative importance of each risk factor.
[0052] Dynamic regularization and optimization control: Bayesian optimization and other methods are used to fine-tune the regularization parameters of the Lasso model to ensure the rationality and optimality of the feature selection process and improve the generalization ability and stability of the risk indication parameters generated by the model.
[0053] Optimized risk indicator parameter generation: MLP is used to perform in-depth analysis on the selected key features, outputting a quantitative, multi-dimensional risk feature vector. Furthermore, the contribution of the risk indicator features generated by the model is interpreted through the SHAP method, thereby clarifying the relative importance of each input feature in the generation of risk parameters. This provides more reliable and comprehensive data support for clinical decision-making and assists in risk stratification and personalized management. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating a machine learning-based method for analyzing risk factors associated with missed miscarriages, as provided in an embodiment of the present invention.
[0055] Figure 2 This is a schematic diagram of the correlation missing value imputation algorithm (covariance matrix and imputation process) of a machine learning-based method for analyzing risk factors related to missed miscarriage provided in an embodiment of the present invention.
[0056] Figure 3 This is a schematic diagram of an interpretable feature screening algorithm for a machine learning-based method for analyzing risk factors related to missed miscarriages, provided in an embodiment of the present invention.
[0057] Figure 4 This is a system modular flowchart of a machine learning-based method for analyzing risk factors related to missed miscarriages, provided in an embodiment of the present invention. Detailed Implementation
[0058] This application aims to utilize artificial intelligence technology to provide intermediate data for clinical risk assessment of missed abortion, assisting clinicians in more comprehensively assessing risk factors, reducing over-reliance on subjective experience and the limitations of single methods such as ultrasound imaging; this application is committed to optimizing data modeling capabilities to more effectively capture the multi-factor nonlinear interaction relationships that lead to missed abortion.
[0059] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0060] Example 1
[0061] Embodiment 1 of this invention proposes a novel machine learning-based method for analyzing risk factors associated with missed miscarriage, and constructs a modular system for analyzing risk factors associated with missed miscarriage, including the following core steps (see...). Figure 1 ):
[0062] ① Obtain the patient's physical examination data;
[0063] ② Standardize the existing data to construct integrated and normalized data;
[0064] ③ Missing value imputation is performed based on the missing value imputation method proposed in this invention, which relates to the statistical distribution of features and the correlation of dynamic features; see [link to invention]. Figure 2 ;
[0065] ④ Based on the completed data after imputation, the interpretable feature filtering algorithm designed in this invention is used to extract feature associations, filter important risk indicator features, and construct a key dataset containing important risk indicator features; see Figure 3 ;
[0066] ⑤ Based on the key dataset, a set of quantitative risk feature vector combinations related to the risk of missed miscarriage is output through machine learning model processing.
[0067] in,
[0068] ① Modular architecture of machine learning-based risk factor analysis method for missed miscarriage: The system breaks down the data processing flow related to missed miscarriage into independent functional modules, such as data acquisition, standardization, missing value imputation, feature screening, and risk indicator parameter generation. It supports module-level upgrades and flexible configuration, significantly improving the system's scalability and adaptability to different medical data.
[0069] ② Imputation of Missing Values Based on Relationships: This invention proposes a method for imputing missing values based on dynamic relationships between features. Based on the covariance matrix and Gaussian conditional distribution, a feature association network is constructed to dynamically generate imputed values. This method is particularly suitable for processing complex missing patterns in medical data. Its core formula is as follows:
[0070]
[0071] in, Let be the covariance matrix of the eigenset, and α be the regularization parameter, which effectively avoids matrix singularity problems.
[0072] ③ Interpretable Feature Selection Algorithm: This invention designs an interpretable feature selection mechanism that combines tree models and weighted Lasso optimization. This mechanism first uses a set of tree models (such as random forests, gradient boosting trees, etc.) to evaluate the global importance score of each feature for identifying the risk of missed miscarriages, performing preliminary screening. Subsequently, a weighted Lasso optimization algorithm is used to refine the selection and weight allocation of the preliminarily screened features. Its objective function is:
[0073]
[0074] Where N is the number of training samples, y is the training label vector, A is the risk assessment matrix constructed from the training data, w is the feature weight vector to be solved, D is the diagonal weighted matrix in the dimension of feature quantity, used to apply differential sparsity constraints, and α is a hyperparameter controlling the regularization strength.
[0075] Key risk indicator features are automatically screened using sparse constraints, ensuring they are non-negative to enhance model interpretability and prevent features with opposite directions from interfering with risk indications. Bayesian optimization is used for tuning to ensure appropriate regularization strength. This two-stage method first utilizes the non-linear feature capture capability of the tree model for preliminary screening, followed by weighted Lasso for more refined feature selection and weight allocation. The non-negative constraint on the weights and the variance-based dynamic adjustment enhance the stability and interpretability of the screening results.
[0076] Example 2
[0077] Embodiment 2 of the present invention provides a machine learning-based system for analyzing risk factors associated with missed miscarriages, implemented based on the method of Embodiment 1, such as... Figure 4 As shown: Detailed description of system modules:
[0078] Module 1: Data Acquisition Module
[0079] Input: Structured or semi-structured data such as patient electronic medical records (EMR) and laboratory biochemical indicators.
[0080] Module 2: Standardization Processing Module
[0081] Input: The loaded original dataset (N is the number of samples, P is the number of features).
[0082] Processing flow:
[0083] ① De-identification processing: Sensitive information such as patient identity is de-identified or removed to ensure that data processing complies with ethical and regulatory requirements.
[0084] ② Normalization: Apply Min-Max standardization to numerical features to eliminate dimensional differences and obtain normalized data X for each data dimension after integration. ′ The formula is:
[0085]
[0086] Among them, X min and X max These are the minimum and maximum values of a certain data dimension after noise removal; X represents the data corresponding to that data dimension.
[0087] If X max =X minIn this case, Z-score standardization is used instead.
[0088] ③ Label desensitization: The analysis reference results are encrypted to ensure data privacy compliance.
[0089] Module 3: Missing Value Imputation Module
[0090] Input: Standardized dataset.
[0091] Processing flow:
[0092] ① Calculate the feature mean vector μ and covariance matrix Σ. Based on the missing data and the correlation between features, determine the target missing feature f and its related observed feature subset F.
[0093] ② For the missing feature f,
[0094] Based on the observed feature subset F, the interpolation value Z is generated by the following formula. f :
[0095]
[0096] in, Z is the covariance matrix of the eigenset. oF The observed eigenvalue matrix is denoted by α, which is a regularization parameter used to avoid matrix singularity issues.
[0097] ③ Merge the imputed data with the complete data to generate the final complete dataset.
[0098] Module 4: Feature Filtering Module
[0099] Input: The preprocessed complete dataset (N is the number of samples, P is the number of features).
[0100] Processing flow:
[0101] ① Tree model set selection: Generate decision trees and calculate the global importance score of each feature j.
[0102]
[0103] Where t represents the node number with feature j as the splitting condition, GiniGain(j,t) represents the Gini index improvement brought about by splitting node t with feature j, ∑ represents the summation of all nodes that meet the condition, and T represents the total number of nodes.
[0104] The global importance scores of all features are ranked, and the top Q important features with the highest importance scores are selected.
[0105] ② Weighted Lasso Optimization and Selection: The initially selected Top-Q features are input into a weighted Lasso model. By minimizing the objective function, a subset of key features that are most indicative of the risk of missed miscarriage is further selected, constructing a low-dimensional key risk feature dataset. Bayesian optimization is used to fine-tune α to ensure that the regularization constraints are reasonable.
[0106] Module 5: Risk Assessment Modeling Module
[0107] Input: Low-dimensional key risk feature dataset X s .
[0108] Processing flow:
[0109] ① Model Training: The MLP (Multilayer Perceptron) algorithm is used to perform deep modeling on the aforementioned low-dimensional key risk feature dataset. The model consists of a set of fully connected neural network layers, including an input layer, several hidden layers, and an output layer. The input layer receives the processed key feature vectors, and the hidden layers perform feature transformation and combination learning through non-linear activation functions (such as ReLU) to extract deep-level feature association patterns. During the model training phase, a supervised learning strategy is adopted, introducing historical label data (whether missed abortions occurred) as the learning objective. The loss function is set to binary cross-entropy loss to optimize the model's ability to capture risk-related structures, as shown in the following formula:
[0110]
[0111] Among them, y i This is a true label (a code indicating whether a missed miscarriage occurred). The risk-related score output by the model measures the degree of non-linear correlation between input features and the risk of missed miscarriage. After training, the model no longer directly outputs binary labels during the inference phase. Instead, it extracts the activation results of the hidden layer preceding the output layer as a risk indicator feature vector related to missed miscarriage. This vector has fixed dimensions and a clear structure, capable of expressing the interaction patterns and weight structures of different features in the overall risk assessment. This vector is not used as a basis for clinical diagnosis; it is only used to represent the compressed projection results of high-dimensional features in the training context, for doctors or researchers to analyze and refer to, in order to assist in non-diagnostic work such as risk stratification, feature attribution, and population difference analysis.
[0112] ② Model / Evaluation: The model is divided into a training set (80%) and a test set (20%). The stability and effectiveness of the risk feature vectors generated by the model are evaluated using the SHAP (SHapley Additive ex Planations) method, rather than directly evaluating the accuracy, recall, F1 score, and other indicators used for analysis reference in case risk assessment.
[0113] It is worth noting that in the embodiments of the above system, the modules included are divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional module are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0114] Example 3
[0115] Embodiments of the present invention may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, it can implement the various steps in the above method embodiments.
[0116] Example 4
[0117] Embodiments of the present invention may also provide a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the various steps in the above method embodiments can be implemented.
[0118] Innovation points:
[0119] Modular architecture: The process flow of retained data is broken down into independent functional modules, which improves the system's flexibility and scalability.
[0120] Correlational missing value imputation: A dynamic missing value imputation method based on covariance matrix and Gaussian conditional distribution is proposed. It utilizes the intrinsic correlation between features for imputation. Compared with traditional methods, it can better preserve the true structure of data, significantly improve data integrity and the reliability of subsequent analysis, and is especially suitable for complex medical data.
[0121] Interpretable Feature Screening: A two-stage feature screening mechanism combining weighted Lasso and tree models is designed and implemented. This mechanism not only effectively screens out key features that have a significant indicative role in the risk of missed miscarriages, but also improves the interpretability and robustness of the screening results through weight design (such as non-negativity constraints and dynamic adjustment based on feature variance) and optimization algorithms (such as Bayesian optimization parameter tuning), which helps to reveal potential risk factors.
[0122] Machine learning model application and risk indicator parameter generation: MLP is used to process the selected key features, generating quantitative, multi-dimensional risk feature vector combinations to provide more refined, data-driven auxiliary information for feature-based risk assessment. The model's analytical capabilities are comprehensively evaluated using the SHAP method.
[0123] The risk factor analysis methods, systems, and products provided by this invention are only used to extract characteristic relationships and risk factors related to missed abortion from multidimensional clinical data. The generated quantitative results are only used for clinical research or doctors' reference and are not used for direct judgment or diagnosis of diseases, nor do they constitute any medical diagnostic conclusions.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A machine learning-based method for analyzing risk factors associated with missed miscarriage, comprising: Step 1: Standardize the obtained physical examination data to be analyzed to obtain integrated normalized data; Step 2: For missing values in the normalized data, use a missing value imputation method based on feature statistical distribution and dynamic feature correlation to imputation and obtain complete data; Step 3: Use an interpretable feature filtering algorithm to extract feature associations from the complete data, and filter to construct key data containing important risk indicator features; Step 4: Input the key data into the pre-established and trained risk factor analysis model to obtain the quantitative feature vector combination of missed miscarriage risk factors; The risk factor analysis model is a multilayer perceptron model.
2. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, Step 1 includes: The process involves de-identifying physical examination data and constructing well-organized data with aligned features. Based on the degree of data loss in the neat data, remove noisy data with a missing rate greater than a set value; The Min-Max normalization method was used to eliminate dimensional differences, resulting in normalized data X for each data dimension after integration. ′ : Among them, X min and X max These are the minimum and maximum values of a certain data dimension after noise removal; X represents the data corresponding to that data dimension. When X max =X min In this case, Z-score standardization is used instead.
3. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, Step 2 includes: Calculate the feature mean vector μ and covariance matrix Σ, and determine the target missing feature f and its related observed feature subset F based on the missing data and the correlation between features; For the missing feature f, the imputation value Z is generated according to the conditions of the observed feature subset F by the following formula. f : in, Z is the covariance matrix of the eigenset. oF The observed eigenvalue matrix is denoted by α, which is a regularization parameter used to avoid matrix singularity issues.
4. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, Step 3 includes: Generate a decision tree, process the features in the complete data, and calculate the global importance score I of each feature j. j : Where t represents the node number with feature j as the splitting condition, GiniGain(j,t) represents the Gini index improvement brought about by splitting node t with feature j, ∑ represents the summation of all nodes that meet the condition, and T represents the total number of nodes. Sort all features by their global importance scores and select the top Q most important features. A weighted Lasso optimization algorithm is used to jointly optimize and screen the top Q important features to construct key data containing important risk indicator features.
5. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, The method also includes a training step for a feature selection algorithm, the optimization objective of which is: Where N is the number of training samples, y is the training label vector, A is the risk assessment matrix constructed from the training data, w is the feature weight vector to be solved, D is the diagonal weighted matrix in the dimension of feature quantity, used to apply differential sparsity constraints, and α is a hyperparameter controlling the regularization strength.
6. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, The risk factor analysis model employs a multilayer perceptron model, comprising an input layer, several hidden layers, and an output layer; wherein, The input layer receives the processed key feature vector; The hidden layer uses a non-linear activation function to perform feature transformation and combination learning, which is used to extract deep-level feature association patterns; The output layer extracts the activation result of the previous hidden layer as a risk indicator feature vector related to missed miscarriage.
7. The method for analyzing risk factors related to missed miscarriage based on machine learning according to claim 1, characterized in that, The method also includes a training step for the risk factor analysis model, which adopts a supervised learning strategy, introduces historical label data on whether missed abortions have occurred as the learning target, and sets the loss function to binary cross-entropy loss to optimize the risk factor analysis model's ability to capture risk-related structures.
8. A machine learning-based system for analyzing risk factors associated with missed miscarriage, characterized in that, include: The standardization processing module is used to standardize the acquired physical examination data to be analyzed, and obtain the integrated normalized data. The missing value imputation module is used to impute missing values in normalized data using a missing value imputation method based on feature statistical distribution and dynamic feature correlation to obtain complete data; The feature filtering module is used to extract feature associations from complete data using an interpretable feature filtering algorithm, and to filter and construct key data containing important risk indicator features. and The risk factor analysis module is used to input key data into a pre-established and trained risk factor analysis model to obtain a combination of quantitative feature vectors of risk factors for missed miscarriages; the risk factor analysis model is a multilayer perceptron model.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method described in any one of claims 1-7.
Citation Information
Patent Citations
Method for establishing risk model for predicting spontaneous abortion by prenatal screening markers of maternal serum in early pregnancy period
CN111653354A