Postmenopausal osteoporosis risk assessment system based on big data
Through big data technology, multi-dimensional data of postmenopausal women are collected and processed, and a risk assessment system is built, which solves the shortcomings of existing assessment methods, and accurately assesses and personalized interventions for postmenopausal osteoporosis risks, reduces fracture risks, and improves quality of life.
Patent Information
- Application Number
- CN202510855210.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing postmenopausal osteoporosis risk assessment methods rely on single-dimensional indicators or clinical symptoms judgment, resulting in insufficient comprehensiveness and accuracy of the evaluation results, making it difficult to achieve early prevention and precise intervention.
A postmenopausal osteoporosis risk assessment system based on big data was designed. Through data preprocessing, feature extraction and risk assessment modules, multi-dimensional data is collected for cleaning and normalization, features related to osteoporosis risk are extracted, risk assessment models are constructed, and personalized intervention suggestions are provided.
Accurate prediction of postmenopausal osteoporosis risks is achieved, comprehensive and accurate assessment results are provided, and early prevention and intervention is supported, fracture risk is reduced, and quality of life is improved.
Smart Images

Figure CN120356677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a postmenopausal osteoporosis risk assessment system based on big data. Background Art
[0002] Postmenopausal osteoporosis is a common skeletal disease among middle-aged and elderly women. Its incidence rate is continuously rising with the aggravation of population aging. Due to the decline of ovarian function and the sharp drop of estrogen level in postmenopausal women, bone metabolism imbalance occurs, with bone resorption being greater than bone formation, thus triggering osteoporosis. Osteoporosis can reduce the bone mass of patients, damage the bone microstructure, increase the risk of fractures, seriously affect the quality of life of patients, and bring a heavy economic burden to families and society. At present, the methods for assessing the risk of postmenopausal osteoporosis mainly rely on single-dimensional indicators or clinical symptom judgments, resulting in insufficient comprehensiveness and accuracy of the assessment results and making it difficult to achieve early prevention and precise intervention. Summary of the Invention
[0003] The purpose of the present invention is to solve the above problems and design a postmenopausal osteoporosis risk assessment system based on big data.
[0004] The present invention provides a postmenopausal osteoporosis risk assessment system based on big data, which includes: A data preprocessing module, which is used to collect multi-dimensional data of postmenopausal women, perform data cleaning on the collected data, and perform normalization processing on numerical data and one-hot encoding processing on categorical data to obtain preprocessed data; A feature extraction module, which is used to extract features related to the risk of postmenopausal osteoporosis from the preprocessed data to obtain target feature data; A risk assessment module, which is used to build a risk assessment model based on the target feature data, and input the original data of the postmenopausal women to be evaluated after data preprocessing and feature extraction into the risk assessment model, and output the risk assessment result of the women to be evaluated suffering from osteoporosis; A result visualization module, which is used to visualize the risk assessment result and provide corresponding intervention suggestions according to different risk levels.
[0005] Optionally, in the first implementation manner of the present invention, the data preprocessing module includes: A cleaning sub-module, which is used to collect multi-dimensional data of postmenopausal women, and apply a causal perception missing value filling network to clean the multi-dimensional data and process the missing values, where the multi-dimensional data includes structured data and unstructured data; An integration sub-module, which is used to apply adversarial adaptive normalization processing to numerical data and dynamic graph embedding encoding to categorical data to generate context-aware embedding vectors, and integrate them to obtain preprocessed data.
[0006] Optionally, in the second implementation manner of the present invention, the cleaning sub-module includes: Collect multi-dimensional data containing missing values, construct a data set matrix, where the positions of the missing values are marked by a mask matrix; Run a causal discovery algorithm to analyze the causal relationships between data variables, generate a causal graph, where nodes represent variables and edges represent causal directions; Input the data set matrix, the mask matrix, and the causal graph into the generator of the causality-aware missing value imputation network. The generator constructs a conditional probability distribution based on the causal graph, applies a random mask to the multi-dimensional data to create partially observed data, generates imputation candidate values through the GAIN network, and then uses an attention mechanism to integrate causal information and adjust the imputation value distribution, and outputs the imputed data set.
[0007] Optionally, in the third implementation manner of the present invention, the integration sub-module includes: Divide the numerical data into different domains according to the source devices, input the numerical data of each domain into a feature extractor, and the feature extractor processes the input data, learns to generate domain-invariant feature representations, and outputs normalized numerical data; Construct a co-occurrence matrix of categorical variables, count the co-occurrence frequencies between different categorical values, and construct a dynamic graph based on the co-occurrence frequencies, where nodes represent different values of categorical variables and the weights of edges represent the co-occurrence intensities between nodes; Input the dynamic graph into a graph attention network. The initial feature of each node in the graph attention network is a one-hot encoded vector. Calculate the attention weights between nodes through graph attention layers, aggregate the information of neighbor nodes within the context window, capture the relationships between nodes from different perspectives through the multi-head attention mechanism, update the node embeddings, and output the context-aware embedding vectors of categorical variables.
[0008] Optionally, in the fourth implementation manner of the present invention, the feature extraction module includes: A construction sub-module for arranging the preprocessed data in a time series to obtain nodes and constructing a spatio-temporal graph structure; An introduction sub-module for introducing an LSTM unit for each node in the time dimension. The LSTM unit learns the long-term dependencies of the data in the time series through memory cells and gating mechanisms, and outputs a hidden state containing time features; An application sub-module for applying graph convolution operations in the spatial dimension to aggregate the features of each node and its neighbor nodes, aggregate the information of neighbor nodes by weighted summation, update the features of the current node, and obtain the node representations integrating spatio-temporal features after multiple layers of graph convolution operations. A grouping sub-module for grouping the node representation data that fuses spatio-temporal features according to whether osteoporosis is present, and constructing a contingency table for each feature, where the rows of the contingency table represent different value categories of the feature and the columns represent the categories of whether the disease is present; A screening sub-module for screening out features that have an impact on osteoporosis risk assessment through the chi-square test method based on the actual observed frequencies in the contingency table to obtain target feature data.
[0009] Optionally, in the fifth implementation manner of the present invention, each dimension data under each time point in the spatio-temporal graph structure corresponds to a node, and the attributes of the node include various data indicators at that time point.
[0010] Optionally, in the sixth implementation manner of the present invention, the screening sub-module includes: For each feature, the corresponding chi-square statistic value is obtained through the chi-square test method. The chi-square statistic value of each feature is compared with a preset threshold. If it is greater than the preset threshold, it is determined that the current feature has no significant association with osteoporosis risk and is excluded. If it is less than the preset threshold, it is determined that there is a significant association between the current feature and osteoporosis risk and is retained.
[0011] Optionally, in the seventh implementation manner of the present invention, the risk assessment module includes: An input sub-module for preprocessing and feature extraction of the original data of postmenopausal women to be evaluated and then inputting it into the risk assessment model; A generation sub-module for generating a preliminary risk score through the MedKG-Transformer layer of the risk assessment model; A prediction sub-module for inputting spatio-temporal features and the preliminary risk score into the neural controlled differential equation layer to predict the dynamic change of risk over time; An integration sub-module for integrating the outputs of the MedKG-Transformer layer and the neural controlled differential equation layer to generate a final risk assessment result.
[0012] Optionally, in the eighth implementation manner of the present invention, the construction process of the risk assessment model includes: Construct a medical knowledge graph based on entities and relationships related to bone metabolism obtained from the target feature data, perform embedded representation on the entities and relationships in the medical knowledge graph, and map them to a low-dimensional vector space; Use the embedding of the medical knowledge graph as the initial value of the Key-Value matrix, use the target feature data as the Query vector, calculate the similarity score between Query and Key in the attention mechanism of the MedKG-Transformer layer, and normalize the similarity score using the softmax function; Aggregate Values using the normalized attention weights to generate context-aware feature representations. Perform a non-linear transformation on the context features through the feed-forward network of the MedKG-Transformer layer to generate the final feature representations. Map the features to the risk prediction space through the output layer and output the preliminary risk scores; Construct a neural control differential equation layer, and output the risk prediction values at different time points through the neural control differential equation layer to finally obtain the risk assessment model.
[0013] Optionally, in the ninth implementation manner of the present invention, the process of the neural control differential equation layer includes: Input the time-series features of the patient into the neural control differential equation layer, and initialize the state variables as the currently known health indicators; Solve the differential equation through a numerical integration method to predict the state variables at future time points. At each time step, input the current state variables and control variables into the neural network to calculate the state change rate. According to the predicted future state variables, calculate the dynamic change curve of the risk over time and output the risk prediction values at different time points.
[0014] In the technical solution provided by the present invention, multi-dimensional data of postmenopausal women is collected, the collected data is subjected to data cleaning processing, and the numerical data is normalized and the categorical data is one-hot encoded to obtain the preprocessed data; extract the features related to the risk of postmenopausal osteoporosis from the preprocessed data to obtain the target feature data; construct a risk assessment model based on the target feature data, and input the original data of the postmenopausal women to be evaluated after data preprocessing and feature extraction into the risk assessment model to output the risk assessment result of the women to be evaluated for osteoporosis; visualize the risk assessment result and provide corresponding intervention suggestions according to different risk levels; by collecting multi-dimensional data of postmenopausal women, covering basic information, medical history information, examination and test information, lifestyle information, family history information, etc., all factors affecting osteoporosis risk are comprehensively considered. Compared with traditional assessment methods, the assessment result is more accurate and comprehensive. Using big data technology for preprocessing and feature engineering of massive data can effectively extract key features, remove redundant information, improve data quality and model training efficiency. At the same time, constructing an assessment model can automatically learn the complex relationships and patterns in the data to achieve accurate prediction of osteoporosis risk and meet the needs of personalized assessment; providing targeted intervention suggestions according to the assessment results helps to achieve early prevention and intervention of postmenopausal osteoporosis, reduce the risk of complications such as fractures, and improve the quality of life of patients, with important clinical application value and social and economic benefits. Description of the Drawings
[0015] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the following detailed description of the preferred embodiments. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.
[0016] Figure 1 Schematic structural diagram of a postmenopausal osteoporosis risk assessment system based on big data provided by an embodiment of the present invention; Figure 2 Schematic structural diagram of a feature extraction module provided by an embodiment of the present invention; Figure 3 Schematic structural diagram of a risk assessment module provided by an embodiment of the present invention. Detailed implementation manners
[0017] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of the present invention are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that illustrated or described herein. In addition, the terms "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0018] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to Figure 1 Schematic structural diagram of a postmenopausal osteoporosis risk assessment system based on big data provided by an embodiment of the present invention. The system includes: A data preprocessing module, configured to collect multi-dimensional data of postmenopausal women, perform data cleaning processing on the collected data, perform normalization processing on numerical data, and perform one-hot encoding processing on categorical data to obtain preprocessed data; A feature extraction module, configured to extract features related to postmenopausal osteoporosis risk from the preprocessed data to obtain target feature data; A risk assessment module, configured to construct a risk assessment model based on the target feature data, and input the original data of the postmenopausal women to be evaluated after data preprocessing and feature extraction into the risk assessment model, and output a risk assessment result of the women to be evaluated for osteoporosis; A result visualization module, configured to visualize the risk assessment result and provide corresponding intervention suggestions according to different risk levels.
[0019] In this embodiment, the data preprocessing module includes: A cleaning sub-module, which is used to collect multi-dimensional data of postmenopausal women, apply a causal-aware missing value imputation network to clean the multi-dimensional data and process the missing values, where the multi-dimensional data includes structured data and unstructured data; An integration sub-module, which is used to apply adversarial adaptive normalization processing to numerical data and apply dynamic graph embedding coding to categorical data to generate context-aware embedding vectors, and integrate to obtain preprocessed data.
[0020] In this embodiment, the cleaning sub-module includes: collecting multi-dimensional data containing missing values, constructing a dataset matrix, where the positions of the missing values are marked by a mask matrix; running a causal discovery algorithm to analyze the causal relationships between data variables and generate a causal graph, where the nodes represent variables and the edges represent the causal directions; inputting the dataset matrix, the mask matrix and the causal graph into the generator of the causal-aware missing value imputation network, the generator constructs a conditional probability distribution based on the causal graph, applies a random mask to the multi-dimensional data to create partially observed data, generates imputation candidate values through the GAIN network, and then uses an attention mechanism to integrate causal information and adjust the imputation value distribution, and outputs the imputed dataset.
[0021] In this embodiment, the integration sub-module includes: dividing the numerical data into different domains according to the source devices, inputting the numerical data of each domain into a feature extractor, the feature extractor processes the input data, learns to generate a domain-invariant feature representation, and outputs the normalized numerical data; constructing a co-occurrence matrix of categorical variables, counting the co-occurrence frequencies between different categorical values, and constructing a dynamic graph based on the co-occurrence frequencies, where the nodes represent different values of the categorical variables and the weights of the edges represent the co-occurrence intensities between the nodes; inputting the dynamic graph into a graph attention network, the initial feature of each node in the graph attention network is a one-hot encoded vector, calculates the attention weights between the nodes through a graph attention layer, aggregates the neighbor node information within the context window, captures the relationships between the nodes from different perspectives through a multi-head attention mechanism, updates the node embeddings, and outputs the context-aware embedding vectors of the categorical variables.
[0022] In this embodiment, please refer to Figure 2 , the feature extraction module includes: A construction sub-module, which is used to arrange the preprocessed data in a time series to obtain nodes and construct a spatio-temporal graph structure; An introduction sub-module, which is used to introduce an LSTM unit for each node in the time dimension. The LSTM unit learns the long-term dependencies of the data in the time series through memory cells and gating mechanisms, and outputs a hidden state containing time features; The application sub-module is used to aggregate the features of each node and its neighbor nodes through graph convolution operations in the spatial dimension, aggregate the information of neighbor nodes by weighted summation, update the features of the current node, and obtain the node representation that fuses spatio-temporal features after multiple graph convolution operations. The grouping sub-module is used to group the node representation data that fuses spatio-temporal features according to whether the subject has osteoporosis, and construct a contingency table for each feature, where the rows of the contingency table represent different value categories of the feature and the columns represent the categories of whether the subject is diseased. The screening sub-module is used to screen out the features that have an impact on osteoporosis risk assessment through the chi-square test method based on the actual observed frequencies in the contingency table, and obtain the target feature data.
[0023] In this embodiment, each dimension data at each time point in the spatio-temporal graph structure corresponds to a node, and the attributes of the node include various data indicators at that time point.
[0024] In this embodiment, the screening sub-module includes: obtaining the corresponding chi-square statistic value for each feature through the chi-square test method, comparing the chi-square statistic value of each feature with a preset threshold, if it is greater than the preset threshold, it is determined that the current feature has no significant association with osteoporosis risk and is excluded, if it is less than the preset threshold, it is determined that there is a significant association between the current feature and osteoporosis risk and is retained.
[0025] In this embodiment, please refer to Figure 3 , the risk assessment module includes: The input sub-module is used to preprocess and extract features from the original data of postmenopausal women to be evaluated and input them into the risk assessment model. The generation sub-module is used to generate a preliminary risk score through the MedKG-Transformer layer of the risk assessment model. The prediction sub-module is used to input the spatio-temporal features and the preliminary risk score into the neural control differential equation layer to predict the dynamic change of risk over time. The integration sub-module is used to integrate the outputs of the MedKG-Transformer layer and the neural control differential equation layer to generate the final risk assessment result.
[0026] In this embodiment, the construction process of the risk assessment model includes: constructing a medical knowledge graph based on entities and relationships related to bone metabolism obtained from target feature data, performing embedding representation on the entities and relationships in the medical knowledge graph, and mapping them to a low-dimensional vector space; using the embedding of the medical knowledge graph as the initial value of the Key-Value matrix, taking the target feature data as the Query vector, calculating the similarity score between the Query and the Key in the attention mechanism of the MedKG-Transformer layer, and normalizing the similarity score using the softmax function; aggregating the Value using the normalized attention weights to generate a context-aware feature representation, performing a non-linear transformation on the context features through the feed-forward network of the MedKG-Transformer layer to generate the final feature representation, mapping the features to the risk prediction space through the output layer, and outputting a preliminary risk score; constructing a neural controlled differential equation layer, and outputting risk prediction values at different time points through the neural controlled differential equation layer to finally obtain the risk assessment model.
[0027] In this embodiment, the process of the neural controlled differential equation layer includes: inputting the time series features of the patient into the neural controlled differential equation layer, and initializing the state variable as the currently known health indicators; solving the differential equation through a numerical integration method to predict the state variable at a future time point. At each time step, input the current state variable and the control variable into the neural network, calculate the rate of change of the state, and calculate the dynamic change curve of the risk over time according to the predicted future state variable, and output the risk prediction values at different time points.
[0028] In this embodiment, multi-dimensional data of postmenopausal women is collected, specifically including: Basic information: age, height, weight, menopause age, marital status, education level, etc.; Medical history information: previous fracture history, chronic disease history (such as diabetes, thyroid disease, rheumatic disease, etc.), medication history (such as glucocorticoid use situation, etc.); Examination and test information: bone density test data, blood biochemical indicators (such as blood calcium, blood phosphorus, alkaline phosphatase, estrogen level, etc.), imaging examination results (such as bone-related information shown by X-ray, CT, etc.); Lifestyle information: smoking history, drinking history, exercise amount, eating habits (such as calcium intake, vitamin D intake, etc.), sunshine duration, etc.; Family medical history information: the prevalence of osteoporosis or fractures among family members.
[0029] In this embodiment, multi-dimensional data containing missing values is collected to construct a dataset matrix, where the positions of the missing values are marked by a mask matrix. The causal discovery algorithm (such as NOTEARS) is run to analyze the causal relationships between data variables, and a causal graph is generated, where nodes represent variables and edges represent causal directions. The generator and discriminator of the GAIN framework are initialized: the generator receives the original data, the mask matrix, and the causal graph as inputs, and the discriminator distinguishes between real observed values and generated imputed values. The generator generates imputed values through the following steps: applying a random mask to the original data to create partially observed data, constructing a conditional probability distribution P(X∣Pa(X)) (Pa(X) represents the parent nodes of X in the causal graph) based on the causal graph, inputting the partially observed data and causal constraints into a neural network to generate imputed candidate values, integrating causal information through an attention mechanism to adjust the distribution of imputed values, and the discriminator attempts to distinguish between real observed values and generated imputed values. The parameters of the generator and discriminator are optimized through adversarial training to minimize the loss function, ensuring that the generated imputed values conform to the constraints of the causal graph. The trained generator is used to fill in the missing values in the full-scale data, and a complete dataset is output. Medical institution nodes participating in federated learning are organized, and each node holds local multi-modal datasets (such as structured bone density data, unstructured medical record texts). Feature vectors are extracted from the structured data, and pre-trained models such as BERT are used to extract text embeddings Vt from the unstructured text data. A federated learning framework is constructed: the central server maintains global model parameters, and each node trains the model based on local data without sharing the original data. The intra-modal contrast loss is defined: ensuring that different modal representations of the same patient are close in the embedding space. The inter-modal contrast loss is defined: maximizing the distance between different patient representations through negative sample pairs. The gradients updated by each node are federally aggregated to optimize the global model. The central server broadcasts the current global model parameters, each node calculates the gradients based on local data, applies the contrast learning loss function, and each node encrypts and uploads the gradients to the central server. The central server aggregates the gradients and updates the global model parameters. Using the trained global model, the multi-modal data of each node is mapped to a unified embedding space to achieve data integration. The numerical data is divided into different domains according to the source device or collection center, and a network architecture including a feature extractor, a domain classifier, and a task predictor is constructed: the feature extractor maps the input data to a feature space, and the domain classifier attempts to identify which domain the features come from; the task predictor performs the target prediction task (such as disease diagnosis). An adversarial training process is designed: the feature extractor learns to generate domain-invariant feature representations, and the domain classifier attempts to distinguish features from different domains, and adversarial training is achieved through a Gradient Reversal Layer. Forward propagation: calculating the feature representation, domain prediction, and task prediction; Backward propagation: simultaneously minimizing the task loss and the domain classification loss. The trained feature extractor is used to normalize all numerical data to eliminate the distribution shift caused by device differences.Construct a co-occurrence matrix for categorical variables, count the co-occurrence frequencies between different categorical values, and construct a dynamic graph based on the co-occurrence matrix: nodes represent different values of categorical variables, and the weight of an edge represents the co-occurrence strength between nodes; Initialize the Graph Attention Network (GAT): the initial feature of each node is a one-hot encoded vector, and the graph attention layer calculates the attention weights between nodes and aggregates neighbor node information; Design a context-aware embedding process: define the context window of a node, consider the influence of adjacent categorical variables, and within each context window, update the node embedding through the graph attention mechanism, applying the multi-head attention mechanism to capture the relationships between nodes from different perspectives; Iteratively update the node embedding: calculate the attention weights of each node towards its neighbors, aggregate neighbor node information based on the attention weights, and update the node embedding through a non-linear transformation; Use the updated node embedding as the final representation of the categorical variable.
[0030] In this embodiment, basic feature extraction: multi-scale spatio-temporal graph convolutional network (MST-GCN): Organize the preprocessed data, arrange the multi-dimensional data of the patient (age, BMI, hormone level, etc.) in time series, construct nodes, and each dimension data at each time point corresponds to a node. The attributes of the node include various data indicators at this time point; construct a spatio-temporal graph structure. In the spatial dimension, according to the biological or medical association relationship between data indicators, determine the edge connection relationship between nodes. For example, if there is an association between hormone level and bone density, then a connection is established between the nodes representing the two. In the time dimension, connect the nodes of the same patient at different time points in sequence to form a link in the time series, thus constructing a complete spatio-temporal graph; in the time dimension, introduce an LSTM unit for each node. Take the attribute data of each time point node as the input of the LSTM unit. The LSTM unit learns the long-term dependence relationship of the data in the time series through memory cells and gating mechanisms, and outputs a hidden state containing time features; in the spatial dimension, perform graph convolution operations. According to the constructed graph structure, aggregate the features of each node and its neighbor nodes. The graph convolution kernel slides on the graph and aggregates the information of neighbor nodes by means of weighted summation, etc., to update the features of the current node, so as to capture the interaction effect between biomarkers. After multiple layers of graph convolution operations, obtain the node representations that fuse spatio-temporal features, and these node representations are the extracted basic features; causal feature selection: reinforcement learning-driven causal discovery (RL-Causal): Take the preprocessed data as the input, clarify all the feature sets in the data, and regard each feature as a candidate option; initialize the deep reinforcement learning agent, and set the state space, action space and reward function of the agent. The state space is the currently selected feature subset and the remaining candidate features; the action space is the operation of selecting one or more features from the remaining candidate features; the reward function is designed based on the backdoor criterion. When the selected feature subset satisfies the backdoor criterion, that is, it can effectively block the interference of confounding factors on the causal relationship, give the agent a positive reward, otherwise give a negative reward; the agent selects an action from the action space according to the current state according to the set strategy (such as the ε-greedy strategy), that is, selects one or more features to add to the current feature subset; according to the selected feature subset, judge whether it satisfies the backdoor criterion and calculate the reward value according to the reward function. At the same time, update the state of the agent, that is, update the selected feature subset and the remaining candidate features; let the agent continuously perform action selection and state update. Through multiple iterations, gradually select the feature subset that satisfies the backdoor criterion, and this feature subset is the selected causal feature; biomarker enhancement: metabolomics pathway activation analysis (MetPA-Activation): Collect the metabolomics data of the patient, including information such as the concentration of various metabolites.Meanwhile, obtain the pathway information related to bone metabolism from the KEGG database and construct a prior knowledge graph, which contains the association relationships between metabolites and pathways, as well as the interaction relationships between pathways; perform pathway enrichment analysis on metabolomics data and map metabolites to the pathways in the prior knowledge graph. By calculating the enrichment degree of metabolites in each pathway, determine which pathways have high activity potential in the current data; quantify the pathway activity score. According to the importance of metabolites in the pathway (such as the position of metabolites in the pathway, the degree of participation in reactions, etc.) and the enrichment degree, calculate an activity score for each pathway. For example, adopt the weighted summation method to comprehensively calculate the relevant indicators of metabolites to obtain the pathway activity score; use the calculated pathway activity score as a new feature and add it to the original data feature set to complete the enhancement of biomarkers; Temporal feature modeling: Time2Vec++ Period-aware embedding: Extract time-related features from the preprocessed data, especially the key time feature of the number of years after menopause. Divide the number of years after menopause at certain time intervals (such as months, years) to form a time series; improve the Time2Vec algorithm. On the basis of the original Time2Vec algorithm, introduce a function that can explicitly model periodicity. For example, use trigonometric functions (sine function, cosine function) to simulate the periodic changes of physiological processes such as calcium absorption, and use the time series as the input variable of these functions; through the improved Time2Vec++ algorithm, convert the time features at each time point into embedding vectors. The embedding vectors not only contain the numerical information of the time points, but also encode the periodic influence information of the number of years after menopause through the periodic changes of trigonometric functions. These embedding vectors are the extracted temporal features, and they are fused with other features to obtain the final target feature data for analysis.
[0031] In this embodiment, the node representation data integrating spatio-temporal features are grouped according to whether osteoporosis is present (divided into a diseased group and a non-diseased group). For each feature, a contingency table is constructed, where the rows represent different value categories of the feature (e.g., age can be divided into multiple age intervals), and the columns represent the categories of disease presence or absence. Based on the actual observed frequencies in the contingency table, calculations are performed using the chi-square statistic formula, and each feature will obtain a corresponding chi-square statistic value, which reflects the degree of association between the feature and the osteoporosis risk. The larger the value, the stronger the correlation between the feature and the osteoporosis risk. A significance level threshold (e.g., 0.05) is set, and the chi-square statistic value of each feature is compared with this threshold. Combining with the degrees of freedom (determined according to the number of rows and columns of the contingency table), the corresponding p-value is calculated by referring to the chi-square distribution critical value table or using statistical software. If the p-value of a feature is less than the set significance level threshold, it is considered that there is a significant association between this feature and the osteoporosis risk, and it is retained; otherwise, if the p-value is greater than the threshold, it is determined that the association between this feature and the osteoporosis risk is not significant, and it is removed from the feature set. Through this screening process, redundant and irrelevant features are removed, the data dimension is reduced, and the target feature data finally used for model training is obtained, thereby improving the training efficiency of the subsequent model and the accuracy of postmenopausal osteoporosis risk assessment.
[0032] In this embodiment, entities related to bone metabolism (such as diseases, biomarkers, drugs) and relationships (such as "affect", "participate") are extracted from medical knowledge bases such as UMLS to construct a medical knowledge graph; the entities and relationships in the knowledge graph are embedded and represented, mapped to a low-dimensional vector space, and common methods include knowledge graph embedding algorithms such as TransE and RotatE; the Transformer architecture is initialized, and its attention mechanism is specially designed; the knowledge graph embedding is used as the initial value of the Key-Value matrix, and these matrices will remain fixed or be fine-tuned during the inference process; the patient feature vector is used as the Query vector and is dynamically generated during attention calculation; the target feature data (such as multi-omics features, temporal features) extracted through feature engineering is used as the input to construct the patient feature vector; in the attention layer of MedKG-Transformer, the similarity score between the Query (patient feature) and the Key (knowledge graph embedding) is calculated, and the softmax function is applied to normalize the similarity score; the normalized attention weights are used to aggregate the Value (knowledge graph embedding) to generate a context-aware feature representation; the context features are non-linearly transformed through the feed-forward network of the Transformer to generate the final feature representation; the features are mapped to the risk prediction space through the output layer (such as a fully connected layer) to output the preliminary risk score; the temporal features of the patient (such as the number of years after menopause, regular bone density measurement values) are represented as a time series function; a neural controlled differential equation production layer is constructed, including: State variables: representing the health status of the patient (such as bone density, hormone levels) Control variables: representing external interventions or measurable factors (such as calcium intake, exercise level) Vector field function: parameterized by a neural network, describing the rate of change of state variables over time The temporal features of the patient are input into the neural controlled differential equation model, and the state variables are initialized to the currently known health indicators; the differential equation is solved through numerical integration methods (such as the Euler method, Runge-Kutta method) to predict the state variables at future time points; at each time step, the current state variables and control variables are input into the neural network to calculate the rate of change of the state; based on the predicted future state variables, the dynamic change curve of the risk over time is calculated, and the risk prediction values at different time points are output.
[0033] In this embodiment, after processing the relevant data of postmenopausal women to be evaluated according to the steps of the above-mentioned data collection, preprocessing, and feature engineering, the data is input into the trained model. The model outputs the risk probability of the woman suffering from osteoporosis. According to the set risk threshold, the risk levels are divided into low risk, medium risk, and high risk. For example, a risk probability less than 0.3 is low risk, 0.3 - 0.7 is medium risk, and greater than 0.7 is high risk.
[0034] In this embodiment, the evaluation results are fed back to doctors and patients in an intuitive manner, and corresponding intervention suggestions are provided according to different risk levels. For the low-risk population, it is recommended to maintain a healthy lifestyle and regularly conduct bone density tests; for the medium-risk population, in addition to lifestyle interventions, it is recommended to appropriately supplement calcium and vitamin D under the guidance of a doctor; for the high-risk population, it is recommended to seek medical attention in a timely manner for further examinations and treatments, and drug treatments and other measures may be required.
[0035] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above-mentioned embodiments. The above-mentioned embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A postmenopausal osteoporosis risk assessment system based on big data, characterized in that, The system includes: A data preprocessing module, which is used to collect multi-dimensional data of postmenopausal women, clean the collected data, and perform normalization processing on numerical data and one-hot encoding processing on categorical data to obtain preprocessed data; A feature extraction module, which is used to extract features related to the risk of postmenopausal osteoporosis from the preprocessed data to obtain target feature data; A risk assessment module, which is used to construct a risk assessment model based on the target feature data, and input the original data of the postmenopausal women to be evaluated after data preprocessing and feature extraction into the risk assessment model, and output the risk assessment result of the women to be evaluated for osteoporosis; A result visualization module, which is used to visualize the risk assessment result and provide corresponding intervention suggestions according to different risk levels.
2. The postmenopausal osteoporosis risk assessment system based on big data according to claim 1, characterized in that, The data preprocessing module includes: A cleaning sub-module, which is used to collect multi-dimensional data of postmenopausal women, and clean the multi-dimensional data by applying a causal-aware missing value imputation network to process missing values, where the multi-dimensional data includes structured data and unstructured data; An integration sub-module, which is used to apply adversarial adaptive normalization processing to numerical data and dynamic graph embedding encoding to categorical data, generate context-aware embedding vectors, and integrate to obtain preprocessed data.
3. The postmenopausal osteoporosis risk assessment system based on big data according to claim 2, characterized in that, The cleaning sub-module includes: Collect multi-dimensional data containing missing values and construct a dataset matrix, where the positions of the missing values are marked by a mask matrix; Run a causal discovery algorithm to analyze the causal relationship between data variables and generate a causal graph, where nodes represent variables and edges represent the causal direction; Input the dataset matrix, mask matrix, and causal graph into the generator of the causal-aware missing value imputation network. The generator constructs a conditional probability distribution based on the causal graph, applies a random mask to the multi-dimensional data to create partial observed data, generates imputation candidate values through the GAIN network, and then uses an attention mechanism to integrate causal information and adjust the imputation value distribution to output the imputed dataset.
4. The postmenopausal osteoporosis risk assessment system based on big data according to claim 2, characterized in that The integration sub-module includes: Divide the numerical data into different domains according to the source device, input the numerical data of each domain into a feature extractor, and the feature extractor processes the input data, learns to generate domain-invariant feature representations, and outputs normalized numerical data; Construct a co-occurrence matrix of categorical variables, count the co-occurrence frequencies between different categorical values, and construct a dynamic graph based on the co-occurrence frequencies, where nodes represent different values of categorical variables and the weights of edges represent the co-occurrence intensity between nodes; Input the dynamic graph into a graph attention network. The initial feature of each node in the graph attention network is a one-hot encoded vector. Calculate the attention weights between nodes through the graph attention layer, aggregate neighbor node information within the context window, capture the relationships between nodes from different perspectives through the multi-head attention mechanism, update the node embeddings, and output the context-aware embedding vectors of categorical variables.
5. The postmenopausal osteoporosis risk assessment system based on big data according to claim 1, characterized in that The feature extraction module includes: A construction sub-module, which is used to arrange the preprocessed data in a time series to obtain nodes and construct a spatio-temporal graph structure; An introducing sub-module is used to introduce LSTM units for each node in the time dimension. The LSTM units learn the long-term dependencies of data in the time series through memory cells and gating mechanisms, and output hidden states containing time features. An applying sub-module is used to aggregate the features of each node and its neighbor nodes through graph convolution operations in the spatial dimension. By means of weighted summation, the information of neighbor nodes is aggregated to update the features of the current node. After multiple layers of graph convolution operations, a node representation integrating spatio-temporal features is obtained. A grouping sub-module is used to group the node representation data integrating spatio-temporal features according to whether osteoporosis is present. A contingency table is constructed for each feature, where the rows of the contingency table represent different value categories of the feature and the columns represent the categories of whether the disease is present. A screening sub-module is used to screen out the features that have an impact on osteoporosis risk assessment through the chi-square test method based on the actual observed frequencies in the contingency table, and obtain the target feature data.
6. The postmenopausal osteoporosis risk assessment system based on big data according to claim 5, characterized in that In the spatio-temporal graph structure, each dimension data under each time point corresponds to a node, and the attributes of the node include various data indicators at that time point.
7. The postmenopausal osteoporosis risk assessment system based on big data according to claim 5, characterized in that, The screening sub-module includes: For each feature, the corresponding chi-square statistic value is obtained through the chi-square test method. The chi-square statistic value of each feature is compared with a preset threshold. If it is greater than the preset threshold, it is determined that the association between the current feature and osteoporosis risk is not significant and the feature is excluded. If it is less than the preset threshold, it is determined that there is a significant association between the current feature and osteoporosis risk and the feature is retained.
8. The postmenopausal osteoporosis risk assessment system based on big data according to claim 1, wherein, The risk assessment module includes: An input sub-module is used to preprocess and extract features from the original data of postmenopausal women to be evaluated and then input them into the risk assessment model. A generating sub-module is used to generate a preliminary risk score through the MedKG-Transformer layer of the risk assessment model. A predicting sub-module is used to input the spatio-temporal features and the preliminary risk score into the neural controlled differential equation layer to predict the dynamic change of risk over time. An integrating sub-module is used to integrate the outputs of the MedKG-Transformer layer and the neural controlled differential equation layer to generate the final risk assessment result.
9. The postmenopausal osteoporosis risk assessment system based on big data according to claim 1, wherein The construction process of the risk assessment model includes: Construct a medical knowledge graph based on the entities and relationships related to bone metabolism obtained from the target feature data, perform embedding representation on the entities and relationships in the medical knowledge graph, and map them to a low-dimensional vector space. Use the embedding of the medical knowledge graph as the initial value of the Key-Value matrix, and use the target feature data as the Query vector. In the attention mechanism of the MedKG-Transformer layer, calculate the similarity score between the Query and the Key, and apply the softmax function to normalize the similarity score. Use the normalized attention weights to aggregate the Value to generate a context-aware feature representation. Non-linearly transform the context features through the feed-forward network of the MedKG-Transformer layer to generate the final feature representation, and map the features to the risk prediction space through the output layer to output the preliminary risk score. Construct a neural controlled differential equation layer, output risk prediction values at different time points through the neural controlled differential equation layer, and finally obtain a risk assessment model.
10. The postmenopausal osteoporosis risk assessment system based on big data according to claim 9, characterized in that, The process of the neural controlled differential equation layer includes: Input the time series features of the patient into the neural controlled differential equation layer, and initialize the state variables as the currently known health indicators; Solve the differential equation by numerical integration method to predict the state variables at future time points. At each time step, input the current state variables and control variables into the neural network, calculate the state change rate, calculate the dynamic change curve of risk over time according to the predicted future state variables, and output the risk prediction values at different time points.
Citation Information
Cited By
Multi-source gene expression quantity processing method and related product
CN121366634A
Multi-mode osteoporosis layered early warning system
CN121687524A