Real-time health risk prediction method and system based on dynamic knowledge graph
Through the layered federated learning and lightweight deployment of dynamic knowledge graphs, privacy and real-time problems in medical data sharing are solved, cross-institutional health risk prediction is achieved, and prediction accuracy and real-timeness are improved.
Patent Information
- Application Number
- CN202510254929.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art has problems such as data silos and privacy conflicts, lag in updates and insufficient real-time performance in the medical field, especially when data sharing of multi-institutions is difficult to take into account the health risk predictions that are both privacy, real-time and accurate.
Using a hierarchical federated learning framework based on dynamic knowledge graphs, cross-institutional data fusion and real-time health risk prediction are achieved through differential privacy and noise- robust gradient differential dynamic update mechanisms, combining lightweight deployment and edge computing.
It improves the privacy, real-time and accuracy of health risk prediction, reduces communication overhead, and is suitable for cross-institutional medical and nursing data privacy protection and real-time risk prediction scenarios.
Smart Images

Figure CN120280136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and more particularly, to a real-time health risk prediction method and system based on a dynamic knowledge graph. Background Art
[0002] With the continuous progress of science and technology, the service methods in all walks of life have undergone earth-shaking changes. More and more intelligent devices have replaced manual services to provide people with more efficient and convenient intelligent services, and the medical field has always been a part that has received much attention. At the same time, with the popularization of Internet of Things (IoT) terminal devices in the medical field, various wearable body sensors and medical devices are applied to medical services. These IoT terminal devices can collect a large amount of medical data and upload it to a remote cloud server through the network for data processing and analysis, providing accurate and rapid medical services for patients. In recent years, with the rapid development of big data and artificial intelligence technologies, machine learning models, especially large models, have shown significant performance advantages in processing complex data analysis tasks. Artificial intelligence has achieved success in fields such as healthcare. For example, well-trained machine learning models can assist medical staff in providing faster and more accurate diagnoses and treatments for diseases. These models usually require a large amount of data for training, but the data sets often contain sensitive information, such as personal identity information, health records, or financial data, etc. For example, Patent CN119480112A provides a dynamic health adaptive monitoring method and system using artificial intelligence, which relates to the technical field of health monitoring, including: through multi-source heterogeneous health data, using a hierarchical diffusion probability model to denoise and combining an improved attention autoencoder to extract preliminary features, then learning discriminative features through a contrastive learning network, inputting the discriminative features into an adaptive wavelet neural network for multi-scale feature decomposition, and then inputting them into a dynamic graph neural network to obtain a health feature tensor, using a hierarchical memory enhancement network to obtain temporal features, and combining a dynamic routing capsule network and a causal inference network to generate a health knowledge graph, inputting the temporal features and the health knowledge graph into a deep Bayesian network for health risk assessment, using a hierarchical reinforcement learning network, combining a multi-agent collaborative optimization system and a deep deterministic policy gradient network, and obtaining an adaptive monitoring scheme through a model predictive control network. Although a health warning system based on a knowledge graph is proposed, the problem of multi-institutional data privacy is not solved; Patent CN119358035A provides a user opinion privacy protection method and system based on federated learning and differential privacy, which relates to the technical field of privacy protection, including the server receiving a client request, partitioning the data set and distributing it; the client using a differential privacy machine learning algorithm for local training and perturbing the model parameters with Laplace noise; the server aggregating the parameters using secret sharing and homomorphic encryption technologies to update the global model; the client receiving a privacy protection policy, performing privacy processing on the user opinion and then sending it to the server for analysis; although hierarchical federated learning is used to integrate medical data, edge computing is not combined, resulting in a serious inference delay problem.
[0003] In summary, there are several problems in the application of the existing training methods in the field of digital healthcare:
[0004] 1. Data islands and privacy conflicts: Strict authorization is required for data sharing between medical institutions. Centralized knowledge graphs rely on the transmission of raw data, violating privacy regulations;
[0005] 2. Lag in updates: Static knowledge graphs are difficult to capture the dynamic changes in the health status of the elderly (such as the deterioration of chronic diseases, sudden falls);
[0006] 3. Lack of real-time performance: The cloud inference mode is affected by network latency and cannot meet the millisecond-level response requirements for emergency scenarios such as falls and myocardial infarctions.
[0007] How to conduct health risk prediction that takes into account privacy, real-time performance, and accuracy is an urgent problem to be solved at present. Summary of the Invention
[0008] The purpose of the present invention is to provide a real-time health risk prediction method and system based on a dynamic knowledge graph, which can improve the privacy, real-time performance, and accuracy of health risk prediction.
[0009] The present invention provides a real-time health risk prediction method based on a dynamic knowledge graph, including the following steps: S1: Preprocess heterogeneous medical and elderly care data to obtain preprocessed data; S2: According to the preprocessed data, use a hierarchical federated learning framework to obtain a global dynamic knowledge graph; S3: Based on a noise-robust gradient difference dynamic update mechanism, update the local knowledge graph embedding models and the global model of each institution; S4: According to the global dynamic knowledge graph, the local knowledge graph embedding models of each institution, and the global model, perform lightweight deployment to obtain a lightweight graph model, and use the lightweight graph model to perform health risk prediction.
[0010] Further, step S2 specifically includes: S21: According to the preprocessed data, train the local knowledge graph embedding models of each institution, add Laplace noise that satisfies differential privacy to the gradients, and use the SMPC protocol to encrypt and aggregate the parameters of the local knowledge graph embedding models of each institution in the hierarchical federated learning framework to obtain initial global model parameters; S22: According to the initial global model parameters, use the output distribution of the local knowledge graph embedding models of each institution as the teacher signal, and use the loss function to minimize the KL divergence loss to optimize the prediction distribution of the global model to obtain a global dynamic knowledge graph.
[0011] Further, step S21 specifically includes: According to the preprocessed data, train the local knowledge graph embedding models of each institution, add Laplace noise that satisfies differential privacy to the gradients, and use the SMPC protocol to encrypt and aggregate the parameters of the local knowledge graph embedding models of each institution in the hierarchical federated learning framework to obtain initial global model parameters, as shown in the formula:
[0012]
[0013] Δf = max‖g i ‖1,
[0014] where, is the gradient after adding Laplace noise; g i is the original gradient; Lap(·) represents Laplace distribution sampling, and Laplace noise is added to the gradient for differential privacy noise injection to meet the differential privacy constraint; Δf represents the sensitivity of the model gradient, and the sensitivity is the maximum L1 norm of each local model gradient, and ∈ is the privacy budget.
[0015] Furthermore, the above loss function, such as the formula:
[0016]
[0017] where, L KL is the loss function, (h, r, t) represents the head entity-relation-tail entity triple in the knowledge graph, P global and are the predicted probability distributions of the global model and the i-th local model, respectively.
[0018] Furthermore, step S3 specifically includes: S31: Obtain the differences between each local gradient and the global gradient according to each institution's local knowledge graph embedding model and the global dynamic knowledge graph; S32: Smooth the noise interference using the moving window mean method according to the differences between each local gradient and the global gradient to obtain a difference metric; S33: Dynamically adjust the data drift threshold using Bayesian optimization. When the difference metric is greater than the data drift threshold, or the average value of the F1 scores of the validation sets in the last 3 rounds drops by more than a preset ratio, update the global model; otherwise, do not update the global model and only update each institution's local knowledge graph embedding model.
[0019] Furthermore, step S31 specifically includes: Obtain the differences between each local gradient and the global gradient according to each institution's local knowledge graph embedding model and the global dynamic knowledge graph, such as the formula:
[0020]
[0021] where, is the difference between the local gradient and the global gradient of the current i-th institution, is the local gradient of the i-th institution, and the global gradient is the gradient parameter saved by the current global model, and ∥·∥2 represents the L2 norm.
[0022] Furthermore, step S32 specifically includes: Smooth the noise interference using the moving window mean method according to the differences between each local gradient and the global gradient to obtain a difference metric, such as the formula:
[0023]
[0024] where δ is the difference metric; W is the window size, is the gradient difference in the k-th round of training of the i-th institution; t represents the number of training rounds.
[0025] Furthermore, the above-mentioned dynamic adjustment of the data drift threshold using Bayesian optimization includes: defining the search space of the data drift threshold δ threshold : δ threshold ∈[0.05, 0.15], determining the upper and lower bounds based on the 95% quantile of the historical gradient difference distribution; using the Gaussian process as the surrogate model, selecting the radial basis function as the kernel function, and the expected improvement as the acquisition function, and performing Bayesian optimization using the iterative termination condition and the objective function to obtain the data drift threshold.
[0026] Furthermore, the above-mentioned lightweight deployment includes: calculating the triplet confidence using the embedding similarity, pruning the redundant edges with a confidence less than the preset confidence to obtain the pruned global model; compressing the pruned global model using 8-bit symmetric quantization, and dynamically adjusting the scaling factor s through the calibration set to minimize the KL divergence error, as shown in the formula:
[0027]
[0028] where L quant is the KL divergence error, and P PFP32 (x) and P INT8 (x) are the floating-point and quantized distributions respectively.
[0029] The present invention also provides a real-time health risk prediction system based on a dynamic knowledge graph. The system includes the following modules: a multi-source data acquisition module configured to preprocess the heterogeneous medical and elderly care data to obtain the preprocessed data; a federated knowledge graph construction module configured to obtain the global dynamic knowledge graph using a hierarchical federated learning framework according to the preprocessed data; a dynamic update module configured to update the local knowledge graph embedding models and the global model of each institution based on a noise-robust gradient difference dynamic update mechanism; and an edge inference module configured to perform lightweight deployment according to the global dynamic knowledge graph, the local knowledge graph embedding models of each institution, and the global model to obtain a lightweight graph model, and use the lightweight graph model to perform health risk prediction.
[0030] Implementing the real-time health risk prediction method and system based on the dynamic knowledge graph provided by the present invention has the following beneficial effects:
[0031] The present invention utilizes a federated knowledge distillation mechanism to achieve multi-institutional knowledge fusion through model parameter distillation while protecting the privacy of the original data, thereby solving the problem of data silos. In particular, the innovative application of "differential privacy DP + knowledge distillation KD" adopted by the present invention in hierarchical federated learning of knowledge graphs ensures that the increase or decrease of a single data sample will not significantly affect the model output through noise injection (Lap(·)), meeting the approximate differential privacy protection requirements of (∈,δ)-DP. The minimization of KL divergence forces the global model to learn the "soft label" knowledge of the local models, avoiding the direct sharing of the original data or sensitive parameters, thereby achieving knowledge distillation optimization. While solving the problem of multi-source data heterogeneity, through the dual-engine drive of "federated learning + edge computing", a dynamic knowledge graph is constructed to achieve cross-domain data privacy protection and real-time health risk prediction.
[0032] The present invention utilizes a dynamic update mechanism of "L2 norm threshold + differential privacy" to solve the problems of inaccurate data drift detection and insufficient privacy protection in the prior art. At the same time, real-time health risk prediction is achieved through edge lightweight inference, forming a complete technical loop. Its dynamic event-triggered update mechanism measures the data distribution shift through the L2 norm of the gradient difference and triggers the global model synchronization only when the shift exceeds the threshold. Through experiments, it is verified that the communication overhead is reduced by 60% compared with the fixed-period update. Its core idea is data drift detection and dynamic update triggering. Among them, data drift detection indirectly reflects the changes in the data distribution of each institution by comparing the gradient direction differences between the local model and the global model. Dynamic update triggering means that when δ exceeds the preset threshold, it indicates that the data distribution of some institutions has significantly deviated from the global model (such as the outbreak of a new disease or the change of the device acquisition mode), and the global model synchronization needs to be triggered; otherwise, local independent updates are maintained. Its theoretical support lies in the gradient similarity hypothesis in federated learning. In an ideal situation, the gradient directions of each local model should approach the global model, and the change in the data distribution will cause the gradient direction to deviate. This phenomenon has been used in the literature to detect the impact of non-independent and identically distributed (Non-IID) data. Among them, lightweight detection indicators are used, taking advantage of the high efficiency of L2 norm calculation and its ability to intuitively reflect vector differences. Compared with most studies in the prior art that judge data drift through statistics (such as mean / variance) or model performance (such as accuracy decline), but need to share data or labels, violating privacy requirements, the present invention uses the gradient difference as a proxy indicator, without exposing the original data, and is directly related to model updates at the same time. Through the linkage of the L2 norm threshold and differential privacy, the balance between model dynamic updates and the risk of privacy leakage is achieved.
[0033] The present invention uses 8-bit quantization and model pruning techniques to reduce the size of the graph model by 75%, supporting the deployment of low-computing-power devices, thereby achieving edge lightweight inference. Its edge lightweight inference ensures real-time performance and forms a closed loop with the update mechanism of the dynamic knowledge graph.
[0034] The technical advantages of the present invention are as follows: First, lightweight computing, only the gradient difference needs to be calculated, without transmitting the original data or complex statistics, which is suitable for resource-constrained edge devices; second, privacy protection, the gradient itself has been processed by adding noise with differential privacy (DP) to avoid leakage of sensitive information; third, adaptive update: experiments show that compared with fixed-period updates (such as once a day), the communication overhead is reduced by an average of 60% (supported by experimental data); through experiments, it is shown that the present invention can resist the membership inference attack (MIA) with a success rate of <5% in terms of privacy (compared with 15% of traditional federated learning); in terms of real-time performance, the end-to-end delay from data input to warning output is <500 ms; in terms of accuracy, the F1-score of fall risk prediction reaches 92.3% (compared with 89.7% of the centralized model); it can be seen that the present invention has obvious technical synergy: 1. Synergy between SMPC and differential privacy: SMPC protects the parameter transmission process, and differential privacy protects the parameter content, providing double protection for data privacy; 2. Collaborative optimization of pruning and quantization: Pruning reduces parameter redundancy, and quantization reduces storage and computing overhead, jointly compressing the model to 10% - 20% of the original size; 3. Closed-loop of dynamic threshold and edge inference: Data drift detection triggers model update, and lightweight inference ensures real-time response, forming an adaptive health monitoring system;
[0035] In summary, the present invention realizes cross-institutional privacy data fusion through hierarchical federated learning, combines dynamic event-triggered updates to reduce communication overhead, and deploys a lightweight inference engine at the edge to meet the real-time warning requirements, which can improve the privacy, real-time performance, and accuracy of health risk prediction, and is applicable to cross-institutional medical and elderly care data privacy protection and real-time risk prediction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0037] Figure 1 is the flowchart of the real-time health risk prediction method based on dynamic knowledge graph provided by the present invention;
[0038] Figure 2 is the overall architecture diagram of the real-time health risk prediction system based on dynamic knowledge graph provided by the present invention;
[0039] Figure 3 is the relationship diagram between pruning rate and accuracy provided by the present invention;
[0040] Figure 4 is the privacy budget and leakage risk curve graph provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific implementation manners of the present invention will now be described in detail with reference to the drawings.
[0042] Figure 1 Shows a schematic diagram of the real-time health risk prediction method based on a dynamic knowledge graph in this embodiment. In this embodiment, the real-time health risk prediction method based on a dynamic knowledge graph includes the following steps:
[0043] S1: Preprocess the heterogeneous medical and elderly care data to obtain the preprocessed data;
[0044] In an exemplary embodiment, the sources of heterogeneous medical and elderly care data include hospital data sources, elderly care data sources, and wearable devices; the preprocessing includes data desensitization, entity alignment, and multimodal fusion;
[0045] In an exemplary embodiment, for data desensitization, a privacy protection technology based on k-anonymization is adopted to ensure that individual data cannot be traced; for entity alignment, a joint embedding learning method based on a cross-modal graph neural network is used to align medical entities of different institutions; for multimodal fusion, text, images, and sensor data are fused through an attention mechanism to generate standardized structured data;
[0046] As an exemplary embodiment, in step S1, a cross-domain multimodal medical and elderly care data acquisition layer is constructed to receive heterogeneous data from hospitals, nursing homes, and wearable devices. The heterogeneous data includes structured electronic health records (EHRs), time-series sensor data, and unstructured nursing texts; for data desensitization, a privacy protection technology based on k-anonymization is adopted to ensure that individual data cannot be traced; for entity alignment, a joint embedding learning method based on a cross-modal graph neural network (Cross-MM-GNN) is used to align medical entities of different institutions (such as "blood pressure" and "BP"), and the alignment accuracy is ≥95%; for multimodal fusion, text, images, and sensor data are fused through an attention mechanism to generate standardized structured data;
[0047] S2: According to the preprocessed data, use a hierarchical federated learning framework to obtain a global dynamic knowledge graph;
[0048] In an exemplary embodiment, step S2 specifically includes:
[0049] S21: According to the preprocessed data, train the local knowledge graph embedding models of each institution, add Laplace noise that satisfies differential privacy to the gradients, and use the SMPC protocol to encrypt and aggregate the parameters of the local knowledge graph embedding models of each institution in the hierarchical federated learning framework to obtain the initial global model parameters;
[0050] In an exemplary embodiment, step S21 specifically includes: training each institution's local knowledge graph embedding model based on the preprocessed data, adding Laplace noise that satisfies differential privacy to the gradients, and using the SMPC protocol to encrypt and aggregate the parameters of each institution's local knowledge graph embedding model in a hierarchical federated learning framework to obtain initial global model parameters, as shown in the formula:
[0051]
[0052] Δf=max‖g i ‖1,
[0053] where, is the gradient after adding Laplace noise; g i is the original gradient; Lap(·) represents Laplace distribution sampling, and Laplace noise is added to the gradients for differential privacy noise injection to satisfy differential privacy constraints; Δf represents the sensitivity of the model gradients, and the sensitivity is the maximum L1 norm of the gradients of each local model, and ∈ is the privacy budget;
[0054] As an exemplary embodiment, the privacy budget is not greater than 1.0, that is, ∈≤1.0; in step S21, each institution locally trains a knowledge graph embedding model (such as TransE), adds Laplace noise that satisfies differential privacy (∈≤1.0) to the gradients, and uses the SMPC protocol to encrypt and aggregate the local parameters in a hierarchical architecture (regional node → global node) to generate initial global model parameters;
[0055] S22: According to the initial global model parameters, using the output distribution of each institution's local knowledge graph embedding model as the teacher signal, and using the loss function, minimize the KL divergence loss to optimize the prediction distribution of the global model to obtain the global dynamic knowledge graph;
[0056] In an exemplary embodiment, the loss function is as shown in the formula:
[0057]
[0058] where, L KL is the loss function, (h, r, t) represents the head entity-relation-tail entity triple in the knowledge graph, P global and are the prediction probability distributions of the global model and the i-th local model respectively;
[0059] As an exemplary embodiment, in step S2, a horizontal hierarchical federated learning framework is adopted, and the local knowledge graphs of each institution are aggregated under privacy protection through Secure Multi-Party Computation (SMPC) to generate a global dynamic knowledge graph; wherein the hierarchical federated learning framework includes a three-layer architecture of an edge node layer, a local server layer, and a global aggregation center; Laplace noise is added to the training gradients of the local knowledge graph embedding models of each institution to satisfy the differential privacy budget ∈ ≤ 1.0, and the KL divergence between the global model and the outputs of each local model is minimized through knowledge distillation technology to achieve knowledge transfer under privacy protection; first, local model training and differential privacy noise injection are performed. When training the knowledge graph embedding model in a local institution (such as a hospital or a nursing home), Laplace noise is added to the gradients to satisfy the differential privacy constraint; it should be noted that Secure Multi-Party Computation (SMPC) allows multiple participants to collaboratively complete computational tasks without revealing local data; in federated learning, SMPC is mainly used to protect the parameter aggregation process and prevent the leakage of gradients or model parameters;
[0060] S3: Update the local knowledge graph embedding models and the global model of each institution based on a noise-robust gradient difference dynamic update mechanism;
[0061] In an exemplary embodiment, step S3 specifically includes:
[0062] S31: Obtain the differences between the local gradients and the global gradients based on the local knowledge graph embedding models of each institution and the global dynamic knowledge graph;
[0063] In an exemplary embodiment, step S31 specifically includes: obtaining the differences between the local gradients and the global gradients based on the local knowledge graph embedding models of each institution and the global dynamic knowledge graph, as shown in the formula:
[0064]
[0065] wherein, is the difference between the local gradient and the global gradient of the current i-th institution, is the local gradient of the i-th institution, and the global gradient is the gradient parameter saved by the current global model, and ∥·∥2 represents the L2 norm;
[0066] S32: Smooth the noise interference using the moving window mean method according to the differences between the local gradients and the global gradients to obtain a difference metric;
[0067] In an exemplary embodiment, step S32 specifically includes: according to the differences between the local gradients and the global gradient, using the moving window mean method to smooth the noise interference to obtain a difference metric, as shown in the formula:
[0068]
[0069] where δ is the difference metric; W is the window size, is the gradient difference in the k-th round of training of the i-th institution; t represents the number of training rounds;
[0070] As an exemplary embodiment, in this embodiment, the window size W = 5;
[0071] S33: Dynamically adjust the data drift threshold using Bayesian optimization. When the difference metric is greater than the data drift threshold, or when the mean value of the F1 scores of the validation set in the last 3 rounds drops by more than a preset ratio, update the global model; otherwise, do not update the global model, but only update the local knowledge graph embedding models of each institution;
[0072] In an exemplary embodiment, dynamically adjusting the data drift threshold using Bayesian optimization includes:
[0073] Define the search space of the data drift threshold δ threshold : δ threshold ∈[0.05, 0.15], and determine the upper and lower bounds based on the 95% quantile of the historical gradient difference distribution;
[0074] Use the Gaussian process as the surrogate model, select the radial basis function as the kernel function, and the expected improvement as the acquisition function. Use the iterative termination condition and the objective function to perform Bayesian optimization to obtain the data drift threshold;
[0075] As an exemplary embodiment, in step S33, dynamically adjusting the data drift threshold using Bayesian optimization includes:
[0076] (1) Define the search space: The search space of the threshold δ threshold is defined as δ threshold ∈[0.05, 0.15], and determine the upper and lower bounds based on the 95% quantile (δ = 0.1) of the historical gradient difference distribution;
[0077] (2) Bayesian optimization process:
[0078] 1) Surrogate model: Adopt the Gaussian process (GP) as the surrogate model, and select the radial basis function (RBF) as the kernel function:
[0079]
[0080] Among them, l is the length scale parameter, and the initial value of l is 0.1;
[0081] 2) Acquisition function: Use the Expected Improvement (EI) as the acquisition function to balance exploration and exploitation:
[0082] EI(x) = E[max(f(x) - f(x + ), 0]],
[0083] where f(x + ) is the current optimal objective function value;
[0084] 3) Iteration termination condition: The maximum number of iterations is 50 rounds; or the change in the objective function < 1% (that is, the absolute value of the difference in the objective function values between two consecutive iterations < 0.01);
[0085] 4) The relevant pseudocode is as follows:
[0086]
[0087]
[0088] 5) Objective function: The objective function is defined as maximizing the model performance and constraining the communication frequency:
[0089]
[0090] where λ = 0.3 is the penalty coefficient;
[0091] The dynamic update trigger condition is "δ > δ threshold or the average F1-score of the validation set in the last 3 rounds drops by ≥ 5%", triggering the global model update, otherwise only updating the local model;
[0092] The average F1-score of the validation set in the last 3 rounds refers to evaluating the output of the model using the validation set, the local knowledge graph embedding models of each institution, and the global dynamic knowledge graph to obtain the F1-score, and taking the average of the F1-scores in the last 3 rounds;
[0093] S4: According to the global dynamic knowledge graph, the local knowledge graph embedding models of each institution, and the global model, perform lightweight deployment to obtain a lightweight graph model, and use the lightweight graph model for health risk prediction;
[0094] In an exemplary embodiment, the lightweight deployment includes:
[0095] Using the embedding similarity, calculate the confidence of triples, prune the redundant edges with a confidence less than the preset confidence, and obtain the pruned global model;
[0096] Compress the pruned global model using 8-bit symmetric quantization, dynamically adjust the scaling factor s through the calibration set, and minimize the KL divergence error, as shown in the formula:
[0097]
[0098] where L quant is the KL divergence error, and P PFP32 (x) and P INT8 (x) are the floating-point and quantized distributions respectively;
[0099] In an exemplary embodiment, the preset confidence is 0.7;
[0100] As an exemplary embodiment, in step S4, perform lightweight deployment on the global model, including:
[0101] S41: Calculate the triple confidence C(h,r,t)=cos(h,t) based on the embedding similarity, and prune the redundant edges with a confidence < 0.7 (experimental results show that when the pruning rate is 40%, the accuracy loss ≤ 3%);
[0102] S42: Compress the model using 8-bit symmetric quantization (INT8), dynamically adjust the scaling factor s through the calibration set, and minimize the KL divergence error:
[0103]
[0104] where s is the scaling factor, and PFP32 and PINT8 are the floating-point and quantized distributions respectively;
[0105] S43: Deploy the lightweight model on an edge device (such as NVIDIA Jetson Nano) to achieve real-time prediction of health risks with an end-to-end latency < 500ms.
[0106] As another exemplary embodiment, perform 8-bit integer quantization and redundant edge pruning on the global model; first perform redundant edge pruning (confidence < 0.7) on the global model, and then perform 8-bit symmetric quantization on the pruned model to reduce the accuracy loss. Using the symmetric quantization method, map the FP32 parameters to the INT8 range (-128 to 127), and fine-tune the quantization error through the calibration set;
[0107] Next, introduce the pruning criteria and quantization method of this embodiment; for the pruning criteria, use structured channel pruning and select the channels to be retained based on the weight importance score; the importance score is as shown in the formula:
[0108]
[0109] The above formula represents the sum of the absolute values of the weights within the channel; generally remove the k% channels with the lowest scores (such as k = 30%);
[0110] Gradient significance pruning refers to optimizing the pruning decision by combining gradient information, as shown in the formula:
[0111]
[0112] The quantization method and calibration of this embodiment are introduced below; this embodiment uses symmetric quantization (INT8), and its calibration range is to determine the scaling factor according to the maximum value max(|w|) statistically calculated from the training data, as shown in the formula:
[0113]
[0114] Among them, the quantization formula is as follows:
[0115] w int8 = round(w fp32 )·s,
[0116] The dequantization is as shown in the formula:
[0117]
[0118] Calibration optimization uses KL divergence to minimize the distribution difference before and after quantization, as shown in the formula:
[0119] min s KL(P fp32 ||P int8 );
[0120] In an exemplary embodiment, edge lightweight inference, i.e., real-time optimization, is implemented in the following ways: One is to adopt model compression technology. Regarding pruning, its principle is: by removing redundant neurons or connections in the neural network (such as channels with weights close to zero), the number of model parameters and the amount of computation are reduced. Its implementation method is: structured pruning: removing the entire convolutional kernel or channel to maintain a hardware-friendly computational structure; unstructured pruning: removing individual weight parameters, which requires support from a sparse computing library; balancing compression rate and accuracy. In an experimental example, channel pruning is performed on the ResNet-18 model, removing 30% of the convolutional channels: the number of parameters is reduced by 45% (from 11.7M to 6.4M); the inference speed is increased by 50% (CPU latency is reduced from 120ms to 60ms); the accuracy loss is controlled within 3% (Top-1 accuracy is reduced from 70.2% to 68.5%); For quantization, its principle is: converting model parameters and activation values from 32-bit floating-point numbers (FP32) to a low-precision format (such as 8-bit integer INT8) to reduce memory occupancy and computational overhead. The implementation method is: static quantization: calibrating the quantization range offline, suitable for deploying fixed models; dynamic quantization: dynamically adjusting quantization parameters during runtime to adapt to changes in input data; balancing compression rate and accuracy: the formula for quantization bit width selection is optimized based on the signal-to-noise ratio (SNR). In an experimental example, quantizing the FP32 model to INT8 reduces the model size by 75% (from 100MB to 25MB); the inference speed is increased by 4 times (GPU latency is reduced from 50ms to 12ms); the accuracy loss ≤ 1% (classification task accuracy is reduced from 92.1% to 91.5%).
[0121] It should be noted that the principle of knowledge distillation is: using a large teacher model (such as BERT) to guide the training of a small student model, and transferring knowledge through soft labels to improve the performance of the small model. Considering the balance between compression rate and accuracy, in an experimental example, distilling BERT-base (110M parameters) into TinyBERT (14M parameters) reduces the model size by 87%; the inference speed is increased by 6 times (from 200ms to 33ms); the accuracy loss is 8% (GLUE benchmark score is reduced from 78.5 to 72.3);
[0122] The following introduces the real-time inference process. First, dynamic knowledge graph reception and preprocessing are carried out, including data reception: the edge node receives the dynamic knowledge graph (in JSON format) sent by the central server through the MQTT / HTTP protocol; the local cache mechanism includes: cache policy: adopting the LRU (Least Recently Used) algorithm to retain frequently accessed sub-graphs; cache capacity: dynamically adjusted according to the device memory limit (e.g., medical sensors retain the latest 100 health events); data preprocessing is carried out, including: entity alignment: matching local sensor data with knowledge graph entities (such as patient ID association); feature normalization: normalizing numerical features (such as heart rate, blood pressure) (mean-variance normalization); then lightweight model loading and inference are carried out, including: model loading: loading the pre-compressed lightweight model from local storage (such as the ONNX format model after pruning + quantization); memory occupancy optimization, using memory mapping technology to avoid loading all model parameters; parallel computing optimization: GPU acceleration, using CUDA cores for parallel convolution and matrix operations (such as batch size = 16); multi-threaded inference, splitting tasks on the CPU side into multiple threads (such as OpenMP parallelization); inference execution: input, the preprocessed dynamic knowledge graph fragment (such as patient real-time vital sign data); output, health risk prediction results (such as the probability of a heart attack).
[0123] In an exemplary embodiment, real-time monitoring of a medical edge device is carried out. The device is a portable electrocardiogram monitor (edge node), and the task is to predict the risk of arrhythmia in patients in real time; in the model deployment stage, a lightweight CNN model (size 8MB) after pruning (removing 40% of the channels) and INT8 quantization is used; during the inference process, real-time electrocardiogram signals (sampling rate 500Hz) are received, and the latest 10 seconds of data are retained through the local cache; GPU-accelerated inference (NVIDIA Jetson Nano), with a latency of 15ms; the results are transmitted to the doctor's terminal through BLE, with a latency of 5ms; the performance metrics are: end-to-end latency: 20ms; accuracy: 95% (compared with 97% of the original model); device temperature: ≤45°C (no overheating risk); it can be seen that the following technical effects are achieved: efficient compression, a combination strategy of pruning + quantization + distillation, achieving a model size reduction of more than 80% with controllable accuracy loss; low-latency inference, GPU acceleration and multi-threaded optimization, with an end-to-end latency ≤50ms; resource adaptability, dynamic caching and memory mapping technology, adapting to different edge hardware resources.
[0124] As an exemplary embodiment, in step S4, a lightweight knowledge graph inference module is deployed at the edge side, and a temporal graph convolutional network (TGCN) is used to predict health risks for real-time sensor data, outputting a risk level and a warning signal. This mainly includes: performing 8-bit integer quantization (INT8) and redundant edge pruning (confidence < 0.7) on the global knowledge graph embedding to generate a lightweight graph model; deploying a real-time warning interface based on the MQTT protocol, and when the risk score Rt exceeds the dynamic calibration threshold θ, pushing an alarm message to the caregiver's terminal, with the response delay verified to be less than 500 ms through actual measurement.
[0125] This embodiment provides a real-time health risk prediction system based on a dynamic knowledge graph. The system includes the following modules:
[0126] A multi-source data acquisition module, configured to: preprocess heterogeneous medical and elderly care data to obtain preprocessed data; a federated knowledge graph construction module, configured to: use a hierarchical federated learning framework to obtain a global dynamic knowledge graph based on the preprocessed data; a dynamic update module, configured to: update the local knowledge graph embedding models of each institution and the global model based on a noise-robust gradient difference dynamic update mechanism; an edge inference module, configured to: perform lightweight deployment based on the global dynamic knowledge graph, the local knowledge graph embedding models of each institution, and the global model to obtain a lightweight graph model, and use the lightweight graph model to predict health risks.
[0127] As an exemplary embodiment, the real-time health risk prediction system based on a dynamic knowledge graph can also be implemented in the following way. In this embodiment, the overall architecture of the real-time health risk prediction system based on a dynamic knowledge graph is as Figure 2 shown. The edge node layer is responsible for local model training, the local server layer performs noise injection and gradient upload, and the global aggregation center realizes parameter aggregation through secure multi-party computing. The system specifically includes: a multi-source data acquisition module: used to interface with the hospital HIS system, the elderly care home management platform, and the wearable device API; a federated knowledge graph construction module: used to support local model training, global parameter aggregation, and conflict resolution; an edge inference engine: integrated in edge computing devices (such as Raspberry Pi, edge servers) to perform real-time risk prediction; a visualization interaction interface: providing functions such as health status graph display, risk history backtracking, and warning log management.
[0128] In an exemplary embodiment, the above real-time health risk prediction system based on a dynamic knowledge graph can be implemented in the following way; in this embodiment, the collaboration between hierarchical federated learning and knowledge distillation includes:
[0129] (1) Secure Multi-Party Computation (SMPC): Hierarchically encrypt and aggregate local model parameters through secure multi-party computation protocols to ensure the privacy of parameter transmission; regional nodes adopt Shamir secret sharing sharding (threshold t = 3), and the global node restores the complete parameters through the Beaver triple protocol. In this embodiment, the communication overhead is reduced by 40% (from 1200 MB / month to 720 MB / month);
[0130] (2) Use the noisy output distribution of each local model as the teacher signal, and optimize the global model prediction distribution by minimizing the KL divergence loss, as shown in the formula:
[0131]
[0132] In this embodiment, the privacy leakage risk after noisy distillation is reduced from 15% to 2%, and the F1-score is increased by 3.2% (89.7% → 92.3%);
[0133] (3) Noise injection and secure aggregation: Add Laplace noise (∈ = 0.5) to the gradient, combined with Rényi differential privacy (RDP) noise (α = 2, ∈ = 0.3, δ = 10 -5 ), to ensure that the total privacy budget ∈ total = 0.8 ≤ 1.0; where α ∈ (0, 1) ∪ (1, +∞) is the order parameter, used to control the sensitivity to distribution differences;
[0134] In this embodiment, the noise-robust dynamic update mechanism includes:
[0135] (1) Gradient difference measurement: Calculate the L2 norm difference between each local gradient and the global gradient , as shown in the formula:
[0136]
[0137] (2) Sliding window smoothing: Adopt the sliding window mean method with a window size W = 5 to smooth the noise interference and calculate the difference measurement, as shown in the formula:
[0138]
[0139] (3) Bayesian optimization for dynamic adjustment: Maximize the model performance and constrain the communication frequency, as shown in the formula:
[0140]
[0141] where λ = 0.3 is the penalty coefficient; the maximum number of iterations is 50 rounds, or the change in the objective function < 1%;
[0142] (4) Trigger condition priority: When the average value of the F1-score of the validation set in the last 3 rounds drops by ≥ 5%, the global model update is immediately triggered, and the end-to-end delay < 200 ms;
[0143] In this embodiment, the edge lightweight deployment includes:
[0144] (1) Redundant edge pruning: Dynamically prune based on the triple confidence C(h,r,t) = cos(h,t), and the threshold 0.7 achieves the optimal balance among the pruning rate (40%), the accuracy loss (≤ 2.8%), and the critical path retention rate (98%);
[0145] (2) INT8 quantization: Dynamically adjust the scaling factor s through the calibration set to minimize the KL divergence error between the FP32 and INT8 distributions:
[0146]
[0147] In this embodiment, the number of model parameters after quantization is reduced by 45% (1.2M → 0.66M), the inference speed is increased by 4 times (60 ms → 15 ms), and the end-to-end delay < 30 ms.
[0148] In an exemplary embodiment, the above real-time health risk prediction system based on a dynamic knowledge graph can be implemented in the following manner; in this embodiment, the collaborative process of federated learning and knowledge distillation is as follows: In the process of cross-institutional data alignment and privacy enhancement, it includes:
[0149] (1) Cross-modal entity alignment and collaborative k-anonymization: Adopt a cross-institutional coordination mechanism: Each medical institution (A / B / C) adopts a unified equivalence class division rule, defines core sensitive fields (such as age ± 5 years, blood pressure ± 10 mmHg) as the anonymization benchmark to ensure semantic consistency of cross-institutional data after k-anonymization (k = 5); for example, the data of "hypertensive" patients from different institutions are uniformly divided into the equivalence class of age: 50 - 60 years, systolic blood pressure: 140 - 160 mmHg, age: 50 - 60 years; Conduct anti-re-identification attack tests: Simulate re-identification attacks through a generative adversarial network (GAN). In this embodiment, after injecting 10% adversarial samples, the re-identification probability of the secondarily desensitized data increases from 0.8% to 1.2%, which is significantly lower than the baseline method (12% → 15%);
[0150] (2) Local training and privacy budget management: Inject gradient noise: Each institution locally trains the TransE model, and Laplace noise (∈ = 0.5) is added to the gradient. The sensitivity Δf is defined as the maximum value of the L1 norm of the gradient: Δf = max‖g i‖1. The noise intensity is Lap(Δf / ∈), ensuring that the privacy leakage risk of a single training is ≤ 1%; cumulative privacy budget control is carried out: regarding the gradient noise, the privacy budget per single training ∈gradient = 0.5; regarding the knowledge distillation noise, the privacy budget per single distillation ∈distillation = 0.3; regarding the total budget, after 100 rounds of training, the total privacy budget is calculated by the Moments Accountant method: ∈total = ∈gradient + ∈distillation = 0.5 + 0.3 = 0.8 ≤ 1.0; regarding the compliance proof, the differential privacy constraint (∈ ≤ 1.0) is satisfied;
[0151] In this embodiment, hierarchical secure aggregation and privacy-enhanced distillation include:
[0152] (1) Hierarchical federated aggregation and Byzantine fault tolerance: Regional node secure aggregation is performed: The parameters of each institution are sharded through Shamir secret sharing (threshold t = 3), and the regional nodes use the BFT-SMaRt protocol to achieve Byzantine fault-tolerant aggregation, tolerating f = 1 malicious nodes (such as parameter tampering); experiments show that after introducing BFT-SMaRt, the communication delay increases by 12% (from 230ms to 258ms), but the success rate of malicious attacks drops from 18% to 0%; Global aggregation optimization is carried out: The global node restores the complete model parameters based on the intermediate parameters through the Beaver triple protocol, and the calculation efficiency is increased by 15% (from 320ms to 272ms);
[0153] (2) Differential privacy knowledge distillation: Teacher signal noise injection: Rényi differential privacy (RDP) noise (α = 2, ∈ = 0.3, δ = 10 -5 ) is added to the output distribution of the local model, and the noise intensity σ = 0.1, satisfying the (∈, δ)-differential privacy constraint; Global model optimization: The global model is optimized by minimizing the KL divergence loss: Experiments show that the privacy leakage risk after noise injection distillation drops from 15% to 2%, and the F1-score is increased by 3.2% (89.7% → 92.3%), as shown in Table 1:
[0154] Table 1: Privacy-performance balance verification table
[0155]
[0156] This embodiment conducts dynamic update threshold calibration experiments, including:
[0157] (1) Forced trigger mechanism and response time verification: The dynamic update trigger condition is "F1-score decrease first", that is, when the average value of the F1-score of the validation set in the recent 3 rounds decreases by ≥ 5%, the global model update is immediately triggered, regardless of whether the gradient difference (δ) reaches the threshold; If only δ > δ thresholdIf the F1-score does not decrease, update as needed to avoid ineffective communication; avoid deterioration of model performance caused by noise interference or gradient difference delay, and prioritize ensuring the prediction accuracy of high-risk medical events; simulate the real-time monitoring data stream of myocardial infarction patients, inject sudden abnormal signs (such as ST segment elevation, sudden drop in heart rate), and test the end-to-end delay from the detection of F1-score decrease to model update. The forced trigger mechanism compresses the response time to <200ms, while the F1-score increases by 2.9%, meeting the requirements of the first-aid scenario, as shown in Table 2;
[0158] Table 2: Comparison Table of Delay and F1 Score
[0159] Scenario Trigger condition End-to-end delay F1-score Myocardial infarction warning (forced trigger) F1 decrease ≥ 5% 182 ms 93.10% Myocardial infarction warning (non-forced) δ > 0.1 and F1 decrease ≥ 5% 245 ms 90.20%
[0160] (2) Optimization by dynamic adjustment method: Perform objective function and Bayesian optimization: Maximize model performance and constrain the communication frequency. The objective function is defined as: where λ = 0.3 is the penalty coefficient to balance accuracy and efficiency; experiments show that within the search space δ threshold ∈
[0161] [0.05, 0.15], through 50 rounds of iteration of the Gaussian process surrogate model, it converges to the optimal threshold δ threshold = 0.08, as shown in Table 3;
[0162] Table 3: Performance Comparison Table
[0163] Indicator Historical data method (δ = 0.1) Dynamic adjustment method (λ = 0.3) Communication frequency Once a day Once every three days on average Communication overhead (MB / month) 1200 480(↓60%) F1-score (normal) 89.70% 92.3% (loss ≤ 3.4%) F1-score (myocardial infarction warning) 90.20% 93.10% End-to-end delay (forced trigger) 245 ms 182 ms
[0164] The multi-scenario robustness test situation is shown in Table 4;
[0165] Table 4: Multi-scenario Robustness Test Situation Table
[0166]
[0167]
[0168] Extreme scenario test: In an environment with high noise (σ = 0.2) and network jitter (packet loss rate 10%), the forced trigger mechanism still maintains an end-to-end delay <200ms and a false trigger rate ≤ 12%;
[0169] Verification in resource-constrained environment: In the low-power mode of Jetson Nano, the dynamic adjustment method maintains an F1-score ≥ 90% and reduces the communication frequency to once every 3 days;
[0170] The effects of lightweight deployment are as follows:
[0171] (1) Pruning Threshold Optimization and Critical Path Protection: Dynamically prune redundant edges based on triple confidence (C(h,r,t) = cos(h,t)), and optimize the confidence threshold through ablation experiments; the threshold of 0.7 achieves the best balance among pruning rate (40%), precision loss (≤2.8%), and critical edge retention rate (98%); set a lower confidence limit (C≥0.8) for high-risk medical-related edges (such as "hypertension → myocardial infarction") to force the retention of such edges. Experiments show that this mechanism reduces the critical edge loss rate from 15% to 2%; the data of the threshold ablation experiment is shown in Table 5;
[0172] Table 5: Data Table of Threshold Ablation Experiment
[0173]
[0174] Knowledge Graph Topological Structure Analysis: Based on medical expert annotations, define 10 high-risk paths (such as "hypertension → myocardial infarction") and evaluate the integrity of the paths after pruning; experiments show that the threshold of 0.7 achieves the best balance between pruning efficiency and the retention of critical medical associations, as shown in Table 6;
[0175] Table 6: Comparison Table of Threshold, Pruning Rate, and Critical Path Retention Rate
[0176] Threshold Pruning rate Critical path retention rate 0.6 52% 92% 0.7 40% 98% 0.8 30% 99%
[0177] (2) Stability Test of the Quantization Model under Extreme Data
[0178] Extreme Data Injection Experiment: Inject 10% abnormal sign data (such as systolic blood pressure > 200 mmHg, heart rate < 30 beats / min) into the test set to evaluate the robustness of the quantization model; the results show that KL divergence calibration performs more stably under extreme data, and the precision loss is significantly lower than that of the maximum calibration method (1.4% vs. 8.7%), as shown in Table 7;
[0179] Table 7: Stability Test Table of the Quantization Model under Extreme Data
[0180] Model version F1-score of normal data F1-score of abnormal data Accuracy loss FP32 (original) 92.30% 85.60% 6.70% INT8 (KL calibration) 89.10% 84.20% 1.40% INT8 (maximum calibration) 88.50% 79.80% 8.70%
[0181] Dynamic Range Adjustment Strategy: Adopt the 99.9% percentile truncation method to avoid the amplification of quantization errors caused by extreme values (such as systolic blood pressure > 200 mmHg); the calibration set contains 5% abnormal data to ensure the dynamic adaptability of the scaling factor s;
[0182] (3) The performance comparison of joint optimization is shown in Table 8;
[0183] Table 8: Performance Comparison Table of Joint Optimization
[0184] Indicator Original model (FP32) Pruned + quantized model (INT8) Number of parameters 1.2M 0.66M(↓45%) Inference speed (ms / sample) 60 15 (↑ 4 times) F1-score (normal) 92.30% 89.1% (loss ≤ 3.4%) F1-score (abnormal) 85.60% 84.2% (loss ≤ 1.4%) End-to-end delay 75 ms 30 ms Critical path retention rate 100% 98%
[0185] (4) INT8 Quantization: Dynamically adjust the scaling factor s through the calibration set, and the objective function is to minimize the KL divergence error between the FP32 and INT8 distributions: L quant = ∑ x P PFP32 (x) log(P PFP32 (x) / P INT8 (x)), Experiments show that when L quant < 0.1, the accuracy loss of the quantized model is controllable (≤ 3.4%); Adopt percentile calibration (99.9% truncation) to avoid the amplification of quantization errors caused by extreme values (such as abnormal blood pressure values); Inject 10% abnormal sign data (such as systolic blood pressure > 200 mmHg or heart rate < 30 beats per minute) into the test set, as shown in Table 9;
[0186] Table 9: Comparison Table of F1 Scores of Different Models
[0187] Model version F1-score of normal data F1-score of abnormal data FP32 (original) 92.30% 85.60% INT8 (KL calibration) 89.10% 84.2% (loss ≤ 1.4%) INT8 (maximum calibration) 88.50% 79.8% (loss ≥ 5.8%)
[0188] The KL divergence calibration performs more stably under abnormal data, and the accuracy loss is significantly lower than that of the maximum calibration method, as shown in Table 10;
[0189] Table 10: Comparison Table of Joint Optimization Performance
[0190]
[0191]
[0192] The relationship diagram between the pruning rate and accuracy is as Figure 3 shown, and the curve of privacy budget and leakage risk is as Figure 4 shown; In this embodiment, the inference speed is increased by 4 times (60 ms → 15 ms) on Jetson Nano, and the energy consumption is reduced by 50%, adapting to resource-constrained medical edge devices.
[0193] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. These all belong to the protection scope of the present invention.
Claims
1. A real-time health risk prediction method based on a dynamic knowledge graph, characterized in that, It includes the following steps: S1: Preprocess the heterogeneous medical and elderly care data to obtain preprocessed data; S2: Based on the preprocessed data, use a hierarchical federated learning framework to obtain a global dynamic knowledge graph; S3: Based on a noise-robust gradient difference dynamic update mechanism, update the local knowledge graph embedding models and the global model of each institution; S4: Based on the global dynamic knowledge graph, the local knowledge graph embedding models and the global model of each institution, perform lightweight deployment to obtain a lightweight graph model, and use the lightweight graph model to predict health risks.
2. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 1, wherein Step S2 specifically includes: S21: Based on the preprocessed data, train the local knowledge graph embedding models of each institution, add Laplace noise that satisfies differential privacy to the gradients, and use the SMPC protocol to encrypt and aggregate the parameters of the local knowledge graph embedding models of each institution in the hierarchical federated learning framework to obtain initial global model parameters; S22: Based on the initial global model parameters, use the output distribution of the local knowledge graph embedding models of each institution as the teacher signal, and use the loss function to minimize the KL divergence loss to optimize the prediction distribution of the global model to obtain a global dynamic knowledge graph.
3. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 2, wherein Step S21 specifically includes: Based on the preprocessed data, train the local knowledge graph embedding models of each institution, add Laplace noise that satisfies differential privacy to the gradients, and use the SMPC protocol to encrypt and aggregate the parameters of the local knowledge graph embedding models of each institution in the hierarchical federated learning framework to obtain initial global model parameters, as shown in the formula: Among them, is the gradient after adding Laplace noise; g i is the original gradient; Lap(·) represents Laplace distribution sampling, and Laplace noise is added to the gradient for differential privacy noise injection to meet the differential privacy constraint; Δf represents the sensitivity of the model gradient, and the sensitivity is the maximum L1 norm of each local model gradient, and ∈ is the privacy budget.
4. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 2, characterized in that The loss function, as shown in the formula: Among them, L KL is the loss function, (h, r, t) represents the head entity-relation-tail entity triple in the knowledge graph, and P global and are the predicted probability distributions of the global model and the i-th local model, respectively.
5. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 1, wherein, Step S3 specifically includes: S31: Based on the local knowledge graph embedding models and the global dynamic knowledge graph of each institution, obtain the differences between the local gradients and the global gradients; S32: Based on the differences between the local gradients and the global gradients, use the moving window mean method to smooth the noise interference to obtain a difference metric; S33: Use Bayesian optimization to dynamically adjust the data drift threshold. When the difference metric is greater than the data drift threshold, or the average F1 score of the validation set in the last 3 rounds drops by more than a preset ratio, update the global model; otherwise, do not update the global model, and only update the local knowledge graph embedding models of each institution.
6. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 5, wherein Step S31 specifically includes: Based on the local knowledge graph embedding models and the global dynamic knowledge graph of each institution, obtain the differences between the local gradients and the global gradients, as shown in the formula: where, is the difference between the local gradient and the global gradient of the current i-th institution, is the local gradient of the i-th institution, and the global gradient is the gradient parameter saved by the current global model, and ∥·∥2 represents the L2 norm.
7. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 5, wherein Step S32 specifically includes: Based on the differences between the local gradients and the global gradients, use the moving window mean method to smooth the noise interference to obtain a difference metric, as shown in the formula: where δ is the difference metric; W is the window size, is the gradient difference of the i-th institution in the k-th round of training; t represents the number of training rounds.
8. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 5, wherein The use of Bayesian optimization to dynamically adjust the data drift threshold includes: Define the data drift threshold δ threshold Search space: δ threshold ∈[0.05, 0.15], and determine the upper and lower bounds based on the 95% quantile of the historical gradient difference distribution; Using a Gaussian process as a surrogate model, selecting a radial basis function as the kernel function, and expected improvement as the acquisition function, and using the iterative termination condition and the objective function to perform Bayesian optimization to obtain the data drift threshold.
9. The real-time health risk prediction method based on a dynamic knowledge graph according to claim 1, wherein The performance of lightweight deployment includes: Using embedding similarity, calculate the confidence of triples, and prune redundant edges with a confidence less than the preset confidence to obtain a pruned global model; Compress the pruned global model using 8-bit symmetric quantization, dynamically adjust the scaling factor s through the calibration set, and minimize the KL divergence error, as shown in the formula: Among them, L quant is the KL divergence error, and P PFP32 (x) and P INT8 (x) are the floating-point and quantized distributions respectively.
10. A real-time health risk prediction system based on a dynamic knowledge graph, characterized in that, The system includes the following modules: A multi-source data acquisition module, configured to: preprocess the heterogeneous medical and elderly care data to obtain preprocessed data; A federated knowledge graph construction module, configured to: obtain a global dynamic knowledge graph using a hierarchical federated learning framework based on the preprocessed data; A dynamic update module, configured to: update the local knowledge graph embedding models and the global model of each institution based on a noise-robust gradient difference dynamic update mechanism; An edge inference module, configured to: perform lightweight deployment based on the global dynamic knowledge graph, the local knowledge graph embedding models of each institution, and the global model to obtain a lightweight graph model, and use the lightweight graph model for health risk prediction.
Citation Information
Patent Citations
Health risk prediction method and device based on joint learning
CN112289448A
Multi-center medical diagnosis knowledge graph representation learning method and system
CN113434626A
Federal learning-based privacy protection system for realizing medical data
CN115563650A
Physical medical data fusion privacy protection method based on cloud and mist architecture longitudinal federal learning
CN116595584A
Federal learning-based model training method and device
CN116957103A
Cited By
Low-altitude multi-agent cooperative control method and system
CN120523098A
Middle number merchant credit scoring method and device, electronic equipment and storage medium
CN120563213A
Esophageal cancer neoadjuvant chemotherapy and immunization effect prediction method and system
CN120600344A
Nuclear power station risk information dual robust differential privacy defense method and device based on Hodges-Lehmann
CN120724483A
Trusted federal learning method and equipment for smart home health monitoring system
CN120745755A