A multi-modal hb a1 c fluctuation prediction method based on an adversarial representation graph fusion

By combining adversarial representation learning with hierarchical graph fusion mechanism, the problem of information interference in multimodal health data is solved, achieving stable prediction and improved accuracy under small sample conditions, which is applicable to diabetes follow-up management and blood glucose control in medical IoT.

CN122135937APending Publication Date: 2026-06-02TIANJIN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-02-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Health data from different modalities differ significantly in statistical distribution, scale, and time structure. Direct splicing or simple weighted fusion can easily cause information interference, leading to unstable predictions. In particular, overfitting or output degradation is more likely to occur in small sample scenarios.

Method used

We employ adversarial representation learning and hierarchical graph fusion mechanism. Through adversarial training, we align the embedding distributions of each modality and construct a hierarchical graph fusion network in a unified embedding space. We explicitly model the interaction relationships of unimodality, bimodality, and multimodality. By combining modality attention and similarity modulation weights, we generate multimodal fusion representations.

Benefits of technology

It reduces modal differences, improves the accuracy and stability of predictions, outputs interpretable modal weights and interaction contributions, provides a basis for personalized intervention, and is suitable for follow-up management of medical IoT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135937A_ABST
    Figure CN122135937A_ABST
Patent Text Reader

Abstract

This invention provides a multimodal HbA1c improvement prediction method based on adversarial representation learning and hierarchical graph fusion. Addressing the challenges of large modal differences in multi-source heterogeneous health data from the Internet of Things (IoT) for healthcare, the susceptibility to interference from direct splicing, and the difficulty in modeling cross-modal interactions, a two-stage prediction framework is constructed. First, through adversarial training between the encoder and modality discriminator, modalities such as personal information, lifestyle, medical testing, and multi-day nutritional intake are mapped to a unified modality-invariant embedding space. Reconstruction constraints and binary classification constraints are then combined to improve fidelity and discriminativity. Subsequently, a hierarchical graph fusion network is constructed to progressively model unimodal and bimodal interactions, highlighting complementary information through similarity modulation, and outputting the HbA1c improvement / no improvement prediction result. This method possesses advantages such as structural interpretability and strong robustness, making it suitable for diabetes follow-up management and personalized intervention decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent analysis of medical and health data and chronic disease management technology. Specifically, it relates to a multimodal HbA1c fluctuation prediction method and device based on adversarial representation learning and hierarchical graph fusion mechanism, which is particularly suitable for diabetes follow-up management, blood glucose control assessment and personalized intervention decision support in medical Internet of Things scenarios. Background Technology

[0002] HbA1c reflects an individual's average blood glucose level over the past 2-3 months and is an important indicator for assessing diabetes treatment efficacy and risk stratification. With the development of wearable devices, mobile follow-up, and health management platforms, multi-source health data, including personal information, lifestyle data, medical tests, and dynamic nutritional intake, can be continuously collected. However, different modalities of data exhibit significant differences in statistical distribution, scale, and temporal structure. Direct splicing or simple weighted fusion can easily lead to intermodal information interference and make it difficult to explicitly characterize the interactions between multiple factors. In small sample scenarios, overfitting or output degradation is more likely to occur, resulting in unstable predictions. Therefore, how to reduce modal differences and model multimodal interactions in a structured manner has become an important research direction for multimodal prediction tasks in healthcare. Summary of the Invention

[0003] The purpose of this invention is to propose a multimodal HbA1c fluctuation prediction method and device based on adversarial representation learning and hierarchical graph fusion mechanism. This method can reduce the differences in the distribution of different health data modes and explicitly model the interaction relationships of single-modality, dual-modality and multimodality, thereby improving the accuracy, stability and interpretability of HbA1c improvement prediction. It is suitable for medical IoT follow-up management scenarios.

[0004] Step S1 (Data Construction and Labeling): Construct a health monitoring dataset containing multi-source inputs from the Internet of Things for Medical Care, and collect personal information modalities. Lifestyle modalities Medical testing modalities and multi-day nutrient intake dynamic mode Static modalities are uniformly encoded as feature vectors, while dynamic nutrient intake is aligned daily to form a time series matrix of length T. HbA1c is calculated based on initial and follow-up data. And based on the improvement threshold Generate binary classification labels y: when The time mark is for improvement Otherwise, it will be marked as not improved. This forms a training and validation (or testing) sample set.

[0005] Step S2 (Multimodal Encoding and Unified Representation): Construct encoders for each modality, mapping static modality vectors to fixed-length embedding representations; for dynamic nutrient intake sequences... A sequence encoder is used to extract fixed-length nutrient representation vectors. And it is mapped to an embedding space of a unified dimension in a consistent manner with other modalities, thus obtaining the embeddings of each modality. ,in .

[0006] Step S3 (Adversarial Representation Learning and Modality Alignment): Construct an adversarial representation learning module, introduce a modality discriminator into the unified embedding space, perform modality category discrimination on each modality embedding, and enable the encoder to learn modality-invariant embedding representations through adversarial training; at the same time, introduce reconstruction constraints to maintain information fidelity and introduce task classification constraints to enhance embedding discriminability, thereby reducing the adverse effects of modality distribution differences on improving HbA1c prediction.

[0007] Step S4 (Hierarchical Graph Fusion and Cross-Modal Interaction Modeling): Construct a hierarchical graph fusion network in a unified embedding space, using single-modal embedding... As a single-modal node representation; for any two modes Constructing a bimodal interactive representation Furthermore, the interaction intensity is modulated by modal similarity to highlight complementary information and suppress redundant information; the single-modal representation and the bimodal interaction representation are aggregated to obtain the multimodal fusion representation u.

[0008] Step S5 (Improving Prediction and Training / Inference): Input the multimodal fusion representation u into the classification prediction head, and output the improved probability. And obtain the predicted label The model is trained end-to-end using a joint loss function, which includes adversarial loss, reconstruction loss, and binary classification loss (BCE). After training, the model is used to improve HbA1c state prediction for new samples.

[0009] Furthermore, the adversarial representation learning module employs adversarial training, reconstruction constraints, and binary classification constraints to learn the modality-invariant embedding space, and combines hierarchical graph fusion to achieve cross-modal interaction. Its computation and loss function are defined as follows: (1) (2) (3) (4) (5) (6) (7) (8) (9) (10) (11) (12) (13) (14) Furthermore, the optional implementation of the method can be configured according to the data source and deployment environment of the actual medical IoT platform.

[0010] The present invention has the following beneficial effects: 1. Reduce modal differences and improve fusion robustness: Align different modal embedding distributions through adversarial representation learning mechanism, so that multi-source health data can be fused in a unified embedding space, reducing the adverse impact of inter-modal distribution differences on prediction performance.

[0011] 2. Explicitly modeling multi-factor interactions and improving prediction accuracy: Through a hierarchical graph fusion structure, progressively modeling single-modal, dual-modal, and multi-modal interaction relationships, effectively uncovering the synergistic effects of factors such as lifestyle, nutritional intake, and medical testing on HbA1c changes.

[0012] 3. Output interpretable weight and contribution information: Through modal attention and similarity modulated weighting mechanism, output modal weights and interaction node contributions to provide interpretable evidence for clinical analysis and personalized intervention.

[0013] 4. Applicable to medical IoT follow-up management: This invention supports unified modeling of static and dynamic modalities, and can maintain stable prediction under conditions of small sample size and multi-source heterogeneous data. It is suitable for applications such as health assessment, risk warning and intervention recommendation for diabetic patients. Attached Figure Description

[0014] Figure 1 This is a flowchart of the HbA1c fluctuation prediction method based on adversarial representation learning and hierarchical graph fusion according to the present invention.

[0015] Figure 2 This is a functional block diagram of the multimodal HbA1c fluctuation prediction device of the present invention. Detailed Implementation

[0016] like Figure 1 and Figure 2 As shown, a multimodal HbA1c fluctuation prediction method based on adversarial representation graph fusion includes the following steps: Step S1 (Data Construction and Labeling): Construct a health monitoring dataset containing multi-source inputs from the Internet of Things for Medical Care, and collect personal information modalities. Lifestyle modalities Medical testing modalities and multi-day nutrient intake dynamic mode Static modalities are uniformly encoded as feature vectors, while dynamic nutrient intake is aligned daily to form a time series matrix of length T. HbA1c is calculated based on initial and follow-up data. And based on the improvement threshold Generate binary classification labels y: when The time mark is for improvement Otherwise, it will be marked as not improved. This forms a training and validation (or testing) sample set.

[0017] Step S2 (Multimodal Encoding and Unified Representation): Construct encoders for each modality, mapping static modality vectors to fixed-length embedding representations; for dynamic nutrient intake sequences... A sequence encoder is used to extract fixed-length nutrient representation vectors. And it is mapped to an embedding space of a unified dimension in a consistent manner with other modalities, thus obtaining the embeddings of each modality. ,in .

[0018] Step S3 (Adversarial Representation Learning and Modality Alignment): Construct an adversarial representation learning module, introduce a modality discriminator into the unified embedding space, perform modality category discrimination on each modality embedding, and enable the encoder to learn modality-invariant embedding representations through adversarial training; at the same time, introduce reconstruction constraints to maintain information fidelity and introduce task classification constraints to enhance embedding discriminability, thereby reducing the adverse effects of modality distribution differences on improving HbA1c prediction.

[0019] Step S4 (Hierarchical Graph Fusion and Cross-Modal Interaction Modeling): Construct a hierarchical graph fusion network in a unified embedding space, using single-modal embedding... As a single-modal node representation; for any two modes Constructing a bimodal interactive representation Furthermore, the interaction intensity is modulated by modal similarity to highlight complementary information and suppress redundant information; the single-modal representation and the bimodal interaction representation are aggregated to obtain the multimodal fusion representation u.

[0020] Step S5 (Improving Prediction and Training / Inference): Input the multimodal fusion representation u into the classification prediction head, and output the improved probability. And obtain the predicted label The model is trained end-to-end using a joint loss function, which includes adversarial loss, reconstruction loss, and binary classification loss (BCE). After training, the model is used to improve HbA1c state prediction for new samples.

[0021] The original input data includes at least the following modalities: personal information modality, lifestyle modality, medical testing modality, and dynamic nutrition intake time series modality. Each static modality is cleaned, encoded, and standardized before being input into its corresponding encoder; the dynamic nutrition intake sequence is sequence encoded to obtain a summary representation before being input into its corresponding encoder. Adversarial representation learning aligns the distribution of source modalities to the target modality through a discriminator, and constrains the embedding space through reconstruction loss and classification loss. The hierarchical graph fusion network uses modality attention weights and similarity modulation weights to generate unimodal, bimodal, and multimodal interactive representations, and outputs the final fused representation for classification prediction.

[0022] The training data can be used for model training and evaluation using stratified partitioning or cross-validation. Model performance can be compared and verified using metrics such as AUC, F1, accuracy, and recall. During training, an alternating update strategy is used to update the encoder / decoder, discriminator, and classifier parameters to ensure stable convergence of adversarial training.

[0023] Specifically, the model first obtains unified embedding representations for each modality in the adversarial representation learning module; then, in the graph fusion modeling module, it generates interaction representations according to a hierarchical structure and outputs a fusion representation for predicting HbA1c improvement status. By outputting modality weights and interaction contributions, combinations of health factors that contribute significantly to the prediction results can be identified, providing decision-making references for subsequent personalized interventions and follow-up management.

[0024] During the testing phase, the multimodal inputs of new subjects, after undergoing the same preprocessing, are fed into the trained model, which outputs predictions of HbA1c improvement / no improvement. When the output probability exceeds a preset threshold, a health alert or intervention suggestion can be triggered; simultaneously, the output modal weights and interaction contributions are used to interpret the basis of the predictions.

[0025] Correspondingly, the multimodal HbA1c fluctuation prediction device based on adversarial representation graph fusion includes a processor and a memory. The memory stores program instructions, and the processor calls the program instructions in the memory to implement the following modules: Data preprocessing module: Cleans, encodes, standardizes, and sequences personal information, lifestyle, medical testing, and dynamic nutritional intake data to form model input; Adversarial representation learning module: Based on encoder-discriminator adversarial training and combined with reconstruction constraints and classification constraints, a unified modality-invariant embedding representation is learned; Graph Fusion Modeling Module: Constructs a hierarchical graph fusion network, progressively generates unimodal, bimodal, and multimodal interactive representations, and outputs the fused representation; Classification prediction module: Outputs prediction results of HbA1c improvement / no improvement and their confidence levels based on the fusion representation; Interpretable output module: Outputs the modal weights, interaction node contributions, or key feature contributions to assist medical analysis and intervention decisions; Optimization module: Executes the joint loss function to optimize model parameters and completes model training and inference.

[0026] The above description is merely a preferred embodiment of the present invention and should not be construed as limiting the scope of the present invention. All equivalent changes and modifications made in accordance with the scope of the patent application and the contents of the specification of the present invention should still fall within the scope of the patent of the present invention.

Claims

1. A multimodal HbA1c fluctuation prediction method based on adversarial representation graph fusion, comprising the following steps: Step S1 (Data Construction and Labeling): Construct a health monitoring dataset containing multi-source inputs from the Internet of Things for Medical Care, and collect personal information modalities. Lifestyle modalities Medical testing modalities and multi-day nutrient intake dynamic mode Static modalities are uniformly encoded as feature vectors, while dynamic nutrient intake is aligned daily to form a time series matrix of length T. HbA1c is calculated based on initial and follow-up data. And based on the improvement threshold Generate binary classification labels y: when The time mark is for improvement Otherwise, it will be marked as not improved. This forms a training and validation (or testing) sample set. Step S2 (Multimodal Encoding and Unified Representation): Construct encoders for each modality and map static modality vectors to fixed-length embedded representations; Dynamic nutrient intake sequence A sequence encoder is used to extract fixed-length nutrient representation vectors. And it is mapped to an embedding space of a unified dimension in a consistent manner with other modalities, thus obtaining the embeddings of each modality. ,in . Step S3 (Adversarial Representation Learning and Modality Alignment): Construct an adversarial representation learning module, introduce a modality discriminator into the unified embedding space, perform modality category discrimination on each modality embedding, and enable the encoder to learn modality-invariant embedding representations through adversarial training; at the same time, introduce reconstruction constraints to maintain information fidelity and introduce task classification constraints to enhance embedding discriminability, thereby reducing the adverse effects of modality distribution differences on improving HbA1c prediction. Step S4 (Hierarchical Graph Fusion and Cross-Modal Interaction Modeling): Construct a hierarchical graph fusion network in a unified embedding space, using single-modal embedding... As a single-modal node representation; for any two modes Constructing a bimodal interactive representation Furthermore, the interaction intensity is modulated by modal similarity to highlight complementary information and suppress redundant information; the single-modal representation and the bimodal interaction representation are aggregated to obtain the multimodal fusion representation u. Step S5 (Improving Prediction and Training / Inference): Input the multimodal fusion representation u into the classification prediction head, and output the improved probability. And obtain the predicted label The model is trained end-to-end using a joint loss function, which includes adversarial loss, reconstruction loss, and binary classification loss (BCE). After training, the model is used to improve HbA1c state prediction for new samples.

2. The method according to claim 1, wherein the multimodal data includes at least the following three modal channels: personal information modality: basic information such as age, gender, and body mass index; lifestyle modality: dietary habits, exercise frequency, sleep duration, etc.; medical testing modality: pre-study HbA1c and related test indicators; and dynamic nutritional intake modality: a continuous T-day nutritional intake vector sequence. The above modal data are uniformly projected to the same feature space during the embedding stage, and modality sources are represented by modality identifier encoding.

3. The method according to claim 2, wherein the embedding form of the adversarial representation learning module is: for the input of any modality m encoder Map it to a k-dimensional embedding , where k is the uniform embedding dimension; for dynamic nutrient intake sequences First, the summary representation is obtained through sequence encoding. Then mapped to .

4. The method according to claim 3, characterized in that: Adversarial training includes adversarial losses between the discriminator D and each modal encoder. The discriminator D is used to distinguish the modal sources of different modal embeddings. Each modal encoder makes it difficult for the discriminator D to distinguish the embedding sources through adversarial training, thereby achieving alignment of different modal embeddings distributed in a unified embedding space.

5. The method according to claim 4, characterized in that: The adversarial representation learning module introduces modality reconstruction constraints and task classification constraints, where the reconstruction constraints are implemented through the decoder. make Approximate reconstruction input To minimize the reconstruction loss; the task classification constraint uses classifier C to predict the label of the embedding representation to minimize the classification loss, thereby ensuring the discriminability of the embedding space.

6. The method according to claim 5, characterized in that: The overall training loss function consists of a weighted sum of adversarial loss, reconstruction loss, and classification loss, and the parameters of the encoder / decoder, discriminator, and classifier are updated in an alternating manner.

7. The method according to claim 6, characterized in that: Hierarchical graph fusion networks include single-modal layers, bimodal layers, and multimodal layers; the single-modal layer outputs the modality weights through a weight calculation mechanism. The dual-modal layer generates interaction node representations for any two modalities through fusion mapping. The multimodal layer further integrates the single-modal representation and the bimodal interactive representation to obtain the final multimodal fused representation u.

8. The method according to claim 7, characterized in that: The importance of bimodal interaction nodes is jointly determined by modal weights and modal similarity. Modal similarity is calculated using a normalized inner product, and the weights are modulated by similarity to suppress redundant interactions between highly similar modalities and highlight complementary interactions.

9. A multimodal HbA1c fluctuation prediction method according to any one of claims 1 to 8, characterized in that: The method is applicable to medical IoT follow-up management scenarios and is used for long-term blood glucose control assessment and intervention recommendations for patients with type 2 diabetes or impaired glucose tolerance.

10. A multimodal HbA1c fluctuation prediction device, characterized in that: It includes a processor and a memory, the memory storing program instructions, and the processor executing the program instructions to implement the following functional modules: Data preprocessing module: Cleans, encodes, standardizes, and sequences personal information, lifestyle, medical testing, and dynamic nutritional intake data to form model input; Adversarial representation learning module: Based on encoder-discriminator adversarial training and combined with reconstruction constraints and classification constraints, a unified modality-invariant embedding representation is learned; Graph Fusion Modeling Module: Constructs a hierarchical graph fusion network, progressively generates unimodal, bimodal, and multimodal interactive representations, and outputs the fused representation; Classification prediction module: Outputs prediction results of HbA1c improvement / no improvement and their confidence levels based on the fusion representation; Interpretable output module: Outputs the modal weights, interaction node contributions, or key feature contributions to assist medical analysis and intervention decisions; Optimization module: Executes the joint loss function to optimize model parameters and completes model training and inference.