Metabolic syndrome risk prediction method and system based on multi-modal data

By constructing a three-branch parallel network model and a feature alignment layer, and combining an attention mechanism for feature extraction and weighted fusion of multimodal data, the problem of strong heterogeneity of multimodal data is solved, and high-precision prediction and dynamic assessment of metabolic syndrome risk are achieved.

CN121528533APending Publication Date: 2026-02-13SHANGHAI TONGJI HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511706106.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing methods for predicting the risk of metabolic syndrome based on multimodal data ignore the diagnostic value of imaging data and the long-term cumulative impact of time-series lifestyle data. They cannot effectively handle the heterogeneity of multimodal data, resulting in a one-sided approach that emphasizes indicators while neglecting etiology, and their multimodal data processing capabilities are insufficient.

Method used

A three-branch parallel network model is constructed using a deep learning framework. Multimodal data features are extracted through feature alignment layers and modality adaptation optimization terms. Feature weighting and fusion are performed by combining an attention mechanism. Finally, a multimodal risk prediction neural network model is constructed to achieve cross-modal data matching and risk prediction.

Benefits of technology

It achieves high-precision, multi-dimensional dynamic prediction of metabolic syndrome risk, improves the stability and generalization ability of feature extraction, enhances the robustness of the model in complex health data environments, and supports real-time quantitative output and level determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528533A_ABST
    Figure CN121528533A_ABST
Patent Text Reader

Abstract

The invention discloses a metabolic syndrome risk prediction method and system based on multi-modal data, and relates to the technical field of risk prediction. The method comprises the following steps: S1, data and a data set are classified and labeled; s2, preprocessing the data, extracting a feature parameter set of the data, and performing cross-modal data matching; s3, performing weighted fusion on cross-modal data matching results, and forming a metabolic syndrome risk characteristic parameter set; s4, constructing a multi-modal risk prediction model to obtain a metabolic syndrome risk prediction result; and S5, verifying the prediction result, and generating a metabolic syndrome risk prediction report. According to the method, coupling modeling of metabolic syndrome multi-modal data fusion analysis and a risk prediction mechanism is introduced, so that the problems of high isomerism and high fusion difficulty of multi-modal data are solved, and the stability and generalization ability of feature extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk prediction technology, and more particularly to a method and system for predicting the risk of metabolic syndrome based on multimodal data. Background Technology

[0002] Metabolic syndrome, a chronic metabolic disease characterized by insulin resistance, abdominal obesity, dyslipidemia, and elevated blood pressure, has become a major challenge in global public health. It is not only a leading cause of cardiovascular disease but also significantly increases the risk of complications such as stroke and kidney disease, broadly impacting diverse scenarios including community health management, clinical diagnosis and treatment, and public health prevention and control. For example, community health service centers need to identify high-risk groups early through risk prediction to prevent the disease from progressing to an irreversible stage.

[0003] Accurate prediction of metabolic syndrome risk is not only a key means to meet patients' needs for early detection and intervention, helping high-risk groups to adjust their lifestyles in a timely manner and reduce the probability of disease occurrence, but also provides clinical decision support for medical institutions. It is also an important tool for optimizing the public health service system. Through group risk stratification, it can guide the allocation of prevention and control resources to key populations, reduce the waste of medical resources caused by indiscriminate screening, and provide data basis for the research and development of metabolic syndrome-related drugs and the design of health management products, avoiding the failure of intervention measures or over-treatment due to inaccurate risk assessment.

[0004] However, existing methods and systems for predicting metabolic syndrome risk based on multimodal data rely heavily on structured clinical indicators such as blood glucose, blood lipids, and blood pressure in practical applications, neglecting the diagnostic value of imaging data and the long-term cumulative impact of time-series lifestyle data on metabolic risk. They fail to cover the dual pathogenic logic of metabolic syndrome's pathological mechanisms and environmental triggers, resulting in a one-sided approach to risk prediction methods that emphasize indicators while neglecting etiology. Furthermore, their multimodal data processing capabilities are insufficient. If a small amount of non-clinical data is introduced, it is often handled through simple splicing or static weight fusion, which fails to address the heterogeneity of imaging, clinical, and time-series data, leading to ineffective alignment of data semantics. Currently, no effective solutions have been proposed to address these issues in the relevant technologies. Summary of the Invention

[0005] In response to the problems in related technologies, this invention proposes a method and system for predicting the risk of metabolic syndrome based on multimodal data, so as to overcome the above-mentioned technical problems existing in the existing related technologies.

[0006] To achieve the above objectives, the specific technical solution adopted by the present invention is as follows: According to one aspect of the present invention, a method for predicting the risk of metabolic syndrome based on multimodal data is provided, comprising the following steps: S1. Obtain multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and classify and label the multimodal data of patients with metabolic syndrome. S2. Preprocess the multimodal data of metabolic syndrome patients after classification and labeling, extract the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on the deep learning framework, and perform cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. As a preferred embodiment, the preprocessing of the classified and labeled multimodal data of metabolic syndrome patients, the extraction of syndrome feature parameter sets from the multimodal data of metabolic syndrome patients based on a deep learning framework, and the cross-modal data matching of the syndrome feature parameter sets with historical metabolic syndrome datasets include the following steps: S21. Perform targeted preprocessing of multimodal data based on the classification and labeling of multimodal data from patients with metabolic syndrome; S22. Construct a multimodal feature extraction network model based on a deep learning framework, and use the multimodal feature extraction network model to extract standardized image datasets, structured clinical datasets, and time-series life datasets of patients with metabolic syndrome. As a preferred embodiment, the step of constructing a multimodal feature extraction network model based on a deep learning framework and using this model to extract standardized image datasets, structured clinical datasets, and time-series lifestyle datasets of patients with metabolic syndrome includes the following steps: S221. A three-branch parallel network and feature alignment layer architecture are set up using a deep learning framework, and a modality adaptation optimization term is introduced for the three-branch parallel network to build a multimodal feature extraction network model. As a preferred embodiment, the method of constructing a multimodal feature extraction network model by setting up a three-branch parallel network and a feature alignment layer architecture using a deep learning framework, and introducing a modality adaptation optimization term for the three-branch parallel network, includes the following steps: S2211. A three-branch parallel network is built based on a deep learning framework, which includes an improved residual neural network architecture for the imaging branch, a deep multilayer perceptron architecture for the clinical branch, and a Transformer architecture for the temporal branch. S2212. Set a feature alignment layer at the output of the three-branch parallel network, and map the output features of each branch to the feature space of the same dimension through the projection matrix, and calculate the intramodal feature similarity matrix. S2213. Preset modal adaptation optimization strategy and bias to construct adaptation loss function, combine the intra-modal feature similarity matrix with modal adaptation optimization strategy and bias to construct adaptation loss function, and form the total loss function of multimodal feature extraction network; S2214. The parameters of the three-branch parallel network and the feature alignment layer are iteratively updated according to the total loss function to form a multimodal feature extraction network model.

[0007] S222. Input the multimodal data of patients with metabolic syndrome into a multimodal feature extraction network model for feature classification and extraction to obtain standardized image datasets, structured clinical datasets, and time-series life datasets. S223. Validate and output the standardized image dataset, structured clinical dataset, and time-series life dataset.

[0008] S23. Integrate the standardized image dataset, structured clinical dataset, and time-series lifestyle dataset to obtain a set of characteristic parameters for the syndrome; S24. The cosine similarity algorithm is used to calculate the sample matching degree between the syndrome feature parameter set and the historical metabolic syndrome dataset, and a cross-modal association mapping table is established based on the sample matching degree to perform cross-modal data matching.

[0009] S3. A pre-defined medical database for metabolic syndrome and multimodal feature weighting rules are used to weight and fuse cross-modal data matching results based on the multimodal feature weighting rules, and a set of metabolic syndrome risk feature parameters is formed by combining the medical database for metabolic syndrome. As a preferred embodiment, the preset metabolic syndrome medical database and multimodal feature weight allocation rules, which weight and fuse cross-modal data matching results based on the multimodal feature weight allocation rules, and combine them with the metabolic syndrome medical database to form a metabolic syndrome risk feature parameter set, include the following steps: S31. Define clinical diagnostic criteria, imaging feature thresholds, and life risk factor benchmarks to construct a medical database for metabolic syndrome, and set multimodal feature weight allocation rules based on attention mechanisms; S32. Extract the data modal feature parameters of cross-modal data matching, and perform weighted calculation on the data modal feature parameters based on the multimodal feature weight allocation rule to obtain the multimodal weight parameter set; As a preferred embodiment, the step of extracting data modal feature parameters for cross-modal data matching and weighting the data modal feature parameters based on multimodal feature weight allocation rules to obtain a multimodal weight parameter set includes the following steps: S321. Extract the separated image modality feature parameters, clinical modality feature parameters, and time-series lifestyle modality feature parameters from the cross-modal data matching results; S322. Calculate the set of single-modal feature weight coefficients for separating image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters based on the multimodal feature weight allocation rule; As a preferred embodiment, the calculation of the single-modal feature weight coefficient set based on the multimodal feature weight allocation rule for separating image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters includes the following steps: S3221. Extract image feature thresholds, clinical diagnostic criteria, and life risk factor benchmarks from the metabolic syndrome medical database as weight calculation constraints. S3222. Construct a single-modal feature weight calculation model based on the attention mechanism. Input the separated image modality feature parameters, clinical modality feature parameters, and time series life modality feature parameters into the single-modal feature weight calculation model to obtain the separated image weight parameters, clinical modality weight parameters, and time series life weight parameters. S3223. Based on the weight calculation constraints, the weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are corrected, and the corrected weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are integrated to obtain a set of single-modality feature weight coefficients.

[0010] S323. Perform dimensional consistency verification on the single-modal feature weight coefficients, remove abnormal weight parameters, and integrate them to form a multimodal weight parameter set.

[0011] S33. Combine the multimodal weight parameter set with the metabolic syndrome medical database for data correction, and align the feature dimensions of the corrected multimodal weight parameter set to form a metabolic syndrome risk feature parameter set.

[0012] S4. Construct a multimodal risk prediction model based on the attention mechanism, input the set of metabolic syndrome risk feature parameters into the multimodal risk prediction model for training, and obtain the metabolic syndrome risk prediction results. As a preferred embodiment, the step of constructing a multimodal risk prediction model based on an attention mechanism, and training the model by inputting a set of metabolic syndrome risk feature parameters to obtain the metabolic syndrome risk prediction result, includes the following steps: S41. Construct a multimodal risk prediction neural network model based on the attention mechanism, and introduce a modal difference penalty term into the multimodal risk prediction neural network model; As a preferred embodiment, the construction of a multimodal risk prediction neural network model based on an attention mechanism, and the introduction of a modality difference penalty term into the multimodal risk prediction neural network model, includes the following steps: S411. Constructing a multimodal feature input layer for a multimodal risk prediction neural network model based on an attention mechanism; S412. Set up a cross-modal attention interaction layer for the multimodal risk prediction neural network model, and generate an intermodal attention matrix through an attention mechanism; S413. Preset the modal difference penalty function, construct the risk prediction output layer to integrate the inter-modal attention matrix and the modal difference penalty function, and form a multimodal risk prediction neural network model.

[0013] S42. Divide the set of risk characteristic parameters for metabolic syndrome into a training set and a test set, and iteratively train the multimodal risk prediction neural network model. S43. Input the set of risk feature parameters of metabolic syndrome into the trained multimodal risk prediction neural network model to obtain three types of risk probability values, and filter the three types of risk probability values. Use the filtered three types of risk probability values ​​as the risk prediction results of metabolic syndrome.

[0014] S5. Validate the risk prediction results of metabolic syndrome and generate a metabolic syndrome risk prediction report based on the validated risk prediction results.

[0015] According to another aspect of the present invention, a metabolic syndrome risk prediction system based on multimodal data is provided, the system comprising: a parameter acquisition and annotation module, a parameter processing and matching module, a feature weighted fusion module, a model generation and training module, and a validation report generation module; The parameter acquisition and annotation module is used to acquire multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and to classify and annotate the multimodal data of patients with metabolic syndrome. The parameter processing and matching module is used to preprocess the multimodal data of metabolic syndrome patients after classification and labeling. It extracts the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on a deep learning framework, and performs cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. The feature weighted fusion module is used to preset the metabolic syndrome medical database and multimodal feature weight allocation rules, perform weighted fusion on the cross-modal data matching results based on the multimodal feature weight allocation rules, and combine the metabolic syndrome medical database to form a set of metabolic syndrome risk feature parameters; The model generation and training module is used to build a multimodal risk prediction model based on the attention mechanism. The risk feature parameter set of metabolic syndrome is input into the multimodal risk prediction model for training, and the risk prediction result of metabolic syndrome is obtained. The validation report generation module is used to validate the metabolic syndrome risk prediction results and generate a metabolic syndrome risk prediction report based on the validated metabolic syndrome risk prediction results.

[0016] The beneficial effects of this invention are as follows: 1. This invention introduces coupled modeling of multimodal data fusion analysis and risk prediction mechanisms for metabolic syndrome, and constructs an intelligent assessment system covering the entire process from data acquisition, feature extraction, cross-modal matching to risk prediction. This achieves high-precision, multi-dimensional dynamic prediction of metabolic syndrome risk. Furthermore, by integrating standardized imaging, structured clinical information, and time-series lifestyle data, a three-branch parallel network architecture is constructed and a modality adaptation mechanism is introduced to solve the problems of strong heterogeneity and high fusion difficulty of multimodal data, and improve the stability and generalization ability of feature extraction.

[0017] 2. This invention introduces a feature alignment layer and modality adaptation optimization strategy, and combines an attention mechanism to perform weighted fusion of features from different modalities. It dynamically calculates the feature weight coefficient set to ensure that the influence of each modality in the final risk assessment is dynamically allocated based on actual performance. This avoids the limitations of traditional methods where feature weights are manually set and evaluation standards are singular, further enhancing the model's ability to adapt to different populations and life scenarios. Furthermore, by setting modality difference penalty terms and an attention interaction mechanism, it enhances the ability to suppress potential conflicts or redundant information between modalities, improving robustness in complex and ever-changing health data environments. At the same time, by constructing an iteratively trainable risk prediction neural network, it achieves real-time quantitative output and level determination of metabolic syndrome risk, and automatically generates prediction reports, contributing to the automation and intelligence of the medical decision-making process. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart of a method for predicting the risk of metabolic syndrome based on multimodal data according to an embodiment of the present invention; Figure 2 This is a system block diagram of a metabolic syndrome risk prediction system based on multimodal data according to an embodiment of the present invention.

[0020] In the picture: 1. Parameter acquisition and annotation module; 2. Parameter processing and matching module; 3. Feature weighted fusion module; 4. Model generation and training module; 5. Validation report generation module. Detailed Implementation

[0021] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0022] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0023] According to embodiments of the present invention, a method and system for predicting the risk of metabolic syndrome based on multimodal data are provided.

[0024] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. According to one embodiment of the present invention, such as... Figure 1 As shown, the metabolic syndrome risk prediction method based on multimodal data according to an embodiment of the present invention includes the following steps: S1. Obtain multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and classify and label the multimodal data of patients with metabolic syndrome. Specifically, in terms of multimodal data acquisition, imaging data (such as body composition analysis, abdominal CT, and ultrasound images, used to extract features such as visceral fat area and liver fat infiltration) are collected from clinical diagnosis and treatment scenarios; structured clinical data (indicators such as blood glucose, blood lipids, blood pressure, body mass index, and waist circumference are retrieved from the electronic medical record system, and patient age, gender, past medical history, smoking and drinking history are recorded simultaneously); and time-series lifestyle data (daily exercise duration and sleep quality are collected through wearable devices, and calorie intake and dietary structure information are obtained by combining the patient's dietary log). The historical metabolic syndrome labeled dataset is obtained by screening complete data of previously diagnosed patients (including diagnosis and treatment records and follow-up results) from the hospital's medical record management system, or by referencing metabolic syndrome-related data with diagnostic labels from public medical databases, and clearly labeling the three states of confirmed, suspected, and healthy and the corresponding key indicator thresholds. During classification and labeling, the labels are categorized by data type: imaging data are labeled with pathological feature categories (such as excessive visceral fat or normal liver fat), clinical data are labeled with abnormal indicator categories (such as elevated fasting and / or 2-hour postprandial blood glucose, elevated glycated hemoglobin, or normal triglycerides / high-density lipoprotein cholesterol), and time-series lifestyle data are labeled with behavioral risk categories (such as insufficient daily exercise and high-sugar diet). The labeling adopts a combination of manual review by medical staff and AI pre-labeling tools, and finally cross-validation is used to ensure the accuracy of the labeling, forming a multimodal classification and labeling dataset for model training.

[0025] S2. Preprocess the multimodal data of metabolic syndrome patients after classification and labeling, extract the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on the deep learning framework, and perform cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. In this embodiment of the application, the preprocessing of the classified and labeled multimodal data of metabolic syndrome patients, the extraction of syndrome feature parameter set from the multimodal data of metabolic syndrome patients based on a deep learning framework, and the cross-modal data matching of the syndrome feature parameter set with historical metabolic syndrome datasets include the following steps: S21. Perform targeted preprocessing of multimodal data based on the classification and labeling of multimodal data from patients with metabolic syndrome; Specifically, during image data preprocessing, the format is first standardized according to the annotations (e.g., excessive visceral fat or normal liver fat), and then the pathological regions corresponding to the annotations are cropped in a targeted manner (e.g., if the annotation is related to visceral fat, the core abdominal region is cropped). Then, grayscale normalization is performed to eliminate device differences, and data enhancement such as rotation and flipping is used to preserve the pathological features corresponding to the annotations, while removing artifacts irrelevant to the annotations (e.g., metallic foreign bodies in CT scans). For structured clinical data preprocessing, referring to the annotations (e.g., elevated fasting blood glucose or normal triglycerides), multiple imputation is used to complete the missing values ​​of the annotation-related indicators (e.g., if the blood glucose annotation is missing, blood glucose data is completed first). Outliers are judged based on the annotations (indicators with elevated annotations exceeding the threshold are reasonable data and do not need to be removed), and then the dimensions are eliminated. One-hot encoding is performed on the variables corresponding to the classification annotations such as smoking history.

[0026] Preprocessing of time-series lifestyle data requires combining annotations (such as insufficient daily exercise and high-sugar diet) with a unified time granularity (such as aggregation by day), filling in missing data for the relevant time periods of the annotations (such as filling in exercise records for the corresponding dates if exercise annotations are missing), smoothing outliers that are not related to the annotations (such as occasional excessive exercise data) with moving averages, and finally extracting the time-series features corresponding to the annotations (such as calculating the average weekly exercise duration corresponding to the insufficient exercise annotations).

[0027] S22. Construct a multimodal feature extraction network model based on a deep learning framework, and use the multimodal feature extraction network model to extract standardized image datasets, structured clinical datasets, and time-series life datasets of patients with metabolic syndrome. In this embodiment of the application, the step of constructing a multimodal feature extraction network model based on a deep learning framework and using the multimodal feature extraction network model to extract standardized image datasets, structured clinical datasets, and time-series lifestyle datasets of patients with metabolic syndrome includes the following steps: S221. A three-branch parallel network and feature alignment layer architecture are set up using a deep learning framework, and a modality adaptation optimization term is introduced for the three-branch parallel network to build a multimodal feature extraction network model. In this embodiment of the application, the step of using a deep learning framework to set up a three-branch parallel network and a feature alignment layer architecture, and introducing a modality adaptation optimization term to the three-branch parallel network to construct a multimodal feature extraction network model includes the following steps: S2211. A three-branch parallel network is built based on a deep learning framework, which includes an improved residual neural network architecture for the imaging branch, a deep multilayer perceptron architecture for the clinical branch, and a Transformer architecture for the temporal branch. Specifically, a network-shared input layer is built to receive preprocessed multimodal data, and then three dedicated branches for imaging, clinical, and time series are constructed. The parameters of each branch are initialized independently, but the output feature dimensions are unified (e.g., all are set to 256 dimensions).

[0028] The imaging branch employs an improved residual neural network architecture, based on the ResNet50 model. The size of the first-layer convolutional kernel is adjusted from 7×7 to 3×3 to reduce parameter redundancy. Dilated convolutions are inserted in the 3rd and 4th residual blocks to enhance the ability to capture subtle pathological features such as visceral fat and liver fat in abdominal CT / ultrasound images. At the same time, a squeeze-excited attention module is added after each residual block to enhance the extraction of pathologically relevant features (such as grayscale features of fat regions) through channel weight redistribution. Finally, the image modal feature vector is output through a global average pooling layer.

[0029] The clinical branch adopts a deep multilayer perceptron architecture. The input layer receives standardized structured clinical data (such as 12-15 dimensional indicators such as blood glucose, blood lipids, and blood pressure), and sets 3-5 hidden layers (with the number of neurons being 128, 64, and 32 respectively). Batch normalization layers are inserted between the hidden layers to suppress overfitting. To adapt to the correlation of clinical indicators, residual connections (connections across 1 layer) are added in the penultimate layer to alleviate the gradient vanishing problem. The clinical modality feature vector is output through a fully connected layer.

[0030] The temporal branch adopts the Transformer architecture. The input layer transforms temporal lifestyle data (such as daily exercise, diet, sleep, etc.) into a 512-dimensional temporal vector through the embedding layer, and adds sinusoidal positional encoding to preserve temporal order information. The encoder part sets up 2-3 encoder blocks, each containing an 8-head self-attention mechanism (capturing the correlation between lifestyle data at different time periods, such as the lag correlation between diet and blood glucose the next day) and a 1-layer feedforward network (1024 hidden layer dimensions). The training is stabilized by layer normalization, and then the temporal modality feature vector is output by mean pooling.

[0031] S2212. Set a feature alignment layer at the output of the three-branch parallel network, and map the output features of each branch to the feature space of the same dimension through the projection matrix, and calculate the intramodal feature similarity matrix. Specifically, the feature alignment layer receives initial feature vectors from the imaging branch (improved residual neural network), the clinical branch (deep multilayer perceptron), and the temporal branch (each branch has a pre-defined output dimension of 256 to avoid differences in basic dimensions). Each branch is configured with an independent learnable projection matrix (all matrices are set to 256×256). Matrix multiplication maps the initial features of each branch to the same standardized feature space. For example, the initial features of the imaging branch focus on pathological structural information, and the projection matrix adjusts the weights through training, transforming them into vectors that are semantically compatible with clinical and temporal features. The initial features of the clinical branch focus on numerical correlations, and the projection matrix strengthens the feature dimensions directly related to metabolic risk. The initial features of the temporal branch contain time-dependent information, and the projection matrix optimizes the mapping relationship between time features and risk correlation. Ultimately, the mapped features of all three branches are N×256 dimensions (N being the number of samples), ensuring spatial consistency.

[0032] Subsequently, the intramodal feature similarity matrix is ​​calculated: for the features mapped to each branch, the cosine similarity algorithm is used to construct N×N similarity matrices for three modalities: imaging, clinical, and time series (each element in the matrix represents the degree of feature similarity between two corresponding samples in that modality). For example, the imaging modality similarity matrix can measure the similarity of pathological features such as visceral fat and liver fat among different patients; the clinical modality matrix focuses on the correlation of indicators such as blood glucose, blood lipids, and blood pressure; and the time series modality matrix reflects the consistency of lifestyle behavior patterns. During the calculation process, outlier samples (such as outliers with similarity values ​​below 0.1) need to be removed to ensure that the matrix can truly reflect the distribution pattern of sample features within the same modality. This provides data support for the subsequent introduction of modality adaptation optimization terms and the construction of the total loss function, while also verifying the mapping effect of the projection matrix to avoid intramodal feature distortion due to spatial mapping bias.

[0033] S2213. Preset modal adaptation optimization strategy and bias to construct adaptation loss function, combine the intra-modal feature similarity matrix with modal adaptation optimization strategy and bias to construct adaptation loss function, and form the total loss function of multimodal feature extraction network; Specifically, a pre-defined modality adaptation optimization strategy is developed, with the core focus on setting targets for feature consistency within modalities. For three modalities—imaging, clinical, and time-series—feature similarity thresholds are set based on the standards of the metabolic syndrome medical database (e.g., visceral fat feature similarity for the imaging modality, correlation of blood glucose, blood lipids, and blood pressure for the clinical modality, and similarity of lifestyle behavior patterns for the time-series modality). The thresholds are 0.7 for the imaging modality, 0.65 for the clinical modality, and 0.6 for the time-series modality. The deviation is defined as the difference between the element of the feature similarity matrix within each modality and the corresponding threshold, thereby quantifying the degree of deviation between the features within the modality and the target consistency.

[0034] Next, the adaptation loss function is constructed: for each modality, the sum of squared deviations between the elements of the similarity matrix and the preset threshold is calculated using mean squared error, and then the mean value is taken to obtain the single-modality adaptation loss. The adaptation loss is then weighted and fused with the feature extraction loss of the three-branch network. The total weight of the adaptation loss is set to 0.3 (the weight of each single-modality adaptation loss is 0.1), and the total weight of the feature extraction loss is set to 0.7 (the weight of each branch extraction loss is approximately 0.23).

[0035] S2214. The parameters of the three-branch parallel network and the feature alignment layer are iteratively updated according to the total loss function to form a multimodal feature extraction network model.

[0036] Specifically, parameter initialization is performed on the network layer parameters of the imaging branch (improved residual neural network), clinical branch (deep multilayer perceptron), and temporal branch. Then, the projection matrix of the feature alignment layer is initialized with random normal distribution and the standard deviation is limited to ensure that the initial parameter distribution is adapted to the data characteristics of each branch and to avoid gradient explosion or vanishing in the early stage of training.

[0037] The training components are then configured, and the Adam optimizer is selected (the learning rate is initially set to 1e-4, dynamically adjusted using a cosine annealing strategy, and gradually reduced to 1e-6 in the later stages of iteration). The total loss function is used as the optimization objective. The preprocessed classification-labeled multimodal data is divided into training and validation sets in a 7:3 ratio, with a batch size of 32 or 64 (adjusted according to hardware computing power). The training set data is input into a three-branch network. After feature extraction, alignment layer mapping, and similarity matrix calculation, the loss value for this round is obtained by substituting it into the total loss function. Then, the gradient of the total loss with respect to each parameter (branch network weights, projection matrix elements) is calculated through backpropagation. The optimizer updates the parameters according to the gradient descent principle, while introducing L2 regularization to suppress overfitting. The loss is evaluated using the validation set every 5 rounds. If the validation set loss does not decrease for 3 consecutive rounds, early stopping is triggered to avoid overtraining. The above iterative process is repeated (usually 50-100 rounds) until the training set loss stabilizes and converges. The final parameters (including branch network weights and alignment layer projection matrix) are saved, thus forming a network model that can be used for multimodal feature extraction.

[0038] S222. Input the multimodal data of patients with metabolic syndrome into a multimodal feature extraction network model for feature classification and extraction to obtain standardized image datasets, structured clinical datasets, and time-series life datasets. Specifically, preprocessed image data (such as abdominal CT and ultrasound images) are converted into tensor format (such as 3×224×224), structured clinical data (such as blood glucose, blood lipid, and blood pressure indicators) are organized into one-dimensional numerical vectors, and time-series lifestyle data (such as exercise and diet records) are converted into time-series matrices (such as time steps × feature number).

[0039] Subsequently, the three types of data are input into the corresponding branches. Imaging data is fed into the improved residual neural network branch, where pathological features such as visceral fat distribution and liver steatosis are extracted through convolution, residual blocks, and attention modules. Clinical data is input into the deep multilayer perceptron branch, where correlation features between indicators (such as the correlation between blood glucose and insulin resistance) are captured through fully connected layers and residual connections. Time-series lifestyle data are processed through a self-attention mechanism and a feedforward network to extract time-dependent features of diet and exercise (such as the pattern of insufficient weekly exercise).

[0040] After the output features of each branch are standardized by the feature alignment layer, they are integrated according to modality. The image features are integrated into a standardized image dataset (including pathological feature vectors and corresponding annotations), the clinical features form a structured clinical dataset (including indicator association features), and the time series features form a time series life dataset (including behavioral pattern features). The three types of datasets output in the end have unified dimensions and standardized formats, and can be directly used for subsequent cross-modal data matching.

[0041] S223. Validate and output the standardized image dataset, structured clinical dataset, and time-series life dataset.

[0042] Specifically, format and dimensionality checks are performed to verify whether the image dataset is a uniform tensor format (such as an N×256-dimensional feature vector), whether the clinical dataset is a standardized numerical matrix, and whether the time series dataset is a regular time series feature sequence. This ensures that the three types of datasets have matching dimensions and no missing / outlier values ​​(the missing rate must be <5%, and outliers are verified a second time using the Z-score method).

[0043] Further verification of feature validity: The image dataset was compared with the original image annotations to confirm that the extracted pathological features (such as visceral fat features) were consistent with the annotated pathological regions; the clinical dataset was verified through medical database validation to confirm that the indicator correlation features (such as the correlation between blood glucose, blood lipids and blood pressure) conformed to the clinical diagnostic logic; the time series dataset was combined with life logs to check that the behavioral pattern features (such as the lack of exercise pattern) were consistent with the records.

[0044] Simultaneously, 20% of the samples are randomly selected and input into the model to verify the stability of feature extraction (the coefficient of variation of features extracted from the same data multiple times is <10%). After successful verification, the data is output in the standard format, along with a verification report (including data quality indicators and conclusions on feature validity) to ensure that the dataset can be directly used for subsequent cross-modal data matching.

[0045] S23. Integrate the standardized image dataset, structured clinical dataset, and time-series lifestyle dataset to obtain a set of characteristic parameters for the syndrome; Specifically, based on the output of the previous feature alignment layer, it ensures that the sample feature dimensions of the three datasets are consistent (all are 256-dimensional feature vectors). If there is a slight dimensional deviation, it is fine-tuned by learning the mapping matrix so that each patient sample corresponds to a 256-dimensional feature vector in the imaging, clinical and temporal modalities, thus achieving one-to-one matching of sample feature dimensions.

[0046] Subsequently, a multimodal weight coefficient set (previously generated based on attention mechanism and medical database constraints) was introduced to assign appropriate weights to imaging, clinical, and temporal features respectively (e.g., clinical features weight 0.4, imaging 0.3, and temporal features weight 0.3, in line with clinical diagnostic logic). The three types of single-modal features were then fused into a 256-dimensional comprehensive feature vector through weighted summation, preserving the core information of each modality (e.g., the pathological diagnostic value of clinical indicators, anatomical features of images, and behavioral inducements of temporal features).

[0047] Finally, the fusion features were optimized by normalization to eliminate dimensional differences and Pearson correlation analysis to remove redundant features with a correlation of less than 0.3 with metabolic syndrome. The result was a set of syndrome feature parameters with the sample integrated feature vector labeled with the state as the core structure. Each sample corresponds to one integrated record, which can be directly used for cross-modal matching with historical datasets.

[0048] S24. The cosine similarity algorithm is used to calculate the sample matching degree between the syndrome feature parameter set and the historical metabolic syndrome dataset, and a cross-modal association mapping table is established based on the sample matching degree to perform cross-modal data matching.

[0049] Specifically, a unified data dimension ensures that the feature parameter set of the syndrome (including a 256-dimensional comprehensive feature vector) is consistent with the feature dimensions of the historical metabolic syndrome dataset, with each sample corresponding to one standardized feature vector.

[0050] Subsequently, the sample matching degree is calculated. For each sample to be matched in the syndrome feature parameter set, the cosine similarity is calculated one by one with all samples in the historical dataset. The matching degree value in the interval [-1, 1] is obtained by dividing the vector dot product by the product of the two vector magnitudes (the closer the value is to 1, the more similar the sample features). The focus is on capturing the comprehensive correlation between imaging pathological features, clinical indicators, and temporal behavioral patterns.

[0051] Finally, a cross-modal association mapping table is established, a matching degree threshold is set (e.g., ≥0.7), highly similar sample pairs are screened, and the ID of the sample to be matched, the ID of the historical sample, the matching degree, and the corresponding cross-modal feature association items (e.g., the association between imaging fat features and clinical blood lipid indicators of historical samples) are recorded in the table to form a structured mapping relationship. Through this table, the accurate association between new samples and historical samples in multimodal features is realized, providing a basis for subsequent weighted fusion and ensuring the accuracy and traceability of cross-modal data matching.

[0052] S3. A pre-defined medical database for metabolic syndrome and multimodal feature weighting rules are used to weight and fuse cross-modal data matching results based on the multimodal feature weighting rules, and a set of metabolic syndrome risk feature parameters is formed by combining the medical database for metabolic syndrome. In this embodiment of the application, the preset metabolic syndrome medical database and multimodal feature weight allocation rules, which weight and fuse cross-modal data matching results based on the multimodal feature weight allocation rules, and combine them with the metabolic syndrome medical database to form a metabolic syndrome risk feature parameter set, include the following steps: S31. Define clinical diagnostic criteria, imaging feature thresholds, and life risk factor benchmarks to construct a medical database for metabolic syndrome, and set multimodal feature weight allocation rules based on attention mechanisms; Specifically, the clinical diagnostic criteria refer to international standards, specifying thresholds for indicators such as waist circumference (≥90cm for men, ≥85cm for women), fasting blood glucose ≥6.1mmol / L, and / or 2h PG ≥7.8mmol / L, blood pressure ≥130 / 85mmHg, TG ≥1.7mmol / L, and HDL-C <1.04mmol / L, and the rule that a diagnosis is made if three or more of these indicators are abnormal. Imaging characteristic thresholds are based on clinical consensus, setting the abdominal CT visceral fat area at ≥80cm². 2 The pathological criteria for determining liver fat infiltration (CT value ≤ 40 HU) are established. Based on epidemiological studies, risk thresholds for lifestyle risk factors are determined, including daily exercise < 30 minutes, high-sugar diet (daily added sugar ≥ 25g), and sleep < 6 hours. These three criteria are stored in a structured manner to form a medical database containing standard names, indicator thresholds, and clinical significance.

[0053] When setting weight allocation rules based on the attention mechanism, an attention calculation module is first constructed, and image, clinical, and temporal life features are input. The attention score of each modality feature is calculated through a fully connected layer. Then, medical library constraints are introduced to assign higher initial attention weights to clinical indicators that meet the diagnostic criteria (such as elevated blood glucose) and image features that reach the pathological threshold (such as excessive visceral fat). Finally, the weights are dynamically adjusted through backpropagation so that the clinical modality weight (about 0.4) is higher than that of image (0.3) and temporal (0.3), ensuring that the weight allocation conforms to the medical logic of clinical diagnosis as the core and image pathology and life causes as secondary factors, thus forming a multimodal feature weight allocation rule.

[0054] S32. Extract the data modal feature parameters of cross-modal data matching, and perform weighted calculation on the data modal feature parameters based on the multimodal feature weight allocation rule to obtain the multimodal weight parameter set; In this embodiment of the application, the step of extracting data modal feature parameters for cross-modal data matching and performing weighted calculations on the data modal feature parameters based on multimodal feature weight allocation rules to obtain a multimodal weight parameter set includes the following steps: S321. Extract the separated image modality feature parameters, clinical modality feature parameters, and time-series lifestyle modality feature parameters from the cross-modal data matching results; Specifically, the effective matching pairs (the sample to be matched and the historical sample with a matching degree ≥ 0.7) in the cross-modal association mapping table are located, and the complete multimodal feature package of the two types of samples is extracted from the mapping table fields. This feature package includes the original feature vectors and associated annotations of the three modalities of image, clinical and time series.

[0055] Next, features are split according to modal attributes. When separating image modal feature parameters, 256-dimensional pathological feature vectors (such as visceral fat distribution and liver fat infiltration features) output by the improved residual neural network are extracted from the feature package, and image annotations of historical samples (such as visceral fat excess labels) are simultaneously associated. When extracting clinical modal feature parameters, indicator-related feature vectors output by the deep multilayer perceptron (such as blood glucose-lipid interaction features) are selected, and abnormal indicator items are labeled in conjunction with medical database diagnostic standards. When extracting time-series lifestyle modal feature parameters, time-dependent feature vectors (such as weekly average exercise duration and dietary structure sequence features) are obtained, and behavioral risk categories (such as high-sugar diet) are labeled according to the corresponding lifestyle risk factor benchmarks.

[0056] Finally, the consistency between the three types of parameters and the modality is verified (e.g., image features only contain pathology-related vectors, with no clinical indicators mixed in). The parameters are organized according to sample ID, modality feature vector, and associated annotation format to form independent sets of separated image, clinical, and time-series life modality feature parameters, ensuring that the parameters can be directly used for subsequent weight calculations.

[0057] S322. Calculate the set of single-modal feature weight coefficients for separating image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters based on the multimodal feature weight allocation rule; In this embodiment of the application, the calculation of the single-modal feature weight coefficient set based on the multimodal feature weight allocation rule to separate image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters includes the following steps: S3221. Extract image feature thresholds, clinical diagnostic criteria, and life risk factor benchmarks from the metabolic syndrome medical database as weight calculation constraints. Specifically, the medical database has a structured storage module, in which the clinical diagnostic standards module stores existing standards, the imaging pathology threshold module records the basis for judging imaging features, and the life risk benchmark module includes thresholds related to behavioral factors, and the target information is accurately located according to the module path.

[0058] Next, core constraints are extracted by category. Diagnostic criteria are extracted from the clinical module, such as waist circumference (≥90cm for men, ≥85cm for women), fasting blood glucose ≥6.1mmol / L and / or 2h PG ≥7.8mmol / L, blood pressure ≥130 / 85mmHg, TG ≥1.7mmol / L, HDL-C <1.04mmol / L, and other indicator thresholds, as well as the rule that a diagnosis is made if three or more of these indicators are abnormal. Pathological thresholds are extracted from the imaging module, such as visceral fat area ≥80cm² on abdominal CT. 2 The criteria for determining fatty liver include a CT value of ≤40HU for liver fat infiltration and an ultrasound fatty liver grade of ≥2. Risk factor benchmarks are extracted from the lifestyle module, such as behavioral thresholds like daily exercise <30 minutes, daily added sugar intake ≥25g, and sleep duration <6 hours.

[0059] Finally, the extracted information is standardized into a constraint table, which labels the constraint type, threshold range, and clinical significance (e.g., clinical constraint - blood pressure ≥130 / 85 mmHg - risk of hypertension). This ensures that when calculating the weights, the constraint table is used to directly determine whether the features meet the criteria (if the clinical indicators exceed the criteria, the corresponding weights are increased), thus forming an accurate basis for weight calculation constraints.

[0060] S3222. Construct a single-modal feature weight calculation model based on the attention mechanism. Input the separated image modality feature parameters, clinical modality feature parameters, and time series life modality feature parameters into the single-modal feature weight calculation model to obtain the separated image weight parameters, clinical modality weight parameters, and time series life weight parameters. Specifically, a shared basic framework is built, including an input layer (adapted to the feature dimensions of each modality), two fully connected encoding layers (with a hidden dimension of 128), and an attention computation layer. The parameters of each branch are trained independently to adapt to modal differences.

[0061] During input processing, image modality feature parameters (256-dimensional pathological feature vector), clinical modality feature parameters (256-dimensional indicator association vector), and time-series lifestyle modality feature parameters (256-dimensional behavioral pattern vector) are separated and connected to their respective branches: the image branch adds a channel attention submodule after the coding layer to enhance the weight sensitivity of key pathological features such as visceral fat area; the clinical branch embeds prior importance of indicators (e.g., the initial weight of blood glucose indicators is higher than that of other indicators); the time-series branch adds a time decay factor to increase the attention ratio of recent lifestyle behavioral features.

[0062] In the attention calculation stage, each branch generates feature attention scores through a query-key-value mechanism: the encoded feature vectors are mapped to query, key, and value matrices respectively, the correlation between features is calculated by the dot product of the query and the key, and the attention weight matrix is ​​obtained after normalization. Then, it is multiplied by the value matrix to output weighted features. Finally, a single fully connected layer outputs weight parameters with the same dimension as the input features (e.g., the imaging branch outputs 256-dimensional weight parameters, corresponding to the importance ratio of each pathological feature). The model outputs separate image weight parameters (reflecting the contribution of each image pathological feature), clinical modality weight parameters (reflecting the diagnostic weight of each clinical indicator), and time-series lifestyle weight parameters (characterizing the influence weight of each lifestyle behavior feature). All three types of parameters are normalized to ensure that the sum of the weights is 1, providing standardized input for subsequent correction based on medical database constraints.

[0063] S3223. Based on the weight calculation constraints, the weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are corrected, and the corrected weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are integrated to obtain a set of single-modality feature weight coefficients.

[0064] Specifically, targeted adjustments are made to the weight parameters of each modality. When adjusting the image weight parameters, the image feature thresholds from the medical database (such as visceral fat area ≥80cm²) are called up. 2 For high-risk characteristics, the weight parameters of corresponding pathological features (such as fat distribution characteristics) are verified. If the weight is lower than 0.1 (the minimum proportion of important features is preset), it is increased by a multiplication factor (e.g., 1.5 times). At the same time, the weight of non-threshold related features (such as irrelevant tissue imaging features) is reduced (multiplied by 0.8). The clinical weight parameter is adjusted according to the clinical diagnostic criteria. The weight of the related features of core indicators that meet the abnormal threshold (such as fasting blood glucose ≥6.1mmol / L) is strengthened (e.g., increased from 0.05 to 0.08). The weight of the related features of secondary indicators that do not meet the diagnostic criteria (such as uric acid values ​​within the normal range) is appropriately reduced. The temporal lifestyle weight parameter is adjusted according to the lifestyle risk factor benchmark. The weight of high-risk behavioral characteristics such as daily exercise <30 minutes and high sugar diet is increased with a punitive adjustment (e.g., increased by 20% when the weight is lower than 0.06). The weight of low-risk behavioral characteristics remains stable.

[0065] After correction, cross-modal consistency verification is performed to ensure that the total weight of core diagnostic indicators in the clinical modality (such as the sum of the weights of the first 5 important indicators) is not less than 60% of the total weight of the modality, the total weight of pathological threshold-related features in the imaging modality is not less than 50%, and the total weight of high-risk behavioral features in the temporal modality is not less than 40%. If the standards are not met, adjustments are made retrospectively. Finally, the corrected weight parameters of imaging (256 dimensions), clinical (256 dimensions), and temporal lifestyle (256 dimensions) are integrated according to modality type, feature index, and weight value format to form a single-modal feature weight coefficient set. This not only preserves the differentiated importance of each modality feature, but also ensures that the weight allocation conforms to medical logic through constraint correction, providing a reliable basis for subsequent multimodal weighted fusion.

[0066] S323. Perform dimensional consistency verification on the single-modal feature weight coefficients, remove abnormal weight parameters, and integrate them to form a multimodal weight parameter set.

[0067] Specifically, a dimensional consistency check is performed to verify whether the dimensions of the weight coefficients for the three single modalities of imaging, clinical, and time-series life are consistent (all are 256 dimensions). The correspondence between the feature index of each modality and the original feature parameters is checked (e.g., the 32nd dimension of imaging corresponds to the visceral fat feature). This ensures that there are no missing, misaligned, or redundant dimensions. If there are dimensional deviations, they are corrected by zero-filling or truncation.

[0068] Next, abnormal weight parameters are removed. The effective range of weight values ​​is set to [0, 1]. Outliers outside the range are screened out. Based on the constraints of the medical database, abnormal parameters with weights <0.05 for clinical core indicators (such as blood glucose) and weights <0.03 for imaging pathology threshold-related features are removed. Extreme weight values ​​in each modality are identified by the Z-score method (threshold ±3) (such as a sudden increase in the weight of a certain behavioral feature in the time-series lifestyle modality to 0.5) and replaced with the mean of the same modality.

[0069] Finally, a multimodal weight parameter set is formed: the verified data are stored in association according to the sample ID, image weight vector, clinical weight vector, and time-series life weight vector structure. Each sample corresponds to three categories of 256-dimensional weight parameters, and verification labels are added (such as consistent dimensions and no outliers) to ensure that the parameter set can be directly used for subsequent multimodal feature weighted fusion, providing a standardized weight basis for the construction of risk feature parameter set.

[0070] S33. Combine the multimodal weight parameter set with the metabolic syndrome medical database for data correction, and align the feature dimensions of the corrected multimodal weight parameter set to form a metabolic syndrome risk feature parameter set.

[0071] Specifically, data correction involves calling upon clinical diagnostic criteria (such as fasting blood glucose and blood pressure thresholds), imaging feature thresholds (such as visceral fat area standards), and time-series lifestyle risk benchmarks from the medical database. The weight parameters are verified modally. In the clinical modality, if the weight associated with core diagnostic indicators (such as blood glucose) is lower than 0.08 (the minimum weight for important indicators as determined by the medical database), it is increased to the 0.08-0.12 range through linear adjustment. In the imaging modality, parameters with weights associated with pathological thresholds (such as fat infiltration features) lower than 0.05 are increased by 15%-20%. In the time-series modality, for features with high-risk behaviors (such as lack of exercise) with weights less than 0.06, they are corrected to 0.06-0.09 based on the risk benchmark to ensure that the weight allocation aligns with medical logic.

[0072] Next, feature dimension alignment is carried out to construct a shared semantic mapping matrix, which maps the corrected image, clinical, and time-series weight parameters (all 256 dimensions) to a unified metabolic risk semantic space. For example, the 32nd dimension of the three modalities corresponds to the associated risk dimensions of visceral fat, blood glucose, and exercise. The semantic consistency of the dimensions is checked by cosine similarity (similarity ≥ 0.8 is considered as qualified alignment), and if the standard is not met, the mapping matrix is ​​fine-tuned.

[0073] Finally, the risk feature parameter set for metabolic syndrome is integrated and stored in the following structure: sample ID, corrected image weight vector, clinical weight vector, time-series weight vector and risk-related label. Each sample corresponds to one record containing 768 dimensions (256×3) of weight parameters, ensuring that the parameters can be directly input into the subsequent risk prediction model, providing medically adapted feature support for accurate prediction.

[0074] S4. Construct a multimodal risk prediction model based on the attention mechanism, input the set of metabolic syndrome risk feature parameters into the multimodal risk prediction model for training, and obtain the metabolic syndrome risk prediction results. In this embodiment of the application, the step of constructing a multimodal risk prediction model based on an attention mechanism, inputting a set of metabolic syndrome risk feature parameters into the multimodal risk prediction model for training, and obtaining the metabolic syndrome risk prediction result includes the following steps: S41. Construct a multimodal risk prediction neural network model based on the attention mechanism, and introduce a modal difference penalty term into the multimodal risk prediction neural network model; In this embodiment of the application, the construction of a multimodal risk prediction neural network model based on an attention mechanism, and the introduction of a modality difference penalty term into the multimodal risk prediction neural network model, includes the following steps: S411. Constructing a multimodal feature input layer for a multimodal risk prediction neural network model based on an attention mechanism; Specifically, a three-channel input interface is set up to receive imaging risk features (256 dimensions), clinical risk features (256 dimensions), and time-series life risk features (256 dimensions) from the set of metabolic syndrome risk feature parameters. Each channel is connected to an independent embedding layer, and the features are mapped to a 512-dimensional semantic space through linear transformation (enhancing feature expression ability). At the same time, modality label embedding (such as image label 0, clinical label 1, and time-series label 2) is added to distinguish modality attributes.

[0075] Next, a cross-modal attention sublayer is introduced: the three types of embedded features are used as query, key, and value matrices, respectively, and attention scores between different modal features are calculated (e.g., the attention of clinical features to image features). Cross-modal attention weights are obtained through normalization, and then weighted fusion is performed to generate preliminary interactive features. At the same time, the intramodal self-attention branch is retained to capture the internal correlation of similar features (e.g., the mutual influence of different clinical indicators).

[0076] Finally, the cross-modal and intra-modal attention outputs are integrated through a splicing layer, and the feature distribution is stabilized through a LayerNormalization layer to output a 1024-dimensional fused feature vector. This not only preserves the unique risk information of each modality, but also strengthens the key intermodal correlations (such as the interaction between imaging fat features and clinical blood glucose indicators) through the attention mechanism, providing input rich in multimodal correlation information for the subsequent risk prediction layer.

[0077] S412. Set up a cross-modal attention interaction layer for the multimodal risk prediction neural network model, and generate an intermodal attention matrix through an attention mechanism; Specifically, the 512-dimensional feature vectors of the three categories of images, clinical, and time series output from the input layer are mapped to query, key, and value matrices respectively through independent linear projection layers (the dimensions after projection are all 512×64, where 64 is the dimension split corresponding to the number of attention heads), ensuring that the query and key dimensions of different modalities are consistent and satisfying the cross-modal computing conditions.

[0078] Next, the interaction logic is designed using a pairwise interaction plus global integration mode. First, three sets of pairwise cross-modal interactions are implemented: image clinical, image time series, and clinical time series. Taking the correlation between image query and clinical key calculation as an example, the numerical bias caused by excessive dimensionality is weakened by scaling the dot product attention formula. Then, the image clinical modality attention weight matrix (64×64 dimensions, elements represent the correlation strength between image features and clinical features) is obtained after normalization. Similarly, the other two sets of pairwise interaction weight matrices are generated.

[0079] Finally, the three pairs of interaction matrices are integrated into a global intermodal attention matrix (192×192 in dimension, covering the interaction relationships of all modal features) through a splicing layer. High-value elements in the matrix correspond to strongly correlated feature pairs (such as visceral fat features in imaging and clinical triglyceride indicators), providing accurate intermodal correlation basis for subsequent feature weighted fusion. At the same time, residual connections are added at the end of the layer to superimpose the output of the interaction layer with the original input features, avoiding the loss of feature information.

[0080] S413. Preset the modal difference penalty function, construct the risk prediction output layer to integrate the inter-modal attention matrix and the modal difference penalty function, and form a multimodal risk prediction neural network model.

[0081] Specifically, the variance and mean of the imaging, clinical, and temporal modal features are calculated, and a penalty threshold is set (e.g., when the variance of a certain modality exceeds 1.5 times the average variance of the three modalities, a penalty is triggered). The function adopts a regularization variant based on the differences in modal features. When the differences in modal features are too large, the loss weight of the corresponding modality is increased through this function, forcing the model to balance multimodal information.

[0082] When constructing the risk prediction output layer, the weights of the intermodal attention matrix (reflecting the strength of modal association) are multiplied by the fusion features of each modality to obtain a 1024-dimensional global fusion feature; the fusion feature is then input into the output layer consisting of two fully connected layers (with 256 and 64 neurons respectively).

[0083] Simultaneously, the modality difference penalty function is embedded as a regularization term into the output layer loss function. The total loss is equal to the cross-entropy loss (risk prediction loss) plus the penalty value. Attention weights and penalty coefficients are optimized synchronously through backpropagation. Finally, the cross-modal attention interaction layer, modality difference penalty function and risk prediction output layer are integrated. After training, a multimodal risk prediction neural network model is formed, which ensures that the prediction utilizes modal correlation information while suppressing prediction bias caused by modality differences.

[0084] S42. Divide the set of risk characteristic parameters for metabolic syndrome into a training set and a test set, and iteratively train the multimodal risk prediction neural network model. Specifically, when dividing the risk feature parameter set for metabolic syndrome, it is necessary to consider both data balance and randomness. The parameter set should be divided into a training set and a test set in a 7:3 ratio. Stratified sampling should be used to ensure the proportion of positive and negative samples for metabolic syndrome in the two sets (e.g., if the positive rate in the original dataset is 30%, then the positive rate in both the training set and the test set should remain at 30%) to avoid class bias. Before dividing, the parameter set should be randomly shuffled according to the sample ID (to exclude the interference of the order of time series or image data), and the consistency of feature dimensions should be verified (all are 768-dimensional weight parameters) to ensure that the input format is compatible with the model.

[0085] When iteratively training a multimodal risk prediction neural network model, the model parameters are first initialized (He initialization is used for the weights of the attention layer and output layer), and the total loss function is the cross-entropy loss plus the modality difference penalty. The training process is as follows: in each round, the training set is input into the model in batches of 32, risk prediction values ​​are generated through forward propagation, the total loss is calculated, and the gradient is obtained through backpropagation. The optimizer updates the parameters of the cross-modal attention layer, output layer, etc. Simultaneously, overfitting is suppressed. Every 5 rounds, the model's prediction accuracy and loss value are evaluated using the test set. If the loss on the test set does not decrease for 3 consecutive rounds, early stopping is triggered. This iteration is repeated for 50-80 rounds until the training set loss converges (the difference in loss between adjacent rounds is <1e-5) and the test set performance is stable, thus completing the model training.

[0086] S43. Input the set of risk feature parameters of metabolic syndrome into the trained multimodal risk prediction neural network model to obtain three types of risk probability values, and filter the three types of risk probability values. Use the filtered three types of risk probability values ​​as the risk prediction results of metabolic syndrome.

[0087] Specifically, the parameter set is converted into a model-compatible Tensor (preserving 768 feature dimensions), and input into the multimodal risk prediction neural network model in batches of 32. Through forward propagation, three types of risk probability values ​​are output, corresponding to imaging modality-dominated risks (such as pathological feature-related risks), clinical modality-dominated risks (such as abnormal indicator-related risks), and time-series lifestyle modality-dominated risks (such as behavioral pattern-related risks). The probability values ​​range from 0 to 1 (the higher the value, the higher the risk).

[0088] The selection of the three risk probability values ​​requires a dual standard: first, setting a confidence threshold (e.g., ≥0.3 to exclude low-confidence noise results); second, referring to the diagnostic logic of the medical database. If the clinical modality risk probability is ≥0.6 (the core diagnostic criterion), it should be retained first and its weight should be strengthened. If the difference between the three probabilities exceeds 0.4 (e.g., imaging 0.8, clinical 0.3, time series 0.2), the bias should be corrected by weighted harmonics (the weights are taken from the multimodal weight parameter set).

[0089] Finally, the three types of probability values ​​that have passed the consistency verification after screening are integrated, and the weighted average (clinical weight 0.4, imaging weight 0.3, time series weight 0.3) is taken as the comprehensive risk probability. According to medical standards, it is divided into high risk (≥0.7), medium risk (0.3-0.7), and low risk (<0.3), and the three original probabilities and screening criteria are attached to form a complete metabolic syndrome risk prediction result.

[0090] S5. Validate the risk prediction results of metabolic syndrome and generate a metabolic syndrome risk prediction report based on the validated risk prediction results.

[0091] Specifically, the model's predictive performance is calculated using a reserved test set (30%). Key metrics include accuracy (≥85%), recall (≥80%, ensuring fewer missed diagnoses of high-risk patients), and AUC (≥0.88, reflecting risk discrimination ability). The accuracy of predictions for the three modalities (imaging, clinical, and time series) is compared to verify the advantages of multimodal fusion. Then, 20% of the test set samples are extracted, and the prediction results are compared with the manual diagnostic reports from clinicians to calculate the concordance rate (≥82%). The clinical diagnostic matching degree of high-risk samples (prediction probability ≥0.7) is verified to avoid model misjudgment.

[0092] The generated prediction report should present key information in a structured manner: the homepage should include the patient's basic information (ID, age, gender); the core section should display the comprehensive risk level (high / medium / low) and the original risk probabilities of the three modalities (e.g., clinical 0.72, imaging 0.68, time series 0.55), with validation evidence (e.g., model AUC=0.91, 84% concordance rate with clinical diagnosis); and the conclusion should include clinical recommendations (e.g., high-risk patients should have their blood glucose and lipid levels checked within one month, and adjust their exercise and diet), while also indicating the prediction confidence level (e.g., based on multimodal fusion prediction, confidence level 92%), ensuring that the report is both scientific and clinically instructive, facilitating doctors to develop intervention plans based on actual situations.

[0093] According to another aspect of the invention, such as Figure 2 As shown, a metabolic syndrome risk prediction system based on multimodal data is provided. The system includes: parameter acquisition and annotation module 1, parameter processing and matching module 2, feature weighted fusion module 3, model generation and training module 4, and validation report generation module 5. Parameter acquisition and annotation module 1 is used to acquire multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and to classify and annotate the multimodal data of patients with metabolic syndrome. The parameter processing and matching module 2 is used to preprocess the multimodal data of metabolic syndrome patients after classification and labeling, extract the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on the deep learning framework, and perform cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. Feature weighted fusion module 3 is used to preset the metabolic syndrome medical database and multimodal feature weight allocation rules, perform weighted fusion on cross-modal data matching results based on the multimodal feature weight allocation rules, and combine the metabolic syndrome medical database to form a set of metabolic syndrome risk feature parameters; Model generation and training module 4 is used to build a multimodal risk prediction model based on the attention mechanism. The risk feature parameter set of metabolic syndrome is input into the multimodal risk prediction model for training to obtain the risk prediction result of metabolic syndrome. The verification report generation module 5 is used to verify the metabolic syndrome risk prediction results and generate a metabolic syndrome risk prediction report based on the verified metabolic syndrome risk prediction results.

[0094] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting the risk of metabolic syndrome based on multimodal data, characterized in that, Includes the following steps: S1. Obtain multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and classify and label the multimodal data of patients with metabolic syndrome. S2. Preprocess the multimodal data of metabolic syndrome patients after classification and labeling, extract the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on the deep learning framework, and perform cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. S3. A pre-defined medical database for metabolic syndrome and multimodal feature weighting rules are used to weight and fuse cross-modal data matching results based on the multimodal feature weighting rules, and a set of metabolic syndrome risk feature parameters is formed by combining the medical database for metabolic syndrome. S4. Construct a multimodal risk prediction model based on the attention mechanism, input the set of metabolic syndrome risk feature parameters into the multimodal risk prediction model for training, and obtain the metabolic syndrome risk prediction results. S5. Validate the risk prediction results of metabolic syndrome and generate a metabolic syndrome risk prediction report based on the validated risk prediction results.

2. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 1, characterized in that, The preprocessing of the classified and labeled multimodal data of metabolic syndrome patients, the extraction of syndrome feature parameter sets from the multimodal data of metabolic syndrome patients based on a deep learning framework, and the cross-modal data matching of the syndrome feature parameter sets with historical metabolic syndrome datasets include the following steps: S21. Perform targeted preprocessing of multimodal data based on the classification and labeling of multimodal data from patients with metabolic syndrome; S22. Construct a multimodal feature extraction network model based on a deep learning framework, and use the multimodal feature extraction network model to extract standardized image datasets, structured clinical datasets, and time-series life datasets of patients with metabolic syndrome. S23. Integrate the standardized image dataset, structured clinical dataset, and time-series lifestyle dataset to obtain a set of characteristic parameters for the syndrome; S24. The cosine similarity algorithm is used to calculate the sample matching degree between the syndrome feature parameter set and the historical metabolic syndrome dataset, and a cross-modal association mapping table is established based on the sample matching degree to perform cross-modal data matching.

3. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 1, characterized in that, The preset metabolic syndrome medical database and multimodal feature weight allocation rules, based on which cross-modal data matching results are weighted and fused, and combined with the metabolic syndrome medical database to form a metabolic syndrome risk feature parameter set, include the following steps: S31. Define clinical diagnostic criteria, imaging feature thresholds, and life risk factor benchmarks to construct a medical database for metabolic syndrome, and set multimodal feature weight allocation rules based on attention mechanisms; S32. Extract the data modal feature parameters of cross-modal data matching, and perform weighted calculation on the data modal feature parameters based on the multimodal feature weight allocation rule to obtain the multimodal weight parameter set; S33. Combine the multimodal weight parameter set with the metabolic syndrome medical database for data correction, and align the feature dimensions of the corrected multimodal weight parameter set to form a metabolic syndrome risk feature parameter set.

4. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 1, characterized in that, The process of constructing a multimodal risk prediction model based on an attention mechanism, and training the model by inputting a set of metabolic syndrome risk feature parameters to obtain the metabolic syndrome risk prediction results, includes the following steps: S41. Construct a multimodal risk prediction neural network model based on the attention mechanism, and introduce a modal difference penalty term into the multimodal risk prediction neural network model; S42. Divide the set of risk characteristic parameters for metabolic syndrome into a training set and a test set, and iteratively train the multimodal risk prediction neural network model. S43. Input the set of risk feature parameters of metabolic syndrome into the trained multimodal risk prediction neural network model to obtain three types of risk probability values, and filter the three types of risk probability values. Use the filtered three types of risk probability values ​​as the risk prediction results of metabolic syndrome.

5. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 2, characterized in that, The process of constructing a multimodal feature extraction network model based on a deep learning framework and using this model to extract standardized image datasets, structured clinical datasets, and time-series lifestyle datasets from patients with metabolic syndrome includes the following steps: S221. A three-branch parallel network and feature alignment layer architecture are set up using a deep learning framework, and a modality adaptation optimization term is introduced for the three-branch parallel network to build a multimodal feature extraction network model. S222. Input the multimodal data of patients with metabolic syndrome into a multimodal feature extraction network model for feature classification and extraction to obtain standardized image datasets, structured clinical datasets, and time-series life datasets. S223. Validate and output the standardized image dataset, structured clinical dataset, and time-series life dataset.

6. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 3, characterized in that, The process of extracting cross-modal data matching data modal feature parameters and weighting these parameters based on multimodal feature weight allocation rules to obtain a multimodal weight parameter set includes the following steps: S321. Extract the separated image modality feature parameters, clinical modality feature parameters, and time-series lifestyle modality feature parameters from the cross-modal data matching results; S322. Calculate the set of single-modal feature weight coefficients for separating image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters based on the multimodal feature weight allocation rule; S323. Perform dimensional consistency verification on the single-modal feature weight coefficients, remove abnormal weight parameters, and integrate them to form a multimodal weight parameter set.

7. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 4, characterized in that, The construction of the multimodal risk prediction neural network model based on the attention mechanism, and the introduction of a modality difference penalty term into the multimodal risk prediction neural network model, includes the following steps: S411. Constructing a multimodal feature input layer for a multimodal risk prediction neural network model based on an attention mechanism; S412. Set up a cross-modal attention interaction layer for the multimodal risk prediction neural network model, and generate an intermodal attention matrix through an attention mechanism; S413. Preset the modal difference penalty function, construct the risk prediction output layer to integrate the inter-modal attention matrix and the modal difference penalty function, and form a multimodal risk prediction neural network model.

8. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 5, characterized in that, The process of constructing a multimodal feature extraction network model using a deep learning framework, setting up a three-branch parallel network and a feature alignment layer architecture, and introducing a modality adaptation optimization term for the three-branch parallel network includes the following steps: S2211. A three-branch parallel network is built based on a deep learning framework, which includes an improved residual neural network architecture for the imaging branch, a deep multilayer perceptron architecture for the clinical branch, and a Transformer architecture for the temporal branch. S2212. Set a feature alignment layer at the output of the three-branch parallel network, and map the output features of each branch to the feature space of the same dimension through the projection matrix, and calculate the intramodal feature similarity matrix. S2213. Preset modal adaptation optimization strategy and bias to construct adaptation loss function, combine the intra-modal feature similarity matrix with modal adaptation optimization strategy and bias to construct adaptation loss function, and form the total loss function of multimodal feature extraction network; S2214. The parameters of the three-branch parallel network and the feature alignment layer are iteratively updated according to the total loss function to form a multimodal feature extraction network model.

9. The method for predicting the risk of metabolic syndrome based on multimodal data according to claim 6, characterized in that, The calculation of the single-modal feature weight coefficient set based on the multimodal feature weight allocation rule for separating image modal feature parameters, clinical modal feature parameters, and time-series lifestyle modal feature parameters includes the following steps: S3221. Extract image feature thresholds, clinical diagnostic criteria, and life risk factor benchmarks from the metabolic syndrome medical database as weight calculation constraints. S3222. Construct a single-modal feature weight calculation model based on the attention mechanism. Input the separated image modality feature parameters, clinical modality feature parameters, and time series life modality feature parameters into the single-modal feature weight calculation model to obtain the separated image weight parameters, clinical modality weight parameters, and time series life weight parameters. S3223. Based on the weight calculation constraints, the weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are corrected, and the corrected weight parameters of the separated images, the weight parameters of the clinical modality, and the weight parameters of the time series life are integrated to obtain a set of single-modality feature weight coefficients.

10. A metabolic syndrome risk prediction system based on multimodal data for implementing the metabolic syndrome risk prediction method based on multimodal data according to any one of claims 1-9, characterized in that, The system includes: a parameter acquisition and annotation module, a parameter processing and matching module, a feature weighting and fusion module, a model generation and training module, and a validation report generation module; The parameter acquisition and annotation module is used to acquire multimodal data of patients with metabolic syndrome and historical labeled datasets of metabolic syndrome, and to classify and annotate the multimodal data of patients with metabolic syndrome. The parameter processing and matching module is used to preprocess the multimodal data of metabolic syndrome patients after classification and labeling. It extracts the syndrome feature parameter set of the multimodal data of metabolic syndrome patients based on a deep learning framework, and performs cross-modal data matching between the syndrome feature parameter set and the historical metabolic syndrome dataset. The feature weighted fusion module is used to preset the metabolic syndrome medical database and multimodal feature weight allocation rules, perform weighted fusion on the cross-modal data matching results based on the multimodal feature weight allocation rules, and combine the metabolic syndrome medical database to form a set of metabolic syndrome risk feature parameters; The model generation and training module is used to build a multimodal risk prediction model based on the attention mechanism. The risk feature parameter set of metabolic syndrome is input into the multimodal risk prediction model for training, and the risk prediction result of metabolic syndrome is obtained. The validation report generation module is used to validate the metabolic syndrome risk prediction results and generate a metabolic syndrome risk prediction report based on the validated metabolic syndrome risk prediction results.