Cerebral stroke risk prediction method and system based on multi-modal data fusion
Through deep learning technology and self-attention mechanism, multimodal data is integrated to predict stroke risk, solving the problem of existing methods relying on a single data source, achieving higher prediction accuracy and personalized support.
Patent Information
- Application Number
- CN202510061793.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-06
AI Technical Summary
Existing stroke risk prediction methods rely on a single physiological data and cannot effectively utilize multimodal data, such as image data, gene data and time series data, resulting in insufficient prediction accuracy.
Deep learning technology and self-attention mechanism are adopted to integrate medical imaging data, gene data, electronic health records and biomarker data, and optimize feature fusion through feature embedding, multimodal fusion and collaborative training, and achieve dynamic prediction and real-time updates.
It improves the accuracy and personalization of stroke risk prediction, can dynamically adjust the prediction model, adapt to the clinical scenarios of different patients, and provides timely and accurate risk assessment support.
Smart Images

Figure CN119943401A_ABST
Abstract
Description
Technical Field
[0003] The present invention relates to artificial intelligence and deep learning technology, and in particular to a method and system for predicting stroke risk based on multimodal data fusion.
[0004] The project on which this invention is based: Shanxi Province Key R&D Plan "Research and Application of Key Technologies for Predicting Risk of Stroke Based on Big Data Analysis". The dataset used is from the Second Hospital of Shanxi Medical University. The dataset covers imaging data, electronic health records (EHR) and other related data. Background Art
[0006] Stroke is one of the leading causes of death and long-term disability worldwide, usually caused by blockage or rupture of cerebral blood vessels. Traditional stroke risk prediction methods, such as the Framingham risk score model and the SCORE risk score model, mainly rely on single physiological data (such as blood pressure, cholesterol levels, etc.) for evaluation. These methods have significant limitations: they ignore the importance of multimodal information such as imaging, genetic information, and time series data, and cannot fully capture multidimensional pathological changes and their complex associations, thereby limiting the accuracy of prediction.
[0007] In recent years, multimodal data fusion has become an important research direction to improve the accuracy of stroke risk prediction. Deep learning technology has demonstrated excellent performance in modeling medical imaging, genetic data, and time series data, but it still faces challenges. Multimodal data comes from various sources, and there are problems such as dimensional differences and inconsistent annotations, which makes feature extraction and fusion more complicated. In addition, traditional weighted fusion methods are difficult to adapt to the dynamic changes of multimodal data and cannot flexibly respond to the differences in the importance of each data modality in different patients.
[0008] Therefore, designing a method that can efficiently integrate multimodal data features and perform real-time predictions on dynamic changes has important theoretical value and practical significance. Summary of the invention
[0010] The technical problem to be solved by the present invention is to provide a stroke risk prediction method and system based on multimodal data fusion in response to the shortcomings of the existing technology; by integrating various types of data and using advanced deep learning technology and self-attention mechanism to optimize feature fusion, it can not only improve the accuracy of prediction, but also provide personalized and timely support for clinical decision-making, thereby improving the patient's prognosis and quality of life.
[0011] To achieve the above object, the present invention provides the following technical solutions:
[0012] A method for predicting stroke risk based on multimodal data fusion, comprising the following steps:
[0013] A1 Data collection and preprocessing steps: Collect stroke-related multimodal data from different data sources such as medical imaging data, genetic data, electronic health records (EHR), and biomarker data, and preprocess and standardize the collected data;
[0014] A2 Feature embedding and multimodal fusion steps: feature extraction and embedding of each modality data through deep learning network, where medical image data uses convolutional neural network (CNN) to extract structural information, time series data uses recurrent neural network (RNN) to mine time series features, gene data uses embedded method to map gene variation information to low-dimensional vector space, biomarker data uses fully connected network to extract potential features, and then uses feature embedding module to map different modality features to a unified feature space to achieve fusion;
[0015] A3 Self-attention mechanism optimizes feature fusion steps: Based on the self-attention mechanism and cross-modal attention mechanism, weighted processing is performed on the features of each modality. First, the feature vectors or feature matrices obtained by deep learning model processing of each modality are mapped to feature representation, and the query matrix (Q), key matrix (K), and value matrix (V) are constructed to calculate the self-attention weights. The cross-modal attention mechanism is introduced to query, construct keys and values between different modalities, and calculate feature similarity to weightedly fuse the features of each modality.
[0016] A4 Multimodal collaborative training steps: Use shared and private network structures for collaborative training. The shared part designs the network structure to learn cross-modal common features, optimizes cross-modal feature fusion through parameter initialization, loss function design, parameter optimization and back propagation, and cross-modal loss function design. The private part designs a separate network module for each modality, and uses the weighted sum of their respective loss functions as the final loss function to optimize the extraction of proprietary information for each modality.
[0017] A5 Dynamic prediction and real-time update steps: After the model training is completed, a dynamic prediction mechanism is set up to combine real-time data to conduct stroke risk assessment, and incremental learning and online learning methods are used to update model parameters. The prediction results are adjusted based on the comparison between the new data and the original prediction results to achieve tracking of changing trends in patients' health status and real-time update of risk prediction results.
[0018] In step A1, the medical imaging data includes: CT images of stroke patients, each image has a size of 512×512 and a single-channel grayscale format; preprocessing includes resizing to 128×128, denoising (using a Gaussian filter), and normalization; genetic data includes gene mutation information related to stroke, with a dimension of the number of samples N×20,000 gene loci; an embedded method is used to map these high-dimensional sparse information to a low-dimensional feature space; EHR data (electronic health record): the time series data in the EHR data includes changes in 10 physiological indicators (including: blood pressure, blood sugar, heart rate, blood oxygen saturation, body temperature, respiratory rate, intracranial pressure, body mass index (BMI), electrolyte balance, and intake and output) recorded once every hour for 24 hours, which means that the data covers a whole day, and these key physiological parameters are measured and recorded every hour to form a continuous time series data stream. Time series features were mined through recurrent neural networks (RNNs); biomarker data: a total of 8 features (including: C-reactive protein (CRP), interleukin-6 (IL-6), tumor necrosis factor α (TNF-α), homocysteine, blood glucose, triglycerides, high-density lipoprotein cholesterol (HDL-C) and adiponectin), and potential risk features were extracted using a fully connected network.
[0019] In step A1, in the data collection and preprocessing step, the medical image data preprocessing includes denoising, normalization and format unification of image data to improve data quality and usability, so as to facilitate the subsequent deep learning network to accurately extract features.
[0020] Medical imaging data preprocessing includes:
[0021] Denoising: Use Gaussian filter to reduce noise interference.
[0022] Normalization: Map pixel values to the range of 0-1 to improve the stability of model training; the formula is as follows:
[0023] in, is the original image data, is the normalized data, and are the minimum and maximum pixel values respectively;
[0024] Gene data preprocessing: Use embedding methods to map gene mutation information into a low-dimensional vector space; the formula is as follows: ;
[0025] in, is the embedding matrix, is a sparse gene mutation vector, For bias.
[0026] EHR data preprocessing:
[0027] Extract time series features, such as physiological parameters such as blood pressure and blood sugar, and perform missing value filling and outlier processing. Time series data can be filled with missing values by interpolation. The formula is as follows:
[0028] Biomarker data preprocessing: Data normalization: Same as the normalization method for medical imaging data.
[0029] In step A2, for medical imaging data, the convolutional layer of the convolutional neural network (CNN) adopts a specific convolution kernel size and step size setting to accurately capture key imaging information such as vascular structure, thereby improving the accuracy and effectiveness of feature extraction of imaging data; for time series data, a recurrent neural network (RNN) is used to capture the temporal features in the patient's historical health records and extract dynamic change patterns; for genetic data, an embedding method is used to map high-dimensional sparse gene mutation information to a low-dimensional feature space, retaining mutation characteristics and sequence correlation; for biomarker data, the potential risk features of biological signals are deeply mined based on a fully connected network, and the multi-dimensional information of multimodal data is integrated through the embedding of a unified feature space to provide accurate feature expression for subsequent risk prediction.
[0030] In the processing of medical imaging data, a convolutional neural network (CNN) was used with a 3x3 convolution kernel and a step size of 1 to accurately capture key information such as vascular structure in CT images. The input image was resized to 128×128 pixels and subjected to Gaussian filtering, denoising and normalization. For time series data, a standard recurrent neural network (RNN) with 2 layers and 256 hidden units in each layer was used to mine the temporal features in the electronic health record (EHR). The input shape was [number of samples N, 24 hours, 10 physiological indicators] to extract dynamic change patterns. The genetic data was mapped to a low-dimensional vector by setting the 20,000-dimensional gene locus information to a 128-dimensional embedding space, and the weights were initialized with a uniform distribution. The biomarker data was extracted using a two-layer fully connected network. The first layer had 64 nodes and the nonlinear expression was increased by the ReLU activation function. The second layer was output to a 128-dimensional feature space for fusion with other modal data, and a dropout ratio of 0.5 was applied to prevent overfitting. These specific configurations work together to achieve a more accurate prediction of stroke risk.
[0031] Medical imaging data: Convolutional neural network (CNN) is used to extract structural information. The convolution kernel size and step size of the convolution layer are set to k×k and s, and the formula is: ,in, is the convolution kernel weight, is the bias term, is the activation function (such as ReLU).
[0032] Time series data: Mining time series features through recurrent neural networks (RNN). The update formula of RNN is: h t = σ ( W h [ h t − 1 , x t ] + b h ) ,in, is the hidden state, is the input, is the weight matrix, is the bias term.
[0033] Genetic data: Genetic variation information is mapped to a low-dimensional vector space using an embedded method. The formula is:
[0034] Biomarker data: Extract latent features through a fully connected network, the formula is: ,in, is the weight matrix, is the bias term, is the activation function.
[0035] Finally, the features of different modalities are mapped to a unified feature space through the feature embedding module to achieve fusion.
[0036] Feature fusion: Use the feature embedding module to map different modal features to a unified feature space to achieve fusion. Assuming that F1, F2, …, Fn are feature vectors of different modalities, the fused feature F can be expressed as:
[0037] In the step A3, in the step of optimizing feature fusion by the self-attention mechanism, when calculating the self-attention weight, the query matrix (Q), the key matrix (K), and the value matrix (V) are constructed according to a specific linear transformation formula to ensure that the weight calculation accurately reflects the importance of each modal feature, thereby improving the feature fusion effect and risk prediction accuracy.
[0038] Mapping of feature representation: After processing the data of each modality through a deep learning model, a feature vector or feature matrix is obtained. Let Q, K, and V be the query matrix, key matrix, and value matrix, respectively, which are generated from the feature matrix X through linear transformation: , , in, , , is the weight matrix.
[0039] Calculate self-attention weight: Calculate the attention weight A according to the query matrix Q and the key matrix K. The formula is:
[0040] in, is the dimension of the key matrix.
[0041] Weighted summation: Apply the attention weights A to the value matrix V to obtain the weighted feature representation Z:
[0042] Among them, A is the attention weight matrix, V is the value matrix, and Z is the final fused feature.
[0043] In step A4, in the shared part loss function design of the multimodal collaborative training step, the classification loss function in the joint loss function is calculated using cross entropy loss, and the regularization loss is based on or Regularization method, cross-modal loss is measured by contrast loss, and the optimal value of each loss weight hyperparameter is determined based on experimental analysis of a large amount of sample data to ensure that the model effectively learns cross-modal shared features and accurately predicts risks. Shared part model training:
[0044] Parameter initialization: Use the Xavier initialization method to initialize the weights of the fully connected layer:
[0045] ,
[0046] in, and are the number of input and output nodes respectively.
[0047] Loss function design: The joint loss function consists of classification loss, regularization loss and cross-modal loss, and the formula is: .in, Cross entropy loss, yes or Regularization loss, is the contrast loss, and is a hyperparameter.
[0048] Parameter optimization and back propagation: Use the Adam optimizer to update the network parameters of the shared part. The update rule is: .in, is the learning rate, and are the first and second order momentum, is a smoothing term.
[0049] In step A4, in the private part model training of the multimodal collaborative training step, the loss function of each modality private network module can select cross entropy loss or mean square error loss according to the modality characteristics, and the learning rate is adjusted by an adaptive learning rate optimization method according to the sample distribution characteristics to ensure that the proprietary information of each modality is fully extracted and the model performance is optimized.
[0050] Private part model training: Design a network module for each modality separately and optimize it through its independent loss function. The loss function for medical imaging data can be mean square error loss:
[0051] The loss function for genetic data can be cross entropy loss:
[0052]
[0053] The loss function for EHR (Electronic Health Record) data can be the cross entropy loss:
[0054]
[0055] The loss function for biomarker data can be the mean squared error loss:
[0056]
[0057] In step A5, in the dynamic prediction and real-time update step, the incremental learning method determines the scope and frequency of incorporating new data into model training based on the sliding window technology, and adaptively adjusts the model update granularity according to the timeliness and importance of the data, thereby improving the model's efficiency in processing real-time data and the timeliness of risk prediction.
[0058] Incremental learning and online learning: Incremental learning is used to determine the scope and frequency of new data included in model training based on sliding window technology. Every time new data is received, the model weights are quickly updated without having to retrain the entire model. The update formula for incremental learning is:
[0059] .in, is the weight increment calculated based on the new data.
[0060] In step A5, in the dynamic prediction and real-time update step, when the prediction result is updated, the significance of the risk change caused by the new data is determined based on the risk threshold setting and Bayesian decision theory, and the model prediction is corrected in combination with clinical expert experience feedback to enhance the reliability and clinical practicality of the prediction result.
[0061] Prediction update: Whenever new data is received, the system uses the existing training model to make predictions and compares them with previous predictions. If the new data indicates that the patient's risk level has changed significantly, the system automatically adjusts the prediction. Risk threshold setting and Bayesian decision theory are used to determine the significance of risk changes. The formula is:
[0062] in, is the prior probability, is the likelihood function, is the marginal probability.
[0063] Furthermore, the deep learning network uses GPU cluster parallel computing technology to accelerate calculations during training and prediction, dynamically allocates computing resources according to data modal characteristics and network structure, and optimizes overall computing efficiency and model training and prediction speed.
[0064] Furthermore, this method introduces multiple sets of external public stroke datasets for comparative verification during the model training and evaluation stages, adjusts model parameters according to the characteristics of different datasets, and improves the model's generalization ability as well as the accuracy and stability of risk prediction in different clinical scenarios.
[0065] The present invention has the following beneficial effects:
[0066] (1) This prediction method can comprehensively assess the risk of stroke by integrating medical imaging data, genetic data, electronic health records (EHR) and biomarker data, thus overcoming the limitations of a single data source. Through the independent feature extraction and embedding processing of various data types by deep learning models, this method can capture the deep information in each data modality and provide a more detailed and accurate risk assessment. For example, medical imaging data uses CNN to extract vascular structure information, genetic data uses embedded methods to extract mutation information, EHR data uses RNN to process patients' time series health records, and biomarker data uses a fully connected network to extract potential risk signals. These processing methods ensure that the features of each modality can be deeply mined and fused in a unified feature space, thereby achieving a comprehensive and accurate prediction of stroke risk.
[0067] (2) This prediction method optimizes the fusion process of multimodal features by introducing self-attention mechanism and cross-modal attention mechanism. By calculating the relative importance of different modal features, the model can adaptively adjust the contribution ratio of each modality in the final decision to ensure attention to key features. For example, vascular abnormalities in imaging data and specific mutation information in genetic data may have greater weight in the risk assessment of some patients, while in other patients, physiological data (such as blood pressure, blood sugar, etc.) may be the key factor in predicting risk. The combination of self-attention mechanism and cross-modal attention mechanism enables the model to dynamically adjust the prediction model in the ever-changing clinical scenarios, further improving the accuracy and stability of the prediction.
[0068] (3) Through the collaborative training of shared and private network structures, this prediction method can not only extract rich common features from each modality data, but also retain the proprietary information of each modality. The cross-modal feature learning of the shared part ensures the mutual complementation of different data sources, while the private part can be trained specifically for each modality to avoid information loss and improve the robustness of the model. Through collaborative training, the model can handle complex and diverse stroke risk prediction tasks and can adapt itself when new data arrives. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 This is a flow chart of data collection and preprocessing of the present invention;
[0071] Figure 2 This is a flow chart of feature embedding and multimodal fusion of the present invention;
[0072] Figure 3 Optimizing feature fusion flow chart for the self-attention mechanism of the present invention;
[0073] Figure 4 A multi-modal collaborative training diagram of a shared network structure of the present invention;
[0074] Figure 5 A multi-modal collaborative training diagram of the private network structure of the present invention;
[0075] Figure 6 It is a dynamic prediction and real-time update diagram of the present invention; DETAILED DESCRIPTION
[0077] The present invention is described in detail below in conjunction with specific embodiments.
[0078] refer to Figure 1-6, medical imaging data includes CT images of stroke patients, each image is 512×512 in size, single-channel grayscale format; gene data contains gene mutation information related to stroke, with a dimension of N×20000 (number of gene loci); EHR data records 10 physiological indicators of patients, such as blood pressure, blood sugar, and heart rate, with a time span of 24 hours (recorded once every hour); biomarker data includes inflammatory factors and metabolic indicators, with a total of 8 features. In order to input these data into a unified deep learning model, they need to be preprocessed and all features mapped to the same 128-dimensional space.
[0079] Preprocessing of image data: image resizing, denoising and normalization. First, all CT images were resized to a fixed size of 128×128 using OpenCV, which can reduce computational overhead while maintaining the resolution of key structural information such as blood vessels and lesions in the image.
[0080] Subsequently, a Gaussian filter is used for denoising to reduce random noise in the image. Normalization is to scale the pixel values to the range of 0 to 1 so that the model converges more stably. The formula is: .
[0081] The preprocessing of genetic data mainly targets its high dimensionality and sparsity problems. Through the embedded mapping method, the 20,000-dimensional gene mutation matrix is mapped to a 128-dimensional low-dimensional space, while retaining the key information of mutation characteristics and sequence correlation. in , .
[0082] EHR data, including 10 physiological characteristics (such as blood pressure, blood sugar, heart rate, etc.) for 24 consecutive hours. Missing values of time series data are filled by linear interpolation and normalized to the range of 0 to 1. In order to extract the dynamic features of the time series, a single-layer bidirectional RNN model is used. The input dimension is 10 (representing 10 physiological features), and the hidden unit is set to 128 to capture the temporal correlation between features. The hidden state update formula of RNN is as follows: h t = σ ( W h [ h t − 1 , x t ] + b h ) .
[0083] The data of each modality is represented as follows after their respective features are extracted:
[0084] Feature matrix of medical imaging data;
[0085] Feature vector of genetic data;
[0086] The feature matrix of HR data;
[0087] Feature vectors for biomarker data
[0088] After the feature extraction of the above four modal data is completed, the features of all modalities are concatenated together to form a unified fusion vector. The feature dimension of each modality is 128, and the dimension of the fused feature vector is 512.
[0089] The fused features then enter the self-attention module to optimize the feature expression. The self-attention module can dynamically adjust the weight of each modality feature according to its importance, thereby highlighting the key features. Let Q, K, V be the query matrix, key matrix, and value matrix, respectively, which are generated from the feature matrix X through linear transformation: , , . Calculate the attention weights based on the query matrix Q and the key matrix K:
[0090] ,
[0091] Optimized feature representation It is the input of the model and is used for subsequent classification tasks.
[0092] The model training phase uses a collaborative structure of shared and private networks. The shared network learns common features across modalities, such as the correlation between gene mutations and imaging lesion areas, while the private network optimizes features for each modality to avoid information loss.
[0093] The optimization goal of the shared part is to minimize the classification loss (using cross entropy loss) and regularization loss to ensure that the model has good generalization ability for the overall distribution of data. The private part uses mean square error or cross entropy loss to optimize for imaging, gene, time series and biomarker features respectively.
[0094] In the experiment, we randomly divided the dataset, using 80% of the samples for training and the remaining 20% for testing. The purpose of data partitioning is to ensure that the characteristics of multimodal data can be fully learned during model training, while leaving out data that is not involved in training to evaluate the generalization performance of the model. The training and test set partitioning process uses random seeds to ensure the repeatability of the results. The experimental dataset includes medical images, genetic data, EHR data, and biomarker data, with a total of 1,000 samples. Specifically, 800 samples are used for training and 200 samples are used for testing.
[0095] The model was iteratively trained on the training set for 50 epochs, with each epoch traversing the entire training dataset once. During the training process, we used the Adam optimizer to adjust the model parameters, with the initial learning rate set to 0.001, and dynamically adjusted the learning rate according to the performance of the validation set during the training process to accelerate convergence. At the end of each epoch, the loss value and performance indicators of the training set were calculated, and the generalization ability of the model was evaluated on the test set.
[0096] After training is completed, the model enters the dynamic prediction stage, which can receive new data and update the prediction results in real time. To this end, the incremental learning method is used to dynamically adjust the model parameters based on the sliding window technology. The update formula of incremental learning is:
[0097] in, is the weight increment calculated based on the new data.
[0098] Whenever new data is received, the system uses the existing training model to make predictions and compares them with previous predictions. If the new data indicates that the patient's risk level has changed significantly, the system automatically adjusts the prediction. Risk threshold setting and Bayesian decision theory are used to determine the significance of risk changes. The formula is:
[0099]
[0100] Finally, after 50 epochs of training, the model performed well on the test set, with an accuracy of 92.3%, a precision of 89.7%, a recall of 91.5%, and an F1 score of 90.6%. The dynamic prediction module further improved the real-time performance of the model, and the accuracy of the prediction of the newly added samples reached 95%. The specific data is shown in Table 1:
[0101] Table 1 Evaluation indicators
[0102] Evaluation indicators Performance during training Testing phase performance Dynamic prediction performance Accuracy 94.5% 92.3% 95.0% Precision 91.8% 89.7% 93.2% Recall 93.6% 91.5% 94.8% F1 Score 92.7% 90.6% 94.0%
[0103] Experimental results show that this method performs well in the task of stroke risk prediction, significantly improving the accuracy and generalization of prediction. At the same time, the introduction of the dynamic prediction module provides strong support for the model to adapt to the dynamic changes of data in the clinical environment, so that the model can not only process static data, but also continuously optimize the prediction effect in real-time scenarios. This feature is of great significance for clinical applications. For example, it can be used for continuous monitoring of patients, timely detection of risk changes, and support for doctors' intervention decisions.
[0104] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
Claims
1. A stroke risk prediction method based on multimodal data fusion, characterized in that: The following steps are involved: A1 Data collection and preprocessing steps: Collect stroke-related multimodal data from different data sources such as medical imaging data, genetic data, electronic health record EHR and biomarker data, and preprocess and standardize the collected data; A2 Feature embedding and multimodal fusion steps: feature extraction and embedding of each modality data through deep learning network, where medical image data uses convolutional neural network CNN to extract structural information, time series data in EHR data uses recurrent neural network RNN to mine temporal features, gene data uses embedded method to map gene variation information to low-dimensional vector space, biomarker data uses fully connected network to extract potential features, and then uses feature embedding module to map different modality features to a unified feature space for fusion; A3 Self-attention mechanism optimizes feature fusion steps: Based on the self-attention mechanism and cross-modal attention mechanism, weighted processing is performed on the features of each modality. First, the feature vectors or feature matrices obtained by deep learning model processing of each modality are mapped to feature representation, and the query matrix (Q), key matrix (K), and value matrix (V) are constructed to calculate the self-attention weights. The cross-modal attention mechanism is introduced to query, construct keys and values between different modalities, and calculate feature similarity to weightedly fuse the features of each modality. A4 Multimodal collaborative training steps: Use shared and private network structures for collaborative training. The shared part designs the network structure to learn cross-modal common features, optimizes cross-modal feature fusion through parameter initialization, loss function design, parameter optimization and back propagation, and cross-modal loss function design. The private part designs a separate network module for each modality, and uses the weighted sum of their respective loss functions as the final loss function to optimize the extraction of proprietary information for each modality. A5 Dynamic prediction and real-time update steps: After the model training is completed, a dynamic prediction mechanism is set up to combine real-time data to conduct stroke risk assessment, and incremental learning and online learning methods are used to update model parameters. The prediction results are adjusted based on the comparison between the new data and the original prediction results to achieve tracking of changing trends in patients' health status and real-time update of risk prediction results.
2. The prediction method according to claim 1, characterized in that: In step A1, the medical imaging data includes: CT images of stroke patients, each image size is 512×512, single-channel grayscale format; preprocessing includes resizing to 128×128, denoising and normalization; gene data includes gene mutation information related to stroke, the dimension is the number of samples N×20000 gene loci; An embedded method is used to map these high-dimensional sparse information into a low-dimensional feature space; EHR data: The time series data in the EHR data contains the changes in the patient's physiological indicators recorded once every hour within 24 hours, forming a continuous time series data stream; time series features are mined through recurrent neural networks; biomarker data: including C-reactive protein (CRP), interleukin-6 (IL-6), tumor necrosis factor α (TNF-α), homocysteine, blood glucose, triglycerides, high-density lipoprotein cholesterol (HDL-C) and adiponectin, a total of 8 features, and a fully connected network is used to extract potential risk features.
3. The prediction method according to claim 1, characterized in that: In the step A1, in the data collection and preprocessing step, the medical image data preprocessing includes denoising, normalization and format unification of image data to improve data quality and usability, so as to facilitate accurate feature extraction by the subsequent deep learning network; wherein: Medical imaging data preprocessing includes: Denoising: Use Gaussian filter to reduce noise interference; Normalization: Map pixel values to the range of 0-1 to improve the stability of model training; the formula is as follows: in, is the original image data, is the normalized data, and are the minimum and maximum pixel values respectively; Gene data preprocessing: Use embedding methods to map gene mutation information into a low-dimensional vector space; the formula is as follows: ; in, is the embedding matrix, is a sparse gene mutation vector, is bias; EHR data preprocessing: Extract time series features, such as physiological parameters such as blood pressure and blood sugar, and perform missing value filling and outlier processing; time series data can fill missing values through interpolation; the formula is as follows: ; Biomarker data preprocessing: Data standardization: Same as medical imaging data normalization method.
4. The prediction method according to claim 1, characterized in that: In step A2, in the medical image data processing, the convolutional neural network CNN is used with a convolution kernel of 3x3 size and a step size of 1 to accurately capture key information such as vascular structure in CT images. The input image is resized to 128×128 pixels and subjected to Gaussian filtering, denoising and normalization. For time series data, a standard recurrent neural network RNN containing 2 layers and 256 hidden units in each layer is used to mine the temporal features in HER data, and the input shape is [number of samples N, 24 hours, 10 physiological indicators] to extract dynamic change patterns. The gene data is mapped to a low-dimensional vector by setting it to a 128-dimensional embedding space, and the weights are initialized with a uniform distribution. The biomarker data is feature extracted using a two-layer fully connected network, the first layer has 64 nodes and uses a ReLU activation function to increase nonlinear expression, and the second layer outputs to a 128-dimensional feature space for fusion with other modality data, and a dropout ratio of 0.5 is applied to prevent overfitting.
5. The prediction method according to claim 1, characterized in that: In step A3, in the step of optimizing feature fusion by the self-attention mechanism, when calculating the self-attention weight, the query matrix Q, the key matrix K, and the value matrix V are constructed according to a specific linear transformation formula to ensure that the weight calculation accurately reflects the importance of each modal feature, thereby improving the feature fusion effect and risk prediction accuracy; Mapping of feature representation: After processing the data of each modality through the deep learning model, a feature vector or feature matrix is obtained; let Q, K, V be the query matrix, key matrix and value matrix respectively, which are generated from the feature matrix X through linear transformation: , , in, , , is the weight matrix; Calculate self-attention weight: Calculate the attention weight A according to the query matrix Q and the key matrix K. The formula is: ;in, is the dimension of the bond matrix; Weighted summation: Apply the attention weights A to the value matrix V to obtain the weighted feature representation Z: ; Where A is the attention weight matrix, V is the value matrix, and Z is the final fused feature.
6. The prediction method according to claim 1, characterized in that: In step A4, in the shared part loss function design of the multimodal collaborative training step, the classification loss function in the joint loss function is calculated using cross entropy loss, and the regularization loss is based on or Regularization method, cross-modal loss is measured by contrast loss, and the optimal value of each loss weight hyperparameter is determined based on experimental analysis of a large amount of sample data to ensure that the model effectively learns cross-modal shared features and accurately predicts risks; shared part model training: Parameter initialization: Use the Xavier initialization method to initialize the weights of the fully connected layer: , ;in, and are the number of input and output nodes, respectively; Loss function design: The joint loss function consists of classification loss, regularization loss and cross-modal loss, and the formula is: ;in, Cross entropy loss, yes or Regularization loss, is the contrast loss, and is a hyperparameter; Parameter optimization and back propagation: Use the Adam optimizer to update the network parameters of the shared part. The update rule is: ;in, is the learning rate, and are the first and second order momentum, is a smoothing term.
7. The prediction method according to claim 1, characterized in that: In step A4, in the private part model training of the multimodal collaborative training step, the loss function of each modality private network module can select cross entropy loss or mean square error loss according to the modality characteristics, and the learning rate is adjusted by an adaptive learning rate optimization method according to the sample distribution characteristics to ensure that the proprietary information of each modality is fully extracted and the model performance is optimized; Private part model training: Design a network module for each modality separately and optimize it through its independent loss function; the loss function for medical imaging data is mean square error loss: ; The loss function for genetic data is cross entropy loss: ; The loss function for EHR (electronic health record) data can be cross entropy loss: ; The loss function for biomarker data is the mean squared error loss: 。 8. The prediction method according to claim 1, characterized in that: In step A5, in the dynamic prediction and real-time update step, the incremental learning method determines the scope and frequency of incorporating new data into model training based on the sliding window technology, and adaptively adjusts the model update granularity according to the timeliness and importance of the data, thereby improving the model's efficiency in processing real-time data and the timeliness of risk prediction; Incremental learning and online learning: Using incremental learning methods, the scope and frequency of new data included in model training are determined based on sliding window technology; each time new data is received, the model weights are quickly updated without having to retrain the entire model; The update formula for incremental learning is: ;in, is the weight increment calculated based on the new data.
9. The prediction method according to claim 1, characterized in that: In step A5, in the dynamic prediction and real-time update step, when the prediction result is updated, the significance of the risk change caused by the new data is determined based on the risk threshold setting and Bayesian decision theory, and the model prediction is corrected in combination with clinical expert experience feedback to enhance the reliability and clinical practicality of the prediction result; Prediction update: Whenever new data is received, the system uses the existing training model to make predictions and compares them with previous predictions. If the new data indicates that the patient's risk level has changed significantly, the system automatically adjusts the prediction. Risk threshold setting and Bayesian decision theory are used to determine the significance of risk changes. The formula is: ;in, is the prior probability, is the likelihood function, is the marginal probability.
10. A risk prediction system according to any one of the risk prediction methods of claims 1 to 9, characterized in that: The system implements any of the stroke risk prediction methods.
Citation Information
Patent Citations
Self-supervised action recognition method based on cross-modal time sequence contrast learning
CN116721458A
Low-voltage distribution network phase sequence identification method based on Bayesian probability theory
CN116819188A
Tobacco disease and insect pest identification method, medium and system
CN118072251A
Cancer risk prediction method based on deep learning
CN118507048A
Systems and methods for deep orthogonal fusion for multimodal prognostic biomarker discovery
EP4239647A1
Cited By
Rehabilitation training evaluation system for dealing with schizophrenia patients
CN120199504A
Multi-modal confidence coefficient dynamic interactive fusion method for stroke rehabilitation diagnosis
CN120473125A
Multimodal confidence dynamic interactive fusion method for stroke rehabilitation diagnosis
CN120473125B
Typhoon rapid enhancement prediction method based on time-space sequence and multi-modal feature fusion
CN120633957A
A rapid enhancement prediction method for typhoons based on the fusion of spatiotemporal sequences and multimodal features.
CN120633957B