Heart failure prediction system and prediction method based on deep learning
By combining the VGG-16 network and the MLP network to process CT or MRI images and clinical data, efficient and accurate prediction of heart failure is achieved, solving the problems of subjective judgment of doctors and insufficient accuracy of machine learning in existing technologies.
Patent Information
- Application Number
- CN202510869475.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-19
Smart Images

Figure CN120670959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition and analysis, and in particular to a heart failure prediction system and prediction method based on deep learning, which are used to determine whether a patient has heart failure. Background Art
[0002] With the rapid development of artificial intelligence, scientific and technological progress is advancing by leaps and bounds. In the medical field, the increasing maturity of computer technology has provided new ideas for determining whether a patient has heart failure.
[0003] Currently, diagnosis criteria are often based solely on the physician's experience, resulting in qualitative and subjective diagnoses. This can easily lead to misjudgments and significantly reduce judgment efficiency. In this new era, relying on manpower will result in a thankless task. The development of machine learning, deep learning, and multidisciplinary integration has provided relatively mature theoretical support and technical guarantees for the diagnosis of heart failure, and has also yielded many advanced research results. However, traditional medical judgment models suffer from shortcomings such as significant influence from physician subjective factors, poor applicability, and difficulty in quantifying severity. Existing judgment models based on machine learning technology still perform poorly in terms of judgment accuracy. Summary of the Invention
[0004] To address the above issues, the present invention aims to provide a deep learning-based heart failure prediction system and method for determining whether a patient has heart failure. This method primarily utilizes a VGG-16 network to extract features from patient imaging data, then uses an MLP to process clinical data features. Finally, multimodal learning is used to fuse these two approaches for final prediction.
[0005] The present invention provides a heart failure prediction system based on deep learning, comprising: A data acquisition module for acquiring CT or MRI images and clinical data of patients from multiple data sources; A first data preprocessing module extracts first data features of the patient's CT or MRI image using a VGG-16 network; A second data preprocessing module uses an MLP network to process the patient's clinical data to obtain a second data feature; The data fusion processing module fuses the first data features and the second data features output by the first data preprocessing module and the second data preprocessing module through a multimodal learning method to obtain fused feature data, so as to make predictions about the patient based on the fused feature data.
[0006] Clinical data include the patient's age, gender, weight, and underlying disease data.
[0007] The present invention also provides a method for predicting heart failure based on deep learning, comprising the following steps: Train VGG-16 network and MLP network; Extracting first data features of the patient's CT or MRI image using a VGG-16 network; Use the MLP network to process the patient's clinical data to obtain the corresponding second data features; The first data features and the second data features output by the VGG-16 network and the MLP network are fused by a multimodal learning method to obtain fused feature data, so as to make predictions about the patient based on the fused feature data.
[0008] Clinical data include the patient's age, gender, weight, and underlying disease data.
[0009] Optionally, extracting a first data feature of the patient's CT or MRI image using a VGG-16 network includes: Acquire CT or MRI images and perform standardization and normalization; The standardized and normalized CT or MRI image is fed into convolutional and pooling layers. The convolution operation extracts local features by sliding the convolution kernel across the image. Each convolution layer is followed by a ReLU activation function to introduce nonlinearity. The VGG-16 network uses five max pooling layers, which reduce the spatial dimensionality of the feature map. The feature map after convolution and pooling operations will be flattened into a one-dimensional vector and input into the fully connected layer. After processing by the fully connected layer, the first data feature is obtained.
[0010] Optionally, using the MLP network to process the patient's clinical data to obtain the corresponding second data features includes: Obtain patients' clinical data and standardize them; Perform feature selection based on standardized clinical data to screen clinical feature data related to the prediction of heart failure; The screened clinical feature data are input into the trained MLP network to extract the second data feature.
[0011] The deep learning-based heart failure prediction system and prediction method of the present invention combine VGG-16 and MLP. The model can learn the complex associations between imaging and clinical data, and mine more discriminative feature combinations, thereby improving the accuracy of patient health prediction assessment.
[0012] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 This is a workflow diagram of a deep learning-based heart failure prediction system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to illustrate the present invention and are not intended to limit the present invention.
[0015] An embodiment of the present invention provides a heart failure prediction system based on deep learning, comprising: The data acquisition module is used to acquire CT or MRI images and clinical data of patients from multiple data sources.
[0016] The first data preprocessing module uses the VGG-16 network to extract the first data features of the patient's CT or MRI image.
[0017] The second data preprocessing module uses the MLP network to process the patient's clinical data to obtain the second data features.
[0018] The data fusion processing module fuses the first data features and the second data features output by the first data preprocessing module and the second data preprocessing module through a multimodal learning method to obtain fused feature data, so as to make predictions about the patient based on the fused feature data.
[0019] The present invention's deep learning-based heart failure prediction system and method utilizes a combination of VGG-16 and MLP. The model can learn the complex relationships between imaging and clinical data, extracting more discriminative feature combinations, thereby improving the accuracy of patient health prediction assessments. The workflow of each module is described in detail below.
[0020] 1. Data Acquisition Module The data acquisition module can collect CT or MRI images and clinical data from multiple data sources. CT images have relatively high spatial resolution and can clearly display the human anatomy; MRI images have excellent soft tissue contrast and can more finely display the internal soft tissue structure of the human body. It can provide a variety of imaging sequences, such as T1-weighted imaging, T2-weighted imaging, and proton density-weighted imaging, which can provide different tissue information.
[0021] 2. First Data Preprocessing Module In an embodiment of the present invention, the first data preprocessing module uses a VGG-16 network to learn CT or MRI images.
[0022] VGG-16 is a convolutional neural network (CNN) architecture, a deep network structure that has achieved remarkable computer vision performance through its unique structure and training method.
[0023] The VGG-16 network contains 13 convolutional layers, each of which uses a 3x3 kernel and the same padding to maintain spatial resolution. Using small kernels (3x3) instead of larger kernels is a major feature of VGG-16. Its advantage lies in its ability to extract more detailed features and improve the model's expressive power by increasing the number of convolutional layers.
[0024] After the convolutional layer, VGG-16 uses 5 maximum pooling layers (2x2, stride=2) to gradually reduce the spatial dimension, thereby compressing the size of the feature map and reducing computational complexity.
[0025] The final stage of the network consists of three fully connected layers. The first two fully connected layers contain 4096 neurons, and the last fully connected layer contains the same number of neurons as the number of categories (1000 categories for the ImageNet dataset).
[0026] VGG-16 uses ReLU (Rectified Linear Unit) as the activation function. This is because ReLU can effectively avoid the gradient disappearance problem and speed up training compared to traditional Sigmoid or Tanh activation functions.
[0027] The working process of the first data preprocessing module is as follows: First, the patient's CT or MRI image data is fed into the VGG-16 network. These images are typically two-dimensional or three-dimensional medical images, stored as pixels or voxels. It's important to note that before entering the network, the image data typically needs to be normalized. This can be done by resizing the image to the network's required input dimensions (e.g., 224×224 pixels) and normalizing the pixel values to between 0 and 1.
[0028] Next, convolution and pooling layers are applied. Convolution extracts local features by sliding the convolution kernel across the image. Each convolution layer is followed by a ReLU activation function, which introduces nonlinearity and enables the network to learn more complex features. Between convolutional layers, the VGG-16 network uses five max pooling layers. Pooling layers reduce the spatial dimensionality of feature maps, reducing computational effort while preserving important features.
[0029] Finally, the fully connected layer is used. After convolution and pooling, the feature map is flattened into a one-dimensional vector and then input into the fully connected layer.
[0030] The first data feature output by the first data preprocessing module refers to the image features extracted by the VGG-16 network, and its specific data format is a high-dimensional vector. In this embodiment of the present invention, the first data feature after processing by the fully connected layer is a vector of length 4096. This vector contains high-level features extracted from the CT or MRI image, such as edge features, texture features, shape features, etc.
[0031] The embodiment of the present invention uses the VGG-16 network to extract features from the patient's CT or MRI images, and can perform deep learning on CT or MRI images to extract basic image features, such as edges and textures. For example, in CT images, edge features can help distinguish the boundaries between tissues and organs; in MRI images, texture features can reflect subtle structural differences in tissues, such as texture changes in different areas of brain tissue. The VGG-16 network can learn unique feature representations for different tissues and organs in CT and MRI images. Taking liver MRI as an example, VGG-16 can distinguish normal liver tissue, cirrhosis tissue, and liver tumor tissue in terms of signal intensity, morphology, and texture differences, providing a strong basis for subsequent disease diagnosis.
[0032] 3. Second Data Preprocessing Module In the deep learning-based heart failure prediction system of the embodiment of the present invention, the second data preprocessing module uses an MLP network to process the patient's clinical data to obtain a second data feature. The clinical data includes the patient's age, gender, weight, and underlying disease data.
[0033] The Multilayer Perceptron (MLP) is a feedforward neural network consisting of an input layer, hidden layers, and an output layer. Each layer is composed of multiple neurons, and the layers are fully connected. It introduces complexity through nonlinear activation functions, enabling it to handle nonlinear problems.
[0034] In this embodiment of the present invention, an MLP is used to process clinical data such as a patient's age, gender, weight, and underlying diseases for deep learning. For example, older patients may have different patterns in their risk of developing certain diseases. The MLP can learn the relationship between age, a numerical feature, and disease. Therefore, the MLP network is able to effectively fit these complex nonlinear relationships. It can perform nonlinear transformations on the input data in the hidden layer. By learning and integrating data from various dimensions, it can comprehensively consider all aspects of a patient's clinical characteristics.
[0035] The working process of the second data preprocessing module is as follows: First, standardize the data. Continuous variables (such as age, weight, and other clinical data) are standardized or normalized to conform to a specific distribution or range. For example, use StandardScaler to standardize the data to a mean of 0 and a variance of 1. For categorical variables (such as gender and underlying diseases), one-hot encoding or label encoding can be used.
[0036] Next, feature selection is performed. Feature selection methods (such as LASSO) are used to screen for features relevant to heart failure prediction. This helps reduce the number of features and improves model training efficiency and performance. The LASSO feature selection method automatically compresses unimportant feature coefficients to 0 by adding an L1 regularization term to the loss function, thereby achieving feature selection. The data standardized in the previous step is applied to the LASSO feature selection model to obtain the coefficients of each feature. Finally, by eliminating features with a coefficient of 0, important features can be screened. For example, the feature names are ['Age', 'Gender', 'Weight', 'BNP','Hypertension', 'Diabetes', 'Heart Disease', 'BloodPressure', 'Blood Sugar', 'Cholesterol'] and the corresponding feature coefficients are [0.0, 0.0, 0.0,0.7,0.5, -0.3, 0.0, 0.2, 0.0, 0.1]. In this way, the important features ['BNP','Hypertension', 'Diabetes', 'Blood Pressure', 'Cholesterol'] can be filtered out.
[0037] Then, the MLP network is constructed and trained. The MLP network consists of an input layer, a hidden layer, and an output layer. The training is performed using an appropriate optimizer (such as Adam) and a loss function (such as binary cross entropy). The initial learning rate of the Adam optimizer is set to 0.001, and the first-order momentum is used to train the network. Maintain gradient directional stability, second-order momentum Adaptively adjust the learning step size for different feature parameters. When the gradient of a feature is large, the learning rate is automatically reduced; for features with smaller gradients, the learning rate is increased to avoid oscillatory convergence. When using the weighted binary cross-entropy loss function, we apply a weight coefficient α = 2.5 to heart failure-positive samples to address the class imbalance problem in clinical data.
[0038] Finally, the preprocessed clinical data was fed into the trained MLP network to extract the secondary data features. The clinical features, filtered by the LASSO, were then fed into the MLP network, where they underwent nonlinear transformations through two hidden layers (each containing 64 neurons and ReLU activations). The final output was a 64-dimensional vector, serving as the secondary data features. The secondary data features, processed by the MLP network, are high-dimensional feature vectors that reflect the complex, nonlinear relationship between clinical data and heart failure.
[0039] 4. Data fusion processing module The data fusion processing module in the deep learning-based heart failure prediction system of an embodiment of the present invention uses a multimodal learning approach to fuse the first and second data features output by the first and second data preprocessing modules to generate fused feature data. Multimodal learning is a machine learning method that aims to simultaneously process and fuse information from multiple modalities (data types or signals) to improve the model's performance and reasoning capabilities. Each modality represents a different form of input data, such as text, images, audio, video, etc. By combining these different modalities, the model can more comprehensively understand and infer the meaning of the data.
[0040] The working process of the data fusion processing module is as follows: First, the CT or MRI image features extracted by the VGG-16 network and the clinical data features processed by the MLP network are each output as a high-dimensional feature vector, which are the first data feature and the second data feature, respectively.
[0041] Second, we introduce an attention mechanism to dynamically adjust the importance weights of features from different modalities, allowing the model to focus more on features that are more valuable for heart failure prediction. First, we project the clinical feature vector to 4096 dimensions through a fully connected layer, aligning it with the imaging feature dimension. We then concatenate the two modal features to form an 8192-dimensional joint vector. This joint vector then passes through two fully connected layers (8192 dimensions are reduced to 1024 in the first layer and to 2 in the second layer) to output an attention score. This is then normalized using Softmax to obtain the modality weights. Finally, the clinical feature vector and imaging features (both 4096 dimensions) are weighted and summed according to their weights, and then reduced to 256 dimensions through a fully connected layer to generate the final fused feature.
[0042] Third, the above 256-dimensional fusion features are input into the Softmax classifier. First, the fusion features are output through the fully connected layer [z0, z1], which represent the strength of evidence for non-heart failure and heart failure respectively. Then, the Softmax probability conversion is performed, and the formula is When P>0.7, heart failure was diagnosed; when P was between 0.4 and 0.7, it was marked as high risk and follow-up was recommended; when P<0.3, the diagnosis of heart failure was excluded.
[0043] An embodiment of the present invention provides a heart failure prediction system based on deep learning, which can be used to learn and evaluate whether a patient has potential heart failure.
[0044] The embodiment of the present invention also provides a method for predicting heart failure based on deep learning, which includes the following steps: first, training the VGG-16 and MLP networks, then using VGG-16 to extract features from the patient's CT or MRI images, using MLP to process the patient's age, gender, weight, underlying diseases and other clinical data, and finally fusing the two through a merging layer to make a final judgment on whether the patient has heart failure. The complete system workflow diagram is as follows: Figure 1 shown.
[0045] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A heart failure prediction method based on deep learning, characterized in that: The following steps are involved: Train VGG-16 network and MLP network; Extracting first data features of the patient's CT or MRI image using a VGG-16 network; Use the MLP network to process the patient's clinical data to obtain the corresponding second data features; The first data features and the second data features output by the VGG-16 network and the MLP network are fused by a multimodal learning method to obtain fused feature data, so as to make predictions about the patient based on the fused feature data.
2. The heart failure prediction method based on deep learning according to claim 1, characterized in that: Clinical data include the patient's age, gender, weight, and underlying disease data.
3. The heart failure prediction method based on deep learning according to claim 1, characterized in that: The first data features extracted from the patient's CT or MRI image using the VGG-16 network include: Acquire CT or MRI images and perform standardization and normalization; The standardized and normalized CT or MRI image is fed into convolutional and pooling layers. The convolution operation extracts local features by sliding the convolution kernel across the image. Each convolution layer is followed by a ReLU activation function to introduce nonlinearity. The VGG-16 network uses five max pooling layers, which reduce the spatial dimensionality of the feature map. The feature map after convolution and pooling operations will be flattened into a one-dimensional vector and input into the fully connected layer. After processing by the fully connected layer, the first data feature is obtained.
4. The heart failure prediction method based on deep learning according to claim 1, characterized in that: The corresponding second data features obtained by processing the patient's clinical data using the MLP network include: Obtain patients' clinical data and standardize them; Perform feature selection based on standardized clinical data to screen clinical feature data related to the prediction of heart failure; The screened clinical feature data are input into the trained MLP network to extract the second data feature.
5. A heart failure prediction system based on deep learning, characterized in that: The device is used to perform the deep learning-based heart failure prediction method according to any one of claims 1 to 4, comprising: A data acquisition module for acquiring CT or MRI images and clinical data of patients from multiple data sources; A first data preprocessing module extracts first data features of the patient's CT or MRI image using a VGG-16 network; A second data preprocessing module uses an MLP network to process the patient's clinical data to obtain a second data feature; The data fusion processing module fuses the first data features and the second data features output by the first data preprocessing module and the second data preprocessing module through a multimodal learning method to obtain fused feature data, so as to make predictions about the patient based on the fused feature data.
6. The heart failure prediction system based on deep learning according to claim 5, characterized in that: Clinical data include the patient's age, gender, weight, and underlying disease data.