Patient information analysis method and device based on multi-modal fusion, equipment and medium
Through the deep learning model of multimodal fusion, the liver CT images and clinical information are encoded and fused, which solves the problem of insufficient accuracy in cirrhosis diagnosis in the prior art, and achieves early accurate prediction and clinical support for the progress of cirrhosis.
Patent Information
- Application Number
- CN202510403130.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-01
AI Technical Summary
The existing cirrhosis diagnosis methods cannot comprehensively and accurately reflect the complex pathological changes in the progression of cirrhosis, resulting in difficulty in early diagnosis and thus increasing the risk of decompensation period.
The patient information analysis method based on multimodal fusion is adopted, and the liver CT images, radiological feature markers and clinical examination information are encoded and fused through deep learning models to generate cross-modal fusion features for cirrhosis outcome classification.
Achieve more accurate prediction of cirrhosis outcomes, support clinical diagnosis and treatment decisions, and improve the accuracy and reliability of early predictions.
Smart Images

Figure CN120411700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, equipment and medium for analyzing patient information based on multimodal fusion. Background Art
[0002] At present, the incidence of liver cirrhosis has shown a significant upward trend. Patients in the compensated stage of liver cirrhosis often do not show obvious clinical symptoms, which makes early diagnosis complicated. However, according to research data, about 11% of patients with compensated liver cirrhosis progress to the decompensated stage every year. This transformation is accompanied by the emergence of serious complications such as ascites, gastrointestinal bleeding, and hepatic encephalopathy, increasing the risk of patient death. Therefore, early prediction of the progression of liver cirrhosis can timely guide treatment and prevent the progression of compensated liver cirrhosis to the decompensated stage, which is crucial for improving the prognosis of patients.
[0003] At present, scoring systems such as Child-Pugh and the model for end-stage liver disease (MELD), which have been widely used, evaluate the prognosis of liver cirrhosis patients by combining the patients' clinical data and laboratory indicators. However, these methods cannot accurately reflect the changes in the liver's morphology and surface texture during the progression of liver cirrhosis. Computed tomography (CT) can make up for this deficiency. In recent years, as a method for high-throughput extraction of features from medical images, radiomics analysis can quantitatively describe tissue heterogeneity that cannot be visually identified and has been applied to the evaluation of liver cirrhosis and the risk prediction of poor prognosis. At the same time, deep learning methods, in a data-driven manner, autonomously extract relevant features from medical images through convolutional neural networks without the need for artificial design of a feature list, and also show excellent potential in the field of liver diseases. In addition, studies have shown that the human body fat components obtained based on CT images, such as visceral fat and subcutaneous fat, can reflect the liver's metabolic state and histological changes and are significantly correlated with the occurrence of adverse events such as liver cancer and liver decompensation. Therefore, diagnostic methods based on a single parameter or modality have certain limitations, and they may not be able to comprehensively capture the complex pathological changes during the progression of liver cirrhosis.
[0004] Therefore, how to comprehensively and accurately predict the classification results of the outcome of liver cirrhosis to obtain a comprehensive and accurate prediction result of liver cirrhosis has become an urgent problem to be solved. Summary of the Invention
[0005] Based on this, the present invention provides a method, device, computer equipment and medium for analyzing patient information based on multimodal fusion to solve the problem of how to comprehensively and accurately predict whether adverse outcomes occur in liver cirrhosis to obtain a comprehensive and accurate prediction result of liver cirrhosis.
[0006] In a first aspect, an embodiment of the present invention provides a patient information analysis method based on multimodal fusion. The patient information analysis method includes: Obtaining a target liver CT image of a patient, a target radiological feature marker for the target liver CT image, and target clinical examination information corresponding to the patient; Using a feature extraction module in a trained deep learning model to encode the target liver CT image to obtain an image encoding feature, encoding the target radiological feature marker to obtain a radiological encoding feature, and encoding the target clinical examination information to obtain an examination information encoding feature; Using a multimodal feature fusion module in the trained deep learning model to fuse the image encoding feature, the radiological encoding feature, and the examination information encoding feature to obtain a cross-modal fusion feature; Using a classification head module in the trained deep learning model to classify the cross-modal fusion feature for liver cirrhosis outcome to obtain a liver cirrhosis outcome classification result.
[0007] In a second aspect, an embodiment of the present invention provides a patient information analysis device based on multimodal fusion, including: An information acquisition module for obtaining a target liver CT image of a patient, a target radiological feature marker for the target liver CT image, and target clinical examination information corresponding to the patient; A feature extraction module for using a feature extraction module in a trained deep learning model to encode the target liver CT image to obtain an image encoding feature, encoding the target radiological feature marker to obtain a radiological encoding feature, and encoding the target clinical examination information to obtain an examination information encoding feature; A feature fusion module for using a multimodal feature fusion module in the trained deep learning model to fuse the image encoding feature, the radiological encoding feature, and the examination information encoding feature to obtain a cross-modal fusion feature; A classification module for using a classification head module in the trained deep learning model to classify the cross-modal fusion feature for liver cirrhosis outcome to obtain a liver cirrhosis outcome classification result.
[0008] In a third aspect, an embodiment of the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the patient information analysis method based on multimodal fusion described in the first aspect above.
[0009] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method for analyzing patient information based on multimodal fusion described in the first aspect above.
[0010] The beneficial effects of the present invention compared with the prior art are as follows: By obtaining the target liver CT image of the patient, the target radiological feature markers of the target liver CT image, and the target clinical examination information of the corresponding patient, using the feature extraction module in the trained deep learning model to encode the target liver CT image to obtain image encoding features, encoding the target radiological feature markers to obtain radiological encoding features, encoding the clinical examination information to obtain examination information encoding features, using the multimodal feature fusion module in the trained deep learning model to fuse the image encoding features, radiological encoding features, and examination information encoding features to obtain cross-modal fusion features, and using the classification head module in the trained deep learning model to classify the cross-modal fusion features for the liver cirrhosis outcome to obtain the liver cirrhosis outcome classification result. The outcome classification of the fusion features of the patient's liver information is performed through an automated fusion method, so as to more accurately predict the classification result of the liver cirrhosis outcome. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts.
[0012] Figure 1 It is a schematic diagram of the application environment of a method for analyzing patient information based on multimodal fusion provided in Embodiment 1 of the present invention; Figure 2 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 2 of the present invention; Figure 3 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 3 of the present invention; Figure 4 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 4 of the present invention; Figure 5 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 5 of the present invention; Figure 6 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 6 of the present invention; Figure 7 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in the seventh embodiment of the present invention; Figure 8 It is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in the eighth embodiment of the present invention; Figure 9 It is a schematic structural diagram of a device for analyzing patient information based on multimodal fusion provided in the ninth embodiment of the present invention; Figure 10 It is a schematic network structure diagram of a method for analyzing patient information based on multimodal fusion provided in the second embodiment of the present invention; Figure 11 It is a schematic structural diagram of a computer device provided in the tenth embodiment of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] As Figure 1 shown, it is a schematic diagram of an application environment of a method for analyzing patient information based on multimodal fusion provided in the first embodiment of the present invention. Among them, the client and the server are connected for communication. The user can provide conditions, requirements, operation instructions, etc. for analyzing patient information based on multimodal fusion to the server by operating the client. The server is used to execute the method for analyzing patient information based on multimodal fusion of the present invention according to the relevant content sent by the client. Among them, the client includes, but is not limited to, various computer devices such as personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The computer device corresponding to the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0015] As Figure 2 shown, it is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in the second embodiment of the present invention. Among them, the method for analyzing patient information based on multimodal fusion is applied to the Figure 1 server therein. The method for analyzing patient information based on multimodal fusion may include the following steps: Step S201, obtain the target liver CT image of the patient, the target radiological feature markers of the target liver CT image, and the target clinical examination information of the corresponding patient.
[0016] Among them, by using a CT device to scan the entire abdomen of a patient, a CT image of the entire abdomen of the patient is obtained. Based on the CT image of the entire abdomen, a CT image to be processed can be obtained. In this case, the target liver CT image of the patient is the CT image to be processed by the patient, including but not limited to, liver CT image, pancreas CT image, spleen CT image. Using a CT device to obtain an image can ensure that the image is clear, artifact-free, and includes the complete area of the abdomen, and the obtained CT image is preprocessed. For example, for the three-dimensional CT images of the entire abdomen of all patients, the image size is resampled to 64×512×512 by cubic spline interpolation, and the pixel values are normalized to the range [0, 1] using min-max normalization to unify the network input.
[0017] Among them, the target radiological features are the target radiomics features. The target radiological features include but are not limited to the radiological features of the CT images of the patient's liver and spleen. The target radiological features refer to a series of quantitative features extracted from medical image data, used to describe and analyze specific patterns, structures, and morphologies in the images. Image analysis software (such as 3D Slicer, ITK-SNAP, etc.) or automated image processing algorithms (such as deep learning models) can be used to detect and label the target radiological features. The radiological features that need to be labeled include but are not limited to, liver, tumor, cyst, vascular structure. For example, a junior radiologist uses 3D-Slicer software to semi-automatically segment the liver and spleen regions on the CT image, and a senior radiologist checks to determine the region of interest. Subsequently, the pyradiomics toolbox in Python is used to perform radiomics analysis on the region of interest and extract 1781 radiomics features.
[0018] Among them, the target clinical examination information of the patient includes clinical information (gender, age) and laboratory examination indicators (complete blood cell technology, biochemical indicators, etc.) that can be collected from the patient's electronic medical record. The Bone’s FRAX software is used for quantitative analysis of human body fat. The subcutaneous fat volume, visceral fat volume, intermuscular fat volume, visceral fat volume ratio, and intermuscular fat volume ratio are calculated at the level of the patient's L3 vertebra. The collected target clinical examination information will be integrated into a clinical text as the network input.
[0019] Step S202, use the feature extraction module in the trained deep learning model to encode the target liver CT image to obtain image encoding features, encode the target radiological feature labels to obtain radiological encoding features, and encode the target clinical examination information to obtain examination information encoding features.
[0020] Among them, using the feature extraction module in the trained deep learning model, specific modal features are extracted from the CT images by an independent encoder. All encoders are initialized with pre-trained parameters to avoid overfitting when training on private datasets. For example, 3D-ResNet50 is used to extract the image features of the global abdominal region in the whole abdominal CT images, a multi-layer perceptron (MLP) is used to process the target radiological feature markers, and BioBERT designed specifically for biomedical texts and tasks is used to encode the target clinical examination information.
[0021] Step S203: Use the multi-modal feature fusion module in the trained deep learning model to fuse the image encoded features, radiological encoded features, and examination information encoded features to obtain cross-modal fusion features.
[0022] Among them, the multi-modal information fusion module consists of an intra-modal enhancement module and a cross-modal interaction module. The intra-modal enhancement module uses the multi-head self-attention mechanism to enhance the intra-modal information, and the cross-modal interaction module uses the cross-modal attention mechanism to learn the complementary information among the image encoded features, radiological encoded features, and examination information encoded features. The encoded features are input into the multi-modal feature fusion module to obtain cross-modal fusion features. The obtained cross-modal fusion features can be used for various subsequent tasks, such as classification, regression, clustering, etc.
[0023] Step S204: Use the classification head module in the trained deep learning model to classify the cross-modal fusion features for the cirrhosis outcome and obtain the cirrhosis outcome classification result.
[0024] Among them, using the classification head module in the trained deep learning model, that is, the linear layer of the pre-trained 1D DenseNet-121 can be used as the classification head to predict the prognosis of cirrhosis, supplemented by the supervision of the cross-entropy loss function. The cross-modal fusion features are input into the classification head module to obtain the cirrhosis outcome classification result. According to the output of the model, the cirrhosis outcome classification result is interpreted. Assuming this is a binary classification task, a sigmoid activation function is used to output a probability value. In this way, using the classification head module in the trained deep learning model to classify the cross-modal fusion features for the cirrhosis outcome and obtain the cirrhosis outcome classification result, providing support for clinical diagnosis and treatment.
[0025] In the embodiment of the present application, by obtaining the target liver CT image of the patient, the target radiological feature markers of the target liver CT image, and the target clinical examination information of the corresponding patient, using the feature extraction module in the trained deep learning model, encoding the target liver CT image to obtain the image encoding feature, encoding the target radiological feature markers to obtain the radiological encoding feature, encoding the clinical examination information to obtain the examination information encoding feature, using the multi-modal feature fusion module in the trained deep learning model to fuse the image encoding feature, the radiological encoding feature, and the examination information encoding feature to obtain the cross-modal fusion feature, and using the classification head module in the trained deep learning model to classify the cross-modal fusion feature for the cirrhosis outcome to obtain the cirrhosis outcome classification result. By automatically fusing the fusion features of the patient's liver information, the classification result of the cirrhosis outcome can be predicted more accurately.
[0026] For example, Figure 10 as an example, Figure 10 FIG. is a schematic diagram of the network structure of a patient information analysis method based on multi-modal fusion provided in the second embodiment of the present application. Among them, the target liver CT image is input into the image encoder for encoding to obtain the image encoding feature, the target radiological feature markers are input into the radiomics encoder for encoding to obtain the radiological encoding feature, the target clinical examination information is input into the text encoder to obtain the examination information encoding feature, and each encoding feature is input into the intra-modal feature aggregation mechanism (IMAM) in the feature extraction module to enhance the encoding feature information, and the enhanced information is input into the three-modal cross-attention module to integrate various information. At the same time, through the three-modal feature fusion loss module, the loss of the three-modal feature fusion is calculated, and the deep learning model is improved according to the loss. The information integrated by the three-modal cross-attention module is classified into two categories through the classification head file, that is, the cirrhosis outcome is classified into two results: good and bad.
[0027] As Figure 3 shown, FIG. is a schematic flow chart of a patient information analysis method based on multi-modal fusion provided in the third embodiment of the present invention. The training process of the trained deep learning model in step S202 is as follows: Step S301, obtain a training set, where the training set includes the patient information of at least one patient, and the patient information includes the target liver CT image, the target radiological feature markers of the target liver CT image, the target clinical examination information, and the annotation result of the cirrhosis outcome.
[0028] 7]]Among them, the training set should include the following information of at least one patient: the target liver CT image, the radiological feature markers of the target liver CT image, the target clinical examination information, and the annotation result of the cirrhosis outcome, so as to provide data support for the subsequent training process.
[0029] Step S302: Input the target liver CT image, target radiological feature markers, and target clinical examination information in each patient's information into the feature extraction module in a preset deep learning model for encoding, to obtain the first feature corresponding to the target liver CT image, the second feature corresponding to the target radiological feature markers, and the third feature corresponding to the target clinical examination information.
[0030] Among them, three feature extraction modules are constructed, which are respectively used to extract target liver CT image features, target radiological feature markers, and target clinical examination information features. For example, a convolutional neural network (CNN) can be used to extract target liver CT image features to obtain the first feature, and fully connected layers can be used to extract target radiological feature markers and target clinical examination information features to obtain the second feature corresponding to the target radiological feature markers and the third feature corresponding to the target clinical examination information.
[0031] Step S303: Input the first feature, the second feature, and the third feature into the multi-modal feature fusion module in the preset deep learning model for fusion to obtain the first fusion feature, and calculate the fusion loss.
[0032] Among them, features from different sources often have different dimensions. For example, the first feature may be a high-dimensional feature vector, while the second feature and the third feature may be low-dimensional feature vectors from different modalities or different data sources. In order to effectively fuse these features together, the low-dimensional features need to be extended to the same dimension as the high-dimensional features through a feature alignment process. By performing linear or non-linear transformations on the low-dimensional features, more information can be introduced, thereby enhancing the expressiveness of the features. For example, using a fully connected layer (Dense layer) to expand a 256-dimensional feature to 512 dimensions helps the model better capture the potential information in these features, so that feature splicing and fusion can be effectively performed.
[0033] Step S: Input the first fusion feature into the classification head module in the preset deep learning model for liver cirrhosis outcome classification to obtain the first classification result, and calculate the classification task loss based on the first classification result and the annotation result.
[0034] Among them, the classification head module is usually composed of several fully connected layers, and finally outputs the classification result (for example, the probability in a binary classification task). Compile the model to use an appropriate loss function and optimizer. For a binary classification task, binary cross-entropy loss is usually used. Input the fused features and annotation results into the classification head module in the deep learning model for training. In each training batch or validation batch, the deep learning model will automatically calculate the loss. We can obtain the loss value during the training process, or use the validation set to calculate the classification task loss after training is completed. In each training batch, the deep learning model will calculate the classification loss based on the current prediction result and annotation result. Specifically, the deep learning model will use the loss function specified during compilation (for example, binary_crossentropy) to calculate the classification loss.
[0035] Step S305: Update the parameters in the preset deep learning model according to the fusion loss and the classification task loss to obtain an updated deep learning model.
[0036] Among them, the total loss function refers to the weighted sum of the fusion loss and the classification task loss. In deep learning frameworks (such as TensorFlow or PyTorch), backpropagation and parameter update are usually performed automatically. An optimizer (such as Adam) can be defined, and the minimize method of the optimizer is called in each training step. In the training loop, call a function to update the parameters in the preset deep learning model according to the fusion loss and the classification task loss to obtain an updated deep learning model.
[0037] Step S306: Use the updated deep learning model as the preset deep learning model, and return to execute the step of encoding the target liver CT image, the target radiological feature label, and the target clinical examination information in each patient information into the feature extraction module in the preset deep learning model until the total loss meets the preset condition, and the obtained updated deep learning model is the trained deep learning model.
[0038] Among them, use the above-updated deep learning model as the deep learning model again, and execute the above step of encoding the target liver CT image, the target radiological feature label, and the target clinical examination information in each patient information into the feature extraction module in the preset deep learning model respectively, so as to achieve the purpose of loop execution. Through this process, the deep learning model is self-looped and updated, so as to complete the training of the deep learning model.
[0039] In an embodiment of the present application, a training set is obtained to input patient information into a feature extraction module of a deep learning model for encoding, obtaining a first feature, a second feature, and a third feature, and fusing the first feature, the second feature, and the third feature in a multi-modal feature fusion module to obtain a first fusion feature, calculating a fusion loss, and classifying the liver cirrhosis outcome through a classification head module to obtain a first classification result. Calculate the classification task loss based on the first classification result and the annotation result, combine the fusion loss and the classification task loss to update the parameters in the deep learning model, obtain an updated deep learning model, and return the updated deep learning model to execute step S202 until the total loss meets the preset condition, and the obtained updated deep learning model is the trained deep learning model. Thus, by adjusting the model parameters, the value of the loss function is gradually reduced, enabling the model to gradually improve its prediction ability according to the training data.
[0040] As Figure 4 shown, it is a schematic flowchart of a method for analyzing patient information based on multi-modal fusion provided in Embodiment 4 of the present invention. The feature extraction module in step S302 includes an image encoder, a radiomics encoder, and a text encoder. The target liver CT image, the target radiological feature marker, and the target clinical examination information in each patient information are respectively input into the feature extraction module of a preset deep learning model for encoding, obtaining a first feature corresponding to the target liver CT image, a second feature corresponding to the target radiological feature marker, and a third feature corresponding to the target clinical examination information. This step may include the following steps: Step S401, input the target liver CT image into the image encoder for encoding, and output a first feature corresponding to the target liver CT image.
[0041] Among them, inputting the target liver CT image into the image encoder is to convert the target liver CT image into a high-level feature representation. The image encoder is usually a convolutional neural network (CNN), which can extract useful features from the image, such as edges, textures, shapes, etc.
[0042] Step S402, input the target radiological feature marker into the radiomics encoder for encoding, and output a second feature corresponding to the target radiological feature marker.
[0043] Among them, inputting the target radiological feature marker into the radiomics encoder is to convert the target radiological feature marker into a high-level feature representation. The target radiological features are usually some numerical features, such as the size, shape, density of the tumor, etc. The radiomics encoder can further abstract and integrate these features to generate a more informative feature representation.
[0044] Step S403: Input the target clinical examination information into the text encoder for encoding, and output the third feature corresponding to the target clinical examination information.
[0045] Among them, inputting the target clinical examination information into the text encoder is to convert the clinical examination information into a high-level feature representation. Clinical examination information is usually some text data, such as the patient's medical history, laboratory test results, vital signs, etc. The text encoder can process these text information to generate a more informative feature representation.
[0046] In this embodiment, by inputting the target liver CT image, the target radiomics feature label, and the target clinical examination information into the image encoder, the radiomics encoder, and the text encoder respectively for encoding, the first feature, the second feature, and the third feature are output, thus providing the basic data for the subsequent first feature fusion and the first classification result.
[0047] As Figure 5 shown, it is a schematic flowchart of a patient information analysis method based on multimodal fusion provided in Embodiment 5 of the present invention. The multimodal feature fusion module in step S303 includes a linear unit, a feature aggregation unit, and a fusion unit. Inputting the first feature, the second feature, and the third feature into the multimodal feature fusion module of the preset deep learning model in step S303 for fusion to obtain the first fusion feature and calculate the fusion loss may include the following steps: Step S501: Input the first feature, the second feature, and the third feature into the linear model for linear transformation respectively, to obtain the transformed first feature, the transformed second feature, and the transformed third feature.
[0048] Among them, after inputting the first feature, the second feature, and the third feature into the linear model, the transformed first feature, the transformed second feature, and the transformed third feature are output. The purpose of this step is to convert the original features into a more suitable form for fusion through linear transformation. The linear transformation can be implemented by a linear model (such as a fully connected layer), which maps the input features to a new space for subsequent feature fusion and aggregation.
[0049] Step S502: Input the transformed first feature, the transformed second feature, and the transformed third feature into the feature aggregation unit for intra-feature aggregation respectively, to obtain the aggregated first feature, the aggregated second feature, and the aggregated third feature.
[0050] Among them, the feature aggregation unit inputs the transformed first feature, the transformed second feature, and the transformed third feature, and outputs the aggregated first feature, the aggregated second feature, and the aggregated third feature. This step further integrates and optimizes each feature through intra-feature aggregation, making it more representative and informative for subsequent feature fusion. The feature aggregation unit usually adopts some common aggregation methods, such as attention mechanism, pooling operation, multi-layer perceptron (MLP), etc., for further processing and integration of features.
[0051] Step S503: Input the aggregated first feature, the aggregated second feature, and the aggregated third feature into the fusion unit for fusion to obtain the first fused feature.
[0052] Among them, the fusion unit inputs the aggregated first feature, the aggregated second feature, and the aggregated third feature and outputs the first fused feature. This step integrates features of different modalities through the fusion unit to generate a comprehensive, multi-modal feature representation, providing basic data for subsequent classification tasks. Feature fusion methods can include concatenation, weighted fusion, attention mechanism, and deep fusion network.
[0053] Among them, concatenation can refer to directly concatenating features of different modalities to form a longer feature vector. Weighted fusion can refer to weighted fusion of features of different modalities by learning the weights of each feature. The attention mechanism can refer to calculating the attention weights between features of different modalities to make the model pay more attention to important feature parts. The deep fusion network can refer to gradually fusing features of different modalities through a multi-layer network structure to generate the final fused feature.
[0054] In this embodiment, the first feature, the second feature, and the third feature are processed by a linear model and then sequentially input into the feature aggregation unit and the fusion unit to obtain the first fused feature, thereby providing more comprehensive information and better support for the final classification task.
[0055] Such as Figure 6 As shown, it is a schematic flowchart of a method for analyzing patient information based on multi-modal fusion provided in Embodiment 6 of the present invention. In step S502, inputting the transformed first feature, the transformed second feature, and the transformed third feature into the feature aggregation unit for intra-feature aggregation to obtain the aggregated first feature, the aggregated second feature, and the aggregated third feature may include the following steps: Step S601: The feature aggregation unit respectively extracts the self-modal feature representation, adjacent-modal feature representation, and enhanced feature representation of the first feature, the second feature, and the third feature.
[0056] Among them, the feature aggregation unit further integrates and optimizes features to extract self-modal feature representations, adjacent-modal feature representations, and enhanced feature representations, thereby generating more representative and informative features. The feature aggregation unit usually adopts some complex model structures. For example, multi-layer perceptrons (MLPs), attention mechanisms, graph neural networks, etc., are used to further process and integrate features.
[0057] The self-modal feature representation can refer to extracting the unique feature representation of this modality from the transformed features, retaining the main information of the original features. The adjacent-modal feature representation can refer to extracting the feature representation related to the adjacent modality from the transformed features, capturing the correlation information between different modalities. The enhanced feature representation can refer to further optimizing and integrating features through some enhancement operations (such as multi-layer perceptrons, attention mechanisms, etc.) to generate a more informative feature representation.
[0058] Step S602: Multiply the self-modal feature representation of the first feature by the transpose of the adjacent-modal feature representation of the first feature, and perform normalization processing through the softmax function to obtain the attention weight of the first feature.
[0059] Step S603: Multiply the self-modal feature representation of the second feature by the transpose of the adjacent-modal feature representation of the second feature, and perform normalization processing through the softmax function to obtain the attention weight of the second feature.
[0060] Step S604: Multiply the self-modal feature representation of the third feature by the transpose of the adjacent-modal feature representation of the third feature, and perform normalization processing through the softmax function to obtain the attention weight of the third feature.
[0061] Step S605: Multiply the attention weight of the first feature by the enhanced feature representation of the first feature to obtain the attention score of the first feature. Multiply the attention weight of the second feature by the enhanced feature representation of the second feature to obtain the attention score of the second feature. Multiply the attention weight of the third feature by the enhanced feature representation of the third feature to obtain the attention score of the third feature.
[0062] Among them, multiply the self-modal feature representation of the first feature by the transpose of the adjacent-modal feature representation of the first feature, and perform normalization processing through the softmax function to obtain the attention weight of the first feature. Perform the same operation process for the second feature and the third feature to obtain the attention weights of the second feature and the third feature respectively. Multiply the attention weight of the first feature by the enhanced feature representation of the first feature to obtain the enhanced representation of the first feature. Perform the same operation for the second feature and the third feature to obtain the attention scores of the second feature and the third feature.
[0063] In step S503, the aggregated first feature, the aggregated second feature, and the aggregated third feature are input into a fusion unit for fusion to obtain a first fusion feature, which may include the following steps: Step S606, obtain the initial features of the preset first feature, the second feature, and the third feature.
[0064] Step S607, perform combined calculations on the first feature attention score, the second feature attention score, and the third feature attention score with the initial features of the first feature, the second feature, and the third feature respectively through residual connections to obtain a combined result.
[0065] Step S608, obtain a preset learnable weight matrix, and perform a concatenation calculation on the combined result and the learnable weight matrix to obtain a first fusion feature.
[0066] Among them, through feature fusion, information integration, feature optimization, and multi-modal information capture can be achieved.
[0067] Among them, information integration may refer to fusing the aggregated multi-modal features to further integrate the information of each modality. Feature optimization may refer to retaining the information of the initial features through residual connections and simultaneously adding the optimization of attention scores to generate higher-quality feature representations. Multi-modal information capture may refer to capturing the correlation information between modalities by fusing features of different modalities to generate more representative and informative fusion features.
[0068] Among them, in a deep learning model, the main role of the learnable weight matrix is to be automatically adjusted through the learning process to optimize the performance of the model. The learnable weight matrix refers to the parameter matrix that is continuously updated through optimization algorithms such as gradient descent during the neural network training process. These weight matrices can automatically learn the optimal values during the model training process, enabling the model to better fit the data. Perform a matrix multiplication operation on the combined result and the learnable weight matrix to obtain the first fusion feature.
[0069] In this embodiment, the self-modal feature representation, adjacent-modal feature representation, and enhanced-modal feature representation of the first feature, the second feature, and the third feature are respectively extracted by the feature aggregation unit. Then, the self-modal feature representation of the first feature is multiplied by the transpose of the adjacent-modal representation of the first feature, and then normalized by the softmax function to obtain the first feature attention weight. For the second feature and the third feature, the second feature attention weight and the third feature attention weight are obtained through the same operation. The attention weight of the first feature is multiplied by the enhanced representation of the first feature to obtain the attention score of the first feature. The same operation is performed on the attention weights of the second feature and the third feature to obtain the attention scores of the second feature and the third feature. Through residual connection, the attention scores of the first feature, the second feature, and the third feature are respectively combined and calculated with the initial features of the first feature, the second feature, and the third feature to obtain a combined result, and the combined result is cascaded with a learnable weight matrix to obtain the first fusion feature. Thus, through a series of feature processing and fusion operations, the feature representation is gradually optimized to generate the final first fusion feature.
[0070] As Figure 7 shown, it is a schematic flowchart of a method for analyzing patient information based on multi-modal fusion provided in the seventh embodiment of the present invention. The calculation of the fusion loss in step S303 may include the following steps: Step S701, calculate the alignment loss between the first feature and the second feature based on the tri-modal feature fusion loss function to obtain the first alignment loss result.
[0071] Step S702, calculate the alignment loss between the first feature and the third feature based on the tri-modal feature fusion loss function to obtain the second alignment loss result.
[0072] Step S703, calculate the alignment loss between the second feature and the third feature based on the tri-modal feature fusion loss function to obtain the third alignment loss result.
[0073] Step S704, obtain a preset scalar weight coefficient, and calculate the fusion loss according to the scalar weight coefficient in combination with the first alignment loss result, the second alignment loss result, and the third alignment loss result to obtain the result of the fusion loss.
[0074] Among them, drawing on the idea of the similarity distribution matching (SDM) loss function, a tri-modal feature fusion loss function is used to align features of different modalities. Due to the asymmetry of the SDM loss between different modality pairs, and considering the homogeneity between images and radiomics features at the same time, the tri-modal feature fusion loss can be expressed by the mathematical formula: , where Represents the alignment loss between the first feature and the second feature, i.e., the first alignment loss result. Represents the alignment loss between the first feature and the third feature, i.e., the second alignment loss result. Represents the alignment loss between the second feature and the third feature, i.e., the third alignment loss result. Represents the weight coefficient in scalar form, which is set between [-1, 1] and determines the weight ratio of the alignment loss. Represents the fusion loss.
[0075] In this embodiment, the fusion loss is calculated by computing the first alignment loss result, the second alignment loss result, the third alignment loss result, and combining the weight coefficient in scalar form. By calculating the alignment loss between different features, the deep learning model can learn how to make these features more consistent or similar under a certain metric. For example, in multimodal learning, features of different modalities (such as images, text, audio) may have different representation forms, and calculating the alignment loss can help the model better understand the correlations between these modalities.
[0076] Such as [[ID=1s]] Figure 8 As shown, it is a schematic flowchart of a method for analyzing patient information based on multimodal fusion provided in Embodiment 8 of the present invention. In step S305, according to the fusion loss and the classification task loss, the parameters in the preset deep learning model are updated to obtain an updated deep learning model, which may include the following steps: Step S801, perform a weighted sum calculation based on the result of the fusion loss and the classification task loss to obtain the overall loss.
[0077] Among them, through the above calculation of the fusion loss and combining the classification task loss, the overall loss can be calculated, which can be expressed by the formula as , where represents the overall loss, represents the classification task loss, represents the scalar weight coefficient, represents the fusion loss.
[0078] Step S802, update the parameters in the preset deep learning model according to the overall loss to obtain an updated deep learning model.
[0079] Among them, backpropagation is performed using the overall loss to calculate the gradient of each parameter in the model. The gradient represents the rate of change of the loss function with respect to each parameter. Before each backpropagation, it is usually necessary to clear the previous gradient to avoid gradient accumulation. Calculate the gradient of the overall loss, and use an optimizer to update the parameters of the model according to the calculated gradient. After the parameter update, the parameters of the deep learning model will be adjusted to obtain an updated deep learning model. This updated deep learning model will be used in subsequent training and validation processes in the hope of performing better on the test data.
[0080] In this embodiment, by combining and calculating the weighted sum of the fusion loss and the classification task loss, the overall loss is obtained, and the parameters of the deep learning model are updated to obtain an updated deep learning model. By calculating the overall loss and performing backpropagation, the model can adjust its parameters according to the gradient of the loss function, enabling the model to gradually reduce the loss and thus perform better on the training data.
[0081] As Figure 9 shown, it is a schematic diagram of a patient information analysis device based on multimodal fusion provided in Embodiment 9 of the present invention. The patient information analysis device based on multimodal fusion corresponds one-to-one with the patient information analysis method based on multimodal fusion in the above embodiment. The patient information analysis device based on multimodal fusion includes an information acquisition module 91, a feature extraction module 92, a feature fusion module 93, and a classification module 94. The detailed description of each functional module is as follows: The information acquisition module 91 is used to acquire the target liver CT image of the patient, the target radiological feature markers of the target liver CT image, and the target clinical examination information of the corresponding patient; The feature extraction module 92 is used to use the feature extraction module in the trained deep learning model to encode the target liver CT image to obtain image encoding features, encode the target radiological feature markers to obtain radiological encoding features, and encode the target clinical examination information to obtain examination information encoding features; The feature fusion module 93 is used to use the multimodal feature fusion module in the trained deep learning model to fuse the image encoding features, radiological encoding features, and examination information encoding features to obtain cross-modal fusion features; The classification module 94 is used to use the classification head module in the trained deep learning model to classify the cross-modal fusion features for liver cirrhosis outcome to obtain a liver cirrhosis outcome classification result.
[0082] Optionally, the above feature extraction module 92 includes: A training set acquisition unit for acquiring a training set, which includes patient information of at least one patient. The patient information includes a target liver CT image, a target radiological feature marker for the target liver CT image, target clinical examination information, and an annotation result for the cirrhosis outcome.
[0083] An encoding unit for separately inputting the target liver CT image, the target radiological feature marker, and the target clinical examination information in each patient information into a feature extraction module in a preset deep learning model for encoding, to obtain a first feature corresponding to the target liver CT image, a second feature corresponding to the target radiological feature marker, and a third feature corresponding to the target clinical examination information.
[0084] A fusion unit for inputting the first feature, the second feature, and the third feature into a multi-modal feature fusion module in a preset deep learning model for fusion, to obtain a first fusion feature and calculate a fusion loss.
[0085] An outcome classification unit for inputting the first fusion feature into a classification head module in a preset deep learning model for cirrhosis outcome classification, to obtain a first classification result, and calculating a classification task loss according to the first classification result and the annotation result.
[0086] A parameter update unit for updating the parameters in the preset deep learning model according to the fusion loss and the classification task loss, to obtain an updated deep learning model.
[0087] A loop execution unit for using the updated deep learning model as the preset deep learning model, and returning to execute the step of separately inputting the target liver CT image, the target radiological feature marker, and the target clinical examination information in each patient information into the feature extraction module in the preset deep learning model for encoding, until the updated deep learning model obtained when the total loss meets the preset condition is the trained deep learning model.
[0088] Optionally, the above encoding unit includes: A first feature encoding subunit for inputting the target liver CT image into an image encoder for encoding, and outputting a first feature corresponding to the target liver CT image.
[0089] A second feature encoding subunit for inputting the target radiological feature marker into a radiomics encoder for encoding, and outputting a second feature corresponding to the target radiological feature marker.
[0090] A third feature encoding subunit for inputting the target clinical examination information into a text encoder for encoding, and outputting a third feature corresponding to the target clinical examination information.
[0091] Optionally, the above fusion unit includes: A linear transformation subunit, configured to respectively input the first feature, the second feature, and the third feature into a linear model for linear transformation, so as to obtain the transformed first feature, the transformed second feature, and the transformed third feature.
[0092] An aggregation subunit, configured to respectively input the transformed first feature, the transformed second feature, and the transformed third feature into a feature aggregation unit for intra-feature aggregation, so as to obtain the aggregated first feature, the aggregated second feature, and the aggregated third feature.
[0093] A first fusion feature subunit, configured to input the aggregated first feature, the aggregated second feature, and the aggregated third feature into a fusion unit for fusion, so as to obtain a first fusion feature.
[0094] Optionally, the above-mentioned aggregation subunit includes: A feature representation extraction subt subunit, configured to respectively extract the self-modal feature representation, the adjacent-modal feature representation, and the enhanced feature representation of the first feature, the second feature, and the third feature through a feature aggregation unit.
[0095] A first feature attention weight subt subunit, which performs a dot product operation on the feature vector of the first feature in its own modality and the transposed feature matrix of the adjacent modality, then normalizes the calculation result through a softmax function, and finally generates the attention weight of the first feature.
[0096] A second feature attention weight subt subunit, which performs a dot product operation on the feature vector of the second feature in its own modality and the transposed feature matrix of the adjacent modality, then normalizes the calculation result through a softmax function, and finally generates the attention weight of the second feature.
[0097] A third feature attention weight subt subunit, which performs a dot product operation on the feature vector of the third feature in its own modality and the transposed feature matrix of the adjacent modality, then normalizes the calculation result through a softmax function, and finally generates the attention weight of the third feature.
[0098] An attention score calculation subt subunit, configured to multiply the attention weight of the first feature by the enhanced feature representation of the first feature to obtain the attention score of the first feature, multiply the attention weight of the second feature by the enhanced feature representation of the second feature to obtain the attention score of the second feature, and multiply the attention weight of the third feature by the enhanced feature representation of the third feature to obtain the attention score of the third feature.
[0099] Optionally, the above-mentioned first fusion feature subunit includes: An initial feature acquisition subt subunit, configured to acquire the initial features of the preset first feature, the second feature, and the third feature.
[0100] The combined result calculation sub-unit is used to perform combined calculations on the initial features of the first feature, the initial features of the second feature, and the initial features of the third feature respectively with the first feature attention score, the second feature attention score, and the third feature attention score through residual connections to obtain a combined result.
[0101] The concatenation calculation sub-unit is used to obtain a preset learnable weight matrix and perform concatenation calculations on the combined result and the learnable weight matrix to obtain a first fused feature.
[0102] Optionally, the above fusion unit further includes: The first alignment loss calculation sub-unit is used to calculate the alignment loss between the first feature and the second feature based on the three-modal feature fusion loss function to obtain a first alignment loss result.
[0103] The second alignment loss calculation sub-unit is used to calculate the alignment loss between the first feature and the third feature based on the three-modal feature fusion loss function to obtain a second alignment loss result.
[0104] The third alignment loss calculation sub-unit calculates the alignment loss between the second feature and the third feature based on the three-modal feature fusion loss function to obtain a third alignment loss result.
[0105] The fusion loss calculation sub-unit is used to obtain a preset scalar weight coefficient and perform fusion loss calculations based on the scalar weight coefficient in combination with the first alignment loss result, the second alignment loss result, and the third alignment loss result to obtain a fusion loss result.
[0106] Optionally, the above parameter update unit includes: The overall loss calculation sub-unit is used to perform weighted sum calculations on the fusion loss result and the classification task loss to obtain an overall loss.
[0107] The overall loss parameter update sub-unit is used to update the parameters in a preset deep learning model according to the overall loss to obtain an updated deep learning model.
[0108] For the specific limitations of the patient information analysis device based on multi-modal fusion, reference can be made to the limitations of the patient information analysis method based on multi-modal fusion in the above text, which will not be elaborated here. Each module in the above patient information analysis device based on multi-modal fusion can be implemented in whole or in part through software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above modules.
[0109] Such as Figure 11As shown in the figure, it is a schematic structural diagram of a computer device provided in the tenth embodiment of the present invention. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for analyzing patient information based on multimodal fusion.
[0110] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for analyzing patient information based on multimodal fusion in the above embodiment, for example Figures 2 to 8 As shown in the figure, to avoid repetition, it will not be elaborated here. Or, when the processor executes the computer program, it implements the functions of each module / unit in the embodiment of the evaluation device for project R & D efficiency, for example Figure 9 the function of the information acquisition module 91 shown in the figure. To avoid repetition, it will not be elaborated here.
[0111] In one embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the method for analyzing patient information based on multimodal fusion in the above embodiment, such as Figures 2 to 8 As shown in the figure, to avoid repetition, it will not be elaborated here. Or, when the computer program is executed by the processor, it implements the functions of each module / unit in the above embodiment of the device for analyzing patient information based on multimodal fusion, for example Figure 9 the function of the feature fusion module 93 shown in the figure. To avoid repetition, it will not be elaborated here. The computer-readable storage medium can be non-volatile or volatile.
[0112] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc. [[ID=I]]
[0113] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0114] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.
Claims
1. A patient information analysis method based on multimodal fusion, characterized in that, The patient information analysis method includes: obtaining a target liver CT image of a patient, target radiological feature markings for the target liver CT image, and target clinical examination information corresponding to the patient; using a feature extraction module in a trained deep learning model to encode the target liver CT image to obtain an image encoding feature, encoding the target radiological feature markings to obtain a radiological encoding feature, and encoding the target clinical examination information to obtain an examination information encoding feature; using a multi-modal feature fusion module in the trained deep learning model to fuse the image encoding feature, the radiological encoding feature, and the examination information encoding feature to obtain a cross-modal fusion feature; using a classification head module in the trained deep learning model to classify the cross-modal fusion feature for liver cirrhosis outcome to obtain a liver cirrhosis outcome classification result.
2. The patient information analysis method based on multimodal fusion according to claim 1, wherein The training process of the trained deep learning model is as follows: obtaining a training set, where the training set includes patient information of at least one patient, and the patient information includes a target liver CT image, target radiological feature markings for the target liver CT image, target clinical examination information, and a labeled result for the liver cirrhosis outcome; respectively inputting the target liver CT image, the target radiological feature markings, and the target clinical examination information in each patient information into a feature extraction module in a preset deep learning model for encoding to obtain a first feature corresponding to the target liver CT image, a second feature corresponding to the target radiological feature markings, and a third feature corresponding to the target clinical examination information; inputting the first feature, the second feature, and the third feature into a multi-modal feature fusion module in the preset deep learning model for fusion to obtain a first fusion feature, and calculating a fusion loss; inputting the first fusion feature into a classification head module in the preset deep learning model for liver cirrhosis outcome classification to obtain a first classification result, and calculating a classification task loss according to the first classification result and the labeled result; updating parameters in the preset deep learning model according to the fusion loss and the classification task loss to obtain an updated deep learning model; using the updated deep learning model as the preset deep learning model, and returning to execute the step of respectively inputting the target liver CT image, the target radiological feature markings, and the target clinical examination information in each patient information into a feature extraction module in the preset deep learning model for encoding until the updated deep learning model obtained when the total loss meets a preset condition is the trained deep learning model.
3. The method for analyzing patient information based on multimodal fusion according to claim 2, wherein The feature extraction module includes an image encoder, a radiomics encoder, and a text encoder. Encoding the target liver CT image, the target radiological feature markers, and the target clinical examination information in each patient information into the feature extraction module of a preset deep learning model respectively to obtain a first feature corresponding to the target liver CT image, a second feature corresponding to the target radiological feature markers, and a third feature corresponding to the target clinical examination information includes: Inputting the target liver CT image into the image encoder for encoding, and outputting a first feature corresponding to the target liver CT image; Inputting the target radiological feature markers into the radiomics encoder for encoding, and outputting a second feature corresponding to the target radiological feature markers; Inputting the target clinical examination information into the text encoder for encoding, and outputting a third feature corresponding to the target clinical examination information.
4. The patient information analysis method based on multimodal fusion according to claim 2, wherein The multi-modal feature fusion module includes a linear unit, a feature aggregation unit, and a fusion unit. Fusing the first feature, the second feature, and the third feature into the multi-modal feature fusion module of the preset deep learning model to obtain a first fusion feature, and calculating a fusion loss, includes: Inputting the first feature, the second feature, and the third feature into the linear model respectively for linear transformation to obtain a transformed first feature, a transformed second feature, and a transformed third feature; Inputting the transformed first feature, the transformed second feature, and the transformed third feature into the feature aggregation unit respectively for intra-modal feature aggregation to obtain an aggregated first feature, an aggregated second feature, and an aggregated third feature; Inputting the aggregated first feature, the aggregated second feature, and the aggregated third feature into the fusion unit for fusion to obtain the first fusion feature.
5. The method for analyzing patient information based on multimodal fusion according to claim 4, wherein The step of inputting the transformed first feature, the transformed second feature, and the transformed third feature into the feature aggregation unit respectively for intra-modal feature aggregation to obtain an aggregated first feature, an aggregated second feature, and an aggregated third feature includes: Respectively extracting the self-modal feature representation, adjacent-modal feature representation, and enhanced feature representation of the first feature, the second feature, and the third feature through the feature aggregation unit; Multiplying the self-modal feature representation of the first feature by the transpose of the adjacent-modal feature representation of the first feature, and performing normalization processing through the softmax function to obtain the attention weight of the first feature; Multiplying the self-modal feature representation of the second feature by the transpose of the adjacent-modal feature representation of the second feature, and performing normalization processing through the softmax function to obtain the attention weight of the second feature; Multiplying the self-modal feature representation of the third feature by the transpose of the adjacent-modal feature representation of the third feature, and performing normalization processing through the softmax function to obtain the attention weight of the third feature. Multiply the attention weight of the first feature by the enhanced feature representation of the first feature to obtain the attention score of the first feature, multiply the attention weight of the second feature by the enhanced feature representation of the second feature to obtain the attention score of the second feature, and multiply the attention weight of the third feature by the enhanced feature representation of the third feature to obtain the attention score of the third feature; The step of inputting the aggregated first feature, the aggregated second feature, and the aggregated third feature into the fusion unit for fusion to obtain the first fusion feature includes: Obtain the initial features of the first feature, the second feature, and the third feature preset; Perform combined calculations on the first feature attention score, the second feature attention score, and the third feature attention score with the initial features of the first feature, the second feature, and the third feature respectively through residual connections to obtain a combined result; Obtain a preset learnable weight matrix, and perform a concatenation calculation on the combined result and the learnable weight matrix to obtain the first fusion feature.
6. The method for analyzing patient information based on multimodal fusion according to claim 2, wherein The calculation of the fusion loss includes: Calculate the alignment loss between the first feature and the second feature based on the three-modal feature fusion loss function to obtain the first alignment loss result; Calculate the alignment loss between the first feature and the third feature based on the three-modal feature fusion loss function to obtain the second alignment loss result; Calculate the alignment loss between the second feature and the third feature based on the three-modal feature fusion loss function to obtain the third alignment loss result; Obtain a preset scalar weight coefficient, and calculate the fusion loss according to the scalar weight coefficient in combination with the first alignment loss result, the second alignment loss result, and the third alignment loss result to obtain the result of the fusion loss.
7. The method for analyzing patient information based on multimodal fusion according to claim 6, wherein The step of updating the parameters in the preset deep learning model according to the fusion loss and the classification task loss to obtain an updated deep learning model includes: Perform a weighted sum calculation on the result of the fusion loss and the classification task loss to obtain an overall loss; Update the parameters in the preset deep learning model according to the overall loss to obtain an updated deep learning model.
8. A patient information analysis device based on multimodal fusion, characterized in that, It includes: An information acquisition module for acquiring a target liver CT image of a patient, target radiological feature markers for the target liver CT image, and target clinical examination information corresponding to the patient; A feature extraction module for encoding the target liver CT image using the feature extraction module in the trained deep learning model to obtain an image encoding feature, encoding the target radiological feature markers to obtain a radiological encoding feature, and encoding the target clinical examination information to obtain an examination information encoding feature; A feature fusion module for fusing the image encoding feature, the radiological encoding feature, and the examination information encoding feature using the multi-modal feature fusion module in the trained deep learning model to obtain a cross-modal fusion feature; A classification module, which is used to classify the cross-modal fusion features for liver cirrhosis outcomes by using the classification head module in the trained deep learning model, so as to obtain a liver cirrhosis outcome classification result.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multi-modal fusion-based patient information analysis method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the multi-modal fusion-based patient information analysis method according to any one of claims 1 to 7.