Multi-mode heart failure early screening system based on large model

The large-scale multimodal heart failure early screening system solves the problems of information integration and semantic expression in multimodal heart failure data fusion, realizes early screening and risk prediction of heart failure, and improves the personalization and accuracy of diagnosis and treatment.

CN121260433APending Publication Date: 2026-01-02ZHEJIANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511181520.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies lack efficient information integration methods when fusing multimodal heart failure data, fail to fully model the dependencies and complementarities between modalities, lack high-level interaction mechanisms, and cannot achieve unified semantic expression and personalized and generalized capabilities for heterogeneous data.

Method used

A large-scale model is used for a multimodal early heart failure screening system, including data acquisition and processing, summary generation, multimodal feature extraction, feature measurement and alignment modules. Features are extracted through table encoder, text encoder and deep neural network encoder, and the alignment and fusion of multimodal features are achieved by using bidirectional cross-attention mechanism and adversarial learning strategy.

Benefits of technology

It improves the accuracy of heart failure risk assessment, enables early screening and risk prediction of heart failure, provides personalized support for intelligent assisted diagnosis and treatment, and enhances the complementarity and fusion expression capabilities of multimodal features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121260433A_ABST
    Figure CN121260433A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode heart failure early screening system based on a large model, and the system comprises a data collection and processing module which is used for collecting forms, clinical record texts and electrocardiosignals, related to heart failure, of a patient; the abstract generation module is used for extracting text abstracts from clinical record texts by using a large model and introducing heart failure knowledge constraints; the multi-modal feature extraction module is used for extracting pathological index change features, semantic features and electrocardio waveform change features from tables, clinical record texts and electrocardio signals; the multi-modal feature measurement module is used for obtaining a weight coefficient of each modal based on the multi-modal data; the multi-modal feature alignment module maps the weighted modal features to a shared representation space through a learnable projection layer, and realizes feature fusion through a bidirectional cross attention mechanism and an adversarial learning strategy; and the heart failure early screening module is used for obtaining a heart failure disease early screening prediction result and a risk probability through feature fusion. The heart failure prediction accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of early screening and prediction of heart failure, and in particular to a multi-modal heart failure early screening system based on a large model. BACKGROUND

[0002] Heart failure is a serious chronic cardiovascular disease, and its early identification is of great significance for delaying disease progression, developing intervention strategies and reducing medical burden. Currently, doctors often rely on structured electronic medical records, test results or single type of physiological signals for judgment, and the diagnosis process is complex and easily influenced by subjective experience. With the development of medical informatization, patients will generate multi-modal data such as electronic medical record texts, physiological signals, laboratory indicators and statistical data during diagnosis and treatment, which contains rich pathological information. However, existing methods still have the following problems when facing multi-modal data fusion: (1) The information integration method is relatively rough, and the dependence and complementary relationship between modalities is not fully modeled; (2) There is a lack of efficient alignment mechanism, making it difficult to achieve unified semantic expression of heterogeneous data; (3) The modeling ability of patient individual characteristics and potential health risks is limited, and there is a lack of interpretability and generalization ability.

[0003] Therefore, an intelligent system capable of fusing multi-modal medical data, having good feature expression and knowledge reasoning ability is urgently needed to support early identification and risk prediction of chronic diseases such as heart failure.

[0004] Chinese patent document with publication number CN116451068A discloses a multi-modal data fusion heart failure diagnosis assistance method, including a feature extraction module, a feature encoding module and a classification module. The feature extraction module extracts features from collected electrocardiogram, chest X-ray and text information to generate vectors. The feature encoding module adds the generated vectors to the respective relative position encodings, and inputs the processed data into an encoder network for training. The classification module classifies heart failure, grades the severity, and predicts readmission and mortality rates after patient discharge.

[0005] The Chinese patent document with publication number CN116269426A discloses a twelve-lead ECG auxiliary heart disease multi-modal fusion screening method, which includes a data acquisition and preprocessing module, an electrocardio-ultrasound network training module, an electrocardiogram network training module, and a feature comparison module. The data acquisition and preprocessing module acquires and pre-processes the cardiac ultrasound data and corresponding 12-lead ECG electrocardiogram data of the patient to be detected. The electrocardio-ultrasound network training module is constructed based on the cardiac ultrasound data of the patient to be detected and a cardiac ultrasound LSTM neural network is obtained by training. The electrocardiogram network training module is constructed based on the 12-lead ECG electrocardiogram data of the patient to be detected and an ECG LSTM neural network is obtained by training. The feature comparison module compares the patient features output by the patient feature multi-layer perception with a plurality of example features respectively, and obtains an example feature with the smallest feature difference value.

[0006] The Chinese patent document with publication number CN116469553A discloses a multi-modal heart failure prediction auxiliary method based on an LSTM model and a ResNet50 model, which includes a data preprocessing module, a feature encoding and fusion module, and a diagnosis result prediction module. The data preprocessing module includes preprocessing the data of the patient's electronic health record and the chest X-ray: using ResNet50 to process the chest X-ray and generate a radiology report, and using the LSTM model to process the data in the electronic health record. The data is extracted and a report is generated. The feature encoding and fusion module fuses the extracted information in the chest X-ray and the electronic health record using a text-image embedding network. The diagnosis result prediction module predicts and determines the final diagnosis result.

[0007] The above-mentioned patents generally have a relatively simple multi-modal fusion method, lack higher-level interaction mechanisms, and cannot fully explore the deep correlations between modalities. Secondly, there is a lack of effective alignment mechanism between modal features. Different modalities are not considered in terms of semantic level timing alignment and common feature extraction, which limits the ability of cross-modal collaborative modeling. And without the help of large models in medical knowledge modeling and explanatory aspects, it cannot realize the high personalization and generalization ability of intelligent auxiliary diagnosis and treatment. SUMMARY

[0008] The present application provides a multi-modal heart failure early screening system based on a large model, which can improve the prediction accuracy of heart failure and realize the preliminary screening of heart failure risk.

[0009] A multi-modal heart failure early screening system based on a large model, comprising: A data acquisition and processing module is used to acquire and preprocess three modal data of EHR forms, clinical record texts, and electrocardio signals related to heart failure during the patient's hospitalization period. An abstract generation module uses a heart failure disease knowledge graph as a prompt and a constraint condition, uses a large language model to extract key information closely related to heart failure risk from clinical record text, and generates a text abstract. A multi-modal feature extraction module extracts three modal features through a table encoder, a text encoder, and a deep neural network encoder. The table encoder extracts pathological index change features from the EHR table. The text encoder extracts semantic features of heart failure symptom descriptions from the text abstract. The deep neural network encoder extracts electrocardiogram waveform change features from the electrocardiogram signal. A multi-modal feature measurement module is used to evaluate the quality of the three modal data and obtain modal weight coefficients. A multi-modal feature alignment module weights the three modal features based on the modal weight coefficients and maps the weighted modal features to a shared representation space through a learnable projection layer. Then, through a bidirectional cross-attention mechanism and an adversarial learning strategy, the alignment and fusion of the multi-modal features are realized, and the final fusion features are obtained. A heart failure early screening module inputs the fusion features into a deep learning classification network to output heart failure disease early screening prediction results and risk probabilities.

[0010] In the data acquisition and processing module, the EHR table includes laboratory indicators reflecting heart function status, medication information, and vital signs. The clinical record text includes disease history records describing the evolution of heart failure symptoms, admission records, and discharge summaries. The electrocardiogram signal includes twelve-lead electrocardiogram.

[0011] In the data acquisition and processing module, the preprocessing includes data denoising and missing value interpolation.

[0012] In the abstract generation module, a large language model fine-tuned in the medical field is used to extract the main content related to diagnosis information, past medical history, key test indicators, medication regimen, and treatment effect from the clinical record text under the guidance of the set medical prompt words, and generate a text abstract. The heart failure specific knowledge graph is introduced as a prompt and a constraint condition to cover the core clinical features of the generated abstract.

[0013] In the multi-modal feature extraction module, for EHR table data , there are: ; wherein, is a table encoder that encodes the EHR table data into a feature representation vector , i.e., pathological index change features. For text abstract data , there are: ; wherein, is a text encoder, which encodes the text summary data and maps it into a feature representation vector , i.e., semantic features. For electrocardio signal data , we have: ; wherein, is a signal encoder, which encodes the electrocardio signal data and maps it into a feature representation vector , i.e., electrocardio waveform variation features.

[0014] In the multi-modal feature measurement module, the quality of the multi-modal data is evaluated, and the modal weight coefficient is obtained, and the calculation method is as follows: For EHR table data , the feature integrity index (missing rate) and the average time continuity index between the adjacent two EHR records are calculated: ; ; wherein, is the number of clinical examinations performed irregularly during the patient's hospitalization; represents the th EHR examination.

[0015] The quality score of the table modal is calculated by weighting , wherein , are learnable coefficients.

[0016] ; The weight coefficient of the table modal data is calculated: ; For text summary data , a pre-defined medical entity set is carried out to carry out a medical named entity recognition task to extract an entity set , wherein K is the number of non-repeated entities in the medical entity set, and M is the number of extracted entity sets. The proportion of medical entities appearing in the text is calculated : ; wherein is the number of medical entities covered in the text, is the total number of entities in the knowledge base.

[0017] The weight coefficient of the text modal data is calculated ; signal data , the signal-to-noise ratio is used for quality quantization. The weight coefficient of the signal modal data is calculated.

[0018] In the multi-modal feature alignment module, the three modal features are weighted based on the modal weight coefficient, specifically: ; ; ; Among them, is the weighted pathological index change feature, is the weighted semantic feature, is the weighted electrocardiogram waveform change feature.

[0019] In the multi-modal feature alignment module, the weighted modal features are mapped to the shared representation space through a learnable projection layer, specifically: A set of trainable linear projection parameters 、 、 are designed for each modal feature respectively, and after mapping, they are uniformly projected to the shared space : ; ; ; Among them, denotes the projected and mapped pathological index change feature, is the projected and mapped semantic feature, is the projected and mapped electrocardiogram waveform change feature, , , , which ensures that different modal features have the same dimension and spatial basis.

[0020] In the multi-modal feature alignment module, the working process of the bidirectional cross-attention mechanism is as follows: The three single-modal features 、 、 are respectively calculated by two-way cross-attention, obtaining: the cross-attention vector between text and table 、 , the cross-attention vector between text and signal 、 , and the cross-attention vector between table and signal , ; the obtained six cross attention vectors are fused to obtain a fused feature .

[0021] In the multi-modal feature alignment module, the working process of the adversarial learning strategy is as follows: A discriminator is introduced to discriminate the fused feature from each single-modal feature representation , , Discrimination training is performed, so as to prompt the fused feature to have modal indistinguishability while retaining the original modal information, thereby realizing the promotion of modal collaboration and shared representation.

[0022] Compared with the prior art, the present application has the following beneficial effects: The present application can fully fuse the multi-source heterogeneous medical data such as text, table and physiological signal collected during hospitalization, and summarize and extract the clinical long text by combining a large model, while introducing heart failure knowledge constraints to improve the availability and professionalism of important diagnosis and treatment information. By designing an independent encoder network for each modality and designing a multi-modal data quality evaluation weighting mechanism, introducing a learnable projection layer, a bidirectional attention mechanism and an adversarial training strategy in the multi-modal feature alignment module, the problem of inconsistent feature spaces and semantic misplacement between different modalities is effectively alleviated, and the complementarity and fusion expression ability of multi-modal features are enhanced. In addition, the unified representation after fusion is input into a deep classification model for heart failure early screening prediction, which can significantly improve the recognition accuracy of individuals with high risk of heart failure, provide intelligent support for clinical auxiliary decision-making, and realize the transformation of complex medical data into interpretable early screening results. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 A structure diagram of a multi-modal heart failure early screening system based on a large model is provided.

[0024] Figure 2 A prediction result comparison graph in an embodiment of the present application is provided. DETAILED DESCRIPTION

[0025] The present application will be further described in detail below in conjunction with the drawings and embodiments, and it should be pointed out that the following embodiments are intended to facilitate the understanding of the present application and do not limit the same.

[0026] As shown in Figure 1 , a multi-modal heart failure early screening system based on a large model includes a data acquisition and processing module, a long text information summarization module and a heart failure screening model, wherein the heart failure screening model comprises a multi-modal feature extraction module, a multi-modal feature alignment module and a heart failure early screening module.

[0027] A data collection and preprocessing module is configured to collect multi-modal data related to heart failure during a patient's hospitalization, including table data reflecting heart function status in laboratory indicators, medication information, vital signs, and the like in electronic health records (EHR); clinical long texts containing descriptions of the evolution of heart failure symptoms, such as hospitalization records and discharge summaries; and physiological signal data such as multi-channel electrocardiogram signals (such as twelve-lead electrocardiogram), and perform structured and denoising preprocessing operations, and process through a standardized preprocessing process.

[0028] An abstract generation module based on heart failure knowledge constraints is configured to use an artificially constructed instruction prompt word template to extract structured medical knowledge such as heart failure symptoms, heart function classification, past medical history, and treatment plan from unstructured medical record texts by means of a large language model. Under the guidance of the designed prompt word, a heart failure specific knowledge graph is introduced as a prompt and constraint condition, and a large language model is used to automatically extract and summarize symptoms, diagnoses, and treatment plans from clinical long texts, and to extract main content closely related to heart failure risk, such as dyspnea, fluid retention, and heart function classification, to form a structured text abstract.

[0029] A heart failure screening model is established for preprocessed and summarized multi-modal medical data, and the model is composed of a multi-modal feature extraction module, a multi-modal feature measurement module, a multi-modal feature alignment module, and a heart failure early screening module.

[0030] A multi-modal feature extraction module is configured to construct a special encoder for each of the three modalities for deep semantic representation learning: a table encoder is constructed for EHR structured data to extract pathological indicator trends; a text encoder is constructed for text summary sentences to extract semantic features of heart failure symptom descriptions; and a deep neural network encoder is constructed for electrocardiogram signals to extract electrocardiogram waveform features such as QRS complex morphology and heart rate variability.

[0031] A multi-modal feature measurement module is configured to evaluate the quality of multi-modal data and obtain modality weight coefficients: a feature integrity index (missing rate) and an average time continuity index are calculated for EHR structured data; the proportion of medical entities appearing in the text is calculated for text summary sentences; and a signal-to-noise ratio is used to quantify the quality of electrocardiogram signals. The modality features are multiplied by the calculated weight coefficients to obtain weighted modality features.

[0032] The multi-modal feature alignment module adopts modal projection, fusion attention and adversarial learning mechanism. It is used for alignment between different modal features. Firstly, the three types of modal features are mapped to the same semantic space. Secondly, a bidirectional cross-attention mechanism between modalities is constructed to enhance the information between tables and texts, texts and signals, and signals and tables. Finally, an adversarial learning mechanism is introduced to make the fused features unidentifiable by constructing a modal discriminator, so as to realize more robust multi-modal alignment. The representation difference between modalities is effectively eliminated, and the perception ability of the model to early heart failure is enhanced.

[0033] The heart failure early screening module is used for fusing the aligned multi-modal features, and a lightweight classifier is used to realize the prediction of individual heart failure risk. The output results include whether there is a heart failure risk and a risk score. The fused multi-modal shared representation is input into a deep classification model to realize the prediction of heart failure disease risk and risk level evaluation. The system can assist doctors in identifying potential heart failure patients in the early stage of clinical treatment, and improve the efficiency of diagnosis and treatment and the timing of disease intervention.

[0034] In the embodiment of the application, the data collection is the multi-modal data of the patient, including the electronic health record , clinical record text T, multi-channel electrocardiogram physiological signal . Among them is the number of clinical examinations made irregularly during the patient's hospitalization, corresponding to electronic health records; ts is the number of electrocardiogram examinations made irregularly during the patient's hospitalization. Among them, for any , there is , that is, the length of the collection is fixed, including sampling points of c channels.

[0035] The data preprocessing method is a comprehensive processing method according to the characteristics of different modal data, including data denoising and missing value interpolation. For the electrocardiogram physiological signal of the patient, the band-pass filtering and wavelet transform method is used to remove the baseline drift, power interference and high-frequency noise, and to retain the key waveform features of the electrocardiogram signal such as P wave, QRS complex and T wave, so as to improve the accuracy of subsequent signal analysis; for the missing items that may exist in the EHR table data, the median filling and interpolation filling method based on statistical method is used to complete the missing values, so as to ensure the integrity and time sequence consistency of the table features, and provide reliable input for downstream modeling.

[0036] The abstract generation module based on heart failure knowledge constraint is as follows Figure 1The method is shown, which is a method for semantic extraction and content condensation of long text of patient's clinical records (including medical history records, admission records, discharge summaries, etc.). A large language model fine-tuned in the medical field is used, guided by the set medical prompt words, and a specific knowledge graph of heart failure is introduced as a prompt and constraint condition to extract core content related to diagnosis information, past medical history, key test indicators, medication regimen and treatment effect, etc., and generate structured or abstracted medical text representation to improve the efficiency of subsequent model understanding and utilization of text information. The final text summary . Among them represents a large language model summary generator, represents a carefully designed prompt word, is heart failure knowledge.

[0037] Multi-modal feature extraction, such as Figure 1 As shown, multiple modal-specific encoders are designed to extract and embed features of table health records (EHR), clinical text summary sentences, and multi-channel electrocardiogram physiological signals. The reason for designing special encoders for different modalities is that different types of data have significantly different structural and distribution characteristics. Table health records (EHR) are mainly structured numerical information, clinical text summary sentences have natural language semantic complexity, and electrocardiogram physiological signals are continuous time series signals. Therefore, designing structured data encoders, text encoders, and time series signal encoders can better adapt to the characteristics of each modality, effectively improve the accuracy and semantic expression ability of feature extraction, and provide high-quality bottom representation for subsequent multi-modal fusion. For table modal data , there are: ; Among them is a table encoder that encodes table modal data into feature representation vector . Similarly, for text modal data after LLM summary, there are: ; Among them is a text encoder that encodes text modal data into feature representation vector . Similarly, for signal modal data , there are: ; Among them is a signal encoder that encodes signal modal data into feature representation vector .

[0038] Multi-modal feature measurement, such as Figure 1The quality of the multi-modal data is evaluated and the modal weight coefficient is obtained, and the calculation method is as follows: For EHR table data , we have: ; ; Wherein, is the number of clinical examinations performed at irregular intervals during the patient's hospitalization. is the feature completeness index (missing rate), is the average time continuity index between adjacent two EHR records, denotes the th EHR examination.

[0039] The quality score of the table modality is calculated by weighting. Wherein , are learnable coefficients.

[0040] ; The weight coefficient of the table modality data is calculated as: ; For text summary data , a pre-defined medical entity set is used to perform a medical named entity recognition task to extract the entity set , wherein K is the number of non-repeated entities in the medical entity set, and M is the number of extracted entity sets. The proportion of medical entities appearing in the text is calculated, wherein is the number of medical entities covered in the text, is the total number of entities in the knowledge base.

[0041] ; The weight coefficient of the text modality data is calculated as ; For signal data , the signal-to-noise ratio is used for quality quantification. The weight coefficient of the signal modality data is calculated as .

[0042] For modal features , , , multiply the calculated weight coefficient, and the weighted modal features are: ; ; Multi-modal feature alignment, such as Figure 1 is introduced on the basis of the weighted feature projection representation. The mechanism calculates the attention output for any two modalities A and B, taking the features of modality A as Query and the features of modality B as Key and Value; at the same time, it reversely calculates the attention of modality B to modality A to achieve bidirectional enhancement. Independent linear transformation layers are defined for the two modalities to generate Query, Key, and Value.

[0043] ; ; ; wherein, is a learnable attention parameter, is the embedding dimension learnable by the attention layer. The attention is then calculated as follows. Through dot product operation, modality A calculates the similarity between its query vector and the key vector of modality B to obtain the attention weight distribution of modality B; then the value vector of modality B is weighted and summed using the weight to construct the attention representation of modality B to modality A.

[0044] ; Similarly, the bidirectional cross-attention mechanism further requires the reverse calculation of the attention of modality B to modality A to achieve bidirectional feature enhancement.

[0045] ; ; ; ; wherein, is a learnable attention parameter, is the embedding dimension learnable by the attention layer.

[0046] To achieve more comprehensive information fusion, this module performs pairwise cross-attention calculation on the three modalities of features , , , which specifically includes: the cross-attention vector between text and table, the cross-attention vector between text and signal, the cross-attention vectorcross attention vectors between tables and signals , After bidirectional cross attention of each group of modal pairs, the output representations thereof are fused to obtain an aligned multi-modal joint representation, i.e., an aligned fusion feature: In the multi-modal feature alignment module, the multi-modal feature alignment module further comprises an adversarial learning mechanism. The mechanism performs discriminant training on the fusion feature and each single-modal feature representation , , so as to promote the fusion feature to have modal indistinguishability while retaining the original modal information, thereby improving the modal collaboration and shared representation. Specifically: For the discriminator , the single-modal feature is , wherein . The objective of the discriminator is to distinguish which modal the input feature belongs to, and the loss function is cross-entropy loss: ; wherein is the modal label. On the contrary to the objective of the discriminator, the objective of the modal encoder and the fusion module is to deceive the discriminator so that the fusion feature is indistinguishable in the modal space.

[0047] ; During the training process, the discriminator and the feature fusion module are alternately optimized, forming a game relationship, thereby promoting the fusion feature to have cross-modal consistency and information fusion capability.

[0048] Heart failure early screening, as shown in Figure 1 , comprises a deep learning classification network, which receives the fused multi-modal feature representation as input and outputs a prediction label and a corresponding risk score probability for heart failure early screening.

[0049] As shown in Figure 2 , the method proposed in the present application is compared with two other single-modal methods, and their accuracy indicators are evaluated on the same data set. The method proposed in the present application achieves better performance, proving the beneficial effects of the present application.

[0050] The above embodiments describe the technical solutions and advantages of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modification, supplement and equivalent replacement made within the principle range of the present application shall be included in the protection range of the present application.

Claims

1. A multimodal early heart failure screening system based on a large model, characterized in that, include: The data acquisition and processing module is used to collect three modalities of data related to heart failure during the patient's hospitalization: EHR forms, clinical record texts, and electrocardiogram signals, and to perform preprocessing. The summary generation module uses a heart failure disease knowledge graph as prompts and constraints, and uses a large language model to extract key information closely related to heart failure risk from clinical record texts to generate text summaries. The multimodal feature extraction module extracts features from three modalities: a table encoder, a text encoder, and a deep neural network encoder. Specifically, the table encoder extracts pathological index change features from EHR tables, the text encoder extracts semantic features describing heart failure symptoms from text summaries, and the deep neural network encoder extracts ECG waveform change features from ECG signals. The multimodal feature measurement module is used to assess the quality of three modalities and obtain modality weight coefficients. The multimodal feature alignment module weights the three modal features based on modal weight coefficients, maps the weighted modal features to a shared representation space through a learnable projection layer, and then achieves alignment and fusion of multimodal features through a bidirectional cross-attention mechanism and adversarial learning strategy to obtain the final fused features. The heart failure early screening module inputs fused features into a deep learning classification network and outputs prediction results and risk probabilities for early screening of heart failure.

2. The multimodal early heart failure screening system based on a large model according to claim 1, characterized in that, In the data acquisition and processing module, the EHR table contains laboratory indicators reflecting cardiac function status, medication information, and vital signs; the clinical record text contains a course of medical records describing the evolution of heart failure symptoms, admission records, and discharge summaries; and the electrocardiogram signal includes a twelve-lead electrocardiogram.

3. The multimodal early heart failure screening system based on a large model according to claim 1, characterized in that, In the data acquisition and processing module, the preprocessing includes data denoising and missing value imputation.

4. The multimodal early heart failure screening system based on a large model according to claim 1, characterized in that, In the multimodal feature extraction module, for EHR table data ,have: ; in, It is a table encoder that encodes EHR table data and maps it into feature representation vectors. That is, the characteristics of changes in pathological indicators; Text summarization data ,have: ; in, It is a text encoder that encodes text summary data and maps it into feature representation vectors. That is, semantic features; ECG signal data ,have: ; in, It is a signal encoder that encodes ECG signal data and maps it into feature representation vectors. This refers to the characteristics of changes in electrocardiogram waveforms.

5. The multimodal early heart failure screening system based on a large model according to claim 4, characterized in that, The specific process of the multimodal feature measurement module is as follows: For EHR table data Calculate the feature integrity index The average time continuity index between two adjacent EHR records : ; ; in, This refers to the number of clinical examinations that a patient undergoes intermittently during their hospitalization. Indicates the first Second EHR check.

6. The mass fraction of the table modes is obtained by weighted calculation. ,in , These are learnable coefficients: ; Calculate the weighting coefficients for tabular modal data: ; Text summarization data Predefined medical entity set To carry out medical named entity recognition tasks and extract entity sets Where K is the number of unique entities in the medical entity set, and M is the number of extracted entity sets; calculate the proportion of medical entities appearing in the text. : ; in, It is the number of medical entities covered in the text. This represents the total number of entities in the knowledge base. Calculate the weighting coefficients for text modal data: ; For signal data Using signal-to-noise ratio Perform quality quantization and calculate the weighting coefficients of the signal modal data. .

7. The multimodal early heart failure screening system based on a large model according to claim 5, characterized in that, In the multimodal feature alignment module, the three modal features are weighted based on modality weight coefficients, specifically as follows: ; ; ; in, The weighted changes in pathological indicators are the characteristics. The weighted semantic features The weighted ECG waveform variation characteristics.

8. The multimodal early heart failure screening system based on a large model according to claim 6, characterized in that, In the multimodal feature alignment module, the weighted modal features are mapped to a shared representation space through a learnable projection layer, specifically as follows: Design a set of trainable linear projection parameters for the features of each modality. , , After mapping, the data is uniformly projected onto the shared space. : ; ; ; in, This indicates the characteristics of changes in pathological indicators after projection mapping. These are the semantic features after projection mapping. The characteristics of ECG waveform changes after projection mapping. , , This ensures that different modal features have the same dimension and spatial basis.

9. The multimodal early heart failure screening system based on a large model according to claim 7, characterized in that, In the multimodal feature alignment module, the bidirectional cross-attention mechanism works as follows: For three single-modal features , , Perform pairwise cross-attention calculations to obtain the cross-attention vector between the text and the table. , Cross-attention vector between text and signal , Cross-attention vector between table and signal , ; The six cross-attention vectors are then fused to obtain the fused features. .

10. The multimodal early heart failure screening system based on a large model according to claim 8, characterized in that, In the multimodal feature alignment module, the adversarial learning strategy works as follows: Introducing a discriminator to analyze fused features With each single-modal feature representation , , Perform discriminative training to promote feature fusion. While preserving the original modal information, it possesses modal indiscriminability, thereby enhancing modal collaboration and shared representation.

Citation Information

Patent Citations

  • Twelve-lead ECG-assisted heart disease multi-modal fusion screening method

    CN116269426A

  • Heart failure diagnosis auxiliary method based on multi-modal data fusion

    CN116451068A

  • Multi-mode heart failure prediction auxiliary method based on LSTM model and ResNet50 model

    CN116469553A