Text prompt type heart function intelligent evaluation method based on visual language large model
By constructing a multi-scale semantic system using a large visual language model, the problem of neglecting the physiological interdependence of the four chambers in cardiac function assessment is solved. It achieves accurate mapping from echocardiogram images to EF values, improves the accuracy and interpretability of assessment, adapts to complex clinical cases, and reduces the cost of data annotation and model deployment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV OF FINANCE & ECONOMICS
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for assessing cardiac function neglect the physiological interdependence and synergistic mechanisms among the four chambers of the heart, and fail to establish a coherent mapping system from pixel-level features to clinical semantics, resulting in predictions that lack physiological rationality and interpretability.
A multi-scale semantic system is constructed using a large visual language model. Through hierarchical feature learning and multi-head attention fusion, the system captures the macroscopic dynamic features and microscopic motion patterns of the heart, accurately aligning visual features with professional medical language descriptions to achieve precise mapping from echocardiogram images to EF values.
It significantly improves the accuracy and interpretability of cardiac function assessment, adapts to complex clinical cases, reduces data annotation costs and model deployment difficulty, and achieves accurate prediction of EF values from ultrasound video.
Smart Images

Figure CN122089671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a text-based intelligent assessment method for cardiac function based on a large visual language model, belonging to the fields of Internet and artificial intelligence technology. Background Technology
[0002] Cardiovascular disease remains a leading cause of death worldwide, accounting for approximately one-third of all deaths globally, posing a significant challenge to public health. Against this critical backdrop, accurate assessment of cardiac function is essential for the prevention, diagnosis, and treatment of cardiovascular disease. Among various cardiac imaging modalities, echocardiography is the most widely used clinical technique due to its non-ionizing radiation, real-time dynamic imaging capabilities, portability, cost-effectiveness, and ease of operation. It comprehensively assesses cardiac morphology and functional parameters, including ventricular volume, ejection fraction, and valvular performance, providing crucial evidence for diagnostic and treatment decisions in cardiovascular disease.
[0003] In recent years, deep learning-based methods for echocardiography analysis have made significant progress. The field has evolved from early convolutional neural networks to complex Transformer architectures, and researchers have developed various automated echocardiography (EF) assessment methods. EchoNet Dynamic pioneered the use of an R(2+1)D network for video-based end-to-end automated EF assessment, significantly improving computational efficiency while maintaining high performance. Building on this, EchoCoTr effectively captures spatiotemporal features using the UniFormer architecture and frame sampling techniques. EFNet further advances the field by combining the local feature extraction capabilities of CNNs with the global dependency modeling of Transformers, achieving even higher prediction accuracy.
[0004] However, current methods exhibit two fundamental limitations that hinder their clinical application. First, the main approaches either treat the heart as a monolithic structure or focus only on the left ventricle, thus neglecting the inherent physiological interdependencies and synergistic mechanisms between all four chambers. As a precisely coordinated four-chamber pump, cardiac performance depends heavily on the mechanical coupling between ventricles, atrioventricular synchrony, and comprehensive interventricular hemodynamic regulation circuitry. This overly simplistic modeling paradigm often yields predictions lacking physiological plausibility. Second, contemporary feature extraction and fusion methods operate primarily at the data level, failing to establish a coherent mapping system from pixel-level features to clinical semantics, thereby limiting the interpretability and clinical applicability of the models. Summary of the Invention
[0005] To address these key challenges, this invention proposes a cardiac function assessment system utilizing visual language models and multi-scale semantic fusion. The fundamental innovation of this invention involves modeling the heart as an integrated multi-scale semantic system, explicitly constructing cross-level feature representations through visual language models, ranging from coarse-grained chamber dynamics to fine-grained tissue motion. This approach enables precise alignment of visual features with specialized medical language descriptions within a unified semantic space.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows: a text-based intelligent assessment method for cardiac function based on a large visual language model, the method comprising the following steps:
[0007] Step 1: Construction of the echocardiographic marker dataset;
[0008] Step 2: Division of the four chambers of the heart;
[0009] Step 3: Hierarchical feature learning;
[0010] Step 4: Multi-scale feature fusion;
[0011] Step 5: Prediction of cardiac function (ejection fraction).
[0012] As an improvement to this invention, step 1: Construction of the echocardiographic labeled dataset. This invention collected 10,030 echocardiographic videos from a hospital that did not contain any sensitive personal information of the patients. These echocardiographic videos were processed frame by frame, and each echocardiographic video was processed into its corresponding echocardiographic image sequence according to cardiac time sequence. Then, LabelMe software was used to perform professional four-chamber segmentation labeling on the first frame of each echocardiographic image sequence to obtain the four-chamber segmentation mask image of its first frame. Thus, we have constructed a large echocardiographic image sequence labeled dataset containing the first frame mask.
[0013] As an improvement to this invention, step 2 involves four-chamber segmentation of the heart. Left ventricular ejection fraction (EF) is a key indicator reflecting normal and healthy cardiac function. Currently, most methods focus on segmenting a single ventricle of the left ventricle and then evaluating the EF value. However, these methods are limited by acquisition equipment and methods (such as blurred ventricular boundaries and artifacts in some echocardiographic videos), which cannot guarantee the accuracy of left ventricular volume calculation, leading to bias in EF value prediction. To reduce this problem, this invention uses a large model to extend single-chamber (left ventricle) segmentation methods (such as manual labeling and optical flow methods) to four-chamber (left ventricle, left atrium, right ventricle, right atrium) segmentation methods. This extension has the following advantages:
[0014] (1) It avoids the EF value prediction bias caused by image interference and individual differences.
[0015] (2) It has been expanded from a simple prediction of a single EF value to a more comprehensive assessment of cardiac function, and the medical interpretability of cardiac function prediction is stronger.
[0016] (3) Significantly reduce the redundant time consumed by the label.
[0017] (4) This provides important input for subsequent hierarchical feature learning and multi-scale feature extraction networks.
[0018] First, this invention uses LabelMe software to segment the first frame of the original video and generate corresponding masks. These first-frame masks will serve as reference images for the MedSAM2 large-scale model (a medical image segmentation model developed based on the large medical dataset SAM2). Then, the MedSAM2 large-scale model combines the first-frame mask images of each video to label subsequent frames of that video. The segmented videos of the four chambers generated using the MedSAM2 large-scale model show clear chamber boundaries, minimal artifact interference within the chambers, and a smooth, clear overall video. These high-quality segmented videos will provide an important information foundation for subsequent hierarchical feature learning and multi-scale feature extraction networks.
[0019] As an improvement of this invention, step 3 involves hierarchical feature learning. The high-quality segmented video obtained in step 2 contains rich dynamic information about the dynamic changes of the four chambers. This invention inputs the high-quality four-chamber segmented video into coarse-grained and fine-grained coding layers for hierarchical learning. The coarse-grained coding layer focuses on the macroscopic dynamic feature changes of the heart, while the fine-grained coding layer focuses on the microscopic motion patterns of the heart. To more flexibly adapt to different cardiac information, this invention designs semantic and temporal flows in the coarse-grained coding layer (the semantic flow uses the Ling Shu-7B medical large model to extract deep semantics, and the temporal flow extracts periodic information from the video). Simultaneously, this invention designs visual and dynamic flows in the fine-grained coding layer (the visual flow uses the R3D model to extract spatiotemporal feature fusion information, and the dynamic flow extracts and encodes cardiac information through a feature extractor and a deep coding layer). In this way, this invention constructs a multi-angle feature extraction network spanning four levels: semantic, temporal, visual, and dynamic, and encodes chamber features obtained from different angles through four separate flow channels (semantic flow, temporal flow, visual flow, and dynamic flow). This step can be divided into the following sub-steps:
[0020] Sub-step 3-1, Coarse-grained Coding Layer. The coarse-grained coding layer focuses on the overall macroscopic changes of the four chambers of the heart, globally capturing chamber features such as the overall semantic information of the segmented video of the four chambers and the changes in the cardiac cycle (systole and diastole) of the four chambers. Based on the macroscopic changes of the four chambers, a targeted design was implemented to extract semantic information from the semantic stream and cycle information from the temporal stream. This invention will be described in detail from two aspects: deep semantic extraction of AI diagnostic reports and temporal extraction of cardiac cycles. Sub-step 3-1-1, Semantic Stream Feature Extraction. Semantic stream feature extraction focuses on learning deep semantics from AI diagnostic reports. High-quality segmented video includes the chamber states of the four chambers and multiple cardiac cycles in each frame, with segmented labels for consecutive frames from systole to diastole within each cardiac cycle. This invention quantitatively analyzes the changes in systolic and diastolic area and rate of change of area of four chambers—left atrium (LA), left ventricle (LV), right atrium (RA), and right ventricle (RV)—as well as hemodynamic assessment records for each of the four chambers using text prompts in high-quality segmented video. The measurement data and state content are input into the Lingshu-7B model (an open-source multimodal medical model developed by Alibaba) for reliable deep physiological inference, generating AI diagnostic reports that meet clinical needs. These AI diagnostic reports can be further input into a semantic stream as deep medical semantics, and the BioBERT text feature encoder is used to perform AI diagnostic analysis. in Indicates the encoding process, It is a hidden feature vector. Indicates the feature dimension. and Represents the weight matrix. and It is the deviation vector. Representative approval standardization, It is an activation function. These are the encoded features ultimately obtained from the semantic stream. Sub-step 3-1-2, temporal stream feature extraction. The changes in the four chambers of an echocardiogram video are continuous in time, but the changes in the four chambers are inconsistent at different times. Dividing the four chambers in the echocardiogram video into systolic and diastolic phases is more representative. This invention provides high-quality echocardiogram video sequences. Segmentation was performed to extract and label the systolic phase of echocardiography. and diastolic phase .
[0021] Specifically, firstly, based on manually marked ESV frames Obtain the critical point of systole and diastole Then, query the video sequence. Frame index closest to the half-cycle critical point : Then, using the following constraints, based on the frame index Obtain video sequence The contraction and relaxation phases: Based on the above formula, the systolic phase in each of the four chambers can be obtained. and diastolic phase Sequential state changes This is used as input to the time stream. The cardiac cycle is encoded through a fully connected layer: in ( - dimensional phase characteristics). It is an activation function. Representatives criticized normalization. It is a weight matrix. It is the deviation vector. These are the encoded features ultimately obtained from the time stream.
[0022] Sub-step 3-2, Fine-grained coding layer. The fine-grained coding layer focuses on the microscopic changes in each of the four chambers, performing detailed local mining, such as visual changes in the spatiotemporal fusion features of each chamber, as well as dynamic changes in blood flow direction, chamber phase, and other factors. Based on the individual microscopic changes of the four chambers, a targeted design was developed to extract spatiotemporal information from the visual flow and dynamic change information from the dynamic flow. This invention will be described in detail from two aspects: visual extraction of spatiotemporal fusion and dynamic extraction of chamber change information. Sub-step 3-2-1, Visual flow feature extraction. Visual flow feature extraction focuses on revealing the visual semantic information of spatiotemporal integration. In order to better handle the temporal and spatial relationships of the four ventricles, capture more ventricular details at the visual level, and accurately obtain more clinically valuable cardiac data, this invention selects the spatiotemporal feature fusion model (R3D model). The R3D model is based on the ResNet-18 architecture and is a commonly used 3D CNN model for spatiotemporal feature extraction. It performs well in predicting and evaluating ejection fraction. Therefore, the R3D model is suitable for feature extraction from visually segmented four-chamber videos. The R3D model does not require separate processing of time and space; instead, it extracts features from echocardiograms through 3D spatiotemporal convolution. The 3D convolution kernel processes width, height, and time simultaneously. This processing mode tightly integrates spatiotemporal features, allowing the acquisition of the four chamber boundary positions, valve morphology, and ventricular wall thickness within a single frame, while also considering the changes in chamber systole, diastole, and blood flow direction over time. This invention inputs high-quality segmented echocardiogram video into the R3D model, obtaining eight 3D CNN visual features, including four chambers (left atrium, left ventricle, right atrium, and right ventricle) and two time periods (systole and diastole). These features contain rich temporal and spatial information.
[0023] in, , , , and .
[0024] All acquired 3D CNN features will be used as input to the visual stream and encoded individually. The encoding for each chamber and time segment is as follows: in , It is a weight matrix. It is the deviation vector. This is the encoded feature of a single 3D CNN. The concatenation of the 8 encoded features is as follows: in This represents the final visual stream encoding. Sub-step 3-2-2, Dynamic Flow Feature Extraction. In the fine-grained coding layer, dynamic flow captures the dynamic information of the four chambers as they change, refining the detailed features of chamber motion. Changes in the original chamber area can be transformed into physiologically meaningful representations through dynamic flow. These representations enable multi-level pathological identification, have strong medical interpretability, and further improve the accuracy of cardiac function (EF value) prediction. This invention uses high-quality echocardiographic video segmented from the four-chamber view. In the video, the left atrium (LA) is marked in green, the left ventricle (LV) in red, the right atrium (RA) in blue, and the right ventricle (RV) in yellow. Compared to the original black-and-white blurred imaging mode, the color-segmented echocardiographic video has clearer chamber segmentation. Clear chamber segmentation allows this invention to extract more chamber states, including the area and phase of each chamber in a single frame, as well as dynamic changes over time, such as changes in blood flow direction. This invention refers to the above information as dynamic information and uses it as input to the dynamic flow. Then, the feature extractor performs feature engineering processing. The detailed calculation formula is as follows: I. Basic characteristics: Area of a single-frame chamber: in It is the chamber area. It is the image width. It's a high-resolution image. It is a segmentation mask.
[0025] Chamber stage:
[0026]
[0027] in It is a time frame Phase encoding at the location, Indicates the time frame index. Hemodynamic characteristics:
[0028]
[0029] in Indicates blood flow characteristics, Indicates the type of blood flow pattern. It is a tag encoder. It represents the total number of blood flow categories.
[0030] II. Regional dynamic characteristics:
[0031] Regional velocity variation:
[0032]
[0033] in It is a time frame Area velocity at that location, It is a time frame The area of the cavity at that location, Indicates the time frame index.
[0034] Regional acceleration changes:
[0035]
[0036] in It is a time frame Area acceleration at that point It is a time frame Surface velocity at the location, Indicates the time frame index.
[0037] Area change rate:
[0038]
[0039] in It is a time frame The rate of change of area, It is a tiny constant to avoid division by zero. This indicates the cavity area of the current frame.
[0040] Cumulative area change:
[0041]
[0042] in It is a time frame The cumulative area change, For the first Frame area This represents the average area.
[0043] Localized changes:
[0044]
[0045] in It is a time frame The local area range, For the first The region at the frame, For frame radius, Indicates the frame index. Indicates the maximum area. This represents the minimum area.
[0046] III. Phase characteristics:
[0047] The chamber phase code is shown in the equation above.
[0048] Phase transition:
[0049]
[0050] in Detection time frame Phase transition at the location. This value is 1 when the heart phase changes (from systole to diastole and vice versa), and 0 otherwise.
[0051] Phase position:
[0052]
[0053] in It is a time frame Position within the current phase, Indicates the start frame of the current phase. This indicates the end frame of the current phase.
[0054] IV. Characteristics of blood flow phase interaction:
[0055] The hemodynamic characteristics are shown in the equations above.
[0056] Blood flow phase interaction:
[0057]
[0058] in It is a time frame Blood flow phase interaction characteristics at the location, Indicates blood flow characteristics, It is phase encoding. This represents element-wise multiplication.
[0059] V. Statistical characteristics:
[0060] Smoothed average:
[0061]
[0062] in It is a time frame The smoothed mean at that point, It is a time frame The sliding window at the location, It is the i-th area value within the window. It is the normalization factor.
[0063] Standard deviation:
[0064]
[0065] in It is a time frame The rolling standard deviation at that point Indicates a sliding window. It is the area value within the window. It is the rolling average, while It is the normalization factor.
[0066] Rate of change:
[0067]
[0068] in It is a time frame coefficient of variation, Indicates standard deviation, This represents the average value.
[0069] A complete 16-dimensional dynamic information vector:
[0070]
[0071] This invention uses formulas to process dynamic information into 16-dimensional features through feature engineering. The basic features are three-dimensional (single-frame chamber area, chamber phase, hemodynamic features); the regional dynamic features are five-dimensional (regional velocity change, regional acceleration change, regional rate of change, cumulative area change, local region change); the phase features are three-dimensional (chamber phase, phase transition, phase position); the blood flow phase interaction features are two-dimensional (hemodynamic features, mutual influence of blood flow phases); and the statistical features are three-dimensional (smoothed mean, standard deviation, rate of change). After passing through a feature extractor, the features are processed in a deep coding layer. Since bidirectional LSTM can ensure the integrity of dynamic information, this invention inputs the obtained 16-dimensional features of each chamber (64 dimensions for four chambers) into a BiLSTM layer:
[0072]
[0073] in It is in a hidden state, superscript Indicates the chamber type, Indicates the cell state, subscript Indicates the frame number. It is the maximum sequence length.
[0074] Subsequently, this invention designs a multi-head attention layer, with eight attention heads per ventricle, to monitor the overall characteristics of the ventricle from different perspectives. The calculation formula is as follows:
[0075]
[0076]
[0077] in By stacking all The resulting matrix It is the first Bidirectional LSTM output sequence for each chamber Indicates a cascading operation. It is the output projection matrix. This represents the multi-head attention extraction process. This is the final output.
[0078] Combining the two-cell layer (mean cell and maximal cell), the mean cell focuses on capturing the overall trend of the ventricles, while the maximal cell focuses on the peak characteristics of the heart. To comprehensively assess the information, the calculation formula is as follows:
[0079]
[0080]
[0081] in , .
[0082] Then, the encoding dimension is reduced by using three fully connected layers, calculated as follows:
[0083]
[0084] Obtain the encoding vector of the final dynamic stream .
[0085] As an improvement to this invention, step 4: multi-scale feature fusion, utilizing the semantic stream coding features learned in step 3. Time-stream coding features Visual encoding function and dynamic stream coding characteristics These are then connected to form high-quality coded features that are multi-faceted, multi-level, and highly interpretable. This step is implemented as follows:
[0086]
[0087] And use a multi-head attention mechanism to dynamically assign weights to each encoded feature:
[0088]
[0089] in This represents the complete input encoding features. This indicates the enhanced encoding function after attention processing. This represents the multi-head attention extraction process.
[0090] As an improvement to this invention, step 5: cardiac function (ejection fraction) prediction. Based on the design of the previous four steps, the final encoded features of this invention will contain complete semantic, temporal, visual, and dynamic information, and establish correlations between different levels. Finally, Tanh(·) linear mapping constraints are used, as follows:
[0091]
[0092]
[0093] in It is a two-layer fully connected network, and It is the tanh mapping function.
[0094] Subsequently, inverse normalization is performed to generate a prediction of cardiac function (EF value).
[0095] By employing linear constraints and dimensionality reduction with fully connected layers, this study accurately assesses the unity among dimensional information within features from different levels and the temporal continuity of feature fusion. This approach significantly improves the model's ejection fraction prediction capability while exhibiting strong medical interpretability and robustness to variations in terminology and complex clinical presentations, highlighting its potential for real-world deployment. Overall, this work provides a principled and transparent approach to AI-assisted cardiac ultrasound, laying the foundation for more reliable collaborative diagnostic systems for clinicians.
[0096] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the aforementioned text-based intelligent assessment method for cardiac function based on a large visual language model.
[0097] A storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned text-based intelligent assessment method for cardiac function based on a large visual language model.
[0098] Compared with the prior art, the advantages of the present invention are as follows:
[0099] (1) This invention introduces a text-based intelligent assessment method for cardiac function based on a large visual language model. This invention proposes a novel multi-scale semantic fusion framework. This framework employs a hierarchical feature learning mechanism to simultaneously capture the macroscopic dynamic features and microscopic motion patterns of the heart. Through an attention-based fusion module, complex semantic relationships are established between coarse-grained and fine-grained features. The hierarchical feature learning + attention fusion module achieves complementarity and enhancement of multi-scale features, significantly improving the robustness of feature representation and adapting it to complex clinical echocardiogram data.
[0100] (2) This invention is the first to systematically integrate a visual language model into cardiac ultrasound analysis. Through medical-guided feature embedding, it achieves deep semantic alignment between visual features and clinical terms, significantly improving the interpretability and clinical relevance of the model. This invention achieves deep semantic alignment between visual features and clinical terms through medical-guided feature embedding, realizing feature-level and decision-level interpretability at the technical level. At the same time, medical-guided feature embedding achieves hierarchical enhancement of interpretability, breaking through the "black box" bottleneck of traditional deep learning.
[0101] (3) This invention designs an end-to-end trainable architecture that, while maintaining physiological rationality constraints, ensures the model's prediction accuracy and avoids the "physiological paradox," achieving a precise mapping from raw ultrasound video to EF prediction. Extensive experiments on public datasets show that this method has superior performance compared to state-of-the-art methods, with significant improvements in handling complex clinical cases. The end-to-end trainable architecture of this invention integrates the entire process of "raw echocardiogram video input → EF value and other index output" into a unified model, achieving three major engineering advantages at the technical level: end-to-end training throughout the entire process avoids error propagation and accumulation between step-by-step modules, improving overall prediction accuracy; only one model needs to be deployed in clinical practice, without the need to connect to multiple independent modules, reducing the connection cost of hospital information systems; the end-to-end architecture adopts a hybrid training mode of edge / cloud, which can directly train on raw echocardiogram videos without complex manual preprocessing, significantly reducing the labeling cost of clinical data. The architecture design of this invention meets the requirements of "simplified deployment and low-cost adaptation" for the clinical application of medical AI. Compared with the traditional multi-module architecture, it can be quickly adapted to different information systems in primary hospitals and large tertiary hospitals. Attached Figure Description
[0102] Figure 1 This is an overall model diagram of an embodiment of the present invention;
[0103] Figure 2 This is a flowchart illustrating the method processing of an embodiment of the present invention. Detailed Implementation
[0104] To enhance understanding of the present invention, the invention will be further explained below with reference to specific embodiments.
[0105] Example 1: A text-based intelligent assessment method for cardiac function based on a large visual language model. This method models the heart as a multi-scale semantic system, capturing features from macroscopic chamber dynamics to microscopic tissue motion through a hierarchical feature learning mechanism. By aligning these hierarchical visual features with professional medical descriptions in a shared semantic space, the model achieves deep integration of imaging data and clinical knowledge.
[0106] For the specific model, please refer to [link / details]. Figure 1 The detailed implementation steps are as follows:
[0107] Step 1: Construction of the Echocardiographic Labeling Dataset. This invention collected 10,030 echocardiographic videos from a hospital that did not contain any sensitive patient information. These echocardiographic videos were processed frame by frame, and each video was processed into its corresponding echocardiographic image sequence according to cardiac time sequence. Then, LabelMe software was used to perform professional four-chamber segmentation labeling on the first frame of each echocardiographic image sequence, obtaining the four-chamber segmentation mask image of the first frame. Thus, we constructed a large echocardiographic image sequence labeling dataset containing the first frame mask.
[0108] Step 2: Four-chamber segmentation of the heart. Left ventricular ejection fraction (EF) is a key indicator reflecting the normality and health of cardiac function. Currently, most methods focus on segmenting the left ventricle individually and then evaluating the EF value. However, these methods are limited by acquisition equipment and methods (such as blurred ventricular boundaries and artifacts in some echocardiographic videos), which cannot guarantee the accuracy of left ventricular volume calculation, leading to bias in EF value prediction. To reduce this problem, this invention uses a large model to extend single-chamber (left ventricle) segmentation methods (such as manual labeling and optical flow methods) to four-chamber (left ventricle, left atrium, right ventricle, right atrium) segmentation methods. This extension has the following advantages:
[0109] (1) It avoids the EF value prediction bias caused by image interference and individual differences.
[0110] (2) It has been expanded from a simple prediction of a single EF value to a more comprehensive assessment of cardiac function, and the medical interpretability of cardiac function prediction is stronger.
[0111] (3) Significantly reduce the redundant time consumed by the label.
[0112] (4) This provides important input for subsequent hierarchical feature learning and multi-scale feature extraction networks.
[0113] First, this invention uses LabelMe software to segment the first frame of the original video and generate corresponding masks. These first-frame masks will serve as reference images for the MedSAM2 large-scale model (a medical image segmentation model developed based on the large medical dataset SAM2). Then, the MedSAM2 large-scale model combines the first-frame mask images of each video to label subsequent frames of that video. The segmented videos of the four chambers generated using the MedSAM2 large-scale model show clear chamber boundaries, minimal artifact interference within the chambers, and a smooth, clear overall video. These high-quality segmented videos will provide an important information foundation for subsequent hierarchical feature learning and multi-scale feature extraction networks.
[0114] Step 3: Hierarchical Feature Learning. The high-quality segmented video obtained in Step 2 contains rich dynamic information about the dynamic changes of the four chambers. This invention inputs the high-quality four-chamber segmented video into coarse-grained and fine-grained coding layers for hierarchical learning. The coarse-grained coding layer focuses on the macroscopic dynamic feature changes of the heart, while the fine-grained coding layer focuses on the microscopic motion patterns of the heart. To more flexibly adapt to different cardiac information, this invention designs semantic and temporal flows in the coarse-grained coding layer (the semantic flow uses the Ling Shu-7B medical large model to extract deep semantics, and the temporal flow extracts periodic information from the video). Simultaneously, this invention designs visual and dynamic flows in the fine-grained coding layer (the visual flow uses the R3D model to extract spatiotemporal feature fusion information, and the dynamic flow extracts and encodes cardiac information through a feature extractor and a deep coding layer). In this way, this invention constructs a multi-angle feature extraction network across four levels: semantic, temporal, visual, and dynamic, and encodes chamber features obtained from different angles through four separate flow channels (semantic flow, temporal flow, visual flow, and dynamic flow). The implementation of this step can be divided into the following sub-steps:
[0115] Sub-step 3-1, Coarse-grained Coding Layer. The coarse-grained coding layer focuses on the overall macroscopic changes of the four chambers of the heart, globally capturing chamber features such as the overall semantic information of the segmented video of the four chambers and the changes in the cardiac cycle (systole and diastole) of the four chambers. Based on the macroscopic changes of the four chambers, a targeted design was implemented to extract semantic information from the semantic stream and cycle information from the temporal stream. This invention will be described in detail from two aspects: deep semantic extraction of AI diagnostic reports and temporal extraction of cardiac cycles. Sub-step 3-1-1, Semantic Stream Feature Extraction. Semantic stream feature extraction focuses on learning deep semantics from AI diagnostic reports. High-quality segmented video includes the chamber states of the four chambers and multiple cardiac cycles in each frame, with segmented labels for consecutive frames from systole to diastole within each cardiac cycle. This invention quantitatively analyzes the changes in systolic and diastolic area and rate of change of area of four chambers—left atrium (LA), left ventricle (LV), right atrium (RA), and right ventricle (RV)—as well as hemodynamic assessment records for each of the four chambers using text prompts in high-quality segmented video. The measurement data and state content are input into the Lingshu-7B model (an open-source multimodal medical model developed by Alibaba) for reliable deep physiological inference, generating AI diagnostic reports that meet clinical needs. These AI diagnostic reports can be further input into a semantic stream as deep medical semantics, and the BioBERT text feature encoder is used to perform AI diagnostic analysis. in Indicates the encoding process, It is a hidden feature vector. Indicates the feature dimension. and Represents the weight matrix. and It is the deviation vector. Representative approval standardization, It is an activation function. These are the encoded features ultimately obtained from the semantic stream. Sub-step 3-1-2, temporal stream feature extraction. The changes in the four chambers of an echocardiogram video are continuous in time, but the changes in the four chambers are inconsistent at different times. Dividing the four chambers in the echocardiogram video into systolic and diastolic phases is more representative. This invention provides high-quality echocardiogram video sequences. Segmentation was performed to extract and label the systolic phase of echocardiography. and diastolic phase .
[0116] Specifically, firstly, based on manually marked ESV frames Obtain the critical point of systole and diastole Then, query the video sequence. Frame index closest to the half-cycle critical point : Then, using the following constraints, based on the frame index Obtain video sequence The contraction and relaxation phases: Based on the above formula, the systolic phase in each of the four chambers can be obtained. and diastolic phase Sequential state changes This is used as input to the time stream. The cardiac cycle is encoded through a fully connected layer: in ( - dimensional phase characteristics). It is an activation function. Representatives criticized normalization. It is a weight matrix. It is the deviation vector. These are the encoded features ultimately obtained from the time stream.
[0117] Sub-step 3-2, Fine-grained coding layer. The fine-grained coding layer focuses on the microscopic changes in each of the four chambers, performing detailed local mining, such as visual changes in the spatiotemporal fusion features of each chamber, as well as dynamic changes in blood flow direction, chamber phase, and other factors. Based on the individual microscopic changes of the four chambers, a targeted design was developed to extract spatiotemporal information from the visual flow and dynamic change information from the dynamic flow. This invention will be described in detail from two aspects: visual extraction of spatiotemporal fusion and dynamic extraction of chamber change information. Sub-step 3-2-1, Visual flow feature extraction. Visual flow feature extraction focuses on revealing the visual semantic information of spatiotemporal integration. In order to better handle the temporal and spatial relationships of the four ventricles, capture more ventricular details at the visual level, and accurately obtain more clinically valuable cardiac data, this invention selects the spatiotemporal feature fusion model (R3D model). The R3D model is based on the ResNet-18 architecture and is a commonly used 3D CNN model for spatiotemporal feature extraction. It performs well in predicting and evaluating ejection fraction. Therefore, the R3D model is suitable for feature extraction from visually segmented four-chamber videos. The R3D model does not require separate processing of time and space; instead, it extracts features from echocardiograms through 3D spatiotemporal convolution. The 3D convolution kernel processes width, height, and time simultaneously. This processing mode tightly integrates spatiotemporal features, allowing the acquisition of the four chamber boundary positions, valve morphology, and ventricular wall thickness within a single frame, while also considering the changes in chamber systole, diastole, and blood flow direction over time. This invention inputs high-quality segmented echocardiogram video into the R3D model, obtaining eight 3D CNN visual features, including four chambers (left atrium, left ventricle, right atrium, and right ventricle) and two time periods (systole and diastole). These features contain rich temporal and spatial information.
[0118] in, , , , and .
[0119] All acquired 3D CNN features will be used as input to the visual stream and encoded individually. The encoding for each chamber and time segment is as follows: in , It is a weight matrix. It is the deviation vector. This is the encoded feature of a single 3D CNN. The concatenation of the 8 encoded features is as follows: in This represents the final visual stream encoding. Sub-step 3-2-2, Dynamic Flow Feature Extraction. In the fine-grained coding layer, dynamic flow captures the dynamic information of the four chambers as they change, refining the detailed features of chamber motion. Changes in the original chamber area can be transformed into physiologically meaningful representations through dynamic flow. These representations enable multi-level pathological identification, have strong medical interpretability, and further improve the accuracy of cardiac function (EF value) prediction. This invention uses high-quality echocardiographic video segmented from the four-chamber view. In the video, the left atrium (LA) is marked in green, the left ventricle (LV) in red, the right atrium (RA) in blue, and the right ventricle (RV) in yellow. Compared to the original black-and-white blurred imaging mode, the color-segmented echocardiographic video has clearer chamber segmentation. Clear chamber segmentation allows this invention to extract more chamber states, including the area and phase of each chamber in a single frame, as well as dynamic changes over time, such as changes in blood flow direction. This invention refers to the above information as dynamic information and uses it as input to the dynamic flow. Then, the feature extractor performs feature engineering processing. The detailed calculation formula is as follows: I. Basic characteristics: Area of a single-frame chamber: in It is the chamber area. It is the image width. It's a high-resolution image. It is a segmentation mask.
[0120] Chamber stage:
[0121]
[0122] in It is a time frame Phase encoding at the location, Indicates the time frame index. Hemodynamic characteristics:
[0123]
[0124] in Indicates blood flow characteristics, Indicates the type of blood flow pattern. It is a tag encoder. It represents the total number of blood flow categories.
[0125] II. Regional dynamic characteristics:
[0126] Regional velocity variation:
[0127]
[0128] in It is a time frame Area velocity at that location, It is a time frame The area of the cavity at that location, Indicates the time frame index.
[0129] Regional acceleration changes:
[0130]
[0131] in It is a time frame Area acceleration at that point It is a time frame Surface velocity at the location, Indicates the time frame index.
[0132] Area change rate:
[0133]
[0134] in It is a time frame The rate of change of area, It is a tiny constant to avoid division by zero. This indicates the cavity area of the current frame.
[0135] Cumulative area change:
[0136]
[0137] in It is a time frame The cumulative area change, For the first Frame area This represents the average area.
[0138] Localized changes:
[0139]
[0140] in It is a time frame The local area range, For the first The region at the frame, For frame radius, Indicates the frame index. Indicates the maximum area. This represents the minimum area.
[0141] III. Phase characteristics:
[0142] The chamber phase code is shown in the equation above.
[0143] Phase transition:
[0144]
[0145] in Detection time frame Phase transition at the location. This value is 1 when the heart phase changes (from systole to diastole and vice versa), and 0 otherwise.
[0146] Phase position:
[0147]
[0148] in It is a time frame Position within the current phase, Indicates the start frame of the current phase. This indicates the end frame of the current phase.
[0149] IV. Characteristics of blood flow phase interaction:
[0150] The hemodynamic characteristics are shown in the equations above.
[0151] Blood flow phase interaction:
[0152]
[0153] in It is a time frame Blood flow phase interaction characteristics at the location, Indicates blood flow characteristics, It is phase encoding. This represents element-wise multiplication.
[0154] V. Statistical characteristics:
[0155] Smoothed average:
[0156]
[0157] in It is a time frame The smoothed mean at that point, It is a time frame The sliding window at the location, It is the i-th area value within the window. It is the normalization factor.
[0158] Standard deviation:
[0159]
[0160] in It is a time frame The rolling standard deviation at that point Indicates a sliding window. It is the area value within the window. It is the rolling average, while It is the normalization factor.
[0161] Rate of change:
[0162]
[0163] in It is a time frame coefficient of variation, Indicates standard deviation, This represents the average value.
[0164] A complete 16-dimensional dynamic information vector:
[0165]
[0166] This invention uses formulas to process dynamic information into 16-dimensional features through feature engineering. The basic features are three-dimensional (single-frame chamber area, chamber phase, hemodynamic features); the regional dynamic features are five-dimensional (regional velocity change, regional acceleration change, regional rate of change, cumulative area change, local region change); the phase features are three-dimensional (chamber phase, phase transition, phase position); the blood flow phase interaction features are two-dimensional (hemodynamic features, mutual influence of blood flow phases); and the statistical features are three-dimensional (smoothed mean, standard deviation, rate of change). After passing through a feature extractor, the features are processed in a deep coding layer. Since bidirectional LSTM can ensure the integrity of dynamic information, this invention inputs the obtained 16-dimensional features of each chamber (64 dimensions for four chambers) into a BiLSTM layer:
[0167]
[0168] in It is in a hidden state, superscript Indicates the chamber type, Indicates the cell state, subscript Indicates the frame number. It is the maximum sequence length.
[0169] Subsequently, this invention designs a multi-head attention layer, with eight attention heads per ventricle, to monitor the overall characteristics of the ventricle from different perspectives. The calculation formula is as follows:
[0170]
[0171]
[0172] in By stacking all The resulting matrix It is the first Bidirectional LSTM output sequence for each chamber Indicates a cascading operation. It is the output projection matrix. This represents the multi-head attention extraction process. This is the final output.
[0173] Combining the two-cell layer (mean cell and maximal cell), the mean cell focuses on capturing the overall trend of the ventricles, while the maximal cell focuses on the peak characteristics of the heart. To comprehensively assess the information, the calculation formula is as follows:
[0174]
[0175]
[0176] in , .
[0177] Then, the encoding dimension is reduced by using three fully connected layers, calculated as follows:
[0178]
[0179] Obtain the encoding vector of the final dynamic stream .
[0180] As an improvement to this invention, step 4: multi-scale feature fusion, utilizing the semantic stream coding features learned in step 3. Time-stream coding features Visual encoding function and dynamic stream coding characteristics These are then connected to form high-quality coded features that are multi-faceted, multi-level, and highly interpretable. This step is implemented as follows:
[0181]
[0182] And use a multi-head attention mechanism to dynamically assign weights to each encoded feature:
[0183]
[0184] in This represents the complete input encoding features. This indicates the enhanced encoding function after attention processing. This represents the multi-head attention extraction process.
[0185] Step 5: Cardiac function (ejection fraction) prediction. Based on the design of the previous four steps, the final encoded features of this invention will contain complete semantic, temporal, visual, and dynamic information, and establish correlations between different levels. Finally, Tanh(·) linear mapping constraints are used, as follows:
[0186]
[0187]
[0188] in It is a two-layer fully connected network, and It is the tanh mapping function.
[0189] Subsequently, inverse normalization is performed to generate a prediction of cardiac function (EF value).
[0190] By employing linear constraints and dimensionality reduction with fully connected layers, this study accurately assesses the unity among dimensional information within features from different levels and the temporal continuity of feature fusion. This approach significantly improves the model's ejection fraction prediction capability while exhibiting strong medical interpretability and robustness to variations in terminology and complex clinical presentations, highlighting its potential for real-world deployment. Overall, this work provides a principled and transparent approach to AI-assisted cardiac ultrasound, laying the foundation for more reliable collaborative diagnostic systems for clinicians.
[0191] Based on the same inventive concept, the present invention provides a text-based intelligent assessment method and apparatus for cardiac function based on a large visual language model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the aforementioned text-based intelligent assessment method for cardiac function based on a large visual language model.
[0192] Those skilled in the art will recognize that the embodiments described herein are intended to help readers understand the principles of the invention. It should be understood that the embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the claims of this application.
Claims
1. A text-based intelligent assessment method for cardiac function based on a large visual-language model, characterized in that, The method includes the following steps: Step 1: Construction of the echocardiographic marker dataset; Step 2: Division of the four chambers of the heart; Step 3: Hierarchical feature learning; Step 4: Multi-scale feature fusion; Step 5: Prediction of cardiac function (ejection fraction).
2. The text-based intelligent assessment method for cardiac function based on a large visual-language model according to claim 1, characterized in that, Step 1: Construction of the echocardiogram labeled dataset. 10,030 echocardiogram videos from a hospital were collected without containing any sensitive personal information of the patients. These echocardiogram videos were processed frame by frame, and each echocardiogram video was processed into its corresponding echocardiogram image sequence according to the cardiac time sequence. Then, LabelMe software was used to perform professional four-chamber segmentation labeling on the first frame of each echocardiogram image sequence to obtain the four-chamber segmentation mask image of its first frame.
3. The text-based intelligent assessment method for cardiac function based on a large visual-language model according to claim 2, characterized in that, Step 2: Four-chamber segmentation of the heart. Left ventricular ejection fraction (EF) is a key indicator reflecting whether the heart function is normal and healthy. A large model is used to extend the single-chamber (left ventricle) segmentation method to a four-chamber (left ventricle, left atrium, right ventricle, right atrium) segmentation method. First, the first frame of the original video is segmented using LabelMe software, and corresponding masks are generated. These first-frame masks will serve as reference images for the MedSAM2 large model. Then, the MedSAM2 large model will combine the first-frame mask images of each video to label the subsequent frames of its corresponding video. The segmented video of the four chambers generated using the MedSAM2 large model shows clear segmented chamber boundaries, minimal artifact interference within the chambers, and a smooth and clear overall video.
4. The text-based intelligent assessment method for cardiac function based on a large visual-language model according to claim 2, characterized in that, Step 3: Hierarchical feature learning, which is implemented in the following sub-steps: Sub-step 3-1, coarse-grained coding layer: This layer focuses on the overall macroscopic changes of the four chambers of the heart, globally capturing chamber features. Sub-step 3-1-1, semantic flow feature extraction: This step focuses on learning deep semantics from the AI diagnostic report. High-quality segmented video includes the chamber states of the four chambers in each frame and multiple cardiac cycles. Within each cardiac cycle, consecutive frames from systole to diastole are segmented and labeled. Through textual cues in the high-quality segmented video, the systolic and diastolic area changes and rate of change of the four chambers (left atrium (LA), left ventricle (LV), right atrium (RA), and right ventricle (RV)) are quantitatively analyzed, along with hemodynamic assessment records for each of the four chambers. The measurement data and state content are input into the Lingshu-7B model for reliable deep physiological inference, generating an AI diagnostic report that meets clinical needs. The AI diagnostic report, as deep medical semantics, is further input into the semantic flow. The BioBERT text feature encoder is used to perform AI diagnostic analysis. in Indicates the encoding process, It is a hidden feature vector. Representing feature dimension, and Represents the weight matrix. and It is the deviation vector. Representative approval standardization, It is an activation function. The final encoded features are obtained from the semantic stream. Sub-step 3-1-2, temporal flow feature extraction, shows that the changes in the four chambers in an echocardiogram video are continuous in time, but the changes in the four chambers are inconsistent at different times. Dividing the four chambers in the echocardiogram video into systolic and diastolic phases is more representative for high-quality echocardiogram video sequences. Segmentation was performed to extract and label the systolic phase of echocardiography. and diastolic phase , Specifically, the process begins with manually tagged ESV frames. Obtain the critical point of systole and diastole Then, query the video sequence. Frame index closest to the half-cycle critical point : Then, using the following constraints, based on the frame index Obtain video sequence The contraction and relaxation phases: Based on the above formula, the systolic phase in the four chambers is obtained. and diastolic phase Sequential state changes This is used as input to the time stream, and the cardiac cycle is encoded through a fully connected layer: in ( - dimensional phase characteristics). It is an activation function. Representatives criticized normalization. It is a weight matrix. It is the deviation vector. These are the encoded features ultimately obtained from the time stream. Sub-step 3-2, Fine-grained coding layer: This layer focuses on the microscopic changes in each of the four chambers, performing detailed local analysis. Sub-step 3-2-1, Visual flow feature extraction: This layer emphasizes revealing the spatiotemporally integrated visual semantic information. A spatiotemporal feature fusion model (R3D model) was selected. The R3D model, based on the ResNet-18 architecture, is a commonly used 3D CNN model for spatiotemporal feature extraction. It performs well in predicting and evaluating ejection fraction. Therefore, the R3D model is suitable for feature extraction from visually segmented four-chamber videos. The 3D convolutional kernels simultaneously process width, height, and time. This processing mode closely integrates spatiotemporal features, allowing the acquisition of the four chamber boundary positions, valve morphology, and ventricular wall thickness in a single frame image. It also focuses on the changes in chamber contraction, relaxation, and blood flow direction over time. High-quality echocardiographic segmentation video is input into the R3D model, and eight 3D features are obtained through the R3D model. CNN visual features, including four chambers (left atrium, left ventricle, right atrium, and right ventricle) and two time periods (systole and diastole), contain rich temporal and spatial information. in, , , , and , All obtained 3D CNN features will be used as input to the visual stream and encoded individually, with the encoding for each chamber and time period as follows: in , It is a weight matrix. It is the deviation vector. This is the encoded feature of a single 3D CNN. The connections of the eight encoded features are as follows: in This represents the final visual stream encoding. Sub-step 3-2-2, Dynamic Flow Feature Extraction: In the fine-grained coding layer, the dynamic flow captures the dynamic information of the four chambers during changes, refining the detailed features of chamber motion. The changes in the original chamber area can be transformed into a physiologically meaningful representation through dynamic flow. This information is called dynamic information and is used as the input to the dynamic flow, then enters the feature extractor for feature engineering processing. The detailed calculation formula is as follows: I. Basic characteristics: Single-frame chamber area: in It is the chamber area. It is the image width. It's a high-resolution image. It is a segmentation mask. Chamber stage: in It is a time frame Phase encoding at the location, Indicates time frame index, hemodynamic characteristics: in Indicates blood flow characteristics, Indicates the type of blood flow pattern. It is a tag encoder. It is the total number of blood flow categories. II. Regional dynamic characteristics: Regional velocity variation: in It is a time frame Area velocity at that location, It is a time frame The area of the cavity at that location, Indicates the time frame index. Regional acceleration changes: in It is a time frame Area acceleration at that point It is a time frame Surface velocity at the location, Indicates the time frame index. Area change rate: in It is a time frame The rate of change of area, It is a tiny constant to avoid division by zero. This represents the chamber area of the current frame. Cumulative area change: in It is a time frame The cumulative area change, For the first Frame area This represents the average area. Localized changes: in It is a time frame The local area range, For the first The region at the frame, For frame radius, Indicates the frame index. Indicates the maximum area. Represents the minimum area. III. Phase characteristics: The chamber phase code is shown in the equation above. Phase transition: in Detection time frame The phase transition value is 1 when the heart phase changes (from systole to diastole and vice versa), and 0 otherwise. Phase position: in It is a time frame Position within the current phase, Indicates the start frame of the current phase. This indicates the end frame of the current phase. IV. Characteristics of blood flow phase interaction: The hemodynamic characteristics are shown in the equations above. Blood flow phase interaction: in It is a time frame Blood flow phase interaction characteristics at the location, Indicates blood flow characteristics, It is phase encoding. Indicates element-wise multiplication. V. Statistical characteristics: Smoothed average: in It is a time frame The smoothed mean at that point, It is a time frame The sliding window at the location, It is the i-th area value within the window. It is a normalization factor. Standard deviation: in It is a time frame The rolling standard deviation at that point Indicates a sliding window. It is the area value within the window. It is the rolling average, while It is a normalization factor. Rate of change: in It is a time frame coefficient of variation, Indicates standard deviation, This represents the average value. A complete 16-dimensional dynamic information vector: The dynamic information is processed into 16-dimensional features using feature engineering. The basic features are three-dimensional (single-frame chamber area, chamber phase, hemodynamic features), the regional dynamic features are five-dimensional (regional velocity change, regional acceleration change, regional rate of change, cumulative area change, local regional change), the phase features are three-dimensional (chamber phase, phase transition, phase position), the blood flow phase interaction features are two-dimensional (hemodynamic features, blood flow phase interaction), and the statistical features are three-dimensional (smoothed mean, standard deviation, rate of change). After passing through the feature extractor, the features are processed in a deep coding layer. Since bidirectional LSTM can ensure the integrity of the dynamic information, the obtained 16-dimensional features of each chamber (64 dimensions for four chambers) are input into the BiLSTM layer. in It is in a hidden state, superscript Indicates the chamber type, Indicates the cell state, subscript Indicates the frame number. It is the maximum sequence length. Subsequently, a multi-head attention layer was designed, with 8 attention heads per ventricle to focus on the overall characteristics of the ventricle from different levels. The calculation formula is as follows: in By stacking all The resulting matrix It is the first Bidirectional LSTM output sequence for each chamber Indicates a cascading operation. It is the output projection matrix. This represents the multi-head attention extraction process. This is the final output. Combining the two-cell layer (mean cell and maximal cell), the mean cell focuses on capturing the overall trend of the ventricles, while the maximal cell focuses on the peak characteristics of the heart. To comprehensively judge the information, the calculation formula is as follows: in , , Then, the encoding dimension is reduced by using three fully connected layers, calculated as follows: Obtain the encoding vector of the final dynamic stream .
5. The text-based intelligent assessment method for cardiac function based on a large visual-language model according to claim 4, characterized in that, Step 4: Multi-scale feature fusion, utilizing the semantic stream coding features learned in Step 3. Time-stream coding features Visual encoding function and dynamic stream coding characteristics Connecting these features to form high-quality coding features with multiple perspectives, levels, and strong interpretability is achieved through the following steps: And use a multi-head attention mechanism to dynamically assign weights to each encoded feature: in This represents the complete input encoding features. This indicates the enhanced encoding function after attention processing. This represents the multi-head attention extraction process.
6. The text-based intelligent assessment method for cardiac function based on a large visual-language model according to claim 5, characterized in that, Step 5: Cardiac function (ejection fraction) prediction. Based on the design of the previous four steps, the final encoded features will contain complete semantic, temporal, visual, and dynamic information, and establish correlations between different levels. Finally, Tanh (·) linear mapping constraints are used, as follows: in It is a two-layer fully connected network, and It is the tanh mapping function. Subsequently, inverse normalization is performed to generate a prediction of cardiac function (EF value).
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the text-based intelligent assessment method for cardiac function based on a large visual language model as described in any one of claims 1 to 6.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the text-based intelligent assessment method for cardiac function based on a large visual language model as described in any one of claims 1 to 6.