Intelligent echocardiography heart function accurate evaluation system and method

By integrating the C3D network and graph convolutional network (GCN) model that integrates spatial and temporal feature attention, the problems of spatiotemporal feature extraction and chamber relationship modeling in cardiac function assessment are solved, accurate prediction and non-invasive assessment of ejection fraction are achieved, and the accuracy and efficiency of cardiac health testing are improved.

CN120616602APending Publication Date: 2025-09-12SUN YAT SEN MEMORIAL HOSPITAL SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510710889.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing cardiac function assessment methods have deficiencies in spatiotemporal feature extraction and inter-chamber relationship modeling, resulting in limited assessment accuracy, reliance on manual operation, and high cost.

Method used

The C3D network that integrates spatial feature attention and temporal feature attention mechanisms is used to extract spatiotemporal features, and the relationship between the various chambers of the heart is constructed through the graph convolutional network (GCN) model. Combined with image segmentation and preprocessing technology, accurate prediction of ejection fraction is achieved.

Benefits of technology

It significantly improves the prediction accuracy of ejection fraction EF, reduces human error, lowers assessment costs, and provides a non-invasive method for detecting heart health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120616602A_ABST
    Figure CN120616602A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent echocardiography heart function accurate evaluation system and method, and the system collects a heart image through an ultrasonic instrument, determines the ventricular diastole and contraction period through a data processing module, employs an image segmentation technology to extract the heart image, employs a cubic interpolation method to standardize the size of the image, and achieves the accurate evaluation of the heart function. The system constructs a graph model, spatial and temporal features of a heart image are extracted through a spatial and temporal feature extraction module by using a 3D convolutional neural network of a C3D structure fusing spatial and temporal features, a graph convolutional network module performs feature aggregation and message passing on the image through a multi-layer edge condition convolution ECC layer to obtain feature descriptors fusing all cavity information, and the feature descriptors are used for analyzing the cavity information. The ejection fraction prediction module inputs the characteristics into a full-connection layer for ejection fraction prediction; the invention innovatively provides an intelligent analysis framework based on a cardiac ultrasound dynamic image video, and aims to realize comprehensive and accurate evaluation of cardiac functions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent echocardiographic cardiac function accurate evaluation system and method. Background Art

[0002] Cardiac function assessment is an important task in clinical medicine, involving the comprehensive diagnosis and monitoring of cardiac health status. Traditional methods of cardiac function assessment mainly rely on technologies such as electrocardiogram (ECG), echocardiography, and hemodynamic testing. Although these methods are widely used, they still have some significant problems. First, traditional methods require high professional knowledge of medical staff and require comprehensive interpretation of data from multiple different sources. They are easily affected by human errors and are highly dependent on the operational skills of medical staff. Secondly, the process of cardiac function assessment is often time-consuming and costly, especially some invasive operations, which may cause certain harm to patients. Therefore, the accuracy and operability of traditional methods are limited, and there are certain technical bottlenecks.

[0003] In recent years, with the rapid development of machine learning and deep learning technologies, cardiac function assessment methods based on these technologies have gradually become a hot topic in research and clinical practice. Deep learning methods have been widely applied in the evaluation of echocardiographic data, achieving considerable success. By building efficient models, deep learning can automatically learn underlying patterns from large-scale data, achieving higher accuracy and efficiency than traditional methods. Furthermore, deep learning can automatically extract spatiotemporal features, reducing human error and thus improving the accuracy of cardiac function assessment.

[0004] However, although machine learning and deep learning techniques have shown great potential in cardiac function assessment, current methods still have some shortcomings. First, deep learning models may lose effective spatiotemporal information when processing spatiotemporal feature maps, resulting in a decrease in model performance. The morphological changes of the heart in different cycles have complex spatiotemporal characteristics, and the feature extraction steps of existing models often cannot fully capture these features, thus affecting the accuracy of the assessment. Secondly, the various chambers of the heart (such as the left ventricle (LV), right ventricle (RV), left atrium (LA) and right atrium (RA)) are closely related in cardiac function, but existing cardiac function assessment methods lack in-depth exploration of the relationship between these chambers. The interaction between different chambers is a key factor affecting the overall function of the heart. How to effectively mine and utilize these correlation features remains an important research challenge.

[0005] Therefore, while existing cardiac function assessment methods based on machine learning and deep learning have achieved breakthroughs in some areas, many technical challenges remain to be overcome, particularly in the effective extraction of spatiotemporal features and modeling of inter-chamber relationships. Addressing these challenges will further advance cardiac function assessment technology and potentially lead to more accurate, rapid, and non-invasive methods for monitoring heart health, significantly improving the efficiency of clinical diagnosis and patient treatment. Summary of the Invention

[0006] The purpose of the present invention is to provide an intelligent echocardiography cardiac function accurate assessment system and method, which uses a C3D network that integrates spatial feature attention and temporal feature attention mechanisms to extract efficient and robust spatiotemporal features from echocardiography videos; then constructs a graph convolutional network (GCN) model, defines each cardiac chamber as a node and calculates multiple edge attributes to fully explore the relationship between cardiac chambers; finally, adopts a feature fusion strategy to integrate multi-source information, and realizes accurate prediction of ejection fraction through a fully connected layer, providing clinicians with a reliable decision-making support tool.

[0007] The technical solution adopted by the present invention to solve its technical problem is:

[0008] An intelligent echocardiography cardiac function accurate assessment system, including:

[0009] Ultrasound video acquisition module: used to acquire cardiac images through ultrasound. Each image is acquired at least twice, and at least three consecutive cardiac cycles are acquired each time. The acquired data must include the left ventricular end-diastolic volume (LVEDV) and the left ventricular end-systolic volume (LVESV).

[0010] Data processing module: used to locate the time point T0 when ventricular relaxation begins, the time point T1 when the ventricular volume reaches its maximum and the ventricle begins to contract, and the time point T2 when the ventricular volume reaches its minimum and the ventricle begins to relax again. The video is cropped with [T0, T1] as a diastolic cycle and [T1, T2] as a systolic cycle to obtain a cardiac ultrasound animation;

[0011] Image segmentation module, which uses image segmentation technology to segment the left ventricle LV, right ventricle RV, left atrium LA, and right atrium RA from the cardiac ultrasound dynamic image, and uses cubic interpolation method to standardize the image to a uniform size;

[0012] A graph construction module, configured to define nodes of a graph G, wherein the nodes include a left ventricle LV, a right ventricle RV, a left atrium LA, and a right atrium RA;

[0013] The spatiotemporal feature extraction module uses a 3D convolutional neural network with a C3D structure that integrates spatial feature attention and temporal feature attention mechanisms to extract spatiotemporal features of cardiac images.

[0014] A similarity calculation module is used to calculate similarity indicators such as position similarity, functional similarity, time similarity, mutual information similarity, and blood flow similarity, and calculate the attributes of the edges in the graph;

[0015] The GCN module is used to perform feature aggregation and message passing on the graph G through the GCN. The GCN model contains three edge-conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes.

[0016] The ejection fraction prediction module performs feature fusion on the aggregated node features and inputs them into the fully connected layer for ejection fraction prediction. The mean square error (MSE) is used as the loss function, and the Adam optimization algorithm is used to optimize the model, ultimately outputting the cardiac ejection fraction.

[0017] Another technical problem to be solved by the present invention is to provide a method for accurately evaluating cardiac function using intelligent echocardiography, comprising the following steps:

[0018] Use ultrasound to acquire cardiac images, with each image acquired at least twice and at least three consecutive cardiac cycles each time, to obtain the left ventricular end-diastolic volume (LVEDV) and left ventricular end-systolic volume (LVESV);

[0019] Locate the time point T0 when ventricular relaxation begins, the time point T1 when the ventricular volume reaches its maximum and the ventricle begins to contract, and the time point T2 when the ventricular volume reaches its minimum and the ventricle begins to relax again, and crop the video with [T0, T1] as a diastolic cycle and [T1, T2] as a systolic cycle to obtain a cardiac ultrasound animation;

[0020] Image segmentation technology was used to segment the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA) from cardiac ultrasound images, and the images were normalized to a uniform size using cubic interpolation.

[0021] defining nodes of a graph G, wherein the nodes include a left ventricle LV, a right ventricle RV, a left atrium LA, and a right atrium RA;

[0022] Extract the spatiotemporal features of cardiac images using a 3D convolutional neural network with a C3D structure that integrates spatial and temporal feature attention mechanisms.

[0023] Calculate similarity indices such as positional similarity, functional similarity, temporal similarity, mutual information similarity, and blood flow similarity, and calculate the attributes of edges in the graph;

[0024] The graph G is subjected to feature aggregation and message passing through the graph convolutional network (GCN) model. The GCN model contains three edge-conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes.

[0025] After three layers of ECC feature aggregation and message passing, a feature descriptor that integrates all cavity information is generated. And the generated feature descriptor The fully connected layer is input to predict the ejection fraction. The mean square error (MSE) is used as the loss function, and the Adam optimization algorithm is used to optimize the model, and finally the cardiac ejection fraction is output.

[0026] Preferably, the spatiotemporal features of the cardiac image are extracted by using a 3D convolutional neural network with a C3D structure that integrates the spatial feature attention and temporal feature attention mechanisms. The method for extracting the spatiotemporal features is as follows:

[0027] The spatial attention function f is calculated through operations such as spatial convolution layers and global average pooling. s (X), generates the weight A for each pixel position s , these weights A s Applied to the original feature map X, it implements element-by-element weighted adjustment, namely:

[0028]

[0029] Among them, the original feature map X s is a spatiotemporal feature map extracted from echocardiography video; A s is the spatial attention weight; f s (X s ) is the output feature of the spatial convolution layer, which is the input feature map X s Features obtained after spatial convolution operation; d s is the feature dimension of the spatial dimension; the weighted adjusted feature map X s′ , is to set the spatial attention weight A s Applied to the original feature map X s The results obtained after

[0030] The temporal attention function f is calculated through the temporal convolution layer and global average pooling t (X), generate the weights for each time step, and add these weights A t Application to echocardiographic sequence X t , perform weighted adjustment on the features of each time step, namely:

[0031]

[0032] Among them, echocardiographic sequence Xt , is a sequence of images of the heart at different time points; A t is the temporal attention weight, which is used to adjust the weights of different time steps in the echocardiographic sequence to better capture the dynamic changes of cardiac motion; f t (X t ) is the output feature of the temporal convolution layer, which is the input echocardiogram sequence X t Features obtained after temporal convolution operation; d t is the characteristic dimension of the time dimension; the weighted adjusted echocardiographic sequence X t′ , is to convert the time attention weight A t Applied to the original echocardiographic sequence X t The result obtained after.

[0033] Preferably, the method for calculating the similarity indices of position similarity, functional similarity, time similarity, mutual information similarity and blood flow similarity is:

[0034] Position similarity E 1xm : If the two chambers are adjacent or have direct blood flow connection, the position similarity is 1, otherwise it is 0;

[0035] Functional similarity E 2xm : The similarity between the depth feature vectors of two cavities is measured by the Euclidean distance, and the calculation formula is:

[0036]

[0037] Among them, Z x and Z m are the feature vectors of nodes x and m respectively;

[0038] Temporal similarity E 3xm : The similarity between the main frequencies of the two chambers is calculated by time-frequency analysis. The calculation formula is:

[0039]

[0040] Among them, ST′ x and ST′ m are the frequency features of the spatiotemporal feature maps of nodes x and m after fast Fourier transform;

[0041] Mutual information similarity E 4xm : Mutual information is an indicator that measures the degree of information sharing between two random variables. Its calculation formula is:

[0042]

[0043] Where X and Y are the eigenvectors of the two chambers, p(x, y) is the joint probability distribution, and p(x) and p(y) are the marginal probability distributions;

[0044] Blood flow similarity E 5xm : Use the optical flow method to extract blood flow velocity and direction data at each time point. Calculate the optical flow field by minimizing the brightness change between adjacent frames. For each pair of adjacent frames, calculate the optical flow field. The specific formula is as follows:

[0045] v(x,y,t)=(u(x,y,t),v(x,y,t))

[0046] Among them, u(x, y, t) and v(x, y, t) represent the horizontal and vertical velocity components at position (x, y) and time t, respectively. The velocity and direction information of each pixel is extracted from the optical flow field. The magnitude of the velocity is calculated by the modulus of the vector, and the direction is calculated by the angle of the vector:

[0047]

[0048] The extracted blood flow velocity and direction information is converted into feature vectors.

[0049] As a preferred method, the method for calculating the attributes of the edges in the graph is:

[0050] The final edge attribute E xm is the above position similarity E 1xm , functional similarity E 2xm , time similarity E 3xm , mutual information similarity E 4xm and blood flow similarity E 5xm The synthesis of:

[0051] E xm =E 1xm +E 2xm +E 3xm +E 4xm +E 5xm .

[0052] Preferably, feature aggregation and message passing are performed on the graph G through a graph convolutional network (GCN) model. The GCN model includes three edge conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes.

[0053] The input layer receives the node feature matrix , where d represents the dimension of each node feature. In the first ECC layer, the input feature dimension d is mapped to d1, and the filter generates the network F lUsing edge attributes as input, a weight matrix is ​​generated to weight the features of adjacent nodes, thereby achieving initial feature aggregation and abstraction. The second ECC layer maps the feature dimension d1 to d2, further deepening the abstraction and fusion of features. The third ECC layer maps the feature dimension d2 to d3, generating a more advanced and comprehensive feature representation.

[0054] Each node updates its own feature vector by aggregating the feature vectors of its adjacent nodes and edges. The formula for the ECC operation is:

[0055]

[0056] Among them, F l It is a filter generation network with edge attributes E xm As input, a weight matrix is ​​generated to weight the features of adjacent nodes. l is the learnable weight, b l Is the bias term. The filter generates the network F l It is usually a small neural network that can dynamically generate a weight matrix based on edge attributes. In each layer, the weight matrix w l and the bias term b l Updates are made through backpropagation and optimization algorithms to minimize the loss function.

[0057] As a preference, after feature aggregation and message passing of three ECC layers, a feature descriptor integrating all cavity information is generated. The method is:

[0058] The first ECC layer maps the input feature dimension d to d1 and performs preliminary feature aggregation;

[0059] The second ECC layer maps the feature dimension d1 to d2 to further deepen the abstraction and fusion of features;

[0060] The third ECC layer maps the feature dimension d2 to d3 to generate the final feature descriptor

[0061] As a preference, the generated feature descriptor Input the fully connected layer to predict the ejection fraction. The mean square error (MSE) is used as the loss function. The Adam optimization algorithm is used to optimize the model. The final method to output the cardiac ejection fraction is:

[0062] The mean square error (MSE) is used as the loss function to measure the difference between the predicted value and the true value. The mathematical expression of the loss function is:

[0063]

[0064] Where N is the number of training samples, is the predicted ejection fraction of the ith sample, is the true ejection fraction of the i-th sample, optimized by the Adam algorithm with a learning rate of 10 -4 , and adopts a piecewise constant decay strategy to minimize the loss function, realize model training and optimization, and finally obtain accurate ejection fraction prediction results.

[0065] Another technical problem to be solved by the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements any of the intelligent echocardiography cardiac function accurate assessment methods described above.

[0066] Another technical problem to be solved by the present invention is to provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described intelligent echocardiography methods for accurately evaluating cardiac function.

[0067] The beneficial effects of the present invention are:

[0068] An intelligent echocardiographic cardiac function assessment architecture was constructed, consisting of modules including ultrasound video acquisition, preprocessing, graph construction, graph convolutional network (GCN) model design, feature fusion, and ejection fraction prediction. Ultrasound video acquisition acquires high-quality images. Preprocessing improves data accuracy through cropping, masking, image segmentation, and normalization. Graph construction defines each cardiac chamber as a node and calculates multiple edge attributes to explore inter-chamber relationships. The GCN model design aggregates node features through edge-conditional convolution operations. Feature fusion and the ejection fraction prediction module ultimately achieve accurate ejection fraction prediction. Through joint training, the modules collaborate with each other, significantly improving the accuracy of ejection fraction prediction.

[0069] To address the issues of effective information loss and insufficient mining of correlation features between different chambers in the processing of cardiac ultrasound animations, the present invention has made innovative designs in multiple links. In the preprocessing stage, image segmentation technology is used to accurately segment the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA), providing accurate information for feature extraction; in the graph construction stage, edge attributes are comprehensively defined, including position similarity, functional similarity, temporal similarity, mutual information similarity, blood flow similarity, etc., to fully mine the correlation features between different chambers; in the GCN model design, a C3D network with an attention mechanism is used to extract the spatiotemporal features of ultrasound cardiac videos. The spatial feature attention mechanism highlights the key areas of the heart, and the temporal feature attention mechanism better captures the dynamic changes of cardiac motion, thereby retaining effective information to the greatest extent during the feature extraction process and improving the expressive power of features. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 This is a principle block diagram of an intelligent echocardiography cardiac function accurate assessment system of the present invention;

[0071] Figure 2 The present invention is a flowchart of an intelligent echocardiographic cardiac function accurate assessment method. Specific implementation methods

[0072] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention. The following paragraphs describe the present invention in more detail by way of example with reference to the accompanying drawings. The advantages and features of the present invention will become more apparent from the following description and claims. It should be noted that the drawings are all in a very simplified form and are not in exact proportions. They are only used to facilitate and clearly illustrate the purpose of the embodiments of the present invention.

[0073] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, features defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be directly connected, or indirectly connected through an intermediate medium, or it can be internal communication between two elements.

[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Example

[0075] See Figure 1 As shown, the hardware foundation required for the implementation of the relevant technical solutions of this application can be deployed according to the architecture shown in the figure.

[0076] Although the various methods of the present application are described based on the same concept and thus exhibit commonality, unless otherwise specified, these methods can be independently executed, as those skilled in the art should be aware of.

[0077] First, see Figure 2 The present invention provides an intelligent echocardiographic method for accurately assessing cardiac function. In a typical embodiment, the method includes five main steps: ultrasound video acquisition, ultrasound video preprocessing, graph construction, graph convolutional network (GCN) model design, and feature fusion and ejection fraction prediction. The main process of the method is divided into the following steps:

[0078] Generally speaking, when acquiring images, skilled ultrasound technicians usually ensure that each image is acquired at least twice, and each acquisition covers at least three consecutive cardiac cycles. For patients with arrhythmias, the technical requirements are more stringent, and each acquisition must ensure at least five consecutive cardiac cycles, and the dynamic cardiac ultrasound images must be saved. During the acquisition process, measurements are especially performed through the apical four-chamber view to ensure that the endocardial line is clearly visible and to avoid short-circuiting of the left ventricle. At the end of diastole and end of systole of the left ventricle, the technician should outline along the trajectory of the endocardial line to obtain the left ventricular end-diastolic volume (LVEDV) and the left ventricular end-systolic volume (LVESV). This type of data acquisition method not only helps to provide a more accurate assessment of cardiac function, but also provides doctors with vital clinical information.

[0079] It should be noted that by acquiring data from multiple consecutive cardiac cycles each time and ensuring that each image is acquired at least twice, especially for patients with arrhythmias, where five consecutive cardiac cycles are acquired, image stability and consistency can be effectively improved. This method helps reduce accidental errors, thereby improving the measurement accuracy of left ventricular end-diastolic volume (LVEDV) and left ventricular end-systolic volume (LVESV), and thus enhancing the accuracy of cardiac function assessment.

[0080] During the acquisition process, special emphasis is placed on measuring through the apical four-chamber view, ensuring clear visualization of the endocardium and avoiding left ventricular short-circuiting. This procedure not only ensures image quality but also ensures that important structural details are not missed, laying the foundation for providing high-quality imaging data, thereby helping doctors make more reliable diagnoses.

[0081] Furthermore, the preservation of dynamic ultrasound images provides valuable data for subsequent analysis, research, and review. These images can not only be used for long-term monitoring of current patients, but also provide data support for scientific research, thereby promoting further development and exploration in the field of cardiology.

[0082] The ultrasound video preprocessing steps are as follows:

[0083] In echocardiography, the systolic and diastolic cycles are important indicators for assessing cardiac function, especially changes in myocardial thickness and area. Therefore, accurately cropping the video and selecting the cycle of interest is crucial for the accuracy and reliability of subsequent data. Typically, this process is marked by experienced physicians: First, the physician determines the time point T0 when ventricular relaxation begins. Then, the physician finds the time point T1 when the ventricular volume reaches its maximum and the ventricles begin to contract, and the time point T2 when the ventricular volume reaches its minimum and the ventricles begin to relax again. Thus, [T0, T1] is a diastolic cycle, and [T1, T2] is a systolic cycle.

[0084] After completing video cropping, the present invention performs a masking operation on the cropped image sequence to remove interfering information unrelated to the cardiac structure, such as text, electrocardiogram marks, and spirometers. Since the contrast between the cardiac structure and background noise is low and it is easily affected by noise, the present invention uses image segmentation technology (such as U-Net) to separate the cardiac structure from the background. At the same time, tracking technology is used to track the position and shape changes of the cardiac structure at different time points in real time, thereby improving the accuracy and stability of the data.

[0085] Furthermore, to ensure the feature extraction network can process ultrasound images of varying sizes, and considering the size and resolution differences between different ultrasound devices and operators, this paper uses cubic interpolation to normalize images of varying pixel sizes into a uniform W×H ultrasound animation sequence. In practice, ultrasound devices used by different medical institutions may have varying imaging parameters, resulting in significant differences in pixel size and resolution. Furthermore, the operating habits and settings of different operators may further exacerbate these inconsistencies. If unprocessed images are directly fed into the feature extraction network, the mismatch in size and resolution may prevent the model from effectively learning, thus affecting prediction accuracy.

[0086] Image segmentation technology accurately segments the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA) from cardiac color Doppler ultrasound images, enabling clear acquisition of the boundaries and regional information of each chamber, providing a solid foundation for subsequent feature extraction. The segmented chamber images not only facilitate accurate measurement of the area, perimeter, and shape changes of each chamber, but also provide key data for further analysis of the heart's dynamic function. The application of this segmentation technology enables the extraction of pure and accurate cardiac structural information from complex cardiac color Doppler ultrasound images, eliminating irrelevant interference and thereby improving the reliability and efficiency of the entire analysis process.

[0087] By precisely cropping the video and selecting the period of interest, we ensure more accurate data on the cardiac systolic and diastolic cycles. This process reduces the impact of irrelevant time periods, enabling more precise assessment of cardiac function in subsequent analysis, thus providing physicians with more reliable clinical data.

[0088] Image segmentation techniques are used to separate cardiac structures from the background, and masking operations are used to remove interfering information such as text and ECG markers, significantly reducing the impact of noise. This method ensures clearer and more accurate cardiac structural information extracted from cardiac ultrasound animations, further facilitating feature extraction and dynamic functional analysis.

[0089] Finally, cubic interpolation was used to standardize images collected by different ultrasound devices and operators to a uniform size, effectively addressing the differences in equipment and resolution. This standardization ensured that the feature extraction network could handle ultrasound images of various sizes, thereby improving the model's learning and prediction accuracy, making ultrasound image analysis more adaptable in different medical settings.

[0090] The steps for constructing a graph are as follows:

[0091] First, node definition and feature extraction are performed; then edge attribute calculation is performed.

[0092] The steps for node definition and feature extraction are as follows:

[0093] Node Definition: In the ejection fraction prediction task based on cardiac color Doppler ultrasound, the first step in constructing the graph G is node definition. The heart is a complex pumping organ, with its left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA), each with distinct structural and functional characteristics. These chambers work together to ensure efficient blood circulation. The left ventricle (LV), the heart's primary pumping chamber, is responsible for pumping oxygen-rich blood into the aorta, which then supplies various tissues and organs throughout the body. The right ventricle (RV) pumps deoxygenated blood from the systemic circulation into the pulmonary artery, completing the oxygenation process. The left atrium (LA) receives oxygenated blood from the pulmonary veins and delivers it to the left ventricle; the right atrium (RA) collects deoxygenated blood from the systemic circulation and delivers it to the right ventricle. These chambers occupy specific locations within the heart's anatomy and maintain normal cardiac function through precise contraction and relaxation.

[0094] Therefore, using the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA) as nodes in the graph G not only aligns with the anatomical and physiological characteristics of the heart, but also fully utilizes the rich spatiotemporal information contained in cardiac ultrasound images. This node definition provides a solid foundation for subsequent feature extraction and graph convolutional network (GCN) processing, helping to improve the accuracy and reliability of ejection fraction prediction.

[0095] By using the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA) as nodes in graph G, the structural and functional characteristics of the heart can be effectively reflected. This node configuration closely aligns with the heart's actual anatomy and blood flow patterns, more accurately capturing the dynamic changes of each chamber and providing more representative and efficient data for ejection fraction prediction.

[0096] Cardiac ultrasound animations contain rich spatiotemporal information. Leveraging this information and defining key chambers as nodes effectively captures the inter-chamber relationships and their dynamic behavior over time. This effective use of spatiotemporal information provides more features for graph convolutional networks (GCNs), helping to improve the model's expressiveness and prediction accuracy.

[0097] This node definition not only ensures high consistency between the data and cardiac structure and function, but also lays a solid foundation for subsequent feature extraction and graph convolutional network processing. By accurately dividing the different cardiac chambers into nodes, the G graph can better facilitate information flow and feature integration, thereby improving the effectiveness and reliability of ejection fraction prediction.

[0098] Spatiotemporal feature extraction: Cardiac echocardiography is a non-invasive medical imaging technology that is widely used to assess a patient's heart health. This technology records the movement of the heart, volume changes, and blood flow dynamics through continuous image frames. However, these video data are large and complex, and require precise analysis to extract useful cardiac information, such as cardiac systolic and diastolic characteristics, as well as ventricular wall texture, shape, and motion characteristics. In order to accurately capture the spatiotemporal information of echocardiography and obtain accurate cardiac function data, the present invention adopts a three-dimensional convolutional neural network (3DCNN) with a C3D structure that integrates spatial feature attention and temporal feature attention mechanisms. This network architecture is widely used in the field of video analysis. Its uniqueness lies in that the convolutional layer can consider both the temporal and spatial dimensions, thereby more efficiently processing dynamic images.

[0099] The C3D (Convolutional 3D) network is a deep learning model specifically optimized for medical image video analysis, and is particularly suitable for processing dynamic medical images such as echocardiograms. The core innovation of this network is that its convolution operation can be performed simultaneously on the time axis and the spatial axis, enabling the network to capture the time-varying cardiac structure and functional characteristics in ultrasound image sequences. Specifically, the C3D network introduces three-dimensional convolution kernels in the convolution layer. These convolution kernels not only operate on the spatial dimension (i.e., the width and height of the image), but also slide and calculate on the time dimension (i.e., the frame sequence). This design enables the C3D network to capture the temporal dynamic relationship between frames while extracting the spatial features of the image.

[0100] In each 3D convolutional layer, the convolution kernel operates on each frame in the video and its adjacent frames, extracting structural changes and motion patterns of the heart at different moments. Through this combined spatiotemporal feature extraction approach, the stacked 3D convolutional layers gradually abstract and capture complex spatiotemporal features in the video, such as the heart's contraction and relaxation patterns and blood flow dynamics. These features provide richer and more accurate information for diagnosing heart lesions and assessing cardiac function, making the C3D network valuable in medical image analysis.

[0101] In addition, C3D networks usually combine pooling layers and fully connected layers to further refine features and complete the final classification or regression tasks. The pooling layer helps to reduce the dimensionality of the data while retaining the sensitivity of key information and extracting key spatiotemporal features; while the fully connected layer is responsible for mapping the learned high-level features to specific cardiac function parameters, such as ejection fraction (EF), end-diastolic volume (EDV), and end-systolic volume (ESV). The introduction of the C3D network has significantly improved the ability to extract cardiac function information from echocardiographic videos, making the automatic assessment of cardiac function faster and more accurate, and providing clinicians with a powerful auxiliary diagnostic tool. This automated analysis method can greatly reduce the workload of doctors and improve the consistency and reliability of diagnosis. Unlike traditional C3D networks, the present invention introduces spatial feature attention mechanisms and temporal feature attention mechanisms to enhance the ability to capture spatiotemporal characteristics in echocardiographic video analysis. The spatial attention mechanism enables the model to dynamically adjust the correlation between different pixel positions within each frame, thereby highlighting the key areas of the heart. This is achieved by calculating the spatial attention weight A. s Specifically, we calculate the spatial attention function f through operations such as spatial convolution layers and global average pooling. s (X), generate weights for each pixel position. Then, these weights A s It is applied to the original feature map X to achieve element-by-element weighted adjustment, namely:

[0102]

[0103] Among them, the original feature map X s is a spatiotemporal feature map extracted from echocardiography video; A s is the spatial attention weight; f s (X s ) is the output feature of the spatial convolution layer, which is the input feature map X s Features obtained after spatial convolution operation; d s is the feature dimension of the spatial dimension; the weighted adjusted feature map X s′ , is to set the spatial attention weight A s Applied to the original feature map X s The result obtained after.

[0104] At the same time, the temporal attention mechanism is used to adjust the weights between different time steps to better capture the dynamic changes of cardiac motion. When processing echocardiographic sequences, we calculate the temporal attention weight A t Specifically, operations such as temporal convolutional layers and global average pooling are used to calculate the temporal attention function f t (X), generate the weights for each time step. Then, these weights A t Applied to echocardiographic sequence X t , perform weighted adjustment on the features of each time step, namely:

[0105]

[0106] Among them, echocardiographic sequence X t , is a sequence of images of the heart at different time points; A t is the temporal attention weight, which is used to adjust the weights of different time steps in the echocardiographic sequence to better capture the dynamic changes of cardiac motion; f t (X t ) is the output feature of the temporal convolution layer, which is the input echocardiogram sequence X t Features obtained after temporal convolution operation; d t is the characteristic dimension of the time dimension; the weighted adjusted echocardiographic sequence X t′ , is to convert the time attention weight A t Applied to the original echocardiographic sequence X t The result obtained after.

[0107] The introduction of the spatiotemporal attention mechanism provides a powerful tool for processing cardiac ultrasound moving images. It allows the model to dynamically adjust the weights at different moments or image positions to better focus on key areas in the cardiac image. For example, it can automatically identify the motion trajectory of the heart's ventricular walls to assess cardiac function. In addition, spatiotemporal attention helps reduce redundant information and improve data utilization efficiency, thereby more accurately analyzing cardiac ultrasound moving images. In addition, considering the long-term dependencies and information richness in cardiac ultrasound moving images, the spatiotemporal attention mechanism can help the model better capture these characteristics, improve the accuracy and comprehensiveness of cardiac information, and deeply explore the spatiotemporal characteristics of ultrasound videos, providing more possibilities in the field of medical diagnosis.

[0108] The C3D network simultaneously processes both time and space dimensions in three-dimensional convolutional layers, enabling the model to capture the spatiotemporal dynamic characteristics of the heart, such as contraction and relaxation patterns and ventricular wall motion. This approach better reflects cardiac motion and structural changes, improving the accuracy of cardiac function analysis. By integrating spatial and temporal feature attention mechanisms, the model dynamically adjusts the weights of different time steps and spatial locations, focusing more closely on key areas in cardiac ultrasound images. This enables the model to more accurately extract features closely related to cardiac function, improving the reliability of cardiac health assessments.

[0109] The spatiotemporal attention mechanism automatically identifies and highlights important components of cardiac images, such as the motion trajectory of the ventricular wall, providing more accurate information for clinical diagnosis. This automated analysis significantly reduces physician workload and improves diagnostic consistency and speed. The spatiotemporal attention mechanism helps suppress the influence of redundant information, enabling the model to more efficiently utilize data when processing large and complex cardiac ultrasound animations. This not only improves data utilization efficiency but also enhances sensitivity to important features.

[0110] There are long-term dependencies in cardiac ultrasound animations. The spatiotemporal attention mechanism can help the model better capture these long-term characteristics, thereby establishing effective relationships between different time steps and improving the accuracy and comprehensiveness of cardiac information. By introducing the spatiotemporal attention mechanism, the C3D network can provide more possibilities in medical image analysis, especially in the diagnosis of heart lesions and the assessment of cardiac function, and has broad application potential.

[0111] Edge attributes are calculated as follows:

[0112] Calculate the position similarity (E 1xm): The anatomical positional relationships and blood flow connections between cardiac chambers have a significant impact on cardiac function. If two chambers are adjacent or have a direct blood flow connection, their positional similarity is 1; otherwise, it is 0. For example, the left and right ventricles are adjacent and have a blood flow connection, so their positional similarity is 1; on the other hand, the left ventricle and right atrium are not directly adjacent, so their positional similarity is 0. This similarity helps the model understand the physical and functional adjacency between cardiac chambers.

[0113] Calculate the functional similarity (E 2xm ): Each chamber of the heart performs different functions during the cardiac cycle. For example, the left ventricle is primarily responsible for pumping blood into the aorta, while the right ventricle is responsible for pumping blood into the pulmonary artery. The similarity between the depth feature vectors of two chambers is measured by the Euclidean distance, which is calculated as follows:

[0114]

[0115] Among them, Z x and Z m are the feature vectors of nodes x and m, respectively. Feature vectors can include features such as the area, perimeter, and shape change of a chamber. A smaller Euclidean distance indicates a high degree of functional similarity between the two chambers, while a smaller distance indicates a low degree of functional similarity. For example, the left and right ventricles may share certain similarities in features such as area change during systole and diastole, resulting in a relatively high functional similarity value.

[0116] Calculate the time similarity (E 3xm ): The movements and changes of the heart's chambers during the cardiac cycle have temporal regularity. For example, the contraction and relaxation of the left ventricle are temporally coordinated with the contraction and relaxation of the right ventricle. The similarity between the dominant frequencies of the two chambers is calculated using time-frequency analysis. The calculation formula is:

[0117]

[0118] Among them, ST′ x and ST′ m The frequency features of the spatiotemporal feature maps of nodes x and m after Fast Fourier Transform (FFT) are obtained. Frequency features reflect the dominant motion frequency of a chamber in a time series. If the dominant frequencies of two chambers are similar, their temporal similarity is high; otherwise, their temporal similarity is low. For example, the motion frequencies of the left and right ventricles during the cardiac cycle may be similar, resulting in a high temporal similarity value.

[0119] Calculate mutual information similarity (E 4xm): In cardiac function assessment, mutual information plays a key role in ejection fraction prediction. Mutual information is a measure of the degree of information sharing between two random variables, and its calculation formula is:

[0120]

[0121] Here, X and Y are the eigenvectors of the two chambers, p(x, y) is the joint probability distribution, and p(x) and p(y) are the marginal probability distributions. In cardiac function assessment, mutual information plays a key role in ejection fraction prediction. It can quantify the correlation between cardiac chambers, such as the synchronization of left and right ventricular contraction and relaxation, which directly affects the accuracy of ejection fraction. Mutual information can also identify functional abnormalities, such as asynchronous chamber movement in heart failure, which can change the normal range of ejection fraction. For example, there may be a certain correlation between the area change of the left ventricle and the area change of the right ventricle, and this correlation can be quantified using mutual information. A higher mutual information value indicates a higher degree of shared feature information between the two chambers, and vice versa.

[0122] Calculate blood flow similarity (E 5xm Cardiac ultrasound animations contain rich blood flow information, such as blood flow velocity and direction. By analyzing this information, we can calculate the blood flow similarity between two chambers. For example, the changes in blood flow velocity and direction between the left ventricle and the aorta during the cardiac cycle show a certain similarity, which can reflect the functional coordination between the left ventricle and the aorta. Calculating blood flow similarity helps the model better understand the blood flow relationship between the heart chambers.

[0123] The specific steps are as follows:

[0124] ① Extract blood flow velocity and direction information: From the pre-processed ultrasound dynamic image, the optical flow method is used to extract the blood flow velocity and direction data at each time point. The optical flow method tracks the movement of blood flow by calculating the optical flow field between adjacent frames. The optical flow field can be expressed as the motion vector of each pixel point in time, including speed and direction. Classic optical flow algorithms such as the Horn-Schunck algorithm or the Lucas-Kanade algorithm can be used. These algorithms estimate the optical flow field by minimizing the brightness change between adjacent frames. For each pair of adjacent frames, the optical flow field is calculated. The specific formula is as follows:

[0125] v(x,y,t)=(u(x,y,t),v(x,y,t)) (6)

[0126] Where u(x, y, t) and v(x, y, t) represent the horizontal and vertical velocity components at position (x, y) and time t, respectively. The velocity and direction information of each pixel is extracted from the optical flow field. The magnitude of the velocity can be calculated by the modulus of the vector, and the direction can be calculated by the angle of the vector:

[0127]

[0128] ② Calculate blood flow feature vectors: The extracted blood flow velocity and direction information is converted into feature vectors for subsequent similarity calculations. Feature vectors can include the following features: velocity features, such as mean velocity, peak velocity, velocity standard deviation, and velocity change rate; direction features, such as mean direction, direction standard deviation, and direction change rate; and comprehensive features, such as velocity-direction correlation and blood flow pattern index. Mean velocity reflects the average level of overall blood flow velocity, peak velocity represents the highest velocity in blood flow, velocity standard deviation reflects the degree of velocity fluctuation, and velocity change rate reflects the dynamic changes in blood flow velocity. Mean direction reflects the directional trend of overall blood flow, direction standard deviation reflects the degree of direction fluctuation, and direction change rate reflects the dynamic changes in blood flow direction. Velocity-direction correlation reflects the coordinated changes in velocity and direction. The blood flow pattern index combines velocity and direction features to calculate a comprehensive index reflecting the blood flow pattern.

[0129] In summary, the final edge attribute Exm is a combination of the above similarities, namely:

[0130] E xm =E 1xm +E 2xm +E 3xm +E 4xm +E 5xm (9)

[0131] This comprehensive edge attribute can comprehensively reflect the various relationships between cardiac chambers and provide rich information for subsequent graph convolutional network (GCN) processing.

[0132] By integrating multiple similarity indicators (such as position, function, time, mutual information, and blood flow), the complex relationship between chambers can be fully reflected, rather than being limited to just one aspect. This multi-level information fusion can capture the physical, functional, and dynamic connections between cardiac chambers, thereby improving the performance and robustness of the model; different types of similarity, such as functional similarity and temporal similarity, can more accurately measure the relationship between cardiac chambers. For example, functional similarity can quantify the changing characteristics of the chambers, such as area and shape, by calculating the Euclidean distance, while temporal similarity captures the periodicity and coordination of chamber movement through frequency analysis, further improving the accuracy of the analysis.

[0133] Combining blood flow similarity with other similarity calculation methods (such as optical flow) can effectively extract and process dynamic blood flow information in echocardiograms. This provides rich input data for subsequent graph convolutional networks, helping the model learn more complex spatiotemporal relationships and blood flow patterns. By calculating mutual information, the synchronization and collaboration between chambers can be quantified, especially in the prediction of important indicators such as ejection fraction, which can improve the accuracy of assessment. In particular, in pathological conditions such as heart failure, the asynchrony of chamber movement may affect the ejection fraction, and this method can help capture these subtle changes.

[0134] This method flexibly leverages multiple edge attributes within a graph convolutional network (GCN), providing powerful support for the diagnosis and analysis of heart disease. GCNs can process graph-structured data and, through precise modeling of edge attributes, further enhance the ability to learn and understand complex spatiotemporal relationships. This analysis method, based on spatiotemporal similarity, not only improves the automated assessment of cardiac function but also provides clinicians with an accurate auxiliary diagnostic tool, helping them to more effectively analyze and diagnose heart diseases. It has significant application value in automated assessment and real-time monitoring.

[0135] The steps for designing a graph convolutional network (GCN) model are as follows:

[0136] Network structure: The network structure of the graph convolutional network (GCN) model consists of three edge conditional convolution (ECC) layers. Each ECC layer undertakes the important tasks of feature aggregation and message passing. The input layer receives the node feature matrix Where d represents the dimension of each node feature. In the first ECC layer, the input feature dimension d is mapped to d1, and the filter generates the network F l A weight matrix is ​​generated using edge attributes as input, and the features of adjacent nodes are weighted, thereby achieving preliminary aggregation and abstraction of features. The second ECC layer maps the feature dimension d1 to d2, further deepening the abstraction and fusion of features. The third ECC layer maps the feature dimension d2 to d3, generating a more advanced and comprehensive feature representation. Throughout the network structure, each node updates its own feature vector by aggregating the feature vectors of its adjacent nodes and edges. This process not only retains the important information of the original features but also integrates the features of adjacent nodes, making the final feature descriptor richer and more comprehensive. The weights and biases are continuously updated through backpropagation and optimization algorithms to ensure that the model can gradually learn the optimal feature representation, providing strong support for the final ejection fraction prediction.

[0137] Through multiple edge-conditional convolution (ECC) layers, the graph convolutional network (GCN) model effectively aggregates and fuses features from adjacent nodes and their edge attributes at each layer. Each ECC layer weights the features of adjacent nodes by introducing edge attributes. This not only preserves the feature information of the original nodes but also fully utilizes the relevant information of adjacent nodes in the graph structure, ensuring a comprehensive and rich feature representation. This approach significantly enhances the model's capabilities, enabling each node to perceive a wider range of contextual information through the network structure.

[0138] The output of each ECC layer undergoes feature mapping and abstraction processing to generate more complex and high-level feature representations. As the network layers deepen, feature abstraction and fusion become increasingly in-depth, enabling the capture of complex nonlinear relationships between nodes. This hierarchical feature abstraction process enables GCN to gradually build complex high-level features from simple low-level features, significantly improving the model's learning capabilities. This is particularly true for complex tasks such as ejection fraction prediction, enabling more accurate predictions.

[0139] The model dynamically adjusts weights and biases through backpropagation and optimization algorithms to continuously optimize feature representations. Each node's feature updates are based on the feature vectors of its neighboring nodes and edge attributes, allowing the model to flexibly adapt to data changes and task requirements during training. This dynamic optimization process ensures that GCN can effectively adjust and discover the optimal feature representation for complex and variable cardiac data, providing efficient support for subsequent cardiac function assessment and prediction.

[0140] Edge Conditional Convolution Operation: In the GCN model, the edge conditional convolution (ECC) operation is the core link, responsible for updating the feature vector of the node in each layer. Specifically, each node updates its own feature vector by aggregating the feature vectors of its adjacent nodes and edges. The formula for the ECC operation is:

[0141]

[0142] Among them, F l It is a filter generation network with edge attributes E xm As input, a weight matrix is ​​generated to weight the features of adjacent nodes. l is the learnable weight, b l Is the bias term. The filter generates the network F l It is usually a small neural network that can dynamically generate a weight matrix based on edge attributes. In each layer, the weight matrix w l and the bias term b lUpdates are made through back-propagation and optimization algorithms to minimize the loss function. This process ensures the layer-by-layer abstraction and fusion of node features, enabling the model to capture the complex relationships between different chambers.

[0143] The ECC operation generates a network through filters, dynamically constructing a weight matrix based on edge attributes. This approach enables the model to weight the features of adjacent nodes differently at each layer based on the different edge attributes. This dynamic adjustment mechanism can more accurately reflect the relationships between edges when processing tasks with complex graph structures, thereby improving the model's adaptability and expressiveness on graph data.

[0144] Each node updates its own feature vector by aggregating the feature vectors of its neighboring nodes and edges. This approach allows the model to fully leverage the relationships between nodes in the graph, focusing on both the characteristics of the node itself and the influence of neighboring nodes and their connections. This information aggregation process helps the model abstract features layer by layer, enabling each node to integrate information about its neighbors, leading to a more comprehensive understanding of the complex interactions between nodes and edges in the graph.

[0145] In each layer, node features are updated and optimized layer by layer through weight matrices and bias terms. Through backpropagation and optimization algorithms, the weight matrices and bias terms are continuously adjusted to minimize the loss function. This enables the model to start from low-level features, gradually achieve abstraction and fusion, and then generate more advanced and comprehensive feature representations. Ultimately, these deep features can help the model more effectively capture the complexity of the graph structure and the nonlinear relationships between nodes, thereby significantly enhancing the model's learning ability and adapting to various complex graph tasks.

[0146] The steps of feature fusion and ejection fraction prediction are as follows

[0147] Feature fusion: After three layers of ECC feature aggregation and message passing, a feature descriptor that integrates all cavity information is generated. This process integrates the feature information of different chambers into a unified feature space through layer-by-layer feature abstraction and fusion. Specifically, the first ECC layer maps the input feature dimension d to d1 and performs preliminary feature aggregation. The second ECC layer maps the feature dimension d1 to d2, further deepening the feature abstraction and fusion. The third ECC layer maps the feature dimension d2 to d3 to generate the final feature descriptor. This feature descriptor contains comprehensive information of all chambers and provides a rich feature basis for subsequent ejection fraction prediction.

[0148] Through layer-by-layer feature abstraction and fusion through multi-layer ECC operations, each layer maps the feature space dimensions to a higher space, enabling the features at each layer to deeply capture more complex relationships and patterns. The first layer performs preliminary feature aggregation, the second layer further deepens feature abstraction, and the third layer generates higher-level comprehensive feature descriptors. This layer-by-layer feature fusion mechanism helps the model gradually refine and capture more complex features, helping to improve the model's understanding and prediction capabilities for complex problems.

[0149] This feature fusion method integrates information from different chambers into a unified feature space, enabling the model to comprehensively consider the relationships between each chamber. Feature aggregation at each layer not only updates and fuses local information but also strengthens the unified representation of different chamber features layer by layer. The resulting feature descriptor incorporates comprehensive information from all chambers, providing a richer feature foundation for subsequent tasks (such as ejection fraction prediction) and improving the accuracy and reliability of predictions.

[0150] Through this feature fusion process, the model can fully utilize the relationships and characteristics between different chambers, providing a more accurate and comprehensive feature description for subsequent tasks (such as ejection fraction prediction). This deeply integrated feature descriptor can better reflect the mutual influence between chambers, helping the model capture potential and complex dependencies, ultimately improving the performance and effectiveness of prediction tasks.

[0151] Ejection fraction prediction: After feature fusion, the generated feature descriptor The input is fed into a fully connected layer that maps the feature vector to a scalar value, the ejection fraction prediction value. Specifically, the fully connected layer uses an appropriate activation function (such as ReLU or Sigmoid) to ensure that the output value is within a reasonable range. To train the model, the mean squared error (MSE) is used as the loss function to measure the difference between the predicted value and the true value. The mathematical expression of the loss function is:

[0152]

[0153] Where N is the number of training samples, is the predicted ejection fraction of the ith sample, is the true ejection fraction of the i-th sample. Using the Adam optimization algorithm with a learning rate of 10⁻¹⁴ and a piecewise constant decay strategy to minimize the loss function, the model is trained and optimized, ultimately obtaining an accurate ejection fraction prediction result.

[0154] Using the mean squared error (MSE) as a loss function effectively measures the difference between predicted and true values. By calculating the squared difference between the predicted ejection fraction and the true value, MSE clearly guides model training and optimizes prediction accuracy. Because MSE is sensitive to large errors, it helps minimize large errors during model training, thereby improving the accuracy of the final predictions.

[0155] The Adam optimization algorithm can be used to efficiently optimize the parameters of the model. Adam combines the momentum method and the adaptive learning rate adjustment mechanism. It can adaptively adjust the learning rate of each parameter and has good convergence speed and stability. The learning rate is 10 -4 , and adopts a piecewise constant decay strategy, which can help the model gradually reduce the learning rate during the training process, thereby avoiding oscillation or overfitting in the later stages of training and further improving the prediction accuracy of the model.

[0156] By using appropriate activation functions in the fully connected layer, we can ensure that the output range of the ejection fraction prediction value is reasonable and meets practical application requirements. ReLU helps introduce nonlinear features to the model and enhances its expressiveness, while Sigmoid maps the output to a specific interval (for example, [0, 1]), keeping the output value within a reasonable range and helping the model function stably in various situations.

[0157] Accurate assessment of cardiac function is crucial for the early diagnosis, treatment decision-making, and prognosis of cardiovascular disease. With the rapid development of artificial intelligence (AI) technology in medical imaging, advanced algorithms such as deep learning and graph neural networks have brought new opportunities to echocardiographic analysis. However, traditional methods for assessing cardiac function still face numerous challenges. On the one hand, conventional echocardiographic measurement methods are highly operator-dependent, and differences in experience and skill levels among physicians can lead to inconsistent measurement results. On the other hand, existing analysis techniques often struggle to fully exploit the complex spatiotemporal features and correlations between cardiac chambers when processing cardiac ultrasound moving images, thus limiting the accuracy and comprehensiveness of the assessment. Furthermore, traditional methods are prone to losing valuable information during feature extraction and fail to fully exploit the correlation features between different chambers in echocardiographic videos, such as the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA). To overcome these limitations, the present invention innovatively proposes an intelligent analysis framework based on cardiac ultrasound moving images, aiming to achieve comprehensive and accurate assessment of cardiac function. The framework first uses the C3D network that integrates spatial feature attention and temporal feature attention mechanisms to extract efficient and robust spatiotemporal features from echocardiography videos; then constructs a graph convolutional network (GCN) model, which fully explores the relationships between cardiac chambers by defining each cardiac chamber as a node and calculating multiple edge attributes; finally, a feature fusion strategy is used to integrate multi-source information, and a fully connected layer is used to achieve accurate prediction of ejection fraction, providing clinicians with a reliable decision-making support tool.

[0158] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the intelligent echocardiography cardiac function accurate assessment method as described above is implemented.

[0159] This embodiment also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the intelligent echocardiography cardiac function accurate assessment method as described above is implemented.

[0160] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0161] Those skilled in the art will clearly understand that for the sake of convenience and brevity in description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0162] The above embodiments of the present invention are not intended to limit the scope of protection of the present invention, and the implementation methods of the present invention are not limited thereto. All other modifications, replacements or changes made to the above structures of the present invention based on the above contents of the present invention, in accordance with common technical knowledge and customary means in this field, without departing from the above basic technical ideas of the present invention, should fall within the scope of protection of the present invention.

Claims

1. An intelligent echocardiography cardiac function accurate assessment system, characterized by: Includes: Ultrasound video acquisition module: used to acquire cardiac images through ultrasound. Each image is acquired at least twice, and at least three consecutive cardiac cycles are acquired each time. The acquired data must include the left ventricular end-diastolic volume (LVEDV) and the left ventricular end-systolic volume (LVESV). Data processing module: used to locate the time point T0 when ventricular relaxation begins, the time point T1 when the ventricular volume reaches its maximum and the ventricle begins to contract, and the time point T2 when the ventricular volume reaches its minimum and the ventricle begins to relax again. The video is cropped with [T0, T1] as a diastolic cycle and [T1, T2] as a systolic cycle to obtain a cardiac ultrasound animation; An image segmentation module, which uses image segmentation technology to segment the left ventricle LV, right ventricle RV, left atrium LA, and right atrium RA from the cardiac ultrasound image, and uses cubic interpolation to standardize the image to a uniform size; A graph construction module, configured to define nodes of a graph G, wherein the nodes include a left ventricle LV, a right ventricle RV, a left atrium LA, and a right atrium RA; The spatiotemporal feature extraction module uses a 3D convolutional neural network with a C3D structure that integrates spatial feature attention and temporal feature attention mechanisms to extract spatiotemporal features of cardiac images. A similarity calculation module is used to calculate similarity indicators such as position similarity, functional similarity, time similarity, mutual information similarity, and blood flow similarity, and calculate the attributes of the edges in the graph; The GCN module is used to perform feature aggregation and message passing on the graph G through the GCN. The GCN model contains three edge-conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes. The ejection fraction prediction module performs feature fusion on the aggregated node features and inputs them into the fully connected layer for ejection fraction prediction. The mean square error (MSE) is used as the loss function, and the Adam optimization algorithm is used to optimize the model, ultimately outputting the cardiac ejection fraction.

2. An intelligent echocardiographic cardiac function accurate assessment method, characterized in that: The following steps are involved: Use ultrasound to acquire cardiac images, with each image acquired at least twice and at least three consecutive cardiac cycles each time, to obtain the left ventricular end-diastolic volume (LVEDV) and left ventricular end-systolic volume (LVESV); Locate the time point T0 when ventricular relaxation begins, the time point T1 when the ventricular volume reaches its maximum and the ventricle begins to contract, and the time point T2 when the ventricular volume reaches its minimum and the ventricle begins to relax again, and crop the video with [T0, T1] as a diastolic cycle and [T1, T2] as a systolic cycle to obtain a cardiac ultrasound animation; Image segmentation technology was used to segment the left ventricle (LV), right ventricle (RV), left atrium (LA), and right atrium (RA) from cardiac ultrasound images, and the images were normalized to a uniform size using cubic interpolation. defining nodes of a graph G, wherein the nodes include a left ventricle LV, a right ventricle RV, a left atrium LA, and a right atrium RA; Extract the spatiotemporal features of cardiac images using a 3D convolutional neural network with a C3D structure that integrates spatial and temporal feature attention mechanisms. Calculate similarity indices such as positional similarity, functional similarity, temporal similarity, mutual information similarity, and blood flow similarity, and calculate the attributes of edges in the graph; The graph G is subjected to feature aggregation and message passing through the graph convolutional network (GCN) model. The GCN model contains three edge-conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes. After three layers of ECC feature aggregation and message passing, a feature descriptor that integrates all cavity information is generated. And the generated feature descriptor The fully connected layer is input to predict the ejection fraction. The mean square error (MSE) is used as the loss function, and the Adam optimization algorithm is used to optimize the model, and finally the cardiac ejection fraction is output.

3. The intelligent echocardiographic cardiac function accurate assessment method according to claim 2, characterized in that: The method for extracting spatiotemporal features of cardiac images using a 3D convolutional neural network with a C3D structure that integrates spatial feature attention and temporal feature attention mechanisms is as follows: The spatial attention function f is calculated through operations such as spatial convolution layers and global average pooling. s (X), generates the weight A for each pixel position s , these weights A s Applied to the original feature map X, it implements element-by-element weighted adjustment, namely: X s′ =A s ·X s , Among them, the original feature map X s is a spatiotemporal feature map extracted from echocardiography video; A s is the spatial attention weight; f s (X s ) is the output feature of the spatial convolution layer, which is the input feature map X s Features obtained after spatial convolution operation; d s is the feature dimension of the spatial dimension; the weighted adjusted feature map X s′ , is to take the spatial attention weight A s Applied to the original feature map X s The results obtained after The temporal attention function f is calculated through the temporal convolution layer and global average pooling t (X), generate the weights for each time step, and add these weights A t Application to echocardiographic sequence X t , perform weighted adjustment on the features of each time step, namely: Among them, echocardiographic sequence X t , is a sequence of images of the heart at different time points; A t is the temporal attention weight, which is used to adjust the weights of different time steps in the echocardiographic sequence to better capture the dynamic changes of cardiac motion; f t (X t ) is the output feature of the temporal convolution layer, which is the input echocardiogram sequence X t Features obtained after temporal convolution operation; d t is the characteristic dimension of the time dimension; the weighted adjusted echocardiographic sequence X t′ , is to convert the time attention weight A t Applied to the original echocardiographic sequence X t The result obtained after.

4. The intelligent echocardiographic cardiac function accurate assessment method according to claim 2, characterized in that: The method for calculating the similarity indices of position similarity, functional similarity, temporal similarity, mutual information similarity, and blood flow similarity is: Position similarity E 1xm : If the two chambers are adjacent or have direct blood flow connection, the position similarity is 1, otherwise it is 0; Functional similarity E 2xm : The similarity between the depth feature vectors of two cavities is measured by the Euclidean distance, and the calculation formula is: Among them, Z x and Z m are the feature vectors of nodes x and m respectively; Temporal similarity E 3xm : The similarity between the main frequencies of the two chambers is calculated by time-frequency analysis. The calculation formula is: Among them, ST x ′ and ST m ′ are the frequency features of the spatiotemporal feature maps of nodes x and m after fast Fourier transform; Mutual information similarity E 4xm : Mutual information is an indicator that measures the degree of information sharing between two random variables. Its calculation formula is: Where X and Y are the eigenvectors of the two chambers, p(x,y) is the joint probability distribution, and p(x) and p(y) are the marginal probability distributions; Blood flow similarity E 5xm : Use the optical flow method to extract blood flow velocity and direction data at each time point. Calculate the optical flow field by minimizing the brightness change between adjacent frames. For each pair of adjacent frames, calculate the optical flow field. The specific formula is as follows: v(x,y,t)=(u(x,y,t),v(x,y,t)) Among them, u(x,y,t) and v(x,y,t) represent the horizontal and vertical velocity components at position (x,y) and time t respectively; the speed and direction information of each pixel point is extracted from the optical flow field. The magnitude of the speed is calculated by the modulus of the vector, and the direction is calculated by the angle of the vector: The extracted blood flow velocity and direction information is converted into feature vectors.

5. The intelligent echocardiographic cardiac function accurate assessment method according to claim 4, characterized in that: The method for calculating the attributes of the edges in the graph is: The final edge attribute E xm is the above position similarity E 1xm , functional similarity E 2xm , time similarity E 3xm , mutual information similarity E 4xm and blood flow similarity E 5xm The synthesis of: AND xm =And 1xm +E 2xm +E 3xm +E 4xm +E 5xm 。 6. The intelligent echocardiographic cardiac function accurate assessment method according to claim 2, characterized in that: The graph G is subjected to feature aggregation and message passing through the graph convolutional network (GCN) model. The GCN model includes three edge-conditional convolution (ECC) layers. Each ECC layer updates the feature vector of a node by weighted aggregation of feature information of adjacent nodes. The input layer receives the node feature matrix Where d represents the dimension of each node feature. In the first ECC layer, the input feature dimension d is mapped to d1, and the filter generates the network F l Using edge attributes as input, a weight matrix is ​​generated to weight the features of adjacent nodes, thereby achieving initial feature aggregation and abstraction. The second ECC layer maps the feature dimension d1 to d2, further deepening the abstraction and fusion of features. The third ECC layer maps the feature dimension d2 to d3, generating a more advanced and comprehensive feature representation. Each node updates its own feature vector by aggregating the feature vectors of its adjacent nodes and edges. The formula for the ECC operation is: Among them, F l It is a filter generation network with edge attributes E xm As input, a weight matrix is ​​generated to weight the features of adjacent nodes. l is the learnable weight, b l Is the bias term. The filter generates the network F l It is usually a small neural network that can dynamically generate a weight matrix based on edge attributes. In each layer, the weight matrix w l and the bias term b l Updates are made through backpropagation and optimization algorithms to minimize the loss function.

7. The intelligent echocardiographic cardiac function accurate assessment method according to claim 6, characterized in that: After three layers of ECC feature aggregation and message passing, a feature descriptor that integrates all cavity information is generated. The method is: The first ECC layer maps the input feature dimension d to d1 and performs preliminary feature aggregation; The second ECC layer maps the feature dimension d1 to d2 to further deepen the abstraction and fusion of features; The third ECC layer maps the feature dimension d2 to d3 to generate the final feature descriptor 8. The intelligent echocardiographic cardiac function accurate assessment method according to claim 7, characterized in that: The generated feature descriptor Input the fully connected layer to predict the ejection fraction, use the mean square error (MSE) as the loss function, and use the Adam optimization algorithm to optimize the model. The final method to output the cardiac ejection fraction is: The mean square error (MSE) is used as the loss function to measure the difference between the predicted value and the true value. The mathematical expression of the loss function is: Where N is the number of training samples, is the predicted ejection fraction of the ith sample, is the true ejection fraction of the i-th sample, optimized by the Adam algorithm with a learning rate of 10 -4 , and adopts a piecewise constant decay strategy to minimize the loss function, realize model training and optimization, and finally obtain accurate ejection fraction prediction results.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for accurately evaluating cardiac function by intelligent echocardiography as claimed in any one of claims 2 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the intelligent echocardiography cardiac function accurate assessment method as described in any one of claims 2 to 8 is implemented.