Method for estimating left ventricular motion based on shape attention cascade of cmr cine sequence

CN122367929APending Publication Date: 2026-07-10CAPITAL UNIVERSITY OF MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CAPITAL UNIVERSITY OF MEDICAL SCIENCES
Filing Date
2026-04-09
Publication Date
2026-07-10

Smart Images

  • Figure CN122367929A_ABST
    Figure CN122367929A_ABST
Patent Text Reader

Abstract

This application proposes a left ventricular motion estimation method based on shape attention cascades for CMR cine sequences. This method consists of a base module and a sequence module. The base module uses a shape flow layer to extract boundary features, improving estimation accuracy at the left ventricular boundary. The sequence module uses a motion attention layer to ensure temporal coherence of motion estimation through bidirectional encoding of motion features. Furthermore, this application proposes novel data augmentation methods and loss functions, enhancing the accuracy of left ventricular motion estimation through comprehensive improvements to the training method and model structure. Experimental results show that the proposed method outperforms existing methods in both geometric and clinical metrics. This application improves the accuracy of left ventricular motion estimation, contributing to research on motion estimation and the clinical application of automated cardiac function assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of ventricular motion estimation technology, and in particular to a method and system for left ventricular motion estimation based on shape attention cascade CMRcine sequences. Background Technology

[0002] Cardiac magnetic resonance cine (CMR cine) is a non-invasive imaging technique used to acquire information about cardiac motion. Left ventricular motion estimation in CMR cine sequences is a crucial step in the quantitative assessment of myocardial wall motion, and its accuracy significantly impacts the diagnosis and treatment of myocardial diseases. Generally, motion estimation outputs a displacement field, which reflects the coordinate changes of pixels between consecutive frames. In early studies, tagged magnetic resonance imaging (tagged MRI) was used to track the displacement field of the myocardium. However, due to the lengthy post-processing, tagged MRI sequences are no longer used clinically. Currently, CMR cine is the primary sequence used clinically for left ventricular motion estimation. However, because CMR cine image signals are relatively uniform and lack identifiable feature points, the left ventricular and epiventricular contours are required as markers. Left ventricular motion estimation in CMR cine is achieved by tracking changes in these contours.

[0003] Early studies often employed a two-step segmentation-registration approach for left ventricular motion estimation: first, segmentation methods were used to obtain the left ventricular and epiventricular contours on the frame sequence, with common models including ResUnet, dilated ResUnet, and U-Net with attention mechanisms; then, registration algorithms were applied to achieve motion estimation, with typical methods including B-spline free deformation registration and graph convolutional networks based on contour sampling point registration. Since the feature spaces of left ventricular segmentation and motion tracking are highly correlated, some studies have used a Siamese network architecture to simultaneously learn both segmentation and registration tasks, synergistically improving the accuracy of both.

[0004] The input to a Siamese network architecture is a pair of images, typically with the end-diastolic frame as the reference frame, which is then paired with subsequent frames in the sequence to form the input image pair. Subsequent studies have also employed similar Siamese network architectures, directly learning pixel displacement fields from the image pairs. Considering that cardiac motion dynamics is a complex rhythmic pattern exhibiting a nonlinear trajectory regulated by molecular, electrophysiological, and biophysical processes, to achieve biologically reasonable motion estimation, some studies have proposed regularization terms based on strain tensors, others have used the deformation space reconstructed from finite element models as a priori assumptions about the motion trajectory, and still others have utilized bidirectional recurrent neural networks to obtain the Lagrangian motion field between images. Some studies have improved boundary estimation accuracy by introducing a shape loss function, but this method relies on manually labeled contours across the entire sequence as supervision information. However, in real-world data scenarios, only manually labeled end-diastolic or end-systolic frames are typically available, leading to poor training performance for such fully supervised models. Summary of the Invention

[0005] In view of this, the purpose of this application is to propose a method and system for left ventricular motion estimation based on shape attention cascade CMR cine sequence, which can specifically solve the existing problems.

[0006] To achieve the above objectives, this application proposes a left ventricular motion estimation method based on shape attention cascade CMR cine sequences, comprising: The initial segmentation results of the CMR cine sequence used for training are obtained through the trained motion estimation model and segmentation model. The boundaries of the initial segmentation results are subjected to noise enhancement data processing, and the basic module is trained using the data-processed boundaries to obtain an error map corresponding to the prediction results of the segmentation model and the motion estimation model. Pixel confidence is calculated based on the error map, frame confidence is calculated based on the consistency of the prediction results of the basic module, segmentation model, and motion estimation model, and a loss function is obtained by combining the pixel confidence and frame confidence. The sequence module is then trained using the loss function. After inputting the CMR cine sequence to be inferred into the motion estimation model and the segmentation model to obtain the initial segmentation result of the CMR cine sequence to be inferred, the initial segmentation result is input into the trained sequence module, and the trained sequence module outputs the optimized segmentation result.

[0007] To achieve the above objectives, this application also proposes a left ventricular motion estimation system based on shape attention cascade CMR cine sequences, comprising: The segmentation unit obtains the initial segmentation results of the CMR cine sequence used for training through the trained motion estimation model and segmentation model; The training base module unit performs noise-enhanced data processing on the boundary of the initial segmentation result, and uses the data-processed boundary to train the base module to obtain an error map corresponding to the prediction results of the segmentation model and the motion estimation model. The training sequence module unit calculates pixel confidence based on the error map, calculates frame confidence based on the consistency of the prediction results of the base module, segmentation model, and motion estimation model, and obtains a loss function by combining the pixel confidence and frame confidence. The loss function is then used to train the sequence module. The inference unit inputs the CMR cine sequence to be inferred into the motion estimation model and the segmentation model to obtain the initial segmentation result of the CMR cine sequence to be inferred. Then, the initial segmentation result is input into the trained sequence module, and the trained sequence module outputs the optimized segmentation result.

[0008] In summary, the advantages of this application and the user experience it brings are as follows: the method proposed in this application outperforms existing methods in both geometric and clinical metrics. This application improves the accuracy of left ventricular motion estimation, which helps to advance research on motion estimation problems and the application of automated cardiac function assessment in clinical practice. Attached Figure Description

[0009] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0010] Figure 1 A schematic diagram of the cascaded structure of this application is shown; Figure 2 This diagram illustrates the basic module structure according to an embodiment of the present application. Figure 3 A schematic diagram of a sequence model structure according to an embodiment of this application is shown; Figure 4 The figure shows an example of experimental results according to an embodiment of this application; Figure 5 A schematic diagram of radial and circumferential myocardial strain curves calculated on the left ventricular shape sequence predicted by the method of this application is shown. Detailed Implementation

[0011] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0012] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0013] Due to factors such as motion artifacts and improper imaging parameter settings, CMR cine sequences often exhibit varying degrees of blurring, making it difficult for pixel-registration-based motion estimation methods to achieve accurate displacement field estimation in blurred boundary regions. To address this technical bottleneck, this application proposes a cascaded structure to optimize the motion estimation results of existing methods, thereby correcting estimation errors in boundary regions. A schematic diagram of the structure is shown below. Figure 1 As shown, the system consists of a basic module and a sequence module. The workflow is as follows: During the training phase, initial segmentation results are obtained using the trained motion estimation and segmentation models; noise enhancement is applied to the boundary data to train the basic module to recognize the error map; pixel confidence is calculated from the error map, frame confidence is calculated from the consistency of multi-model prediction results, and the sequence module is trained using a biased loss function; During the inference phase, after obtaining the initial segmentation results through the motion estimation and segmentation models, the sequence module directly outputs the optimized segmentation results. The main contributions of this application are as follows: (1) The basic module extracts key boundary features through the shape flow layer and combines error map-supervised training to improve the model's estimation accuracy of image boundary regions.

[0014] (2) The sequence module uses the motion attention layer to perform bidirectional temporal encoding of motion features to achieve motion estimation with temporal coherence.

[0015] (3) A novel data augmentation method is proposed during the training phase, and a biased loss function is designed to significantly improve the model accuracy.

[0016] Experimental results show that the method proposed in this application can effectively improve the estimation accuracy of the left ventricular and epicardial contours, and its geometric and clinical indices are significantly improved compared with existing methods. p <0.05. Statistical analysis was performed using the Kruskal-Wallis test (KW test).

[0017] 1. Materials and Methods 1.1 Experimental Data The CMR cine data consisted of the Automatic Cardiac Diagnosis Challenge (ACDC) dataset and a private database. The ACDC dataset contained CMR cine data from 150 patients, with 100 used for training and 50 for testing. Each dataset included manually segmented and labeled right ventricle, myocardium, and left ventricle images at both end-diastole and end-systole. The private database comprised CMR cine data from 186 patients collected from hospitals, with 143 used for training and 43 for testing. Left ventricular and epicardial endocardium were manually segmented and labeled by two radiologists with over 10 years of experience.

[0018] The ultrasound data is from the EchoNet-Dynamic database, containing 10,271 apical four-chamber echocardiogram videos, which is the largest echocardiogram dataset to date. Of these, 7,465 samples were used for training, 1,277 for validation, and 1,288 for testing. The end-diastolic and end-systolic images for each data point were provided with segmentation annotations by clinical experts.

[0019] 1.2 Basic Modules The basic module proposed in this application consists of a shared encoder, two decoders, and a shape flow layer, and its structure is as follows: Figure 2 As shown in the diagram. The input to this module includes an image, as well as the prediction results of the segmentation model and the motion estimation model; the output is a fine prediction result of the left ventricular contour, and error maps corresponding to the prediction results of the segmentation model and the motion estimation model, respectively.

[0020] The shared encoder consists of dense blocks of DenseNet-121, the decoder uses dual-attention decoding blocks, and the shape flow layer is implemented using gated shape flow. The gated shape flow model consists of 1×1 convolutional functions, residual block functions, and gated convolutional layers. This represents a 1×1 convolution function. Represents the residual block function. Indicates the first If the gated convolutional layer of the first layer outputs the following, then the output of the next gated convolutional layer will be: (1) in For the feature layer, t For the index of the encoding module, Here, || represents the concatenation of feature layers along the channel dimension. Concatenating the output of the gated shape stream with the Canny edges of the input image produces a shape feature map, which improves the segmentation accuracy of the base module at boundaries.

[0021] Since pixel-wise error correction can effectively improve segmentation quality, this application designs the base model as a multi-task model that can simultaneously output prediction results and error maps. Specifically, the base model uses a shared decoder to output two error maps for pixel-wise error correction of the initial prediction. To represent segmentation labels, use , The error map (obtained by comparing the prediction results with the segmentation labels) is represented using... (Original image) (Prediction by the segmentation model) (The prediction of the motion estimation model) represents the input to the basic module, using (Prediction Results) (error Figure 1 ), (error Figure 2 The output of the basic module is represented by ), and the loss function for the error plot branch is: (2) in For focus loss function, , For pixel-by-pixel error correction of the prediction results, Mean square error, Indicates the weighting coefficient. Use Let represent the cross-entropy loss function. The loss function for the basic module is: (3) This represents the error weight coefficient. Furthermore, this application proposes a novel data augmentation method that effectively improves the robustness of the model by adding noise to the segmentation boundaries. The method comprises four steps: first, performing random dilation or erosion operations on the boundaries using a kernel function of random size; second, adding random pixel perturbations at random locations on the boundaries; third, adding random pixel perturbations at random locations; and finally, adding Gaussian blur using a kernel function of random size.

[0022] 1.3 Sequence Module To ensure the temporal coherence of motion estimation, this application proposes a sequence model. This model takes a CMR cine sequence as input and applies a motion attention layer to encode and reweight the feature maps in both directions. The structure of the sequence model is as follows: Figure 3 As shown, this model can output detailed prediction results for CMR cine sequences.

[0023] The motion attention layer adds an attention mechanism to the bidirectional convolutional long short-term memory (ConvLSTM) layer. For example... Figure 3 As shown, the motion attention layer consists of a bidirectional ConvLSTM, a global average pooling layer, and a bidirectional long short-term memory (LSTM) layer. The bidirectional ConvLSTM layer contains one forward unit and one backward unit, used to encode motion features from two temporal directions. The global average pooling layer reduces the motion features to a single value by averaging all values, generating a one-dimensional attention vector. The bidirectional LSTM layer also contains one forward unit and one backward unit, and its operation is similar to the bidirectional ConvLSTM layer, but it processes a one-dimensional attention vector. It achieves adaptive weighting of the motion features by element-wise multiplying the attention vector with the motion features. When training the sequence model, this application proposes a novel weighted Dice loss function, which evaluates the uncertainty of annotation from two aspects: pixel confidence and frame confidence. Pixel confidence is derived from the error map of the base module; that is, pixels with high error probabilities are assigned smaller weight values, expressed as: (4) in This is the error map of the basic module's predictions. Frame credibility evaluates the consistency between different models' predictions of the same frame, using... This represents the predictions from the basic module, the segmentation model, and the motion estimation model. A high degree of consistency among these predictions indicates high reliability, as expressed below: (5) in and The HD function distance between the two results is used to measure the consistency between the two results at the boundary. and The HD function distance between the two results is used to measure the consistency between the two results at the boundary.

[0024] Based on the above confidence calculation, the weighted Dice loss function is expressed as follows (A and B represent two variables): (6) Because the Dice loss function introduces bias during model training, this application uses an unbiased loss function for manual annotations and a biased loss function for the output of the base model to correct this bias. In summary, the loss function for the sequence model is expressed as: (7) in It is a prediction from a sequence model. It was manually labeled. It is the prediction of the basic model. Represents the cross-entropy loss function. This represents the weighting coefficient.

[0025] 1.4 Model Training The model in this application is implemented based on the PyTorch deep learning framework. The training process uses the Adam optimizer, with an initial learning rate of 0.001 and a Step Decay learning rate decay strategy. In the loss function... Weight parameters All are set to 0.1. Set to 0.5. With the value set to 0.5, model training was completed on four Tesla V100S GPUs.

[0026] 1.5 Evaluation Indicators This application uses geometric and clinical metrics to evaluate model performance. Geometric metrics include the Dice Similarity Coefficient (DICE), Jaccard Distance (JD), Hausdorff Distance (HD), and Average Surface Distance (ASD); clinical metrics include radial strain, circumferential strain, and ejection fraction (EF). The Pearson Correlation Coefficient (PCC) is used to evaluate the consistency between the clinical metrics predicted by the model and the manually labeled results. 2 Results 2.1 Comparison with other methods This application uses a joint segmentation and motion estimation model, a registration neural network based on variational autoencoder regularization, a differential homeomorphic registration neural network, and a cardiac motion learning model incorporating biomechanical information modeling to achieve left ventricular motion estimation. These methods are labeled M2-M5, with M1 representing the segmentation module using only the model. Experimental comparison results on the ACDC database are shown in Table 1, and experimental comparison results on a private database are shown in Table 2. The method in this application achieves significant improvements in all indicators. p<0.05), specifically manifested as follows: the DICE coefficient and JD index significantly improved, both focusing on the accuracy of left ventricular global region estimation. This improvement demonstrates that the method effectively optimizes the regional accuracy of left ventricular motion estimation; the HD and ASD indices decreased. These two indices measure the distance between boundary points, and their reduction directly reflects a significant reduction in boundary region estimation error. Examples of experimental results are shown below. Figure 4 As shown, this intuitively verifies that the method in this application effectively improves the accuracy of left ventricular motion estimation.

[0027] Table 1. Comparison of geometric indicators in the ACDC database

[0028] Table 2 Comparison of Geometric Indices of Private Databases

[0029] Each frame of the private dataset is manually annotated, allowing the calculation of circumferential and radial strain from the geometric deformation of the manually annotated contours. Experimental comparison results are shown in Table 3. The clinical indicators obtained using the method described in this application show higher consistency with those obtained through manual annotation. In 43 test cases, circumferential and radial strain were significantly correlated with the manually annotated data. p The data volumes of <0.01 were 40 and 42, respectively, and the experiment verified the usability of the proposed method to replace manual annotation for automatic assessment of cardiac function.

[0030] Table 3 Comparison of clinical indicators in private databases

[0031] The ACDC database only provides manually annotated end-diastolic and end-systolic images, making it impossible to calculate circumferential and radial strain. However, the ACDC database provides pathological groups, including Right Ventricular Abnormality (RV), Dilated Cardiomyopathy (DCM), Hypertrophic Cardiomyopathy (HCM), Myocardial Infarction (MI), and a Normal Control Group (NOR). This application directly uses model-predicted left ventricular shape sequences to calculate circumferential and radial strain curves for these five pathological groups. Figure 5As shown, compared with the NOR group, the peak radial and circumferential strains of the DCM and MI groups were reduced, and the peak time of the strain curve of the DCM group was also delayed. In addition, the peak circumferential strain of the HCM group was also lower than that of the NOR group. These results verify the clinical reference value of the method proposed in this application.

[0032] The experimental results on the EchoNet-Dynamic database are shown in Table 4. The method in this application shows significant improvements in DICE, JD, and HD metrics. p <0.05). The EchoNet-Dynamic database provides EF annotations. The differences between the EF measurement results of each method and the annotations are shown in Table 5. The root mean square error (RMSE) of the method in this application is significantly reduced, and the PCC is significantly improved, indicating that the method in this application has higher measurement accuracy of EF and stronger consistency with clinical reference values. Table 4. Comparison of geometric metrics for the EchoNet-Dynamic database.

[0033] Table 5 Comparison of clinical indicators in the EchoNet-Dynamic database

[0034] 2.2 Ablation Experiment Results This application proposes a novel data augmentation method, an error map, and a weighted loss function. This section verifies its effectiveness through ablation experiments. As shown in Table 6, the ASD metric is significantly improved after applying data augmentation; as shown in Table 7, the ASD metric is significantly improved after introducing the error map; as shown in Table 8, the weighted loss function significantly improves the DICE and JD metrics by calculating pixel confidence and frame confidence, and quantifying the uncertainty of labels during model training can effectively improve model accuracy.

[0035] Table 6 Data Enhancement Ablation Experiment

[0036] Table 7 Error Chart Ablation Experiment

[0037] Table 8 ablation experiment

[0038] 3. Discussion While pixel-based motion estimation can yield valuable results, it struggles to achieve accurate estimation in areas with blurred boundaries, resulting in inaccurate boundary prediction. To address this issue, this application first proposes a basic module to improve the accuracy of left ventricular boundary region estimation. Subsequently, to ensure the temporal coherence of motion estimation, a sequence model is further proposed. The sequence model takes a sequence of left ventricular images as input and generates temporally coherent motion features through a motion attention layer. During training, the sequence model uses the high-quality output of the basic model as pseudo-labels and employs a biased loss function that fuses pixel and frame credibility, significantly improving the geometric and clinical metrics of motion estimation.

[0039] The method proposed in this application is not only applicable to motion estimation problems in CMR cine sequences, but can also be extended to cardiac ultrasound sequences, and experimental results on the EchoNet-Dynamic database have verified its effectiveness. Beyond cardiac imaging, thanks to its high-precision estimation capability for regions with blurred boundaries and its mechanism for ensuring sequence coherence, the framework of this application can also be extended to other medical imaging scenarios requiring high-precision motion estimation. For example, lung 4D-CT suffers from blurred target area boundaries due to respiratory motion, and dynamic contrast-enhanced MRI of the liver requires the elimination of motion artifacts in multi-phase scanning and accurate capture of lesion boundaries; both scenarios rely on high-precision dynamic motion estimation technology to support clinical applications. Therefore, this method has broad clinical application prospects in the field of medical imaging.

[0040] Currently, the method proposed in this application is limited to motion estimation from two-dimensional cardiac images. However, comprehensive cardiac function assessment relies on accurate three-dimensional shape estimation of the left ventricle. Therefore, this application plans to extend the shape attention cascade mechanism to three-dimensional motion estimation of the left ventricle; simultaneously, it plans to further increase the size and quality of the dataset and conduct more comprehensive experimental comparisons based on different types of heart disease, thereby more comprehensively evaluating the usability of the proposed method in cardiac function assessment.

[0041] 4. Conclusion Accurate left ventricular motion estimation is crucial for clinical diagnosis and treatment. To improve its accuracy, this application proposes a shape attention cascade method, which includes a basic module and a sequence module. The basic module captures key boundary features through a shape flow layer, effectively improving the accuracy of left ventricular boundary estimation. The sequence module uses a motion attention layer to encode motion features bidirectionally, ensuring the temporal coherence of the estimation results. To further optimize model performance, this application also designs novel data augmentation methods and loss functions. Experimental results demonstrate that the proposed method outperforms existing methods in both geometric and clinical metrics. Current research is mainly limited to two-dimensional shape estimation. The next step is to extend the proposed method to three-dimensional left ventricular shape estimation. Furthermore, validation experiments will be conducted on multi-center, multi-disease cardiovascular datasets to more comprehensively evaluate the effectiveness and generalization ability of this method in cardiac function assessment.

[0042] It should be noted that: The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, this application is not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of this application.

[0043] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0044] Similarly, it should be understood that, in order to simplify this application and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this application, various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.

[0045] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0046] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0047] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation system according to the embodiments of this application. This application can also be implemented as a device or system program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0048] It should be noted that the above embodiments are illustrative of this application and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several systems, several of these systems may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0049] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for estimating left ventricular motion in CMR cine sequences based on shape attention cascade, characterized in that, include: The initial segmentation results of the CMR cine sequence used for training are obtained through the trained segmentation model and motion estimation model. The boundaries of the initial segmentation results are subjected to noise enhancement data processing, and the basic module is trained using the data-processed boundaries to obtain an error map corresponding to the prediction results of the segmentation model and the motion estimation model. Pixel confidence is calculated based on the error map, frame confidence is calculated based on the consistency of the prediction results of the basic module, segmentation model, and motion estimation model, and a loss function is obtained by combining the pixel confidence and frame confidence. The sequence module is then trained using the loss function. After inputting the CMR cine sequence to be inferred into the motion estimation model and the segmentation model to obtain the initial segmentation result of the CMR cine sequence to be inferred, the initial segmentation result is input into the trained sequence module, and the trained sequence module outputs the optimized segmentation result.

2. The method according to claim 1, characterized in that, The basic module consists of a shared encoder, two decoders and a shape flow layer. The input of the basic module includes the image and the prediction results of the segmentation model and the motion estimation model. The output is a detailed prediction of the left ventricular contour, and error plots corresponding to the prediction results of the segmentation model and the motion estimation model, respectively.

3. The method according to claim 2, characterized in that, The shared encoder consists of dense blocks of DenseNet-121, the decoder uses dual-attention decoding blocks, and the shape flow layer is implemented using a gated shape flow model.

4. The method according to claim 3, characterized in that, The gated shapeflow model consists of a 1×1 convolutional function, a residual block function, and a gated convolutional layer; using This represents a 1×1 convolution function. Represents the residual block function. Indicates the first If the gated convolutional layer of the first layer outputs the following, then the output of the next gated convolutional layer will be: (1) in For the feature layer, t For the index of the encoding module, The sigmoid function is used, and || represents the concatenation of the feature layers along the channel dimension; the output of the gated shape flow model is concatenated with the Canny edges of the input image to output a shape feature map.

5. The method according to claim 3, characterized in that, The shared decoder outputs two error maps, which are used to perform pixel-by-pixel error correction on the initial prediction results.

6. The method according to claim 1, characterized in that, The noise enhancement data processing includes: Use a kernel function of random size to perform random expansion or erosion operations on the boundary; Add random pixel perturbations at random locations on the boundary; Add random pixel perturbations at random locations; Add Gaussian blur using a kernel function of random size.

7. The method according to claim 4, characterized in that, The sequence model takes the CMR cine sequence as input and applies a motion attention layer to encode and reweight the shape feature map in both directions.

8. The method according to claim 7, characterized in that, The motion attention layer consists of a bidirectional ConvLSTM layer, a global average pooling layer, and a bidirectional LSTM layer; the bidirectional ConvLSTM layer contains a forward unit and a backward unit, used to encode motion features from two temporal directions; The global average pooling layer reduces the motion features to a single value by averaging all values ​​and generates a one-dimensional attention vector. The bidirectional LSTM layer contains a forward unit and a backward unit to process the one-dimensional attention vector. It achieves adaptive weighting of the motion features by multiplying the attention vector element-wise with the motion features.

9. The method according to claim 1, characterized in that, When training the sequence model, the weighted Dice loss function is used, which evaluates the uncertainty of the annotation from two aspects: pixel confidence and frame confidence. The pixel reliability is derived from the error map of the base module and is expressed as: (4) in This is the error map predicted by the basic module; The frame reliability evaluation assesses the consistency between different prediction results from different models for the same frame, using... The predictions of the basic module, the segmentation model, and the motion estimation model are represented as follows: (5) in and The HD function distance between the two results is used to measure the consistency between them at the boundary. and The HD function distance between the two results is used to measure the consistency between them at the boundary. The weighted Dice loss function is then expressed as: (6) A and B represent two variables; The loss function for the sequence model is expressed as: (7) in It is a prediction from a sequence model. It was manually labeled. It is the prediction of the basic model. Represents the cross-entropy loss function. This represents the weighting coefficient.

10. A left ventricular motion estimation system based on shape attention cascade CMR cine sequence, characterized in that, include: The segmentation unit obtains the initial segmentation results of the CMR cine sequence used for training through the trained motion estimation model and segmentation model; The training base module unit performs noise-enhanced data processing on the boundary of the initial segmentation result, and uses the data-processed boundary to train the base module to obtain an error map corresponding to the prediction results of the segmentation model and the motion estimation model. The training sequence module unit calculates pixel confidence based on the error map, calculates frame confidence based on the consistency of the prediction results of the base module, segmentation model, and motion estimation model, and obtains a loss function by combining the pixel confidence and frame confidence. The loss function is then used to train the sequence module. The inference unit inputs the CMR cine sequence to be inferred into the motion estimation model and the segmentation model to obtain the initial segmentation result of the CMR cine sequence to be inferred. Then, it inputs the initial segmentation result into the trained sequence module, and the trained sequence module outputs the optimized segmentation result.