Cardiac ultrasound key point tracking implementation method and device based on pseudo tag, and medium
Through the pseudo-label-based cardiac ultrasound key point tracking method, the key frame and label propagation model are used to optimize the training, which solves the problem of low efficiency in the cardiac ultrasound key point tracking process and realizes efficient and intelligent key point detection.
Patent Information
- Application Number
- CN202510742909.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-30
AI Technical Summary
During the cardiac ultrasound key point tracking process, manual labeling is tedious and time-consuming, and highly dependent on the skills of professional doctors, resulting in low efficiency and insufficient intelligence.
A pseudo-label-based cardiac ultrasound key point tracking method is adopted. By extracting key frames from ultrasound images of the cardiac beat cycle, high-quality pseudo labels are generated. The label propagation model is used to optimize training to improve the efficiency and intelligence of key point detection.
It significantly improves the efficiency and intelligence of cardiac ultrasound key point tracking, fully taps the potential value of unlabeled data, provides solid technical support, and lays the foundation for subsequent cardiac data analysis.
Smart Images

Figure CN120725968A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasonic image analysis, and in particular to a method, device, and medium for implementing cardiac ultrasonic key point tracking based on pseudo labels. Background Art
[0002] Due to the beating process of the heart, the ultrasound images corresponding to the cardiac cycle are constantly changing. The purpose of cardiac ultrasound key point tracking is to track the changes of key points as the heart fluctuates during the heart beat, so as to accurately locate and track the anatomical key points of the heart and support the quantitative analysis of cardiac structure and motion function. Therefore, cardiac ultrasound key point tracking technology is of great significance in evaluating cardiac function and monitoring treatment effects.
[0003] However, a complete cardiac cycle typically consists of approximately 60 frames. Relying on manual annotation for cardiac ultrasound keypoint tracking requires precise annotation of multiple keypoints per frame, a tedious and time-consuming process. Furthermore, cardiac ultrasound annotation relies heavily on specialized physicians with extensive medical knowledge and practical experience. Therefore, improving the efficiency and intelligence of cardiac ultrasound keypoint tracking has become a pressing issue. Summary of the Invention
[0004] The present invention provides a method, device and medium for implementing cardiac ultrasound key point tracking based on pseudo labels, which are used to improve the efficiency and intelligence level of cardiac ultrasound key point tracking.
[0005] In order to solve the above technical problems, the first aspect of the present invention discloses a method for tracking cardiac ultrasound key points based on pseudo labels, the method comprising:
[0006] Extracting a key frame from a set of ultrasound images corresponding to a cardiac cycle, wherein the key frame includes at least one change node in the cardiac cycle, and ultrasound images other than the key frame in the set of ultrasound images are intermediate frames;
[0007] For each key frame, extract the key points on the key frame as key pseudo labels;
[0008] According to the key pseudo-label, combined with the time series relationship between the ultrasound images in the ultrasound image set, key point detection is performed on each of the intermediate frames based on a preset label propagation model, and the key points corresponding to the key pseudo-label on each of the intermediate frames are extracted as intermediate pseudo-labels;
[0009] The label propagation model is optimized and trained according to the intermediate pseudo-label to obtain an optimized label propagation model; wherein the optimized label propagation model is used to track key points of the ultrasound video corresponding to the cardiac cycle.
[0010] A second aspect of the present invention discloses a device for tracking cardiac ultrasound key points based on pseudo labels, the device comprising:
[0011] a key frame extraction module, configured to extract key frames from a set of ultrasound images corresponding to a cardiac cycle, wherein the key frames include at least one change node in the cardiac cycle, and ultrasound images other than the key frames in the set of ultrasound images are intermediate frames;
[0012] A key point detection module is used to extract the key points on each key frame as key pseudo labels;
[0013] a label propagation module, configured to perform key point detection on each intermediate frame based on the key pseudo-label and a time series relationship between the ultrasound images in the ultrasound image set based on a preset label propagation model, and extract key points corresponding to the key pseudo-label on each intermediate frame as intermediate pseudo-labels;
[0014] A model optimization module is used to optimize and train the label propagation model based on the intermediate pseudo-label to obtain an optimized label propagation model; wherein the optimized label propagation model is used to track key points of the ultrasound video corresponding to the cardiac beat cycle.
[0015] A third aspect of the present invention discloses a system for tracking cardiac ultrasound key points based on pseudo labels, the system comprising:
[0016] a memory storing executable program code;
[0017] a processor coupled to the memory;
[0018] The processor calls the executable program code stored in the memory to execute the method for implementing cardiac ultrasound key point tracking based on pseudo labels disclosed in the first aspect of the present invention.
[0019] A fourth aspect of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the method for implementing cardiac ultrasound key point tracking based on pseudo labels disclosed in the first aspect of the present invention.
[0020] Compared with the existing technology, the present invention designs a cardiac ultrasound key point tracking process. First, key frames are extracted from cardiac cycle ultrasound images to ensure that the focus of data processing is on the key frame area; then, a key point detection algorithm is used to generate high-quality pseudo labels online; then, a label propagation model is used to combine high-quality pseudo labels to generate pseudo labels for intermediate frames, fully exploring the potential value of unlabeled data; finally, the generated intermediate frame pseudo labels are used to further optimize the training of the pseudo label propagation model, thereby improving the efficiency and intelligence of cardiac ultrasound key point tracking and providing solid technical support for subsequent cardiac data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 1 is a flow chart of a method for implementing cardiac ultrasound key point tracking based on pseudo labels disclosed in an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of the convex hull search and matching strategy disclosed in an embodiment of the present invention;
[0024] Figure 3 is a flowchart of cardiac ultrasound key point tracking disclosed in an embodiment of the present invention;
[0025] Figure 4 is a flow chart of the two-stage training strategy disclosed in an embodiment of the present invention;
[0026] Figure 5 Schematic diagram of a key tracking network architecture based on label propagation disclosed in an embodiment of the present invention;
[0027] Figure 6 1 is a schematic structural diagram of a pseudo-label-based cardiac ultrasound key point tracking implementation system disclosed in an embodiment of the present invention;
[0028] Figure 7 This is a schematic diagram of a heat map of key points of the long-axis inflow tract of the right atrium and right ventricle disclosed in an embodiment of the present invention;
[0029] Figure 8 Schematic diagram of key point distribution differences under the same view disclosed in an embodiment of the present invention;
[0030] Figure 9 This is a schematic diagram of a heat map in which similar key points disclosed in an embodiment of the present invention share a common heat map;
[0031] Figure 10Schematic diagram of a heat map regression network architecture based on a graph convolutional neural network disclosed in an embodiment of the present invention;
[0032] Figure 11 This is a schematic diagram of the lack of shape constraints of predicted image key points disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0034] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different items, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or end comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or end.
[0035] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0036] Example 1
[0037] See also Figure 1 , Figure 1 : is a flow chart of a method for implementing cardiac ultrasound key point tracking based on pseudo labels disclosed in an embodiment of the present invention. Figure 1 The described method for implementing cardiac ultrasound key point tracking based on pseudo labels can be implemented in a device for implementing cardiac ultrasound key point tracking based on pseudo labels, and the device for implementing cardiac ultrasound key point tracking based on pseudo labels can be integrated into a cloud server or a local server. Figure 1 As shown, the method for implementing cardiac ultrasound key point tracking based on pseudo labels may include the following operations:
[0038] Step 101: extract key frames from a set of ultrasound images corresponding to a cardiac cycle.
[0039] In embodiments of the present invention, a keyframe may include at least one change node in the cardiac cycle, and ultrasound images in the ultrasound image set other than the keyframes are intermediate frames. The heart is a core organ of the human body, located within the chest cavity, conical in shape, and the center of the circulatory system. The heart's working process is based on the cardiac cycle, completing a complete contraction and relaxation cycle, which is divided into the following phases: atrial systole, in which the atrial muscles contract, pushing blood into the ventricles; ventricular systole, in which the ventricular walls contract forcefully, pushing blood into the large arteries (aorta and pulmonary artery), during which the atria relax and accumulate blood for the next cycle; and ventricular diastole, in which the ventricles relax, allowing venous blood to return to the atria, gradually filling the ventricles and preparing for the next contraction. The cardiac cycle can be divided into systole (the phase of blood ejection) and diastole (the phase of cardiac chamber filling), which are regulated by a precisely coordinated conduction system. Optionally, keyframes may be end-diastolic (ED) and end-systolic (ES) frames in the cardiac cycle.
[0040] In this embodiment of the present invention, ED and ES frames are identified. These key frames have significant physiological characteristics during the cardiac cycle. ED and ES frames represent the moments of maximum and minimum left ventricular volume, respectively, during the cardiac cycle. They are essential nodes in cardiac function assessment and are widely used to measure indicators such as ejection fraction, ventricular wall motion, and cardiac chamber volume. By focusing on these key frames, the computational overhead of processing invalid information can be significantly reduced, concentrating resources on frames with the highest diagnostic value.
[0041] Step 102: For each key frame, extract the key points on the key frame as key pseudo labels.
[0042] In embodiments of the present invention, a keypoint detection algorithm based on heatmap regression can optionally be used to generate initial keypoint heatmaps. For example, a network model can be used to identify the initial keypoint locations in the keyframe image and output them as a probability distribution. Correspondingly, these heatmaps are high-quality pseudo-labels generated online.
[0043] In the embodiment of the present invention, other key point detection algorithms may optionally be used to detect key points from key frames.
[0044] Step 103: Based on the key pseudo-labels and the time series relationship between the ultrasound images in the ultrasound image set, key point detection is performed on each intermediate frame based on a preset label propagation model, and the key points corresponding to the key pseudo-labels on each intermediate frame are extracted as intermediate pseudo-labels.
[0045] In this embodiment of the present invention, to fully utilize pseudo-labeled data, a label propagation learning model is employed to generate pseudo-labels for intermediate frames. By integrating pseudo-labeled data from keyframes, the model continuously optimizes the quality of pseudo-labels during the learning process. Specifically, this involves leveraging the high-quality labels from keyframes to guide keypoint detection in intermediate frames. Simultaneously, through iterative pseudo-label generation and optimization, the model's adaptability to unlabeled data is gradually enhanced.
[0046] In an optional embodiment, the generated pseudo-labels for the intermediate frames are systematically stored to form a pseudo-label library. This pseudo-label library covers unlabeled intermediate frames in the cardiac cycle, making up for the lack of data annotation. Leveraging this pseudo-label library, the training process of the label propagation model can be further optimized.
[0047] Step 104: Optimize and train the label propagation model based on the intermediate pseudo labels to obtain an optimized label propagation model.
[0048] In an embodiment of the present invention, an optimized label propagation model is used to track key points in an ultrasound video corresponding to a cardiac cycle.
[0049] In this embodiment of the present invention, by combining annotated data with high-quality pseudo-labels, the model can more accurately learn the geometric relationships and dynamic changes between key points. Finally, the label propagation model can optionally generate key point detection results based on a heatmap regression method. Heatmap regression accurately locates the position of each key point by generating a heatmap of the probability distribution.
[0050] Among them, the principle of the embodiment of the present invention is as follows: The embodiment of the present invention realizes the key point tracking of cardiac ultrasound based on the pseudo-labeling method rather than unsupervised regularization. The main reason is that the pseudo-labeling method can make full use of unlabeled data and quickly expand the amount of training data through online generation and label propagation, thereby significantly improving the performance of the model. In the field of cardiac ultrasound, which has high labeling costs and scarce data, the pseudo-labeling method has the advantages of being simple to implement and easy to expand, and is suitable for capturing the detailed features of anatomical structures. In contrast, although the unsupervised regularization method performs well in improving the generalization ability of the model, it has a high reliance on the design of the regularization strategy, has high computational complexity, and may converge slowly in specific tasks. Therefore, the pseudo-labeling method is more in line with the requirements of this study for high efficiency and application adaptability.
[0051] It can be seen that the embodiment of the present invention designs a cardiac ultrasound key point tracking process, which first extracts key frames from cardiac cycle ultrasound images to ensure that the focus of data processing is on the key frame area; then, a key point detection algorithm is used to generate high-quality pseudo labels online; then, a label propagation model is used to combine high-quality pseudo labels to generate pseudo labels for intermediate frames, fully exploring the potential value of unlabeled data; finally, the generated intermediate frame pseudo labels are used to further optimize the training of the pseudo label propagation model, thereby improving the efficiency and intelligence of cardiac ultrasound key point tracking and providing solid technical support for subsequent cardiac function analysis.
[0052] In an optional embodiment, for each key frame, extracting key points on the key frame as key pseudo labels may include:
[0053] For each key frame, a key point heat map representing key points on the key frame is extracted as a key pseudo label based on a preset heat map regression algorithm;
[0054] Furthermore, according to the key pseudo-label, combined with the time series relationship between the ultrasound images in the ultrasound image set, key point detection is performed on each intermediate frame based on a preset label propagation model, and key points corresponding to the key pseudo-label on each intermediate frame are extracted as intermediate pseudo-labels, which may include:
[0055] For key pseudo-labels, a densely connected convolutional network is used to fully fuse shallow and deep features to obtain heat map features corresponding to the key pseudo-labels. This dense connection mechanism fully integrates shallow and deep features, achieving efficient feature reuse and improving fine-grained representation capabilities. This structure not only enhances the expressive power of heat map features, but also improves the accuracy and robustness of inter-frame matching.
[0056] Based on a preset inter-frame similarity calculation mechanism and in combination with the temporal sequence relationship between ultrasound images in the ultrasound image set, a matching relationship between intermediate frames is constructed;
[0057] For each intermediate frame, based on the heat map features corresponding to the key pseudo-label and the matching relationship between the intermediate frame and other intermediate frames, the key points corresponding to the key pseudo-label on the intermediate frame are extracted as the intermediate pseudo-label.
[0058] In this optional embodiment, an efficient inter-frame similarity calculation mechanism is used to directly construct an inter-frame matching relationship, which greatly reduces computational redundancy and processing complexity. In addition, the inter-frame matching mechanism can fully exploit the continuity of the time series and adapt to the smoothness of the trajectory of key points in the cardiac cycle.
[0059] It can be seen that in order to improve the quality and reliability of pseudo-label generation, this optional embodiment designs an optimized pseudo-label generation strategy based on the continuity of time series and the regularity of cardiac cycles. In the specific implementation, preliminary training is performed on a small amount of labeled data to generate initial pseudo-labels for intermediate frames, and the pseudo-labels are iteratively optimized in combination with the smoothness characteristics of the time series. Ultimately, the improvement in pseudo-label quality not only expands the amount of available training data, but also significantly enhances the model's ability to learn unlabeled data, laying a solid technical foundation for efficient key point tracking.
[0060] In yet another optional embodiment, extracting key frames from a set of ultrasound images corresponding to a cardiac cycle may include:
[0061] Acquiring an ultrasound video corresponding to a cardiac cycle, identifying an initial window region in the ultrasound video whose brightness change satisfies a preset brightness change pattern, and modifying the initial window region based on morphological features corresponding to the cardiac cycle to obtain an ultrasound window corresponding to the ultrasound video;
[0062] According to the ultrasound window corresponding to the ultrasound video, an ultrasound image set corresponding to the ultrasound video is extracted.
[0063] In an embodiment of the present invention, ultrasound window extraction is a key step in ultrasound image analysis. Taking cardiac ultrasound as an example, it can improve the efficiency and accuracy of subsequent analysis by extracting the area containing the cardiac anatomical structure from the image or video and removing background noise. In this embodiment, the area containing the cardiac structure is identified and extracted by processing multiple frames in the ultrasound video. Specifically, first, by performing brightness area detection on the video frame, the part with significant brightness changes is identified, which usually corresponds to the heart area. Optionally, it can also include: using morphological operations to further optimize the extracted area, remove noise and fill holes. Finally, combining multiple frame information, a comprehensive mask is generated to ensure that a stable and accurate ultrasound window is extracted.
[0064] In yet another optional embodiment, for each key frame, extracting key points on the key frame as key pseudo labels may include:
[0065] For each key frame, the key frame is input into the preset heat map prediction model. The heat map prediction model extracts all key points corresponding to the key frame according to the preset geometric features corresponding to the category of the key frame; the boundary path of the geometric distribution of the key points is generated according to the preset geometric features, the key points are sorted according to the preset boundary order, each layer of the heat map is restricted to represent only one key point, a predicted heat map corresponding to each key point is generated, and the key point coordinates are extracted from the predicted heat map based on the preset key point extraction strategy.
[0066] In this optional embodiment, the heat map is a method of characterizing the spatial location of key points through Gaussian distribution.
[0067] In this optional embodiment, key points of the same type are sorted and matched by geometric features to ensure that each type of key points has a fixed sequential relationship. For example, for cardiac ultrasound, based on the characteristics of cardiac ultrasound medical key points, the geometric distribution characteristics of key points are very clear, and the key points are almost sequentially distributed on the anatomical structures of the endocardium and epicardium. The endocardium and epicardium have regular annular or linear distribution patterns in anatomy, and this characteristic provides a natural basis for the sorting and matching of key points. This solution directly limits each layer of heat map to represent only one key point in the heat map generation stage, avoiding the repeated appearance of similar points in the heat map, and thus does not need to rely on complex non-maximum suppression operations. In the post-processing stage, only the maximum coordinates need to be extracted from each heat map, which simplifies the key point selection process.
[0068] In this optional embodiment, further optionally, the above-mentioned extraction of key point coordinates from the predicted heat map based on the preset key point extraction strategy may include: generating a boundary path of the key point geometric distribution according to the preset geometric features corresponding to all key points, and sorting all key points according to the preset boundary order. Specifically, it may include:
[0069] Calculate the convex hull boundaries corresponding to all key points, and generate the boundary path of the geometric distribution of key points based on the convex hull boundaries and preset geometric features; perform convex hull search on the key points of this category according to the preset boundary order, optimize the position of each key point based on the convex hull search results, and sort the key points of this category.
[0070] This optional embodiment, taking cardiac ultrasound as an example, may specifically include:
[0071] First, the image is corrected through key points, and the anatomical structure is adjusted to a unified reference coordinate system to ensure that the anatomical relationship in the image is consistent with the actual structure. Secondly, on the corrected image, the candidate key points are screened and sorted based on the key point matching strategy of convex hull search: First, the candidate key points are extracted and their convex hull boundaries are calculated to generate a boundary path reflecting the geometric distribution; then, the key points are sorted according to the order of the convex hull boundaries. For annular or linear anatomical structures, they are arranged in counterclockwise, clockwise, or from base to apex, respectively. The specific matching strategy is as follows: Figure 2 As shown in Figure 3, the positions of key points are optimized using a convex hull search algorithm to ensure that the key point distribution conforms to the spatial constraints of the anatomical structure.
[0072] In this optional embodiment, the ultrasound image contains multiple views, and the anatomical structures and key point distribution patterns in different views are significantly different. Relying only on the original image features, the model may not fully understand the semantic differences between views, resulting in a decrease in the accuracy of key point detection. To solve this problem, this embodiment embeds view category information into the network, enabling the model to explicitly consider the view type during feature extraction. Therefore, in this embodiment of the present invention, further optionally, the heat map prediction model is obtained by the following steps:
[0073] According to the category of each ultrasound image corresponding to the cardiac cycle, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the cardiac cycle to obtain a heat map prediction model.
[0074] In this optional embodiment, generating a category embedding vector based on the category of each ultrasound image corresponding to the cardiac cycle and embedding the category embedding vector into a preset initial UNet network framework may include:
[0075] Generate a category embedding vector e according to the category of each ultrasound image corresponding to the cardiac cycle v ∈R d , where d is the dimension of the embedding vector;
[0076] The embedding vector e is based on the following formula v The mapping embedding vector e is obtained by mapping the channel dimension C of the feature map output by the convolutional layer of the initial UNet network framework to the fully connected layer of the initial UNet network framework. v1 :
[0077] e v1 =W·e v +b∈R C
[0078] In the above formula, W∈R C×d and b∈R C are the weights and biases of the fully connected layer of the initial UNet network framework;
[0079] Embed the mapping into vector e v1 The class embedding vector is expanded to match the dimension of the feature map and fused with the feature map, thus embedding it into the preset initial UNet network framework. This design explicitly embeds view information into the feature extraction process, enhancing the model's view perception capabilities.
[0080] In this embodiment, keypoint detection plays a crucial and fundamental role in the cardiac ultrasound keypoint tracking task. Keypoint detection accurately locates anatomically significant keypoints in keyframe cardiac ultrasound images, providing high-quality pseudo-labels for subsequent keypoint tracking based on label propagation. Therefore, in an optional embodiment, a phased training strategy is designed for the cardiac ultrasound keypoint detection task. First, a multi-view model is trained, and then its characteristics are leveraged to train a single-view model, improving the performance and robustness of the single-view model.
[0081] Therefore, the above-mentioned heat map prediction model obtained by training the initial UNet network framework based on multiple ultrasound training images corresponding to the cardiac cycle may include:
[0082] Acquire multiple multi-view ultrasound images corresponding to the cardiac cycle at multiple viewpoints, train the initial UNet network framework based on the multi-view ultrasound images, and obtain an intermediate heat map prediction model;
[0083] A plurality of single-view ultrasound images corresponding to a cardiac cycle at a fixed viewing angle are obtained, and an intermediate heat map prediction model is trained based on the single-view ultrasound images to obtain a heat map prediction model.
[0084] In this optional embodiment, in the first stage, a multi-view model is trained. The multi-view model is jointly trained using cardiac ultrasound images from multiple perspectives to fully explore the feature correlations and geometric constraints between different perspectives. Through joint learning of multi-view data, the model can capture the shape consistency and position correlation of the cardiac structure under different perspectives, thereby establishing a more comprehensive geometric representation. The model at this stage is based on the initial UNet network framework and is trained on multi-view input data. Through such multi-view training, the model not only improves the global consistency of key point detection, but also provides high-quality initial weights for the subsequent training of the single-view model.
[0085] In the second stage, the multi-view model (intermediate heatmap prediction model) from the first stage is used as a pre-trained model to further train a single-view model. The single-view model focuses on keypoint detection in a single ultrasound image. Combined with the geometric consistency and shape properties learned from the multi-view model, the single-view model achieves higher-precision keypoint detection from specific viewing angles. Compared to training a single-view model from scratch, this multi-view pre-training strategy significantly accelerates convergence while reducing the risk of overfitting.
[0086] The embodiments of the present invention further discovered that traditional loss functions such as L2 Loss, BCE Loss, Wing Loss, and Adaptive Wing Loss are usually used to measure the difference between predicted key points and labeled key points. These loss functions mainly focus on the error between single points, but may not be sufficient to effectively capture shape information in scenarios where there is a certain geometric relationship between key points. Therefore, based on the specific requirements of the task, the embodiments of the present invention design a shape-aware loss function for heatmap regression to introduce geometric constraints between key points in the optimization process, thereby generating a more structured predicted heatmap.
[0087] Therefore, in another optional embodiment, the aforementioned heat map prediction model obtained by training the initial UNet network framework based on multiple ultrasound training images corresponding to the cardiac cycle may include:
[0088] The initial UNet network framework is trained based on multiple ultrasound training images corresponding to the cardiac cycle. The training process is fed back based on a preset target loss function to obtain a heat map prediction model that has been trained to convergence.
[0089] Among them, the preset target loss function is calculated as follows for each key point:
[0090] Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation;
[0091] The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
[0092] In this optional embodiment, further optionally, determining the loss function corresponding to the key point according to the deviation between the real heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point may include:
[0093] The loss function corresponding to the key point i is calculated according to the following formula
[0094]
[0095] In the above formula, y is the real heat map, is the prediction heat map, E i is the geometric weight, is the cross entropy loss corresponding to the key point i, which is calculated by the following formula
[0096]
[0097] In the above formula, N is the total number of heat maps.
[0098] In this optional embodiment, represents the cross entropy loss calculated for each pixel position, which is used to evaluate the matching degree between the predicted heat map and the true heat map at the pixel level, and the additional geometric weight term E i The importance of the geometric relationship between key points can be quantified by shape-aware constraints. i The loss function is calculated based on the geometric deviation between the extracted keypoint coordinates, reflecting each keypoint's contribution to overall structural consistency. This weighted mechanism allows the loss function to not only focus on the accuracy of local pixel values but also globally optimize the structural relationships between keypoints, thereby better ensuring the anatomical plausibility of the predicted heatmap. This formula unifies the dual optimization objectives of pixel-level and global structural level through weighted accumulation of all keypoints and images.
[0099] Example 2
[0100] The embodiment of the present invention is based on semi-supervised key point tracking technology. By combining a small amount of labeled data with a large amount of unlabeled data, the model is used to learn the potential features of the unlabeled data, effectively improving the accuracy and generalization ability of key point detection. Its main inventive concept is to achieve key point tracking of cardiac ultrasound based on a pseudo-label method rather than unsupervised regularization. The main reason is that the pseudo-label method can make full use of unlabeled data and quickly expand the amount of training data through online generation and label propagation, thereby significantly improving the performance of the model. In the field of cardiac ultrasound, where labeling costs are high and data is scarce, the pseudo-label method has the advantages of being simple to implement and easy to expand, and is suitable for capturing detailed features of anatomical structures. In contrast, although the unsupervised regularization method performs well in improving the generalization ability of the model, it has a high reliance on the design of the regularization strategy, has high computational complexity, and may converge slowly in specific tasks. Therefore, the pseudo-label method is more in line with the requirements of this study for high efficiency and application adaptability.
[0101] The cardiac ultrasound key point tracking process in the embodiment of the present invention is as follows: Figure 3 First, the key frame prediction model is used to extract key frame indexes from cardiac cycle ultrasound images to ensure that the focus of data processing is on the key frame area; then, the heat map-based key point detection algorithm is used to generate the initial heat map and generate high-quality pseudo labels online; then, the label propagation model is used to combine high-quality pseudo labels to generate pseudo labels for intermediate frames, fully exploring the potential value of unlabeled data; finally, the generated intermediate frame pseudo label library is used to further optimize the training of the pseudo label propagation model, significantly improving the accuracy and robustness of cardiac key point detection, providing solid technical support for subsequent cardiac function analysis.
[0102] In an embodiment of the present invention, specifically, the first step of the cardiac ultrasound key point tracking process is to extract key frame indexes from the entire cardiac cycle. Optionally, the ED and ES frames can be accurately identified by using a key frame prediction model. These key frames have significant physiological characteristics in the cardiac cycle. The ED and ES frames represent the maximum and minimum moments of the left ventricular volume in the cardiac cycle, respectively. They are indispensable key nodes in cardiac function assessment and are widely used to measure indicators such as ejection fraction, ventricular wall motion, and cardiac chamber volume. By focusing on these key frames, the computational overhead of processing invalid information can be significantly reduced, and resources can be concentrated on frames with the highest diagnostic value.
[0103] In this embodiment of the present invention, a keypoint detection algorithm based on heatmap regression can optionally be used on the extracted keyframes to generate initial keypoint heatmaps. Its primary goal is to use the network model to identify the initial keypoint locations in the keyframe images and output them as a probability distribution. These heatmaps are high-quality pseudo-labels generated online.
[0104] To fully utilize pseudo-labeled data, embodiments of the present invention employ a label propagation learning model to generate pseudo-labels for intermediate frames. By integrating pseudo-labeled data from keyframes, the model continuously optimizes the quality of pseudo-labels during the learning process. Specifically, this involves leveraging the high-quality labels from keyframes to guide keypoint detection in intermediate frames. Simultaneously, through iterative pseudo-label generation and optimization, the model's adaptability to unlabeled data is gradually enhanced.
[0105] In an embodiment of the present invention, optionally, the generated intermediate frame pseudo labels are systematically stored to form a pseudo label library. This pseudo label library covers the unlabeled intermediate frame data in the cardiac cycle, making up for the lack of data labeling. The pseudo label library is used to further optimize the training process of the label propagation model. By combining labeled data and high-quality pseudo labels, the model can more accurately learn the geometric relationship and dynamic change characteristics between key points. Finally, the key point detection results are generated based on the heat map regression method. Heat map regression accurately locates the position of each key point by generating a heat map of the probability distribution.
[0106] In an optional embodiment, keypoint detection is the first step in keypoint tracking. Accurately predicting the keypoint locations provides accurate initial conditions for subsequent time series tracking, while ensuring geometric consistency throughout the tracking sequence. Only by accurately locating keypoints in each frame can the smoothness and rationality of keypoint trajectories be ensured, thereby reflecting the motion patterns of the heart.
[0107] In this optional embodiment, the task of cardiac ultrasound key frame prediction can be described as a video frame classification problem, aiming to predict the probability of each frame being ED or ES. This task can be formalized as a supervised learning problem, where the model input is the ultrasound video frame and the output is the probability of each frame belonging to the ED or ES category. Let the input be the ultrasound video frame X = {x1, x2, ..., x N}, the model output is each frame x i Classification probability:
[0108] f θ (x i )=P(y i =ED|x i ;θ)∈[0,1](4.1)
[0109] Where P(y i =ED|x i ; θ) is the frame x i The probability of ED is higher, the closer it is to 1, the more likely it is ED; 1-P(y i =ED|x i ; θ) is the frame x i The probability of being ES, the closer it is to 1, the more likely it is ES.
[0110] In this optional embodiment, the pre-processing process of the ultrasound video may include extracting 60 consecutive frames from the video (ensuring that at least one cardiac cycle is covered), cropping each frame, scaling it to 112×112, and normalizing the pixel values to [0,1] or standardizing it to zero mean unit variance. Data enhancement is used to improve data diversity and adapt to different imaging conditions. In addition, to retain time series information, frame indexes or time annotations can be attached. Finally, the processed frame sequence is converted into a tensor format and saved in an efficient storage format for use in model training.
[0111] In this optional embodiment, the model can use a deep network architecture based on Video ResNet_MultiMLP_v2 to extract the spatiotemporal features of the ultrasound video through R2Plus1dStem and multi-layer Bottleneck modules, gradually compressing the temporal and spatial resolution while expanding the number of feature channels. Finally, the model generates an output through fully connected layers and multi-layer perceptrons: a prediction result of the keyframe probability for each frame [1,60].
[0112] In this optional embodiment, the loss used for key frame prediction is a binary cross-classification loss, which is mainly used to optimize the prediction probability of each key frame output by the model. Specifically, for a key frame prediction output of an input batch And the true label y, the loss function L frame Defined as:
[0113]
[0114] Where N is the total number of video frames in the batch, y i is the true label of the i-th frame, is the probability predicted by the model that the frame is a key frame.
[0115] In this optional embodiment, post-processing analyzes the probability distribution of keyframe predictions, using local extremum detection and threshold filtering to identify ED (end-diastole) and ES (end-systole) frames. Specifically, this involves finding local peaks with probabilities close to 1 as ED frames and local valleys close to 0 as ES frames. Valid frames are then filtered based on a set threshold range (0.9 for ED and 0.1 for ES). Ultimately, keyframe extraction from the ultrasound video is complete.
[0116] The accuracy of key frame prediction is used to evaluate the model's ability to correctly classify end-diastolic and end-systolic frames, and is defined as the ratio of all correctly predicted frames to the total number of frames. The specific formula is:
[0117]
[0118] Among them, N correct Indicates the number of frames correctly classified as ED or ES, N total is the total number of key frames in the video.
[0119] In yet another optional embodiment, keypoint detection plays a crucial and fundamental role in the cardiac ultrasound keypoint tracking task. Keypoint detection accurately locates anatomically significant keypoints in keyframe cardiac ultrasound images, providing high-quality pseudo-labels for subsequent keypoint tracking based on label propagation.
[0120] In this optional embodiment, a phased training strategy is designed for the task of cardiac ultrasound key point detection. First, a multi-view model is trained, and then its characteristics are used to train a single-view model to improve the performance and robustness of the single-view model. The specific process is as follows: Figure 4 shown.
[0121] In this optional embodiment, in the first stage, a multi-view model is trained. The multi-view model is jointly trained using cardiac ultrasound images from multiple perspectives to fully explore the feature correlations and geometric constraints between different perspectives. Through joint learning of multi-view data, the model can capture the shape consistency and position correlation of the cardiac structure under different perspectives, thereby establishing a more comprehensive geometric representation. The model at this stage is based on the improved ResUNet in Example 1 of the present invention, and is trained on multi-view input data. Through such multi-view training, the model not only improves the global consistency of key point detection, but also provides high-quality initial weights for the subsequent training of single-view models.
[0122] In the second stage, the multi-view model from the first stage is used as a pre-trained model to further train a single-view model. The single-view model focuses on keypoint detection in a single ultrasound image. Combined with the geometric consistency and shape properties learned from the multi-view model, the single-view model achieves higher-precision keypoint detection from specific viewing angles. Compared to training a single-view model from scratch, this multi-view pre-training strategy significantly accelerates convergence while reducing the risk of overfitting.
[0123] In yet another optional embodiment, the label propagation model is designed as follows:
[0124] The application of keypoint tracking technology based on label propagation in cardiac ultrasound videos provides an innovative solution for keypoint tracking. By utilizing a small amount of predicted data, it achieves efficient labeling while significantly reducing cost and computational complexity. This approach addresses the time-consuming and highly dependent nature of traditional labeling methods. By leveraging pseudo-label generation technology and deep learning models, it overcomes the bottleneck of keypoint tracking in cardiac ultrasound videos.
[0125] This optional embodiment uses STCN as the backbone framework of the key point tracking network based on label propagation. The core advantage of STCN is that it can directly establish pixel-level correspondences between frames without the need to encode mask features for each object separately. Traditional methods often require the extraction of object features frame by frame, while STCN directly constructs inter-frame matching relationships through an efficient inter-frame similarity calculation mechanism, which greatly reduces computational redundancy and processing complexity. At the same time, STCN adopts a lightweight association mechanism, combined with the characteristics of time series data, to provide stable and efficient support for target tracking tasks. In scenarios based on pseudo-label learning, especially for time series data (such as cardiac ultrasound videos), there is a high correlation and smoothness between frames. Through its unique inter-frame matching mechanism, STCN can fully explore the continuity of time series and adapt to the smoothness characteristics of key point trajectories in the cardiac cycle.
[0126] Taking advantage of the strong inter-frame correlation and smooth trajectory characteristics of cardiac ultrasound time series data, STCN fully exploits the continuity of key point motion during the cardiac cycle. This inter-frame matching mechanism not only improves the efficiency of pseudo-label generation, but also ensures the smoothness and logical consistency of key point detection trajectories, adapting to the dynamic changes of the heart.
[0127] However, the STCN basic feature extraction module has certain limitations in its ability to represent complex heat map features. To this end, this optional embodiment further optimizes the feature extraction module and introduces DenseNet as a heat map feature extractor. DenseNet fully integrates shallow features with deep features through a dense connection mechanism, achieving efficient feature reuse and improved fine-grained representation capabilities. This structure not only enhances the expressive power of heat map features, but also improves the accuracy and robustness of inter-frame matching.
[0128] Furthermore, to improve the quality and reliability of pseudo-label generation, this optional embodiment also designs an optimized pseudo-label generation strategy based on the continuity of time series and the regularity of the cardiac cycle. In the specific implementation, initial pseudo-labels for intermediate frames are generated through preliminary training on a small amount of labeled data, and the pseudo-labels are iteratively optimized in combination with the smoothness characteristics of the time series. Ultimately, the improved pseudo-label quality not only expands the amount of available training data but also significantly enhances the model's ability to learn from unlabeled data, laying a solid technical foundation for efficient keypoint tracking.
[0129] In this optional embodiment, the network structure of STCN is mainly developed around inter-frame association and feature extraction, such as Figure 5 As shown. Its core consists of two parts: first, the query frame (Query) and the memory frame (Memory) respectively extract features through independent feature encoders (Key Encoder), among which the features of the memory frame will be further used to establish the spatiotemporal correspondence between frames; second, the association module (Affinity) based on the inter-frame similarity calculation directly establishes a pixel-by-pixel matching relationship between the query frame and the memory frame without the need to encode each object separately. Through the decoder module (Decoder), the inter-frame similarity matrix is converted into a target prediction output. The overall design of STCN achieves efficient inter-frame feature transfer and matching through skip connections and lightweight association mechanisms, which is particularly suitable for target tracking tasks in time series data.
[0130] The heat map feature extraction module is crucial in the key point tracking task. Its goal is to extract high-quality features from the heat map to support the calculation of the subsequent inter-frame association and matching modules. For STCN, the heat map feature extraction module directly determines the effect of inter-frame matching, and its accuracy and robustness play a key role in the performance of the entire model. Heat map information contains fine-grained features such as the spatial position and probability distribution of key points. This high complexity requires the feature extraction module to not only capture shallow spatial details, but also to effectively fuse deep semantic information. Although the traditional lightweight feature extraction module has high computational efficiency, it is prone to information loss or insufficient feature expression when processing complex heat map features, and cannot meet the needs of high-precision inter-frame matching. Therefore, this optional embodiment proposes to replace the lightweight feature extraction module in the original model with DenseNet to improve the feature representation capability in response to the complexity of heat map information.
[0131] In the STCN network structure of this optional embodiment, DenseNet is integrated into the heat map feature extraction module, replacing the traditional lightweight feature extractor. The high-quality features generated by DenseNet provide richer and more detailed feature inputs for the association module of STCN, thereby significantly improving the effect of inter-frame similarity calculation. In addition, the feature fusion capability of DenseNet also enables STCN to better capture the continuity and regularity between frames when processing time series data, thereby optimizing the accuracy of inter-frame matching and the overall performance of the model.
[0132] In the embodiment of the present invention, in order to verify the beneficial effects of the embodiment of the present invention, the following experimental tests are disclosed:
[0133] (1) Experimental verification of cardiac ultrasound key point detection
[0134] The present invention designs a set of ablation experiments to evaluate the performance of the two-stage training strategy in the task of cardiac ultrasound key point detection. In this experiment, all models are trained and validated based on the cardiac ultrasound key point detection dataset introduced in the present invention. The heatmap regression method adopts the key point redefinition method based on set significance in the present invention. The optimizer is Adam, the initial learning rate is set to 0.001, and the batch_size is set to 16. All training configurations in the experiment remain consistent to ensure the fairness of the comparison results. The results of the comparative experiment are shown in the following table.
[0135]
[0136] In multi-view joint training, the model not only learns the independent features of each view, but also improves the consistency of key point positioning by optimizing the shared geometric space. This cross-view geometric inference capability is not achievable by single-view models, and it effectively reduces the error caused by feature ambiguity in certain view angles.
[0137] The performance of single-view models improves significantly in weak view angles, a direct contribution of multi-view pre-training. Single-view models for weak view angles, such as PRVIT and PSAXAO, perform poorly in initial training. This is likely because the feature distribution of these view angles is relatively limited, making it difficult to provide sufficient geometric constraints for deep learning models. However, through pre-training of multi-view models, these view angles benefit from the transfer of global geometric consistency. The multi-view pre-training model gives the single-view model initial shape perception capabilities, allowing it to quickly find reasonable anatomical positioning in the early stages of training even when features are insufficient.
[0138] The combination of viewpoints is crucial for achieving marginal improvements in model performance. The full view + A4C combination achieved the best results, thanks to the crucial role of the A4C view in cardiac anatomy. The distribution of key points in the A4C view encompasses the geometric centers of the major cardiac chambers, and these points effectively enhance the geometric relevance of other viewpoints during multi-view training.
[0139] Overall, the experimental results highlight the effectiveness of the staged training strategy. Multi-view pre-training not only provides a performance foundation for the single-view model, but also enhances robustness through the transfer of geometric consistency.
[0140] (2) Experimental verification of the effectiveness of label propagation methods based on dataset annotation
[0141] Based on the annotation of key frames and key points on key frames in a cardiac ultrasound dataset by professional doctors, the embodiment of the present invention constructs a comparative experiment for key point tracking by generating pseudo labels for intermediate frames. The experiment aims to evaluate the performance of the pseudo labeling method in the task of cardiac ultrasound key point tracking, and explore the impact of key frame annotation on the accuracy and robustness of key point tracking. Specifically, the experiment in the embodiment of the present invention utilizes the annotated end-diastolic and end-systolic frames and their key points in the dataset, combined with a learning strategy based on pseudo labels, to expand the scope of utilization of unlabeled intermediate frames, thereby improving the performance of key point detection.
[0142]
[0143] All models in this experiment were trained and validated based on the cardiac ultrasound key point detection dataset introduced in the embodiments of the present invention. The key point prediction method adopted the heatmap regression method based on set meaning redefinition in the embodiments of the present invention. During the training process, Adam was selected as the optimizer, the initial learning rate was set to 0.001, and the batch size was set to 16 to ensure the stability of the training process and the optimization efficiency. All experimental configurations remained consistent, thus ensuring the fairness and credibility of the comparison results between different models. The experimental results are shown in the table below, which provides a comprehensive quantitative evaluation of the performance of the pseudo-label key point tracking method.
[0144] With fully annotated data, the Full-set STCN achieved the highest detection success rate and lowest mean error. This demonstrates that increasing the amount of annotated data can effectively improve model performance. However, the limited cost of annotation makes this approach difficult to scale in practical applications. Therefore, more efficient use of small amounts of annotated data has become a research priority.
[0145] The STCN model has a high detection success rate. Despite a significant reduction in the amount of labeled data, the model is still able to learn key point features from inter-frame similarities, demonstrating the superiority of its inter-frame association mechanism. Furthermore, the STCN model, which incorporates DenseNet, enhances feature representation capabilities through a dense connection mechanism, resulting in improvements in detection success rate and error. However, due to its increased structural complexity, the performance improvement is limited.
[0146] The introduction of a pseudo-label generation strategy based on cardiac ultrasound significantly improved the performance of STCN. By optimizing the utilization of unlabeled frames, the STCN+modified strategy achieved a detection success rate of 89.65%. This pseudo-label generation strategy leverages the continuity and periodicity of time series to significantly improve the prediction accuracy of unlabeled frames, thereby reducing the reliance on large-scale labeled data.
[0147] The improved STCN model, which combines DenseNet with a pseudo-label generation strategy, achieved the best performance, with a detection success rate of 90.72%. This result not only approaches the performance of fully annotated models but also significantly reduces the need for annotated data. By enhancing feature extraction capabilities and optimizing the utilization of unlabeled data, the improved STCN provides a more efficient and cost-effective solution to heatmap regression methods.
[0148] (3) Verification of the effectiveness of label propagation methods based on pseudo-labels
[0149] The key frame prediction model and key point detection model of the embodiment of the present invention are designed to further optimize the key point annotation of the intermediate frames by generating high-quality pseudo labels, thereby realizing a comparative experiment of key point tracking. The core goal of this experiment is to use high-quality pseudo labels and evaluate the performance of this strategy in the key point tracking task. All models in this experiment are trained and verified based on the cardiac ultrasound key point detection dataset introduced in the embodiment of the present invention. The key point prediction method adopts the heat map regression method based on set meaning redefinition in the embodiment of the present invention. The accuracy of the key frame prediction model in the embodiment of the present invention is shown in the following table.
[0150]
[0151] In this embodiment of the present invention, the experiment first accurately identified the ED and ES frames in cardiac ultrasound videos using a keyframe prediction model. This model, combined with a keypoint detection model, generated precise keypoint annotation heatmaps for these keyframes. Furthermore, by combining the continuity of the time series with the periodicity of the cardiac cycle, high-quality pseudo-labels were generated for the unlabeled intermediate frames. Through this iterative pseudo-label generation process, a complete keypoint tracking solution was constructed.
[0152] During training, all models used the same optimization configuration, including the Adam optimizer, an initial learning rate of 0.001, and a batch size of 16, to ensure stability and consistency in model training. Experimental results systematically evaluated the performance of the pseudo-label generation strategy, demonstrating its advantages and potential in cardiac ultrasound keypoint tracking tasks, as shown in the table below.
[0153]
[0154] Under fully annotated data conditions, the Full-set STCN demonstrated optimal performance, achieving the highest detection success rate and the lowest average point-to-point error. This demonstrates that by fully utilizing annotated data, the model can effectively capture key point features and provide highly accurate predictions. However, the high cost of annotation and the difficulty in acquiring large-scale annotated data make this approach difficult to scale in practical applications. Therefore, effectively utilizing limited annotated data becomes crucial.
[0155] The improved STCN model based on pseudo-labeling achieves excellent performance by introducing a densely connected feature extraction mechanism (such as DenseNet) and a pseudo-label generation strategy based on time series characteristics. In the pseudo-labeling method, the detection success rate reached 87.63%, narrowing the performance gap with the fully labeled model, with only a slight increase in the average error. This strategy effectively expands the training dataset and significantly alleviates the dependence on manual labeling. In particular, the performance of the improved STCN in the pseudo-labeling scenario fully demonstrates the contribution of the improvement in feature extraction capabilities and the optimization of the pseudo-label generation strategy to model accuracy. Compared with the unimproved pseudo-label + STCN method, the improved STCN model improved the detection success rate by about 2% and reduced the average error by 0.28.
[0156] Overall, the pseudo-labeling method combined with the improved STCN provides an efficient and low-cost solution for scenarios with limited labeled data. It not only approaches the performance of fully labeled models but also significantly reduces reliance on manual labeling, providing important technical support and practical value for cardiac ultrasound key point tracking.
[0157] Example 3
[0158] An embodiment of the present invention discloses a device for tracking cardiac ultrasound key points based on pseudo labels, which may include:
[0159] a key frame extraction module, configured to extract key frames from a set of ultrasound images corresponding to a cardiac cycle, wherein the key frame may include at least one change node in the cardiac cycle, and ultrasound images other than the key frame in the set of ultrasound images are intermediate frames;
[0160] A key point detection module is used to extract the key points on each key frame as key pseudo labels;
[0161] A label propagation module is used to detect key points on each intermediate frame based on the key pseudo-label and the time series relationship between the ultrasound images in the ultrasound image set based on a preset label propagation model, and extract the key points corresponding to the key pseudo-label on each intermediate frame as the intermediate pseudo-label;
[0162] The model optimization module is used to optimize the training of the label propagation model based on the intermediate pseudo-labels to obtain an optimized label propagation model; wherein the optimized label propagation model is used to track key points of the ultrasound video corresponding to the cardiac cycle.
[0163] In an optional embodiment, the key point detection module extracts key points on each key frame as key pseudo labels, and the specific operation method may include:
[0164] For each key frame, the key frame is input into the preset heat map prediction model. The heat map prediction model extracts all key points corresponding to the key frame according to the preset geometric features corresponding to the category of the key frame; the boundary path of the geometric distribution of the key points is generated according to the preset geometric features, the key points are sorted according to the preset boundary order, each layer of the heat map is restricted to represent only one key point, a predicted heat map corresponding to each key point is generated, and the key point coordinates are extracted from the predicted heat map based on the preset key point extraction strategy.
[0165] In another optional embodiment, the heat map prediction model is obtained by the following steps:
[0166] According to the category of each ultrasound image corresponding to the cardiac cycle, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the cardiac cycle to obtain a heat map prediction model.
[0167] In yet another optional embodiment, the specific operation of training the initial UNet network framework according to multiple ultrasound training images corresponding to the cardiac cycle to obtain the heat map prediction model may include:
[0168] Acquire multiple multi-view ultrasound images corresponding to the cardiac cycle at multiple viewpoints, train the initial UNet network framework based on the multi-view ultrasound images, and obtain an intermediate heat map prediction model;
[0169] A plurality of single-view ultrasound images corresponding to a cardiac cycle at a fixed viewing angle are obtained, and an intermediate heat map prediction model is trained based on the single-view ultrasound images to obtain a heat map prediction model.
[0170] In another optional embodiment, the key point detection module extracts key points on each key frame as key pseudo labels, and the specific operation method may include:
[0171] For each key frame, a key point heat map representing key points on the key frame is extracted as a key pseudo label based on a preset heat map regression algorithm;
[0172] Furthermore, the label propagation module detects key points on each intermediate frame based on the key pseudo-label and the time series relationship between the ultrasound images in the ultrasound image set based on a preset label propagation model, and extracts key points corresponding to the key pseudo-label on each intermediate frame as the intermediate pseudo-label. The specific operation method may include:
[0173] For key pseudo labels, based on densely connected convolutional networks, shallow features and deep features are fully integrated to obtain heat map features corresponding to key pseudo labels;
[0174] Based on a preset inter-frame similarity calculation mechanism and in combination with the temporal sequence relationship between ultrasound images in the ultrasound image set, a matching relationship between intermediate frames is constructed;
[0175] For each intermediate frame, based on the heat map features corresponding to the key pseudo-label and the matching relationship between the intermediate frame and other intermediate frames, the key points corresponding to the key pseudo-label on the intermediate frame are extracted as the intermediate pseudo-label.
[0176] In yet another optional embodiment, the specific operation of the key frame extraction module to extract key frames from the set of ultrasound images corresponding to the cardiac cycle may include:
[0177] Acquiring an ultrasound video corresponding to a cardiac cycle, identifying an initial window region in the ultrasound video whose brightness change satisfies a preset brightness change pattern, and modifying the initial window region based on morphological features corresponding to the cardiac cycle to obtain an ultrasound window corresponding to the ultrasound video;
[0178] According to the ultrasound window corresponding to the ultrasound video, an ultrasound image set corresponding to the ultrasound video is extracted.
[0179] In yet another optional embodiment, the specific operation of training the initial UNet network framework according to multiple ultrasound training images corresponding to the cardiac cycle to obtain the heat map prediction model may include:
[0180] The initial UNet network framework is trained based on multiple ultrasound training images corresponding to the cardiac cycle. The training process is fed back based on a preset target loss function to obtain a heat map prediction model that has been trained to convergence.
[0181] Among them, the preset target loss function is calculated as follows for each key point:
[0182] Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation;
[0183] The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
[0184] Example 4
[0185] See also Figure 6 , Figure 6 FIG is a schematic diagram of a pseudo-label-based cardiac ultrasound key point tracking system disclosed in an embodiment of the present invention. Figure 6 As shown, the pseudo-label-based cardiac ultrasound key point tracking implementation system may include:
[0186] A memory 201 storing executable program code;
[0187] a processor 202 coupled to the memory 201;
[0188] The processor 202 calls the executable program code stored in the memory 201 to execute the steps of the method for implementing cardiac ultrasound key point tracking based on pseudo labels described in the first or second embodiment of the present invention.
[0189] Example 5
[0190] An embodiment of the present invention discloses a computer storage medium storing computer instructions. When the computer instructions are called, they are used to execute the steps of the method for implementing cardiac ultrasound key point tracking based on pseudo labels described in Embodiment 1 or Embodiment 2 of the present invention.
[0191] Example 6
[0192] Embodiment 6 of the present invention discloses a key point detection method for cardiac ultrasound influence, including the following contents:
[0193] This example uses heatmap regression for ultrasound keypoint detection, rather than direct coordinate regression. This is because heatmap regression exhibits significant advantages in spatial information preservation and model generalization, particularly in cardiac ultrasound images, where keypoints often exhibit significant anatomical structural variation and image noise. Furthermore, coordinate regression struggles to handle the interdependencies between multiple keypoints, making inconsistent predictions for multiple keypoints a problem.
[0194] Heatmap is a method to characterize the spatial location of key points through Gaussian distribution. For a given key point coordinate Its corresponding heat map g i (x;σ i ) is defined as x i The Gaussian function centered at . The formula is:
[0195]
[0196] where σ i Indicates the standard deviation of the Gaussian distribution and controls the diffusion range of the heat map. The heat map value is at the key point x i The probability reaches its peak at , and gradually decays as the distance increases. In this way, the heat map generates a probability distribution for each key point in the image, which can effectively capture local context information and spatial relationships. Figure 7 is an example of a heatmap representation of cardiac ultrasound key points.
[0197] Heatmap post-processing involves accurately extracting keypoint coordinates from the probabilistically distributed heatmap to remove noise and improve localization accuracy. Conventional methods include non-maximum suppression (NMS), which uses a local window to find peaks and suppress non-peak responses to generate a sparse heatmap. The top K maximum values in the heatmap are then sorted and selected as candidate keypoint locations, where K is the number of keypoints represented in a single heatmap.
[0198] However, this example found that existing cardiac ultrasound keypoint detection technologies face the following major challenges: First, existing methods lack optimization of the geometric constraints of cardiac keypoints. The cardiac anatomical structure has clear geometric distribution characteristics, such as the circular or linear distribution of the endothelium and epicardium. However, traditional methods process each keypoint independently, ignoring the geometric relationships between keypoints. This results in insufficient structural consistency and anatomical rationality in the prediction results. Specifically, multiple keypoints of the heart are interrelated. Ignoring these relationships can lead to errors in keypoint prediction and structural irrationality, thus affecting the accuracy of quantitative analysis of cardiac function.
[0199] Secondly, ultrasound image quality is poor, often affected by noise, artifacts, and blurred boundaries. A single frame cannot accurately reflect the anatomical features of key points. Even experienced physicians rely on the classification and time-series deformation information from different views to assist in judgment and compensate for the incompleteness and low quality of single-frame information. Therefore, extracting stable and accurate key information from low-quality ultrasound images is crucial to improving detection accuracy.
[0200] In addition, traditional heatmap regression methods rely on non-maximum suppression combined with Top-K screening to regress multiple key points, but their performance is highly sensitive to the NMS window size, which may lead to missed detections or false detections, especially when adjacent key points are close, and multiple peaks are easily mixed into one.
[0201] In response to these problems, this embodiment proposes a series of optimization schemes. The main inventive concept is: through the geometric redefinition of key points and the design of shape-aware loss functions, the geometric consistency and anatomical rationality of the key point detection results are enhanced; the view embedding module is introduced to explicitly embed view category information into the network to improve the model's adaptability to multi-view data; combined with the previous and next frame inputs of the key frame, the continuity of the time series is used to capture dynamic feature changes, and improve the prediction instability caused by insufficient single-frame information. In addition, in the heat map regression method, it is proposed to optimize key points through spatial distance constraints of candidate points, graph convolutional neural networks, and geometric sorting strategies to optimize the key point selection process, and the reliability of the method is verified through experiments.
[0202] In the embodiment of the present invention, original ultrasound data is first acquired, wherein the original ultrasound data may be a series of ultrasound videos or continuous ultrasound images.
[0203] Optionally, this embodiment can also preprocess the ultrasound data first, the main goal of which is to convert the input cardiac ultrasound key frames into tensors of uniform specifications for easy model processing. Optionally, data enhancement is performed on the input cardiac ultrasound key frames, for example, random rotation, translation, scaling and flipping can be included to simulate diverse deformations and improve the generalization ability of the model. Then, based on the balance between accuracy and performance overhead, all images are adjusted to a uniform 512×512 resolution. Assuming that the original image size is H×W×C, the image is scaled to 512×512×C through interpolation and other methods. Finally, the adjusted image is converted into a tensor form, and the data is standardized through normalization.
[0204] Optionally, data preprocessing in this embodiment can also include data cleaning. This includes verifying the video file path, checking for missing or erroneous dot data, filtering out records that do not meet view and chamber requirements, ensuring the completeness of ED / ES annotations, and verifying the format and location of the annotation data. Furthermore, the cleaning process includes validity checks on image frames and cropped regions to ensure that the data meets training requirements. These cleaning steps provide high-quality input data for subsequent model training, ensuring the reliability of experimental results.
[0205] In this embodiment of the present invention, a heatmap regression-based ultrasound imaging keypoint detection method uses UNet as the backbone network framework. This combines the powerful multi-scale feature extraction capabilities of the ResNet50 encoder with the UNet's skip connection mechanism to effectively integrate high-level semantic features with spatial details in keypoint detection tasks. Simultaneously, a bridging module further compresses and fuses global features, while the decoder gradually upsamples to restore resolution, generating accurate keypoint heatmaps.
[0206] The improved network framework of the embodiment of the present invention combines view embedding and time series information, and realizes accurate key point detection in multi-view and dynamic conditions by introducing dynamic features of previous and next frames and residual enhancement modules. The encoder can use ResNet50 for deep feature extraction, the middle layer integrates global context information, and the decoder combines deconvolution with multi-resolution jump connections to gradually restore spatial resolution and retain detail features. The view embedding module enhances the adaptability of the network to different anatomical perspectives, while the capture of the temporal characteristics of previous and next frames improves the robustness in dynamic scenes and the detection accuracy of blurred boundary areas, ultimately generating a high-resolution heat map, which provides an efficient and accurate solution for dynamic key point detection and quantitative analysis of cardiac ultrasound images.
[0207] There are multiple views in ultrasound images, and the distribution of key points in different views is obviously different. Even for the same key point, its position and direction will change in different views, such as Figure 8 shown. Figure 8 Figure 2 shows ultrasound images from the same view, with keypoints marked with blue dots. Despite the consistent view, keypoint distribution varies across frames. This is likely due to imaging noise, blurred boundaries, and dynamic changes in cardiac structure. Furthermore, accurate keypoint localization is difficult in single-frame ultrasound images with blurred boundaries or noise, especially in dynamic ultrasound sequences. The lack of information in a single frame further exacerbates prediction instability.
[0208] To address these issues, the key inventive concept of this embodiment lies in introducing a view residual embedding module to embed view category information into the model, helping the model learn the semantic differences between views and improving the accuracy of keypoint localization. Furthermore, by adding the frames before and after the keyframe as contextual information, the model's understanding of dynamic ultrasound sequences is enhanced by leveraging the continuity of the time series.
[0209] In an optional embodiment, the view embedding module is designed as follows:
[0210] Ultrasound images contain multiple views, and the anatomical structures and keypoint distribution patterns vary significantly across views. Relying solely on raw image features, the model may not fully understand the semantic differences between views, resulting in reduced keypoint detection accuracy. To address this issue, this optional embodiment introduces a view embedding module. By embedding view category information into the network, the model explicitly considers view type during feature extraction.
[0211] Optionally, the view embedding module converts the view category v∈{1,2,…,N} into an embedding vector e v =Embedding(v)∈R d , where d is the dimension of the embedding vector, generating a high-dimensional representation of the view category. Subsequently, the embedding vector is mapped to the channel dimension C of the feature map through a fully connected layer, and the transformation formula is e v1 =W·e v +b∈R C , where W∈R C×d and b∈R C are the weights and biases of the fully connected layer respectively. For the features output by the convolutional layer Figure X ∈R C×H×W , expand the embedding vector to a shape that matches the feature map, and complete feature fusion by pixel-by-pixel addition. The formula is X1=BN(X)+e v2 , where e v2 =e v1[…,None,None]∈R C×H×W Finally, the fused features can be activated by the ReLU function to generate the final output X out =ReLU(X1). This design explicitly embeds view information into the feature extraction process, enhancing the model’s view perception ability.
[0212] In another optional embodiment, in ultrasound images, a single frame of information is often insufficient to accurately locate keypoints due to blurred boundaries or noise interference. This is especially true in dynamic sequences, where the static features of a single frame may not provide sufficient contextual support. To address this issue, the frames before and after the keyframe are introduced as input. By incorporating the contextual information of the time series, the model can capture dynamic feature changes between consecutive frames, thereby improving the accuracy and robustness of keypoint detection.
[0213] Specifically, given the current key frame I t ∈R H×W×C and its preceding and following frames I t-1 and I t+1 , the input sequence is defined as: I seq =[I t-1 ,I t ,I t+1 ]. Where H and W are the height and width of the image, and C is the number of channels.
[0214] In another optional embodiment, the specific process of optimizing the heat map regression method is discussed as follows:
[0215] In the key point detection task, due to the lack of sequentiality in the annotation data of similar key points, conventional methods usually use a single heat map to represent the positions of all similar key points. Figure 9 As shown, four similar lv_wall key points rely on a heat map representation.
[0216] Typically, multiple keypoints are regressed from a single heatmap through non-maximum suppression combined with Top-K filtering. However, the performance of this method is highly dependent on the hyperparameter selection of the non-maximum suppression region radius. If the radius is set too small, the response areas between keypoints may overlap excessively, making it difficult to distinguish keypoints. If the radius is set too large, multiple peaks may be blended into a single one, affecting the precise localization of keypoints.
[0217] In order to further improve the accuracy and robustness of key point regression, this optional embodiment discloses multiple improvement strategies: a method based on spatial distance constraints of candidate points, a heat map regression based on graph convolutional neural networks, and a method based on geometric meaning to redefine key points.
[0218] In a first optional embodiment, a keypoint regression method based on candidate point spatial distance constraints is disclosed: First, a clustering algorithm (DBSCAN) is used to filter outliers from candidate points, eliminating noise points with unusual distributions. Then, the filtered point set is sorted by confidence and spatial distance constraints are applied. This method ensures that high-confidence keypoints are evenly distributed in space and reduces noise interference.
[0219] In a second optional embodiment, a heatmap regression method based on a graph convolutional neural network is disclosed:
[0220] By using graph convolutional neural networks to dynamically learn key point distribution and selection strategies, we overcome the shortcomings of traditional non-maximum suppression in fixed window hyperparameters, neighboring point differentiation, and global information modeling.
[0221] The network takes a multi-keypoint heatmap as input and first extracts local features through a convolutional module, thereby enhancing the understanding of the area near the keypoints. Subsequently, an attention mechanism is introduced to model the distribution of keypoints from a global perspective, capturing long-range dependencies and generating global semantic features. A linear mapping module is used to reduce the dimensionality of high-dimensional features. The network further extracts a compact feature representation, providing efficient input for subsequent optimization steps.
[0222] The reduced features are fed into a multi-layer graph convolutional network. This module models the spatial distribution and topological relationships between key points through layer-by-layer processing, gradually optimizing the key point prediction results. Ultimately, the network generates accurate key point coordinates and completes the position mapping by aligning them with the original image. The overall network combines the advantages of convolution, attention, and graph convolution, such as Figure 10 As shown, the entire process can be optimized from rough heat maps to precise coordinates.
[0223] In a third optional embodiment, a method for redefining key points based on geometric meaning is disclosed:
[0224] Keypoints of the same type are sorted and matched using geometric features, ensuring a fixed order for each type of keypoint. During the heatmap generation phase, this solution directly restricts each heatmap layer to represent only one keypoint, preventing duplicate occurrences of similar points within the heatmap and eliminating the need for complex non-maximum suppression operations. In post-processing, only the maximum coordinates need to be extracted from each heatmap, simplifying the keypoint selection process.
[0225] Based on the characteristics of key points in cardiac ultrasound medicine, the geometric distribution characteristics of key points are very clear. Key points are distributed almost sequentially on the anatomical structures of the endocardium and epicardium. The endocardium and epicardium have regular circular or linear distribution patterns in anatomy, which provides a natural basis for sorting and matching key points.
[0226] First, the image is corrected through key points, and the anatomical structure is adjusted to a unified reference coordinate system to ensure that the anatomical relationship in the image is consistent with the actual structure. Secondly, on the corrected image, the candidate key points are screened and sorted based on the key point matching strategy of convex hull search: First, the candidate key points are extracted and their convex hull boundaries are calculated to generate a boundary path reflecting the geometric distribution; then, the key points are sorted according to the order of the convex hull boundaries. For annular or linear anatomical structures, they are arranged in counterclockwise, clockwise, or from base to apex, respectively. The specific matching strategy is as follows: Figure 2 As shown in Figure 3, the positions of key points are optimized using a convex hull search algorithm to ensure that the key point distribution conforms to the spatial constraints of the anatomical structure.
[0227] In the heat map generation stage, this method directly limits each layer of heat map to represent only one key point, thus avoiding the repeated distribution of similar points in the heat map. This improvement makes the post-processing stage no longer rely on the complex non-maximum suppression algorithm, and only needs to extract the maximum coordinates in each heat map to determine the key point position p max :
[0228]
[0229] Among them, y i,j Represents the confidence value of pixel (i, j) in the heat map.
[0230] In yet another optional embodiment, the loss function optimization scheme is discussed as follows:
[0231] In the task of detecting key points in cardiac ultrasound, traditional loss functions such as L2 Loss, BCE Loss, Wing Loss, and Adaptive Wing Loss are usually used to measure the difference between predicted key points and annotated key points. These loss functions mainly focus on the error between single points, but may not be sufficient to effectively capture shape information in scenarios where there is a certain geometric relationship between key points. Therefore, based on the specific requirements of the task, this optional embodiment designs a shape-aware loss function (Shape-aware Loss) for heat map regression to introduce geometric constraints between key points in the optimization process, thereby generating a more structured predicted heat map.
[0232] The true heat map y represents the spatial distribution of each key point in the image, which is usually represented by the probability distribution of the key point position in the form of Gaussian distribution, where the center point is the true coordinate of the key point and the value of the surrounding area gradually decays; the predicted heat map It is the result generated by the deep learning model after extracting the features of the input image, which represents the model's probability estimate of the key point location. The difference between them is used to guide model training, and the error between them is minimized by optimizing the loss function, thereby improving the accuracy of key point positioning and the generalization ability of the model.
[0233] L2 Loss is a classic regression loss function used to measure the predicted heatmap The pixel-wise squared error between the heatmap y and the true heatmap y is simple and easy to use and suitable for scenarios with uniform error distribution, but it is sensitive to outliers because the squared error amplifies the impact of the error point. In the task of cardiac ultrasound keypoint detection, L2 Loss can optimize the similarity of heatmaps overall, but lacks a focus on keypoint regions. The formula is as follows, where N is the total number of heatmaps:
[0234]
[0235] BCE Loss is used to measure the probability distribution of each pixel in the predicted heat map belonging to a key point The difference between the probability distribution y and the true probability distribution y. It is suitable for processing situations with sparse key points and a large background, and effectively optimizes pixel classification by calculating the probability difference pixel by pixel. However, BCE Loss ignores the geometric relationship between key points and cannot capture the global structure. The formula is as follows:
[0236]
[0237] Wing Loss is a keypoint localization loss function that optimizes small errors by performing logarithmic scaling on them, while applying linear processing to large errors to reduce the impact of outliers. In keypoint heatmap regression tasks, Wing Loss achieves fine-grained optimization of small errors by reducing the impact of large errors, making it suitable for scenarios requiring high-precision localization. The specific formula is as follows: ω controls the sensitivity of the small error region, ∈ determines the smoothness of the logarithmic function, and C ensures the continuity of the loss.
[0238]
[0239] Adaptive Wing Loss is an improved version of Wing Loss. By dynamically adjusting weights, the model becomes more sensitive to errors in the area near keypoints. Compared to Wing Loss, Wing Loss focuses more on the area near keypoints and dynamically adjusts the weights for different errors, resulting in better performance in keypoint detection tasks with complex backgrounds or sparse distributions. The specific formula is as follows, where α controls the degree of dynamic weighting, and |y| is the heatmap pixel value, which is used to adjust the importance of the keypoint.
[0240]
[0241] In the task of cardiac ultrasound key point detection, traditional loss functions such as L2 Loss, BCE Loss and Wing Loss mainly focus on the pixel-level differences between predicted points and true points. However, they ignore the geometric relationship and global structural information between key points, such as Figure 11 In order to better capture the anatomical structural features of key points, a shape-aware loss function is designed to incorporate structural information into the loss optimization process by introducing geometric relationship constraints between key points.
[0242] In heatmap regression, soft argmax is a differentiable method for extracting the coordinates of predicted key points. By using soft argmax, the location of the maximum value in the heatmap can be smoothly estimated, overcoming the non-differentiable problem of the traditional argmax method. The specific formula is as follows:
[0243]
[0244] Where τ is a temperature parameter that controls the degree of smoothing. When τ is close to 0, the result of soft argmax is close to traditional argmax, that is, almost only the index of the largest element contributes to the result.
[0245] Shape-aware loss uses soft argmax to extract the keypoint coordinates between the predicted and ground-truth heatmaps and calculates the geometric relationship deviation between the keypoints. By introducing this geometric constraint, shape-aware loss not only optimizes the localization error of individual keypoints but also enhances the consistency of the structural relationships between keypoints, thereby generating a predicted heatmap that is more consistent with anatomical features. The specific formula is as follows:
[0246]
[0247] Among them, y i,j represents the true heat map value of the jth key point in the i-th image, is the corresponding predicted heat map value, and the coordinates extracted by soft argmax are used to calculate the geometric distance Dist.
[0248] Traditional loss functions such as BCE Loss mainly target pixel-level predictions, but ignore the geometric relationship between key points. In order to simultaneously optimize the pixel-level accuracy of the heat map and the global geometric relationship between key points, this optional embodiment proposes a new loss function - Shape-aware Binary Cross Entropy Loss (Shape-aware BCE Loss). This loss function combines BCE Loss and Shape-aware Loss. Among them, BCE Loss is used to measure the difference between the predicted probability and the true value of each pixel, while the shape-aware part extracts the coordinates of the predicted key points and the true key points through soft argmax, and calculates geometric constraints to optimize the structural relationship between the key points. The new loss function is defined as:
[0249]
[0250] in, represents the cross entropy loss calculated for each pixel position, which is used to evaluate the matching degree between the predicted heat map and the true heat map at the pixel level, and the additional geometric weight term E i The importance of the geometric relationship between key points is quantified by shape-aware constraints. i It is calculated based on the geometric deviation between keypoint coordinates extracted using soft argmax, reflecting each keypoint's contribution to overall structural consistency. This weighting mechanism enables Shape-aware BCE Loss to not only focus on the accuracy of local pixel values but also globally optimize the structural relationships between keypoints, thereby better ensuring the anatomical plausibility of the predicted heatmap. This formula unifies the dual optimization objectives of pixel-level and global structural level through weighted accumulation of all keypoints and images.
[0251] The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.
Claims
1. A method for tracking cardiac ultrasound key points based on pseudo labels, characterized in that: The method comprises: Extracting a key frame from a set of ultrasound images corresponding to a cardiac cycle, wherein the key frame includes at least one change node in the cardiac cycle, and ultrasound images other than the key frame in the set of ultrasound images are intermediate frames; For each key frame, extract the key points on the key frame as key pseudo labels; According to the key pseudo-label, combined with the time series relationship between the ultrasound images in the ultrasound image set, key point detection is performed on each of the intermediate frames based on a preset label propagation model, and the key points corresponding to the key pseudo-label on each of the intermediate frames are extracted as intermediate pseudo-labels; The label propagation model is optimized and trained according to the intermediate pseudo-label to obtain an optimized label propagation model; wherein the optimized label propagation model is used to track key points of the ultrasound video corresponding to the cardiac cycle.
2. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 1, characterized in that: For each key frame, extracting key points on the key frame as key pseudo labels includes: For each key frame, the key frame is input into a preset heat map prediction model, and the heat map prediction model extracts all key points corresponding to the key frame according to the preset geometric features corresponding to the category of the key frame; a boundary path of the geometric distribution of the key points is generated according to the preset geometric features, and the key points are sorted according to a preset boundary order. Each layer of the heat map is restricted to represent only one key point, and a predicted heat map corresponding to each key point is generated. The key point coordinates are extracted from the predicted heat map based on a preset key point extraction strategy.
3. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 2, characterized in that: The heat map prediction model is obtained by the following steps: According to the category of each ultrasound image corresponding to the cardiac cycle, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the cardiac cycle to obtain a heat map prediction model.
4. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 3, characterized in that: The initial UNet network framework is trained according to a plurality of ultrasound training images corresponding to the cardiac cycle to obtain a heat map prediction model, comprising: Acquire multiple multi-view ultrasound images corresponding to a cardiac cycle at multiple viewpoints, and train the initial UNet network framework based on the multi-view ultrasound images to obtain an intermediate heat map prediction model; A plurality of single-view ultrasound images corresponding to a cardiac cycle at a fixed view angle are acquired, and the intermediate heat map prediction model is trained based on the single-view ultrasound images to obtain a heat map prediction model.
5. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 1, characterized in that: For each key frame, extracting key points on the key frame as key pseudo labels includes: For each key frame, a key point heat map representing key points on the key frame is extracted as a key pseudo label based on a preset heat map regression algorithm; Furthermore, the step of performing key point detection on each intermediate frame based on the key pseudo-label and in combination with the time series relationship between the ultrasound images in the ultrasound image set based on a preset label propagation model, and extracting the key point corresponding to the key pseudo-label on each intermediate frame as the intermediate pseudo-label, includes: For key pseudo labels, based on densely connected convolutional networks, shallow features and deep features are fully integrated to obtain heat map features corresponding to key pseudo labels; Based on a preset inter-frame similarity calculation mechanism and in combination with a time series relationship between the ultrasound images in the ultrasound image set, a matching relationship between the intermediate frames is constructed; For each intermediate frame, based on the heat map features corresponding to the key pseudo labels and the matching relationship between the intermediate frame and other intermediate frames, the key points corresponding to the key pseudo labels on the intermediate frame are extracted as intermediate pseudo labels.
6. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 1, characterized in that: The step of extracting key frames from a set of ultrasound images corresponding to a cardiac cycle includes: Acquiring an ultrasound video corresponding to a cardiac cycle, identifying an initial window region in the ultrasound video whose brightness variation satisfies a preset brightness variation pattern, and modifying the initial window region based on morphological features corresponding to the cardiac cycle to obtain an ultrasound window corresponding to the ultrasound video; According to the ultrasound window corresponding to the ultrasound video, an ultrasound image set corresponding to the ultrasound video is extracted.
7. The method for implementing cardiac ultrasound key point tracking based on pseudo labels according to claim 3, characterized in that: The initial UNet network framework is trained according to a plurality of ultrasound training images corresponding to the cardiac cycle to obtain a heat map prediction model, comprising: Training the initial UNet network framework according to a plurality of ultrasound training images corresponding to a cardiac cycle, providing feedback on the training process based on a preset target loss function, and obtaining a heat map prediction model trained to convergence; The preset target loss function is calculated as follows for each key point: Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation; The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
8. A device for tracking cardiac ultrasound key points based on pseudo labels, characterized in that: The device comprises: a key frame extraction module, configured to extract key frames from a set of ultrasound images corresponding to a cardiac cycle, wherein the key frames include at least one change node in the cardiac cycle, and ultrasound images other than the key frames in the set of ultrasound images are intermediate frames; A key point detection module is used to extract the key points on each key frame as key pseudo labels; a label propagation module, configured to perform key point detection on each intermediate frame based on the key pseudo-label and a time series relationship between the ultrasound images in the ultrasound image set based on a preset label propagation model, and extract key points corresponding to the key pseudo-label on each intermediate frame as intermediate pseudo-labels; A model optimization module is used to optimize and train the label propagation model based on the intermediate pseudo-label to obtain an optimized label propagation model; wherein the optimized label propagation model is used to track key points of the ultrasound video corresponding to the cardiac beat cycle.
9. A system for tracking cardiac ultrasound key points based on pseudo labels, characterized in that: The system includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method for implementing cardiac ultrasound key point tracking based on pseudo labels as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the method for implementing cardiac ultrasound key point tracking based on pseudo labels according to any one of claims 1 to 7.
Citation Information
Cited By
Single sample learning path construction method for medical image key point detection
CN121616845A