Ultrasonic image key point detection method and device based on heat map regression, and medium
By combining the time series information of previous and next frames and view category embedding, the heat map regression method is optimized to solve the accuracy and stability problems of key point detection in low-quality ultrasound images and achieve more efficient key point extraction.
Patent Information
- Application Number
- CN202510742667.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-30
AI Technical Summary
In low-quality ultrasound images, existing heat map regression methods have difficulty in extracting key information stably and accurately, especially under the influence of factors such as noise, artifacts and blurred boundaries, the key point detection accuracy is insufficient.
By introducing the previous and next frames of the key frame as input, combining the contextual information of the time series, using the view category information to embed the network, adopting the improved non-maximum suppression and geometric feature sorting strategy, designing the shape-aware loss function, and optimizing the key point extraction process.
The accuracy and robustness of key point detection are improved, ensuring stable and accurate key point information extraction in ultrasound images, and enhancing the view perception and geometric consistency of the model.
Smart Images

Figure CN120725967A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ultrasonic image analysis, and in particular to a method, device, and medium for detecting key points in ultrasonic images based on heat map regression. Background Art
[0002] Automatic detection of ultrasound image keypoints is of great significance in ultrasound applications, particularly in medicine. It is a key step in medical image analysis and is widely used in areas such as medical image registration, parameter measurement, surgical navigation, and pathological tissue segmentation. The complexity and diversity of medical images make the definition and detection of keypoints challenging. Manual labeling is not only time-consuming, labor-intensive, and costly, but also susceptible to inter- and intra-observer variability. Significant research progress has been made in medical image keypoint detection over the past few decades, primarily categorized into traditional methods and deep learning-based approaches. For example, deep learning methods based on heatmap regression generate pseudo-probability heatmaps and locate keypoints using local or global extrema. While achieving higher accuracy than direct coordinate regression methods, they still face the challenges of computational complexity and high time consumption. In 2019, Wolterink et al. proposed using a three-dimensional convolutional neural network to predict the probability distribution of each position and determine the keypoint location via the local maximum of the probability distribution. In 2019, Payer et al. first used the local appearance component to generate candidate heat maps, and combined the spatial configuration component to model the spatial relationship between key points, and finally completed the prediction of key points by fusing the heat maps of the two through point-by-point multiplication. In 2021, Wang et al. proposed a framework based on multi-task learning for right ventricular end-diastolic and end-systolic frame recognition and anatomical landmark detection in echocardiography. In 2023, Schobs et al. proposed using PHD-Net to generate Gaussian heat maps and classify predictions through quantile binning, combining multiple uncertainty measurement methods to quantify the pseudo-probability distribution of landmark location. Also in 2023, Wan et al. proposed a method that enhances the generation ability of key point heat maps and effectively imposes shape constraints by introducing a soft-hard conversion mechanism of reference heat map information and multi-scale feature fusion. In 2024, Choi et al. further proposed a method that combines heat map gradient updates with shape model constraints, using a two-step iterative correction method to constrain the shape distribution of key points.
[0003] Existing technologies use heatmap regression for ultrasound keypoint detection, rather than direct coordinate regression. This is because heatmap regression exhibits significant advantages in spatial information preservation and model generalization. However, ultrasound image quality is poor, often affected by noise, artifacts, and blurred boundaries. Therefore, extracting stable and accurate key information from low-quality ultrasound images is crucial to improving detection accuracy. Summary of the Invention
[0004] The present invention provides a method, device and medium for detecting key points in ultrasound images based on heat map regression, which are used to extract stable and accurate key information from low-quality ultrasound images.
[0005] In order to solve the above technical problems, the first aspect of the present invention discloses a method for detecting key points in ultrasound images based on heat map regression, the method comprising: According to the temporal sequence of a plurality of ultrasound images corresponding to the target object, a preceding frame image and a succeeding frame image corresponding to each ultrasound image are determined, and each ultrasound image and its corresponding preceding frame image and succeeding frame image are used as an input data; Inputting all the input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points; Extracting key point coordinates from the predicted heat map based on a preset key point extraction strategy; The heat map prediction model is obtained by the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
[0006] As an optional embodiment, in the first aspect of the present invention, all the input data are input into the heat map prediction model to generate a plurality of prediction heat maps for representing the distribution probability of key points, including: Inputting all the input data into the heat map prediction model, the heat map prediction model screens out key points of each category according to the preset geometric features corresponding to the category of each ultrasound image; For each category of key points, a boundary path of the key point geometric distribution is generated according to the preset geometric features corresponding to the category, and the key points of the category are sorted according to the preset boundary order. Each layer of heat map is limited to represent only one key point, and a prediction heat map corresponding to each key point is generated; And, extracting key point coordinates from the predicted heat map based on a preset key point extraction strategy includes: For each of the predicted heat maps, the maximum value coordinates are extracted from the predicted heat map as the key point coordinates corresponding to the predicted heat map.
[0007] As an optional embodiment, in the first aspect of the present invention, for each category of key points, generating a boundary path of the key point geometric distribution according to a preset geometric feature corresponding to the category, and sorting the key points of the category according to a preset boundary order includes: For each category of key points, the convex hull boundary corresponding to the key points of this category is calculated, and the boundary path of the geometric distribution of the key points is generated based on the convex hull boundary and the preset geometric features corresponding to the category; a convex hull search is performed on the key points of this category according to the preset boundary order, the position of each key point is optimized based on the convex hull search results, and the key points of this category are sorted.
[0008] As an optional embodiment, in the first aspect of the present invention, the method further comprises: Acquiring multiple ultrasound videos corresponding to the target object, for each of the ultrasound videos, identifying an initial window region in the ultrasound video whose brightness change satisfies a preset brightness change pattern, and modifying the initial window region based on the morphological characteristics of the target object to obtain an ultrasound window corresponding to each ultrasound video; For each of the ultrasound videos, multiple ultrasound images corresponding to the ultrasound video are extracted according to the ultrasound window corresponding to the ultrasound video.
[0009] As an optional embodiment, in the first aspect of the present invention, the initial UNet network framework is trained according to the multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model, including: Training the initial UNet network framework according to a plurality of ultrasound training images corresponding to the target object, providing feedback on the training process based on a preset target loss function, and obtaining a heat map prediction model trained to convergence; The preset target loss function is calculated as follows for each key point: Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation; The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
[0010] As an optional embodiment, in the first aspect of the present invention, generating a category embedding vector according to the category of each ultrasound image corresponding to the target object, and embedding the category embedding vector into a preset initial UNet network framework includes: Generate a category embedding vector based on the category of each ultrasound image corresponding to the target object ,in is the dimension of the embedding vector; The embedding vector is based on the following formula The mapping embedding vector is obtained by mapping the channel dimension C of the feature map output by the convolutional layer of the initial UNet network framework to the fully connected layer of the initial UNet network framework. :
[0011] In the above formula, and are the weights and biases of the fully connected layer of the initial UNet network framework; Embed the mapping into a vector The vector is expanded to a dimension matching the feature map and feature fused with the feature map, thereby embedding the category embedding vector into a preset initial UNet network framework.
[0012] As an optional embodiment, in the first aspect of the present invention, determining the loss function corresponding to the key point according to the deviation between the real heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point includes: The loss function corresponding to the key point i is calculated according to the following formula :
[0013] In the above formula, y is the real heat map, is the prediction heat map, is the geometric weight, is the cross entropy loss corresponding to the key point i, which is calculated by the following formula :
[0014] In the above formula, N is the total number of heat maps.
[0015] A second aspect of the present invention discloses an ultrasonic image key point detection device based on heat map regression, the device comprising: An input data module is used to determine a preceding frame image and a succeeding frame image corresponding to each ultrasound image according to the temporal sequence of the ultrasound images corresponding to the target object, and to use each ultrasound image and its corresponding preceding frame image and succeeding frame image as an input data; A heat map generation module is used to input all the input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points; A key point extraction module, configured to extract key point coordinates from the predicted heat map based on a preset key point extraction strategy; The heat map prediction model is obtained by the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
[0016] A third aspect of the present invention discloses an ultrasound image key point detection system based on heat map regression, the system comprising: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the ultrasound image key point detection method based on heat map regression disclosed in the first aspect of the present invention.
[0017] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the ultrasound image key point detection method based on heat map regression disclosed in the first aspect of the present invention.
[0018] Compared with the existing technology, the present invention introduces the previous and next frames of the key frame as input, and combines the contextual information of the time series to help the model capture the dynamic feature changes between consecutive frames, thereby improving the accuracy and robustness of key point detection; embedding view category information into the network enables the model to explicitly consider the view type in the feature extraction process, enhancing the model's view perception ability; finally, based on the preset key point extraction strategy, the key point coordinates are extracted from the predicted heat map, thereby extracting stable and accurate key point information in the ultrasound image. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 1 is a flow chart of a method for detecting key points in ultrasound images based on heat map regression disclosed in an embodiment of the present invention; Figure 2 Schematic diagram of the convex hull search and matching strategy disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of a heat map of key points of the long-axis inflow tract of the right atrium and right ventricle disclosed in an embodiment of the present invention; Figure 4 Schematic diagram of key point distribution differences under the same view disclosed in an embodiment of the present invention; Figure 5 This is a schematic diagram of a heat map in which similar key points disclosed in an embodiment of the present invention share a common heat map; Figure 6 Schematic diagram of the candidate point spatial distance constraint method disclosed in an embodiment of the present invention; Figure 7 Schematic diagram of a heat map regression network architecture based on a graph convolutional neural network disclosed in an embodiment of the present invention; Figure 8 This is a schematic diagram of the lack of shape constraints for predicted image key points disclosed in an embodiment of the present invention; Figure 9 This is a structural diagram of an ultrasound image key point detection system based on heat map regression disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different items, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or end comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or end.
[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0024] Example 1 See also Figure 1 , Figure 1 This is a flow chart of a method for detecting key points in ultrasound images based on heatmap regression disclosed in an embodiment of the present invention. Figure 1The ultrasound image key point detection method based on heat map regression described above can be used in an ultrasound image key point detection device based on heat map regression. The ultrasound image key point detection device based on heat map regression can be integrated into a cloud server or a local server. Figure 1 As shown, the ultrasound image key point detection method based on heat map regression may include the following operations: Step 101: Determine the preceding frame image and the succeeding frame image corresponding to each ultrasound image according to the temporal sequence of a plurality of ultrasound images corresponding to the target object, and use each ultrasound image and its corresponding preceding frame image and succeeding frame image as an input data.
[0025] In the embodiment of the present invention, the target object is an object that needs to be analyzed by ultrasound, which can be a human organ, such as a heart, or any object that needs to be analyzed by ultrasound, such as an industrial device, a chemical medium, a mechanical structure, etc.
[0026] In the embodiment of the present invention, in ultrasound images, single-frame information is often insufficient to accurately locate key points due to blurred boundaries or noise interference. Especially in dynamic sequences, the static features of a single frame may not provide sufficient contextual support. To solve this problem, the frames before and after the key frame are introduced as input. By combining the contextual information of the time series, the model is helped to capture the dynamic feature changes between consecutive frames, thereby improving the accuracy and robustness of key point detection. Optionally, given the current key frame and its previous and next frames and , the input sequence is defined as: .in, and are the height and width of the image, Number of channels Step 102: Input all input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points.
[0027] In the embodiment of the present invention, a heat map is a method for characterizing the spatial position of key points through Gaussian distribution.
[0028] Ultrasound images contain multiple views, and the anatomical structures and key point distribution patterns in different views vary significantly. Relying solely on raw image features, the model may not fully understand the semantic differences between views, resulting in a decrease in the accuracy of key point detection. To address this problem, this embodiment embeds view category information into the network, enabling the model to explicitly consider view types during feature extraction. Therefore, in this embodiment of the present invention, the heat map prediction model is obtained through the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
[0029] Step 103: extract key point coordinates from the predicted heat map based on a preset key point extraction strategy.
[0030] In embodiments of the present invention, extracting keypoint coordinates from a predicted heatmap involves precisely extracting keypoint coordinates from a probabilistically distributed heatmap. Optional methods include non-maximum suppression (NMS), which uses a local window to find peaks and suppress non-peak responses to generate a sparse heatmap. The top K maximum values in the heatmap are then sorted and selected as candidate keypoint locations, where K is the number of keypoints represented in a single heatmap.
[0031] It can be seen that in the embodiment of the present invention, by introducing the previous and next frames of the key frame as input and combining the contextual information of the time series, the model is helped to capture the dynamic feature changes between consecutive frames, thereby improving the accuracy and robustness of key point detection; the view category information is embedded in the network, so that the model can explicitly consider the view type in the process of feature extraction, thereby enhancing the model's view perception ability; finally, based on the preset key point extraction strategy, the key point coordinates are extracted from the predicted heat map, thereby extracting stable and accurate key point information in the ultrasound image.
[0032] In the above solution of this embodiment, it is proposed that key points can be extracted from the heat map by using the non-maximum suppression method. This embodiment further explores it as follows: Extracting keypoints from heatmaps using non-maximum suppression is essentially a combination of non-maximum suppression and Top-K maximum value filtering. Non-maximum suppression suppresses non-peak responses in the heatmap, retaining only the local maximum values to reduce noise. Subsequently, the rough coordinates of candidate keypoints are obtained by extracting the top K maximum values from the sparse heatmap. This method can efficiently select multiple keypoints from similar feature points.
[0033] However, for this embodiment, the heat map regression method is highly dependent on hyperparameters. Improper setting of the window size of non-maximum suppression may lead to missed detection or false detection, especially when the distance between adjacent key points is close. Non-maximum suppression tends to merge adjacent peaks into one and cannot correctly distinguish multiple key points. In addition, non-maximum suppression assumes a regular peak distribution, and the processing effect is poor for key point distributions with complex shapes or weak target points.
[0034] To solve the above technical problem, in an optional embodiment, all input data is input into the heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points, which may include: All input data are fed into the heat map prediction model, which then filters out key points for each category based on the preset geometric features corresponding to each category of ultrasound images. For each category of key points, a boundary path of the key point geometric distribution is generated according to the preset geometric features corresponding to the category, and the key points of the category are sorted according to the preset boundary order. Each layer of heat map is limited to represent only one key point, and a prediction heat map corresponding to each key point is generated; And, extract key point coordinates from the predicted heat map based on the preset key point extraction strategy, including: For each predicted heat map, the maximum value coordinates are extracted from the predicted heat map as the key point coordinates corresponding to the predicted heat map.
[0035] In this optional embodiment, an improved strategy is proposed to address the limitations of non-maximum suppression, which sorts and matches key points of the same type through geometric features to ensure that each type of key points has a fixed sequential relationship. For example, for cardiac ultrasound, based on the characteristics of cardiac ultrasound medical key points, the geometric distribution characteristics of key points are very clear, and the key points are almost sequentially distributed on the anatomical structures of the endocardium and epicardium. The endocardium and epicardium have regular annular or linear distribution patterns in anatomy, and this characteristic provides a natural basis for the sorting and matching of key points. This solution directly limits each layer of heat map to represent only one key point in the heat map generation stage, avoiding the repeated appearance of similar points in the heat map, and thus eliminating the need to rely on complex non-maximum suppression operations. In the post-processing stage, only the maximum coordinates need to be extracted from each heat map, simplifying the key point selection process.
[0036] In this optional embodiment, further optionally, for each category of key points, generating a boundary path of the key point geometric distribution according to a preset geometric feature corresponding to the category, and sorting the key points of the category according to a preset boundary order includes: For each category of key points, the convex hull boundary corresponding to the key points of this category is calculated, and the boundary path of the geometric distribution of the key points is generated based on the convex hull boundary and the preset geometric features corresponding to the category; a convex hull search is performed on the key points of this category according to the preset boundary order, the position of each key point is optimized based on the convex hull search results, and the key points of this category are sorted.
[0037] This optional embodiment, taking cardiac ultrasound as an example, may specifically include: First, the image is corrected through key points, and the anatomical structure is adjusted to a unified reference coordinate system to ensure that the anatomical relationship in the image is consistent with the actual structure. Secondly, on the corrected image, the candidate key points are screened and sorted based on the key point matching strategy of convex hull search: First, the candidate key points are extracted and their convex hull boundaries are calculated to generate a boundary path reflecting the geometric distribution; then, the key points are sorted according to the order of the convex hull boundaries. For annular or linear anatomical structures, they are arranged in counterclockwise, clockwise, or from base to apex, respectively. The specific matching strategy is as follows: Figure 2 As shown in Figure 3, the positions of key points are optimized using a convex hull search algorithm to ensure that the key point distribution conforms to the spatial constraints of the anatomical structure.
[0038] In yet another optional embodiment, the method further includes: Acquire multiple ultrasound videos corresponding to the target object, identify, for each ultrasound video, an initial window region in which brightness changes satisfy a preset brightness change pattern, and modify the initial window region based on morphological features of the target object to obtain an ultrasound window corresponding to each ultrasound video; For each ultrasound video, multiple ultrasound images corresponding to the ultrasound video are extracted according to the ultrasound window corresponding to the ultrasound video.
[0039] In an embodiment of the present invention, ultrasound window extraction is a key step in ultrasound image analysis. Taking cardiac ultrasound as an example, it can improve the efficiency and accuracy of subsequent analysis by extracting the area containing the cardiac anatomical structure from the image or video and removing background noise. In this embodiment, by processing multiple frames in the ultrasound video, the area containing the target object structure is identified and extracted. Specifically, first, by performing brightness area detection on the video frame, the part with significant brightness changes is identified, which usually corresponds to the target object area. Optionally, it can also include: using morphological operations to further optimize the extracted area, remove noise and fill holes. Finally, combining multiple frame information, a comprehensive mask is generated to ensure that a stable and accurate ultrasound window is extracted.
[0040] The embodiments of the present invention further discovered that traditional loss functions such as L2 Loss, BCE Loss, Wing Loss, and Adaptive Wing Loss are usually used to measure the difference between predicted key points and labeled key points. These loss functions mainly focus on the error between single points, but may not be sufficient to effectively capture shape information in scenarios where there is a certain geometric relationship between key points. Therefore, based on the specific requirements of the task, the embodiments of the present invention design a shape-aware loss function for heatmap regression to introduce geometric constraints between key points in the optimization process, thereby generating a more structured predicted heatmap.
[0041] Therefore, in another optional embodiment, the aforementioned heat map prediction model is obtained by training the initial UNet network framework based on the multiple ultrasound training images corresponding to the target object, including: The initial UNet network framework is trained based on multiple ultrasound training images corresponding to the target object, and the training process is fed back based on a preset target loss function to obtain a heat map prediction model that has been trained to convergence; Among them, the preset target loss function is calculated as follows for each key point: Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation; The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
[0042] In this optional embodiment, further optionally, determining the loss function corresponding to the key point according to the deviation between the real heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point may include: The key point is calculated according to the following formula i The corresponding loss function :
[0043] In the above formula, y is the real heat map, is the prediction heat map, is the geometric weight, For this key point i The corresponding cross entropy loss is calculated by the following formula :
[0044] In the above formula, N is the total number of heat maps.
[0045] In this optional embodiment, Represents the cross entropy loss calculated for each pixel position, which is used to evaluate the degree of match between the predicted heat map and the true heat map at the pixel level, and the additional geometric weight term The importance of geometric relationships between key points can be quantified through shape-aware constraints. The loss function is calculated based on the geometric deviation between the extracted keypoint coordinates, reflecting each keypoint's contribution to overall structural consistency. This weighted mechanism allows the loss function to not only focus on the accuracy of local pixel values but also globally optimize the structural relationships between keypoints, thereby better ensuring the anatomical plausibility of the predicted heatmap. This formula unifies the dual optimization objectives of pixel-level and global structural level through weighted accumulation of all keypoints and images.
[0046] In yet another optional embodiment, the above-mentioned step of generating a category embedding vector according to the category of each ultrasound image corresponding to the target object and embedding the category embedding vector into a preset initial UNet network framework may include: Generate a category embedding vector based on the category of each ultrasound image corresponding to the target object ,in is the dimension of the embedding vector; The embedding vector is based on the following formula The mapping embedding vector is obtained by mapping the channel dimension C of the feature map output by the convolution layer of the initial UNet network framework to the fully connected layer of the initial UNet network framework. :
[0047] In the above formula, and are the weights and biases of the fully connected layer of the initial UNet network framework; Embed the map into a vector The class embedding vector is expanded to match the dimension of the feature map and fused with the feature map, thus embedding it into the preset initial UNet network framework. This design explicitly embeds view information into the feature extraction process, enhancing the model's view perception capabilities.
[0048] Example 2 Cardiac ultrasound technology is a non-invasive and highly effective medical imaging method. It can accurately examine the structure and function of the heart and great blood vessels, particularly in assessing blood flow changes and cardiac function. The entire cardiac ultrasound analysis process includes the following steps: First, ultrasound video containing at least one cardiac cycle is acquired using cardiac ultrasound equipment. Next, the video is classified into standard cardiac ultrasound sections through view classification, and the end-diastolic (ED) and end-systolic (ES) frames of the cardiac cycle are identified. Key point detection is then performed, and key point tracking is used to analyze their dynamic changes. This data is then used to perform three-dimensional reconstruction of the heart and calculate medical functional indicators (such as ejection fraction and cardiac chamber volume), ultimately generating analysis results.
[0049] Heatmap regression represents the locations of target keypoints as a two-dimensional Gaussian heatmap rather than directly predicting their two-dimensional coordinates. Compared to direct coordinate regression, heatmap regression transforms the high-dimensional regression problem into a pixel prediction problem in image space, reducing the complexity of nonlinear regression. The distributed prediction also demonstrates greater robustness to occlusion and noise.
[0050] This example uses heatmap regression for ultrasound keypoint detection, rather than direct coordinate regression. This is because heatmap regression exhibits significant advantages in spatial information preservation and model generalization, particularly in cardiac ultrasound images, where keypoints often exhibit significant anatomical structural variation and image noise. Furthermore, coordinate regression struggles to handle the interdependencies between multiple keypoints, making inconsistent predictions for multiple keypoints a problem.
[0051] Heatmap is a method to characterize the spatial location of key points through Gaussian distribution. For a given key point coordinate , and its corresponding heat map Defined as The Gaussian function centered at . The formula is:
[0052] in Indicates the standard deviation of the Gaussian distribution and controls the diffusion range of the heat map. The heat map value is at the key point The probability reaches its peak at , and gradually decays as the distance increases. In this way, the heat map generates a probability distribution for each key point in the image, which can effectively capture local context information and spatial relationships. Figure 3 is an example of a heatmap representation of cardiac ultrasound key points.
[0053] Heatmap post-processing aims to accurately extract keypoint coordinates from the probabilistically distributed heatmap to remove noise and improve localization accuracy. Conventional methods include non-maximum suppression (NMS), which uses a local window to find peaks and suppress non-peak responses to generate a sparse heatmap. The top K maximum values in the heatmap are then sorted and selected as candidate keypoint locations, where K is the number of keypoints represented in a single heatmap.
[0054] However, this example found that existing cardiac ultrasound keypoint detection technologies face the following major challenges: First, existing methods lack optimization of the geometric constraints of cardiac keypoints. The cardiac anatomical structure has clear geometric distribution characteristics, such as the circular or linear distribution of the endothelium and epicardium. However, traditional methods process each keypoint independently, ignoring the geometric relationships between keypoints. This results in insufficient structural consistency and anatomical rationality in the prediction results. Specifically, multiple keypoints of the heart are interrelated. Ignoring these relationships can lead to errors in keypoint prediction and structural irrationality, thus affecting the accuracy of quantitative analysis of cardiac function.
[0055] Secondly, ultrasound image quality is poor, often affected by noise, artifacts, and blurred boundaries. A single frame cannot accurately reflect the anatomical features of key points. Even experienced physicians rely on the classification and time-series deformation information from different views to assist in judgment and compensate for the incompleteness and low quality of single-frame information. Therefore, extracting stable and accurate key information from low-quality ultrasound images is crucial to improving detection accuracy.
[0056] In addition, traditional heatmap regression methods rely on non-maximum suppression combined with Top-K screening to regress multiple key points, but their performance is highly sensitive to the NMS window size, which may lead to missed detections or false detections, especially when adjacent key points are close, and multiple peaks are easily mixed into one.
[0057] In response to these problems, this embodiment proposes a series of optimization schemes. The main inventive concept is: through the geometric redefinition of key points and the design of shape-aware loss functions, the geometric consistency and anatomical rationality of the key point detection results are enhanced; the view embedding module is introduced to explicitly embed view category information into the network to improve the model's adaptability to multi-view data; combined with the previous and next frame inputs of the key frame, the continuity of the time series is used to capture dynamic feature changes, and improve the prediction instability caused by insufficient single-frame information. In addition, in the heat map regression method, it is proposed to optimize key points through spatial distance constraints of candidate points, graph convolutional neural networks, and geometric sorting strategies to optimize the key point selection process, and the reliability of the method is verified through experiments.
[0058] In the embodiment of the present invention, original ultrasound data is first acquired, wherein the original ultrasound data may be a series of ultrasound videos or continuous ultrasound images.
[0059] Optionally, this embodiment can use an ultrasound window to process the raw ultrasound data. This process extracts regions containing cardiac anatomical structures from images or videos and removes background noise, thereby improving the efficiency and accuracy of subsequent analysis. In this experiment, multiple frames in an ultrasound video are processed to identify and extract regions containing cardiac structures. Specifically, first, by performing brightness region detection on the video frames, portions with significant brightness changes are identified, which typically correspond to cardiac regions. Then, morphological operations are used to further optimize the extracted regions, remove noise, and fill in holes. Finally, a comprehensive mask is generated by combining information from multiple frames to ensure the extraction of a stable and accurate ultrasound window.
[0060] Optionally, this embodiment can also pre-process the ultrasound data first, the main goal of which is to convert the input cardiac ultrasound keyframes into a tensor of uniform specifications to facilitate model processing. Optionally, data enhancement is performed on the input cardiac ultrasound keyframes, for example, random rotation, translation, scaling and flipping can be included to simulate diverse deformations and improve the generalization ability of the model. Then, based on the balance between accuracy and performance overhead, all images are adjusted to a uniform tensor. resolution, assuming the original image size is , scale the image to Finally, the resized image is converted into a tensor and normalized to standardize the data.
[0061] Optionally, data preprocessing in this embodiment can also include data cleaning. This includes verifying the video file path, checking for missing or erroneous dot data, filtering out records that do not meet view and chamber requirements, ensuring the completeness of ED / ES annotations, and verifying the format and location of the annotation data. Furthermore, the cleaning process includes validity checks on image frames and cropped regions to ensure that the data meets training requirements. These cleaning steps provide high-quality input data for subsequent model training, ensuring the reliability of experimental results.
[0062] Further optional data enhancement may include: The role of data augmentation in ultrasound image analysis is to artificially expand the dataset by performing multiple transformations on the original data. This helps improve the generalization ability of the machine learning model, reduce overfitting, and enhance the robustness of the model. In this embodiment, random rotation, affine transformation, and color jittering are used to simulate changes in real-world conditions, such as ultrasound equipment, ultrasound acquisition techniques, imaging artifacts, and differences in image quality. Specifically, the data augmentation method includes random affine transformation, changing the rotation, translation, and scaling of the image, and applying color jittering to simulate changes in illumination and contrast in ultrasound images, such as the data augmentation methods shown in the table below.
[0063]
[0064] In this embodiment of the present invention, a heatmap regression-based ultrasound imaging keypoint detection method uses UNet as the backbone network framework. This combines the powerful multi-scale feature extraction capabilities of the ResNet50 encoder with the UNet's skip connection mechanism to effectively integrate high-level semantic features with spatial details in keypoint detection tasks. Simultaneously, a bridging module further compresses and fuses global features, while the decoder gradually upsamples to restore resolution, generating accurate keypoint heatmaps.
[0065] The improved network framework of the embodiment of the present invention combines view embedding and time series information, and realizes accurate key point detection in multi-view and dynamic conditions by introducing dynamic features of previous and next frames and residual enhancement modules. The encoder can use ResNet50 for deep feature extraction, the middle layer integrates global context information, and the decoder combines deconvolution with multi-resolution jump connections to gradually restore spatial resolution and retain detail features. The view embedding module enhances the adaptability of the network to different anatomical perspectives, while the capture of the temporal characteristics of previous and next frames improves the robustness in dynamic scenes and the detection accuracy of blurred boundary areas, ultimately generating a high-resolution heat map, which provides an efficient and accurate solution for dynamic key point detection and quantitative analysis of cardiac ultrasound images.
[0066] There are multiple views in ultrasound images, and the distribution of key points in different views is obviously different. Even for the same key point, its position and direction will change in different views, such as Figure 4 shown. Figure 4 Figure 2 shows ultrasound images from the same view, with keypoints marked with blue dots. Despite the consistent view, keypoint distribution varies across frames. This is likely due to imaging noise, blurred boundaries, and dynamic changes in cardiac structure. Furthermore, accurate keypoint localization is difficult in single-frame ultrasound images with blurred boundaries or noise, especially in dynamic ultrasound sequences. The lack of information in a single frame further exacerbates prediction instability.
[0067] To address these issues, the key inventive concept of this embodiment lies in introducing a view residual embedding module to embed view category information into the model, helping the model learn the semantic differences between views and improving the accuracy of keypoint localization. Furthermore, by adding the frames before and after the keyframe as contextual information, the model's understanding of dynamic ultrasound sequences is enhanced by leveraging the continuity of the time series.
[0068] In an optional embodiment, the view embedding module is designed as follows: Ultrasound images contain multiple views, and the anatomical structures and keypoint distribution patterns vary significantly across views. Relying solely on raw image features, the model may not fully understand the semantic differences between views, resulting in reduced keypoint detection accuracy. To address this issue, this optional embodiment introduces a view embedding module. By embedding view category information into the network, the model explicitly considers view type during feature extraction.
[0069] Optionally, views can be embedded into a module by passing the view class Convert to embedding vector ,in The dimension of the embedding vector is used to generate a high-dimensional representation of the view category. Subsequently, the embedding vector is mapped to the channel dimension of the feature map through a fully connected layer. , the transformation formula is ,in and are the weights and biases of the fully connected layer respectively. For the feature map output by the convolutional layer , expand the embedding vector to a shape that matches the feature map, and complete feature fusion by pixel-by-pixel addition. The formula is ,in ,None, None Finally, the fused features can be The activation function generates the final output This design explicitly embeds view information into the feature extraction process, enhancing the model’s view-awareness capability.
[0070] In another optional embodiment, in ultrasound images, a single frame of information is often insufficient to accurately locate keypoints due to blurred boundaries or noise interference. This is especially true in dynamic sequences, where the static features of a single frame may not provide sufficient contextual support. To address this issue, the frames before and after the keyframe are introduced as input. By incorporating the contextual information of the time series, the model can capture dynamic feature changes between consecutive frames, thereby improving the accuracy and robustness of keypoint detection.
[0071] Specifically, given the current keyframe and its previous and next frames and , the input sequence is defined as: .in, and are the height and width of the image, is the number of channels.
[0072] In another optional embodiment, the specific process of optimizing the heat map regression method is discussed as follows: In the key point detection task, due to the lack of sequentiality in the annotation data of similar key points, conventional methods usually use a single heat map to represent the positions of all similar key points. Figure 5 As shown, four similar lv_wall key points rely on a heat map representation.
[0073] Typically, multiple keypoints are regressed from a single heatmap through non-maximum suppression combined with Top-K filtering. However, the performance of this method is highly dependent on the hyperparameter selection of the non-maximum suppression region radius. If the radius is set too small, the response areas between keypoints may overlap excessively, making it difficult to distinguish keypoints. If the radius is set too large, multiple peaks may be blended into a single one, affecting the precise localization of keypoints.
[0074] In order to further improve the accuracy and robustness of key point regression, this optional embodiment discloses multiple improvement strategies: a method based on spatial distance constraints of candidate points, a heat map regression based on graph convolutional neural networks, and a method based on geometric meaning to redefine key points.
[0075] In a first optional embodiment, a key point regression method based on a candidate point spatial distance constraint method is disclosed: Figure 6 As shown in the figure, the candidate points are first filtered out using a clustering algorithm (DBSCAN) to eliminate outliers and abnormally distributed noise points. Then, the filtered points are sorted by confidence and spatial distance constraints are added. This method ensures that high-confidence key points are evenly distributed in space and reduces noise interference.
[0076] In a second optional embodiment, a heatmap regression method based on a graph convolutional neural network is disclosed: By using graph convolutional neural networks to dynamically learn key point distribution and selection strategies, we overcome the shortcomings of traditional non-maximum suppression in fixed window hyperparameters, neighboring point differentiation, and global information modeling.
[0077] The network takes a multi-keypoint heatmap as input and first extracts local features through a convolutional module, thereby enhancing the understanding of the area near the keypoints. Subsequently, an attention mechanism is introduced to model the distribution of keypoints from a global perspective, capturing long-range dependencies and generating global semantic features. A linear mapping module is used to reduce the dimensionality of high-dimensional features. The network further extracts a compact feature representation, providing efficient input for subsequent optimization steps.
[0078] The reduced features are fed into a multi-layer graph convolutional network. This module models the spatial distribution and topological relationships between key points through layer-by-layer processing, gradually optimizing the key point prediction results. Ultimately, the network generates accurate key point coordinates and completes the position mapping by aligning them with the original image. The overall network combines the advantages of convolution, attention, and graph convolution, such as Figure 7As shown, the entire process can be optimized from rough heat maps to precise coordinates.
[0079] In a third optional embodiment, a method for redefining key points based on geometric meaning is disclosed: Keypoints of the same type are sorted and matched using geometric features, ensuring a fixed order for each type of keypoint. During the heatmap generation phase, this solution directly restricts each heatmap layer to represent only one keypoint, preventing the recurrence of similar points within the heatmap and eliminating the need for complex non-maximum suppression operations. In post-processing, only the maximum coordinates need to be extracted from each heatmap, simplifying the keypoint selection process.
[0080] Based on the characteristics of key points in cardiac ultrasound medicine, the geometric distribution characteristics of key points are very clear. Key points are distributed almost sequentially on the anatomical structures of the endocardium and epicardium. The endocardium and epicardium have regular circular or linear distribution patterns in anatomy, which provides a natural basis for sorting and matching key points.
[0081] First, the image is corrected through key points, and the anatomical structure is adjusted to a unified reference coordinate system to ensure that the anatomical relationship in the image is consistent with the actual structure. Secondly, on the corrected image, the candidate key points are screened and sorted based on the key point matching strategy of convex hull search: First, the candidate key points are extracted and their convex hull boundaries are calculated to generate a boundary path reflecting the geometric distribution; then, the key points are sorted according to the order of the convex hull boundaries. For annular or linear anatomical structures, they are arranged in counterclockwise, clockwise, or from base to apex, respectively. The specific matching strategy is as follows: Figure 2 As shown in Figure 3, the positions of key points are optimized using a convex hull search algorithm to ensure that the key point distribution conforms to the spatial constraints of the anatomical structure.
[0082] In the heat map generation stage, this method directly limits each layer of heat map to represent only one key point, thus avoiding the repeated distribution of similar points in the heat map. This improvement makes the post-processing stage no longer dependent on the complex non-maximum suppression algorithm, and only needs to extract the maximum value coordinates in each heat map to determine the key point location. :
[0083] in, Represents a pixel in the heat map Confidence value of .
[0084] In yet another optional embodiment, the loss function optimization scheme is discussed as follows: In the task of detecting keypoints in cardiac ultrasound, traditional loss functions such as L2 Loss, BCE Loss, Wing Loss, and Adaptive Wing Loss are typically used to measure the difference between predicted keypoints and annotated keypoints. These loss functions primarily focus on the error between single points, but may not be sufficient to effectively capture shape information in scenarios where keypoints have certain geometric relationships. Therefore, based on the specific requirements of the task, this optional embodiment designs a shape-aware loss function for heatmap regression to introduce geometric constraints between keypoints during the optimization process, thereby generating a more structured predicted heatmap.
[0085] Real heat map Represents the spatial distribution of each key point in the image, usually in the form of a Gaussian distribution to represent the probability distribution of the key point position, where the center point is the true coordinate of the key point and the value of the surrounding area gradually decays; predict heat map It is the result generated by the deep learning model after extracting the features of the input image, which represents the model's probability estimate of the key point location. With prediction heatmap The difference between them is used to guide model training, and the error between them is minimized by optimizing the loss function, thereby improving the accuracy of key point positioning and the generalization ability of the model.
[0086] L2 Loss is a classic regression loss function used to measure the predicted heatmap and the real heat map It is simple and easy to use and suitable for scenarios with uniform error distribution, but it is sensitive to outliers because the square of the error will amplify the impact of the error point. In the task of detecting key points in cardiac ultrasound, L2 Loss can optimize the similarity of the heat map as a whole, but lacks the focus on the key point area. The formula is as follows, where is the total number of heatmaps:
[0087] BCE Loss is used to measure the probability distribution of each pixel in the predicted heat map belonging to a key point and the true probability distribution It is suitable for processing situations with sparse key points and a large background, and effectively optimizes pixel classification by calculating probability differences pixel by pixel. However, BCE Loss ignores the geometric relationship between key points and cannot capture the global structure. The formula is as follows:
[0088] Wing Loss is a keypoint localization loss function that performs logarithmic scaling on small errors to enhance optimization while applying linear processing to large errors to reduce the impact of outliers. In keypoint heatmap regression tasks, Wing Loss achieves fine-grained optimization of small errors by reducing the impact of large errors, making it suitable for scenarios requiring high-precision localization. The specific formula is as follows: ω controls the sensitivity of the small error region, ϵ determines the smoothness of the logarithmic function, and C ensures the continuity of the loss.
[0089]
[0090] Adaptive Wing Loss is an improved version of Wing Loss. It makes the model more sensitive to errors in the area near the key points by dynamically adjusting the weights. Compared with Wing Loss, Wing Loss pays more attention to the area near the key points and dynamically adjusts the weights of different errors. It performs better in key point detection tasks with complex backgrounds or sparse distribution. The specific formula is as follows, where Controls the degree of change in dynamic weighting, is the heatmap pixel value, which is used to adjust the importance of key points.
[0091]
[0092] In the task of cardiac ultrasound key point detection, traditional loss functions such as L2 Loss, BCE Loss and Wing Loss mainly focus on the pixel-level differences between predicted points and true points. However, they ignore the geometric relationship and global structural information between key points, such as Figure 8 In order to better capture the anatomical structural features of key points, a shape-aware loss function is designed to incorporate structural information into the loss optimization process by introducing geometric relationship constraints between key points.
[0093] In heatmap regression, soft argmax is a differentiable method for extracting the coordinates of predicted key points. By using soft argmax, the location of the maximum value in the heatmap can be smoothly estimated, overcoming the non-differentiable problem of the traditional argmax method. The specific formula is as follows:
[0094] in, is a temperature parameter that controls the degree of smoothness. When it is close to 0, the result of soft argmax is close to traditional argmax, that is, almost only the index of the largest element contributes to the result.
[0095] Shape-aware loss uses soft argmax to extract the keypoint coordinates between the predicted and ground-truth heatmaps and calculates the geometric relationship deviation between the keypoints. By introducing this geometric constraint, shape-aware loss not only optimizes the localization error of individual keypoints but also enhances the consistency of the structural relationships between keypoints, thereby generating a predicted heatmap that is more consistent with anatomical features. The specific formula is as follows: , soft
[0096] in, Indicates the In the image The true heat map value of the key points, is the corresponding predicted heat map value, and the two are soft The extracted coordinates are used to calculate the geometric distance Dist.
[0097] Traditional loss functions such as BCE Loss mainly target pixel-level predictions, but ignore the geometric relationship between key points. In order to simultaneously optimize the pixel-level accuracy of the heat map and the global geometric relationship between key points, this optional embodiment proposes a new loss function - Shape-aware Binary Cross Entropy Loss (Shape-aware BCE Loss). This loss function combines BCE Loss and Shape-aware Loss. Among them, BCE Loss is used to measure the difference between the predicted probability and the true value of each pixel, while the shape-aware part extracts the coordinates of the predicted key points and the true key points through soft argmax, and calculates geometric constraints to optimize the structural relationship between the key points. The new loss function is defined as:
[0098] in, Represents the cross entropy loss calculated for each pixel position, which is used to evaluate the degree of match between the predicted heat map and the true heat map at the pixel level, and the additional geometric weight term The importance of geometric relationships between key points is quantified through shape-aware constraints. It is calculated based on the geometric deviation between keypoint coordinates extracted using soft argmax, reflecting each keypoint's contribution to overall structural consistency. This weighting mechanism enables Shape-aware BCE Loss to not only focus on the accuracy of local pixel values but also globally optimize the structural relationships between keypoints, thereby better ensuring the anatomical plausibility of the predicted heatmap. This formula unifies the dual optimization objectives of pixel-level and global structural level through weighted accumulation of all keypoints and images.
[0099] To verify the effect of this embodiment, this embodiment further discloses an experimental verification process, which is as follows: (1) Verification of the effect of network architecture optimization This embodiment designs a set of comparative experiments to verify the performance of the improved ResUNet network architecture proposed in this embodiment in the task of cardiac ultrasound key point detection. This embodiment selects multiple classic medical image key point detection models and conducts comparative experiments with the network model proposed in this embodiment. During the experimental verification process of this embodiment, all models are trained and verified based on the preset cardiac ultrasound key point detection dataset, the heat map regression method adopts the geometric meaning-based redefinition of key points in this embodiment, and the loss function adopts the one proposed in this embodiment. , Adam is selected as the optimizer, the initial learning rate is set to 0.001, and the batch_size is set to 32. The results of the comparative experiment are shown in the following table.
[0100]
[0101] Experimental results show that the improved ResUNet performs best in the task of cardiac ultrasound keypoint detection, with its detection success rate and average point-to-point error significantly outperforming other comparison networks. The following is an analysis of the reasons for its superior performance and the reasons why other networks performed poorly: (1) SCNNet combines the local fine features of key points with global geometric layout information to improve the model's ability to learn the relative positional relationships of key points, decomposing the complex positioning task into two simpler subtasks: the local component focuses on high-precision but potentially ambiguous candidate point predictions; the spatial component eliminates ambiguity through global anatomical structure constraints. The poor key point detection results of SCNNet are mainly due to the performance limitations brought about by its lightweight design. Specifically, SCNNet uses a fixed 64 filters and 7×7 convolution kernels, combined with simple AvgPool2d downsampling and Upsample upsampling. Although this design significantly reduces the computational complexity and the number of parameters, it results in insufficient multi-scale feature capture capabilities. In addition, the add mode used by the decoder has limited restoration of high-resolution features, and the three-fold downsampling further leads to the loss of detail information, resulting in insufficient modeling capabilities for complex key point areas.
[0102] (2) Stacked Hourglass uses eight stacked hourglass modules. Although it has strong feature extraction capabilities, multiple up- and down-sampling results in the loss of high-resolution features. The decoding stage lacks cross-level feature fusion strategies such as skip connections, resulting in insufficient modeling of global and local information. In addition, its complex network structure requires high computing resources, takes a long time to train, and has low optimization efficiency.
[0103] (3) HRNet performs well in high-resolution feature extraction, but the decoding stage only outputs prediction results through simple convolution and final_layer, lacking a multi-level decoding module and cross-layer feature fusion. Although the multi-branch structure provides multi-scale features, it is not combined with a task-customized module. Therefore, it is not as effective as the improved ResUNet in detecting fine targets such as heart valves.
[0104] (4) FARNet’s decoder achieves multi-scale feature fusion through hand-designed modules, but the modules are complex and difficult to train and optimize. In addition, its upsampling operation is not fully integrated with the convolution operation, resulting in inaccurate feature restoration, especially in the restoration of key areas.
[0105] In order to verify the effectiveness of the improved ResUNet in the key point detection task and the contribution of its various modules, this embodiment designed a series of ablation experiments, gradually introduced the core improved modules and compared the performance. The baseline model is the standard ResUNet, whose structure includes a ResNet-50 encoder, skip connections, and a layer-by-layer upsampling decoder, and the input is a single-frame ultrasound image. Based on this, this embodiment sequentially adds view embedding modules and time series context information, aiming to analyze the contribution of each module to the adaptability of multi-view differences and the stability of dynamic sequence prediction. Finally, the complete improved ResUNet combines view embedding and time series features to form an improved network. The results of the comparative experiment are shown in the following table.
[0106]
[0107] After introducing the view embedding module, the model's keypoint detection success rate increased by approximately 1.5%. This improvement is primarily due to the module embedding view category information into the feature extraction process, enabling the model to explicitly model anatomical differences under multi-view conditions. Through high-dimensional embedding and pixel-by-pixel fusion mechanisms, the model's performance in learning feature distribution patterns is significantly enhanced, effectively improving the perception of complex anatomical structures and the robustness of keypoint localization.
[0108] Introducing time series contextual information improves the model's keypoint detection success rate by approximately 2.5%. By incorporating dynamic features from frames before and after the keyframe, the model can compensate for the limitations of single-frame information due to blurred boundaries and noise interference, significantly improving its ability to capture dynamic changes between consecutive frames. This context-enhanced design improves the stability and accuracy of the model's predictions, making it more robust in dynamic ultrasound scenarios. The fully improved ResUNet combines these two modules, ultimately achieving significant overall performance improvements.
[0109] (2) Verification of the optimization effect of heat map regression method This embodiment designs a set of comparative experiments to evaluate the performance of non-maximum suppression combined with Top-K method, method based on candidate point spatial distance constraint, heat map regression method based on graph convolutional neural network, and method based on geometric redefinition of key points. In this embodiment, all models are trained and verified based on the cardiac ultrasound key point detection dataset introduced in this embodiment. The network architecture adopts the modified ResUNet network of this embodiment, and the loss function adopts the one proposed in this embodiment. , Adam is selected as the optimizer, the initial learning rate is set to 0.001, and the batch_size is set to 32. The results of the comparative experiment are shown in the following table.
[0110]
[0111] Experimental results show that the method based on geometrically redefined keypoints far surpasses other methods in keypoint detection accuracy and robustness, demonstrating particularly superior performance in regularly distributed cardiac ultrasound images. While other methods, such as non-maximum suppression combined with Top-K and methods based on spatial distance constraints for candidate points, perform similarly, all suffer from hyperparameter sensitivity and inability to handle complex distributions.
[0112] (1) Non-maximum suppression combined with the Top-K method extracts key points from heat maps through maximum pooling and local peak screening. This method is highly efficient, but relies on the hyperparameter setting of the non-maximum suppression window size. When adjacent key points are close together, missed detection or false detection is prone to occur. Especially when the key point distribution is complex or is affected by noise, the positioning accuracy and robustness of this method are relatively average.
[0113] (2) The method based on candidate point spatial distance constraint significantly reduces missed detections and incorrect candidate points through DBSCAN clustering and distance constraint, but its detection success rate is slightly lower than that of the non-maximum suppression method. This shows that spatial distance constraint improves the consistency of key point positioning to a certain extent, but the overall performance improvement is limited when dealing with complex distributions and weak target points.
[0114] (3) The detection success rate of the graph convolutional neural network method is 59.70%. Although the graph convolutional neural network can theoretically improve the adaptability to complex key point distributions by dynamically learning candidate point distribution and relationship modeling, in actual results, the high computational complexity may limit the generalization performance of the model.
[0115] (4) The detection success rate of the method based on geometric redefinition of key points is much better than other methods. This significant advantage is attributed to the fact that it limits each layer of heat map to only represent one key point from the heat map generation stage. Through convex hull search and sorting matching strategy, it avoids the reliance on non-maximum suppression and simplifies the post-processing steps. This method fully utilizes the geometric distribution characteristics of key points in the anatomical structure, ensuring the consistency of the key point order and high positioning accuracy.
[0116] Experimental results show that the method based on geometrically redefining key points achieves the best detection success rate and error control. This method effectively avoids duplicate distribution by leveraging the geometric distribution characteristics of key points, making it suitable for locating key points of regular anatomical structures in medical images. Other methods still have some applicability in specific scenarios, but their accuracy and robustness are relatively weak.
[0117] (3) Verification of the effect of loss function optimization This embodiment designs a set of comparative experiments to evaluate the performance of different loss functions in the task of cardiac ultrasound key point detection. All models are trained and verified based on the cardiac ultrasound key point detection dataset introduced in this embodiment. The network architecture adopts the modified ResUNet network of this embodiment. The heat map regression method adopts the key point redefinition method based on set significance of this embodiment. Adam is used as the optimizer, the initial learning rate is set to 0.001, and the batch_size is set to 32. All training configurations in the experiment remain consistent to ensure the fairness of the comparison results. The results of the comparative experiment are shown in the following table.
[0118]
[0119] In the task of cardiac ultrasound key point detection, Shape-aware BCE Loss performs best. It combines classification capabilities and geometric structure perception, ensuring high detection accuracy while optimizing anatomical consistency between key points. It is suitable for tasks with high precision and complex structures.
[0120] In the task of cardiac ultrasound keypoint detection, the performance and adaptability of various loss functions vary significantly, and their performance is closely related to their characteristics. L2 Loss performs the worst, with a slow convergence rate and a low plateau. The curve also exhibits significant fluctuations. This is primarily due to its sensitivity to outliers and noise, as well as its lack of geometric structure awareness between keypoints. This makes it ineffective in handling the fuzzy boundaries and noisy characteristics of cardiac ultrasound images. BCE Loss, with its strong classification capabilities, can quickly distinguish foreground from background. Its curve converges quickly and eventually stabilizes at a high detection success rate with relatively low error. It is suitable for general keypoint detection tasks, especially in scenarios with limited computational resources. AdaptiveWing Loss excels at handling outliers, as shown by its fast convergence rate. However, its limited ability to optimize the details of local keypoints limits further performance improvement. Shape-aware Loss incorporates geometric structure awareness to capture the global anatomical shape of ultrasound images. It is suitable for tasks requiring global shape consistency, such as cardiac atrial septum keypoint detection, but its ability to optimize the error of local keypoints is limited. Shape-aware BCE Loss performed the best, with its curve demonstrating the fastest convergence speed and highest detection success rate. By combining the classification capabilities of BCE with the geometric perception of Shape-aware, this method excels in the presence of blurred boundaries and noise in ultrasound images, ensuring high accuracy in keypoint detection while optimizing anatomical consistency. It is suitable for scenarios requiring extremely high accuracy and consistency in complex anatomical structures, such as keypoint detection in the left ventricle, right ventricle, or valve regions.
[0121] Example 3 An embodiment of the present invention discloses an ultrasound image key point detection device based on heat map regression, which may include: An input data module is used to determine a preceding frame image and a succeeding frame image corresponding to each ultrasound image according to the temporal sequence of the ultrasound images corresponding to the target object, and to use each ultrasound image and its corresponding preceding frame image and succeeding frame image as an input data; A heat map generation module is used to input all input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points; A key point extraction module is used to extract key point coordinates from the predicted heat map based on a preset key point extraction strategy; The heat map prediction model is obtained through the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, and the category embedding vector is embedded in a preset initial UNet network framework. The initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
[0122] In an optional embodiment, the heat map generation module inputs all input data into the heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points. The specific operation method may include: All input data are fed into the heat map prediction model, which then filters out key points for each category based on the preset geometric features corresponding to each category of ultrasound images. For each category of key points, a boundary path of the key point geometric distribution is generated according to the preset geometric features corresponding to the category, and the key points of the category are sorted according to the preset boundary order. Each layer of heat map is limited to represent only one key point, and a prediction heat map corresponding to each key point is generated; Furthermore, the specific operation of the key point extraction module to extract key point coordinates from the predicted heat map based on the preset key point extraction strategy may include: For each predicted heat map, the maximum value coordinates are extracted from the predicted heat map as the key point coordinates corresponding to the predicted heat map.
[0123] In another optional embodiment, the heat map generation module generates, for each category of key points, a boundary path of the geometric distribution of the key points according to the preset geometric features corresponding to the category, and sorts the key points of the category according to the preset boundary order. The specific operation method may include: For each category of key points, the convex hull boundary corresponding to the key points of this category is calculated, and the boundary path of the geometric distribution of the key points is generated based on the convex hull boundary and the preset geometric features corresponding to the category; a convex hull search is performed on the key points of this category according to the preset boundary order, the position of each key point is optimized based on the convex hull search results, and the key points of this category are sorted.
[0124] In yet another optional embodiment, the device may further include: An ultrasound window module is configured to obtain multiple ultrasound videos corresponding to the target object, identify an initial window region in each ultrasound video whose brightness variation satisfies a preset brightness variation pattern, and modify the initial window region based on the morphological characteristics of the target object to obtain an ultrasound window corresponding to each ultrasound video; The image extraction module is used to extract, for each ultrasound video, a plurality of ultrasound images corresponding to the ultrasound video according to the ultrasound window corresponding to the ultrasound video.
[0125] In another optional embodiment, the specific operation of training the initial UNet network framework according to the multiple ultrasound training images corresponding to the target object to obtain the heat map prediction model may include: The initial UNet network framework is trained based on multiple ultrasound training images corresponding to the target object, and the training process is fed back based on a preset target loss function to obtain a heat map prediction model that has been trained to convergence; Among them, the preset target loss function is calculated as follows for each key point: Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation; The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
[0126] In yet another optional embodiment, a specific operation of generating a category embedding vector based on the category of each ultrasound image corresponding to the target object and embedding the category embedding vector into a preset initial UNet network framework may include: Generate a category embedding vector based on the category of each ultrasound image corresponding to the target object ,in is the dimension of the embedding vector; The embedding vector is based on the following formula The mapping embedding vector is obtained by mapping the channel dimension C of the feature map output by the convolution layer of the initial UNet network framework to the fully connected layer of the initial UNet network framework. :
[0127] In the above formula, and are the weights and biases of the fully connected layer of the initial UNet network framework; Embed the map into a vector It is expanded to a dimension matching the feature map and features are fused with the feature map to embed the category embedding vector into the preset initial UNet network framework.
[0128] In another optional embodiment, the specific operation of determining the loss function corresponding to the key point according to the deviation between the real heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point may include: The loss function corresponding to the key point i is calculated according to the following formula :
[0129] In the above formula, y is the real heat map, is the prediction heat map, is the geometric weight, is the cross entropy loss corresponding to the key point i, which is calculated by the following formula :
[0130] In the above formula, N is the total number of heat maps.
[0131] Example 4 See also Figure 9 , Figure 9 FIG is a schematic diagram of the structure of an ultrasound image key point detection system based on heat map regression disclosed in an embodiment of the present invention. Figure 9 As shown, the ultrasound image key point detection system based on heat map regression may include: A memory 201 storing executable program code; a processor 202 coupled to the memory 201; The processor 202 calls the executable program code stored in the memory 201 to execute the steps of the ultrasound image key point detection method based on heat map regression described in the first or second embodiment of the present invention.
[0132] Example 5 An embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps of the ultrasound image key point detection method based on heat map regression described in Example 1 or Example 2 of the present invention.
[0133] Example 6 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the ultrasound image key point detection method based on heat map regression described in Example 1.
[0134] The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.
[0135] Through the detailed description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0136] Finally, it should be noted that the ultrasound image key point detection method, device, and medium based on heat map regression disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for detecting key points in ultrasound images based on heatmap regression, characterized in that: The method comprises: According to the temporal sequence of a plurality of ultrasound images corresponding to the target object, a preceding frame image and a succeeding frame image corresponding to each ultrasound image are determined, and each ultrasound image and its corresponding preceding frame image and succeeding frame image are used as an input data; Inputting all the input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points; Extracting key point coordinates from the predicted heat map based on a preset key point extraction strategy; The heat map prediction model is obtained by the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, the category embedding vector is embedded in a preset initial UNet network framework, and the initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
2. The ultrasound image key point detection method based on heat map regression according to claim 1, characterized in that: The step of inputting all the input data into the heat map prediction model to generate a number of prediction heat maps for representing key point distribution probabilities includes: Inputting all the input data into the heat map prediction model, the heat map prediction model screens out key points of each category according to the preset geometric features corresponding to the category of each ultrasound image; For each category of key points, a boundary path of the key point geometric distribution is generated according to the preset geometric features corresponding to the category, and the key points of the category are sorted according to the preset boundary order. Each layer of heat map is limited to represent only one key point, and a prediction heat map corresponding to each key point is generated; And, extracting key point coordinates from the predicted heat map based on a preset key point extraction strategy includes: For each of the predicted heat maps, the maximum value coordinates are extracted from the predicted heat map as the key point coordinates corresponding to the predicted heat map.
3. The ultrasonic image key point detection method based on heat map regression according to claim 2, characterized in that: For each category of key points, a boundary path of the key point geometric distribution is generated according to the preset geometric features corresponding to the category, and the key points of the category are sorted according to the preset boundary order, including: For each category of key points, the convex hull boundary corresponding to the key points of this category is calculated, and the boundary path of the geometric distribution of the key points is generated based on the convex hull boundary and the preset geometric features corresponding to the category; a convex hull search is performed on the key points of this category according to the preset boundary order, the position of each key point is optimized based on the convex hull search results, and the key points of this category are sorted.
4. The ultrasonic image key point detection method based on heat map regression according to claim 1, characterized in that: The method further comprises: Acquiring multiple ultrasound videos corresponding to the target object, for each of the ultrasound videos, identifying an initial window region in the ultrasound video whose brightness change satisfies a preset brightness change pattern, and modifying the initial window region based on the morphological characteristics of the target object to obtain an ultrasound window corresponding to each ultrasound video; For each of the ultrasound videos, multiple ultrasound images corresponding to the ultrasound video are extracted according to the ultrasound window corresponding to the ultrasound video.
5. The ultrasonic image key point detection method based on heat map regression according to claim 1, characterized in that: The step of training the initial UNet network framework according to the plurality of ultrasound training images corresponding to the target object to obtain a heat map prediction model includes: Training the initial UNet network framework according to a plurality of ultrasound training images corresponding to the target object, providing feedback on the training process based on a preset target loss function, and obtaining a heat map prediction model trained to convergence; The preset target loss function is calculated as follows for each key point: Determine a preset geometric feature corresponding to the category of the key point, calculate a geometric deviation of the key point in the geometric feature based on the geometric feature, and determine a geometric weight corresponding to the key point based on the geometric deviation; The loss function corresponding to the key point is determined based on the deviation between the true heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point.
6. The ultrasonic image key point detection method based on heat map regression according to claim 1, characterized in that: Generating a category embedding vector according to the category of each ultrasound image corresponding to the target object, and embedding the category embedding vector into a preset initial UNet network framework, comprising: Generate a category embedding vector based on the category of each ultrasound image corresponding to the target object ,in is the dimension of the embedding vector; The embedding vector is based on the following formula The mapping embedding vector is obtained by mapping the channel dimension C of the feature map output by the convolutional layer of the initial UNet network framework to the fully connected layer of the initial UNet network framework. : In the above formula, and are the weights and biases of the fully connected layer of the initial UNet network framework; Embed the mapping into a vector The vector is expanded to a dimension matching the feature map and feature fused with the feature map, thereby embedding the category embedding vector into a preset initial UNet network framework.
7. The ultrasound image key point detection method based on heat map regression according to claim 5, characterized in that: Determining the loss function corresponding to the key point according to the deviation between the real heat map and the predicted heat map of the key point and the geometric weight corresponding to the key point includes: The loss function corresponding to the key point i is calculated according to the following formula : In the above formula, y is the real heat map, is the prediction heat map, is the geometric weight, is the cross entropy loss corresponding to the key point i, which is calculated by the following formula : In the above formula, N is the total number of heat maps.
8. An ultrasonic image key point detection device based on heat map regression, characterized in that: The device comprises: An input data module is used to determine a preceding frame image and a succeeding frame image corresponding to each ultrasound image according to the temporal sequence of the ultrasound images corresponding to the target object, and to use each ultrasound image and its corresponding preceding frame image and succeeding frame image as an input data; A heat map generation module is used to input all the input data into a preset heat map prediction model to generate a number of prediction heat maps for representing the distribution probability of key points; A key point extraction module, configured to extract key point coordinates from the predicted heat map based on a preset key point extraction strategy; The heat map prediction model is obtained by the following steps: According to the category of each ultrasound image corresponding to the target object, a category embedding vector is generated, the category embedding vector is embedded in a preset initial UNet network framework, and the initial UNet network framework is trained according to multiple ultrasound training images corresponding to the target object to obtain a heat map prediction model.
9. An ultrasound image key point detection system based on heat map regression, characterized in that: The system includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the ultrasound image key point detection method based on heat map regression as described in any one of claims 1-7.
10. A computer storage medium, characterized in that The computer storage medium stores computer instructions, which, when called, are used to execute the ultrasound image key point detection method based on heat map regression as described in any one of claims 1 to 7.
Citation Information
Cited By
Brain tumor detection method based on attention mechanism and MRI (Magnetic Resonance Imaging) multi-modal fusion
CN121353272A