Ultrasonic cardiogram multi-section unified segmentation method and device, equipment and storage medium
By introducing a priori embedding of sections and cross-branch attention mechanisms in the echocardiography segmentation model, combining SAM large model and convolutional neural network, the problem that echocardiography segmentation model needs to be customized for each section in the existing technology is solved, and a unified segmentation of multiple sections of echocardiography is achieved, which improves segmentation efficiency and accuracy.
Patent Information
- Application Number
- CN202510248466.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, segmentation models of echocardiography need to be customized for each specific section, resulting in additional workload and performance degradation, limiting their clinical application.
By obtaining the high-dimensional hidden space feature vector and cluster center features of echocardiography based on the slicing type encoder, computing the cluster center similarity and constructing a slicing prior embedding, combining SAM large model and convolutional neural network, training is adopted using a cross-branch attention mechanism and local feature fusion adaptation module to generate the target image segmentation mask.
The unified segmentation of multiple sections of echocardiography is realized, which reduces the need for specialized models for different sections, improves segmentation efficiency and accuracy, and can effectively deal with structural differences between different sections and the problems of boundary blur and motion artifacts of echocardiography.
Smart Images

Figure CN120182292A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for unified segmentation of multiple echocardiogram sections. Background Art
[0002] The health of the heart is directly related to the normal operation of the whole body. Once the heart function is damaged, serious problems such as myocardial ischemia and heart failure will be caused, which will further affect the normal functions of various organs of the whole body. China is a country with a high incidence of heart diseases. The incidence of cardiovascular diseases is increasing year by year, and cardiovascular disease deaths have ranked first among the total causes of death. Echocardiogram has become one of the most widely used cardiac examination methods in clinical practice due to its advantages such as non-invasive, low cost, and real-time imaging. It can evaluate the size of the heart cavity, the function of heart valves, and detect various heart diseases such as abnormal movement of the heart wall. In clinical practice, echocardiologists need to scan the heart from different positions and angles to obtain echocardiograms of different sections, and comprehensively observe the structure and function of the heart by combining images of different sections. The main sections include the long-axis section and the short-axis section. Common ones are two-chamber heart (2CH), three-chamber heart (3CH), four-chamber heart (4CH), and parasternal left ventricular short-axis section (Parasternal short axis, PSAX), etc. The analysis of multi-section two-dimensional echocardiogram images plays an important role in measuring the morphology and function of the heart and making a diagnosis. This analysis is based on the interpretation of clinical indicators, which are extracted from low-level image processing such as segmentation and tracking. By accurately segmenting these sections, doctors can obtain more information about heart function and structure, thereby improving the accuracy and comprehensiveness of diagnosis. However, in clinical routine, semi-automatic or manual annotation is still the daily work, but this is very time-consuming and laborious and there are significant differences in user experience.
[0003] At present, many computer-aided methods have achieved the segmentation of key structures of echocardiogram to help doctors screen and diagnose diseases. However, due to significant structural differences between different sections, the segmentation model of echocardiogram needs to be customized for each specific section, which causes additional workload, and these models will show obvious performance degradation when generalized to other sections, so their clinical applications are very limited. Summary of the Invention
[0004] The purpose of the embodiments of this application is to propose a method, device, equipment and storage medium for unified segmentation of multiple echocardiogram sections to solve the problem of additional repetitive work caused by the need to design specialized models for different sections, and improve the efficiency of multi-section segmentation of echocardiogram.
[0005] To solve the above technical problems, the embodiments of this application provide a method for unified segmentation of multiple echocardiogram sections, including:
[0006] Based on the sectional type encoder, obtain the high-dimensional latent space feature vectors and clustering center features of echocardiograms including different sections of the long axis and short axis;
[0007] Calculate the clustering center similarity based on the high-dimensional latent space feature vectors and the clustering center features, and construct the prior shape of the clustering center. Then, construct the sectional prior embedding based on the clustering center similarity and the prior shape of the clustering center;
[0008] Predict the echocardiogram based on the sectional prior embedding to generate a rough segmentation result, and use the rough segmentation result as the dense prompt input of the SAM large model;
[0009] Freeze the backbone structure of the SAM large model, introduce a convolutional neural network, and train based on the dense prompt and point prompt according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate the target image segmentation mask.
[0010] To solve the above technical problems, an embodiment of the present application provides an echocardiogram multi-section unified segmentation device, including:
[0011] A feature vector acquisition unit, configured to obtain the high-dimensional latent space feature vectors and clustering center features of echocardiograms including different sections of the long axis and short axis based on the sectional type encoder;
[0012] A sectional prior embedding construction unit, configured to calculate the clustering center similarity based on the high-dimensional latent space feature vectors and the clustering center features, and construct the prior shape of the clustering center, and then construct the sectional prior embedding based on the clustering center similarity and the prior shape of the clustering center;
[0013] A rough segmentation result generation unit, configured to predict the echocardiogram based on the sectional prior embedding to generate a rough segmentation result, and use the rough segmentation result as the dense prompt input of the SAM large model;
[0014] An image segmentation mask generation unit, configured to freeze the backbone structure of the SAM large model, introduce a convolutional neural network, and train based on the dense prompt and point prompt according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate the target image segmentation mask.
[0015] To solve the above technical problems, a technical solution adopted by the present invention is: provide a computer device, including one or more processors; a memory for storing one or more programs, so that one or more processors implement the echocardiogram multi-section unified segmentation method described in any one of the above.
[0016] To solve the above technical problems, a technical solution adopted by the present invention is: a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned echocardiogram multi-plane unified segmentation method is implemented.
[0017] The embodiments of the present invention provide an echocardiogram multi-plane unified segmentation method, device, equipment and storage medium. Among them, the method includes: obtaining high-dimensional hidden space feature vectors and clustering center features of echocardiograms including different long-axis and short-axis planes based on a plane type encoder; calculating clustering center similarity and constructing a clustering center prior shape based on the high-dimensional hidden space feature vectors and the clustering center features, and constructing a plane prior embedding based on the clustering center similarity and the clustering center prior shape; predicting the echocardiogram based on the plane prior embedding to generate a rough segmentation result, and using the rough segmentation result as the dense prompt input of the SAM large model; freezing the backbone structure of the SAM large model, introducing a convolutional neural network and training based on the dense prompt and point prompt according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate a target image segmentation mask. The embodiments of the present invention solve the problem of the need to design a dedicated model for different planes, resulting in additional repetitive work, and obtain a unified segmentation model capable of segmenting any cardiac ultrasound plane, which is beneficial to improving the efficiency of multi-plane segmentation of echocardiograms. Description of the Drawings
[0018] To more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the following drawings are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of the implementation of the echocardiogram multi-plane unified segmentation method provided by the embodiments of this application;
[0020] Figure 2 It is a schematic diagram of the entire prediction and segmentation process of the echocardiogram multi-plane unified segmentation method provided by the embodiments of this application;
[0021] Figure 3 It is a schematic diagram of the network framework of the segmentation model provided by the embodiments of this application;
[0022] Figure 4 It is a schematic diagram of the prior-based composable learning module provided by the embodiments of this application;
[0023] Figure 5It is a schematic diagram of the CNN-Decoder local feature transfer and fusion module provided by an embodiment of the present application;
[0024] Figure 6 It is a schematic diagram of the segmentation results of key structures in echocardiograms of different cross-sections provided by an embodiment of the present application;
[0025] Figure 7 It is a schematic diagram of the echocardiogram multi-section unified segmentation device provided by an embodiment of the present application;
[0026] Figure 8 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0028] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase does not necessarily refer to the same embodiment at every occurrence in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0029] To enable those skilled in the art of this technology to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0030] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0031] It should be noted that the echocardiogram multi-section unified segmentation method provided by the embodiments of this application is generally executed by a server. Correspondingly, the echocardiogram multi-section unified segmentation device is generally configured in the server.
[0032] Please refer to Figure 1 and Figure 2 , Figure 1 shows a specific implementation manner of the echocardiogram multi-section unified segmentation method, Figure 2It is a schematic diagram of the entire prediction and segmentation process of the multi-plane unified segmentation method for echocardiograms provided by the embodiments of this application.
[0033] It should be noted that if there are substantially the same results, the method of the present invention is not limited to Figure 1 the process sequence shown, and the method includes the following steps:
[0034] S1: Based on the section type encoder, obtain the high-dimensional hidden space feature vectors and clustering center features of echocardiograms including different sections of the long axis and short axis.
[0035] In a specific embodiment, step S1 includes obtaining the echocardiograms including different sections of the long axis and short axis, and preprocessing the echocardiograms; inputting the preprocessed echocardiogram of any section into the section type encoder for encoding to generate the high-dimensional hidden space feature vectors of all images; performing clustering based on the high-dimensional hidden space feature vectors to obtain the clustering center features and clustering results.
[0036] Specifically, obtain the echocardiograms including different sections of the long axis and short axis, and then perform preprocessing such as unifying the image size and pixels of the echocardiograms.
[0037] In a specific embodiment, multi-section echocardiograms and corresponding annotations collected from multiple top-three hospitals are included. There are a total of 41,276 echocardiograms of different sections and corresponding annotations. Among them, the 2CH, 3CH, 4CH, and PSAX sections are 16,694, 4,249, 18,016, and 2,317 ultrasound image-annotation pairs respectively. The 2CH and 4CH are annotated with the left atrium, left ventricle, and myocardium, while the 3CH and PSAX are annotated with the regions of the left ventricle and myocardium. As shown in Table 1 below, this table details the dataset of all echocardiograms used:
[0038] Table 1
[0039] Dataset 2CH 3CH 4CH PSAX Internal Dataset 16694(1,2,3) 4249(2,3) 18016(1,2,3) 2317(2,3)
[0040] What is represented in parentheses in Table 1 are the included annotation categories, where 1, 2, and 3 represent the left atrium, left ventricle, and myocardium respectively.
[0041] In a specific embodiment, obtain the echocardiograms including different sections of the long axis and short axis, where the echocardiograms include the original echocardiogram and the annotated echocardiogram; perform image size unification processing on the original echocardiogram and the annotated echocardiogram; perform pixel unification processing on the size-unified annotated echocardiogram according to different annotation parts.
[0042] Specifically, the echocardiograms of different cross-sections and the corresponding dimensions of the echocardiograms are unified into a preset dimension. For example, the image dimension is unified to 256×256. The annotation pixel unification is specifically as follows: the pixel values corresponding to the left atrium, left ventricle, and myocardium in the heart structure of the annotated echocardiogram are unified to 1, 2, and 3. The left ventricle region is divided from the myocardial annotation. When there is only myocardial annotation data, according to the anatomical structure of the human heart, the region surrounded by the myocardium to form the left ventricle is used to improve the richness of the data. Then, the annotation information in the annotated echocardiogram after pixel unification is obtained, and based on the heart structure category of the annotation information, the annotated echocardiogram after pixel unification is subjected to pixel-level summation, and the pixel summation result is normalized. Among them, the pixel sum result is normalized to [0, 255] to obtain the average annotation m i , where the cross-section category i ∈ {1,..., K}, and K is the total number of cross-section categories.
[0043] Specifically, a ResNet34 network is built as the cross-section type encoder. The preprocessed echocardiograms are divided into training data, validation data, and test data according to a preset ratio. The preset ratio can be set according to the actual situation and is not limited here. In a specific embodiment, the echocardiograms are divided into training data, validation data, and test data according to 8:1:1. The training data is put into the cross-section type encoder for training, and the validation data is used to screen the optimal model, and finally the optimal model is tested on the test data. During the model training process, the ResNet34 network is trained from scratch, and before the training data is input into the cross-section classification model, the training data is subjected to image enhancement, where the image enhancement uses random horizontal rotation, random brightness, and contrast adjustment. Since the classification task is relatively simple, the maximum number of iterations is set to 20, the optimizer is SGD, the initial learning rate is 1e-3, and the selected loss is CE Loss. The performance of the optimal model screened by the final validation set in the test set is 100% accuracy, indicating that the model has learned the mutual relationship between the cross-sections, thus obtaining a trained cross-section type encoder.
[0044] The echocardiograms of the same cross-section are input into the cross-section type encoder for classification processing to generate high-dimensional latent space feature vectors of all images of the same cross-section. For example, the echocardiograms of the 2CH, 3CH, 4CH, and PSAX cross-sections are input into the cross-section type encoder to obtain high-dimensional latent space feature vectors of all images of the 2CH, 3CH, 4CH, and PSAX cross-sections respectively. Clustering is performed based on the high-dimensional latent space feature vectors to obtain the clustering center features and clustering results.
[0045] S2: Calculate cluster center similarity and construct cluster center prior shape based on the high-dimensional latent space feature vector and the cluster center feature, and construct section prior embedding based on the cluster center similarity and the cluster center prior shape.
[0046] In a specific embodiment, step S2 includes: generating the cluster center prior shape based on the clustering results and the image annotation shape information of the echocardiogram; calculating the feature vector of the echocardiogram of any section based on the section type encoder, and calculating the similarity between the feature vector and each cluster center feature to obtain the cluster center similarity; weighting the cluster center prior shapes respectively by the cluster center similarity vector to generate the section prior embedding.
[0047] Specifically, the previous step has generated high-dimensional latent space feature vectors and cluster center features. In the embodiment of the present application, the feature center u of each slice is calculated based on all high-dimensional latent space feature vectors of the same slice. i Then re-analyze each echocardiogram j The feature processing is performed in the input section type encoder to generate a feature vector. The cosine similarity ω(i, j) between the feature vector and the cluster center feature of each section is calculated. The cosine similarity ω(i, j) is the distance between the input echocardiogram and the feature center of each section. It is a quantitative indicator to measure the correlation between the image and the average feature of each section. The cosine similarity ω(i, j) is the cluster center similarity. The cosine similarity is calculated as follows:
[0048]
[0049] Among them, E Lat (I j ) refers to the feature vector obtained by passing the input image through the slice classification network, u i is the feature center, I j Pre-processed echocardiogram.
[0050] S3: Predicting the echocardiogram based on the slice prior embedding to generate a rough segmentation result, and using the rough segmentation result as a dense prompt input of the SAM large model.
[0051] In a specific embodiment, step S3 includes: using the coarse segmentation result as a dense prompt input of the SAM large model, and calculating a loss function based on the coarse segmentation result and the expected true mask to train the convolutional neural network and the model parameters of the cross-branch attention mechanism and the local feature fusion adaptation module.
[0052] Specifically, after calculating the cosine similarity, multiply the cosine similarity with the annotation information of the echocardiogram according to the channels corresponding to the section, and stack them to obtain combinable verification, that is, generate a section prior embedding. The section prior embedding is constructed using the following formula:
[0053] PE j = concat([ω(i,j) × m i ), i = 1,...K;
[0054] where, PE j is the section prior embedding, m i is the average annotation, and ω(i,j) is the cosine similarity.
[0055] S4: Freeze the backbone structure of the SAM large model, introduce a convolutional neural network and train it based on the cross-branch attention mechanism and the local feature fusion adaptation module using the dense prompt and the point prompt to generate a target image segmentation mask.
[0056] In a specific embodiment, the backbone structure of the SAM large model includes an image encoder and a mask decoder; step S4 includes: performing multi-scale fusion of the convolutional neural network branch with the output features of the image encoder through the cross-branch attention mechanism, and adaptively fusing the multi-scale features output by the convolutional neural network with the features output by the mask decoder by the local feature fusion adaptation module; freezing all the parameters of the SAM image encoder and the first two layers of parameters of the mask decoder; updating the parameters of the convolutional neural network, the last three layers of the mask decoder, the cross-branch attention mechanism and the local feature fusion adaptation module based on the cardiac ultrasound image segmentation target, and training with the dense prompt input and the point prompt as the input of the mask decoder to generate the target image segmentation mask.
[0057] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the segmentation model network framework provided by the embodiments of the present application. As shown in Figure 3As shown, the SAM large model includes a SAM architecture, a dense prompt generation component, and a local feature supplementation component. Among them, the SAM architecture includes the image encoder, mask decoder, and prompt encoder of the original Segmentation Anything Model. The layer where it is located is the weight freezing layer, and the weights are taken from the weights of the ViT-B backbone pre-trained on the SA-1B dataset, retaining the extensive features learned. The echocardiogram generates image embeddings through downsampling and image encoder blocks in this main branch. This is weighted and summed with the dense prompt embeddings obtained from the dense prompt generation branch and the local feature embeddings obtained from the local feature branch, and then input into the mask decoder. The mask decoder combines the fused image embeddings and sparse prompt embeddings, and then the result is upsampled to obtain the final prediction result. This result is constrained by the preset Dice Loss and BCE Loss, where the preset annotation is set according to the actual situation.
[0058] The dense prompt generation component mainly learns dense prompts from the input cross-sectional prior embeddings. Specifically, it adjusts the mask of the echocardiogram according to the cross-sectional prior embeddings to generate a rough segmentation result. The local feature supplementation component is used to supplement local features for the SAM architecture, which can effectively solve the problems of blurred echocardiogram boundaries and large motion contrast. The input echocardiogram first passes through three residual blocks, and then through four convolutional blocks. The input of each convolutional block passes the local information to the image encoder of the SAM architecture through the cross-branch attention mechanism. At the same time, the output of each convolutional block passes the local features to the mask decoder through the CNN-Decoder local feature transfer and fusion module. At the same time, in order to adapt to the information fusion of the corresponding layer, the mask decoder adds three more decoder Transformer blocks on the basis of the original SAM framework. Its structure is the same as the first two layers, except that the parameters are learnable.
[0059] Please refer to Figure 4 , Figure 4 the schematic diagram of the prior-based combinable learning module provided by the embodiment of the present application, predicts the mask of the echocardiogram based on the cross-sectional prior embeddings, and obtains the rough segmentation result.
[0060] Specifically, the lightweight U-Net in the SAM large model adjusts the mask of the echocardiogram based on the cross-sectional prior embeddings to obtain the rough segmentation result PCM j . The rough segmentation result PCM j is calculated as follows:
[0061] PCM j = UNet θ (PE j );
[0062] Among them, θ indicates that this is a learnable lightweight network. Table 2 below shows the network details of the lightweight U-Net:
[0063] Table 2
[0064]
[0065] Specifically, in the local feature supplementation component, the echocardiogram is feature-processed through multiple residual blocks and convolutional blocks to generate local features, and the local features are transmitted to the SAM architecture and the initial image embedding. Then, feature fusion is performed based on the local features and the image embedding to generate target fusion features. Next, the target fusion features are feature-decoded by a mask decoder to generate a prediction result.
[0066] Please refer to Figures 5 to 6 , Figure 5 which is a schematic diagram of the CNN-Decoder local feature transmission and fusion module provided by an embodiment of the present application, Figure 6 and which is a schematic diagram of the key structure segmentation results of echocardiograms in different sections provided by an embodiment of the present application.
[0067] In a specific embodiment, the echocardiogram is feature-processed through multiple residual blocks to generate residual features. Among them, the feature processing of each residual block includes downsampling processing, convolutional processing, normalization processing, and activation processing; the residual features are feature-processed through multiple convolutional blocks to generate local features for each convolutional block; the local features of each convolutional block are transmitted to the SAM architecture based on a cross-branch attention mechanism, and the local features of the last convolutional block are incorporated into the initial image embedding with a preset weight to obtain a target image embedding.
[0068] Specifically, the echocardiogram passes through the local feature branch in the local feature supplementation component, which includes three residual blocks and four convolutional blocks. The feature processing of each residual block includes downsampling processing, convolutional processing, normalization processing, and activation processing. After passing through an activation function, the output is combined with the residual to obtain the output, which is the residual feature and can be the result of input Conv2. Then, the result output by the residual block passes through four convolutional blocks, and convolutional processing, normalization processing, and activation processing are performed in each convolutional block. Next, the generated local features are transmitted to the SAM architecture through a cross-branch attention mechanism, and the local features of the last convolutional block are incorporated into the initial image embedding with a preset weight to obtain a target image embedding. Among them, the preset weight is not limited, and the weight can be 0.5.
[0069] Specifically, the target image embedding, local features, and the rough segmentation result are weighted and summed to generate the basic features. Then, feature fusion is performed based on the basic features and local features to obtain the target fusion features, and the fusion features are decoded based on the mask decoder to generate the prediction result, that is, the target image segmentation mask. The formula for the specific fusion process is as follows:
[0070] f F,l = conv 1×1 (concat(f CNN,l ,(f DM-K,l )));
[0071] where f F,l is the feature after fusion at the l-th layer, f CNN,l is the feature at the l-th layer passed from the convolutional block. The output of each layer of the mask decoder Transformer block includes Key and Query, f DM-K,l is the Key in the output of the l-th layer of the mask decoder Transformer block, and another output is f DM-Q,l . The output of the next layer of the mask decoder Transformer block can be expressed by the following formula:
[0072] [f DM-K,l+1 ,f DM-Q,l+1 = DTrans(f F,l ,f DM-Q,l );
[0073] where DTrans is the Transformer block of the mask decoder in the SAM architecture.
[0074] In a specific embodiment, when training the SAM large model, the training period is set to 100, the optimizer is Adam, and the initial learning rate is 1e -4 . The total loss function is:
[0075]
[0076] where is the loss between the final prediction result output by the SAM architecture and the preset standard, is the loss between the rough segmentation result and the preset standard, λ is 0.5, and both losses are composed of Dice Loss and BCELoss with weights of 0.8 and 0.2 respectively. Among them, the preset standard is set according to the actual situation and is not limited here. The preset standard can be the gold standard.
[0077] In another specific embodiment, through the local feature fusion adaptation module, the multi-scale features output by the convolutional neural network are fused with the features output by the mask decoder to generate the fused decoder features, and the fused decoder features are input into the next layer of the mask decoder. Among them, the fusion process includes mask decoder feature dimension transformation, local convolutional feature splicing, convolutional transformation, and dimension transformation.
[0078] In yet another specific embodiment, during the training process of the convolutional neural network branch and the local fusion adaptation module, it further includes: performing image morphological processing and image enhancement processing on the echocardiogram. Among them, the image morphological processing includes random horizontal flipping, random cropping, random rotation, and random affine transformation, and the image enhancement processing includes random Gamma noise addition, random contrast enhancement, and random color jitter.
[0079] As Figure 6 shown, it is a schematic diagram of the segmentation results of the key structures of echocardiograms in different sections. There are a total of 4 sections in this figure, from top to bottom are 2CH, 3CH, 4CH, and PSAX. Each row shows the segmentation results of five preferred examples corresponding to the section. Among them, the green, red, and blue solid lines respectively mark the gold standards of the left atrium, left ventricle, and myocardium, and the light green, red, and blue areas are the predicted results of the left atrium, left ventricle, and myocardium by the segmentation model. It can be seen that even though there are significant structural differences between different sections, and the echocardiogram itself has problems of strong noise and strong motion artifacts, the embodiments of the present application can accurately segment the key structure areas of the heart. For each heart structure in each section of the echocardiogram, the present application can achieve high-quality segmentation of the key heart structures. Dice and IoU mainly measure the regional overlap degree between the predicted results and the gold standards, and HD95 measures the matching degree of the shape boundaries. The average Dice coefficients of this algorithm in the four sections of 2CH, 3CH, 4CH, and PSAX are 89.67%, 87.27%, 90.28%, and 88.26% respectively, and the IoU is 81.87%, 78.01%, 82.67%, and 79.68 respectively, indicating that there is a good overlap degree between the predicted results and the gold standards. The average HD95 in the four sections is only 6.52, 4.54, 5.61, and 4.67 respectively, indicating that the embodiments of the present application can effectively solve the problem of blurred boundaries of echocardiograms. As shown in Table 3 below, it shows the evaluation indexes of each heart structure in each section of the internal test set.
[0080] Table 3
[0081]
[0082]
[0083] In the embodiments of the present application, a high-dimensional latent space feature vector and a clustering center feature of an echocardiogram including different cross-sections of a long axis and a short axis are obtained based on a cross-section type encoder; the clustering center similarity is calculated based on the high-dimensional latent space feature vector and the clustering center feature, and a prior shape of the clustering center is constructed, and a prior embedding of the cross-section is constructed based on the clustering center similarity and the prior shape of the clustering center; a rough segmentation result is generated by predicting the echocardiogram based on the prior embedding of the cross-section, and the rough segmentation result is used as the dense prompt input of the SAM large model; the backbone structure of the SAM large model is frozen, a convolutional neural network is introduced, and training is performed based on the dense prompt and the point prompt according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate a target image segmentation mask. The embodiments of the present invention solve the problem of additional repetitive work caused by the need to design a dedicated model for different cross-sections, and obtain a unified segmentation model capable of segmenting any cardiac ultrasound cross-section, which is beneficial to improving the efficiency of multi-cross-section segmentation of echocardiograms.
[0084] The segmentation model of the key cardiac structures of the multi-cross-section echocardiogram in the embodiments of the present application realizes that a single model can perform high-precision segmentation of the key cardiac structures of the echocardiograms of multiple cross-sections, effectively solving the problems of significant structural differences between different cross-sections and the natural blurring of boundaries and large motion artifacts in echocardiograms. In the face of the significant structural differences and potential associations between echocardiograms of different cross-sections, the embodiments of the present application learn the relationships between cross-sections through a classification network, quantify this potential difference and association with the similarity between the input image and the cross-section center, and combine the average shape of the cross-section annotation to obtain the prior knowledge of the cross-section. Then, the rough segmentation mask is adjusted through a learnable lightweight encoding-decoding network and used as a dense prompt to guide the model to focus on the region of interest, so as to make corresponding adjustments for different cross-sections and different images. In the face of problems such as unclear structural contours and large speckle noise in echocardiograms, the present invention extracts local features such as boundary contours in the image through a local feature supplement branch and supplements them to the main branch of the SAM architecture, and transfers the local features in the encoder to the mask decoder through the local feature fusion and adaptation module to help the model better grasp the local information of the image and achieve high-quality segmentation of echocardiograms.
[0085] Please refer to Figure 7 , as an implementation of the method shown above Figure 1 In the embodiments of the present application, an embodiment of a multi-cross-section unified segmentation device for echocardiograms is provided. This device embodiment corresponds to the method embodiment shown in Figure 1 and can be specifically applied to various electronic devices.
[0086] As Figure 7As shown in the figure, the unified multi-plane segmentation device for echocardiogram in this embodiment includes: a feature vector acquisition unit 51, a plane prior embedding construction unit 52, a rough segmentation result generation unit 53, and an image segmentation mask generation unit 54, where:
[0087] The feature vector acquisition unit 51 is used to obtain high-dimensional hidden space feature vectors and clustering center features of echocardiograms including different planes of long axis and short axis based on a plane type encoder.
[0088] The plane prior embedding construction unit 52 is used to calculate the clustering center similarity and construct the prior shape of the clustering center based on the high-dimensional hidden space feature vectors and the clustering center features, and construct the plane prior embedding based on the clustering center similarity and the prior shape of the clustering center.
[0089] The rough segmentation result generation unit 53 is used to predict the echocardiogram based on the plane prior embedding to generate a rough segmentation result, and use the rough segmentation result as the dense prompt input of the SAM large model.
[0090] The image segmentation mask generation unit 54 is used to freeze the backbone structure of the SAM large model, introduce a convolutional neural network, and train based on the dense prompt and point prompt according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate the target image segmentation mask.
[0091] Further, the feature vector acquisition unit 51 includes:
[0092] A preprocessing unit, which is used to obtain the echocardiogram including different planes of long axis and short axis, and preprocess the echocardiogram.
[0093] An encoding unit, which is used to input the preprocessed echocardiogram of any plane into the plane type encoder for encoding to generate high-dimensional hidden space feature vectors of all images.
[0094] A clustering unit, which is used to cluster based on the high-dimensional hidden space feature vectors to obtain the clustering center features and the clustering result.
[0095] Further, the plane prior embedding construction unit 52 includes:
[0096] A clustering center prior shape construction unit, which is used to generate the clustering center prior shape based on the clustering result and the image annotation shape information of the echocardiogram.
[0097] A similarity calculation unit, which is used to calculate the feature vectors of the echocardiogram of any plane based on the plane type encoder, and calculate the similarity according to the feature vectors and each clustering center feature to obtain the clustering center similarity.
[0098] A weighting unit for weighting the prior shape of the clustering center by the clustering center similarity vector respectively to generate the prior embedding of the cross-section.
[0099] Further, the rough segmentation result generation unit 53 includes:
[0100] A prediction unit for predicting the mask of the echocardiogram based on the prior embedding of the cross-section to obtain the rough segmentation result;
[0101] A loss calculation unit for using the rough segmentation result as the dense prompt input of the SAM large model and calculating a loss function based on the rough segmentation result and the expected true mask to train the model parameters of the convolutional neural network, the cross-branch attention mechanism, and the local feature fusion adaptation module.
[0102] Further, the image segmentation mask generation unit 54 includes:
[0103] An adaptive fusion unit for performing multi-scale fusion of the output features of the convolutional neural network branch and the image encoder through the cross-branch attention mechanism, and adaptively fusing the multi-scale features output by the convolutional neural network and the features output by the mask decoder by the local feature fusion adaptation module;
[0104] A parameter freezing unit for freezing all the parameters of the SAM image encoder and the first two layers of parameters of the mask decoder;
[0105] A parameter update unit for updating the parameters of the convolutional neural network, the last three layers of the mask decoder, the cross-branch attention mechanism, and the local feature fusion adaptation module based on the objective of cardiac echocardiogram image segmentation, and using the dense prompt input and the point prompt as the input of the mask decoder for training to generate the target image segmentation mask.
[0106] Further, the echocardiogram multi-cross-section unified segmentation device further includes:
[0107] A feature fusion unit for fusing the multi-scale features output by the convolutional neural network and the features output by the mask decoder through the local feature fusion adaptation module to generate the fused decoder features, and inputting the fused decoder features into the next layer of the mask decoder, where the fusion process includes mask decoder feature dimension transformation, local convolutional feature splicing, convolutional transformation, and dimension transformation.
[0108] Further, during the training process of the convolutional neural network branch and the local fusion adaptation module, it further includes: an image enhancement unit for performing image morphological processing and image enhancement processing on the echocardiogram, where the image morphological processing includes random horizontal flipping, random cropping, random rotation, and random affine transformation, and the image enhancement processing includes random Gamma noise addition, random contrast enhancement, and random color jitter.
[0109] To solve the above technical problems, an embodiment of the present application further provides a computer device. For details, please refer to Figure 8 , Figure 8 which is the basic structural block diagram of the computer device in this embodiment.
[0110] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are communicatively connected to each other through a system bus. It should be noted that Figure 8 only a computer device 7 with three components, namely a memory 71, a processor 72, and a network interface 73, is shown in
[0111] However, it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0112] The memory 71 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 7. Of course, the memory 71 may also include both the internal storage unit and the external storage device of the computer device 7. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed on the computer device 7, such as the program code of the multi-plane unified segmentation method for echocardiogram. In addition, the memory 71 can also be used to temporarily store various data that have been output or will be output.
[0113] In some embodiments, the processor 72 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to run the program code stored in the memory 71 or process data, such as running the program code of the above multi-plane unified segmentation method for echocardiogram to implement various embodiments of the multi-plane unified segmentation method for echocardiogram.
[0114] The network interface 73 may include a wireless network interface or a wired network interface, and the network interface 73 is generally used to establish a communication connection between the computer device 7 and other electronic devices.
[0115] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing a computer program, and the computer program can be executed by at least one processor to enable at least one processor to execute the steps of a multi-plane unified segmentation method for echocardiogram as described above.
[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0117] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The preferred embodiments of the present application are given in the drawings, but do not limit the scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the protection scope of the present application.
Claims
1. A unified segmentation method for multiple slices of an echocardiogram, characterized in that: include: Based on the slice type encoder, high-dimensional latent space feature vectors and cluster center features of the echocardiogram including different slices of long axis and short axis are obtained; Calculating cluster center similarity and constructing cluster center prior shapes based on the high-dimensional latent space feature vector and the cluster center features, and constructing section prior embedding based on the cluster center similarity and the cluster center prior shapes; Predicting the echocardiogram based on the slice prior embedding to generate a rough segmentation result, and using the rough segmentation result as a dense prompt input of the SAM large model; The backbone structure of the SAM large model is frozen, a convolutional neural network is introduced and trained based on the dense prompts and point prompts according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate a target image segmentation mask.
2. The method for unified segmentation of multiple slices of echocardiography according to claim 1, characterized in that: The high-dimensional latent space feature vector and cluster center feature of the echocardiogram including different long-axis and short-axis sections are obtained based on the section type encoder, including: Acquiring the echocardiogram including different sections of the long axis and the short axis, and preprocessing the echocardiogram; Inputting the preprocessed echocardiogram of any section into the section type encoder for encoding to generate high-dimensional latent space feature vectors of all images; Clustering is performed based on the high-dimensional latent space feature vector to obtain the cluster center features and clustering results.
3. The method for unified segmentation of multiple slices of echocardiography according to claim 2, characterized in that: The calculating of cluster center similarity based on the high-dimensional latent space feature vector and the cluster center feature and constructing a priori shape of the cluster center, and constructing a section prior embedding based on the cluster center similarity and the prior shape of the cluster center, includes: generating the cluster center prior shape based on the clustering result and the image annotation shape information of the echocardiogram; Calculating a feature vector of the echocardiogram of any section based on the section type encoder, and calculating similarity between the feature vector and each cluster center feature to obtain the cluster center similarity; The cluster center prior shapes are weighted respectively by the cluster center similarity vectors to generate the section prior embedding.
4. The method for unified segmentation of multiple slices of echocardiography according to claim 1, characterized in that: The method of predicting the echocardiogram based on the slice prior embedding to generate a rough segmentation result, and using the rough segmentation result as a dense prompt input of the SAM large model, includes: Predicting the mask of the echocardiogram based on the slice prior embedding to obtain the rough segmentation result; The coarse segmentation result is used as a dense prompt input of the SAM large model, and a loss function is calculated based on the coarse segmentation result and the expected true mask to train the model parameters of the convolutional neural network and the cross-branch attention mechanism and the local feature fusion adaptation module.
5. The method for unified segmentation of multiple slices of echocardiography according to claim 1, characterized in that: The SAM large model trunk structure includes an image encoder and a mask decoder; the freezing of the SAM large model trunk structure, introducing a convolutional neural network and training based on the dense prompts and point prompts according to a cross-branch attention mechanism and a local feature fusion adaptation module to generate a target image segmentation mask, including: The convolutional neural network branch is multi-scale fused with the output features of the image encoder through the cross-branch attention mechanism, and the local feature fusion adaptation module is adaptively fused with the multi-scale features output by the convolutional neural network and the features output by the mask decoder; Freeze all parameters of the SAM image encoder and the first two layers of parameters of the mask decoder; The convolutional neural network, the last three layers of the mask decoder, the cross-branch attention mechanism and the local feature fusion adaptation module are updated with parameters based on the cardiac ultrasound image segmentation target, and the dense prompt input and the point prompt are used as inputs of the mask decoder for training to generate the target image segmentation mask.
6. The method for unified segmentation of multiple slices of echocardiography according to claim 4, characterized in that: The method further comprises: Through the local feature fusion adaptation module, the multi-scale features output by the convolutional neural network are fused with the features output by the mask decoder to generate fused decoder features, and the fused decoder features enter the next layer of the mask decoder, wherein the fusion process includes mask decoder feature dimensionality transformation, local convolution feature splicing, convolution change and dimensionality change.
7. The method for unified segmentation of multiple slices of echocardiography according to claim 6, characterized in that: The training process of the convolutional neural network branch and the local fusion adaptation module also includes: performing image morphological processing and image enhancement processing on the echocardiogram, wherein the image morphological processing includes random horizontal flipping, random cropping, random rotation and random affine transformation, and the image enhancement processing includes random Gamma noise addition, random contrast enhancement and random color jittering.
8. An echocardiogram multi-section unified segmentation device, characterized in that: include: A feature vector acquisition unit, used for acquiring high-dimensional latent space feature vectors and cluster center features of the echocardiogram including different long-axis and short-axis sections based on the section type encoder; A section priori embedding construction unit, used to calculate cluster center similarity and construct a cluster center prior shape based on the high-dimensional latent space feature vector and the cluster center feature, and to construct a section priori embedding based on the cluster center similarity and the cluster center prior shape; A rough segmentation result generating unit, configured to predict the echocardiogram based on the slice prior embedding to generate a rough segmentation result, and use the rough segmentation result as a dense prompt input of the SAM large model; The image segmentation mask generation unit is used to freeze the main structure of the SAM large model, introduce a convolutional neural network and train it based on the dense prompts and point prompts according to the cross-branch attention mechanism and the local feature fusion adaptation module to generate a target image segmentation mask.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the multi-slice unified segmentation method of an echocardiogram according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for unified segmentation of multiple slices of an echocardiogram according to any one of claims 1 to 7 is implemented.