Dual-branch Fusion Heart Image Segmentation Method and System Based on Memory Queue
Through the dual-branch fusion method based on memory queue, combined with CNN-ViT dual-branch and improved algorithm, the problem of local features and global context balance in cardiac image segmentation is solved, and high-precision and high-efficiency segmentation results are achieved.
Patent Information
- Application Number
- CN202510360892.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Existing cardiac image segmentation methods are difficult to balance local features with global context, resulting in semantic conflicts and high computational overhead in segmentation results.
Using a dual-branch fusion method based on memory queue, local and global features are extracted and enhanced through CNN-ViT dual-branch and memory queue, and dynamic fusion of features and computational efficiency optimization are achieved through improved Softmax algorithm and sparse matrix compression technology.
It improves the accuracy and clinical rationality of cardiac image segmentation, solves the problem of semantic conflict, reduces calculation overhead, and improves the accuracy and real-timeness of segmentation results.
Smart Images

Figure CN119888238B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image segmentation, and particularly to a dual-branch fusion cardiac image segmentation method and system based on a memory queue. Background Art
[0002] In recent years, deep learning technology has significantly promoted the development of medical image segmentation. Convolutional neural networks (CNNs) have become the mainstream solution due to their advantages of local perception and translational invariance; U-Net and its improved models have achieved high accuracy in cardiac CT / MRI segmentation through the encoder-decoder architecture and skip connections; however, the inductive bias of CNNs limits their ability to capture long-range dependencies, resulting in insufficient global context modeling of the overall cardiac anatomical structure (such as the spatial arrangement of ventricles-atria and myocardial continuity). For example, in cardiac MRI, due to inter-slice motion artifacts or partial volume effects, simply relying on CNNs may produce topological errors (such as ventricular wall breaks) or shape distortions (such as excessive contraction of the ventricular cavity).
[0003] At the same time, Vision Transformers (ViTs) have achieved global context modeling through self-attention mechanisms and have performed well in natural image tasks; ViTs divide images into serialized image patch pieces (patches) and capture long-distance feature correlations through multi-head attention, which is theoretically more suitable for modeling the overall anatomical consistency of the cardiac organ; however, the application of ViTs in medical images faces two major challenges: First, the amount of medical data is limited, and ViTs lack the inherent locality prior of CNNs and are prone to overfitting on small datasets; second, ViTs are less sensitive to local details (such as myocardial edges and small blood vessel branches) and may lose key anatomical information.
[0004] Based on this, there is an urgent need for a new cardiac segmentation framework that deeply synergizes CNNs and ViTs, fuses local features and global context, and simultaneously introduces anatomical prior constraints to improve segmentation accuracy and clinical rationality; but for the segmentation methods of their synergy, there are currently the following problems:
[0005] (1) The similarity metric of medical images needs to simultaneously meet the matching requirements of local regions and global structures. However, local features may have semantic conflicts with the global anatomical structure, making it difficult for the model to balance the weights of the two.
[0006] (2) The current cardiac segmentation results may output abnormal shapes that deviate significantly from the true anatomical structure. For example, the ventricular cavity presents a non-physiological geometric shape, the myocardial wall breaks, or there are topological errors in the vascular network. Such problems are likely to cause misdiagnosis in clinical diagnosis.
[0007] (3) Directly calculating the similarity between all image pairs leads to huge computational overhead. Especially when dealing with large-scale image datasets, each pair of images needs to be compared in detail, and the computational volume increases exponentially with the increase in the number of images, consuming a large amount of time and computational resources. Summary of the Invention
[0008] To solve the above problems, the present disclosure proposes a dual-branch fusion cardiac image segmentation method and system based on a memory queue. Through the CNN-ViT dual-branch and the memory queue, the extraction and enhancement of local features and global features are realized, improving the segmentation accuracy and clinical rationality.
[0009] According to some embodiments, the present disclosure adopts the following technical solutions:
[0010] A dual-branch fusion cardiac image segmentation method based on a memory queue, comprising:
[0011] Input the cardiac image to be segmented into the CNN-ViT dual-branch to generate locally enhanced features and globally enhanced features with anatomical priors;
[0012] Fuse the enhanced local features and global features to obtain image segmentation features;
[0013] Decode the image segmentation features to obtain the final segmentation result;
[0014] Wherein, the anatomical prior enhancement is to cluster and store historical image features to construct a memory queue, select the historical image features most similar to the cardiac image to be segmented from the memory queue, and based on the historical image features, calculate the attention of the extracted local features and global features respectively through the similarity of adjacent windows, and enhance the local features and global features with anatomical priors.
[0015] According to some embodiments, the present disclosure adopts the following technical solutions:
[0016] A dual-branch fusion cardiac image segmentation system based on a memory queue, comprising:
[0017] A feature generation module configured to: input the cardiac image to be segmented into the CNN-ViT dual-branch to generate locally enhanced features and globally enhanced features with anatomical priors;
[0018] A feature fusion module configured to: fuse the enhanced local features and global features to obtain image segmentation features;
[0019] A feature decoding module configured to: decode the image segmentation features to obtain the final segmentation result;
[0020] Among them, the anatomical prior enhancement is to cluster and store historical image features to construct a memory queue, select the historical image features most similar to the heart image to be segmented from the memory queue, and calculate the attention of the extracted local features and global features respectively through the proximity window similarity based on the historical image features, and the local features and global features after anatomical prior enhancement.
[0021] According to some embodiments, the present disclosure adopts the following technical solutions:
[0022] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the dual-branch fusion heart image segmentation method based on a memory queue described above.
[0023] According to some embodiments, the present disclosure adopts the following technical solutions:
[0024] A non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the dual-branch fusion heart image segmentation method based on a memory queue described above is implemented.
[0025] According to some embodiments, the present disclosure adopts the following technical solutions:
[0026] An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the dual-branch fusion heart image segmentation method based on a memory queue described above.
[0027] Compared with the prior art, the beneficial effects of the present disclosure are:
[0028] The dual-branch fusion heart image segmentation method and system based on a memory queue proposed by the present invention, through a multi-module collaborative design and a cross-scale feature fusion mechanism, significantly improve the clinical practicability and algorithm robustness of medical image segmentation in the following aspects:
[0029] 1. Dynamic fusion of local-global features to solve the semantic conflict problem:
[0030] The CNN branch captures detailed features such as myocardial texture and blood vessel edges through the local perception ability of the convolution kernel, while the ViT branch uses the self-attention mechanism to model the overall anatomical structure of the heart. The fusion module adopts an improved Softmax algorithm to calculate the importance weights and non-importance weights of local features in the channel dimension, and aligns the global semantic context in the spatial dimension to achieve a dynamic balance between local details and global structures.
[0031] 2. Anatomical prior enhancement based on clustering and memory queue:
[0032] The K-means algorithm is used to perform anatomical structure clustering on the input data set, and images of the same category are assigned to the same sub-queue in the memory queue. This design enables the model to quickly retrieve the characteristics of similar cases during training and inference, significantly improving the segmentation generalization ability for rare variations. The memory queue divides the images into overlapping windows and calculates the cosine similarity of only adjacent windows instead of full-image matching, reducing the computational complexity while preserving the anatomical structure consistency.
[0033] 3. Optimization of Computational Efficiency and Real-time Performance:
[0034] The CNN and ViT branches of the encoder perform feature interaction after each layer of downsampling, avoiding the problem of misalignment of deep features in the traditional concatenation architecture. The fusion module uses sparse matrix compression technology to sparsify the global attention matrix of ViT into a block diagonal form, greatly improving the processing speed of cardiac images. The memory queue performs category division, reducing the required computational resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of this disclosure. The schematic embodiments and descriptions thereof of this disclosure are used to explain this disclosure and do not constitute an improper limitation of this disclosure.
[0036] Figure 1 It is the overall structure diagram of the cardiac image segmentation model for Embodiment 1.
[0037] Figure 2 It is the flow chart of anatomical prior enhancement for Embodiment 1.
[0038] Figure 3 It is the internal structure diagram of the fusion module for Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0040] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs.
[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0042] Example 1
[0043] In an embodiment of the present disclosure, a dual-branch fusion cardiac image segmentation method based on a memory queue is provided. By combining the CNN-ViT dual branches, the combination of local features and global features is realized, and the segmentation accuracy and clinical rationality are improved. The specific steps are as follows:
[0044] Step S1: Input the cardiac image to be segmented into the CNN-ViT dual branches to generate locally enhanced features and globally enhanced features with anatomical priors;
[0045] Step S2: Fuse the enhanced local features and global features to obtain image segmentation features;
[0046] Step S3: Decode the image segmentation features to obtain the final segmentation result;
[0047] Among them, the anatomical prior enhancement is to cluster and store the historical image features to construct a memory queue, select the historical image features most similar to the cardiac image to be segmented from the memory queue, and calculate the attention of the extracted local features and global features respectively through the proximity window similarity based on the historical image features, so as to obtain the locally enhanced features and globally enhanced features with anatomical priors.
[0048] As an embodiment, the dual-branch fusion cardiac image segmentation method based on a memory queue of the present disclosure is divided into two major steps: memory queue construction and cardiac image segmentation, which are described separately below:
[0049] I. Memory queue construction
[0050] The memory queue is used to store the feature patterns of different anatomical variations and contains several sub-queues. The image features of the same type of images are stored in each sub-queue. Therefore, the construction method is to use the K-means algorithm to perform anatomical structure clustering on the historical image dataset, and assign the image features of the same type of images to the same sub-queue of the memory queue. The specific steps for constructing the memory queue are as follows:
[0051] (1) Cardiac image clustering:
[0052] First, traverse the cardiac image files in the specified directory, load all cardiac images and flatten them into one-dimensional vectors and store them in a list. Here, the directory stores the pre-collected historical image dataset.
[0053] Then, use K-means (assuming the number of clusters is N) to perform unsupervised clustering on the cardiac images to generate the class labels corresponding to each image.
[0054] Finally, according to the clustering results, the images of the same category are grouped, and all the images in the same group belong to the same category label, providing a basis for data division for the memory queue in subsequent tasks.
[0055] (2)Construct sub-queues:
[0056] Perform window division and window feature extraction on each image in the same group: The image is divided into windows. In this embodiment, , after dividing the windows, use a feature extractor to extract the image features of each window to obtain window features , where , and finally obtain window features .
[0057] Store all the window features obtained from the images in the same group into a sub-queue, and all the sub-queues form a memory queue.
[0058] In this embodiment, the CNN memory queue for the CNN branch and the ViT memory queue for the ViT branch are constructed respectively by the above method. During the construction of the CNN memory queue, the CNN branch in the following cardiac image segmentation model is used as the feature extractor to extract local features as window features, and then the queue is constructed; during the construction of the ViT memory queue, the ViT branch in the following cardiac image segmentation model is used as the feature extractor to extract global features as window features, and then the queue is constructed.
[0059] II. Cardiac Image Segmentation
[0060] Use a cardiac image segmentation model to perform image segmentation on the cardiac image to be segmented. As Figure 1 shown, the cardiac image segmentation model includes an encoder, a fusion module, and a decoder, which will be described in detail below:
[0061] 1. Encoder
[0062] The encoder adopts a CNN-ViT dual-branch to generate locally enhanced features and globally enhanced features after anatomical prior enhancement. The CNN-ViT dual-branch includes a parallel CNN branch and a ViT branch. The CNN branch includes a number of consecutive CNN encoders and a CNN memory queue enhancement module, and the ViT branch includes a number of consecutive ViT encoders and a ViT memory queue enhancement module.
[0063] In the CNN branch, the input cardiac image is divided into a number of windows. Here, the division method in the construction of the memory queue is adopted, that is, divided into windows, and subsequent feature extraction and feature enhancement are performed in each window. Specifically:
[0064] After the window is divided, the window image enters four consecutive CNN encoders. Each CNN encoder contains convolutional layers for downsampling, gradually increasing the number of channels of the image and gradually decreasing the image size, and extracting local features of different scales of the image. , , is the original input image. The CNN encoder is expressed by the formula:
[0065]
[0066] where, is convolution, are the stride and kernel size of the convolution.
[0067] Finally, the local features of each window are obtained. , and the local features of all windows constitute the local features of the heart image. where is the th window's local feature, that is, .
[0068] The local features of the heart image enter the CNN memory queue enhancement module, and anatomical prior enhancement is performed by calculating the similarity between each adjacent window.
[0069] In the ViT branch, the input heart image is divided into several image patch patches. The division method is the same as that in the construction of the memory queue, and the patches play the role of windows in the CNN branch; the divided patch images enter four consecutive ViT encoders. Each ViT encoder contains a patch merging module, a self-attention (SA)-based transformer module, and a feed-forward neural network (FFN) module to realize the downsampling of the heart image features and extract global features of different scales of the image. (i ∈ (0, 1, 2, 3, 4)), is the original input image. The ViT encoder is expressed by the formula:
[0070]
[0071]
[0072]
[0073]
[0074] where, is the linear layer, is the concatenation of tensors, is a mask. The specific operation is to separately extract the features of the 0-3rd patches in each 2*2 patch group in which are the query, key, and value respectively, is the weight matrix of the query, key, and value, is a feed-forward neural network.
[0075] Finally, the global features of each patch are obtained and the global features of all patches constitute the global features of the heart image where is the feature of the th patch, that is .
[0076] The global features of the heart image enter the ViT memory queue enhancement module. By calculating the similarity between each neighboring window, anatomical prior enhancement is performed.
[0077] Both the CNN memory queue enhancement module and the ViT memory queue enhancement module use their respective memory queues to perform anatomical prior enhancement on the input local features and global features respectively. The number of sub-queues in the memory queue is determined by the number of class labels obtained from the clustering of the heart image. Each sub-queue of each class will only save and update the window features with this class label. Below, taking the CNN memory queue enhancement module as an example, as Figure 2 shown, the specific enhancement process is described as follows:
[0078] (1) Before segmenting the heart image, for the heart image to be segmented, class prediction is performed, and the corresponding sub-queue is selected from the memory queue according to the predicted class.
[0079] For each input image to be segmented, the image at the head of each sub-memory queue is selected respectively, and these images are all flattened into one-dimensional vectors. The Euclidean distance between the image to be segmented and the image at the head of each sub-queue is calculated. The class corresponding to the image with the smallest Euclidean distance obtained is the class of the image to be segmented. After feature extraction by the CNN / ViT encoder, the image to be segmented will enter the sub-memory queue corresponding to its class for feature enhancement.
[0080] (2) From the selected sub-queue, using the similarity of features, the historical image feature most similar to the input feature is selected.
[0081] First, for each window local feature in the local feature it is compared with all the window local features stored in the sub-queue of the same classPerform pairwise cosine similarity calculations and select the window local features that are most similar to them to form the historical image features .
[0082] (3)Based on the selected historical image features , perform attention calculations on the local features respectively through the similarity of neighboring windows
[0083] Using the attention mechanism, perform attention calculations on each window feature through the similar historical features of neighboring windows to obtain the enhanced window local features , and finally obtain the enhanced local features .
[0084] Specifically, taking the window as an example, taking the current window feature as q, the corresponding window local features of the 8 neighboring windows of this window in are used as k and v to perform attention calculations, and the calculation results are then passed through a feed-forward neural network to obtain the anatomical prior enhanced features of the current image .
[0085] To ensure the diversity and integrity of the queue images, the window features before enhancement are stored in the memory queue, that is, the queue is updated and used as the historical image of the feature map that enters the memory queue next. The enhanced local features and the global features will be fused in the fusion module
[0086] 2. Fusion Module
[0087] The local features of different scales extracted by the CNN encoder, together with the enhanced local features, will all enter the fusion module and be fused with the global features of the corresponding ViT encoder to obtain image segmentation features of different scales. Therefore, the fusion module includes five fusion units, and each unit respectively performs the fusion of the corresponding scale features. The structure of the fusion unit is as Figure 3 shown, specifically as follows
[0088] (1)Perform spatial attention calculations on the local features
[0089]
[0090] (2)Perform channel attention calculations on the global features
[0091]
[0092] (3) The local features and the global features are subjected to the Hadamard product operation together:
[0093]
[0094] Among them, represents the Hadamard product.
[0095] (4) , and enter the reSoftmax module together.
[0096] The main structure of the reSoftmax module is cross-attention , and the cross-attention is similar to the general attention mechanism. The difference is that the reSoftmax function is added in the use of the activation function. The form of the reSoftmax function is as follows:
[0097]
[0098] Among them, is the exponential function with base e, is the i-th element of the vector z, where i ∈ (1, 2, ..., N).
[0099] The structure of the reSoftmax module is as follows:
[0100]
[0101]
[0102]
[0103]
[0104] Among them, are the query, key, and value respectively, is the feed-forward neural network, is the weight matrix of the query, is the weight matrix of the key and value, is the dimension of
[0105] The common role of Softmax is to highlight the parts with larger weights in the attention matrix, which are also the parts with higher similarity. reSoftmax can simultaneously pay attention to the parts with smaller weights, so that no part is missed when the two feature maps are fused, achieving better performance.
[0106] 3. Decoding module
[0107] The decoding module includes four decoders, which are used to decode the image segmentation features at different scales to obtain the final segmentation result. Specifically:
[0108] The fusion result of the last layer in the encoder enters the decoder. The main structure of the decoder is the upsampling module. After passing through each layer of the decoder, the expansion of the image size and the reduction of the number of channels are realized. At the same time, the results of the first few fusion units will be connected by residual connection with the corresponding size to achieve feature recovery and multi-scale fusion; finally, the feature image restored to the original size passes through the segmentation prediction head to obtain the predicted result, which is expressed by the formula:
[0109]
[0110] Finally, the feature image restored to the original size passes through the segmentation prediction head to obtain the predicted result.
[0111] The performance of the heart segmentation model is evaluated by some evaluation metrics, including: DICE coefficient, 95% Hausdorff distance (95HD), and mean intersection over union (mIoU). The following are the definitions and calculation methods of these metrics:
[0112] The DICE coefficient (also known as the F1 score) is a metric for calculating the similarity between two samples and is commonly used to evaluate the accuracy of image segmentation. The DICE coefficient calculation formula is:
[0113] ;
[0114] Among them, represents the predicted segmentation result, represents the true segmentation label, is the number of pixels in the intersection of the two, and are the number of pixels of the predicted result and the true label respectively.
[0115] The mean intersection over union (mIoU) is another important metric for evaluating the performance of the model in the image segmentation task. It calculates the ratio of the average intersection to the union between the predicted result and the true label:
[0116] ;
[0117] Among them, is the number of categories, and are the predicted segmentation result and the true label for the th class respectively.
[0118] The Hausdorff distance is an index for calculating the maximum deviation between two sets of point sets. In the field of image segmentation, the 95% Hausdorff distance (95HD) is used to evaluate the accuracy of the segmentation boundary, excluding the highest 5% of abnormal extreme points in the calculation process and retaining the distance distribution characteristics of the remaining 95% of point pairs. The calculation formula is:
[0119] ;
[0120] ;
[0121] where is the Euclidean distance from point to point , and means intercepting the first 95% of the numerical range in the set of minimum distances and taking its upper limit value as the final result.
[0122] Generating and saving the segmentation result map and displaying relevant evaluation indicators can visually show the segmentation effect of the model, helping doctors make diagnosis and treatment decisions; at the same time, by providing a quantitative evaluation of the model performance, it helps to further optimize and adjust the model, thereby improving its accuracy and practicality.
[0123] Example 2
[0124] In an embodiment of the present disclosure, a dual-branch fusion cardiac image segmentation system based on a memory queue is provided, including:
[0125] A feature generation module, configured to: input the cardiac image to be segmented into the CNN-ViT dual-branch to generate local features and global features enhanced with anatomical priors;
[0126] A feature fusion module, configured to: fuse the enhanced local features and global features to obtain image segmentation features;
[0127] A feature decoding module, configured to: decode the image segmentation features to obtain the final segmentation result;
[0128] wherein, the anatomical prior enhancement is to cluster and store historical image features to construct a memory queue, select the historical image features most similar to the cardiac image to be segmented from the memory queue, and based on the historical image features, calculate the attention of the extracted local features and global features respectively through the proximity window similarity, and the local features and global features enhanced with anatomical priors.
[0129] Example 3
[0130] In one embodiment of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the dual-branch fusion cardiac image segmentation method based on a memory queue as described above.
[0131] Embodiment 4
[0132] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium is used to store computer instructions which, when executed by a processor, implement the dual-branch fusion cardiac image segmentation method based on a memory queue as described above.
[0133] Embodiment 5
[0134] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes to implement the dual-branch fusion cardiac image segmentation method based on a memory queue as described above.
[0135] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0136] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate computer-implemented processing, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0137] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A dual-branch fusion cardiac image segmentation method based on a memory queue, characterized in that: include: The heart image to be segmented is input into the dual branches of CNN-ViT to generate local features and global features enhanced by anatomical priors; The enhanced local features and global features are integrated to obtain image segmentation features; Decode the image segmentation features to obtain the final segmentation result; The anatomical prior enhancement is to cluster and store historical image features to construct a memory queue, select the historical image features that are most similar to the heart image to be segmented from the memory queue, and based on the historical image features, perform attention calculation on the extracted local features and global features respectively through the similarity of adjacent windows, and the local features and global features after anatomical prior enhancement; The CNN-ViT dual branch includes a CNN branch and a ViT branch connected in parallel; The CNN branch includes a plurality of consecutive CNN encoders and a CNN memory queue enhancement module, each CNN encoder includes a convolutional layer; The ViT branch includes several consecutive ViT encoders and a ViT memory queue enhancement module, each ViT encoder includes a patch merging module, a self-attention-based transformer module and a feedforward neural network module; The clustering and storage of historical image features is to use the K-means algorithm to perform anatomical structure clustering on the historical image data set, and to distribute the image features of the same type of images to the same sub-queue of the memory queue.
2. The dual-branch fusion cardiac image segmentation method based on memory queue according to claim 1, characterized in that: The steps of selecting the historical image features most similar to the heart image to be segmented from the memory queue are as follows: Predict the category of the heart image to be segmented, and select the corresponding sub-queue from the memory queue according to the predicted category; From the selected sub-queues, the historical image features that are most similar to the input features are selected using feature similarity.
3. The dual-branch fusion cardiac image segmentation method based on memory queue according to claim 1, characterized in that: The neighboring window similarity is to divide the image into several windows, and perform attention calculation on the features of the current window through similar historical features of the neighboring windows.
4. The dual-branch fusion cardiac image segmentation method based on memory queue according to claim 1, characterized in that: The fusion adopts an improved Softmax algorithm to calculate the importance weights and non-importance weights of local features in the channel dimension and align the global semantic context in the spatial dimension.
5. A dual-branch fusion cardiac image segmentation system based on a memory queue, characterized in that: include: The feature generation module is configured to: input the heart image to be segmented into the CNN-ViT dual branch to generate local features and global features enhanced by anatomical priors; The feature fusion module is configured to: fuse the enhanced local features and global features to obtain image segmentation features; The feature decoding module is configured to: decode the image segmentation features to obtain a final segmentation result; The anatomical prior enhancement is to cluster and store historical image features to construct a memory queue, select the historical image features that are most similar to the heart image to be segmented from the memory queue, and based on the historical image features, perform attention calculation on the extracted local features and global features respectively through the similarity of adjacent windows, and the local features and global features after anatomical prior enhancement; The CNN-ViT dual branch includes a CNN branch and a ViT branch connected in parallel; The CNN branch includes a plurality of consecutive CNN encoders and a CNN memory queue enhancement module, each CNN encoder includes a convolutional layer; The ViT branch includes several consecutive ViT encoders and a ViT memory queue enhancement module, each ViT encoder includes a patch merging module, a self-attention-based transformer module and a feedforward neural network module; The clustering and storage of historical image features is to use the K-means algorithm to perform anatomical structure clustering on the historical image data set, and to distribute the image features of the same type of images to the same sub-queue of the memory queue.
6. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the memory queue-based dual-branch fusion cardiac image segmentation method according to any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the memory queue-based dual-branch fusion cardiac image segmentation method according to any one of claims 1 to 6 is implemented.
8. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the dual-branch fusion cardiac image segmentation method based on the memory queue as described in any one of claims 1-4.
Citation Information
Patent Citations
Heart image segmentation method based on hierarchical space-time Transform
CN117893546A
Similarity memory prior network segmentation method and system for medical image segmentation
CN119313901A