Semi-supervised cardiac image segmentation method based on Mamba-Transform double structure and contrast learning
By adopting the semi-supervised method of Mamba-Transformer dual structure and contrast learning in cardiac image segmentation, the problem that the existing technology of central heart image segmentation is difficult to achieve high accuracy and robustness is solved, reducing the dependence on labeled data, and achieving a more efficient and accurate cardiac image segmentation effect.
Patent Information
- Application Number
- CN202510229500.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to achieve high accuracy and robustness in cardiac image segmentation, and at the same time, it is costly and time-consuming to obtain high-quality labeled data, resulting in large-scale labeled data acquisition becoming a bottleneck.
A semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrast learning is adopted. By constructing teacher models and student models, the memory mechanism and contrast learning strategy are used to reduce the dependence on large amounts of labeled data.
A more refined and accurate heart image segmentation result is achieved, which improves the robustness and computing efficiency of the model, significantly improves the segmentation accuracy, and reduces the dependence on labeled data.
Smart Images

Figure CN120147330A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of medical image processing, and specifically is a semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrast learning. Background Art
[0002] Cardiovascular disease has become one of the major diseases threatening human health. Cardiac function assessment plays a vital role in the early diagnosis, formulation of treatment plans and efficacy monitoring of cardiovascular disease, and the structure and morphology of the heart are important bases for evaluating cardiac function. Cardiac CT images provide effective cardiac physiological characteristics and anatomical information. However, the complex anatomical structure of the heart, morphological changes caused by different diseases, and image noise and tissues with high similarity (such as lungs and blood vessels) all increase the difficulty of the heart segmentation task.
[0003] Traditional image segmentation networks, such as convolutional neural networks (CNN) and ViT (Vision Transformer), have achieved certain results in medical image segmentation, but they still cannot meet the requirements of high accuracy and robustness for clinical applications. In addition, high-quality medical image annotation requires experienced experts, and the annotation process is time-consuming and costly, so obtaining large-scale high-quality annotated data has become a bottleneck. Summary of the invention
[0004] In view of the deficiencies in the prior art, the technical problem that the present invention intends to solve is to provide a semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning.
[0005] The present invention solves the technical problem by adopting the following technical solution:
[0006] A semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning, comprising the following steps:
[0007] Step 1: Obtain the data set and preprocess it;
[0008] Step 2: Construct a teacher model and a student model, both of which include an MTSeg encoder and a VNet decoder. The input image enters the MTSeg encoder after a linear layer and position embedding operation. The MTSeg encoder contains multiple MTSeg encoding modules, and the VNet decoder contains multiple decoding modules. The output features of the MTSeg encoding module are projected and then jump-connected with the corresponding decoding modules. The output feature vector of the VNet decoder is normalized to obtain the segmentation result.
[0009] The MTSeg encoding module includes an improved Mamba branch and a Transformer branch. The output feature vectors of the two branches are fused through a cross-attention layer to obtain the output feature vector of the MTSeg encoding module. The improved Mamba branch includes a cascaded MS module and DRFB module. In the MS module, the input feature vector passes through a linear layer and then enters two branches. In one branch, it sequentially passes through a linear layer, an FFT module, and a SiLU activation layer. In the other branch, it sequentially passes through a linear layer, an ACFF module, an SS2D module, and a linear layer. After the output feature vectors of the two branches are element-wise multiplied, they are element-wise added to the input feature vector of the MS module. The resulting feature vector sequentially passes through layer normalization, a multi-layer perceptron, and a ReLu activation layer to obtain the output feature vector of the MS module.
[0010] In the FFT module, the input feature vector is transformed to the frequency domain through Fourier transform to obtain a frequency-domain feature vector. A frequency selection mask is generated, and the frequency-domain feature vector is multiplied by the frequency selection mask to obtain a low-frequency feature vector. The remaining part of the frequency-domain feature vector is the high-frequency feature vector. After the low-frequency feature vector and the high-frequency feature vector are weighted and fused, they are then activated to obtain the output feature vector of the FFT module.
[0011] In the ACFF module, the input feature vector enters two branches. In one branch, it sequentially passes through a convolutional layer, a fully connected layer, and a GELU activation layer. In the other branch, it sequentially passes through a convolutional layer and a SiLu activation layer. The output feature vectors of the two branches are element-wise added to obtain the output feature vector of this module.
[0012] In the DRFB module, the input feature vector enters two branches. In one branch, a convolutional operation is performed. In the other branch, it is element-wise multiplied by a filter weight vector generated by a learnable filter. After the output feature vectors of the two branches are element-wise added, they sequentially pass through layer normalization and a ReLu activation function to obtain the output feature vector of this module.
[0013] Step 3: Use the memory bank mechanism to jointly train the teacher model and the student model, and use the trained student model as the segmentation model for three-dimensional cardiac image segmentation.
[0014] Furthermore, the MTSeg encoder includes 8 MTSeg encoding modules, and the VNet decoder includes 4 decoding modules. The output features of the second, fourth, sixth, and eighth MTSeg encoding modules are projected and then skip-connected to the corresponding decoding modules.
[0015] Further, for the Transformer branch, after the input feature vector passes through the normalization and multi-head attention layers, it is connected with itself through a residual connection to obtain Feature A; after Feature A passes through the normalization and multi-layer perceptron, it is connected with itself through a residual connection to obtain the output feature vector of the Transformer branch.
[0016] Further, the decoding module includes two upsampling layers and a 3D convolutional layer connected in sequence.
[0017] Further, the model training process is divided into three stages:
[0018] The first stage: Use the student model to perform supervised learning on the labeled data and calculate the supervised loss; for the unlabeled data, use the teacher model to generate pseudo-labels and calculate the unsupervised loss;
[0019] The second stage: Use the teacher model to extract feature vectors from the labeled data, screen the feature vectors, and store the feature vectors that meet the prediction accuracy requirements in the memory bank;
[0020] The third stage: Use the feature vectors in the memory bank for contrastive learning, that is, input the training set into the student model for feature extraction. The output feature vector of the student model decoder passes through two multi-layer perceptrons to obtain the feature vector set P. The process is expressed as:
[0021] P = g θ (q θ (f θ -(x))) (2)
[0022] In the formula, x represents the training set, f θ -(x) represents the output feature vector of the student model decoder, and g θ , q θ represent multi-layer perceptrons;
[0023] Divide the feature vectors in the feature vector set P into various category subspaces according to the category to obtain the feature vector sets P 1 ,..., P c ,..., P J , P c = {p c} represents the feature vector set of category c, p c represents the feature vector of category c, and J represents the number of categories; let Z c = {z c} represents the feature vector set of category c stored in the memory bank, and z c represents the feature vector of category c stored in the memory bank; calculate the feature vector p according to the following formula c and z cCosine similarity between:
[0024]
[0025] Calculate the feature vector p according to the following formula c and z c Distance between:
[0026]
[0027] In the formula, Represents the attention weight of the feature vectors p c and z c The calculation methods of both are the same, where Is calculated by the following formula:
[0028]
[0029] In the formula, Represents the number of feature vectors of the set P c , S c,θ (·) represents the attention mechanism, p i Represents the feature vector formed by the i-th dimension of all feature vectors in the set P c ;
[0030] Calculate the contrastive loss according to the following formula:
[0031]
[0032] In the formula, Represents the number of feature vectors of the set Z c .
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] 1. The MTseg encoder integrates the advantages of the Transformer network in capturing long-range dependencies and global information, as well as the powerful capabilities of the Mamba network in multi-scale and local information modeling, realizing the effective integration of long-range dependencies and global information, multi-scale and local information, improving the feature extraction ability of the model for complex images, and thus obtaining more refined and accurate segmentation results. By introducing the cross-attention mechanism, the model can effectively fuse the feature information from different branches, using the features of one branch as queries and the features of the other branch as keys and values, thereby enhancing the interaction between local and global information and improving the accuracy of feature fusion and the overall performance of the model. The MTseg encoding module enables information transfer and integration between different feature spaces through cross-modal feature fusion, further enhancing the model's ability to capture complex features and improving its performance in the segmentation task.
[0035] In the improved Mamba branch, the feature vectors of different frequency components are fused in the frequency domain through the FFT module, optimizing the feature representation, enhancing the model's robustness to noise and its ability to capture detailed information. The ACFF module enhances the model's feature expression and computational efficiency through multi-path feature fusion and non-linear activation, especially suitable for processing high-dimensional data.
[0036] 2. The VNet decoder has strong 3D data modeling capabilities. Using three-dimensional convolutions can more efficiently process the spatial information of three-dimensional cardiac images. Compared with the traditional U-Net, VNet better maintains the spatial consistency between voxels and is more suitable for the segmentation task of three-dimensional medical images.
[0037] 3. To reduce the dependence on a large amount of labeled data, a memory bank mechanism is introduced. The high-confidence features extracted by the teacher model are stored in the memory bank. The contrastive learning strategy can better utilize labeled and unlabeled data, significantly improving the segmentation accuracy, accelerating the model convergence, and enhancing the intra-class consistency and inter-class distinguishability. Description of the Drawings
[0038] Figure 1 Structural diagrams of the student model and the teacher model;
[0039] Figure 2 Framework diagram of semi-supervised training. Detailed Implementation Manner
[0040] Specific embodiments are given below in conjunction with the drawings. The specific embodiments are only used to introduce the technical solutions of the present invention in detail and do not limit the protection scope of this application.
[0041] The present invention provides a semi-supervised cardiac image segmentation method based on the Mamba-Transformer dual structure and contrastive learning (referred to as the method, see Figures 1 - 2 ), including the following steps:
[0042] Step 1: Obtain a number of three-dimensional cardiac images as a data set, and preprocess the data set, including adjusting the resolution and data augmentation;
[0043] Step 2: Construct a teacher model and a student model. Both models adopt the same network structure, which includes an MTSeg encoder and a VNet decoder. The input image generates a feature vector through a linear layer and a positional embedding operation, and this feature vector is input into the MTSeg encoder for encoding. The MTSeg encoder contains multiple MTSeg encoding modules, and the VNet decoder contains multiple decoding modules. After the output features of the MTSeg encoding module pass through a projection operation, they are skip-connected to the corresponding decoding modules. This embodiment contains 8 MTSeg encoding modules and 4 decoding modules. After the output features of the second, fourth, sixth, and eighth MTSeg encoding modules pass through a projection operation, they are skip-connected to the corresponding decoding modules. The output feature vector of the VNet decoder undergoes a normalization operation to obtain the segmentation result.
[0044] The MTSeg encoding module contains an improved Mamba branch and a Transformer branch, forming a Mamba-Transformer dual-branch structure. The output feature vectors of the improved Mamba branch and the Transformer branch are fused through a cross-attention layer to obtain the output feature vector of the MTSeg encoding module. The MTSeg encoding module realizes the effective fusion of global information and local detail information, overcoming the deficiencies of a single branch in long-range dependency modeling and local detail information capture. The improved Mamba branch includes a cascaded MS module and a dynamic residual filtering module (DRFB). The MS module is used to enhance the feature extraction and computational efficiency of the model, and the dynamic residual filtering module DRFB is used to enhance the expressive power and training stability of the model. In the MS module, the input feature vector enters two branches after passing through a linear layer. In one branch, it sequentially passes through a linear layer, a feature optimization module (FFT), and a SiLU activation layer. In the other branch, it sequentially passes through a linear layer, an adaptive convolutional feature fusion module (ACFF), an SS2D module, and a linear layer. The output feature vectors of the two branches are element-wise multiplied and then element-wise added to the input feature vector of the MS module. The resulting feature vector passes through layer normalization, a multi-layer perceptron MLP, and a ReLu activation layer in sequence to obtain the output feature vector of the MS module.
[0045] In the feature optimization module FFT, the input feature vector is transformed to the frequency domain through Fourier transform to obtain the frequency-domain feature vector; a frequency selection mask is generated, and the frequency-domain feature vector is multiplied by the frequency selection mask to obtain the low-frequency feature vector, and the remaining part of the frequency-domain feature vector is the high-frequency feature vector; after weighted fusion of the low-frequency feature vector and the high-frequency feature vector, and then through activation processing, the output feature vector of the feature optimization module is obtained. The low-frequency feature vector usually contains the main structural information of the image, while the high-frequency feature vector contains detail and edge information. Therefore, fusing the high- and low-frequency feature vectors helps capture the detail information in the image, optimizes the representation of the input data, enables the model to better extract and learn features, and thus improves the segmentation accuracy.
[0046] In the ACFF module, the input feature vector enters two branches. In one branch, it sequentially passes through a convolutional layer, a fully connected layer, and a GELU activation layer. In the other branch, it sequentially passes through a convolutional layer and a SiLu activation layer. The output feature vectors of the two branches are added element-wise to obtain the output feature vector of this module.
[0047] In the DRFB module, the input feature vector enters two branches. In one branch, convolutional operations are performed. In the other branch, it is multiplied element-wise with the filter weight vector generated by a learnable filter. After the output feature vectors of the two branches are added element-wise, they sequentially pass through layer normalization and the ReLu activation function to obtain the output feature vector of this module. The initial values of the filter weight vectors are all 1.
[0048] For the Transformer branch, the input feature vector passes through normalization and a multi-head attention layer, and then a residual connection is made with itself to obtain feature A; feature A passes through normalization and a multi-layer perceptron, and then a residual connection is made with itself to obtain the output feature vector of the Transformer branch.
[0049] A cross-attention mechanism is introduced in the MTSeg encoding module, aiming to better fuse the output feature vectors of the improved Mamba branch and the Transformer branch. The cross-attention mechanism uses the multi-head attention mechanism. By taking the output feature vector of one branch as the query, and the output feature vectors of the other branch as the key and value, it effectively enhances the interaction between local information and global information.
[0050] The decoding module contains two upsampling layers and a 3D convolutional layer connected in sequence; different from the traditional U-Net, the VNet decoder introduces 3D convolutional operations, which can make full use of the spatial information of three-dimensional images.
[0051] Step 3: jointly train the teacher model and the student model using the memory bank mechanism, and use the trained student model as a segmentation model for segmenting three-dimensional heart images;
[0052] The model training process is divided into three stages: In the first stage, the student model is used to supervise the labeled data and calculate the supervision loss L label ; For unlabeled data, use the teacher model to generate pseudo labels and calculate the unsupervised loss L unlabel , used to guide the optimization of the student model; supervision loss L label and the unsupervised loss L unlabel They all use a combination of dice loss and cross entropy loss, calculated in voxel form by the following formula:
[0053]
[0054] In the formula, I represents the number of voxels, J represents the number of categories, and Y i,j and G i,j Represents the probability and one-hot encoding truth value of the j-th class at voxel i;
[0055] The second stage: Use the teacher model to extract feature vectors from the labeled data, screen the feature vectors based on the prediction accuracy, and store the feature vectors that meet the requirements into the memory bank. The memory bank is used to store feature vectors of different categories. Combined with the comparative learning strategy, the memory bank effectively enhances the feature differentiation ability by shortening the distance between feature vectors of the same category and increasing the distance between feature vectors of different categories, significantly improving the segmentation accuracy of the model;
[0056] The third stage: Contrastive learning is performed based on the memory bank to optimize the consistency of intra-class features and the separability of inter-class features; the training set is input into the student model for feature extraction, and the output feature vector of the student model decoder passes through two multi-layer perceptrons to obtain the feature vector set P. The process is expressed as:
[0057] P = g θ (q θ (f θ -(x))) (2)
[0058] In the formula, x represents the training set, f θ -(x) represents the output feature vector of the student model decoder, g θ ,q θ represents a multi-layer perceptron;
[0059] Divide the feature vectors in the feature vector set P into each category subspace according to the category, and obtain the feature vector set P of each category 1 ,...,P c ,...,P J, P c = {p c} represents the set of feature vectors of class c, and p c represents the feature vector of class c; Z c = {z c} represents the set of feature vectors of class c stored in the memory bank, and z c represents the feature vector of class c stored in the memory bank; Calculate the cosine similarity between the feature vector p c and z c according to the following formula:
[0060]
[0061] Calculate the distance between the feature vector p c and z c according to the following formula:
[0062]
[0063] In the formula, represents the attention weights of the feature vectors p c and z c , and the calculation methods of both are the same, where is calculated by the following formula:
[0064]
[0065] In the formula, represents the number of feature vectors in the set P c , S c,θ (·) represents the attention mechanism, and p i represents the feature vector formed by the i-th dimension of all feature vectors in the set P c ;
[0066] Calculate the contrastive loss L contr according to the following formula:
[0067]
[0068] In the formula, represents the number of feature vectors in the set Z c .
[0069] Improve the discriminative ability of the model through the contrastive loss L contr , realize the collaborative learning of the student-teacher model on labeled and unlabeled data, make full use of the unlabeled data, and improve the performance and generalization ability of the model.
[0070] Based on the MMWHS dataset, the ratio of labeled data to unlabeled data is 10% and 90%. Using the Dice similarity coefficient (DSC) as the evaluation metric, a comparison is made with existing mainstream models. The segmentation results are shown in Table 1.
[0071] Table 1 Segmentation Results of Different Models on the MMWHS Dataset
[0072]
[0073] As shown in Table 1, the segmentation model of the present invention demonstrates significant advantages in the segmentation task of five parts of the heart (two ventricles, two atria, and myocardium). Compared with existing mainstream models, the average DSC of the five parts achieves performance improvements of 5.86%, 9.03%, and 0.6% respectively, verifying the effectiveness of the segmentation model. Thus, it can be seen that the MTSeg encoder with a double-branch structure can effectively improve the segmentation accuracy of the model. In particular, the improved Mamba branch can better capture local information, and the cross-attention mechanism is used to complement the global attention of the Transformer branch, enhancing the model's feature extraction ability. By using the method of contrastive learning, it can better extract the features of cardiac images and enhance the segmentation ability of the model.
[0074] The sources of the above existing mainstream models are as follows:
[0075] 1. Luo X, Liao W, Chen J, et al. Efficient semi-supervised gross target volume of nasopharyngeal carcinoma segmentation via uncertainty rectified pyramid consistency[C] / / Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part II 24. Springer International Publishing, 2021: 318-329.
[0076] 2.Wang X,Wu Z,Lian L,et al.Debiased learning from naturally imbalanced pseudo-labels[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2022:14647-14657.
[0077] 3.Wang H,Li X.Dhc:Dual-debiased heterogeneous co-training framework for class-imbalanced semi-supervised medical image segmentation[C] / / International conference on medical image computing and computer-assisted intervention.Cham:Springer Nature Switzerland,2023:582-591.
[0078] Matters not described in the present invention are applicable to the prior art.
Claims
1. A semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning, characterized in that: The following steps are involved: Step 1: Obtain the data set and preprocess it; Step 2: Construct a teacher model and a student model, both of which include an MTSeg encoder and a VNet decoder. The input image enters the MTSeg encoder after a linear layer and position embedding operation. The MTSeg encoder contains multiple MTSeg encoding modules, and the VNet decoder contains multiple decoding modules. The output features of the MTSeg encoding module are projected and then jump-connected with the corresponding decoding modules. The output feature vector of the VNet decoder is normalized to obtain the segmentation result. The MTSeg encoding module includes an improved Mamba branch and a Transformer branch, and the output feature vectors of the two branches are fused through a cross attention layer to obtain the output feature vector of the MTSeg encoding module; the improved Mamba branch includes a serially connected MS module and a DRFB module, in the MS module, the input feature vector passes through a linear layer and then enters the two branches, in one branch it passes through a linear layer, an FFT module and a SiLU activation layer in sequence, and in the other branch it passes through a linear layer, an ACFF module, an SS2D module and a linear layer in sequence, the output feature vectors of the two branches are multiplied element by element, and then added element by element with the input feature vector of the MS module, and the feature vector obtained by the addition passes through layer normalization, a multi-layer perceptron and a ReLu activation layer in sequence to obtain the output feature vector of the MS module; In the FFT module, the input feature vector is converted to the frequency domain through Fourier transform to obtain the frequency domain feature vector; a frequency selection mask is generated, and the frequency domain feature vector is multiplied by the frequency selection mask to obtain the low-frequency feature vector, and the remaining part of the frequency domain feature vector is the high-frequency feature vector; the low-frequency feature vector and the high-frequency feature vector are weightedly fused and then activated to obtain the output feature vector of the FFT module; In the ACFF module, the input feature vector enters two branches. In one branch, it passes through the convolution layer, the fully connected layer, and the GELU activation layer in sequence. In the other branch, it passes through the convolution layer and the SiLu activation layer in sequence. The output feature vectors of the two branches are added element by element to obtain the output feature vector of the module. In the DRFB module, the input feature vector enters two branches, performs convolution operation in one branch, and multiplies element-by-element with the filter weight vector generated by the learnable filter in the other branch. The output feature vectors of the two branches are added element-by-element, and then pass through layer normalization and ReLu activation function in turn to obtain the output feature vector of the module; Step 3: Use the memory bank mechanism to jointly train the teacher model and the student model, and use the trained student model as a segmentation model for the segmentation of three-dimensional heart images.
2. The semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning according to claim 1, characterized in that: The MTSeg encoder includes 8 MTSeg encoding modules, and the VNet decoder includes 4 decoding modules. The output features of the second, fourth, sixth, and eighth MTSeg encoding modules are projected and then jump-connected with the corresponding decoding modules.
3. The semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning according to claim 1 or 2, characterized in that: For the Transformer branch, the input feature vector is normalized and passed through a multi-head attention layer, and then residually connected to itself to obtain feature A; feature A is normalized and passed through a multi-layer perceptron, and then residually connected to itself to obtain the output feature vector of the Transformer branch.
4. The semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning according to claim 3 is characterized in that: The decoding module includes two upsampling layers and a 3D convolutional layer connected in sequence.
5. The semi-supervised cardiac image segmentation method based on Mamba-Transformer dual structure and contrastive learning according to claim 1, characterized in that: The model training process is divided into three stages: Phase 1: Use the student model to perform supervised learning on labeled data and calculate the supervised loss; for unlabeled data, use the teacher model to generate pseudo labels and calculate the unsupervised loss; The second stage: use the teacher model to extract feature vectors from the labeled data, screen the feature vectors, and store the feature vectors that meet the prediction accuracy requirements into the memory bank; The third stage: Use the feature vectors in the memory bank for comparative learning, that is, input the training set into the student model for feature extraction. The output feature vector of the student model decoder passes through two multi-layer perceptrons to obtain the feature vector set P. The process is expressed as: P=g θ (q θ (f θ -(x))) (2) In the formula, x represents the training set, f θ -(x) represents the output feature vector of the student model decoder, g θ ,q θ represents a multi-layer perceptron; The feature vectors in the feature vector set P are divided into subspaces of each category according to their categories, and the feature vector sets P1,...,P of each category are obtained. c ,...,P J , P c ={p c } represents the feature vector set of category c, p c represents the feature vector of category c, J represents the number of categories; let Z c ={z c } represents the set of feature vectors of category c stored in the memory bank, z c Represents the feature vector of category c stored in the memory bank; the feature vector p is calculated according to the following formula c With z c The cosine similarity between: The eigenvector p is calculated according to the following formula c With z c Distance between: In the formula, Denotes the feature vector p c and z c The attention weights are calculated in the same way, where Calculated by the following formula: In the formula, Denotes the set P c The number of eigenvectors, S c, θ(·) represents the attention mechanism, p i Denotes the set P c The eigenvector formed by the i-th dimension of all eigenvectors in; The contrast loss is calculated according to the following formula: In the formula, Represents the set Z c The number of eigenvectors.
Citation Information
Patent Citations
Transform and CNN interaction-based semi-supervised medical image segmentation method
CN116258695A
Cardiac MRI segmentation method and system based on U-Net and Transform fusion improvement
CN116823850A
Semi-supervised medical image segmentation method based on collaborative comparative learning and mixed disturbance
CN117975017A
Semantic segmentation method based on Transform model
CN118608787A
Unet medical image segmentation method based on multi-attention mechanism improvement
CN119399228A
Cited By
Target-oriented video semantic communication system based on visual model
CN120529084A
Unmanned aerial vehicle GPS spoofing attack detection method and device based on comparative learning and Mama
CN120559680A
UAV GPS spoofing attack detection method and equipment based on contrastive learning and Mamba
CN120559680B
Semi-supervised semantic segmentation system, method and equipment and storage medium
CN121392839A
A semi-supervised semantic segmentation system, method, device and storage medium
CN121392839B