Undersampling abdominal liver tumor segmentation method based on CNN-Transform and double-coding cascade fusion
By using the undersampling method of CNN-Transformer and dual-coding cascade fusion in abdominal liver tumor segmentation, the problem of excessive MRI data acquisition time is solved, and efficient abdominal liver tumor segmentation is achieved, which significantly improves segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510326254.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art relies on full-sampled MRI data in abdominal liver tumor segmentation, resulting in too long MRI data acquisition time, patient discomfort and motion artifacts generated, and the full-sampling process is time-consuming and affects image reconstruction.
The undersampled abdominal liver tumor segmentation method based on CNN-Transformer and dual-coding cascade fusion is adopted. Through the joint learning strategy of shared parameters, the MRI reconstruction task is aggregated with the segmentation task to improve the segmentation performance in undersampled MRI scenarios.
It significantly improves the performance of abdominal liver tumor segmentation in undersampled MRI scenarios, reduces data acquisition time, and improves segmentation accuracy and robustness.
Smart Images

Figure CN120182604A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of abdominal MRI imaging, and in particular, relates to an undersampled abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion with a joint learning mechanism based on shared parameters. Background Art
[0002] Magnetic resonance imaging (MRI), as a non-ionizing radiation imaging technique, can generate high-resolution images. However, the further development of MRI technology is currently mainly restricted by two major factors. On the one hand, due to the long acquisition time of k-space data, it often causes discomfort to patients and may lead to the generation of motion artifacts. On the other hand, in research fields such as abdominal liver tumor segmentation, the research usually relies on fully sampled abdominal MRI images. This process not only requires a large amount of time for manual annotation, but also often ignores the potential impact of the image reconstruction process on subsequent analysis tasks. Summary of the Invention
[0003] The purpose of the present invention is to introduce an undersampled abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion for the problem of the overly long acquisition time of MRI in the current segmentation that usually relies on fully sampled MRI data. This method is specifically designed for segmenting abdominal liver tumors from undersampled MRI images, and by adopting a joint learning strategy with shared parameters, it cleverly aggregates the reconstruction task and the segmentation task, significantly improving the performance of abdominal liver tumor segmentation in the undersampled MRI scenario.
[0004] To achieve the above object, the present invention includes the following steps:
[0005] 1) Obtain fully sampled multi-coil k-space data of the abdomen, perform multiple operations on the fully sampled multi-coil k-space data through coil sensitivity decoding and encoding, two-dimensional Fourier transform operator, two-dimensional inverse Fourier transform operator, and an undersampling operator that selects acquisition positions from the entire multi-coil k-space grid to obtain undersampled multi-coil k-space data, and label the abdominal liver tumors to obtain various category labels. Finally, form a training set, a validation set, and a test set with the undersampled multi-coil k-space data and the corresponding various category labels;
[0006] 2) Design a joint network model with dual-encoding cascade fusion based on the CNN-Transformer network;
[0007] 3) Construct a joint loss function Loss of the joint network model;
[0008] 4) Use the training set and validation set obtained in step 1) and the joint loss function obtained in step 3) to solve the optimal parameters of the joint network model;
[0009] 5) For the test set obtained in step 1), use the optimal parameters of the joint network model obtained in step 4) to predict the abdomen.
[0010] In the said step 1), to obtain the full-sampled multi-coil k-space data of the abdominal liver, perform multiple operations on the full-sampled multi-coil k-space data through coil sensitivity decoding and encoding, two-dimensional Fourier transform operator, two-dimensional inverse Fourier transform operator, and undersampling operator for selecting the acquisition positions from the entire multi-coil k-space grid to obtain the undersampled multi-coil k-space data, and label the abdominal liver tumors to obtain various category labels. Finally, the specific method of forming the training set, validation set, and test set from the undersampled multi-coil k-space data and the corresponding various category labels is as follows:
[0011] Obtain the full-sampled multi-coil k-space data of the abdomen from the magnetic resonance imaging instrument, and through coil sensitivity decoding S -1 and two-dimensional inverse Fourier transform operator F -1 to obtain the full-sampled multi-coil image data represents the complex number field, N = N h ×N w represents the number of pixels in the two-dimensional image, N h and N w are the height and width of the two-dimensional image respectively; subsequently, obtain the undersampled multi-coil k-space data through the undersampling operator P, coil sensitivity encoding S, and two-dimensional Fourier transform operator F is defined as Y = PFSx, c th The undersampled k-space data of the coil is expressed as c is the number of coils, C is the total number of coils, M represents the number of undersampled points of the single-coil k-space data, M << N. Finally, form the training set, validation set, and test set with the corresponding various category labels.
[0012] The joint network model in step 2) consists of a reconstruction model and a segmentation model, and adopts a joint learning strategy to segment the undersampled input tumor image to segment the specific tumor size. The reconstruction model and the segmentation model share the same U-shaped network architecture. This U-shaped network architecture introduces a novel cascaded fusion module FN to interact and fuse the image feature information obtained through different encoding networks to enhance the feature representation, a lightweight CNN-Transformer hybrid module. While comprehensively integrating the spatial interaction information, it uses the self-attention mechanism to selectively focus on the token subset, not only maintaining the integrity of the global context information but also significantly reducing the computational complexity. In addition, the Mix-Pool module is introduced to perform pooling fusion on the feature maps in the same layer, reducing noise interference while retaining the significant features of multiple poolings, and improving the generalization ability and robustness of the model;
[0013] A. Mix-Pool Module:
[0014] The input feature tensor X is symmetrically channel-separated to obtain X p and X c For X p and X c First, perform convolutions with kernel sizes of 1×1 and 3×3 respectively; for X p Then, perform channel separation again, and respectively perform max-pooling operations with a window size of 2×2 and a stride of 2, and average-pooling operations with a window size of 3×3 and a stride of 2, and finally perform feature fusion;
[0015] B. Lightweight CNN-Transformer Hybrid Module:
[0016] In the lightweight CNN-Transformer hybrid module, the convolutional part synthesizes the convolutional stream by mixing different convolutional kernels to capture the local information inside the image, and the Transformer part has been lightweight optimized, introducing sparse self-attention to capture remote feature dependencies;
[0017] C. Cascade Fusion Module FN:
[0018] The FN module consists of multi-scale convolutional kernels, sigmoid function, and relu function, and is used to efficiently integrate the image features input by two different encoding layers;
[0019] D. Joint Learning Strategy:
[0020] This method integrates MRI reconstruction and segmentation into a unified framework, and adopts a shared parameter method, enabling the model to utilize the detailed information learned during the reconstruction process to improve the segmentation performance of undersampled MRI data.
[0021] In step 3), the joint loss function is defined as consisting of two parts, including the L1 norm and the cross-entropy function; the L1 norm is one of the most commonly used metrics for evaluating the performance of deep learning-based medical reconstruction, and is expressed by the mathematical formula:
[0022]
[0023] where represents the reconstructed image, x i represents the original image, and T represents the number of images in the dataset;
[0024] The cross-entropy function is used to measure the similarity between two distributions, and is expressed by the mathematical formula:
[0025]
[0026] where, Represents the predicted probability for each class, y i Represents the true label value;
[0027] The combined loss function Loss of the network is a loss function that is the sum of the L1 norm and the cross-entropy function:
[0028] Loss = L1 + λL seg
[0029] where λ is the combined loss weight parameter.
[0030] In step 4), to solve for the optimal parameters of the combined network model, the Adam optimizer in deep learning is used. The training set and validation set generated in step 1) are used for network training and validation, and the optimal parameter set θ° is obtained by minimizing the combined loss function Loss in step 3). θ° represents the set of network parameters when the undersampling classification effect of the combined network model is optimal.
[0031] In step 5), using the test set obtained in step 1) and the optimal parameters of the combined network model obtained in step 4), network prediction of abdominal tumors is performed. The network prediction process can be expressed as:
[0032]
[0033] where f overall (·) represents the prediction process of the entire network, is the network prediction result of abdominal tumors.
[0034] Compared with the prior art, the present invention has the following outstanding technical effects:
[0035] Based on the lightweight CNN-Transformer hybrid as the basic framework, the present invention uses the cascaded fusion method of the dual coding structure to perform interactive feature fusion on the image feature information of different coding layers, realizing the deep integration and utilization of multiple information, thereby greatly improving the abdominal liver tumor segmentation performance in the undersampling MRI scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the overall network architecture diagram of the present invention;
[0037] Figure 2 is the comparison diagram of the segmentation results of the present invention and different networks. DETAILED DESCRIPTION OF THE INVENTION
[0038] The following embodiments will further illustrate the present invention in conjunction with the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0039] Embodiment 1: As Figure 1-2 shown, a method for undersampled abdominal liver tumor segmentation based on CNN-Transformer and dual-encoding cascade fusion of the present invention includes the following steps:
[0040] The first step: Obtain the fully sampled multi-coil k-space data of the abdomen, perform multiple operations on the fully sampled multi-coil k-space data through coil sensitivity decoding and encoding, two-dimensional Fourier transform operator, two-dimensional inverse Fourier transform operator, and undersampling operator for selecting acquisition positions from the entire multi-coil k-space grid to obtain undersampled multi-coil k-space data, and label the abdominal liver tumors to obtain various category labels. Finally, the specific method for forming the training set, validation set, and test set from the undersampled multi-coil k-space data and the corresponding various category labels is as follows:
[0041] Obtain the fully sampled multi-coil k-space data of the abdomen from a magnetic resonance imaging instrument, and obtain the fully sampled multi-coil image data through coil sensitivity decoding S -1 and two-dimensional inverse Fourier transform operator F -1 where represents the complex domain, N = N h ×N w represents the number of pixels in the two-dimensional image, N h and N w are the height and width of the two-dimensional image respectively; subsequently, obtain the undersampled multi-coil k-space data through the undersampling operator P, coil sensitivity encoding S, and two-dimensional Fourier transform operator F defined as Y = PFSx, c th The undersampled k-space data of the coil is expressed as c is the number of coils, C is the total number of coils, M represents the number of undersampled points of the single-coil k-space data, M << N, and finally form the training set, validation set, and test set with the corresponding various category labels.
[0042] Step 2: Design a joint network model for double-coding cascaded fusion based on the CNN-Transformer network: The joint network model consists of a reconstruction model and a segmentation model, and adopts a joint learning strategy to segment the undersampled input tumor images to obtain the specific tumor size. The reconstruction model and the segmentation model share the same U-shaped network architecture, which introduces a novel cascaded fusion module FN to interact and fuse the image feature information obtained through different coding networks to enhance the feature representation. The lightweight CNN-Transformer hybrid module comprehensively integrates the spatial interaction information and uses the self-attention mechanism to selectively focus on the token subsets, not only maintaining the integrity of the global context information but also significantly reducing the computational complexity. In addition, the Mix-Pool module is introduced to perform pooling fusion on the feature maps in the same layer, reducing noise interference while retaining the significant features of multiple poolings, and improving the generalization ability and robustness of the model;
[0043] A. Mix-Pool module:
[0044] The input feature tensor X is symmetrically channel-separated to obtain X p 、X c , for X p 、X c First, perform convolutions with kernel sizes of 1×1 and 3×3 respectively; X p Then perform channel separation again, and perform max-pooling operations with window sizes of 2×2 and a stride of 2 and average-pooling operations with window sizes of 3×3 and a stride of 2 respectively, and finally perform feature fusion;
[0045] B. Lightweight CNN-Transformer hybrid module:
[0046] The lightweight CNN-Transformer hybrid module captures the local information inside the image by synthesizing the convolution stream through mixing different convolution kernels in the convolution part, and the Transformer part is lightweight optimized to introduce sparse self-attention to capture the remote feature dependence information;
[0047] C. Cascaded fusion module FN:
[0048] The FN module consists of multi-scale convolution kernels, sigmoid functions, and relu functions, and is used to efficiently integrate the image features input from two different coding layers;
[0049] D. Joint learning strategy:
[0050] This method integrates MRI reconstruction and segmentation into a unified framework and adopts a shared parameter method, enabling the model to utilize the detailed information learned during the reconstruction process to improve the segmentation performance of undersampled MRI data.
[0051] Step 3: Construct the combined loss function Loss of the combined network model: The combined loss function is defined as consisting of two parts, including the L1 norm and the cross-entropy function; the L1 norm is one of the most commonly used metrics for evaluating the performance of medical reconstruction based on deep learning, and is expressed by the mathematical formula:
[0052]
[0053] where represents the reconstructed image, x i represents the original image, and T represents the number of images in the dataset;
[0054] The cross-entropy function is used to measure the similarity between two distributions, and is expressed by the mathematical formula:
[0055]
[0056] where represents the predicted probability of each class, and y i represents the true label value;
[0057] The combined loss function Loss of the network is the loss function that is the sum of the L1 norm and the cross-entropy function:
[0058] Loss = L1 + λL seg
[0059] where λ is the combined loss weight parameter.
[0060] Step 4: Solve for the optimal parameters of the combined network model. Use the Adam optimizer in deep learning, and use the training set and validation set generated in step 1) to train and validate the network. The optimal parameter set θ° is obtained by minimizing the combined loss function Loss in step 3). θ° represents the set of network parameters when the undersampling classification effect of the combined network model is optimal.
[0061] Step 5: Use the test set obtained in step 1) and the optimal parameters of the combined network model obtained in step 4) to perform network prediction for abdominal liver tumors. The network prediction process can be expressed as:
[0062]
[0063] where f overall (·) represents the prediction process of the entire network, is the network prediction result for abdominal tumors.
[0064] The Unet, Transunet, and the method proposed in this invention were tested on the test dataset, and the segmentation results are shown in Table 1. Compared with the comparison methods Unet and Transunet, the average Dice value of the segmentation of the method of this invention increased by 8.00% and 6.77% respectively, the segmentation accuracy of the liver increased by 5.48% and 7.63% respectively, and the accuracy of tumor segmentation reached a higher level, increasing by 18.39% and 12.60% respectively. The experimental results prove that this invention has greatly improved the segmentation accuracy of undersampled abdominal liver tumors and has great potential in disease prediction.
[0065] Table 1 Comparison of segmentation effects of different segmentation methods
[0066]
[0067] Figure 2 For the segmentation result effect diagram. Among them, (a) is the undersampled magnetic resonance image, (b) is the ground truth label, (c) is the segmentation result diagram of the network of this invention, (d) is the segmentation network result diagram of Unet, and (e) is the segmentation result diagram of Transunet. By comparing the result diagrams, it can be more accurately found that the method of this invention is superior to the other two methods in the processing of segmentation details, demonstrating the excellent performance of the combined network of the method of this invention and having great potential in the field of medical disease segmentation.
[0068] This invention is based on a lightweight CNN-Transformer hybrid as the basic framework, and uses the method of cascading and fusing double coding structures to perform interactive feature fusion on the image feature information of different coding layers, realizing the deep integration and utilization of multiple information, thus greatly improving the abdominal tumor segmentation performance in the undersampled MRI scenario.
[0069] The above is only the specific implementation manner of this invention, but the protection scope of this invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by this invention should be covered within the protection scope of this invention.
Claims
1. An undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion, characterized by: The steps include: 1) Obtaining full-sampled multi-coil k-space data of the abdomen, performing multiple operations on the full-sampled multi-coil k-space data through coil sensitivity decoding and encoding, a two-dimensional Fourier transform operator, a two-dimensional Fourier inverse transform operator, and an undersampling operator that selects acquisition positions from the entire multi-coil k-space grid to obtain undersampled multi-coil k-space data, and annotating the abdominal liver tumor to obtain various category labels, and finally forming the undersampled multi-coil k-space data and the corresponding various category labels into a training set, a validation set, and a test set; 2) Design a joint network model based on dual-encoding cascade fusion of CNN-Transformer network; 3) Construct the joint loss function Loss of the joint network model; 4) Using the training set and validation set obtained in step 1) and the joint loss function Loss obtained in step 3), solve the optimal parameters of the joint network model; 5) For the test set obtained in step 1), the optimal parameters of the joint network model obtained in step 4) are used to predict abdominal liver tumors.
2. The undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion according to claim 1, characterized in that: The specific method of step 1) is: The full-sample multi-coil k-space data of the abdomen are acquired from the magnetic resonance imaging instrument, and the S -1 and the two-dimensional inverse Fourier transform operator F -1 Get fully sampled multi-coil image data Represents the complex field, N = N h ×N w Represents the number of pixels in a two-dimensional image, N h and N w are the height and width of the two-dimensional image respectively; then the under-sampling multi-coil k-space data is obtained through the under-sampling operator P, coil sensitivity encoding S and two-dimensional Fourier transform operator F Defined as Y = PFSx,c th The undersampled k-space data of the coil is expressed as c=1,2,...,C, c is the number of coils, C is the total number of coils, M represents the number of under-sampling points of single coil k-space data, M<<N, and finally the training set, validation set and test set are formed with the corresponding category labels.
3. The undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion according to claim 1, characterized in that: In the step 2), the joint network model is composed of a reconstruction model and a segmentation model, and a joint learning strategy is used to segment the undersampled input tumor image to segment the specific tumor size. The reconstruction model and the segmentation model share the same U-shaped network architecture. The U-shaped network architecture introduces a novel cascade fusion module FN, which interactively fuses the image feature information obtained through different encoding networks to enhance the feature representation. The lightweight CNN-Transformer hybrid module uses a self-attention mechanism to selectively focus on a token subset while fully integrating spatial interaction information. In addition, a Mix-Pool module is introduced to perform pooling fusion on feature maps in the same layer. A.Mix-Pool module: The input feature tensor X is symmetrically separated into channels to obtain X p , X c , for X p , X c First, convolution kernel sizes are 1×1 and 3×3 respectively; p Channel separation is performed again, and the maximum pooling operation with a window size of 2×2 and a stride of 2 and the average pooling operation with a window size of 3×3 and a stride of 2 are performed respectively, and finally feature fusion is performed; B. Lightweight CNN-Transformer hybrid module: Lightweight CNN-Transformer hybrid module: The convolution part captures the local information inside the image by mixing different convolution kernels to synthesize the convolution flow. The Transformer part is lightweight optimized and introduces sparse self-attention to capture long-range feature dependency information. C. Cascade fusion module FN: The FN module consists of multi-scale convolution kernels, sigmoid functions, and relu functions, which are used to integrate the image features input from two different encoding layers; D. Joint learning strategy: We integrate MRI reconstruction and segmentation into a unified framework using a shared parameter approach, which enables the model to exploit detailed information learned during reconstruction to improve the segmentation performance of undersampled MRI data.
4. The undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion according to claim 1, characterized in that: The joint loss function Loss in step 3) is defined to consist of two parts, including the L1 norm and the cross entropy function; the L1 norm is expressed in a mathematical formula as follows: in represents the reconstructed image, x i represents the original image, T represents the number of images in the dataset; The cross entropy function is used to measure the similarity between two distributions and is expressed as follows: in, represents the predicted probability of each class, y i Represents the true label value; The joint loss function Loss is the loss function of the sum of the L1 norm and the cross entropy function: Loss=L1+λL seg Among them, λ is the joint loss weight parameter.
5. The undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion according to claim 1, characterized in that: In step 4), the optimal parameters of the joint network model are solved by using the Adam optimizer in deep learning, and the network is trained and verified using the training set and verification set generated in step 1). The optimal parameter set θ° is obtained by minimizing the joint loss function Loss in step 3), where θ° represents the set of network parameters when the undersampling segmentation effect of the joint network model is optimal.
6. The undersampling abdominal liver tumor segmentation method based on CNN-Transformer and dual-encoding cascade fusion according to claim 5, characterized in that: Using the test set obtained in step 1), and the optimal parameters of the joint network model obtained in step 4), network prediction of abdominal tumors is performed. The network prediction process can be expressed as: Among them, f overall (·) represents the prediction process of the entire network, Y is the undersampled multi-coil k-space data, Network prediction results for abdominal tumors.