A VMAT Dose Verification Method and System Based on the Swin Transformer Network
By adopting a Swin Transformer network-based method in VMAT dose verification, multi-scale features are extracted using encoder and decoder, and the deep network extraction capability is enhanced through improved bottleneck parts, the problems of low manual verification efficiency and poor prediction effect in the existing technology are solved, and more efficient and more accurate dose verification prediction is achieved.
Patent Information
- Application Number
- CN202510443346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the prior art, artificial three-dimensional dose verification is inefficient, deep learning models based on convolution/pooling operations have poor dose prediction effect, and are insensitive to changes in three-dimensional dose gradients.
The VMAT dose verification method based on the Swin Transformer network is adopted to extract multiple-scale shallow global correlation features through the encoder. The improved bottleneck part uses the last layer of ResNet for deep network feature extraction, and deep feature extraction is performed through the decoder. Finally, feature fusion is performed through the jump connection part to improve prediction accuracy.
It effectively improves the prediction effect of VMAT dose verification, enhances the sensitivity to changes in three-dimensional dose gradients, and solves the problems of low manual verification efficiency and poor prediction effect of deep learning models.
Smart Images

Figure CN119963952B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a VMAT dose verification method and system based on a Swin Transformer network. Background Art
[0002] With the progress of modern technology, radiotherapy has gradually become an indispensable important means in the treatment of malignant tumors. Volumetric modulated arc therapy (VMAT) has become the core technology of precise radiotherapy by precisely coordinating the dynamic multi-leaf collimator and gantry rotation, while improving the dose conformity of the target area and effectively protecting the organs at risk.
[0003] However, the high complexity of the volumetric modulated arc therapy (VMAT) technology poses severe challenges to its dose verification: First, the efficiency of manual three-dimensional dose verification is seriously out of touch with clinical needs and it is difficult to meet the daily clinical verification needs of dozens of cases; Second, currently, deep learning models generally using convolutional / pooling operations (CP) and using two-dimensional or one-dimensional information input to predict dose verification. Due to the loss of spatial dose features and the locality of CP, these models have limitations in the long-term dependence of dose verification prediction, and the prediction effect is not ideal. The lack of sensitivity of this type of method to three-dimensional dose gradient changes makes it difficult to effectively identify potential clinical errors. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a VMAT dose verification method and system based on a Swin Transformer network, aiming to solve the technical problems of low efficiency of manual three-dimensional dose verification, poor prediction effect of dose verification based on deep learning models using convolutional / pooling operations, and insensitivity to three-dimensional dose gradient changes in the prior art.
[0005] On the one hand, the present invention provides a VMAT dose verification method based on a Swin Transformer network. The VMAT dose verification method based on a Swin Transformer network includes:
[0006] Obtain the CT images, corresponding radiotherapy dose images, and VMAT dose verification images of historical users, and construct a VMAT dose verification model based on a Swin Transformer network. The VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a skip connection part;
[0007] Preprocess the input CT images and corresponding radiotherapy dose images of the target user and segment them into several patches;
[0008] Extract multi-scale shallow global correlation features from the patch input encoder part. The encoder part includes several levels of feature extraction groups;
[0009] Input the encoded feature map output by the last-level feature extraction group into the improved bottleneck part to obtain the bottleneck feature map. The improved bottleneck part includes replacing consecutive Swin Transformer Blocks with the last layer of the ResNet network;
[0010] Input the bottleneck feature map into the decoder part for deep feature extraction. The decoder part includes several levels of feature fusion groups arranged in a stacked manner;
[0011] Input the different-scale shallow global correlation features extracted by each level of feature extraction group in the encoder part into the corresponding-level feature fusion group in the decoder part through the skip connection part for multi-scale feature fusion to obtain the fused feature map;
[0012] Input the fused feature map into the patch projection layer to output the VMAT dose verification prediction map.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: Through a VMAT dose verification method based on the Swin Transformer network provided by the present invention, the prediction effect can be effectively improved. Specifically, the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a skip connection part. Multi-size shallow global correlation features are extracted through the encoder part, and then deep network feature extraction is performed through the improved bottleneck part. Then, deep feature extraction is performed through the decoder part. Finally, the multi-size shallow global correlation features of the encoder are fused with the deep features of the decoder part through the skip connection part and output, improving feature extraction, thereby improving the prediction accuracy. The improved bottleneck part includes replacing consecutive Swin Transformer Blocks with the last layer of the ResNet network, adopting the characteristics of the residual network to enhance the extraction ability in the deep network, and further improving the prediction accuracy. In addition, the improved bottleneck part reduces the network parameters, can effectively improve the performance of the model, and thus improves the prediction accuracy, thereby solving the technical problems of generally low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolutional / pooling operations, and insensitivity to three-dimensional dose gradient changes.
[0014] According to one aspect of the above technical solution, the last layer network of the ResNet includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence. The first residual block layer includes a first residual part and a skip connection branch arranged in parallel. The first residual part and the skip connection branch are connected by pixel-by-pixel addition and output. The first residual part includes a second convolutional block and a BN layer. The second residual block layer and the third residual block layer both include a second residual part and a skip connection branch arranged in parallel. The second residual part and the skip connection branch are connected by pixel-by-pixel addition and output. The second residual part includes a second convolutional block, a BN layer, a Relu activation layer, a second convolutional block, and a BN layer connected in sequence. The first residual block layer is connected to the second residual block layer, and the second residual block layer is connected to the third residual block layer through Relu activation layers.
[0015] According to one aspect of the above technical solution, the last layer network of the ResNet further includes a first convolutional block and a reshape layer. The first convolutional block and the reshape layer are connected in sequence before the first residual block layer. The first convolutional block, a Relu activation layer, and the reshape layer are connected in sequence after the third residual block layer.
[0016] According to one aspect of the above technical solution, the encoder part further includes a linear embedding layer connected before several levels of feature extraction groups. The several levels of feature extraction groups include a first-level feature extraction group, a second-level feature extraction group, and a third-level feature extraction group connected in sequence. The first-level feature extraction group, the second-level feature extraction group, and the third-level feature extraction group all include consecutive Swin Transformer Blocks and a patch merging layer. The consecutive Swin Transformer Blocks are a W-Transformer Block and an SW-Transformer Block connected in sequence.
[0017] According to one aspect of the above technical solution, the several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and consecutive Swin Transformer Blocks connected in sequence. The third-level feature fusion group is connected to the patch projection layer through the patch expansion layer.
[0018] According to one aspect of the above technical solution, the step of inputting the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part into the corresponding level of feature fusion group in the decoder part through the skip connection part for multi-scale feature fusion to obtain a fused feature map specifically includes:
[0019] In the first-level feature extraction group, consecutive Swin Transformer Blocks are connected to consecutive Swin Transformer Blocks in the third-level feature fusion group through skip connection parts. In the second-level feature extraction group, consecutive Swin Transformer Blocks are connected to consecutive Swin Transformer Blocks in the second-level feature fusion group through skip connection parts. In the third-level feature extraction group, consecutive Swin Transformer Blocks are connected to consecutive Swin Transformer Blocks in the first-level feature fusion group through skip connection parts.
[0020] According to one aspect of the above technical solution, the W-Transformer Block includes a W layer block and an M layer block connected in sequence. The W layer block includes a skip connection branch and a first branch arranged in parallel. The skip connection branch and the first branch are connected by pixel-by-pixel superposition to output. The first branch includes an LN layer and a W-MSA layer connected in sequence. The M layer block includes a skip connection branch and an M branch arranged in parallel. The skip connection branch and the M branch are connected by pixel-by-pixel superposition to output. The M branch includes an LN layer and an MLP layer connected in sequence.
[0021] According to one aspect of the above technical solution, the SW-Transformer Block includes an SW layer block and an M layer block connected in sequence. The SW layer block includes a skip connection branch and a second branch arranged in parallel. The skip connection branch and the second branch are connected by pixel-by-pixel superposition to output. The second branch includes an LN layer and an SW-MSA layer connected in sequence.
[0022] According to one aspect of the above technical solution, the loss function of the VMAT dose verification model based on the Swin Transformer network is the L1 function, and the optimizer is the Adam function.
[0023] Another aspect of the present invention is to provide a VMAT dose verification system based on the Swin Transformer network for implementing the above VMAT dose verification method based on the Swin Transformer network. The system includes:
[0024] A model construction module for obtaining CT images of historical users and corresponding radiotherapy doses and verification dose images, and constructing a VMAT dose verification model based on the Swin Transformer network. The VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a skip connection part;
[0025] An image preprocessing module for preprocessing the input CT image and corresponding radiotherapy dose image of the target user and segmenting them into several patches;
[0026] A shallow feature extraction module for inputting the patches into the encoder part to extract multi-scale shallow global correlation features. The encoder part includes several levels of feature extraction groups;
[0027] A deep network feature extraction module for inputting the encoded feature map output by the last level of feature extraction group into the improved bottleneck part to obtain a bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of the ResNet network;
[0028] A deep feature extraction module for inputting the bottleneck feature map into the decoder part for deep feature extraction. The decoder part includes several levels of feature fusion groups arranged in a stacked manner;
[0029] A feature fusion module for inputting the different-scale shallow global correlation features extracted by each level of feature extraction group in the encoder part into the corresponding level of feature fusion group in the decoder part through the skip connection part for multi-scale feature fusion to obtain a fused feature map;
[0030] A prediction output module for inputting the fused feature map into the patch projection layer and outputting a VMAT dose verification prediction map. Description of the Drawings
[0031] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0032] Figure 1 It is a flowchart of the VMAT dose verification method based on the Swin Transformer network in the first embodiment of the present invention;
[0033] Figure 2 It is a schematic structural diagram of the VMAT dose verification model based on the Swin Transformer network in the first embodiment of the present invention;
[0034] Figure 3 It is a schematic structural diagram of the continuous Swin Transformer Block in the first embodiment of the present invention;
[0035] Figure 4 It is a schematic structural diagram of the improved bottleneck part in the first embodiment of the present invention;
[0036] Figure 5 It is a VMAT dose verification prediction map of different prediction methods in the first embodiment of the present invention;
[0037] Figure 6 This is the structural block diagram of the VMAT dose verification system based on the Swin Transformer network in the second embodiment of the present invention;
[0038] Explanation of the symbols of the components in the attached drawings:
[0039] Model construction module 100, image preprocessing module 200, shallow feature extraction module 300, deep network feature extraction module 400, deep feature extraction module 500, feature fusion module 600, prediction output module 700. Specific implementation manners
[0040] To make the objectives, features, and advantages of the present invention more obvious and understandable, the specific implementation manners of the present invention will be described in detail below with reference to the attached drawings. Several embodiments of the present invention are shown in the attached drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0041] Embodiment 1
[0042] Please refer to Figures 1 - 5 , which shows a VMAT dose verification method provided by the first embodiment of the present invention. The method includes steps S10 - S16:
[0043] Step S10, obtain the CT images, corresponding radiotherapy dose images, and VMAT dose verification images of historical users, and construct a VMAT dose verification model based on the Swin Transformer network. The VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a skip connection part;
[0044] Among them, the radiotherapy dose image is an image used to display the dose distribution in the Treatment Planning System (TPS), that is, the TPS dose image.
[0045] The VMAT dose verification image is an image used to display the actual received dose distribution of the patient in the three-dimensional dose verification system (Dolphin-Compass), that is, the TCD dose image.
[0046] Therefore, three-dimensional dose verification is an important link in radiotherapy, which is used to evaluate the accuracy and safety of radiotherapy plans and ensure that the irradiation dose actually received by the patient during the treatment is consistent with the planned dose.
[0047] Step S11: Preprocess the input CT image and corresponding radiotherapy dose image of the target user, and segment them into several patches.
[0048] Among them, in order to perform more efficient feature extraction and analysis on the CT image and the corresponding radiotherapy dose image, the CT image and the corresponding radiotherapy dose image are divided into many smaller, overlapping or non-overlapping patches.
[0049] Specifically, perform Patch Partition processing on the input CT image and corresponding radiotherapy dose image of the target user, and segment the CT image and the corresponding radiotherapy dose image into several patches.
[0050] Step S12: Input the patches into the encoder part to extract multi-scale shallow global correlation features. The encoder part includes several levels of feature extraction groups.
[0051] Among them, the encoder part also includes a linear embedding layer connected before several levels of feature extraction groups. The linear embedding layer of the encoder part tokenizes the data of the patches to achieve a C-dimensional representation of size H / 4×W / 4, where H is the height of the image, W is the width of the image, and C is the number of channels of the image. Consecutive Swin Transformer Blocks (windowed transformer module × 2) are used for feature extraction, and the PatchMerging layer performs downsampling, connecting the four segmented parts together, reducing the resolution of the patches to half of the original, and unifying the dimension to twice the original.
[0052] Furthermore, the several levels of feature extraction groups include a first-level feature extraction group, a second-level feature extraction group, and a third-level feature extraction group connected in sequence. The first-level feature extraction group, the second-level feature extraction group, and the third-level feature extraction group all include consecutive Swin Transformer Blocks and a PatchMerging layer. The consecutive Swin Transformer Blocks are a W-Transformer Block (window-based multi-head self-attention transformer module) and an SW-Transformer Block (shifted window-based multi-head self-attention transformer module) connected in sequence.
[0053] Therefore, after the patches are input into the linear embedding layer and the windowed transformer module × 2 in the first-level feature extraction group, the first-level shallow feature map is extracted, and the feature dimension is , and then the feature dimension becomes , then input the windowed transformer module in the secondary extraction group ×2 to extract the secondary shallow feature map, and the feature dimension is , then change the feature dimension to through the patch merging layer, and finally input the windowed transformer module in the tertiary extraction group ×2 to extract the tertiary shallow feature map, and the feature dimension is , and finally the feature dimension through the patch merging layer is .
[0054] It can be understood that, as Figure 2 shown, the encoder part includes a linearly embedded layer, a windowed transformer module ×2, a patch merging layer, a windowed transformer module ×2, a patch merging layer, a windowed transformer module ×2, and a patch merging layer connected in sequence.
[0055] Further, the continuous Swin Transformer Block is a W-TransformerBlock (window-based multi-head self-attention transformer module) and an SW-Transformer Block (moving window-based multi-head self-attention transformer module) connected in sequence. As Figure 3 shown, the W-Transformer Block includes a W layer block and an M layer block connected in sequence. The W layer block includes a skip connection branch and a first branch arranged in parallel. The skip connection branch and the first branch are connected by pixel-by-pixel superposition to output. The first branch includes an LN layer (layer normalization layer) and a W-MSA layer (window multi-head self-attention mechanism layer) connected in sequence. Among them, the LN layer is a normalization process in the layer dimension to avoid the problem of gradient disappearance. The W-MSA layer independently calculates the data of each window to save a large amount of calculation.
[0056] The M layer block includes a skip connection branch and an M branch arranged in parallel. The skip connection branch and the M branch are connected by pixel-by-pixel superposition to output. The M branch includes an LN layer and an MLP layer connected in sequence. Among them, the MLP layer is a multi-layer perceptron layer for feature mapping.
[0057] Similarly, the SW-Transformer Block includes an SW layer block and an M layer block connected in sequence. The SW layer block includes a skip connection branch and a second branch arranged in parallel. The skip connection branch and the second branch are connected by pixel-by-pixel superposition to output. The second branch includes an LN layer and an SW-MSA layer (moving window multi-head self-attention mechanism layer) connected in sequence. Among them, the calculation of the SW-MSA layer is to solve the problem that the W-MSA layer independently calculates each window, resulting in a lack of information exchange between windows. Specifically, by moving the window of the SW-MSA layer down and to the right by half of the window size, and then calculating the output of the W-MSA layer model of the moving window, the information interaction between windows is realized.
[0058] Among them, the calculation formulas of the models of the W-MSA layer, the SW-MSA layer, and the MLP layer are as follows:
[0059] (1)
[0060] (2)
[0061] (3)
[0062] (4)
[0063] (5)
[0064] Among them, in formulas (1)-(4), and respectively represent the outputs of the models of the th SW-MSA layer or W-MSA layer and the MLP layer, and respectively represent the outputs of the models of the th SW-MSA layer or W-MSA layer and the MLP layer. In formula (5), respectively represent the query matrix, the key matrix, and the value matrix. represents the number of patches in a window, while represents the dimensional information of the query matrix or the key matrix. Since the axis values of the relative positions in the model are all within the range of , a smaller bias matrix needs to be parameterized as , and B is the value extracted from .
[0065] In the continuous Swin Transformer Block, the input data first passes through the LN layer. The LN layer here is similar in function to the BN layer commonly used in computer vision (CV). Both normalize the activation values of the previous layer to a certain extent to avoid the problem of gradient disappearance. The difference between the LN layer and the BN layer lies in the dimension for normalization. The LN layer normalizes in the layer dimension, while the BN layer (batch normalization layer) normalizes in the batch dimension. In the field of natural language processing (NLP), the batch size of the network is usually smaller than that in the field of computer vision (CV), which makes the BN layer perform worse than the LN layer. Therefore, the LN layer is widely used in the continuous Swin Transformer Block.
[0066] The calculation formula of the LN layer is as follows:
[0067] (6)
[0068] Among them, is the input value, is the output value, denotes the mean value of denotes the variance of is a very small constant used to avoid the case of a zero denominator, and are learnable parameters.
[0069] After passing through the LN layer, the data is input into the W-MSA layer or the SW-MSA layer. Compared with the multi-head self-attention mechanism layer (MSA layer), the W-MSA layer can significantly reduce the computational complexity by independently calculating the data of each window. For an input image of size , assuming that each window contains patches of size , the computational complexity formulas of the MSA layer and the W-MSA layer are shown in Formulas (7) and (8) respectively:
[0070] (7)
[0071] (8)
[0072] Among them, and are the computational complexities of the MSA layer and the W-MSA layer respectively, is the number of feature channels, represents the side length of the patches in the window, and represent the height and width of the input image respectively.
[0073] Although the W-MSA layer reduces the computational complexity, it leads to a lack of information exchange between windows. To solve this problem, it is necessary to add the SW-MSA layer calculation in the subsequent blocks. By moving the window down and to the right by half of the window size and calculating the model output of the W-MSA layer for the moving window again, information interaction between windows is achieved. Therefore, the W-MSA layer and the SW-MSA layer need to appear in pairs. For this reason, the number of blocks in consecutive Swin Transformer Blocks is usually even.
[0074] Step S13: Input the encoded feature map output by the last-level feature extraction group into the improved bottleneck part to obtain the bottleneck feature map. The improved bottleneck part includes replacing the consecutive Swin Transformer Blocks with the last layer network of ResNet;
[0075] It should be noted that the inherent characteristics of the Swin Transformer network enable it to process image features stably and at a relatively high resolution, accurately meeting the requirements for more fine-grained and globally consistent prediction of image features in dense prediction tasks. The bottleneck part of the traditional VMAT dose verification model based on the Swin Transformer network is the continuous Swin Transformer Block, but it cannot converge in the deep network and has poor feature extraction ability. Therefore, in this embodiment, the last layer network of ResNet is used to replace the continuous Swin Transformer Block. The last layer network of ResNet will not reduce the feature extraction ability as the network depth increases, which will enhance the feature extraction of the encoded feature map in the deep network and overcome the problem of decreased feature extraction ability caused by poor network depth convergence. At the same time, the last layer network of ResNet reduces the network parameters by nearly 40%, which can effectively improve the performance of the model and the efficiency of prediction.
[0076] In this embodiment, the characteristics of the residual network block in the last layer network of ResNet are used to solve the problem of network degradation caused by the deepening of the network layer, so as to calculate the parameters of the thousand-layer network, improve the accuracy of model prediction, effectively improve the extraction ability of the encoded feature map, and at the same time, as the network depth increases, the performance of the last layer network of ResNet will not decline, and the resolution and feature dimension of the encoded feature map remain unchanged.
[0077] Among them, the last layer network of ResNet includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence. The encoded feature map is subjected to feature extraction in the deep network through the last layer network of ResNet to obtain a bottleneck feature map.
[0078] Specifically, the first residual block layer includes a first residual part and a skip connection branch arranged in parallel. The first residual part and the skip connection branch are connected in parallel through pixel-by-pixel addition. The first residual part includes a second convolutional block and a BN layer (batch normalization layer). The second residual block layer and the third residual block layer both include a second residual part and a skip connection branch arranged in parallel. The second residual part and the skip connection branch are connected in parallel through pixel-by-pixel addition. The second residual part includes a second convolutional block, a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolutional block, and a BN layer (batch normalization layer) connected in sequence. The first residual block layer is connected to the second residual block layer, and the second residual block layer is connected to the third residual block layer through a Relu activation layer (rectified linear unit activation layer).
[0079] In addition, the last layer of the ResNet also includes a first convolutional block and a reshape layer (shape transformation layer). The first convolutional block and the reshape layer (shape transformation layer) are sequentially connected before the first residual block layer, and the first convolutional block, a Relu activation layer (rectified linear unit activation layer), and a reshape layer (shape transformation layer) are sequentially connected after the third residual block layer.
[0080] Among them, the first convolutional block is a convolutional layer with a kernel size of 1×1 and 512 channels to reduce the dimension of the feature vector, and the second convolutional block is a convolutional layer with a kernel size of 3×3 and 512 channels.
[0081] It can be understood that, as Figure 4 shown, the improved bottleneck part includes a reshape layer (shape transformation layer), a first convolutional block (1×1 convolutional kernel, 512), a second convolutional block (3×3 convolutional kernel, 512), a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolutional block (3×3 convolutional kernel, 512), a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolutional block (3×3 convolutional kernel, 512), a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolutional block (3×3 convolutional kernel, 512), a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolutional block (3×3 convolutional kernel, 512), a BN layer (batch normalization layer), a first convolutional block (1×1 convolutional kernel, 512), a Relu activation layer (rectified linear unit activation layer), and a reshape layer (shape transformation layer) connected in sequence.
[0082] Step S14: Input the bottleneck feature map into the decoder part for deep feature extraction. The decoder part includes several levels of feature fusion groups arranged in a stacked manner;
[0083] Among them, the several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and consecutive Swin Transformer Blocks connected in sequence. The third-level feature fusion group is connected to the patch projection layer through the patch expansion layer.
[0084] Furthermore, upsampling is performed through the patch expansion layer (Patch Expanding) to double the resolution of the input features while halving the feature dimension.
[0085] It can be understood that the improved bottleneck part improves the feature extraction ability. The obtained bottleneck feature map does not change the feature dimension and is denoted as , the bottleneck feature map is input into the patch expansion layer in the first-level feature fusion group, and the feature dimension becomes , and then use the window transformer module ×2 in the first-level feature extraction group to extract the first-level deep feature map, the feature dimension is Then, after the patch expansion layer in the secondary feature fusion group, the feature dimension becomes , and then use the window transformer module ×2 of the secondary feature extraction group to extract the secondary deep feature map, the feature dimension is Then, after the patch expansion layer of the three-level feature fusion group, the feature dimension becomes , and then use the window transformer module ×2 in the three-level extraction group to extract the three-level deep feature map, the feature dimension is , and finally restored to full resolution through the patch expansion layer, expressed as .
[0086] like Figure 2 As shown, the decoder part includes a patch expansion layer, a window transformer module × 2, a patch expansion layer, a window transformer module × 2, a patch expansion layer, a window transformer module × 2, and a patch expansion layer, which are connected in sequence.
[0087] Step S15, inputting the shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part to the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map;
[0088] Among them, through the jump connection part, the multi-scale shallow global correlation features of the encoder part are fused with the deep features of the decoder part to enhance the feature extraction capability. It can better retain the edge details of the prediction area. At the same time, the upsampling features and downsampling features are fused to reduce the spatial information loss problem caused by downsampling.
[0089] Specifically, Figure 2 As shown, the continuous Swin Transformer Block in the first-level feature extraction group is connected to the continuous Swin TransformerBlock in the third-level feature fusion group through a skip connection part (Skip Connection), that is, the first-level shallow feature map and the third-level deep feature map are feature fused.
[0090] Similarly, the continuous Swin Transformer Block in the secondary feature extraction group is partially connected to the continuous Swin Transformer Block in the secondary feature fusion group through a jump connection, that is, the secondary shallow feature map and the secondary deep feature map are feature fused.
[0091] Similarly, in the third-level feature extraction group, consecutive Swin Transformer Blocks are connected to consecutive Swin Transformer Blocks in the first-level feature fusion group through the skip connection part, that is, the third-level shallow feature map is fused with the first-level deep feature map.
[0092] Step S16: Input the fused feature map into the patch projection layer to output the VMAT dose verification prediction map.
[0093] Among them, the fused feature map output by the last patch expansion layer, with a feature dimension of , is input into the patch projection layer (Patch Projection) to obtain a VMAT dose verification prediction map with a size of . The method of this embodiment not only has good prediction accuracy, but also shows good robustness and generalization ability.
[0094] In addition, in order to train the model better and more accurately and adapt to the VMAT dose verification prediction task, the loss function of the VMAT dose verification model based on the Swin Transformer network is replaced with the L1 function, and the optimizer is replaced with the Adam function.
[0095] Furthermore, in order to further explore the optimal solution for network training, a weighted combination of the L1 and L2 loss functions is used to train the network.
[0096] Experiments and Results
[0097] Dataset: 307 patients who received volumetric modulated arc therapy (VMAT) treatment between 2018 and 2021. 80% of the dataset is used as the training set, and the remaining 20% is used as the test set. Table 1 shows the clinical characteristics of the patients.
[0098] Table 1:
[0099]
[0100] Training parameters: The proposed VMAT dose verification model based on the Swin Transformer network is implemented in PyTorch and trained and tested on an NVIDIA GeForce RTX 1080 GPU with 11GB of memory, benefiting from CUDA acceleration. The Adam optimizer is selected, the L1 function is used as the loss function, and the base learning rate is 1e-5. The maximum number of training epochs is set to 200.
[0101] Experimental setup: To verify the effectiveness of the method predicted in this embodiment, it was compared with other prediction networks using the same test set: UNet, Res-UNet, TransQA, Swin-UNet, Swin-UNet+ResNet. Three quantitative metrics were used to evaluate the performance of the prediction model, including the Structural Similarity Index (SSIM), Mean Absolute Error (MAE), and Root Mean Square Error (RMSE).
[0102] Among them, UNet: A classic encoding and decoding network.
[0103] Res-UNet: A combination of UNet and the Residual Network (ResNet).
[0104] TransQA: A combination of the Transformer based on the self-attention mechanism and the improved UNet.
[0105] Swin-UNet: A traditional VMAT dose verification model based on the Swin Transformer network, with the bottleneck part being consecutive Swin Transformer Blocks.
[0106] Table 2:
[0107]
[0108] It can be clearly seen from Table 2 that the method of this embodiment has the best prediction results and has significant improvements compared with the three methods of UNet, Res-UNet, and TransQA. Compared with UNet, Res-UNet, and TransQA, in terms of SSIM, it has increased by 0.11, 0.082, and 0.064 respectively. In terms of MAE, it has decreased by 59.9%, 48.9%, and 35.0% respectively. In terms of RMSE, it has decreased by 65.9%, 43.7%, and 34.7% respectively. Compared with the traditional Swin-UNet, this embodiment has improved in terms of SSIM, MAE, and RMSE.
[0109] To more intuitively compare the differences between different prediction methods, the predicted verification dose distributions of several major cancer sites were selected for display, namely the head and neck (H&N), abdomen (Abdomen), and chest (Chest). Among them, the true dose (GT) distribution image is the VMAT dose verification map of the clinical analysis and evaluation plan, which is used as the true reference data; as Figure 5 shown, it can be clearly seen from the VMAT dose verification prediction map that the results obtained by the method of this embodiment are better, especially showing a better fitting effect in the high-dose region and details, indicating that it is closer to the VMAT dose verification map of the clinical analysis and evaluation plan.
[0110] To study the prediction performance of each method for specific cancer sites, we conducted tests for three major cancer sites (head and neck, abdomen, and chest) respectively, and compared the results of each method, as shown in Table 3.
[0111] Table 3:
[0112]
[0113] From the comparison of three metrics: Structural Similarity Index (SSIM), Mean Absolute Error (MAE), and Root Mean Square Error (RMSE), the VMAT dose verification prediction results of all methods for the chest are more accurate than those for the head and neck and abdomen. This may be because the structure of the chest is simpler than that of the head and neck and abdomen, making it easier for the network to extract features. In addition, since the proportion of patients with head and neck cases in the dataset is the largest, the network is more comprehensively trained for head and neck cases, making its VMAT dose verification prediction results second only to those of the chest. Generally speaking, the method of this embodiment has achieved the best VMAT dose verification prediction accuracy for these three cancer sites. This shows that the network of this embodiment exhibits the best performance in terms of various shape and texture differences.
[0114] To conduct ablation experiments to study the influence of various components and important parameter settings on the experimental results, and explore the performance differences between two structures: Swin-UNet and the method of this embodiment. The results are shown in Table 4.
[0115] Table 4:
[0116]
[0117] It can be seen from Table 4 that in this embodiment, the last network layer of ResNet is used to replace the bottleneck part of the traditional Swin-UNet, reducing the memory occupancy of the trained model file by nearly 40%. This shows that the architecture of this embodiment not only reduces redundant parameters, has a lower time complexity, but also has a slight improvement in performance.
[0118] To further conduct ablation experiments to study the influence of various components and important parameter settings of the skip connection part in this embodiment on the experimental results, and study the influence of the number of skip connection parts on the network performance. Among them, the number of skip connection parts being 0 and 3 respectively represents no skip connection branch and three skip connection branches. When the number of skip connection parts is set to 2, it includes skip connection branches at the 1 / 16 and 1 / 8 positions. When the number of skip connection parts is 1, it only includes a skip connection branch at the 1 / 16 position. The results are shown in Table 5.
[0119] Table 5:
[0120]
[0121] As can be seen from Table 5, when the number of skip connection parts of the neural network is 3, the prediction accuracy of VMAT dose verification is the highest.
[0122] Compared with the prior art, the VMAT dose verification method based on the Swin Transformer network provided in this embodiment has the beneficial effects that: through the VMAT dose verification method based on the Swin Transformer network provided by the present invention, the three-dimensional dose error prediction effect can be effectively improved. Specifically, the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a skip connection part. The multi-scale shallow global correlation features are extracted through the encoder part, and then the deep network features are extracted through the improved bottleneck part. Then, the deep features are extracted through the decoder part. Finally, the multi-scale shallow global correlation features of the encoder are fused with the deep features of the decoder part through the skip connection part to output, improving feature extraction, thereby improving the prediction accuracy. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer network of ResNet, adopting the characteristics of the residual network to enhance the extraction ability in the deep network, and further improving the prediction accuracy. In addition, the improved bottleneck part reduces the network parameters, can effectively improve the performance of the model, and thus improves the prediction accuracy, thereby solving the technical problems of generally low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolution / pooling operations and insensitivity to three-dimensional dose gradient changes.
[0123] Embodiment 2
[0124] Please refer to Figure 6 , which shows a VMAT dose verification system based on the Swin Transformer network provided by the second embodiment of the present invention. The system includes:
[0125] A model construction module 100, configured to obtain the CT images, corresponding radiotherapy doses and verification dose images of historical users, and construct a VMAT dose verification model based on the Swin Transformer network. The VMAT dose verification prediction model includes an encoder part, an improved bottleneck part, a decoder part and a skip connection part;
[0126] An image preprocessing module 200, configured to preprocess the input CT images and corresponding radiotherapy dose images of the target user and segment them into several patches;
[0127] A shallow feature extraction module 300, configured to input the patches into the encoder part to extract multi-scale shallow global correlation features. The encoder part includes several levels of feature extraction groups;
[0128] The deep network feature extraction module 400 is configured to input the encoded feature map output by the last-level feature extraction group into the improved bottleneck part to obtain a bottleneck feature map. The improved bottleneck part includes replacing consecutive Swin Transformer Blocks with the last layer of the ResNet network;
[0129] The deep feature extraction module 500 is configured to input the bottleneck feature map into the decoder part for deep feature extraction. The decoder part includes several levels of feature fusion groups arranged in a stacked manner;
[0130] The feature fusion module 600 is configured to input the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part into the corresponding level of feature fusion group in the decoder part through the skip connection part for multi-scale feature fusion to obtain a fused feature map;
[0131] The prediction output module 700 is configured to input the fused feature map into the patch projection layer and output a VMAT dose verification prediction map.
[0132] Compared with the prior art, the VMAT dose verification system based on the Swin Transformer network provided in this embodiment has the beneficial effects that: through the VMAT dose verification system based on the Swin Transformer network provided by the present invention, the three-dimensional dose error prediction effect can be effectively improved. Specifically, multi-size shallow global correlation features are extracted by the shallow feature extraction module, then deep network feature extraction is performed by the deep network feature extraction module, then deep feature extraction is performed by the deep feature extraction module, and finally the multi-size shallow global correlation features of the encoder and the deep features of the decoder part are fused and output by the feature fusion module to improve feature extraction, thereby improving the prediction accuracy. The improved bottleneck part of the deep network feature extraction module includes replacing consecutive Swin Transformer Blocks with the last layer of the ResNet network, adopting the characteristics of the residual network to enhance the extraction ability in the deep network, and further improving the prediction accuracy. In addition, the improved bottleneck part reduces the network parameters, can effectively improve the performance of the model, and thus improves the prediction accuracy, thereby solving the technical problems of generally low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolution / pooling operations, and insensitivity to three-dimensional dose gradient changes.
[0133] The technical features of each of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0134] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0135] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A VMAT dose verification method based on Swin Transformer network, characterized in that: The method comprises: Obtaining historical users' CT images, corresponding radiotherapy dose images and VMAT dose verification images, and constructing a VMAT dose verification model based on a SwinTransformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a skip connection part; The input CT image of the target user and the corresponding radiotherapy dose image are preprocessed and segmented into several patches; The patch is input into the encoder part to extract multi-scale shallow global correlation features, and the encoder part includes several levels of feature extraction groups; The encoded feature map output by the last level feature extraction group is input into the improved bottleneck part to obtain the bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet. Inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; The shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part are input into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The fused feature map is input into the patch projection layer, and a VMAT dose verification prediction map is output.
2. The VMAT dose verification method based on the Swin Transformer network according to claim 1, characterized in that: The last layer of the ResNet network includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence, the first residual block layer includes a first residual part and a jump connection branch arranged in parallel, the first residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the first residual part includes a second convolution block and a BN layer, the second residual block layer and the third residual block layer both include a second residual part and a jump connection branch arranged in parallel, the second residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the second residual part includes a second convolution block, a BN layer, a Relu activation layer, a second convolution block, and a BN layer connected in sequence; the first residual block layer and the second residual block layer, the second residual block layer and the third residual block layer are all connected through a Relu activation layer.
3. The VMAT dose verification method based on the Swin Transformer network according to claim 2, characterized in that: The last layer of the ResNet network also includes a first convolution block and a reshape layer. The first convolution block and the reshape layer are connected in sequence before the first residual block layer, and the first convolution block, the Relu activation layer, and the reshape layer are connected in sequence after the third residual block layer.
4. The VMAT dose verification method based on the Swin Transformer network according to claim 1, characterized in that: The encoder part also includes a linear embedding layer connected before several levels of feature extraction groups, and the several levels of feature extraction groups include a primary feature extraction group, a secondary feature extraction group, and a tertiary feature extraction group connected in sequence. The primary feature extraction group, the secondary feature extraction group, and the tertiary feature extraction group all include continuous Swin Transformer Block and a patch merging layer, and the continuous Swin Transformer Block is a W-Transformer Block and a SW-Transformer Block connected in sequence.
5. The VMAT dose verification method based on the Swin Transformer network according to claim 4 is characterized in that: The several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and a continuous Swin Transformer Block connected in sequence. The third-level feature fusion group is connected to the patch projection layer through a patch expansion layer.
6. The VMAT dose verification method based on the Swin Transformer network according to claim 5, characterized in that: The steps of inputting the shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map specifically include: The continuous Swin Transformer Block in the primary feature extraction group is partially connected to the continuous Swin Transformer Block in the tertiary feature fusion group through a jump connection, the continuous Swin Transformer Block in the secondary feature extraction group is partially connected to the continuous Swin Transformer Block in the secondary feature fusion group through a jump connection, and the continuous Swin Transformer Block in the tertiary feature extraction group is partially connected to the continuous Swin Transformer Block in the primary feature fusion group through a jump connection.
7. The VMAT dose verification method based on the Swin Transformer network according to claim 4, characterized in that: The W-Transformer Block includes a W layer block and an M layer block connected in sequence, the W layer block includes a jump connection branch and a first branch arranged in parallel, the jump connection branch and the first branch are connected and output by pixel-by-pixel superposition, the first branch includes an LN layer and a W-MSA layer connected in sequence, the M layer block includes a jump connection branch and an M branch arranged in parallel, the jump connection branch and the M branch are connected and output by pixel-by-pixel superposition, and the M branch includes an LN layer and an MLP layer connected in sequence.
8. The VMAT dose verification method based on the Swin Transformer network according to claim 7, characterized in that: The SW-Transformer Block includes a SW layer block and an M layer block connected in sequence, the SW layer block includes a jump connection branch and a second branch arranged in parallel, the jump connection branch and the second branch are connected and output by pixel-by-pixel superposition, and the second branch includes an LN layer and a SW-MSA layer connected in sequence.
9. The VMAT dose verification method based on Swin Transformer network according to claim 1, characterized in that: The loss function of the VMAT dose verification model based on the Swin Transformer network is the L1 function, and the optimizer is the Adam function.
10. A VMAT dose verification system based on Swin Transformer network, characterized in that: A system for implementing the VMAT dose verification method based on a Swin Transformer network according to any one of claims 1 to 9, the system comprising: A model building module, used to obtain CT images of historical users and corresponding radiotherapy doses and verification dose images, and to build a VMAT dose verification model based on a Swin Transformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a jump connection part; An image preprocessing module is used to preprocess the input CT image of the target user and the corresponding radiotherapy dose image and segment them into several patches; A shallow feature extraction module, for extracting multi-scale shallow global correlation features from the patch input encoder part, the encoder part includes several levels of feature extraction groups; A deep network feature extraction module is used to input the encoded feature map output by the last level feature extraction group into the improved bottleneck part to obtain a bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet; A deep feature extraction module, used for inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; A feature fusion module is used to input the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part to the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The prediction output module is used to input the fused feature map into the patch projection layer and output the VMAT dose verification prediction map.
Citation Information
Patent Citations
Medical image segmentation method based on Swin Transform and CNN parallel network
CN117351030A
Nonlinear medical image segmentation method based on SMTK-UNet model
CN119762775A