VMAT dose verification method and system based on Swin Transform network
Through the VMAT dose verification method based on Swin Transformer network, multi-scale features are extracted and deep network feature extraction combined with ResNet's deep network feature extraction, the problems of low manual verification efficiency and poor prediction of deep learning models in the prior art are solved, and more efficient and accurate dose verification prediction is achieved.
Patent Information
- Application Number
- CN202510443346.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the prior art, artificial three-dimensional dose verification is inefficient, deep learning models based on convolution/pooling operations have poor dose prediction effect, and are insensitive to changes in three-dimensional dose gradients.
The VMAT dose verification method based on the Swin Transformer network is adopted to extract the shallow global correlation features of multi-scale through the encoder. The improved bottleneck part uses the last layer of ResNet for deep network feature extraction, and deep feature extraction and multi-scale feature fusion are performed through the decoder and jump connection part to output the VMAT dose verification prediction map.
It effectively improves the prediction effect, enhances feature extraction ability, improves the prediction accuracy and various models' performance, and solves the problems of low manual verification efficiency and poor prediction effect of deep learning models.
Smart Images

Figure CN119963952A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a VMAT dose verification method and system based on a Swin Transformer network. Background Art
[0002] With the advancement of modern science and technology, radiotherapy has gradually become an indispensable and important means of treating malignant tumors. Volumetric modulated radiation therapy (VMAT) has become the core technology of precision radiotherapy, which improves the dose conformity of the target area while effectively protecting the endangered organs through the precise coordination of dynamic multi-leaf grating and gantry rotation.
[0003] However, the high complexity of volumetric modulated radiotherapy (VMAT) technology has led to severe challenges in its dose verification: first, the efficiency of manual three-dimensional dose verification is seriously out of touch with clinical needs, and it is difficult to meet the clinical verification needs of dozens of cases per day; second, deep learning models with convolution / pooling operations (CP) are generally used, and two-dimensional or one-dimensional information input is used to predict dose verification. Due to the loss of spatial dose characteristics and the locality of CP, these models have limitations in the long-term dependence of dose verification prediction, and the prediction effect is not ideal. This type of method is not sensitive enough to changes in three-dimensional dose gradients, making it difficult to effectively identify potential clinical errors. Summary of the invention
[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a VMAT dose verification method and system based on the Swin Transformer network, aiming to solve the technical problems in the prior art that the efficiency of manual three-dimensional dose verification is low, the deep learning model based on convolution / pooling operations has poor verification dose prediction effect and is insensitive to three-dimensional dose gradient changes.
[0005] One aspect of the present invention is to provide a VMAT dose verification method based on a Swin Transformer network, the VMAT dose verification method based on a Swin Transformer network comprising: Obtaining historical user's CT images, corresponding radiotherapy dose images and VMAT dose verification images, and constructing a VMAT dose verification model based on the Swin Transformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a skip connection part; The input CT image of the target user and the corresponding radiotherapy dose image are preprocessed and segmented into several patches; The patch is input into the encoder part to extract multi-scale shallow global correlation features, and the encoder part includes several levels of feature extraction groups; The encoded feature map output by the last level feature extraction group is input into the improved bottleneck part to obtain the bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet. Inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; The shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part are input into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The fused feature map is input into the patch projection layer, and a VMAT dose verification prediction map is output.
[0006] Compared with the prior art, the beneficial effect of the present invention is that: a VMAT dose verification method based on a SwinTransformer network provided by the present invention can effectively improve the prediction effect. Specifically, the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a jump connection part. The encoder part extracts multi-scale shallow global correlation features, and then the improved bottleneck part performs deep network feature extraction. Then the decoder part performs deep feature extraction. Finally, the multi-scale shallow global correlation features of the encoder and the deep features of the decoder part are fused and output through the jump connection part to improve feature extraction, thereby improving the accuracy of prediction. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet, using the characteristics of the residual network to enhance the extraction ability in the deep network, thereby improving the accuracy of prediction. In addition, the improved bottleneck part reduces the network parameters, which can effectively improve the various performances of the model, thereby improving the accuracy of prediction, thereby solving the common technical problems of low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolution / pooling operations, and insensitivity to three-dimensional dose gradient changes.
[0007] According to one aspect of the above technical solution, the last layer of the ResNet network includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence, the first residual block layer includes a first residual part and a jump connection branch arranged in parallel, the first residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the first residual part includes a second convolution block and a BN layer, the second residual block layer and the third residual block layer both include a second residual part and a jump connection branch arranged in parallel, the second residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the second residual part includes a second convolution block, a BN layer, a Relu activation layer, a second convolution block, and a BN layer connected in sequence; the first residual block layer and the second residual block layer, the second residual block layer and the third residual block layer are all connected through a Relu activation layer.
[0008] According to one aspect of the above technical solution, the last layer of the ResNet network also includes a first convolution block and a reshape layer. The first convolution block and the reshape layer are connected in sequence before the first residual block layer, and the first convolution block, the Relu activation layer, and the reshape layer are connected in sequence after the third residual block layer.
[0009] According to one aspect of the above technical solution, the encoder part also includes a linear embedding layer connected before several levels of feature extraction groups, and the several levels of feature extraction groups include a primary feature extraction group, a secondary feature extraction group, and a tertiary feature extraction group connected in sequence. The primary feature extraction group, the secondary feature extraction group, and the tertiary feature extraction group all include continuous Swin Transformer Block and a patch merging layer, and the continuous Swin Transformer Block is a W-Transformer Block and a SW-Transformer Block connected in sequence.
[0010] According to one aspect of the above technical solution, the several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group that are connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and a continuous Swin Transformer Block that are connected in sequence. The third-level feature fusion group is connected to the patch projection layer through a patch expansion layer.
[0011] According to one aspect of the above technical solution, the shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part are input into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain the fused feature map, which specifically includes: The continuous Swin Transformer Block in the primary feature extraction group is partially connected to the continuous Swin Transformer Block in the tertiary feature fusion group through a jump connection, the continuous Swin Transformer Block in the secondary feature extraction group is partially connected to the continuous Swin Transformer Block in the secondary feature fusion group through a jump connection, and the continuous Swin Transformer Block in the tertiary feature extraction group is partially connected to the continuous Swin Transformer Block in the primary feature fusion group through a jump connection.
[0012] According to one aspect of the above technical solution, the W-Transformer Block includes a W layer block and an M layer block connected in sequence, the W layer block includes a jump connection branch and a first branch arranged in parallel, the jump connection branch and the first branch are connected and output by pixel-by-pixel superposition, the first branch includes an LN layer and a W-MSA layer connected in sequence, the M layer block includes a jump connection branch and an M branch arranged in parallel, the jump connection branch and the M branch are connected and output by pixel-by-pixel superposition, and the M branch includes an LN layer and an MLP layer connected in sequence.
[0013] According to one aspect of the above technical solution, the SW-Transformer Block includes a SW layer block and an M layer block connected in sequence, the SW layer block includes a jump connection branch and a second branch arranged in parallel, the jump connection branch and the second branch are connected and output by pixel-by-pixel superposition, and the second branch includes an LN layer and a SW-MSA layer connected in sequence.
[0014] According to one aspect of the above technical solution, the loss function of the VMAT dose verification model based on the Swin Transformer network is an L1 function, and the optimizer is an Adam function.
[0015] Another aspect of the present invention is to provide a VMAT dose verification system based on a Swin Transformer network, which is used to implement the above-mentioned VMAT dose verification method based on a Swin Transformer network, and the system comprises: A model building module, used to obtain CT images of historical users and corresponding radiotherapy doses and verification dose images, and to build a VMAT dose verification model based on a Swin Transformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a jump connection part; An image preprocessing module is used to preprocess the input CT image of the target user and the corresponding radiotherapy dose image and segment them into several patches; A shallow feature extraction module, for extracting multi-scale shallow global correlation features from the patch input encoder part, the encoder part includes several levels of feature extraction groups; A deep network feature extraction module is used to input the encoded feature map output by the last level feature extraction group into the improved bottleneck part to obtain a bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet. A deep feature extraction module, used for inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; A feature fusion module is used to input the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part to the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The prediction output module is used to input the fused feature map into the patch projection layer and output the VMAT dose verification prediction map. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 Flow chart of the VMAT dose verification method based on the Swin Transformer network in Embodiment 1 of the present invention; Figure 2 Schematic diagram of the structure of the VMAT dose verification model based on the Swin Transformer network in the first embodiment of the present invention; Figure 3 It is a structural diagram of the continuous Swin Transformer Block in the first embodiment of the present invention; Figure 4 This is a schematic structural diagram of the improved bottleneck part in the first embodiment of the present invention; Figure 5 It is a VMAT dose verification prediction diagram of different prediction methods in Example 1 of the present invention; Figure 6 It is a structural block diagram of a VMAT dose verification system based on a Swin Transformer network in Embodiment 2 of the present invention; Component symbol description: Model building module 100, image preprocessing module 200, shallow feature extraction module 300, deep network feature extraction module 400, deep feature extraction module 500, feature fusion module 600, prediction output module 700. DETAILED DESCRIPTION
[0017] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings. Several embodiments of the present invention are shown in the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0018] Embodiment 1 See also Figure 1-Figure 5 , which shows a VMAT dose verification method based on a Swin Transformer network provided by a first embodiment of the present invention, the method comprises steps S10-S16: Step S10, obtaining the CT images of historical users, the corresponding radiotherapy dose images and VMAT dose verification images, and constructing a VMAT dose verification model based on the Swin Transformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a jump connection part; Among them, the radiotherapy dose image is an image used to display the dose distribution in the radiotherapy planning system (TPS), namely the TPS dose image.
[0019] The VMAT dose verification image is an image used in the three-dimensional dose verification system (Dolphin-Compass) to display the patient's actual received dose distribution, namely the TCD dose image.
[0020] Therefore, three-dimensional dose verification is an important part of radiotherapy, which is used to evaluate the accuracy and safety of radiotherapy plans and ensure that the actual radiation dose received by the patient during treatment is consistent with the planned dose.
[0021] Step S11, preprocessing the input CT image of the target user and the corresponding radiotherapy dose image, and dividing them into a number of patches; In order to perform more efficient feature extraction and analysis on the CT image and the corresponding radiotherapy dose image, the CT image and the corresponding radiotherapy dose image are divided into many smaller patches that overlap or do not overlap with each other.
[0022] Specifically, the input CT image of the target user and the corresponding radiotherapy dose image are processed by patch partitioning, and the CT image and the corresponding radiotherapy dose image are divided into a number of patches.
[0023] Step S12, inputting the patch into an encoder part to extract multi-scale shallow global correlation features, the encoder part includes several levels of feature extraction groups; Among them, the encoder part also includes a linear embedding layer connected before several levels of feature extraction groups. The linear embedding layer (Linear Embedding) of the encoder part (Encoder) tokenizes the patch data to achieve a C-dimensional representation of size H / 4×W / 4, where H is the height of the image, W is the width of the image, and C is the number of channels of the image. Continuous Swin Transformer Block (windowed transformer module × 2) is used for feature extraction, and the patch merging layer (PatchMerging) performs downsampling, connects the four segmented parts together, reduces the resolution of the patch to half of the original, and unifies the dimension to twice the original.
[0024] Furthermore, the several levels of feature extraction groups include a primary feature extraction group, a secondary feature extraction group, and a tertiary feature extraction group that are connected in sequence, and the primary feature extraction group, the secondary feature extraction group, and the tertiary feature extraction group all include continuous Swin Transformer Block and a patch merging layer, and the continuous Swin Transformer Block is a W-Transformer Block (window-based multi-head self-attention transformer module) and a SW-Transformer Block (moving window-based multi-head self-attention transformer module) connected in sequence.
[0025] Therefore, after the patch is input into the linear embedding layer and the window transformer module ×2 in the first-level feature extraction group, the first-level shallow feature map is extracted, and the feature dimension is , and then the feature dimension is transformed into , and then input the window transformer module ×2 in the secondary extraction group to extract the secondary shallow feature map, the feature dimension is , and then the feature dimension is transformed into Finally, the window transformer module ×2 in the three-level extraction group is input to extract the three-level shallow feature map, and the feature dimension is , and finally the feature dimension of the patch merging layer is .
[0026] Understandable, such as Figure 2As shown, the encoder part includes a linear embedding layer, a window transformer module × 2, a patch merging layer, a window transformer module × 2, a patch merging layer, a window transformer module × 2, and a patch merging layer, which are connected in sequence.
[0027] Furthermore, the continuous Swin Transformer Block is a W-TransformerBlock (window-based multi-head self-attention transformer module) and a SW-Transformer Block (moving window-based multi-head self-attention transformer module) connected in sequence, such as Figure 3 As shown, the W-Transformer Block includes a W layer block and an M layer block connected in sequence, the W layer block includes a jump connection branch and a first branch arranged in parallel, the jump connection branch and the first branch are connected and output by pixel-by-pixel superposition, the first branch includes an LN layer (layer normalization layer) and a W-MSA layer (window multi-head self-attention mechanism layer) connected in sequence, wherein the LN layer is a normalization process on the layer dimension to avoid the gradient disappearance problem, and the W-MSA layer independently calculates the data of each window to save a lot of calculations.
[0028] The M-layer block includes a jump connection branch and an M branch arranged in parallel, the jump connection branch and the M branch are connected and output by pixel-by-pixel superposition, and the M branch includes an LN layer and an MLP layer connected in sequence, wherein the MLP layer is a multi-layer perceptron layer for feature mapping.
[0029] Similarly, the SW-Transformer Block includes a SW layer block and an M layer block connected in sequence, the SW layer block includes a jump connection branch and a second branch arranged in parallel, the jump connection branch and the second branch are connected and output by pixel-by-pixel superposition, and the second branch includes a LN layer and a SW-MSA layer (moving window multi-head self-attention mechanism layer) connected in sequence. Among them, the calculation of the SW-MSA layer is to solve the problem of lack of information exchange between windows caused by the independent calculation of each window by the W-MSA layer. Specifically, the window of the SW-MSA layer is moved downward and to the right by half the window size, and the W-MSA layer model output of the moving window is calculated again to realize information interaction between windows.
[0030] Among them, the calculation formulas of the W-MSA layer model, the SW-MSA layer model and the MLP layer model are as follows: (1) (2) (3) (4) (5) Among them, in formula (1)-formula (4), and Respectively represent The output of the model of the SW-MSA layer or W-MSA layer and the model of the MLP layer, and Respectively represent The output of the model of the SW-MSA layer or W-MSA layer and the model of the MLP layer, in formula (5), denote the query matrix, key matrix, and value matrix respectively. represents the number of patches in a window, and Indicates the dimension information of the query matrix or key matrix. Since the axis values of the relative positions in the model are all Therefore, a smaller deviation matrix needs to be parameterized as , B is from The value extracted from .
[0031] In the continuous Swin Transformer Block, the input data first passes through the LN layer. The LN layer here is similar to the BN layer commonly used in computer vision (CV). Both normalize the activation values of the previous layer to a certain extent to avoid the gradient vanishing problem. The difference between the LN layer and the BN layer is the dimension in which the normalization is performed. The LN layer normalizes on the layer dimension, while the BN layer (batch normalization layer) normalizes on the batch dimension. In the field of natural language processing (NLP), the batch size of the network is usually smaller than that in the field of computer vision (CV), which makes the BN layer less effective than the LN layer. Therefore, the LN layer is widely used in the continuous Swin Transformer Block.
[0032] The calculation formula of LN layer is as follows: (6) in, is the input value, is the output value, express The mean of express The variance of . is a very small constant used to avoid the denominator being zero. and are learnable parameters.
[0033] After the LN layer, the data is input to the W-MSA layer or SW-MSA layer. Compared with the multi-head self-attention mechanism layer (MSA layer), the W-MSA layer can significantly reduce the amount of calculation by independently calculating the data of each window. The input image is assumed to be of size The computational complexity formulas for the patch, MSA layer and W-MSA layer are shown in formulas (7) and (8) respectively: (7) (8) in, and are the computational complexity of the MSA layer and the W-MSA layer, is the number of feature channels, represents the side length of the patch in the window, and Represent the height and width of the input image respectively.
[0034] Although the W-MSA layer reduces the amount of computation, it results in a lack of information exchange between windows. To solve this problem, the SW-MSA layer calculation must be added to the subsequent block. By moving the window down and to the right by half the window size, the W-MSA layer model output of the moving window is calculated again to achieve information exchange between windows. Therefore, the W-MSA layer and the SW-MSA layer need to appear in pairs. It is for this reason that the number of blocks in a continuous Swin Transformer Block is usually an even number.
[0035] Step S13, inputting the encoded feature map output by the last level feature extraction group into the improved bottleneck part to obtain a bottleneck feature map, wherein the improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet; It should be noted that the inherent characteristics of the Swin Transformer network enable it to process image features at a stable and relatively high resolution, accurately meeting the needs for finer-grained and globally consistent prediction of image features in dense prediction tasks. The bottleneck of the traditional VMAT dose verification model based on the Swin Transformer network is the continuous Swin Transformer Block, but it cannot converge in a deep network and has poor feature extraction capabilities. Therefore, in this embodiment, the last layer of ResNet is used to replace the continuous Swin Transformer Block. The last layer of ResNet will not reduce the feature extraction capability as the network depth increases, which will enhance the feature extraction of the encoded feature map in the deep network, overcoming the problem of decreased feature extraction capability caused by poor network depth convergence. At the same time, the last layer of ResNet reduces the network parameters by nearly 40%, which can effectively improve the performance of the model and improve the efficiency of prediction.
[0036] In this embodiment, the characteristics of the residual network block in the last layer of ResNet are used to solve the network degradation problem caused by the deepening of the network layers, so that the parameters of the thousand-layer network can be calculated, the accuracy of the model prediction can be improved, and the ability to extract the encoded feature map can be effectively improved. At the same time, as the network depth increases, the performance of the last layer of ResNet will not decrease, and the resolution and feature dimension of the encoded feature map will remain unchanged.
[0037] Among them, the last layer of the ResNet network includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence. The last layer of the ResNet network is used to perform feature extraction on the encoded feature map in a deep network to obtain a bottleneck feature map.
[0038] Specifically, the first residual block layer includes a first residual part and a jump connection branch arranged in parallel, the first residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the first residual part includes a second convolution block and a BN layer (batch normalization layer), the second residual block layer and the third residual block layer both include a second residual part and a jump connection branch arranged in parallel, the second residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the second residual part includes a second convolution block, a BN layer (batch normalization layer), a Relu activation layer (rectified linear unit activation layer), a second convolution block, and a BN layer (batch normalization layer) connected in sequence; the first residual block layer and the second residual block layer, as well as the second residual block layer and the third residual block layer are connected through a Relu activation layer (rectified linear unit activation layer).
[0039] In addition, the last layer of the ResNet network also includes a first convolution block and a reshape layer (shape transformation layer). The first convolution block and the reshape layer (shape transformation layer) are connected in sequence before the first residual block layer, and the first convolution block, the Relu activation layer (rectified linear unit activation layer), and the reshape layer (shape transformation layer) are connected in sequence after the third residual block layer.
[0040] Among them, the first convolution block is a convolution layer with a convolution kernel size of 1×1 and a channel number of 512 to reduce the dimension of the feature vector, and the second convolution block is a convolution layer with a convolution kernel size of 3×3 and a channel number of 512.
[0041] Understandable, such as Figure 4 As shown, the improved bottleneck part includes the reshape layer (shape transformation layer), the first convolution block (1×1 convolution kernel, 512), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the Relu activation layer (rectified linear unit activation layer), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the Relu activation layer (rectified linear unit activation layer), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the Relu activation layer (rectified linear unit activation layer), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the Relu activation layer (rectified linear unit activation layer), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the Relu activation layer (rectified linear unit activation layer), the second convolution block (3×3 convolution kernel, 512), the BN layer (batch normalization layer), the first convolution block (1×1 convolution kernel, 512), the Relu activation layer (rectified linear unit activation layer), and the reshape layer (shape transformation layer).
[0042] Step S14, inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; Among them, the several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group that are connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and a continuous Swin Transformer Block that are connected in sequence. The third-level feature fusion group is connected to the patch projection layer through a patch expansion layer.
[0043] Furthermore, upsampling is performed through the patch expansion layer to double the resolution of the input features and halve the feature dimension.
[0044] It can be understood that the improved bottleneck part improves the feature extraction capability, and the obtained bottleneck feature map does not change the feature dimension and is expressed as , the bottleneck feature map is input into the patch expansion layer in the first-level feature fusion group, and the feature dimension becomes , and then use the window transformer module ×2 in the first-level feature extraction group to extract the first-level deep feature map, the feature dimension is Then, after the patch expansion layer in the secondary feature fusion group, the feature dimension becomes , and then use the window transformer module ×2 of the secondary feature extraction group to extract the secondary deep feature map, the feature dimension is Then, after the patch expansion layer of the three-level feature fusion group, the feature dimension becomes , and then use the window transformer module ×2 in the three-level extraction group to extract the three-level deep feature map, the feature dimension is , and finally restored to full resolution through the patch expansion layer, expressed as .
[0045] like Figure 2 As shown, the decoder part includes a patch expansion layer, a window transformer module × 2, a patch expansion layer, a window transformer module × 2, a patch expansion layer, a window transformer module × 2, and a patch expansion layer, which are connected in sequence.
[0046] Step S15, inputting the shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; Among them, through the jump connection part, the multi-scale shallow global correlation features of the encoder part are fused with the deep features of the decoder part to enhance the feature extraction capability. It can better retain the edge details of the prediction area. At the same time, the upsampling features and downsampling features are fused to reduce the spatial information loss problem caused by downsampling.
[0047] Specifically, Figure 2 As shown, the continuous Swin Transformer Block in the first-level feature extraction group is connected to the continuous Swin TransformerBlock in the third-level feature fusion group through a skip connection part (Skip Connection), that is, the first-level shallow feature map and the third-level deep feature map are feature fused.
[0048] Similarly, the continuous Swin Transformer Block in the secondary feature extraction group is partially connected to the continuous Swin Transformer Block in the secondary feature fusion group through a jump connection, that is, the secondary shallow feature map and the secondary deep feature map are feature fused.
[0049] Similarly, the continuous Swin Transformer Block in the three-level feature extraction group is partially connected to the continuous Swin Transformer Block in the first-level feature fusion group through a jump connection, that is, the three-level shallow feature map and the first-level deep feature map are feature fused.
[0050] Step S16, inputting the fused feature map into the patch projection layer, and outputting a VMAT dose verification prediction map.
[0051] Among them, the fused feature map output by the last patch expansion layer has a feature dimension of , input into the patch projection layer (Patch Projection), and get a size of The VMAT dose verification prediction diagram shows that the method of this embodiment not only has good prediction accuracy, but also shows good robustness and generalization ability.
[0052] In addition, in order to better and more accurately train the model and adapt to the VMAT dose verification prediction task, the loss function of the VMAT dose verification model based on the SwinTransformer network was replaced with the L1 function, and the optimizer was replaced with the Adam function.
[0053] Furthermore, in order to further explore the optimal solution for network training, a weighted combination of L1 and L2 loss functions is used to train the network.
[0054] Tests and Results Dataset: 307 patients treated with volumetric modulated radiation therapy (VMAT) between 2018 and 2021. 80% of the dataset was used as a training set, and the remaining 20% was used as a test set. Table 1 shows the clinical characteristics of the patients.
[0055] Table 1:
[0056] Training parameters: The proposed VMAT dose validation model based on the Swin Transformer network was implemented in PyTorch and trained and tested on an NVIDIA GeForce RTX 1080 GPU with 11GB memory, benefiting from CUDA acceleration. The Adam optimizer was selected, using the L1 function as the loss function and a base learning rate of 1e-5. The maximum number of training epochs was set to 200.
[0057] Experimental setup: To verify the effectiveness of the prediction of the method in this example, the same test set was used to compare it with other prediction networks: UNet, Res-UNet, TransQA, Swin-UNet, Swin-UNet+ResNet. Three quantitative indicators were used to evaluate the performance of the prediction model, including structural similarity index (SSIM), mean absolute error (MAE), and root mean square error (RMSE).
[0058] Among them, UNet: classic encoding and decoding network.
[0059] Res-UNet: Combination of UNet and Residual Network (ResNet).
[0060] TransQA: Combining the self-attention mechanism-based Transformer with the improved UNet.
[0061] Swin-UNet: A traditional VMAT dose verification model based on the Swin Transformer network, where the bottleneck is the continuous Swin Transformer Block.
[0062] Table 2:
[0063] It can be clearly seen from Table 2 that the method of this embodiment has the best prediction results, which is significantly improved compared with the three methods of UNet, Res-UNet, and TransQA. Compared with UNet, Res-UNet, and TransQA, in terms of SSIM, it is improved by 0.11, 0.082, and 0.064, respectively. In terms of MAE, it is reduced by 59.9%, 48.9%, and 35.0%, respectively. In terms of RMSE, it is reduced by 65.9%, 43.7%, and 34.7%, respectively. Compared with the traditional Swin-UNet, this embodiment has improved in SSIM, MAE, and RMSE.
[0064] In order to more intuitively compare the differences between different prediction methods, the predicted verification dose distributions of several major cancer sites were selected for display, namely head and neck (H&N), abdomen (Abdomen), and chest (Chest). The true dose (GT) distribution images are VMAT dose verification images of the clinical analysis and evaluation plan, which are used as real reference data; Figure 5 As shown, it can be clearly seen from the VMAT dose verification prediction diagram that the method of this embodiment produces better results, especially showing better fitting effects in high-dose areas and details, which indicates that it is closer to the VMAT dose verification diagram of the clinical analysis evaluation plan.
[0065] To investigate the prediction performance of each method for specific cancer sites, we tested three major cancer sites (head and neck, abdomen, and chest) separately and compared the results of each method, as shown in Table 3.
[0066] Table 3:
[0067] From the comparison of the three indicators of structural similarity index (SSIM), mean absolute error (MAE) and root mean square error (RMSE), all methods have more accurate VMAT dose verification prediction results for the chest than for the head and neck and abdomen. This may be because the structure of the chest is simpler than that of the head and neck and abdomen, making it easier for the network to extract features. In addition, since the proportion of patients with head and neck cases in the dataset is the largest, the network is more comprehensively trained on head and neck cases, making its VMAT dose verification prediction results second only to the chest. Overall, the method of this embodiment achieved the best VMAT dose verification prediction accuracy in all three cancer sites. This shows that the network of this embodiment performs best in various shape and texture differences.
[0068] In order to study the impact of various components and important parameter settings on the experimental results, an ablation experiment was conducted to explore the performance differences between the two structures: Swin-UNet and the method of this embodiment. The results are shown in Table 4.
[0069] Table 4:
[0070] It can be concluded from Table 4 that this embodiment replaces the bottleneck part of the traditional Swin-UNet with the last network layer of ResNet, which reduces the memory usage of the trained model file by nearly 40%. This shows that the architecture of this embodiment not only reduces redundant parameters, has lower time complexity, but also slightly improves performance.
[0071] In order to further study the influence of various components and important parameter settings of the jump connection part in this embodiment on the experimental results, an ablation experiment is conducted to study the influence of the number of jump connection parts on the network performance, where the number of jump connection parts is 0 and 3, respectively, indicating no jump connection branches and three jump connection branches. When the number of jump connection parts is set to 2, it includes jump connection branches at the 1 / 16 and 1 / 8 positions. When the number of jump connection parts is 1, it only includes the jump connection branch at the 1 / 16 position, and the results are shown in Table 5.
[0072] Table 5:
[0073] It can be seen from Table 5 that when the number of jump connection parts of the neural network is 3, its VMAT dose verification prediction accuracy is the highest.
[0074] Compared with the prior art, the VMAT dose verification method based on the Swin Transformer network provided in this embodiment has the beneficial effect that: the VMAT dose verification method based on the Swin Transformer network provided by the present invention can effectively improve the three-dimensional dose error prediction effect. Specifically, the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a jump connection part. The encoder part extracts multi-scale shallow global correlation features, and then the improved bottleneck part performs deep network feature extraction. Then, the decoder part performs deep feature extraction. Finally, the multi-scale shallow global correlation features of the encoder and the deep features of the decoder part are fused and outputted through the jump connection part, thereby improving feature extraction and thus improving prediction accuracy. The improved bottleneck part includes the continuous Swin Transformer Block is replaced with the last layer of ResNet, and the characteristics of the residual network are adopted to enhance the extraction ability in the deep network, thereby improving the accuracy of the prediction. In addition, the improved bottleneck part reduces the network parameters, which can effectively improve the performance of the model, thereby improving the accuracy of the prediction, thus solving the common technical problems of low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolution / pooling operations, and insensitivity to three-dimensional dose gradient changes.
[0075] Embodiment 2 See also Figure 6 , which is a VMAT dose verification system based on a Swin Transformer network provided by a second embodiment of the present invention, and the system comprises: A model building module 100 is used to obtain CT images of historical users and corresponding radiotherapy doses and verification dose images, and to build a VMAT dose verification model based on a Swin Transformer network, wherein the VMAT dose verification prediction model includes an encoder part, an improved bottleneck part, a decoder part, and a jump connection part; An image preprocessing module 200 is used to preprocess the input CT image of the target user and the corresponding radiotherapy dose image, and divide them into a number of patches; A shallow feature extraction module 300, for inputting the patch into an encoder part to extract multi-scale shallow global correlation features, the encoder part including several levels of feature extraction groups; A deep network feature extraction module 400 is used to input the encoded feature map output by the last level feature extraction group into an improved bottleneck part to obtain a bottleneck feature map, wherein the improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet; A deep feature extraction module 500, used for inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; A feature fusion module 600 is used to input the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part to the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The prediction output module 700 is used to input the fused feature map into the patch projection layer and output a VMAT dose verification prediction map.
[0076] Compared with the prior art, the VMAT dose verification system based on the Swin Transformer network provided in this embodiment has the beneficial effect that: the VMAT dose verification system based on the Swin Transformer network provided by the present invention can effectively improve the three-dimensional dose error prediction effect, specifically, extracting multi-scale shallow global correlation features through a shallow feature extraction module, then performing deep network feature extraction through a deep network feature extraction module, then performing deep feature extraction through a deep feature extraction module, and finally, fusion output of the multi-scale shallow global correlation features of the encoder and the deep features of the decoder part through a feature fusion module, thereby improving feature extraction and thus improving the accuracy of prediction, and the improved bottleneck part of the deep network feature extraction module includes converting the continuous Swin Transformer Block is replaced with the last layer of ResNet, and the characteristics of the residual network are adopted to enhance the extraction ability in the deep network, thereby improving the accuracy of the prediction. In addition, the improved bottleneck part reduces the network parameters, which can effectively improve the performance of the model, thereby improving the accuracy of the prediction, thus solving the common technical problems of low efficiency of manual three-dimensional dose verification, poor dose prediction effect of deep learning models based on convolution / pooling operations, and insensitivity to three-dimensional dose gradient changes.
[0077] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0078] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0079] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A VMAT dose verification method based on Swin Transformer network, characterized in that: The method comprises: Obtaining historical users' CT images, corresponding radiotherapy dose images and VMAT dose verification images, and constructing a VMAT dose verification model based on a SwinTransformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part and a skip connection part; Preprocess the input target user's CT image and the corresponding radiotherapy dose image and segment them into several patches; The patch is input into the encoder part to extract multi-scale shallow global correlation features, and the encoder part includes several levels of feature extraction groups; The encoded feature map output by the last level feature extraction group is input into the improved bottleneck part to obtain the bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet. Inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; The shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part are input into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The fused feature map is input into the patch projection layer, and a VMAT dose verification prediction map is output.
2. The VMAT dose verification method based on the Swin Transformer network according to claim 1, characterized in that: The last layer of the ResNet network includes a first residual block layer, a second residual block layer, and a third residual block layer connected in sequence, the first residual block layer includes a first residual part and a jump connection branch arranged in parallel, the first residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the first residual part includes a second convolution block and a BN layer, the second residual block layer and the third residual block layer both include a second residual part and a jump connection branch arranged in parallel, the second residual part and the jump connection branch are connected and output by pixel-by-pixel superposition, the second residual part includes a second convolution block, a BN layer, a Relu activation layer, a second convolution block, and a BN layer connected in sequence; the first residual block layer and the second residual block layer, the second residual block layer and the third residual block layer are all connected through a Relu activation layer.
3. The VMAT dose verification method based on the Swin Transformer network according to claim 2, characterized in that: The last layer of the ResNet network also includes a first convolution block and a reshape layer. The first convolution block and the reshape layer are connected in sequence before the first residual block layer, and the first convolution block, the Relu activation layer, and the reshape layer are connected in sequence after the third residual block layer.
4. The VMAT dose verification method based on the Swin Transformer network according to claim 1, characterized in that: The encoder part also includes a linear embedding layer connected before several levels of feature extraction groups, and the several levels of feature extraction groups include a primary feature extraction group, a secondary feature extraction group, and a tertiary feature extraction group connected in sequence. The primary feature extraction group, the secondary feature extraction group, and the tertiary feature extraction group all include continuous Swin Transformer Block and a patch merging layer, and the continuous Swin Transformer Block is a W-Transformer Block and a SW-Transformer Block connected in sequence.
5. The VMAT dose verification method based on the Swin Transformer network according to claim 4 is characterized in that: The several levels of feature fusion groups include a first-level feature fusion group, a second-level feature fusion group, and a third-level feature fusion group connected in sequence. The first-level feature fusion group, the second-level feature fusion group, and the third-level feature fusion group all include a patch expansion layer and a continuous Swin Transformer Block connected in sequence. The third-level feature fusion group is connected to the patch projection layer through a patch expansion layer.
6. The VMAT dose verification method based on the Swin Transformer network according to claim 5, characterized in that: The steps of inputting the shallow global correlation features of different scales extracted by each level feature extraction group in the encoder part into the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map specifically include: The continuous Swin Transformer Block in the primary feature extraction group is partially connected to the continuous Swin Transformer Block in the tertiary feature fusion group through a jump connection, the continuous Swin Transformer Block in the secondary feature extraction group is partially connected to the continuous Swin Transformer Block in the secondary feature fusion group through a jump connection, and the continuous Swin Transformer Block in the tertiary feature extraction group is partially connected to the continuous Swin Transformer Block in the primary feature fusion group through a jump connection.
7. The VMAT dose verification method based on the Swin Transformer network according to claim 4, characterized in that: The W-Transformer Block includes a W layer block and an M layer block connected in sequence, the W layer block includes a jump connection branch and a first branch arranged in parallel, the jump connection branch and the first branch are connected and output by pixel-by-pixel superposition, the first branch includes an LN layer and a W-MSA layer connected in sequence, the M layer block includes a jump connection branch and an M branch arranged in parallel, the jump connection branch and the M branch are connected and output by pixel-by-pixel superposition, and the M branch includes an LN layer and an MLP layer connected in sequence.
8. The VMAT dose verification method based on the Swin Transformer network according to claim 7, characterized in that: The SW-Transformer Block includes a SW layer block and an M layer block connected in sequence, the SW layer block includes a jump connection branch and a second branch arranged in parallel, the jump connection branch and the second branch are connected and output by pixel-by-pixel superposition, and the second branch includes an LN layer and a SW-MSA layer connected in sequence.
9. The VMAT dose verification method based on Swin Transformer network according to claim 1, characterized in that: The loss function of the VMAT dose verification model based on the Swin Transformer network is the L1 function, and the optimizer is the Adam function.
10. A VMAT dose verification system based on Swin Transformer network, characterized in that: A system for implementing the VMAT dose verification method based on a Swin Transformer network according to any one of claims 1 to 9, the system comprising: A model building module, used to obtain CT images of historical users and corresponding radiotherapy doses and verification dose images, and to build a VMAT dose verification model based on a Swin Transformer network, wherein the VMAT dose verification model includes an encoder part, an improved bottleneck part, a decoder part, and a jump connection part; An image preprocessing module is used to preprocess the input CT image of the target user and the corresponding radiotherapy dose image and segment them into several patches; A shallow feature extraction module, for extracting multi-scale shallow global correlation features from the patch input encoder part, the encoder part includes several levels of feature extraction groups; A deep network feature extraction module is used to input the encoded feature map output by the last level feature extraction group into the improved bottleneck part to obtain a bottleneck feature map. The improved bottleneck part includes replacing the continuous Swin Transformer Block with the last layer of ResNet; A deep feature extraction module, used for inputting the bottleneck feature map into a decoder part for deep feature extraction, the decoder part including several levels of feature fusion groups arranged in a stacked manner; A feature fusion module is used to input the shallow global correlation features of different scales extracted by each level of feature extraction group in the encoder part to the feature fusion group of the corresponding level in the decoder part through the jump connection part to perform multi-scale feature fusion to obtain a fused feature map; The prediction output module is used to input the fused feature map into the patch projection layer and output the VMAT dose verification prediction map.
Citation Information
Patent Citations
Medical image segmentation method based on Swin Transform and CNN parallel network
CN117351030A
Sea-land segmentation method based on deep learning model
CN117994657A
Tongue body segmentation method based on improved UNet
CN118038058A
CT image segmentation method and system based on improved Swinin-Unet
CN119151963A
Automatic medical image segmentation method of U-shaped network based on channel mixing and self-attention mechanism
CN119180960A
Cited By
Real-time detection method for severity of diseases of various fruit leaves
CN120580439A