Multi-direction dynamic feature fusion method based on deep learning
By adopting a multi-directional dynamic feature fusion method in semantic segmentation of remote sensing images, combining the receptive field module and the cross-window context interaction module, the problem of fusion of small target objects and cross-scale features is solved, and the accuracy of segmentation and recognition capabilities in complex scenarios are improved.
Patent Information
- Application Number
- CN202411771445.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-04
Smart Images

Figure CN119919764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing image technology, and in particular to a multi-directional dynamic feature fusion method based on deep learning. Background Art
[0002] With the continuous advancement of computer technology and satellite remote sensing technology, semantic segmentation of remote sensing images has become an important research direction in remote sensing image processing. By performing semantic segmentation on remote sensing images, each pixel in the image can be attributed to a specific ground object category at the pixel level, such as buildings, roads, vegetation, and water bodies. This technology is widely used in disaster monitoring, land use planning, urban construction, and environmental monitoring.
[0003] Existing semantic segmentation methods are usually based on deep learning techniques, especially the combination of convolutional neural networks (CNNs) and Transformer models. Convolutional neural networks capture detailed features through local receptive fields, while Transformer models are able to model global feature correlations. These methods have made significant progress in the segmentation of large target objects. However, for the recognition of small target objects and cross-scale feature fusion problems, existing hybrid models still face certain challenges.
[0004] Although the existing convolutional neural network and Transformer combined model can effectively extract local and global features, due to the different feature extraction mechanisms of the two, feature fusion may lead to information loss or feature imbalance. In addition, when processing multi-scale features, existing methods usually adopt serial or parallel structures. This method often ignores the mutual guidance of cross-scale features, resulting in unsatisfactory segmentation of small target objects. The similarities between specific categories (such as buildings and roads, trees and vegetation) also increase the difficulty of segmentation, especially in complex remote sensing image scenes.
[0005] In order to improve the accuracy of semantic segmentation of remote sensing images, especially the recognition of small objects in complex scenes, a technical solution that can achieve effective dynamic complementary fusion between local and global features is urgently needed. In addition, the existing methods do not adequately consider the contribution of features of different scales when fusion, which puts forward new requirements for improving the segmentation performance of cross-scale features.
[0006] Therefore, a multi-directional dynamic feature fusion method based on deep learning is proposed, which can effectively combine the advantages of convolutional neural networks and Transformer, and is the key to solving the current bottleneck of remote sensing image semantic segmentation technology. Summary of the invention
[0007] The present invention provides a multi-directional dynamic feature fusion method based on deep learning to solve the problem of segmenting confusing similar objects and small targets in remote sensing urban images.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is: A multi-directional dynamic feature fusion method based on deep learning, comprising the following steps: S1. Prepare remote sensing semantic segmentation dataset; S2. Construct a receptive field module and use convolution kernels of different sizes to process features, so that different convolution kernels focus on different feature subspaces, increase the diversity of features, and better process information and details of different scales; S3. Build a cross-window contextual interaction module, which consists of two parts: window attention and cross-window attention. Window attention focuses on local features, while cross-window attention focuses on global features, which makes up for the shortcomings of the Vision Transformer architecture in building cross-scale attention. S4. Construct a multi-directional dynamic feature fusion method, combining the receptive field module and the cross-window context interaction module.
[0009] In the above S2, when constructing the receptive field module, i.e., the RFB module, the following sub-steps are adopted: Create three parallel feature extraction branches, each of which uses a different convolutional structure to process the input features; The first branch branch0 extracts local features. It first compresses the input feature map through 1x1 convolution, and then uses 3x3 convolution to extract local features. The expansion rate of the convolution kernel is set to 1 to obtain the feature map X0. The second branch branch1 extracts the features of the first scale. It first compresses the channels of the input feature map through 1x1 convolution, then performs normal 3x3 convolution operation to keep the extraction of local features, and finally uses 3x3 dilated convolution with the dilation rate set to 3 to further process the feature map to obtain feature map X1. The third branch branch2 extracts features of the second scale, which is larger than the first scale. It first performs channel compression on the input feature map through 1x1 convolution, then performs normal 5×5 convolution operation, and finally uses 3x3 dilated convolution with the dilation rate set to 5 to further process the feature map, further expand the receptive field, capture global information, and obtain feature map X2. The features extracted by the three branches are spliced in the channel dimension to obtain feature map X3. Then use 1x1 convolution to fuse the concatenated features to reduce the dimension of the output feature map X3 of the three branches to obtain the feature map X4; Pass the input X through the shortcut path to create a tensor with the same dimension as the main path output, and obtain the convolved feature short; Add the fused feature X4 and the feature short obtained through the shortcut path to achieve residual connection and obtain the feature map X5; Apply the ReLU activation function and return the fused feature map X6.
[0010] In the above S3, when constructing the cross-window context interaction module, i.e., the IOWTB module, the following sub-steps are adopted: The input feature map F0 passes through the BN layer to obtain the standardized feature map F1; Input F1 into the IOWTBAttation module, perform window division, and obtain the feature map F2; The output F2 of the attention mechanism is randomly discarded through drop_path to enhance the regularization effect of the model, thereby avoiding overfitting and obtaining the feature map F3; Add the F0 attention output feature F3 to form a new feature map F4; The feature map after residual connection is input into the BN layer for re-standardization to obtain the standardized feature map F5; Through MLP, feature map F5 is nonlinearly expanded and transformed to obtain feature map F6; Use DropPath to further enhance the regularization effect and obtain feature map F7; The feature map F4 is added to the feature map F7 to form the final output, and the final feature map F8 is obtained.
[0011] In the above S3, the IOWTBAttation module is constructed using the following sub-steps: The IOWTBAttation module consists of two parts: window attention WA and cross-window attention CWA. WA focuses on the local information within the window, while CWA focuses on the global information across windows, which makes up for the shortcomings of the Vision Transformer architecture in building cross-scale attention. WA is a branch based on window-based multi-head attention to capture local context information. The CWA operation can capture long-distance dependencies and obtain global attention features. The feature map F0 is rearranged and divided into blocks by branch1 and branch2, and the ID sequence is divided using the window segmentation operation, so that the features in each local window become an independent block, and the feature maps F1 and F2 are obtained. The feature graph F1 is calculated through the qkv convolutional layer to obtain the query Q, key K and value V features, and obtain Q1, K1, V1, as shown in formula (1): Q,K,V=X , , ; (1) The qkv convolutional layer is a 1x1 convolutional layer that receives dim channel input and outputs 3*dim channels; The feature map F2 calculates the query Q, key K and value V features through the qkv1 linear layer. The qkv1 linear layer is also used to generate qkv, which is the processing of local features. The linear layer is more suitable for processing small local features, while the convolution is more suitable for processing spatial information. Another qkv is generated through formula (1) to obtain Q2, K2, V2; Perform matrix dot product operation on q1 and k1, calculate the similarity between the query and the key in each window, and get the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large, and get the final score Dots0, as shown in formula (2): Attention_Scores ; (2) Among them, Q and K are the feature matrices of query Q1 and key K1, query Q2 and key K2 respectively. represents the transpose of the key matrix, is the dimension of query and key; Perform matrix dot product operation on Q2 and K2, calculate the similarity between the query and the key in each window, and obtain the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large. Apply formula (2) to obtain the final score Dots2. The attention score Dots0 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots1; Apply softmax to the attention score Dots1 and normalize it to get the attention weight attn0; Apply the attention weight attn0 to the value V to obtain the weighted attention result attn1, as shown in formula (3); +B)V; (3) The attention score Dots2 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots3; Apply softmax to the attention score Dots3 and normalize it to get the attention weight attn2; Apply the attention weight attn2 to the value V and use formula (3) to get the weighted attention result attn3; Rearrange the attention results attn1 and attn3 to the shape of the original input, crop and remove the redundant padding, restore to the original feature map size, add the global attention and local attention results, and fuse the features of different scales to obtain the feature map F3; The input feature map F3 undergoes deep convolution, processes each channel of the input feature map independently, captures spatial information, and performs a 1x1 convolution operation on all channels at each position, that is, linearly combines them in the channel dimension, mixes information from different channels, and finally normalizes the output to maintain training stability, obtaining the feature map F4.
[0012] In the above S4, a multi-directional dynamic feature fusion method is constructed, using the following sub-steps: Parse the image dataset through the Swin Transformer Block 1 module to obtain the feature map res1; Input the feature map res1 into the Swin Transformer Block 2 module for parsing to obtain the feature map res2; Input feature map res2 into Swin Transformer Block 3 module for parsing to obtain feature map res3; Input feature map res3 into Swin Transformer Block 4 module for parsing to obtain feature map res4; First, res3_w1 is activated by ReLU, then normalized to generate a weight weight; then the weight is applied to res3 and res4 processed by upsampling and convolution, and finally further processed by RFB module to get res5; res2_w1 is activated and normalized by ReLU to generate weight weight; then the weight is applied to res2 and res5 after upsampling, and processed by RFB module to get res6; Perform the same ReLU and normalization operations on res1_w1 to generate weight weight; then fuse res1 with the upsampled res6, and further extract features through the IOWTB module to obtain res7; Use res2_w2 to generate weight weight, and weighted fusion of res2, res6 and res7 after downsampling, and further extract features through IOWTB module to obtain res8; Use res3_w2 to generate weights, perform weighted fusion on res3, res5 and downsampled res8, and further extract features through the IOWTB module to obtain res9; Use res4_w2 to generate weights, weight the downsampled res4 and res9, and process them through the IOWTB module to get the final res10; Finally, the four feature maps res7, res8, res9, and res10 are fused to obtain res11; The feature map res11 is passed through the segmentation head to get the final output.
[0013] The present invention provides a multi-directional dynamic feature fusion method based on deep learning, which has the following beneficial effects: 1. The receptive field module is introduced, and convolution kernels of different sizes are used to process features, so that different convolution kernels focus on different feature subspaces, increasing the diversity of features and better processing information and details of different scales. RFB is customized to enhance the model's feature extraction and generalization capabilities in local areas, which helps to have a more detailed understanding of complex spatial patterns.
[0014] 2. A cross-window contextual interaction module is introduced, which consists of two parts: window attention and cross-window attention. Window attention focuses on local features, and cross-window attention focuses on global features, which makes up for the shortcomings of the Vision Transformer architecture in building cross-scale attention. The cross-window contextual interaction module enhances the model's ability to interact with cross-window information and model long-range dependencies. This comprehensive feature processing framework is crucial to solving the complex dynamic problems of remote sensing image segmentation.
[0015] 3. A multi-directional dynamic feature fusion method based on deep learning is proposed, which dynamically merges the local attention provided by the receptive field module with the global attention promoted by the cross-window context interaction module. This innovative fusion method takes advantage of the complementary advantages of different receptive fields, enabling the network to pay more attention to the discriminative features in similar categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 is a flow chart of the present invention; Figure 2 It is a schematic diagram of the structure of the receptive field RFB module; Figure 3 It is a schematic diagram of the structure of building a cross-window context interaction module IOWTB; Figure 4 yes Figure 2 The structural diagram of the IOWTBAttation module; Figure 5 It is a structural schematic diagram of the feature fusion method proposed in the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the present invention clearer, the following content will systematically and completely describe the specific technical solutions of the present invention in combination with the drawings provided according to the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0018] Embodiment 1: like Figure 1 As shown in , a multi-directional dynamic feature fusion method based on deep learning includes the following steps: Step 1: Prepare remote sensing semantic segmentation dataset; Step 2: Construct a receptive field module and use convolution kernels of different sizes to process features, so that different convolution kernels focus on different feature subspaces, increase the diversity of features, and better process information and details of different scales.
[0019] Step 3: Construct a cross-window contextual interaction module, which consists of two parts: window attention and cross-window attention. Window attention focuses on local features, while cross-window attention focuses on global features, which makes up for the shortcomings of the Vision Transformer architecture in building cross-scale attention.
[0020] Step 4: Construct a multi-directional dynamic feature fusion method combining the receptive field module and the cross-window context interaction module.
[0021] like Figure 2 As shown, in step 2, when constructing the receptive field module, that is, the RFB module, the following sub-steps are adopted: Create three parallel feature extraction branches, each of which uses a different convolutional structure to process the input features; The first branch branch0 extracts local features. It first compresses the input feature map through 1x1 convolution, and then uses 3x3 convolution to extract local features. The expansion rate of the convolution kernel is set to 1 to obtain the feature map X0. The second branch, branch1, extracts larger-scale features. It first compresses the input feature map through 1x1 convolution, then performs normal 3x3 convolution to keep the local features extracted, and finally uses 3x3 dilated convolution with the dilation rate set to 3 to further process the feature map to obtain feature map X1. The third branch, branch2, extracts features of a larger scale. It first compresses the input feature map through a 1x1 convolution, then performs a normal 5×5 convolution operation, and finally uses a 3x3 dilated convolution with a dilation rate of 5 to further process the feature map, further expanding the receptive field and capturing global information to obtain feature map X2. The features extracted by the three branches are concatenated in the channel dimension to obtain the feature map X3; Then use 1x1 convolution to fuse the concatenated features to reduce the dimension of the output feature map X3 of the three branches to obtain the feature map X4; Pass the input X through the shortcut path to create a tensor with the same dimension as the main path output, and obtain the convolved feature short; Add the fused feature X4 and the feature short obtained through the shortcut path to achieve residual connection and obtain the feature map X5; Apply the ReLU activation function and return the fused feature map X6.
[0022] like Figure 2 As shown, in step 3, when building a cross-window context interaction module, that is, an IOWTB module, the following sub-steps are adopted: The input feature map F0 passes through the BN layer to obtain the standardized feature map F1; Input F1 into the IOWTBAttation module, perform window division, and obtain the feature map F2; The output F2 of the attention mechanism is randomly discarded through drop_path to enhance the regularization effect of the model, thereby avoiding overfitting and obtaining the feature map F3; Add the F0 attention output feature F3 to form a new feature map F4; The feature map after residual connection is input into the BN layer for re-standardization to obtain the standardized feature map F5; Through MLP, feature map F5 is nonlinearly expanded and transformed to obtain feature map F6; Use DropPath to further enhance the regularization effect and obtain feature map F7; Add the feature map F4 and the feature map F7 to form the final output, and obtain the final feature map F8; like Figure 3 As shown, in step 3, the IOWTBAttation module is constructed, using the following sub-steps: The OWTBAttation module consists of two parts: window attention WA and cross-window attention CWA. WA focuses on the local information within the window, while CWA focuses on the global information across windows, which makes up for the deficiency of the Vision Transformer architecture in building cross-scale attention. WA is a branch based on window-based multi-head attention to capture local context information. The CWA operation can capture long-distance dependencies and obtain global attention features. The feature map F0 is rearranged and divided into blocks by branch1 and branch2, and the ID sequence is divided using the window segmentation operation, so that the features in each local window become an independent block, and the feature maps F1 and F2 are obtained. The feature graph F1 is calculated through the qkv convolutional layer to obtain the query Q, key K and value V features, and obtain Q1, K1, V1, as shown in formula (1): Q,K,V=X , , ; (1) The qkv convolutional layer is a 1x1 convolutional layer that receives dim channel input and outputs 3*dim channels; The feature map F2 calculates the query Q, key K and value V features through the qkv1 linear layer. The qkv1 linear layer is also used to generate qkv, which is the processing of local features. The linear layer is more suitable for processing small local features, while the convolution is more suitable for processing spatial information. Another qkv is generated through formula (1) to obtain Q2, K2, V2; Perform matrix dot product operation on q1 and k1, calculate the similarity between the query and the key in each window, and get the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large, and get the final score Dots0, as shown in formula (2): Attention_Scores ; (2) Among them, Q and K are the feature matrices of query Q1 and key K1, query Q2 and key K2 respectively. represents the transpose of the key matrix, is the dimension of query and key; Perform matrix dot product operation on Q2 and K2, calculate the similarity between the query and the key in each window, and obtain the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large. Apply formula (2) to obtain the final score Dots2. The attention score Dots0 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots1; Apply softmax to the attention score Dots1 and normalize it to get the attention weight attn0; Apply the attention weight attn0 to the value V to obtain the weighted attention result attn1, as shown in formula (3); +B)V; (3) The attention score Dots2 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots3; Apply softmax to the attention score Dots3 and normalize it to get the attention weight attn2; Apply the attention weight attn2 to the value V and use formula (3) to get the weighted attention result attn3; Rearrange the attention results attn1 and attn3 to the shape of the original input, crop and remove the redundant padding, restore to the original feature map size, add the global attention and local attention results, and fuse the features of different scales to obtain the feature map F3; The input feature map F3 is subjected to deep convolution, which processes each channel of the input feature map independently, captures spatial information, and performs a 1x1 convolution operation on all channels at each position, that is, linear combination in the channel dimension, mixing information from different channels, and finally normalizing the output to maintain training stability. Feature map F4 is obtained; In step 4, a multi-directional dynamic feature fusion method is constructed using the following sub-steps: Parse the image dataset through the Swin Transformer Block 1 module to obtain the feature map res1; Input the feature map res1 into the Swin Transformer Block 2 module for parsing to obtain the feature map res2; Input feature map res2 into Swin Transformer Block 3 module for parsing to obtain feature map res3; Input feature map res3 into Swin Transformer Block 4 module for parsing to obtain feature map res4; First, res3_w1 is activated by ReLU, then normalized to generate a weight. Then the weight is applied to res3 and res4 after upsampling and convolution, and finally further processed by RFB module to get res5. res2_w1 is activated and normalized by ReLU to generate weights. The weights are then applied to res2 and res5 after upsampling, and processed by the RFB module to obtain res6. Perform the same ReLU and normalization operations on res1_w1 to generate weight weight. Then res1 is fused with the upsampled res6, and features are further extracted through the IOWTB module to obtain res7. Use res2_w2 to generate weight weight, and weighted fusion of res2, res6 and res7 after downsampling, and further extract features through IOWTB module to obtain res8; Use res3_w2 to generate weights, perform weighted fusion on res3, res5 and downsampled res8, and further extract features through the IOWTB module to obtain res9; Use res4_w2 to generate weights, weighted fusion of downsampled res4 and res9t, and process them through the IOWTB module to get the final res10; Finally, the four feature maps res7, res8, res9, and res10 are fused to obtain res11; The feature map res11 is passed through the segmentation head to get the final output.
[0023] This paper not only solves the common problems of inter-class similarity and small object segmentation in the field of remote sensing image semantic segmentation, but also designs a multi-directional dynamic feature fusion method based on deep learning. Through this method, we are able to effectively improve the model's ability in modeling short-range and long-range interactions, thereby extracting highly discriminative features necessary to distinguish closely related classes and identify small objects.
[0024] Example: 1) Parameter settings We implemented all experiments using PyTorch on a single NVIDIA RTX3090 GPU and used AdamW as the optimizer with a learning rate of 0.0006. We used SoftCrossEntropyLoss and DiceLoss as the joint loss function and used the backpropagation index to measure the gap between the 2D segmentation map and the true value we obtained. For each method, we used F1, mIOU, and OA as evaluation metrics: Precision: ; (4) Recall: ; (5) F1 value: F1 = ; (6) Mean Intersection Over Union (mIOU): mIOU = ; (7) Overall Accuracy (OA): ; (8) in, , , and They represent the proportion of positive samples correctly predicted as positive, positive samples incorrectly predicted as negative, negative samples correctly predicted as negative, and negative samples incorrectly predicted as positive for a specific category. The overall accuracy (OA) is the overall accuracy calculated for all categories.
[0025] 2) Experimental results In order to verify the effectiveness of this solution, further explanation is given below in combination with experimental data.
[0026] Extensive ablation experiments are conducted on the ISPRS Vaihingen and ISPRS Potsdam datasets to verify the effectiveness and robustness of our method. The same test-time augmentation strategy and auxiliary loss are used in all ablation studies to ensure fair comparison. We conduct three sets of experiments: 1) Base represents a network that uses Swin-s as the encoder and directly upsamples features. 2) Base+RFB+RFB represents a network that uses RFB in both paths of the decoder. 3) Base+RFB+IOWTB represents a network that replaces the bottom-up path of the decoder with IOWTB.
[0027] Table 1
[0028] The results of the ablation study are shown in Table 1. Compared with the base model, Base+RFB+RFB improves the F1 score by 1.88%, OA (overall accuracy) by 0.27%, and mIoU (mean intersection over union) by 2.52% on the Vaihingen dataset. On the Potsdam dataset, the average F1 is improved by 0.77%, OA by 0.1%, and mIoU by 1.35%. These improvements indicate that the segmentation ability is enhanced by using a multi-directional feature interaction method in the decoder. Compared with Base+RFB+RFB, Base+RFB+IOWTB improves the F1 score by 1.71%, OA by 0.7%, and mIoU by 2.73% on the Vaihingen dataset. On the Potsdam dataset, the average F1, OA, and mIoU are improved by 1.3%, 1.02%, and 2.18%, respectively. This shows that the interactive fusion method between CNN and Transformer significantly improves the effect, highlighting the advantages of the model in short-distance and long-distance complementary modeling.
Claims
1. A multi-directional dynamic feature fusion method based on deep learning, characterized in that: The following steps are involved: S1. Prepare remote sensing semantic segmentation dataset; S2, build a receptive field module and use convolution kernels of different sizes to process features, so that different convolution kernels focus on different feature subspaces; S3. Build a cross-window contextual interaction module, which consists of two parts: window attention and cross-window attention. Window attention focuses on local features, while cross-window attention focuses on global features. S4. Construct a multi-directional dynamic feature fusion method, combining the receptive field module and the cross-window context interaction module.
2. According to the multi-directional dynamic feature fusion method based on deep learning described in claim 1, it is characterized in that: In the above S2, when constructing the receptive field module, i.e., the RFB module, the following sub-steps are adopted: Create three parallel feature extraction branches, each of which uses a different convolutional structure to process the input features; The first branch branch0 extracts local features. It first compresses the input feature map through 1x1 convolution, and then uses 3x3 convolution to extract local features. The expansion rate of the convolution kernel is set to 1 to obtain the feature map X0. The second branch branch1 extracts the features of the first scale. It first compresses the channels of the input feature map through 1x1 convolution, then performs normal 3x3 convolution operation to keep the extraction of local features, and finally uses 3x3 dilated convolution with the dilation rate set to 3 to further process the feature map to obtain feature map X1. The third branch branch2 extracts the features of the second scale, which is larger than the first scale. It first compresses the channels of the input feature map through 1x1 convolution, then performs a normal 5×5 convolution operation, and finally uses a 3x3 dilated convolution with a dilation rate of 5 to further process the feature map, further expand the receptive field, capture global information, and obtain feature map X2; The features extracted by the three branches are concatenated in the channel dimension to obtain the feature map X3; Then use 1x1 convolution to fuse the concatenated features to reduce the dimension of the output feature map X3 of the three branches to obtain the feature map X4; Pass the input X through the shortcut path to create a tensor with the same dimension as the main path output, and obtain the convolved feature short; Add the fused feature X4 and the feature short obtained through the shortcut path to achieve residual connection and obtain the feature map X5; Apply the ReLU activation function and return the fused feature map X6.
3. According to the multi-directional dynamic feature fusion method based on deep learning described in claim 2, it is characterized in that: In the above S3, when constructing the cross-window context interaction module, i.e., the IOWTB module, the following sub-steps are adopted: The input feature map F0 passes through the BN layer to obtain the standardized feature map F1; Input F1 into the IOWTBAttation module, perform window division, and obtain the feature map F2; The output F2 of the attention mechanism is randomly discarded through drop_path to enhance the regularization effect of the model, thereby avoiding overfitting and obtaining the feature map F3; Add the F0 attention output feature F3 to form a new feature map F4; The feature map after residual connection is input into the BN layer for re-standardization to obtain the standardized feature map F5; Through MLP, feature map F5 is nonlinearly expanded and transformed to obtain feature map F6; Use DropPath to further enhance the regularization effect and obtain feature map F7; The feature map F4 is added to the feature map F7 to form the final output, and the final feature map F8 is obtained.
4. According to the multi-directional dynamic feature fusion method based on deep learning described in claim 3, it is characterized in that: In the S3, the IOWTBAttation module is constructed using the following sub-steps: The IOWTBAttation module consists of two parts: window attention WA and cross-window attention CWA. WA focuses on the local information within the window, while CWA focuses on the global information across windows. WA is a branch based on window-based multi-head attention to capture local context information. The CWA operation can capture long-distance dependencies and obtain global attention features. The feature map F0 is rearranged and divided into blocks by branch1 and branch2, and the ID sequence is divided using the window segmentation operation, so that the features in each local window become an independent block, and the feature maps F1 and F2 are obtained. The feature graph F1 is calculated through the qkv convolutional layer to obtain the query Q, key K and value V features, and obtain Q1, K1, V1, as shown in formula (1): Q,K,V=X , , ;(1) The qkv convolutional layer is a 1x1 convolutional layer that receives dim channel input and outputs 3*dim channels; The feature map F2 calculates the query Q, key K and value V features through the qkv1 linear layer. The qkv1 linear layer is also used to generate qkv, which is the processing of local features. The linear layer is more suitable for processing small local features, while the convolution is more suitable for processing spatial information. Another qkv is generated through formula (1) to obtain Q2, K2, V2; Perform matrix dot product operation on q1 and k1, calculate the similarity between the query and the key in each window, and get the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large, and get the final score Dots0, as shown in formula (2): Attention_Scores ;(2) Among them, Q and K are the feature matrices of query Q1 and key K1, query Q2 and key K2 respectively. represents the transpose of the key matrix, is the dimension of query and key; Perform matrix dot product operation on Q2 and K2, calculate the similarity between the query and the key in each window, and obtain the attention score. Multiply the attention score by the scaling factor scale to prevent the attention value from being too large. Apply formula (2) to obtain the final score Dots2. The attention score Dots0 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots1; Apply softmax to the attention score Dots1 and normalize it to get the attention weight attn0; Apply the attention weight attn0 to the value V to obtain the weighted attention result attn1, as shown in formula (3); +B)V;(3) The attention score Dots2 is added with the relative position bias relative_position_bias to correct the attention score in each window to obtain Dots3; Apply softmax to the attention score Dots3 and normalize it to get the attention weight attn2; Apply the attention weight attn2 to the value V and use formula (3) to get the weighted attention result attn3; Rearrange the attention results attn1 and attn3 to the shape of the original input, crop and remove the redundant padding, restore to the original feature map size, add the global attention and local attention results, and fuse the features of different scales to obtain the feature map F3; The input feature map F3 undergoes deep convolution, processes each channel of the input feature map independently, captures spatial information, and performs a 1x1 convolution operation on all channels at each position, that is, linearly combines them in the channel dimension, mixes information from different channels, and finally normalizes the output to maintain training stability, obtaining the feature map F4.
5. According to the multi-directional dynamic feature fusion method based on deep learning as described in claim 4, it is characterized in that: In the above S4, a multi-directional dynamic feature fusion method is constructed, using the following sub-steps: Parse the image dataset through the Swin Transformer Block 1 module to obtain the feature map res1; Input the feature map res1 into the Swin Transformer Block 2 module for parsing to obtain the feature map res2; Input feature map res2 into Swin Transformer Block 3 module for parsing to obtain feature map res3; Input feature map res3 into Swin Transformer Block 4 module for parsing to obtain feature map res4; First, res3_w1 is activated by ReLU, then normalized to generate a weight weight; then the weight is applied to res3 and res4 processed by upsampling and convolution, and finally further processed by RFB module to get res5; res2_w1 is activated and normalized by ReLU to generate weight weight; then the weight is applied to res2 and res5 after upsampling, and processed by RFB module to get res6; Perform the same ReLU and normalization operations on res1_w1 to generate weight weight; then fuse res1 with the upsampled res6, and further extract features through the IOWTB module to obtain res7; Use res2_w2 to generate weight weight, and weighted fusion of res2, res6 and res7 after downsampling, and further extract features through IOWTB module to obtain res8; Use res3_w2 to generate weights, perform weighted fusion on res3, res5 and downsampled res8, and further extract features through the IOWTB module to obtain res9; Use res4_w2 to generate weights, weight the downsampled res4 and res9, and process them through the IOWTB module to get the final res10; Finally, the four feature maps res7, res8, res9, and res10 are fused to obtain res11; The feature map res11 is passed through the segmentation head to get the final output.
Citation Information
Patent Citations
Medical image segmentation method fusing multi-scale features and multi-attention mechanism based on Swin Transform
CN116416434A
Global and local feature reconstruction network-based medical image segmentation method
US20230274531A1